Dynamic path control method, device, medium and product for cross-domain AI computing power network
Patent Information
- Application Number
- CN202610759150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-05-29
AI Technical Summary
[0004]本申请的一个目的是提供一种面向跨域AI算力网络的动态路径控制方法、设备、介质及产品,至少用以解决现有技术中跨域AI算力网络的路径控制机制无法综合感知节点硬件能力、AI通信模式差异以及多维无损网络状态,导致路径选择死板、跨域传输存在合规风险、拥塞恢复缓慢且易引发网络震荡的问题
[0016]第四方面,本申请的一些实施例还提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现如上所述方法的步骤。
Smart Images

Figure CN122316966B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computing power networks, and in particular to a dynamic path control method, device, medium, and product for cross-domain AI computing power networks. Background Technology
[0002] With the rapid development of large-scale model training, distributed inference, and multi-regional collaborative AI tasks, enterprise-level AI infrastructure has gradually evolved from interconnection within a single data center to collaborative interconnection across racks, clusters, regions, and even partner computing domains. In such cross-domain AI computing network scenarios, different computing nodes need to frequently exchange gradient parameters, model shards, feature data, and inference requests. Especially in high-performance AI lossless network environments based on RDMA (Remote Direct Memory Access) and RoCE (RDMA network based on converged Ethernet), higher requirements are placed on network latency, bandwidth stability, congestion control, and path reliability. At the same time, different AI services also have significantly differentiated communication semantic characteristics. For example, aggregated communication focuses more on global synchronization efficiency, parameter synchronization focuses more on reliability and convergence performance, while cross-domain inference services focus more on tail latency stability and cross-domain reachability. Therefore, traditional unified and static path control methods are no longer sufficient to meet the dynamic network scheduling needs of complex AI business scenarios.
[0003] In existing technologies, path control schemes for AI computing networks typically only address a single aspect of node access, link scheduling, QoS optimization, or multi-cluster service discovery. Most employ static registration and one-off path decision mechanisms, lacking the ability to generate structured paths tailored to AI communication patterns, and also lacking a closed-loop control mechanism that integrates access description, hierarchical topology, link telemetry, and runtime reconfiguration. Furthermore, existing solutions often fail to incorporate hardware-level parameters such as accelerator card type, RDMA capability, RoCE version, and PFC integrity (priority flow control integrity) into the path candidate filtering logic. Path reconfiguration is often triggered solely by single metrics like latency or packet loss, failing to accurately reflect the true congestion state in lossless AI networks. This can easily lead to path oscillations, slow recovery, and cross-domain compliance risks, making it difficult to adapt to the stable transmission requirements of dynamically changing cross-domain AI computing networks. Summary of the Invention
[0004] One objective of this application is to provide a dynamic path control method, device, medium, and product for cross-domain AI computing power networks, at least to solve the problems that the path control mechanism of cross-domain AI computing power networks in the prior art cannot comprehensively perceive the node hardware capabilities, differences in AI communication modes, and multi-dimensional lossless network status, resulting in rigid path selection, compliance risks in cross-domain transmission, slow congestion recovery, and easy network instability.
[0005] To achieve the above objectives, some embodiments of this application provide the following aspects:
[0006] This application provides a dynamic path control method for cross-domain AI computing power networks, the method comprising:
[0007] Obtain the status information of the computing power nodes to be connected and generate an access description;
[0008] Based on the access description, a hierarchical topology diagram including a node layer, a cluster layer, and a region layer is constructed.
[0009] Determine the AI communication mode corresponding to the data stream to be transmitted;
[0010] The hierarchical topology graph is constrained and pruned based on cross-domain compliance constraints and hardware compatibility constraints.
[0011] Based on the AI communication mode and the cropped hierarchical topology graph, a corresponding candidate path structure is generated.
[0012] The candidate path structure is scored based on network state parameters and task context parameters to determine the target path, and path forwarding control of the corresponding service data flow is executed based on the target path.
[0013] During the operation of the target path, telemetry indicators are collected for joint determination, and local or global path reconstruction is performed based on the joint determination results.
[0014] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.
[0015] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the steps of the method described above.
[0016] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.
[0017] Compared with related technologies, the solution provided in this application constructs a hierarchical topology diagram including a node layer, a cluster layer, and a regional layer. Combined with hardware capability parameters, cross-domain compliance constraints, and AI communication modes described in the access description, candidate paths are structurally generated and dynamically scored, enabling different AI business scenarios to match corresponding path topologies. Furthermore, by associating different communication modes such as aggregated communication, parameter synchronization, model fragmentation transmission, and cross-domain inference services with path structures such as ring, tree, aggregation, chain, and sparse fully connected paths, the data transmission efficiency, link utilization, and path matching accuracy of cross-domain AI tasks can be effectively improved, thereby enhancing the business adaptability and network collaboration capabilities of the cross-domain AI computing network.
[0018] Furthermore, by introducing a multi-dimensional joint judgment mechanism based on AI lossless network-specific indicators such as ECN labeling ratio, congestion queue depth, P95 latency, and packet loss rate, and combining it with local path reconstruction and global path reconstruction control strategies, dynamic adaptive adjustment of runtime paths is achieved. Simultaneously, through a pre-path keep-alive mechanism and an atomic update mechanism for path identifiers, path switching and recovery can be quickly completed in scenarios of link anomalies, node failures, or cross-domain congestion, thereby effectively reducing the probability of path oscillations and improving the stability, reliability, and fault recovery efficiency of cross-domain AI computing power networks. Attached Figure Description
[0019] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0020] Figure 1 A flowchart illustrating a dynamic path control method for cross-domain AI computing power networks, provided as an exemplary embodiment of this disclosure;
[0021] Figure 2 A flowchart illustrating another dynamic path control method for cross-domain AI computing power networks provided as an exemplary embodiment of this disclosure;
[0022] Figure 3 A candidate path structure diagram in a dynamic path control method for cross-domain AI computing power networks provided as an exemplary embodiment of this disclosure;
[0023] Figure 4 A flowchart illustrating the joint decision-making mechanism in a dynamic path control method for cross-domain AI computing power networks, provided as an exemplary embodiment of this disclosure;
[0024] Figure 5 An exemplary structural diagram of the electronic device provided for some embodiments of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] Figure 1 An exemplary flowchart of a dynamic path control method for cross-domain AI computing power networks, provided as an exemplary embodiment of this disclosure, is shown below. The method includes:
[0027] S101. Obtain the status information of the computing power node to be connected and generate an access description.
[0028] Specifically, such as Figure 2 As shown, during the process of a computing node accessing a cross-domain AI computing power network, the operational status information of the node to be accessed can be collected through a node-side agent program, control plane management component, resource orchestration platform, or network controller. This status information may include basic node identity information, tenant information, domain affiliation information, geographic affiliation information, computing resource information, network interconnection capability information, link reachability information, and node operational health status information. Specifically, the computing resource information may include processor architecture, acceleration device type, number of acceleration devices, GPU memory capacity, memory resource utilization, and task load status; the network interconnection capability information may include network bandwidth capability, low-latency communication capability, RDMA support capability, switching protocol compatibility capability, and network forwarding capability; and the link reachability information may include an accessible domain list, an accessible cluster list, candidate relay node information, and egress node information.
[0029] After obtaining the aforementioned state information, the multi-dimensional state parameters can be uniformly structured to form a standardized access description. This access description can be organized using key-value pairs, object descriptions, tagging, or graph structures to support unified parsing and invocation between different network components. Furthermore, the access description can be dynamically updated; when node operating status, link status, or network policies change, the corresponding access description can be regenerated or incrementally updated to ensure the control plane can perceive changes in the network environment in real time. In some implementations, a valid time parameter or version identifier can be set for the access description to prevent long-term retention of historical state information from causing path decision failures.
[0030] S102. Based on the access description, construct a hierarchical topology diagram including a node layer, a cluster layer, and a regional layer.
[0031] Specifically, after obtaining the access descriptions corresponding to multiple computing nodes, a multi-level topology can be constructed based on the connection relationships between nodes, cluster affiliation relationships, and cross-domain interconnection relationships. The node layer can be used to describe the fine-grained physical connection relationships between nodes, and the topology edges in the node layer can be associated with attributes such as link bandwidth, link latency, link jitter, link packet loss rate, queue status, and link health status. The cluster layer can be used to logically aggregate multiple nodes to form a cluster-level abstract topology, and the edge attributes in the cluster layer can include cluster egress capacity, aggregated bandwidth capacity, and cluster-level congestion status. The region layer can be used to describe the interconnection relationships between different regions, different management domains, or different operational domains, and the edge attributes in the region layer can include cross-domain hop count, cross-domain transmission cost, cross-domain reliability, and policy restriction information.
[0032] In constructing a hierarchical topology graph, link state parameters in the node layer can be mapped to the cluster and regional layers through bottom-up aggregation. For example, the bandwidth utilization, congestion level, or latency status of multiple underlying links can be aggregated to form abstract state parameters for the corresponding layer, thereby improving the efficiency of topology management in large-scale cross-domain network scenarios. Furthermore, mapping relationships can be established between different layers to support cross-layer coordinated scheduling during path generation. For instance, when the regional layer detects cross-domain congestion, it can be further mapped to the corresponding cluster or node layer to locate specific abnormal link areas. By constructing a hierarchical topology graph with hierarchical abstraction capabilities, the path calculation complexity in ultra-large-scale AI computing networks can be reduced while ensuring global network visibility.
[0033] S103. Determine the AI communication mode corresponding to the data stream to be transmitted.
[0034] Specifically, upon receiving the data stream to be transmitted, the business behavior characteristics corresponding to the data stream can be analyzed first to identify the corresponding AI communication mode. These business behavior characteristics may include data exchange direction, number of participating nodes, data synchronization frequency, data packet size, communication timing characteristics, and interaction relationships between nodes. For example, when there is periodic parameter synchronization behavior among multiple training nodes, it can be identified as a collective communication mode or parameter synchronization mode; when there is sequential data transmission behavior among multiple model stages, it can be identified as a model fragmentation transmission mode; and when there is dynamic routing request distribution behavior among multiple inference nodes, it can be identified as a cross-domain inference service mode.
[0035] Furthermore, the network requirements for different AI communication modes differ significantly. For example, aggregated communication modes typically have high requirements for link bandwidth consistency and multi-node synchronization performance; parameter synchronization modes usually focus more on link reliability and the load balancing capabilities of aggregation nodes; model fragmentation transmission modes usually focus more on stable transmission capabilities between stages; and cross-domain inference service modes focus more on tail latency stability and cross-domain service reachability. Therefore, after identifying the AI communication modes, differentiated path control strategies can be generated based on the corresponding communication modes, making the path control process more in line with the needs of actual AI business scenarios.
[0036] S104. Perform constraint trimming on the hierarchical topology graph based on cross-domain compliance constraints and hardware compatibility constraints.
[0037] Specifically, before generating candidate paths, the hierarchical topology graph can first undergo constraint filtering and topology pruning. The cross-domain compliance constraints can include data sovereignty restrictions, geographic access restrictions, tenant isolation policies, security domain isolation policies, and cross-domain access permission policies. For example, in some business scenarios, specific types of data can be restricted to flowing only within a specified geographic area, or data traffic between different tenants can be restricted from sharing the same network resource area.
[0038] The hardware compatibility constraints can include node network protocol compatibility, RDMA capability compatibility, RoCE protocol version compatibility, switch capability compatibility, and acceleration device type matching. For example, for data streams using RDMA direct transmission, nodes that do not have RDMA communication capabilities can be removed from the candidate topology; for service scenarios requiring low-latency aggregated communication, nodes and links that do not meet lossless network conditions or have insufficient link capabilities can be further filtered out.
[0039] During constraint pruning, nodes, links, or path segments that do not meet the conditions can be deleted, masked, downgraded, or isolated, thereby forming an effective topology that meets the policy requirements. By performing constraint pruning in advance, the search space in the subsequent path generation stage can be effectively reduced, while the participation of illegal paths in the path scoring and scheduling process can be reduced, improving the overall path control efficiency and path legality.
[0040] S105. Generate corresponding candidate path structures based on the AI communication mode and the cropped hierarchical topology graph.
[0041] Specifically, after topology pruning, candidate path structures can be generated within the effective topology range by combining the identified AI communication patterns. Different AI communication patterns exhibit different data exchange behaviors, therefore different path organization methods can be used to generate candidate paths. For example, for multi-node synchronous training scenarios, data exchange path structures suitable for broadcasting, aggregation, or global synchronization can be generated; for model pipelined execution scenarios, chain-like path structures suitable for stage-sequential transmission can be generated; and for cross-domain inference service scenarios, path structures supporting dynamic request distribution and multi-node forwarding can be generated.
[0042] During path generation, different paths can be combined and filtered based on node reachability, link health status, cross-domain costs, and resource usage. Furthermore, topological constraints can be imposed on the path structure during path generation, such as limiting the maximum number of cross-domain hops in a path, limiting the number of shared links in a path, or limiting the proportion of high-risk links reused, to reduce path conflict risks and resource contention risks.
[0043] Furthermore, in some implementations, a target path and a backup path can be generated simultaneously. The backup path can maintain a certain degree of resource or risk isolation from the main path to enable rapid path switching in case of main path anomalies, thereby improving the continuous operation capability of cross-domain AI services.
[0044] S106. The candidate path structure is scored based on network state parameters and task context parameters to determine the target path, and path forwarding control of the corresponding service data flow is executed based on the target path.
[0045] Specifically, after generating multiple candidate path structures, the candidate paths can be comprehensively scored based on the corresponding network state parameters and task context parameters. The network state parameters may include link bandwidth, link latency, link jitter, link reliability, congestion level, cross-domain transmission cost, link health status, and link load, etc.; the task context parameters may include model size, training stage, task priority, node size, service type, task real-time requirements, and resource consumption, etc.
[0046] Subsequently, different candidate paths can be quantitatively evaluated based on a preset scoring model. For example, a weighted approach can be used to comprehensively calculate link quality, cross-domain cost, reliability, and policy matching degree to obtain a comprehensive score for the corresponding candidate path. In some implementations, the weights of different scoring parameters can be dynamically adjusted according to the AI communication mode. For example, for the aggregated communication mode, the weight corresponding to the link quality parameter can be increased; for the cross-domain inference service mode, the weights corresponding to reliability and cross-domain cost can be increased, thereby making the path selection results more in line with the actual needs of different AI business scenarios.
[0047] Furthermore, in some implementations, the scoring model can be dynamically corrected or adaptively optimized by combining historical path operation results, telemetry feedback information, or task execution results to improve the accuracy and stability of subsequent path scoring processes. Finally, the candidate path with the best scoring result or that meets preset conditions can be selected as the target path, and the corresponding path information can be sent to the network forwarding device or control node.
[0048] S107. During the operation of the target path, telemetry indicators are collected for joint determination, and local path reconstruction or global path reconstruction is performed based on the joint determination results.
[0049] Specifically, after the target path is put into operation, the network status corresponding to the target path can be continuously monitored during operation. The telemetry indicators may include parameters such as link latency, link jitter, link packet loss rate, queue depth, ECN marking ratio, link health status, and service-side tail latency. The above indicators can be collected in real time through switch telemetry interfaces, node detection components, link monitoring modules, or service feedback modules.
[0050] Subsequently, the current path status can be comprehensively determined based on the joint analysis results of multiple telemetry indicators. For example, multiple indicators can be normalized, and a comprehensive status result can be generated based on weighted analysis, threshold comparison, or time window statistics. When partial congestion, local failure, or node anomalies are detected in some links, local path reconstruction can be performed only on the abnormal path segments to reduce the impact on the overall service path; when large-scale link anomalies, cross-domain network failures, or the current topology no longer meets service requirements are detected, global path calculation can be re-performed, and a new target path can be generated.
[0051] In the above implementation, a dynamic path control mechanism for cross-domain AI computing networks is achieved by unifying and coordinating access description generation, hierarchical topology construction, AI communication pattern recognition, cross-domain constraint pruning, candidate path structure generation, dynamic scoring decision-making, and runtime path reconstruction. Compared with traditional static path control schemes, this mechanism not only combines AI communication semantics, hardware interconnection capabilities, and cross-domain compliance constraints for path generation and dynamic scheduling, but also combines runtime telemetry information for real-time perception and dynamic reconstruction of path status. This improves the data transmission efficiency, link stability, and fault recovery capabilities of cross-domain AI tasks, while reducing the probability of path congestion, path oscillation, and cross-domain transmission incompatibility.
[0052] Furthermore, in one embodiment, the access description includes at least one of the following:
[0053] Node identifier, domain affiliation information, accelerator type, number of accelerators, RDMA capability identifier, RoCE protocol version, link bandwidth information, and cross-domain reachability information.
[0054] Specifically, during the process of a computing power node accessing a cross-domain AI computing power network, the node to be accessed can first be discovered, and its node identity information, domain affiliation information, computing power capability information, interconnection capability information, topology location information, cross-domain reachability information, and candidate relay information can be collected to generate a corresponding access description. The node identity information can be used to uniquely identify the corresponding computing power node. The domain affiliation information can include fields such as tenant_id, domain_id, region_id, and cluster_id, used to describe the node's tenant, management domain, region, and cluster. The computing power capability information can include fields such as accelerator_type, accelerator_count, and memory_size, used to describe the type of acceleration device (GPU, NPU, or DPU), the number of acceleration devices, and the memory capacity of the node. The interconnection capability information can include fields such as rdma_capable, roce_version, and link_speed_gbps, used to describe whether the node has RDMA direct transmission capability, the corresponding RoCE protocol version, and the link transmission rate.
[0055] Furthermore, the cross-domain reachability information may include a `reachable_domains` field, representing the range of target domains accessible to the current node; the candidate relay information may include a `candidate_relay_nodes` field, representing the set of candidate relay nodes corresponding to the current node; and the topology egress information may include an `egress_points` field, representing the egress node information of the current node during cross-domain transmission. Additionally, the access description may include an `error_rate` field, a `profile_version` field, and a `ttl_s` field. The `error_rate` field can represent the error rate information corresponding to a node or link, the `profile_version` field can represent the version information corresponding to the current access description, and the `ttl_s` field can represent the effective time length of the current access description, thereby ensuring the timeliness of the access status information.
[0056] In one embodiment, the `accelerator_type` field identifies the type of acceleration card corresponding to the node, including different types of heterogeneous acceleration devices such as GPUs, NPUs, and DPUs; the `rdma_capable` field identifies whether the current node has RDMA low-latency direct transmission communication capability; and the `roce_version` field identifies the RoCE protocol version corresponding to the node, including different protocol types such as RoCEv1 and RoCEv2. By uniformly describing the above hardware capability parameters, basic data support can be provided for subsequent path candidate filtering and communication compatibility determination.
[0057] Furthermore, in some implementations, the access description is collected and reported by a node agent deployed on the node side or a control plane acquisition component. Specifically, during new node access, changes in node capabilities, changes in link topology, and periodic status checks, the status information of the corresponding node can be re-collected, and an access description update operation can be performed. For example, when the acceleration device type corresponding to a node changes, the link rate changes, or the domain to which the node belongs migrates, the corresponding access description can be regenerated and synchronously updated to the control plane to ensure that the status information used in subsequent hierarchical topology construction and path decision-making processes remains real-time and valid.
[0058] Furthermore, in some implementations, path candidate filtering can be performed based on the hardware capability information in the access description. For example, during the generation of a collection communication path, the accelerator type of the node can be checked for consistency based on the accelerator_type field to avoid a decrease in synchronization efficiency due to differences in communication protocols between different types of acceleration devices; when generating an RDMA communication path, nodes that do not support RDMA communication capabilities can be filtered based on the rdma_capable field; when generating a RoCE network path, the RoCE protocol version can be checked for compatibility based on the roce_version field to avoid communication anomalies caused by mixed RoCEv1 and RoCEv2 paths.
[0059] Furthermore, in some implementations, the subsequent path scoring process can be dynamically adjusted based on the link capability information and error rate information in the access description. For example, when the link corresponding to a node has a higher link speed or a lower error rate, the link quality score or reliability score of the corresponding candidate path can be increased; when a node has a high error rate or restricted cross-domain exit, the overall score of the corresponding path can be decreased, thereby improving the matching degree between the path generation result and the actual network state.
[0060] In the above implementation, by introducing node identity information, domain affiliation information, heterogeneous acceleration capability information, RDMA communication capability information, RoCE protocol version information, and cross-domain reachability information into the access description, not only can the completeness and accuracy of node state description in cross-domain AI computing power networks be improved, but also a unified data foundation can be provided for subsequent hierarchical topology construction, hardware compatibility filtering, cross-domain path generation, and dynamic path scheduling, thereby improving the accuracy of path generation, communication compatibility, and network operation stability in cross-domain AI task scenarios.
[0061] Furthermore, in one embodiment, the hardware compatibility constraint includes at least one of the following:
[0062] Constraints include accelerator type consistency, RDMA capability, RoCE protocol version compatibility, and lossless network configuration integrity.
[0063] Specifically, before generating candidate paths based on the hierarchical topology map, compatibility checks can be performed on the nodes and links in the topology based on the hardware capability information in the access description, thereby filtering out nodes, links, or path segments that do not meet the communication conditions. Since cross-domain AI computing power networks typically contain various heterogeneous computing devices such as GPUs, NPUs, and DPUs, and the RDMA capabilities, RoCE protocol versions, and lossless network configurations of different nodes may also differ, the lack of compatibility constraint filtering can easily lead to problems such as inconsistent communication protocols, mismatched network capabilities, or missing lossless transmission capabilities in the paths, thus affecting the data transmission efficiency and operational stability of AI tasks.
[0064] Furthermore, the accelerator type consistency constraint can be used to verify the consistency of node accelerator types in aggregated communication scenarios. Specifically, after identifying that the current data stream belongs to aggregated communication mode, device type matching can be performed on the communication endpoint nodes in the candidate path based on the accelerator_type field in the access description. For example, in multi-node synchronization scenarios such as AllReduce or AllGather, it can be required that all nodes in the candidate path use the same type of GPU or NPU device to avoid decreased synchronization efficiency or communication anomalies due to differences in communication protocols, computing power, or data exchange mechanisms between different heterogeneous acceleration devices. Nodes or paths that do not meet the accelerator type consistency condition can be directly excluded from the candidate path set.
[0065] Furthermore, the RDMA capability constraint can be used to verify the RDMA direct transmission capability of nodes in the candidate path. Specifically, based on the `rdma_capable` field in the access description, it can be determined whether the nodes in the candidate path have RDMA low-latency communication capability. When the AI communication mode corresponding to the data stream to be transmitted requires RDMA direct transmission communication, nodes that do not have RDMA capability can be removed from the candidate path, thereby avoiding protocol degradation or additional protocol conversion in the data stream. By filtering with RDMA capability, the data transmission efficiency in large-scale AI training scenarios can be improved, and the resource consumption caused by CPU participation in data copying can be reduced.
[0066] Furthermore, the RoCE protocol version compatibility constraint can be used to verify the compatibility of RoCE protocol versions in candidate paths. Specifically, based on the `roce_version` field in the access description, a consistency analysis can be performed on the RoCE protocol type corresponding to the nodes in the candidate path. The RoCE protocol version can include different types such as RoCEv1 and RoCEv2. Because different RoCE protocol versions differ in network encapsulation methods, routing mechanisms, and cross-domain transmission capabilities, RoCEv1 and RoCEv2 typically do not have direct interoperability. During path generation, paths with mixed RoCE protocol versions can be filtered to avoid data transmission failures or path anomalies due to protocol incompatibility.
[0067] Furthermore, the lossless network configuration integrity constraint can be used to verify the lossless network capabilities of candidate paths. Specifically, during low-latency AI communication based on RoCE networks, all switch ports in the path can be required to enable PFC mechanisms to ensure the network has complete lossless transmission capabilities. Therefore, when generating candidate paths, the switch port configurations in the path can be verified based on the switch configuration status, link lossless capability status, or lossless network configuration identifier in the access description. When it is detected that some switch ports in the path do not have PFC mechanisms enabled or the corresponding links do not have lossless transmission conditions, the corresponding path can be excluded from the candidate path set to avoid communication retransmission, synchronization blocking, or training performance degradation due to link packet loss in AI training scenarios.
[0068] In some implementations, the hardware compatibility constraints can be executed in conjunction with cross-domain compliance constraints. Specifically, a first-stage pruning of the hierarchical topology can be performed based on cross-domain compliance constraints such as data sovereignty restrictions (matching allowed domain label lists based on data source and type), tenant isolation policies (e.g., the same tenant path must pass through a dedicated network slice or VPC), and compliance domain restrictions (based on a hybrid filtering of whitelist inclusion and blacklist exclusion based on structured compliance domain label vectors). Subsequently, a second-stage hardware compatibility filtering is performed on the remaining topology based on accelerator type consistency constraints, RDMA capability constraints, RoCE protocol version compatibility constraints, and lossless network configuration integrity constraints, thereby gradually narrowing the candidate path search range and improving path generation efficiency and path legitimacy.
[0069] Furthermore, in some implementations, the candidate path scores can be dynamically adjusted based on hardware compatibility constraints. For example, when all nodes in a path meet the RDMA direct transmission conditions, have completely identical RoCE protocol versions, and the corresponding links possess full lossless network capabilities, the link quality score and reliability score of the corresponding candidate path can be increased. When there are some nodes with weak compatibility or low cross-domain interconnection capabilities in a path, the overall score of the corresponding path can be reduced, thereby improving the matching degree between the path selection results and the actual network operating capabilities.
[0070] In the above implementation, by introducing accelerator type consistency constraints, RDMA capability constraints, RoCE protocol version compatibility constraints, and lossless network configuration integrity constraints, it is possible not only to filter out nodes and links that do not meet the communication conditions in advance during the candidate path generation stage, but also to improve the communication compatibility, low-latency transmission capability, and lossless network stability in cross-domain AI computing power networks, thereby reducing the risk of communication anomalies in AI training and cross-domain inference scenarios, and improving the overall network transmission efficiency and path operation reliability.
[0071] In one embodiment, constructing a hierarchical topology diagram comprising a node layer, a cluster layer, and a region layer based on the access description includes:
[0072] The node layer is constructed based on the connection relationship between the computing power nodes to be connected and the physical links;
[0073] The cluster layer is constructed based on the aggregation relationship between multiple computing power nodes to be connected, and a cluster-level congestion summary is generated based on the underlying link status.
[0074] The regional layer is constructed based on the cross-domain connection relationship between multiple clusters, and cross-domain cost information is generated based on the cross-domain hop count, link bandwidth, and link type.
[0075] Specifically, after obtaining the access descriptions corresponding to multiple computing power nodes, a node layer can be constructed first based on the connection relationship between the computing power nodes to be connected and the physical links. This node layer can be used to describe the fine-grained network interconnection relationship between nodes. The topology nodes in the node layer can correspond to specific computing power nodes, switching nodes, or relay nodes, while the topology edges in the node layer can correspond to the physical link connection relationship between nodes. Furthermore, the link edge attributes in the node layer can include information such as link bandwidth, link latency, link jitter, link packet loss rate, link health status, queue depth, and link load status, thereby reflecting the real-time operating status of the underlying network links. By constructing the node layer, a fine-grained description of the underlying physical interconnection relationship of the cross-domain AI computing power network can be achieved, providing basic topology support for subsequent path generation and link status awareness.
[0076] Furthermore, after constructing the node layer, a cluster layer can be built based on the aggregation relationships between multiple computing power nodes to be connected. Specifically, multiple nodes can be logically aggregated according to their cluster_id, management domain, or resource pool to form a corresponding cluster-level abstract topology. The topology nodes in the cluster layer can correspond to different clusters, and the topology edges can be used to represent the interconnection relationships between different clusters. Compared with the node layer, the cluster layer focuses more on cluster-level network capabilities and the network operation status after aggregation, thereby reducing the path calculation complexity in large-scale AI computing power network scenarios.
[0077] In one embodiment, a corresponding cluster-level congestion summary can be generated based on the underlying physical link status. Specifically, a single-link congestion index can first be calculated for each physical link corresponding to the cluster egress. The single-link congestion index can be calculated using the following formula:
[0078]
[0079] in, This represents the current link bandwidth utilization. This is the normalized baseline value corresponding to the link bandwidth utilization. The proportion of ECN tags corresponding to the current link. The congestion determination threshold corresponding to the ECN marking ratio; This represents the percentage of the current link queue depth relative to the buffer. This is the normalized baseline value corresponding to the queue depth; This is the ratio of the current link latency to the baseline latency. α is the normalized threshold corresponding to link latency; α, β, γ, and δ are the corresponding weighting coefficients. Using this method, link bandwidth utilization, congestion status, and latency status can be uniformly mapped to the corresponding congestion index.
[0080] Furthermore, after obtaining the congestion indices corresponding to multiple egress links, P95 percentile aggregation can be performed on the congestion indices of all egress links within the cluster to generate the corresponding cluster-level congestion summary. Specifically, the following formula can be used:
[0081]
[0082] By adopting the P95 percentile aggregation method, the impact of a small number of transient abnormal links on the overall cluster congestion status assessment results can be reduced, thereby improving the stability and accuracy of cluster-level network status description.
[0083] Furthermore, in some implementations, the congestion summary update can employ a combination of periodic and event-triggered methods. For example, the link status can be updated periodically according to a preset time period; simultaneously, when a change in the congestion index corresponding to any link is detected to exceed a preset threshold, a congestion summary update can be triggered immediately, thereby improving the real-time performance of network status awareness. Additionally, a debouncing interval can be set to prevent frequent changes in link status within a short period from causing continuous updates to topology attributes.
[0084] After constructing the cluster layer, a regional layer can be further constructed based on the cross-domain connectivity between multiple clusters. Specifically, the regional layer can be used to describe the cross-domain interconnection relationships between different regions, different management domains, or different operational domains. The topology nodes in the regional layer can correspond to different regions or different cross-domain network areas, and the topology edges in the regional layer can correspond to cross-domain links, cross-domain leased lines, or public network interconnection links, etc.
[0085] Furthermore, at the regional level, corresponding cross-domain cost information can be generated based on cross-domain hop count, link bandwidth, and link type. Specifically, the cross-domain cost can be quantified and calculated using the following formula:
[0086]
[0087] in, Indicates the number of cross-domain hops. Indicates the link bandwidth utilization ratio. This indicates the link type coefficient (e.g., a leased line can be set to 0.5, a VPN to 1.0, and the Internet to 2.0). Indicates the cost of differentiated management domains. , , as well as These represent the weighting coefficients of the corresponding parameters. Through this method, the hop count, bandwidth resource usage, link type differences, and cross-management domain cost differences in cross-domain paths can be comprehensively considered to form cross-domain cost information for subsequent path scoring and path selection. In one embodiment, =0.3、 =0.3、 =0.25、 =0.15.
[0088] Furthermore, in some implementations, different link types can correspond to different link type coefficients. For example, leased links, VPN links, and Internet links can each correspond to different link type coefficients to reflect the differences in stability, security, and transmission capacity between different types of cross-domain links. By introducing cross-domain cost information, the subsequent path control process can consider not only link quality but also cross-domain transmission costs and cross-domain resource consumption.
[0089] In the above implementation, by constructing a hierarchical topology graph consisting of a node layer, a cluster layer, and a regional layer, and introducing link state attributes, cluster-level congestion summaries, and cross-domain cost information at different levels, it is possible not only to improve the topology abstraction and state awareness capabilities in cross-domain AI computing power networks, but also to provide a unified topology foundation for subsequent candidate path generation, cross-level path evaluation, and dynamic path scheduling, thereby improving the path calculation efficiency, network state awareness accuracy, and path control stability in large-scale cross-domain AI computing power network scenarios.
[0090] In one embodiment, determining the AI communication mode corresponding to the data stream to be transmitted includes:
[0091] The AI communication mode is determined based on the data exchange behavior characteristics corresponding to the data stream to be transmitted.
[0092] The AI communication modes include at least the collection communication mode, parameter synchronization mode, model fragment transmission mode, and cross-domain inference service mode.
[0093] In one embodiment, determining the AI communication mode corresponding to the data stream to be transmitted includes: determining the AI communication mode based on the data exchange behavior characteristics corresponding to the data stream to be transmitted; wherein, the AI communication mode includes at least a collection communication mode, a parameter synchronization mode, a model fragmentation transmission mode, and a cross-domain inference service mode.
[0094] Specifically, upon receiving the data stream to be transmitted, the data exchange behavior characteristics corresponding to the data stream can be analyzed to identify the corresponding AI communication mode. These data exchange behavior characteristics may include the direction of data interaction between nodes, the number of participating nodes, the scale of data transmission, synchronization frequency, communication timing relationships, data broadcasting behavior, data aggregation behavior, and cross-domain access behavior. By identifying these behavioral characteristics, the data exchange needs corresponding to different AI business scenarios can be distinguished, thereby providing a communication semantic basis for subsequent path structure generation and path scoring processes.
[0095] In one embodiment, generating the corresponding candidate path structure based on the AI communication mode and the cropped hierarchical topology graph includes:
[0096] When the AI communication mode is a collection communication mode, a ring path structure or a tree path structure is generated;
[0097] When the AI communication mode is parameter synchronization mode, a convergence path structure or a fan-out path structure is generated;
[0098] When the AI communication mode is the model fragmentation transmission mode, a chain path structure is generated;
[0099] When the AI communication mode is the cross-domain inference service mode, a sparse fully connected path structure is generated.
[0100] Specifically, such as Figure 3 As shown, after filtering for cross-domain compliance constraints and hardware compatibility constraints, a corresponding candidate path structure can be generated based on the trimmed hierarchical topology graph. Since the data exchange behavior varies significantly across different AI communication modes, corresponding data exchange topologies can be generated for different communication modes to improve the matching degree between network resources and business communication needs.
[0101] In one embodiment, when the AI communication mode is a ensemble communication mode, a ring-shaped path structure or a tree-shaped path structure can be generated. Specifically, when the corresponding communication scenario is AllReduce (global reduction synchronization), a logical ring structure with connected nodes can be generated according to the synchronization relationship between multiple training nodes, allowing multiple nodes to perform gradient propagation and gradient aggregation sequentially, thereby reducing the bandwidth bottleneck problem caused by the central aggregation node. Furthermore, in the process of generating the ring-shaped path structure, node connection relationships with high link bandwidth consistency, small link latency differences, and high link reliability can be prioritized to reduce the overall synchronization performance degradation caused by weak links.
[0102] Furthermore, when the aggregated communication mode corresponds to AllGather (global aggregate broadcast) or ReduceScatter (reduction distribution) scenarios, a tree-like path structure can be generated. Specifically, multi-level broadcast trees or aggregation trees can be constructed based on the hierarchical relationship between nodes, enabling hierarchical broadcasting or aggregation relationships between upper-level nodes and lower-level nodes, thereby reducing data broadcasting overhead and link duplication issues in large-scale node scenarios. Furthermore, during the generation of the tree-like path structure, the number of child nodes corresponding to a single node can be limited to avoid excessive load on some intermediate nodes leading to local link congestion.
[0103] In one embodiment, when the AI communication mode is parameter synchronization mode, a convergent path structure or a fan-out path structure can be generated. Specifically, during the gradient push phase, a convergent path structure can be generated, causing the data flows corresponding to multiple worker nodes to converge towards the parameter server entry node, thereby supporting centralized gradient synchronization in distributed training scenarios. Furthermore, when generating the convergent path structure, exchange nodes with high entry bandwidth capabilities and low congestion levels can be preferentially selected as aggregation entry points to reduce network hotspot issues caused by centralized parameter synchronization.
[0104] Furthermore, during the parameter fetching phase, a fan-out path structure or a multicast distribution structure can be generated, enabling the parameter server to broadcast the updated model parameters to multiple worker nodes simultaneously. By adopting a fan-out path structure, the number of repeated link forwardings can be reduced, and the parameter synchronization efficiency in large-scale training scenarios can be improved.
[0105] In one embodiment, when the AI communication mode is a model-sharded transmission mode (pipeline execution), a chain-like path structure can be generated. Specifically, in pipelined parallel training scenarios or model-sharded execution scenarios, a sequential chain-like data transmission structure can be generated based on the execution order relationship between multiple Stage nodes, so that the output of the previous Stage node serves as the input data for the next Stage node. Furthermore, during the generation of the chain-like path structure, node connections with higher link stability, lower latency fluctuations, and better link continuity can be prioritized to reduce data blocking issues caused by link fluctuations during pipeline execution.
[0106] In one embodiment, when the AI communication mode is a cross-domain inference service mode, a sparse fully connected path structure can be generated. Specifically, in the parallel MoE (Mixture of Experts) model scenario, different tokens may need to be dynamically routed to different expert nodes, so the data exchange relationship between different nodes usually has dynamic changing characteristics. Based on this, the connection relationship between different nodes can be dynamically established according to the token routing results, so that some nodes form an on-demand data exchange structure, rather than a fixed fully connected structure. By generating a sparse fully connected path structure, invalid link occupation can be reduced, and network resource utilization and dynamic forwarding efficiency in cross-domain inference scenarios can be improved.
[0107] Furthermore, in some implementations, during the generation of candidate path structures, candidate paths can be further screened by incorporating path structure constraints. For example, the maximum number of cross-domain hops in a path can be limited, the number of shared risk links can be limited, the number of duplicate nodes in a path can be limited, or the participation ratio of highly congested links can be limited, thereby reducing the risk of resource contention and link conflict among candidate paths.
[0108] Furthermore, in some implementations, a corresponding backup path can be generated simultaneously for the target path. Specifically, the degree of link overlap between the main path and the backup path can be constrained based on the degree of association of shared risk links, so that the backup path maintains risk isolation from the main path as much as possible. When the main path experiences link anomalies, node failures, or cross-domain congestion, it can quickly switch to the corresponding backup path, thereby improving the continuity of cross-domain AI service operations and fault recovery capabilities.
[0109] In the above implementation, by generating corresponding candidate path structures for different AI communication modes, it is possible not only to improve the matching degree between the path structure and the semantics of AI business communication, but also to reduce problems such as link resource conflicts, centralized network hotspots and cross-domain transmission bottlenecks, thereby improving the data transmission efficiency, network resource utilization and business operation stability in cross-domain AI computing power networks.
[0110] In one embodiment, scoring the candidate path structure based on network state parameters and task context parameters includes:
[0111] The configuration includes initial scoring weights for link quality, cross-domain cost, reliability, and policy constraints;
[0112] The initial scoring weights are dynamically adjusted based on task context parameters;
[0113] The candidate path structures are scored and ranked using adjusted scoring weights to determine the target path.
[0114] Specifically, after generating multiple candidate path structures, different candidate paths can be comprehensively scored based on network state parameters and task context parameters to determine the target path that meets the current AI business needs. The network state parameters can include link bandwidth, link latency, link jitter, link packet loss rate, link health status, link congestion level, and cross-domain transmission cost, etc.; the task context parameters can include model parameter size, training stage, node size, number of cross-domain requests, task type, and business real-time requirements, etc. By jointly analyzing the path operation status and task operation context, the matching degree between the path scoring results and actual business needs can be improved.
[0115] In one embodiment, initial scoring weights, including link quality, cross-domain cost, reliability, and policy constraints, can be configured first. Specifically, different initial weight vectors can be pre-set based on the network requirement characteristics corresponding to different AI communication modes. The scoring term Q(p) for link quality can reflect the bandwidth capacity, latency, jitter, and packet loss in the candidate path; the scoring term C(p) for cross-domain cost can reflect the cross-domain hop count, bandwidth resource usage, and link type differences in the path; the scoring term R(p) for reliability can reflect link stability, link health, and failure risk; and the scoring term P(p) for policy constraints can reflect the degree of matching between the path and tenant isolation policies, data sovereignty restrictions, and compliance domain restrictions.
[0116] Furthermore, in some embodiments, the comprehensive score of the candidate path can be calculated using the following scoring formula:
[0117]
[0118] in, This indicates the scoring weight corresponding to link quality. This indicates the scoring weight corresponding to the cross-domain cost. This indicates the scoring weight corresponding to reliability. This represents the scoring weight corresponding to the policy constraints. Using this method, the overall performance of different candidate paths can be uniformly and quantitatively evaluated.
[0119] In one embodiment, different AI communication modes can correspond to different initial scoring weight configurations. For example, the aggregation communication mode typically focuses more on link bandwidth consistency and synchronization efficiency, thus increasing the scoring weight corresponding to link quality; the parameter synchronization mode typically focuses more on link reliability and parameter synchronization stability, thus increasing the scoring weight corresponding to reliability; and the cross-domain inference service mode focuses more on cross-domain latency and request success rate, thus simultaneously increasing the scoring weight corresponding to cross-domain cost and reliability. By configuring differentiated initial weights for different AI communication modes, the adaptability of the path scoring process to AI business semantics can be improved.
[0120] Furthermore, in some implementations, different AI communication modes can use different initial weight vectors, and the scoring weights can be dynamically adjusted based on task context parameters through a three-level adaptive mechanism.
[0121] Specifically, different initial scoring weights can be configured for the collection communication stream, parameter synchronization stream, model sharding pull stream, and cross-domain inference service stream.
[0122] Link quality weights corresponding to aggregated communication flows It can be set to 0.40, cross-domain cost weight. It can be set to 0.15, reliability weight. It can be set to 0.20, which is the policy constraint weight. It can be set to 0.25;
[0123] Link quality weights corresponding to parameter synchronization streams It can be set to 0.30, cross-domain cost weight. It can be set to 0.15, reliability weight. It can be set to 0.30, which is the policy constraint weight. It can be set to 0.25;
[0124] Link quality weights corresponding to model sharded pull streams It can be set to 0.45, cross-domain cost weight. It can be set to 0.15, reliability weight. It can be set to 0.15, which is the policy constraint weight. It can be set to 0.25;
[0125] Link quality weights corresponding to cross-domain inference service flows It can be set to 0.20, cross-domain cost weight. It can be set to 0.30, reliability weight. It can be set to 0.30, which is the policy constraint weight. It can be set to 0.20.
[0126] Furthermore, in the context of aggregated communication, since the overall communication time during the AllReduce communication process is usually determined by the link corresponding to the slowest node, the scoring weight corresponding to link quality can be increased to enhance the path scoring process's ability to perceive the link bottleneck effect. In the context of parameter synchronization, since gradient loss may affect the convergence stability of subsequent model training, the scoring weight corresponding to reliability can be increased. In the context of cross-domain inference services, since inference services usually have high requirements for tail latency and request success rate, the scoring weight corresponding to cross-domain cost and reliability can be increased simultaneously.
[0127] Furthermore, in some implementations, the scoring weights can be subjected to a first-level adaptive adjustment based on the number of model parameters φ. Specifically, the link quality weight and reliability weight can be dynamically adjusted using a non-linear mapping method based on the number of model parameters φ. The link quality weight can be adjusted using the following formula:
[0128]
[0129] Where φ represents the number of parameters in the current model. Here, λ represents the parameters of the baseline model, and λ represents the link quality weight adjustment coefficient (λ = 0.05-0.10). Furthermore, the reliability weight can be adjusted using the following formula:
[0130]
[0131] Where μ represents the reliability weight adjustment coefficient (μ=0.03-0.05). This method can increase the focus on link quality and reliability in the path scoring process as the model size increases.
[0132] Furthermore, in some implementations, a second-level adaptive adjustment of the scoring weights can be performed based on the training phase. Specifically, during the training warm-up phase, the scoring weights corresponding to link quality and reliability can be increased to improve network stability in the early stages of training; during the steady-state training phase, the initial weight configuration can be restored; and during the training decay phase, the scoring weights corresponding to link quality can be further increased to reduce the impact of link fluctuations on training results in the later stages of training.
[0133] Furthermore, in some implementations, the change in scoring weights during the training phase switching process can be transitioned using an exponential smoothing method, specifically calculated using the following formula:
[0134]
[0135] in, This indicates the target weight corresponding to the target stage. This represents the current weight corresponding to the current stage, and τ represents the sampling period parameter corresponding to the smooth transition (τ = 3-5 sampling periods). By using the above method, the impact of sudden changes in scoring weights during training stage switching on path scheduling stability can be reduced.
[0136] Furthermore, in some implementations, a third-level adaptive adjustment of the scoring weights can be performed based on the cluster size and the number of cross-domain requests. For example, as the number of GPU nodes K increases, the scoring weights corresponding to link quality can be increased to reduce the impact of bottleneck links in large-scale aggregated communication scenarios; as the number of domains N_domain increases, the scoring weights corresponding to cross-domain costs and policy constraints can be increased to reduce cross-domain communication costs and cross-domain compliance risks.
[0137] Furthermore, in some implementations, the scoring weights can be dynamically optimized in a closed loop based on a reinforcement learning mechanism. Specifically, a reinforcement learning mechanism can be used, with the actual performance of the path as the reward signal, to continuously optimize and adjust the scoring weights. The reward signal can include parameters such as actual path latency, tail latency stability, link congestion level, service success rate, and path reconstruction frequency. During the cold start phase, a rule-based adaptive weight formula can be used for initialization, and the process can gradually switch to a reinforcement learning-driven dynamic optimization mechanism, thereby improving the adaptability and long-term optimization capability of the path scoring process in complex cross-domain AI network scenarios.
[0138] Furthermore, in some implementations, for the aggregated communication mode, a short-board evaluation method can be used for the link quality item Q(p). Specifically, since the overall communication time in the AllReduce scenario is usually limited by the slowest link, the link quality can be calculated using the following formula:
[0139]
[0140] in, to This indicates the bandwidth capacity of each link in the candidate path. This indicates the target bandwidth requirement for the current task. By using the weakest link bandwidth as the basis for link quality assessment, the accuracy of path scoring results in aggregated communication scenarios can be improved.
[0141] In one embodiment, after dynamically adjusting the scoring weights, the adjusted scoring weights can be used to score and rank multiple candidate path structures. Specifically, a comprehensive score value can be calculated for each candidate path, and sorting can be performed according to the scoring results. Subsequently, the candidate path with the highest score or that meets preset conditions can be selected as the target path, and path control information can be sent to the corresponding network device.
[0142] Furthermore, in some implementations, the scoring weights can be continuously optimized by combining historical path operation results, runtime telemetry feedback results, or task execution results. For example, the corresponding scoring weights can be adaptively adjusted based on the actual path transmission performance, tail delay stability, or path reconstruction frequency, thereby improving the accuracy and stability of subsequent path scoring processes.
[0143] In the above implementation, by introducing a multi-dimensional scoring mechanism that includes link quality, cross-domain cost, reliability, and policy constraints, and by dynamically adjusting the scoring weights in conjunction with task context parameters such as model parameter size, training stage, cluster size, and number of cross-domain operations, it is possible not only to improve the ability of the path scoring process to perceive AI business semantics and network status, but also to improve the accuracy and stability of target path selection results under different AI business scenarios, thereby improving link utilization, data transmission efficiency, and business operation reliability in cross-domain AI computing power networks.
[0144] Furthermore, in one embodiment, the scoring of the candidate path structure based on network state parameters and task context parameters further includes:
[0145] Based on historical link status data, short-term predictions are made on link latency and ECN marking rate, and corresponding link degradation probabilities are generated based on the prediction results.
[0146] A prediction penalty term is applied to the scoring results corresponding to the candidate paths based on the link degradation probability.
[0147] Construct a path interference matrix between different candidate paths, and generate corresponding path resource contention penalty terms based on the path interference matrix;
[0148] Based on the predicted penalty term and the path resource competition penalty term, the candidate path structure is comprehensively scored and adjusted.
[0149] Specifically, in some implementations, in order to improve the ability of the candidate path scoring process to perceive future network state change trends, a link state prediction mechanism and a multi-path resource contention perception mechanism can be introduced when performing candidate path scoring. This avoids the path scoring results from relying solely on the current network state, which could lead to rapid congestion or performance degradation of the path during subsequent operation.
[0150] In one embodiment, short-term predictions of link latency and ECN marking rate can be made based on historical link status data, and the corresponding link degradation probability can be generated based on the prediction results. Specifically, historical operational status data of links corresponding to candidate paths can be continuously collected. This historical link status data may include sequences of link latency changes, sequences of ECN marking ratio changes, sequences of link queue depth changes, and sequences of link packet loss rate changes. Subsequently, a lightweight time series prediction model can be used based on the historical link status data to predict the link status change trend within a preset time range in the future.
[0151] Furthermore, in some implementations, short-term predictions can be made of link latency and ECN marking rate within the range of 5 to 60 seconds in the future to identify potential link congestion trends or link performance degradation trends in advance. For example, when the prediction results show that the link latency will continue to increase or the ECN marking rate will continue to increase, it can be determined that the current link has a potential congestion risk, and a link degradation probability can be generated based on the corresponding prediction results. The link degradation probability can be used to represent the possibility that the target link will experience performance degradation, link congestion, or service quality degradation within a future time window.
[0152] Furthermore, in some implementations, accuracy verification can be performed on the link state prediction results. For example, the prediction results can be compared with the subsequent actual link states, and the reliability of the current prediction model can be evaluated based on the prediction error. When the prediction accuracy is lower than a preset benchmark value, the influence weight of the corresponding prediction result in the path scoring can be automatically reduced, or the prediction penalty mechanism can be directly turned off to avoid low-precision prediction results interfering with the path scoring process.
[0153] In one embodiment, a prediction penalty term can be applied to the scoring results corresponding to candidate paths based on the link degradation probability. Specifically, when the link degradation probability of a candidate path is high, the overall scoring result corresponding to that candidate path can be downweighted, thereby reducing the probability that a path that may experience congestion or performance degradation will be selected in the future. Furthermore, in some implementations, the strength of the prediction penalty term can be dynamically adjusted according to the magnitude of the link degradation probability. For example, when the link degradation probability continues to increase, the prediction penalty weight of the corresponding path can be gradually increased to improve the path scoring process's ability to avoid potentially risky paths.
[0154] In one embodiment, a path interference matrix can be constructed between different candidate paths, and a corresponding path resource competition penalty term can be generated based on the path interference matrix. Specifically, during the simultaneous execution of multiple AI tasks, different candidate paths may share the same link resources, exchange node resources, or cross-domain exit resources, resulting in resource competition between different paths. Based on this, the degree of resource overlap between different candidate paths can be analyzed, and a corresponding path interference matrix can be constructed.
[0155] Furthermore, in some embodiments, the path interference matrix can be constructed using the following formula:
[0156]
[0157] in, Representing a path The corresponding set of links, Representing a path The corresponding set of links, Representing a path With path The degree of resource interference between different candidate paths can be quantified using the above method, thereby assessing the risk of resource competition between different paths.
[0158] Furthermore, in some implementations, when the resource interference between two candidate paths is high, a path resource contention penalty term can be applied to the corresponding candidate paths to reduce the probability that paths sharing high-risk link resources are selected simultaneously. For example, when multiple training jobs execute AllReduce communication simultaneously, if multiple candidate paths pass through the same set of high-load links at the same time, it may lead to serious link contention problems in the aggregated communication scenario. By introducing a path resource contention penalty term, the probability of link resource conflicts between different AI tasks can be reduced, and the overall network resource utilization balance can be improved.
[0159] In one embodiment, the candidate path structure can be comprehensively scored and adjusted based on the predicted penalty term and the path resource contention penalty term. Specifically, the predicted penalty term and the path resource contention penalty term can be jointly introduced into the comprehensive scoring model based on the original path scoring result to generate the final path scoring result. Furthermore, in some implementations, the influence weights corresponding to different penalty terms can be dynamically adjusted according to the current network load status, AI task scale, or cross-domain resource consumption, thereby improving the matching degree between the path scoring result and the actual network operating status.
[0160] Furthermore, in some implementations, the prediction penalty mechanism and the path resource contention penalty mechanism can be dynamically optimized by incorporating runtime telemetry feedback results. For example, the link state prediction model parameters, path resource contention weights, or path interference matrix update cycles can be adjusted based on historical path performance results, thereby improving prediction accuracy and resource contention awareness in subsequent path scoring processes.
[0161] In the above embodiments, by introducing a link state prediction mechanism and a multi-path resource contention awareness mechanism, not only can the predictive ability of the candidate path scoring process to predict future link state change trends be improved, but the risk of link resource contention in multiple concurrent AI task scenarios can also be reduced, thereby improving the path stability, link utilization and overall network scheduling efficiency of the cross-domain AI computing network.
[0162] In one embodiment, performing local path reconstruction or global path reconstruction based on the joint determination result includes:
[0163] A joint determination is made based on the telemetry indicators to obtain a comprehensive determination result corresponding to the link congestion status.
[0164] When the comprehensive judgment result meets the preset conditions, a partial replacement is performed on a portion of the path segment in the target path;
[0165] If local path reconstruction fails or the network state meets the conditions for global reconstruction, the target path is recalculated globally.
[0166] Specifically, after the target path is put into operation, network telemetry metrics and service operation metrics along the corresponding path can be continuously collected to monitor the current path's operational status in real time. These telemetry metrics may include parameters such as link latency, link jitter, link packet loss rate, ECN marking ratio, congestion queue depth, link health status, service-side tail latency, and service success rate. These metrics can be acquired in real time through switch telemetry interfaces, node monitoring components, link detection modules, or control plane status acquisition modules, and then uploaded to the path control module for unified analysis.
[0167] In one embodiment, a joint determination can be made based on the telemetry indicators to obtain a comprehensive determination result corresponding to the link congestion state. Specifically, different telemetry indicators can be collected and normalized within a sliding time window. The normalized indicator corresponding to the ECN label ratio can be calculated using the following formula:
[0168]
[0169] The normalized metric corresponding to queue depth can be calculated using the following formula:
[0170]
[0171] The normalized index corresponding to the P95 latency can be calculated using the following formula:
[0172]
[0173] The normalized metric corresponding to the packet loss rate can be calculated using the following formula:
[0174]
[0175] in, This indicates the proportion of ECN tags corresponding to the current link. Indicates the queue depth corresponding to the current link. Indicates the maximum depth of the queue. This indicates the P95 latency corresponding to the current link. Indicates the reference P95 delay. This represents the packet loss rate of the current link. By using the above method, telemetry metrics from different dimensions can be uniformly mapped to a standardized range, thereby improving comparability in subsequent joint judgment processes.
[0176] Furthermore, in some implementations, a weighted voting joint determination can be performed based on multiple normalized indicators to generate a corresponding comprehensive determination result. Specifically, the following formula can be used to jointly calculate multiple telemetry indicators:
[0177]
[0178] in, This represents the weighting coefficient corresponding to the proportion of ECN tags. This represents the weight coefficient corresponding to the queue depth. This represents the weighting coefficient corresponding to the P95 delay. This represents the weighting coefficients corresponding to the packet loss rate. In one embodiment, W1=0.30, W2=0.25, W3=0.25, and W4=0.20. Since the ECN labeling ratio can reflect the congestion trend in the RoCE network relatively early, the weight of the ECN indicator can be increased; queue depth and P95 latency can reflect the degree of continuous link congestion and the change in service-side latency, so they can be set with the same weight; although the packet loss rate can directly reflect network anomalies, it usually belongs to the later stages of congestion, so its corresponding weight can be relatively low. By jointly analyzing multiple telemetry indicators, the accuracy of path congestion identification results can be improved, and the problem of misjudgment caused by anomalies in a single indicator can be reduced.
[0179] In one embodiment, the link congestion status can be classified based on a comprehensive judgment result. Specifically, such as... Figure 4 As shown, when the comprehensive judgment result S is less than 0.3, the current path status can be determined as Level 0 normal state; when the comprehensive judgment result S is in the range of 0.3 to 0.5, it can be determined as Level 1 warning state, and continuous monitoring is performed on the corresponding path; when the comprehensive judgment result S is in the range of 0.5 to 0.6, it can be determined as Level 2 intervention state, and local link adjustment or pre-path detection is performed in advance; when the comprehensive judgment result S is greater than or equal to 0.6, it can be determined as Level 3 reconstruction state, and the path reconstruction process is initiated. By adopting a hierarchical triggering mechanism, differentiated path control strategies can be implemented according to different congestion levels, thereby avoiding the direct triggering of global path reconstruction problems by minor link fluctuations.
[0180] Furthermore, in some implementations, conflict arbitration can be performed by combining the priority relationships between different indicators. For example, when there is a conflict between the judgment results corresponding to the packet loss rate indicator and the ECN indicator, the judgment result corresponding to the higher priority indicator can be given priority to improve the accuracy of congestion identification. In one embodiment, the priority order of different indicators can be set as follows: the packet loss rate indicator has a higher priority than the ECN labeling ratio indicator, the ECN labeling ratio indicator has a higher priority than the queue depth indicator, and the queue depth indicator has a higher priority than the P95 latency indicator. By introducing a multi-indicator priority arbitration mechanism, the misjudgment problem caused by short-term abnormal fluctuations in different telemetry indicators can be reduced, and the stability of path state identification in complex cross-domain AI network scenarios can be improved.
[0181] In one embodiment, when the comprehensive judgment result meets preset conditions, partial replacement can be performed on a portion of the target path. Specifically, when a fault or congestion is detected to affect only a local link, local node, or local cross-domain exit in the target path, partial path reconstruction can be performed only on the abnormal path segment without recalculating the entire target path. For example, link-level replacement can be performed on the faulty link, node-level replacement on the faulty relay node, segment-level replacement on the abnormal segment path, or domain-level replacement on the abnormal cross-domain exit. By adopting a partial path reconstruction approach, the service disturbance caused by global path recalculation can be reduced, and the path recovery time can be shortened.
[0182] Furthermore, in some implementations, the legality of the replaced path can be verified before performing local path replacement. For example, it can be checked whether the replaced path meets conditions such as cross-domain compliance constraints, hardware compatibility constraints, path acyclicity constraints, and bandwidth resource constraints. When the replaced path does not meet the corresponding constraints, the current local path replacement operation can be terminated, and the global path reconstruction process can be initiated.
[0183] In one embodiment, when local path reconstruction fails or the network state meets the conditions for global reconstruction, the target path can be recalculated globally. Specifically, when widespread link congestion, large-scale node failures, cross-domain network anomalies, policy configuration changes, or local path replacement failures are detected, the global candidate path generation, path scoring, and target path selection processes can be re-executed based on the current hierarchical topology to generate a new global target path. By performing global path reconstruction, the path recovery capability and service continuity capability of cross-domain AI computing power networks under complex and abnormal scenarios can be improved.
[0184] Furthermore, in some implementations, an atomic update mechanism for path identifiers can be used during the path reconstruction process. Specifically, based on the path identifier structure composed of Path-ID, Version, and SID-List, new version path information can be generated first, and then switching instructions can be sent to the corresponding network nodes, thereby avoiding the problem of inconsistent states of some nodes during the path switching process.
[0185] Furthermore, in some implementations, anti-oscillation control mechanisms can be introduced during path reconstruction. For example, EWMA smoothing can be used to smooth the original telemetry indicators to filter out instantaneous jitter signals; a dual-threshold determination method can be used to distinguish between entry and exit thresholds to reduce the probability of frequent switching; and a minimum hold time limit and a reconstruction frequency penalty mechanism can be used to suppress repeated path switching behavior in a short period of time, thereby improving the overall path control stability.
[0186] In the above implementation, by performing joint judgment based on multi-dimensional telemetry indicators and combining local path reconstruction and global path reconstruction mechanisms, it is possible not only to improve the accuracy of link congestion identification in cross-domain AI computing power networks, but also to achieve dynamic path recovery in scenarios of link anomalies, local faults, and cross-domain network anomalies, thereby reducing the probability of path oscillations, shortening service recovery time, and improving network stability and service continuity in cross-domain AI task scenarios.
[0187] Furthermore, in one embodiment, performing local path reconstruction or global path reconstruction includes one or more of the following:
[0188] Perform link-level replacement on the faulty link;
[0189] Perform node-level replacement on the faulty relay node;
[0190] Perform path segment replacement on a portion of the target path;
[0191] Perform domain export replacement for cross-domain exports.
[0192] Specifically, when link congestion, link anomalies, node failures, or cross-domain egress anomalies are detected at runtime, the corresponding path reconstruction method can be selected based on the scope and type of the anomaly to reduce its impact on the overall AI service transmission process. By adopting a hierarchical and granular path reconstruction mechanism, frequent global path recalculation can be avoided in local anomaly scenarios, thereby improving path recovery efficiency and network operational stability.
[0193] In one embodiment, when a partial physical link in the target path is detected to be interrupted, congested, experiencing abnormal latency, or exhibiting a continuously increasing packet loss rate, link-level replacement can be performed on the faulty link. Specifically, the corresponding abnormal link can be located in the current target path, and an alternative link corresponding to the abnormal link can be searched based on the current hierarchical topology map to bypass the path segment corresponding to the faulty link. Furthermore, during the link-level replacement process, alternative links with higher bandwidth capacity, lower latency, and more stable health status can be prioritized, thereby reducing the impact of local link failures on the overall communication efficiency of the AI task.
[0194] In one embodiment, when a relay node in the target path is detected to have failed, has abnormal load, reduced forwarding capability, or abnormal network capability, a node-level replacement can be performed on the faulty relay node. Specifically, the faulty relay node in the current target path can first be located, and the corresponding candidate relay node set can be determined based on the candidate_relay_nodes field in the access description. Subsequently, a replacement relay node that meets the current communication constraints can be selected from the candidate relay node set, and the connection relationship between the relay node and the preceding and following path segments can be re-established to restore the data forwarding capability of the corresponding path.
[0195] Furthermore, in some implementations, during the node-level replacement process, compatibility checks can be performed on the candidate relay node's RDMA capabilities, RoCE protocol version, link bandwidth capabilities, and cross-domain reachability to avoid the replaced relay node not meeting the network capability requirements corresponding to the current AI communication scenario.
[0196] In one embodiment, when the link area corresponding to a portion of the target path experiences persistent local congestion, cross-domain latency anomalies, or routing forwarding anomalies, path segment replacement can be performed on that portion of the target path. Specifically, the target path can be divided into multiple path segments, and the target path segments with abnormal states can be identified. Subsequently, only the forwarding paths corresponding to the abnormal path segments can be recalculated, while other normal path segments continue to maintain their original forwarding states, thereby reducing service disruptions caused by global path switching.
[0197] Furthermore, in some implementations, the path segment replacement can be performed based on the path segment information corresponding to the path identifier. For example, when using an SRv6 SID list or segmented routing mechanism for path control, only the SID segment corresponding to the abnormal path segment can be replaced, without modifying all forwarding paths in the complete path. By adopting a path segment-level local replacement method, path recovery time can be shortened and control plane overhead during large-scale path reconstruction can be reduced.
[0198] In one embodiment, when an anomaly in the cross-domain egress link, cross-domain egress congestion, cross-domain policy changes, or an anomaly in the target domain network status is detected, a domain egress replacement can be performed on the cross-domain egress. Specifically, based on the current geographic layer topology, a new cross-domain egress node or cross-domain boundary link corresponding to the target domain can be selected, and a new cross-domain forwarding path can be established to bypass the currently abnormal cross-domain egress area. Furthermore, during the domain egress replacement process, the cross-domain cost, link reliability, data sovereignty restrictions, and tenant isolation policies corresponding to the new cross-domain egress can be jointly verified to ensure that the replaced cross-domain path still meets the cross-domain compliance requirements of the current business scenario.
[0199] Furthermore, in some implementations, after performing link-level replacement, node-level replacement, path segment replacement, or domain egress replacement, path validity verification and path keep-alive detection can be performed on the replaced path. For example, it can be verified whether the replaced path meets acyclic constraints, bandwidth constraints, cross-domain constraints, and hardware compatibility constraints, and the availability of the replaced path can be verified through BFD probing or link keep-alive detection. When the detection results meet preset conditions, the formal path switching operation is then performed.
[0200] Furthermore, in some implementations, when local path reconstruction fails, the replaced path cannot meet business needs, or the scope of abnormal impact exceeds a preset threshold, a global path reconstruction process can be further triggered to regenerate candidate paths, score paths, and select target paths for the entire target path, thereby improving the overall network recovery capability under complex abnormal scenarios.
[0201] In the above implementation, by introducing multi-granularity path reconstruction mechanisms such as link-level replacement, node-level replacement, path segment replacement, and domain exit replacement, it is possible not only to improve the recovery efficiency of abnormal paths in cross-domain AI computing power networks, but also to reduce global business disturbances in local abnormal scenarios, thereby improving data transmission stability, network reliability, and business continuity capabilities in AI task scenarios.
[0202] Furthermore, in one embodiment, after performing local path reconstruction or global path reconstruction, the method further includes:
[0203] The telemetry indicators are smoothed by an exponentially weighted moving average to obtain smoothed path status indicators.
[0204] The smoothed path status index is determined by a dual threshold based on the entry threshold and the exit threshold.
[0205] Apply a penalty decay limit to the corresponding path based on the number of historical path reconstructions.
[0206] After the path reconstruction is completed, a minimum hold time limit is imposed on the reconstructed target path.
[0207] Specifically, during the operation of cross-domain AI computing networks, the link status, switching queue status, and service traffic status may experience momentary fluctuations. Therefore, if path reconstruction is triggered solely based on a single telemetry result, it can easily lead to frequent path switching, network status oscillations, and decreased service stability. Based on this, in some implementations, a multi-layered anti-oscillation stabilization control mechanism can be introduced after performing local or global path reconstruction to reduce the impact of momentary anomalies on the path control process and improve the overall stability of the path reconstruction process.
[0208] In one embodiment, the telemetry metrics can first be smoothed using an exponentially weighted moving average to obtain smoothed path status metrics. Specifically, smoothing calculations can be performed on telemetry metrics such as link latency, link jitter, ECN tagging ratio, queue depth, and packet loss rate to reduce the impact of short-term instantaneous fluctuations or sampling spikes on the path status determination results. Further, in some embodiments, the telemetry metrics can be processed using an exponentially weighted moving average using the following formula:
[0209]
[0210] in, This represents the telemetry index value corresponding to the current sampling period. This represents the smoothing result corresponding to the previous period. This represents the current smoothed index result, where α represents the smoothing coefficient. In one embodiment, α can be set to 0.3 to reduce the interference of instantaneous outliers on the overall path state determination result while ensuring the response speed to state changes.
[0211] Furthermore, after obtaining the smoothed path status indicators, a dual-threshold determination can be performed on the smoothed path status indicators based on an entry threshold and an exit threshold. Specifically, an abnormal path entry threshold can be set separately. and abnormal exit threshold Among them, when the path status indicator exceeds the entry threshold... When the path status indicator drops to the exit threshold, it can be determined that the current path has entered an abnormal state; If the following occurs, it can be determined that the current path has returned to normal.
[0212] Furthermore, in some implementations, an exit threshold is defined. The calculation can be performed in the following way:
[0213]
[0214] in, The threshold attenuation coefficient is shown. By setting the exit threshold lower than the entry threshold, the frequent path switching problem caused by repeated fluctuations in the path state around the threshold can be avoided. Furthermore, different attenuation coefficients can be used in different AI communication modes. For example, in the case of aggregated communication, a lower attenuation coefficient can be used due to the high requirements for link stability; while in the case of cross-domain inference services, a relatively higher attenuation coefficient can be used to improve the responsiveness to latency changes.
[0215] In one embodiment, a penalty decay limit can be applied to a corresponding path based on the number of historical path reconstructions. Specifically, the number of historical reconstructions of the target path within a preset time window can be recorded, and a penalty value can be added to the corresponding path after each path reconstruction. When a path is frequently reconstructed in a short period of time, the penalty value of the corresponding path can be gradually increased, thereby reducing the probability of triggering path switching again. Furthermore, in some implementations, the BGP damping mechanism can be adopted, setting a half-life for the penalty value and gradually decaying the corresponding penalty value over time. When the accumulated penalty value exceeds a preset suppression threshold, the corresponding path can be temporarily prohibited from participating in the path reconstruction process again within a preset time range, thereby avoiding the network being in a state of frequent oscillations for a long time.
[0216] Furthermore, in some implementations, the half-life corresponding to the penalty value can be set to 30 to 60 seconds, and the suppression threshold can be set to 2000. When the penalty value corresponding to a path exceeds the suppression threshold, the path switching capability of the corresponding path can be temporarily frozen, and the status of the corresponding link can be continuously monitored; when the penalty value decays to below the recovery threshold, the corresponding path can be allowed to participate in the path scheduling process again.
[0217] In one embodiment, after path reconstruction is completed, a minimum hold time limit can be imposed on the reconstructed target path. Specifically, after path switching is completed, the current target path can be restricted from performing reconstruction operations again within a preset time window, thereby avoiding continuous path switching problems caused by short-term fluctuations in link status. Furthermore, different minimum hold times can be configured for different AI communication scenarios. For example, in aggregated communication scenarios, since the multi-node synchronization process has high requirements for path stability, a longer hold time can be set; in inference service scenarios, a relatively shorter hold time can be set to improve path scheduling flexibility.
[0218] Furthermore, in some implementations, the minimum hold time can be used in conjunction with the cooldown window. For example, the minimum hold time can be compared with the current path cooldown time, and the larger value can be taken as the final path hold time to further reduce the risk of path oscillation.
[0219] Furthermore, in some implementations, the exponentially weighted moving average smoothing, dual-threshold determination, penalty decay limit, and minimum hold time limit can be executed serially. Specifically, the original telemetry index acquisition, EWMA smoothing, dual-threshold state determination, penalty decay check, and path reconstruction can be executed sequentially, with a minimum hold time limit applied after path reconstruction is completed. By employing a multi-layered serial anti-oscillation control mechanism, the stability of the path state determination results can be improved, and the impact of frequent path reconstruction on the stability of AI service operation can be reduced.
[0220] In the above implementation, by introducing an exponentially weighted moving average smoothing mechanism, a dual threshold judgment mechanism, a penalty decay limit mechanism, and a minimum hold time limit mechanism, it is possible not only to reduce the impact of instantaneous link fluctuations or short-term congestion on path control results, but also to reduce network oscillation problems caused by frequent path switching, thereby improving the path stability, service continuity, and operational reliability of cross-domain AI computing power networks.
[0221] In one embodiment, the method further includes:
[0222] Generate at least one preliminary path for the target path;
[0223] The preliminary path is determined based on the degree of association between the target path and the preliminary path with shared risk links;
[0224] During operation, a keep-alive check is performed on the prepared path;
[0225] When an anomaly is detected in the target path, the system switches to the corresponding backup path.
[0226] Specifically, after determining the target path, at least one backup path can be generated for the target path to improve the service recovery capability of the cross-domain AI computing power network in scenarios of link anomalies, node failures, or cross-domain congestion. The backup path can represent a data transmission path used to replace the current service traffic when the target path fails. Furthermore, the backup path can be established simultaneously with the target path, or it can be dynamically generated during the operation of the target path. By pre-maintaining available backup paths, the recovery latency caused by recalculating the path after a path anomaly can be reduced.
[0227] In one embodiment, a corresponding backup path can be determined based on the degree of shared risk link correlation between the target path and the backup path. Specifically, the overlap of links, nodes, switching equipment, or cross-domain egress points between the target path and the candidate backup path can be analyzed to assess the degree of shared risk correlation between the two paths. Here, a shared risk link can represent link resources, switching equipment resources, or cross-domain egress resources that multiple paths commonly rely on. When there is a high degree of shared risk correlation between the target path and the backup path, a failure in the shared resource area may cause both the primary path and the backup path to fail simultaneously.
[0228] Therefore, in some implementations, candidate paths with a high SRLG (Shared Risk Link Group) separation degree from the target path can be preferentially selected as backup paths, thereby reducing the risk coupling between the target path and the backup path. Furthermore, in some implementations, the comprehensive score corresponding to the backup path can be required to be no less than a preset proportion of the comprehensive score of the target path, so as to ensure that the backup path still has good transmission and service carrying capacity after the handover.
[0229] In one embodiment, keep-alive checks can be performed on the prepared path during operation. Specifically, the link status, node status, and cross-domain connectivity status corresponding to the prepared path can be periodically checked using link probe messages, BFD (Bidirectional Forwarding Detection) mechanisms, or path liveness detection mechanisms to confirm the availability of the prepared path in real time. Furthermore, in some implementations, millisecond-level BFD probing can be used to perform fast keep-alive checks on the prepared path, thereby improving the rapid switching capability in abnormal scenarios.
[0230] Furthermore, in some implementations, the backup path resources can be maintained using either a hot standby mode or a cold standby mode. In hot standby mode, corresponding bandwidth and forwarding resources can be reserved in advance for the backup path, enabling it to quickly take over service traffic during a switchover. In cold standby mode, only the path information and keep-alive status of the backup path are maintained, and the full bandwidth resources are dynamically allocated during the actual switchover, thus reducing resource waste caused by long-term network resource occupation.
[0231] In one embodiment, when a link anomaly, link congestion, node failure, cross-domain egress anomaly, or path detection failure is detected on the target path, the system can switch to the corresponding backup path. Specifically, the anomaly determination of the target path status can be based on runtime telemetry metrics or path detection results; when the comprehensive determination result meets the path switching conditions, data forwarding corresponding to the current target path can be stopped, and service traffic can be switched to the corresponding backup path.
[0232] Furthermore, in some implementations, an atomic update mechanism for path identifiers can be employed during the path switching process. Specifically, new path version information for the corresponding reserve path can be generated first, and then a unified switching command can be issued to the corresponding network nodes to avoid inconsistencies in path states caused by premature switching of some nodes. After the switch is completed, the operational status of the new path can be continuously monitored to confirm the stability of service operation after the reserve path switch.
[0233] Furthermore, in some implementations, when the target path returns to normal or the network status once again meets the original path's operating conditions, a decision can be made on whether to perform a path reversal based on the current service status. For example, after the target path has continuously recovered to a stable state and meets the minimum hold time condition, service traffic can be switched back to the original target path, thereby restoring the optimal path's operating state.
[0234] In the above implementation, by introducing a pre-path generation mechanism, a shared risk correlation analysis mechanism, and a pre-path keep-alive detection mechanism, not only can the path disaster recovery capability and fault recovery capability in the cross-domain AI computing power network be improved, but the risk of service interruption in the scenario of link anomaly or node failure can also be reduced, thereby improving the data transmission stability, service continuity operation capability and cross-domain network reliability in the AI task scenario.
[0235] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0236] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices.
[0237] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 5An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0238] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.
[0239] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0240] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).
[0241] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.
[0242] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.
[0243] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0244] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0245] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0246] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0247] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0248] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0249] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0250] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural.
[0251] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A dynamic path control method for cross-domain AI computing power networks, characterized in that, The method includes: Obtain the status information of the computing power nodes to be connected and generate an access description; Based on the access description, a hierarchical topology diagram including a node layer, a cluster layer, and a region layer is constructed. Determine the AI communication mode corresponding to the data stream to be transmitted; The hierarchical topology graph is constrained and pruned based on cross-domain compliance constraints and hardware compatibility constraints. Based on the AI communication mode and the cropped hierarchical topology graph, a corresponding candidate path structure is generated. The candidate path structure is scored based on network state parameters and task context parameters to determine the target path, and path forwarding control of the corresponding service data flow is executed based on the target path. During the operation of the target path, telemetry indicators are collected for joint determination, and local path reconstruction or global path reconstruction is performed based on the joint determination results; The determination of the AI communication mode corresponding to the data stream to be transmitted includes: The AI communication mode is determined based on the data exchange behavior characteristics corresponding to the data stream to be transmitted. The AI communication modes include aggregated communication mode, parameter synchronization mode, model fragmentation transmission mode, and cross-domain inference service mode. When the AI communication mode is a collection communication mode, a ring path structure or a tree path structure is generated; When the AI communication mode is parameter synchronization mode, a convergence path structure or a fan-out path structure is generated; When the AI communication mode is the model fragmentation transmission mode, a chain path structure is generated; When the AI communication mode is the cross-domain inference service mode, a sparse fully connected path structure is generated.
2. The method according to claim 1, characterized in that, The construction of a hierarchical topology diagram based on the access description, comprising a node layer, a cluster layer, and a region layer, includes: The node layer is constructed based on the connection relationship between the computing power nodes to be connected and the physical links; The cluster layer is constructed based on the aggregation relationship between multiple computing power nodes to be connected, and a cluster-level congestion summary is generated based on the underlying link status. The regional layer is constructed based on the cross-domain connection relationship between multiple clusters, and cross-domain cost information is generated based on the cross-domain hop count, link bandwidth, and link type.
3. The method according to claim 1, characterized in that, The scoring of the candidate path structure based on network state parameters and task context parameters includes: The configuration includes initial scoring weights for link quality, cross-domain cost, reliability, and policy constraints; The initial scoring weights are dynamically adjusted based on task context parameters; The candidate path structures are scored and ranked using adjusted scoring weights to determine the target path.
4. The method according to claim 1, characterized in that, The process of performing local or global path reconstruction based on the joint decision result includes: A joint determination is made based on the telemetry indicators to obtain a comprehensive determination result corresponding to the link congestion status. When the comprehensive judgment result meets the preset conditions, a partial replacement is performed on a portion of the path segment in the target path; If local path reconstruction fails or the network state meets the conditions for global reconstruction, the target path is recalculated globally.
5. The method according to claim 1, characterized in that, The method further includes: Generate at least one preliminary path for the target path; The preliminary path is determined based on the degree of association between the target path and the preliminary path with shared risk links; During operation, a keep-alive check is performed on the prepared path; When an anomaly is detected in the target path, the system switches to the corresponding backup path.
6. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 5.
7. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Task scheduling method and device, medium and program product
CN121387489A
Adaptive network topology dynamic reconstruction method and system based on deep reinforcement learning
CN121396801A