Intelligent computing center path allocation method, data transmission method and network system

CN121283926BActive Publication Date: 2026-08-21NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511416834.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-08-21
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

这种设计方式虽然在技术实现上较为简便,但在实际应用中会引发严重的链路拥塞,链路拥塞将导致数据传输延迟增加、传输成功率下降,严重影响网络性能

Benefits of technology

[0012]在本申请实施例中,基于AI训练数据流的时序特征以及每个集合通信库在同一通信周期仅选择一种通信算法驱动数据流传输的特性,本申请实施例预先获取能够标识数据流的数据流标识,通过将同一集合通信域内不同通信算法对应的源IP地址相同的数据流标识分配同一上行链路、不同通信算法对应的目的IP地址相同的数据流标识分配同一下行链路,使得在实际通信时,同一链路上不同通信算法驱动的数据流不会同时传输,降低链路拥塞风险,并为后续相同通信算法的数据流标识预留空闲链路;以及,通过将同一集合通信域内相同通信算法对应的源IP地址或目的IP地址相同的数据流标识分配不同链路,进行多路径负载分担,确保基于相同通信算法驱动的数据流分散传输,进一步降低单链路拥塞风险,以期提升智能计算中心网络的通信效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121283926B_ABST
    Figure CN121283926B_ABST
Patent Text Reader

Abstract

The application provides an intelligent computing center path allocation method, a data transmission method and a network system. The technical solution of the application is based on the time sequence characteristics of the AI training data flow and the characteristics that each set communication library only selects one communication algorithm to drive data flow transmission in the same communication cycle, obtains a data flow identifier capable of identifying the data flow and a communication algorithm corresponding to the data flow identifier, performs path planning based on the communication algorithm, so that in actual communication, data flows driven by different communication algorithms on the same path will not be transmitted at the same time, the link congestion risk is reduced, and an idle link is reserved for subsequent data flows of the same communication algorithm; by allocating different links to data flow identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm in the same set communication domain, multi-path load sharing is performed, the dispersion transmission of data flows driven by the same communication algorithm is ensured, the single-link congestion risk is further reduced, and the communication efficiency of the intelligent computing center network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a path allocation method, data transmission method and network system for an intelligent computing center. Background Technology

[0002] An Artificial Intelligence Data Center (AIDC), often simply called an AIDC, is an infrastructure specifically designed to provide powerful computing power and data storage for Artificial Intelligence (AI) applications. To ensure efficient data transmission over the network, current AIDCs typically employ a network scheduling scheme based on a collective communication library. The core idea of ​​this scheme is that the collective communication library first reports the communication relationships between each computing node, then selects the data transmission path based on these relationships, ultimately achieving load balancing of business traffic.

[0003] However, this scheme treats each reported communication relationship as an independent entity, allocating paths based on the source IP (Internet Protocol) address to the destination IP address within each relationship to distribute traffic and achieve load balancing. While this design is relatively simple to implement technically, it can lead to severe link congestion in practical applications. Link congestion increases data transmission latency, reduces transmission success rate, and seriously impacts network performance. Summary of the Invention

[0004] In view of this, this application provides a path allocation method, a data transmission method, and a network system for an intelligent computing center to improve the network performance of the intelligent computing center.

[0005] Specifically, this application is implemented through the following technical solution:

[0006] According to a first aspect of the embodiments of this specification, a path allocation method for an intelligent computing center is provided, applied to a network controller, comprising: receiving communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center; each communication relationship information reported by any computing node includes at least: a data flow identifier and a communication algorithm corresponding to the data flow identifier; the data flow identifier is characterized at least by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs; allocating the same uplink to data flow identifiers with the same source IP address corresponding to different communication algorithms under the same set communication domain, and allocating the same downlink to data flow identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; allocating different links to data flow identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set communication domain.

[0007] According to a second aspect of the embodiments of this specification, a data transmission method for an intelligent computing center is provided, applied to a computing node, comprising: reporting communication relationship information to a network controller for a currently participating AI training task; each communication relationship information reported by the computing node includes at least: a data stream identifier and a communication algorithm corresponding to the data stream identifier; the data stream identifier is characterized at least by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data stream identifier; the set communication domain to which the data stream identifiers in each communication relationship information reported by the computing node belong is the same as the set communication domain to which the computing node belongs; the network controller allocates the same uplink for data stream identifiers with the same source IP address corresponding to different communication algorithms under the same set communication domain, and allocates the same downlink for data stream identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; allocates different links for data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set communication domain; and forwards the data streams corresponding to each data stream identifier according to the links allocated by the network controller for each data stream identifier.

[0008] According to a third aspect of the embodiments of this specification, a network system for an intelligent computing center is provided, including a network controller for executing the method of the first aspect; a computing cluster including multiple computing nodes, any computing node being used to execute the method of the second aspect; and a switch network including multiple switches, any switch being used to receive routing forwarding information corresponding to a data flow identifier issued by the network controller, and when a service packet matches the data flow identifier corresponding to the routing forwarding information, forwarding the service packet to the next-hop IP address in the routing forwarding information.

[0009] According to a fourth aspect of the embodiments of this specification, a network controller for an intelligent computing center is provided, comprising: an information receiving unit, configured to receive communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center; any communication relationship information reported by any computing node includes at least: a data flow identifier and a communication algorithm corresponding to the data flow identifier; the data flow identifier is characterized by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs; and a path allocation unit, configured to allocate the same uplink for data flow identifiers with the same source IP address corresponding to different communication algorithms under the same set communication domain, and allocate the same downlink for data flow identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; and allocate different links for data flow identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set communication domain.

[0010] According to a fifth aspect of the embodiments of this specification, a computing device for an intelligent computing center is provided, comprising: an uplink information transmission unit, configured to report communication relationship information to a network controller for a currently participating AI training task; any communication relationship information reported by the computing node includes at least: a data stream identifier and a communication algorithm corresponding to the data stream identifier; the data stream identifier is characterized by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data stream identifier; the set communication domain to which the data stream identifiers in each communication relationship information reported by the computing node belong is the same as the set communication domain to which the computing node belongs; the network controller allocates the same uplink link for data stream identifiers with the same source IP address corresponding to different communication algorithms under the same set communication domain, and allocates the same downlink link for data stream identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; and allocates different links for data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set communication domain; and an downlink information transmission unit, configured to forward the data streams corresponding to each data stream identifier according to the links allocated by the network controller for each data stream identifier.

[0011] According to a sixth aspect of the embodiments of this specification, an electronic device is provided, including a processor; and a computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method described in the first or second aspect.

[0012] In this embodiment, based on the temporal characteristics of AI training data streams and the characteristic that each collective communication library selects only one communication algorithm to drive data stream transmission in the same communication cycle, this embodiment pre-obtains data stream identifiers that can identify data streams. By assigning data stream identifiers with the same source IP address corresponding to different communication algorithms within the same collective communication domain to the same uplink and data stream identifiers with the same destination IP address corresponding to different communication algorithms to the same downlink, this ensures that data streams driven by different communication algorithms will not be transmitted simultaneously on the same link during actual communication, reducing the risk of link congestion and reserving idle links for subsequent data stream identifiers with the same communication algorithm. Furthermore, by assigning data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm within the same collective communication domain to different links, multi-path load balancing is performed to ensure that data streams driven by the same communication algorithm are transmitted in a distributed manner, further reducing the risk of single-link congestion, thereby improving the communication efficiency of the intelligent computing center network. Attached Figure Description

[0013] Figure 1 This is a schematic diagram illustrating a path allocation result in an exemplary embodiment of this application;

[0014] Figure 2 This is a schematic diagram of the system architecture of an intelligent computing center shown in an exemplary embodiment of this application;

[0015] Figure 3 This is a flowchart illustrating an exemplary embodiment of a path allocation method for an intelligent computing center.

[0016] Figure 4 This is a schematic diagram illustrating another path allocation result in an exemplary embodiment of this application;

[0017] Figure 5 This is a flowchart illustrating an exemplary embodiment of a data transmission method for an intelligent computing center.

[0018] Figure 6 This is a block diagram of a network system for an intelligent computing center, as illustrated in an exemplary embodiment of this application.

[0019] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment of this application;

[0020] Figure 8 This is a block diagram illustrating a network controller according to an exemplary embodiment of this application;

[0021] Figure 9 This is a block diagram illustrating a computing node in an exemplary embodiment of this application. Detailed Implementation

[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0025] As the scale of intelligent computing center clusters continues to expand, the amount of data transmission between nodes within the cluster is growing exponentially. Each data stream has unique characteristics, including content purpose, data size, bandwidth requirements, and latency requirements. These differentiated characteristics can easily lead to interference between data streams, causing localized link congestion, resulting in network performance degradation and even data loss. Although the network bandwidth of devices and server network interface card (NIC) bandwidth within the computing cluster continues to increase, relying solely on bandwidth expansion is no longer sufficient to meet the demands of explosive business growth. Link congestion is becoming increasingly prominent in intelligent computing center networks, especially in AI large-scale model training scenarios, where it is even more pronounced.

[0026] The training data for large AI models in intelligent computing centers is characterized by complex traffic structures and high bandwidth requirements for individual data streams. Traditional multi-path load balancing strategies (such as hash algorithms) are ineffective in this scenario, leading not only to significantly reduced network utilization but also to increased link congestion risks and packet loss. Particularly in scenarios where computing nodes transmit data based on RDMA (Remote Direct Memory Access), a packet loss rate of 0.1% can result in a 50% performance degradation. Therefore, optimizing network load balancing is of paramount importance.

[0027] Performance evaluation of network traffic for AI training tasks using multi-dimensional metrics (such as congestion count, latency, bandwidth utilization, and cache queue) revealed a significant periodicity in the data flow transmitted between computing nodes. This indicates that the network traffic for AI training tasks is highly predictable. Therefore, the intelligent computing center can pre-plan transmission links for each computing node's data flow from a global perspective, aiming to achieve load balancing and reduce the probability of link congestion through a global network traffic scheduling mechanism.

[0028] In traditional network traffic scheduling scenarios, the network controller performs intelligent path planning based on communication information reported by the aggregated communication library and the network topology information of the intelligent computing center. The communication information includes, for example, a five-tuple of data such as source IP address, destination IP address, source port, destination port, protocol type, source QP (Queue Pair), and destination QP. The network topology information includes all device nodes (e.g., switches, GPU servers) and link data of the intelligent computing center. This intelligent path planning strategy only focuses on the basic transmission requirements of a single data flow from the source computing node to the destination computing node, ignoring which aggregated communication algorithm the data flow is based on for transmission.

[0029] However, in AI training scenarios, different set communication algorithms have different traffic transmission logic and node interaction modes. For example, the traffic of the Ring communication algorithm needs to be transmitted sequentially node by node along the "ring topology" composed of computing nodes. The transmission link of a single data stream is fixed and the interaction between nodes has strong temporal correlation. On the other hand, the traffic of the Tree communication algorithm is based on the "tree topology". It broadcasts data from the root node to the leaf node (or aggregates data from the leaf node to the root node). The range of nodes covered by a single stream is wider and the transmission timing is completely different from that of the Ring communication algorithm.

[0030] Traditional intelligent path planning strategies fail to consider differences in communication algorithms, treating all communication processes driven by different algorithms as independent, unrelated traffic transmission processes for link allocation. This easily leads to different links being assigned to communication relationships driven by different communication algorithms. Consequently, newly added communication relationships are assigned to links already carrying communication relationships because all feasible link resources are already occupied. This results in multiple communication relationships based on the same algorithm and with the same source IP address needing to forward multiple data streams simultaneously through the same uplink link, and communication relationships based on the same algorithm and with the same destination IP address needing to forward multiple data streams simultaneously through the same downlink link, thus causing link congestion.

[0031] by Figure 1 Taking the aforementioned scenario as an example, the nodes in the computing cluster communicate via... Figure 1 The Spine-Leaf network shown enables node communication. Figure 1 In the Spine-Leaf network structure shown, one compute node is connected to one Leaf, and one Leaf is connected to multiple Spine nodes simultaneously. Assume that during the same aggregate communication process, the aggregate communication library deployed on compute node A reports first. Figure 1 Communication relationships 1 and 2 have different source IP addresses and the same destination IP address. Although communication relationships 1 and 2 correspond to different communication algorithms, the intelligent path planning strategy used by the network controller does not take the characteristics of the communication algorithm as input, resulting in these two communication relationships being assigned different Spine nodes. This causes the links Spine1→Leaf2 and Spine2→Leaf2 to be occupied in advance. When the aggregate communication library deployed on compute node C reports communication relationship 3, if the destination node corresponding to this communication relationship is compute node B, since all downlinks from Spine to Leaf2 have already been assigned communication relationships, and the weights of the two links are the same, the network controller will assign the downlink (Spine1→Leaf2) to communication relationship 3. This results in the downlink from Spine1 to Leaf2 being assigned two communication relationships based on the Ring communication algorithm. When the aggregate communication library transmits data streams through service messages, the data streams between compute node A and compute node C will be transmitted based on the downlink from Spine1 to Leaf2 within one communication cycle. Since AI large model training data has the characteristic of high bandwidth requirements for a single data stream, this can easily lead to link congestion.

[0032] Based on this, this technical solution uses communication algorithms as the core basis for path planning, optimizes network scheduling strategies, reduces link congestion, and thus improves the performance and resource utilization of aggregated communication. To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, the system architecture of the intelligent computing center in the embodiments of this application is described below:

[0033] like Figure 2 As shown, the intelligent computing center in this embodiment includes a network controller, multiple computing nodes, and multiple network devices. The multiple computing nodes can be devices such as terminals and servers, or devices such as network cards and GPUs (Graphics Processing Units). This embodiment does not specifically limit them.

[0034] In this embodiment, each computing node is configured with components such as a Collective Communications Library (CCL), a GPU network interface card (NIC), and an Agent. A Collective Communications Library (CCL) is a software architecture used to achieve efficient data communication in a distributed computing environment. Common CCLs include NCCL and XCCL. In practical applications, the appropriate CCL can be selected and configured based on the hardware environment and business requirements of the intelligent computing center. NCCL, short for NVIDIA Collective Communications Library, is a CCL developed by NVIDIA and supports various common collection operations such as All_Reduce, Broadcast, and All_Gather. XCCL, short for Xilinx Collective Communications Library, is a CCL developed by Xilinx. An Agent is an application running on the GPU NIC, responsible for communicating with the GPU NIC, receiving NIC notifications and uploading them to the network controller. It can also relay configurations issued by the network controller to the NIC through the Agent.

[0035] When an AI training task starts, the AI ​​training framework (e.g., a job scheduling system) triggers all GPU network cards participating in this AI training task to perform an All_Reduce global initialization operation, for example... Figure 2Each GPU network interface card (NIC) on computing nodes A1, A2, A3, and A4 participating in the first AI training task performs an All_Reduce global initialization operation. Similarly, each GPU NIC on computing nodes B1, B2, B3, and B4 participating in the second AI training task also performs an All_Reduce global initialization operation. Metadata is automatically reported to the Agent via the aggregated communication library running on each GPU NIC. The Agent then reports this metadata to the network controller, allowing the network controller to aggregate the metadata and establish a global GPU topology view to determine the range of GPUs participating in each AI training task. This metadata includes a global aggregated communication domain identifier, the host identifier where the GPU resides, and the GPU Bus ID, where the GPU Bus ID is the GPU bus identifier.

[0036] Because AI training tasks often employ various parallel strategies such as data parallelism (DP) and pipeline parallelism (PP), the requirements for computing node collaboration change dynamically during the training phase corresponding to different parallel strategies. For example, in the data parallel phase, all computing nodes participating in the same batch of data training need to achieve global gradient convergence and parameter synchronization. Therefore, the ensemble communication library configured on the computing nodes receives parallel strategy instructions from the AI ​​training framework in real time and dynamically creates ensemble communication domains according to the computing node grouping rules specified in the instructions. For example, computing nodes A1 and A2 are grouped into the same ensemble communication domain to participate in the same batch of data training for the first AI training task; computing nodes A3 and A4 are grouped into another ensemble communication domain to participate in another batch of data training for the first AI training task. In this way, by defining the communication scope and rules between computing nodes within the ensemble communication domain, the accuracy and efficiency of data interaction during parallel training are ensured.

[0037] During the collective communication domain construction phase, the collective communication library also reports intra-domain communication information to the network controller in real time. This intra-domain communication information includes, for example, the sub-communication domain identifier CommID and communication relationship information. The CommID is used to identify the collective communication domain to which the communication relationship information belongs, and the communication relationship information is used to indicate the data flow transmission requirements of the computing node, providing a basis for the network controller to perform global network resource scheduling and path planning.

[0038] The network controller delineates the boundaries of aggregated communication domains based on the global GPU topology view and CommID. For each aggregated communication domain, it performs intelligent path planning based on the network topology information of the intelligent computing center, the global GPU topology view, and the communication relationship information within each aggregated communication domain. The resulting path planning is then distributed to [the relevant network controller]. Figure 2The Spine-Leaf switch network in the example. When a compute node switches to data communication mode and begins transmitting data streams, the switch can forward the data streams sent by the compute node according to the path planning results.

[0039] The embodiments described in this specification will now be described in detail.

[0040] Figure 3 This is a schematic flowchart illustrating a path allocation method for an intelligent computing center, as shown in an exemplary embodiment of this application. The path allocation method is executed by the network controller of the intelligent computing center. Figure 3 As shown, method 300 includes at least the following steps S310 to S320:

[0041] Step S310: Receive communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center network.

[0042] In this context, any communication relationship information reported by any computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs.

[0043] A single AI training task refers to a computational task initiated to train a large AI model (such as an LLM or CV model). It typically includes parallel strategies such as data parallelism and pipeline parallelism. Each parallel strategy requires multiple computing nodes to execute collaboratively to ensure that communication resources are allocated on demand. LLM stands for Large Language Model, a deep learning-based artificial intelligence model trained on large-scale text data. It possesses the ability to understand and generate human language, answer questions, and perform text creation and logical reasoning. Typical examples include the GPT series and the LLaMA series. CV stands for Computer Vision Model, an image or video artificial intelligence model. Its core capabilities include image classification, object detection, image segmentation, face recognition, and pose estimation. Typical examples include ResNet, the YOLO series, and Mask R-CNN.

[0044] A ensemble communication domain refers to a logical grouping of all computing nodes participating in subtasks to achieve communication purposes such as gradient aggregation and parameter synchronization for parallel strategies. Each group is uniquely identified by a CommID. Computing nodes within this domain frequently exchange data, while computing nodes outside the domain do not participate, thus avoiding resource waste. Specifically, AI training tasks are divided into subtasks corresponding to various parallel strategies.

[0045] A data flow identifier is used to identify a specific data flow. Each data flow identifier corresponds to a set of transmission relationships from the source compute node to the destination compute node, determined by the source IP address and the destination IP address based on the communication algorithm. For each data flow identifier, if it corresponds to the Ring communication algorithm, the destination IP address is the IP address of the next or previous node of the source compute node in the "ring topology"; if it corresponds to the Tree communication algorithm, the destination IP address is the IP address of the root or leaf node of the source compute node in the "tree topology". The source compute node is the compute node corresponding to the source IP address in the data flow identifier. Understandably, in the communication relationship information reported by a computing node, the Ring communication algorithm should correspond to two data stream identifiers, and the destination IP addresses corresponding to these two data stream identifiers should correspond to the next node and the previous node of the computing node in the "ring topology", respectively. In contrast, the Tree communication algorithm should correspond to multiple data stream identifiers, and the destination IP addresses corresponding to these multiple data stream identifiers should correspond to the root node and each leaf node of the computing node in the "tree topology".

[0046] Communication algorithms refer to algorithms supported by a collection of communication libraries for data transmission, such as the Ring communication algorithm and the Tree communication algorithm, to determine the transmission logic and timing characteristics of the data stream.

[0047] In one example, when computing nodes participating in the same AI training task receive parallel strategy instructions from the AI ​​training framework, each computing node's collective communication library dynamically creates a collective communication domain and generates collective communication information. This collective communication information is identified by CommID and includes all communication algorithms supported by the computing node and the identifiers of each data stream corresponding to each communication algorithm. The aforementioned collective communication domain can be understood as a subdomain of the global collective communication domain for this AI training task.

[0048] Each computing node reports the generated aggregated communication information to the network controller. The network controller receives the aggregated communication information reported by each computing node and obtains all data stream identifiers within the aggregated communication domain. It then associates these data stream identifiers with computing nodes belonging to the same aggregated communication domain using CommID, facilitating path allocation for each aggregated communication domain by the network controller. Unlike related technologies, this embodiment synchronously reports the correspondence between data identifiers and communication algorithms to the network controller. This ensures that subsequent path planning accurately matches the communication characteristics of the AI ​​training task, avoiding link congestion risks caused by data streams with the same source or destination IP address and corresponding to the same communication algorithm being transmitted through the same link.

[0049] Step S320: Assign the same uplink to data stream identifiers with the same source IP address and the same downlink to data stream identifiers with the same destination IP address corresponding to different communication algorithms under the same set of communication domains; assign different links to data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set of communication domains.

[0050] During actual communication, the collective communication library selects one of the communication algorithms supported by the computing node to perform data stream transmission. If the user has specified a specific communication algorithm, the collective communication library will perform data stream transmission based on the specified communication algorithm; if the user has not specified a communication algorithm, the collective communication library will estimate the communication time of different communication algorithms according to the following formula and select the algorithm with the shortest communication time to perform data stream transmission.

[0051] Time = Latency + nBytes / algo_bw

[0052] In the above formula, latency represents the initial communication delay, the specific value of which depends on the hardware and the communication algorithm; algo_bw represents the effective bandwidth of the communication algorithm, the specific width of which depends on the hardware topology and the communication algorithm; and nBytes represents the total amount of data to be transmitted. Latency and algo_bw can be obtained during the initialization phase of the ensemble communication library, while nBytes is determined when the ensemble communication function is called. This ensemble communication function can be a gradient aggregation function, parameter synchronization function, etc.

[0053] In other words, for data stream identifiers corresponding to different communication algorithms within the same communication domain, regardless of whether the user specifies the communication algorithm, only one data stream of that communication algorithm is transmitted during actual communication. Data streams of different communication algorithms are completely mutually exclusive in the time dimension, preventing situations where data streams of different communication algorithms simultaneously occupy a link. Therefore, this step allocates the same uplink link to data stream identifiers with the same source IP address corresponding to different communication algorithms within the same communication domain, and allocates the same downlink link to data stream identifiers with the same destination IP address corresponding to different communication algorithms within the same communication domain. This avoids situations where some links are overloaded while others are idle, and it also reserves idle links for subsequent data streams of the same communication algorithm, fully utilizing the potential of network resources and improving resource utilization efficiency.

[0054] Meanwhile, data streams with the same source IP address or destination IP address corresponding to the same communication algorithm within the same aggregated communication domain will all initiate transmission synchronously within the same communication cycle after the aggregated communication library selects that algorithm. If these data streams are transmitted through the same link, it will cause the link bandwidth to be superimposed and trigger congestion. Therefore, this step assigns different uplink links to data streams with the same source IP address corresponding to the same communication algorithm within the same aggregated communication domain, and assigns different downlink links to data streams with the same destination IP address corresponding to the same communication algorithm within the same aggregated communication domain. This can ensure that data streams with the same algorithm are transmitted in a distributed manner through multi-path load balancing, avoid the risk of single-link congestion, and meet the requirements of low packet loss and high bandwidth in RDMA transmission scenarios.

[0055] It is worth noting that the links described in this step include uplinks and downlinks. An uplink refers to the path segment from the source computing node to the core device in the network, and a downlink refers to the path segment from the network core device to the target computing node. Figure 2 Taking the Spine-Leaf network as an example, the uplink refers to the link from the source computing device → Leaf → Spine, and the downlink refers to the link from Spine → Leaf → destination computing device.

[0056] like Figure 3As can be seen from the path allocation method shown, this solution is based on the timing characteristics of the AI training data stream and the characteristic that each collective communication library selects only one communication algorithm to drive the data stream transmission in the same communication cycle. In the embodiments of the present application, a data stream identifier capable of identifying the data stream is pre-obtained. By allocating the data stream identifiers with the same source IP address corresponding to different communication algorithms within the same collective communication domain to the same uplink, and allocating the data stream identifiers with the same destination IP address corresponding to different communication algorithms to the same downlink, when actual communication occurs, the data streams driven by different communication algorithms on the same link will not be transmitted simultaneously, reducing the risk of link congestion, and reserving idle links for the subsequent data stream identifiers of the same communication algorithm; and, by allocating the data stream identifiers with the same source IP address or the same destination IP address corresponding to the same communication algorithm within the same collective communication domain to different links, multi-path load sharing is performed to ensure that the data streams driven by the same communication algorithm are dispersed for transmission, further reducing the risk of single-link congestion, so as to improve the communication efficiency of the intelligent computing center network.

[0057] In some embodiments, after receiving the communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center network, method 300 further includes: classifying the data stream identifiers within the same collective communication domain according to their corresponding communication algorithm types, obtaining a data stream identifier table corresponding to each communication algorithm, and obtaining structured data about the data stream identifiers according to the data stream identifier table, CommID, and the communication algorithm. This structured data is, for example, a three-layer mapping structure with CommID as the top-level key and the algorithm-data stream identifier table as the secondary mapping, that is, <CommID, Map<algorithm, List<data stream identifier>>>.

[0058] In this embodiment, by performing domain boundary division on the original communication relationship information, data aggregation based on the communication algorithm dimension, and structured processing of the data stream identifiers, it is ensured that the scheduling of data streams in different collective communication domains is independent of each other, providing data support for subsequent differential path allocation based on the communication algorithm.

[0059] Specifically, in the embodiments of the present application, referring to Figure 2 , the detailed process of path allocation in the intelligent computing center is as follows:

[0060] 1. Network information collection: The network controller actively collects the network topology information of the intelligent computing center and stores it as structured information, providing basic support for subsequent path planning. The network topology information includes information such as the whole network topology, links, and the bandwidth occupancy of each link.

[0061] 2. AI training task trigger: When the AI training framework starts an AI training task, it triggers network communication requirements.

[0062] 3. Reporting of communication relationship information: The collective communication library reports the communication relationship information of the AI ​​training task to the network controller. The network controller performs in-depth analysis of the reported communication relationship information and performs structured parsing of the communication relationship information based on the collective communication domain, providing data support for subsequent path optimization.

[0063] In some embodiments, compute nodes can report information to the network controller using a combination of hierarchical reporting and dynamic awareness mechanisms. For example, during the global initialization phase at the start of an AI training task, all GPU network cards participating in the same AI training task are triggered to perform an All_Reduce initialization operation. At this time, the collective communication libraries running on all GPUs automatically report metadata to the Agent, which then reports this metadata to the network controller. This allows the network controller to aggregate this metadata and establish a global GPU topology view. During the dynamic construction phase of the collective communication domain, each GPU reports communication relationship information within the domain in real time (e.g., CommID, the communication algorithm type supported by the GPU network card (Ring / Tree), source / destination IP addresses, source / destination QPs, etc.).

[0064] 4. Path Planning and Traffic Balancing: The network controller, based on data flow identifiers within the same communication domain and combined with network topology information, assigns the same path to source data flows corresponding to different communication algorithms within the same communication domain, and assigns different paths to data flows with the same source or destination IP address corresponding to the same communication algorithm within the same communication domain. This ensures balanced load distribution of data flows across links, avoids congestion, and optimizes resource utilization. It can dynamically adjust path planning results based on real-time network conditions and task requirements, flexibly responding to network topology changes and traffic fluctuations, ensuring the adaptability and robustness of network scheduling.

[0065] Continuing with the Spine-Leaf network as an example, such as Figure 4 As shown, based on the path planning of this application embodiment, data flow identifier 1 and data flow identifier 2 are assigned to the same link, while data flow identifier 1 and data flow identifier 3 are assigned to different links. Since data flow identifier 1 and data flow identifier 2 are driven by different communication algorithms, even if these two data flow identifiers are assigned to the same link, their corresponding data flows will not be transmitted simultaneously in the time dimension, thus preventing network congestion.

[0066] In some embodiments, after obtaining the path planning result in step S320, the network controller may perform end-side scheduling or network scheduling, wherein:

[0067] End-side scheduling refers to the network controller dynamically calculating the source port number corresponding to each data flow identifier based on path planning results and the current routing algorithm of the switch (such as a hash algorithm). The calculated source port number is then sent to the aggregated communication library of the corresponding computing node. The aggregated communication library dynamically modifies the source port number of the service packet on the end side (i.e., the source computing node). By adjusting the source port number, the hash result of the service packet carrying the data flow in the switch is changed, thereby guiding the data flow to be forwarded strictly according to the link allocated by the network controller. This achieves precise traffic scheduling without modifying the switch configuration or network architecture, simplifying the implementation process and reducing deployment and maintenance costs.

[0068] Network-side scheduling refers to the network controller generating routing and forwarding information for each switch on a given link based on the data flow identifier, and then sending this information to the corresponding switches. This routing and forwarding information is uniquely identified by the data flow identifier. The switches configure table entries based on the received routing and forwarding information, and then look up and forward service packets that match the data flow identifier to control the data flow carried by the service packets along the links allocated by the network controller. By configuring the table entries of the switches, the stability and accuracy of data flow scheduling are determined, redundant hops and link congestion are reduced, and the efficiency of aggregated communication is improved, thereby accelerating the completion of AI training tasks.

[0069] The following describes in detail the terminal-side scheduling and network-side scheduling with reference to specific examples.

[0070] For end-side scheduling, in some embodiments, method 300 further includes: for each data flow identifier, determining the source port number corresponding to the data flow identifier based on the link to which the data flow identifier is allocated and the routing parameters corresponding to the data flow identifier; the routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier; and distributing the source port number corresponding to the data flow identifier to the computing node whose IP address is the source IP address corresponding to the data flow identifier, so that the computing node carries the source port number when sending service packets to the destination IP address corresponding to the data flow identifier in the future.

[0071] During service packet forwarding, the switch determines the outgoing interface of the service packet on the current switch based on a pre-defined routing algorithm. Taking a hash algorithm as an example, the current switch performs a hash calculation based on the incoming interface and the five-tuple data of the service packet, and sends the service packet through the outgoing interface corresponding to the hash result. In RDMA communication scenarios, the source IP and destination IP of the service packet need to remain constant after the communication connection is established, and the destination port number of the service packet based on the RDMA protocol is fixed at 4791. Therefore, the end-side scheduler can change the hash result of the packet in the switch by adjusting the source port number in the five-tuple data, thereby controlling the selection of the outgoing interface of the switch and guiding the service packet carrying the data stream to be transmitted along the link allocated by the network controller.

[0072] It is worth noting that in some examples, the data flow identifier can also be represented by the source IP address, source port number, destination IP address, destination port number, source QP, and destination QP of the compute node. That is, the compute node uses the five-tuple data as the data flow identifier and reports it to the network controller. When the network controller performs end-side scheduling, it recalculates the source port number corresponding to each data flow identifier and updates the source port number in the data flow identifier with the new source port number. Then, it sends the updated data flow identifier to the compute node that reported the data flow identifier. This allows the compute node to generate a service packet carrying the new source port number based on the updated data flow identifier. This changes the hash result of the packet in the switch, thereby controlling the selection of the switch's outgoing interface and guiding the service packet to be transmitted along the link allocated by the network controller.

[0073] In some embodiments, determining the source port number corresponding to the data flow identifier includes: for each data flow identifier, based on the link to which the data flow identifier is allocated, sending the source IP address and destination IP address corresponding to the data flow identifier, as well as the target ingress interface of the device, to any device on the link, so that the device calculates the source port number corresponding to each egress interface of the device based on the source IP address and destination IP address corresponding to the data flow identifier, and the target ingress interface of the device, and reports the source port number corresponding to each egress interface of the device to the network controller; the network controller determines the source port number corresponding to the egress interface that matches the link allocated to the data flow identifier as the source port number corresponding to the data flow identifier. The target ingress interface of the device is the ingress interface through which the device receives the service packet corresponding to the data flow identifier, and the network controller can obtain the target ingress interface based on the network topology information of the intelligent computing center.

[0074] In one example, the network controller also sends the range of source port numbers supported by the compute node corresponding to the IP address of the data stream identifier to the source device on the link. This allows the source device to calculate only the outgoing interfaces corresponding to each source port number within the specified range, instead of traversing and calculating all outgoing interfaces of the source device, thus saving computational resources. Here, the source device refers to the Leaf device that establishes a connection with the compute node.

[0075] In practical applications, different manufacturers or models of devices (e.g., switches) may use different hash algorithms when selecting outgoing interfaces. Directly pre-setting the source port number by the network controller may result in data flows failing to forward along the assigned path due to algorithm mismatch. Therefore, in this embodiment, the device calculates the correspondence between source port numbers and outgoing interface numbers based on its own configured hash algorithm to obtain the source port number corresponding to each data flow identifier. Of course, in other embodiments, the network controller can also calculate the source port number corresponding to each data flow identifier.

[0076] For network-side scheduling, in some embodiments, method 300 further includes: for each data flow identifier, based on the link to which the data flow identifier is allocated, controlling each device on the link to record the routing and forwarding information corresponding to the data flow identifier; the routing and forwarding information recorded by any device is used to indicate forwarding packets to the destination IP address corresponding to the data flow identifier, and the routing and forwarding information recorded by any device includes the device's outgoing interface and / or next-hop IP address, wherein the device's outgoing interface and next-hop IP address can be obtained through the network topology information of the intelligent computing center.

[0077] When AI training scenarios have high requirements for communication reliability, routing information can be configured to include the device's outgoing interface and next-hop IP address. If the routing information only records the outgoing interface, and the next-hop device connected to that outgoing interface (e.g., a Spine switch) fails, the switch will still send service packets from the specified outgoing interface, resulting in packet loss. However, when the routing information also records the next-hop IP address, the device can verify the next-hop status in real time through IP reachability detection. If the next hop is unreachable, path recalculation can be triggered and the outgoing interface updated, avoiding continuous packet transmission to the faulty node.

[0078] In one example, controlling each device on the link to record the routing and forwarding information corresponding to the data flow identifier includes: sending the routing and forwarding information corresponding to the data flow identifier to each device on the link.

[0079] In some application scenarios, the network controller distributes path planning results to the corresponding devices (e.g., switches) via Traffic Matrix flow tables. For each data flow identifier, a corresponding Traffic Matrix flow table is generated. This Traffic Matrix flow table includes the outgoing interfaces and next-hop IP addresses of each device on the path assigned to that data flow identifier. Because Traffic Matrix flow table forwarding has higher priority than routing forwarding, switches prioritize forwarding service packets according to the paths corresponding to the Traffic Matrix flow tables. The Traffic Matrix flow table refers to a technology provided by switches based on ACL (Access Control Lists) resources specifically for intelligent computing centers, offering superior traffic forwarding guidance compared to routing forwarding.

[0080] This application also provides a data transmission method for an intelligent computing center. For example... Figure 5 As shown, Figure 5 This is a schematic flowchart illustrating a data transmission method in an intelligent computing center according to an exemplary embodiment of this application. This data transmission method is executed by computing nodes in the intelligent computing system. Figure 5 As shown, method 500 includes at least the following steps S510 to S520:

[0081] Step S510: For the currently participating AI training task, report the communication relationship information to the controller;

[0082] The communication relationship information reported by the computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by the computing node belongs is the same as the set communication domain to which the computing node belongs; the network controller allocates the same uplink to data flow identifiers with the same source IP address corresponding to different communication algorithms under the same set communication domain, and allocates the same downlink to data flow identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; and allocates different links to data flow identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set communication domain.

[0083] Step S520: Forward the data stream corresponding to each data stream identifier according to the link allocated by the network controller for each data stream identifier.

[0084] In some embodiments, method 500 further includes: receiving a source port number corresponding to each data flow identifier reported by the controller, wherein the source port number is determined by the controller based on the link allocated to the data flow identifier and the routing parameters corresponding to the data flow identifier, and the routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier; correspondingly, step S520 above forwards the data flow corresponding to each data flow identifier according to the link allocated by the network controller for each data flow identifier, including: sending a service packet carrying the source port number corresponding to the data flow identifier to the IP address corresponding to each data flow identifier of the computing node, so that the service packet is transmitted along the path corresponding to the data flow identifier.

[0085] In some embodiments, step S520 above forwards the data stream corresponding to each data stream identifier according to the link allocated by the network controller for each data stream identifier, including: sending a service packet to the IP address corresponding to the target data stream identifier of the computing node based on the target data stream identifier corresponding to the communication algorithm currently selected by the computing node.

[0086] This application also provides a network system for an intelligent computing center. For example... Figure 6 As shown, Figure 6 This is a block diagram of a network system for an intelligent computing center, as illustrated in an exemplary embodiment of this application. Figure 6 As shown, the network system includes:

[0087] Network controller 610 is used to execute the steps of method 300;

[0088] The computing cluster includes multiple computing nodes 621, 622, and 623, each of which is used to execute the steps of method 500.

[0089] The switch network includes multiple switches 631, 632, 633, 634, and 635. Any switch is used to receive routing forwarding information corresponding to the data flow identifier issued by the network controller, and when a service packet matches the data flow identifier corresponding to the routing forwarding information, forwards the service packet to the next-hop IP address in the routing forwarding information.

[0090] in, Figure 6 Only three compute nodes and five switches are shown. In practical applications, those skilled in the art can set the number of compute nodes and switches as needed and flexibly construct the topology of the switch network.

[0091] The network system in this embodiment supports the complex network communication requirements of large-scale AI training tasks, can adapt to intelligent computing center environments of different sizes and network topologies, and provides a highly scalable network scheduling solution for distributed computing.

[0092] Figure 7 This is a schematic diagram of an electronic device illustrated in this specification according to an exemplary embodiment. Please refer to... Figure 7 At the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, memory 708, a hardware acceleration device 710, and non-volatile memory 712, and may also include other hardware required for its functions. One or more embodiments of this application can be implemented in software, for example, the processor 702 reads the corresponding computer program from the non-volatile memory 712 into the memory 708 and then runs it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0093] Figure 8 This is a block diagram illustrating an exemplary embodiment of the present application of a network controller, which can be applied to, for example... Figure 7 The electronic device shown implements the technical solution of this application. The network controller may include a first information receiving unit 810 and a path allocation unit 820, wherein:

[0094] The first information receiving unit 810 is used to receive communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center; wherein, any communication relationship information reported by any computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs;

[0095] The path allocation unit 820 is used to allocate the same uplink to data stream identifiers with the same source IP address and the same downlink to data stream identifiers with the same destination IP address corresponding to different communication algorithms under the same set of communication domains; and to allocate different links to data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set of communication domains.

[0096] In some embodiments, the network controller further includes an information sending unit 830, and the path allocation unit 820 includes a port number acquisition module;

[0097] The port number acquisition module is used to determine the source port number corresponding to each data flow identifier based on the link allocated to the data flow identifier and the routing parameters corresponding to the data flow identifier; the routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier.

[0098] The information sending unit 830 sends the source port number corresponding to the data stream identifier to the computing node whose IP address is the source IP address corresponding to the data stream identifier, so that the computing node carries the source port number when sending service packets to the destination IP address corresponding to the data stream identifier in the future.

[0099] In some embodiments, the path allocation unit 820 further includes an information calculation module;

[0100] The information calculation module is used to control each device on the path to record the routing and forwarding information corresponding to each data flow identifier based on the link to which the data flow identifier is allocated; the routing and forwarding information recorded by any device is used to indicate the forwarding of packets to the destination IP address corresponding to the data flow identifier, and the routing and forwarding information recorded by any device includes the device's outgoing interface and / or next-hop IP address.

[0101] Figure 9 This is a block diagram illustrating a computing device according to an exemplary embodiment of this application, the computing device being applicable to, for example... Figure 7 The illustrated electronic device implements the technical solution of this application. The computing device includes an uplink information transmission unit 910 and an downlink information transmission unit 920, wherein:

[0102] The uplink transmission unit 910 is used to report communication relationship information to the controller for the currently participating AI training task. Each communication relationship information reported by the computing node includes at least: a data stream identifier and the corresponding communication algorithm. The data stream identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data stream identifier. The set communication domain to which the data stream identifiers in each communication relationship information reported by the computing node belong is the same as the set communication domain to which the computing node belongs. The network controller allocates the same uplink link to data stream identifiers with the same source IP address corresponding to different communication algorithms within the same set communication domain, and allocates the same downlink link to data stream identifiers with the same destination IP address corresponding to different communication algorithms within the same set communication domain. It also allocates different links to data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm within the same set communication domain.

[0103] The downlink information transmission unit 920 is used to forward the data stream corresponding to each data stream identifier according to the link allocated by the network controller for each data stream identifier.

[0104] In some embodiments, the computing device further includes a second information receiving unit 930;

[0105] The second information receiving unit 930 is used to receive the source port number corresponding to each data flow identifier reported by the network controller. The source port number is determined by the controller based on the link allocated to the data flow identifier and the routing parameters corresponding to the data flow identifier. The routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier. Correspondingly, the information downlink sending unit 920 is used to send service packets carrying the source port number corresponding to the data flow identifier to the IP address corresponding to each data flow identifier of the computing node, so that the service packets are transmitted along the link corresponding to the data flow identifier.

[0106] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0107] Accordingly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above embodiments.

[0108] Accordingly, embodiments of this application also provide a computer program product configured to perform the methods described in any of the above embodiments.

[0109] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0110] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0111] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0112] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0113] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0114] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0115] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0116] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0117] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A path allocation method for an intelligent computing center, characterized in that, Applied to a network controller, the method includes: The system receives communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center; any communication relationship information reported by any computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is characterized by at least the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs; Data streams from different communication algorithms within the same communication domain are mutually exclusive in the time dimension. Based on this, data streams with the same source IP address corresponding to different communication algorithms within the same communication domain are assigned the same uplink, and data streams with the same destination IP address are assigned the same downlink. Data streams with the same source IP address or destination IP address corresponding to the same communication algorithm within the same communication domain are assigned different links. The uplink refers to the path segment from the source computing node to the core device in the network, and the downlink refers to the path segment from the network core device to the target computing node.

2. The method according to claim 1, characterized in that, The method further includes: For each data flow identifier, the source port number corresponding to the data flow identifier is determined based on the link to which the data flow identifier is allocated and the routing parameters corresponding to the data flow identifier; the routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier. The source port number corresponding to the data flow identifier is sent to the computing node whose IP address is the source IP address corresponding to the data flow identifier, so that the computing node carries the source port number when sending service packets to the destination IP address corresponding to the data flow identifier in the future.

3. The method according to claim 1, characterized in that, The method further includes: For each data flow identifier, based on the link to which the data flow identifier is assigned, control each device on that link to record the routing and forwarding information corresponding to the data flow identifier; the routing and forwarding information recorded by any device is used to indicate the forwarding of packets to the destination IP address corresponding to the data flow identifier, and the routing and forwarding information recorded by any device includes the device's outgoing interface and / or next-hop IP address.

4. The method according to claim 3, characterized in that, The control of each device on the link to record the routing and forwarding information corresponding to the data stream identifier includes: For each device on the link, send the routing and forwarding information corresponding to the data flow identifier to each device.

5. A data transmission method for an intelligent computing center, characterized in that, Applied to compute nodes, the method includes: For the current AI training task, the computing node reports communication relationship information to the network controller. Each communication relationship information reported by the computing node includes at least: a data flow identifier and the corresponding communication algorithm. The data flow identifier is represented by at least the source IP address of the computing node and the destination IP address determined based on the corresponding communication algorithm. The set communication domain to which the data flow identifiers in each communication relationship information reported by the computing node belong is the same as the set communication domain to which the computing node belongs. Data flows of different communication algorithms within the same set communication domain are mutually exclusive in the time dimension. Based on this, the network controller allocates the same uplink to data flow identifiers with the same source IP address and the same downlink to data flow identifiers with the same destination IP address corresponding to different communication algorithms within the same set communication domain; and allocates different links to data flow identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm within the same set communication domain. The uplink refers to the path segment from the source computing node to the core device in the network, and the downlink refers to the path segment from the core device in the network to the target computing node. The network controller forwards the data streams corresponding to each data stream identifier according to the links assigned to each data stream identifier.

6. The method according to claim 5, characterized in that, The method further includes: The network controller receives the source port number corresponding to each data flow identifier reported by the network controller. The source port number is determined by the network controller based on the link allocated to the data flow identifier and the routing parameters corresponding to the data flow identifier. The routing parameters include at least the source IP address and / or destination IP address corresponding to the data flow identifier. The forwarding of data flows corresponding to each data flow identifier according to the link allocated by the network controller for each data flow identifier includes: Send a service packet carrying the source port number corresponding to the data flow identifier to the IP address corresponding to each data flow identifier of the computing node, so that the service packet is transmitted along the path corresponding to the data flow identifier.

7. A network system for an intelligent computing center, characterized in that, include: A network controller for performing the method according to any one of claims 1 to 4; A computing cluster comprising multiple computing nodes, wherein any computing node is used to execute the method of claim 5 or 6; A switch network includes multiple switches, each of which receives routing information corresponding to a data flow identifier issued by the network controller, and forwards the service packet to the next-hop IP address in the routing information when the service packet matches the data flow identifier corresponding to the routing information.

8. A network controller for an intelligent computing center, characterized in that, include: The information receiving unit is used to receive communication relationship information reported by each computing node participating in the same AI training task in the intelligent computing center. Any communication relationship information reported by any computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by any computing node belongs is the same as the set communication domain to which the computing node belongs; The path allocation unit, based on the mutual exclusion of data streams with different communication algorithms under the same set of communication domains in the time dimension, is used to allocate the same uplink to data stream identifiers with the same source IP address and the same downlink to data stream identifiers with the same destination IP address corresponding to different communication algorithms under the same set of communication domains; and to allocate different links to data stream identifiers with the same source IP address or destination IP address corresponding to the same communication algorithm under the same set of communication domains. The uplink refers to the path segment from the source computing node to the core device in the network, and the downlink refers to the path segment from the network core device to the target computing node.

9. A computing device for an intelligent computing center, characterized in that, include: The information uplink sending unit is used to report communication relationship information to the network controller for the currently participating AI training task; Each communication relationship information reported by the computing node includes at least: a data flow identifier and the communication algorithm corresponding to the data flow identifier; the data flow identifier is represented by the source IP address of the computing node and the destination IP address determined based on the communication algorithm corresponding to the data flow identifier; the set communication domain to which the data flow identifier in each communication relationship information reported by the computing node belongs is the same as the set communication domain to which the computing node belongs; data flows of different communication algorithms under the same set communication domain are mutually exclusive in the time dimension. Based on this, the network controller allocates the same uplink to data flow identifiers with the same source IP address and the same downlink to data flow identifiers with the same destination IP address corresponding to different communication algorithms under the same set communication domain; and allocates different links to data flow identifiers with the same source IP address or the same destination IP address corresponding to the same communication algorithm under the same set communication domain; wherein, the uplink refers to the path segment from the source computing node to the core device in the network, and the downlink refers to the path segment from the core device in the network to the target computing node; The downlink information transmission unit is used to forward the data stream corresponding to each data stream identifier according to the link allocated by the network controller for each data stream identifier.

10. An electronic device, characterized in that, include: processor; as well as A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information generation method and device, information adjustment method and device, equipment, medium and automatic driving vehicle

    CN114906172A

  • Link distribution method, device, equipment and medium

    CN118233356A