A method for managing traffic and a cloud cluster device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NEW H3C TECH CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-29
Smart Images

Figure CN122120214A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of communication technology, and in particular to a traffic control method and a cloud cluster device. Background Technology
[0002] With the rapid development of large-scale AI models, the number of parameters has reached trillions, requiring GPU interconnects of tens of thousands of cards for model pre-training. This results in massive network scale and complex connections in intelligent computing centers. The massive manual deployment of configurations is time-consuming and error-prone, and network adjustments become increasingly difficult as business needs change. Data shows that in the traditional Spine-Leaf architecture of intelligent computing centers, traffic congestion between Spine switches leads to a 30% drop in computing power. To adapt to the development of AIGC and meet these challenges, data center internal and interconnect technologies need continuous innovation and upgrading to provide more powerful, reliable, and efficient computing, storage, and communication capabilities. To adapt to the business flow model of AI intelligent computing centers, the industry has introduced cloud cluster switches. These switches replace main control boards, network boards, and interface boards with individual box-type switches, adopting a distributed architecture. They replace the original backplane bus connection with a management network, integrating a single device through a cloud cluster management platform, achieving excellent scalability. Rapid network expansion is achieved by adding device management modules, demonstrating strong scalability. Strong fault isolation capability: DDC has excellent fault isolation capabilities. If a module fails, the device management module isolates it promptly, preventing impact on the operation of the entire network. High flexibility: The cloud cluster switch has a flexible structure and can be configured and deployed flexibly according to the size and needs of the intelligent computing center.
[0003] The three components of a cloud cluster switch are NCC, NCF, and NCP, which correspond to the main control board, switching network board, and interface board of a traditional chassis switch, respectively.
[0004] It is evident that cloud cluster switches have addressed a series of challenges in network operation and maintenance. However, cloud cluster solutions still have some shortcomings in handling AI service communication traffic models. AIGC service traffic may be affected by user activity, event-driven processes, and content generation requests, leading to more frequent and irregular peak traffic. This places higher demands on the traffic model, requiring it to be able to quickly adapt to and handle sudden traffic surges. While PKTC-based load balancing and DGSQ global scheduling technologies can effectively regulate and distribute traffic under stable conditions, network congestion can still occur in short periods under abnormal scenarios such as micro-bursts and link failures. In such cases, backpressure mechanisms are still needed to suppress traffic transmission from the source. Traditional PFC or FC are point-to-point local backpressure technologies; once triggered, they can spread throughout the entire network, causing network storms and other problems. Summary of the Invention
[0005] To overcome the problems existing in related technologies, this specification provides a traffic control method and a cloud cluster device.
[0006] According to a first aspect of the embodiments of this specification, a method for controlling traffic is provided, the method comprising: The cloud cluster controller (NCC) in the cloud cluster switch establishes communication with the collective communication library (UCCL) of each server in the server cluster; the NCC monitors the bandwidth utilization of each server. When the NCC monitors that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs on other servers other than the target server that the target server is busy, so that the UCCLs on other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
[0007] Among them, the NCC receives the communication topology sent by the UCCL in each server and integrates it into the overall communication topology; The overall communication topology is sent to the UCCL in each server.
[0008] The overall communication topology includes some or all of the following information: GPU number, RANK number, RDMA network card number, IP address, Server number, NCP number, and port number.
[0009] The methods used by the NCC to monitor the bandwidth utilization of the target server include: NCC monitors the interface occupancy rate between each NCP and its corresponding server, and identifies the server whose interface occupancy rate reaches the first threshold as the target server.
[0010] The notification that the UCCL target server is busy among servers other than the target server includes: A first notification message is sent to the UCCLs of servers other than the target server. The first notification message carries the target address information of the target server so that the other servers can identify that the target server corresponding to the target address information is busy after receiving the first notification message.
[0011] The step of adjusting the UCCL in servers other than the target server to reduce the rate at which traffic is sent to the target server includes: This allows UCCLs in other servers to filter out target traffic destined for the target server from the overall communication topology; Adjust the rate at which the target traffic is reduced.
[0012] As can be seen from the above embodiments, by establishing communication between the NCC and UCCL, when the NCC senses that the target server is busy, it can notify the UCCLs in each server to adjust and reduce the traffic sent to the target server, so that the target server can return from a busy state to a non-busy state. In other words, the NCC can actively notify the UCCL of the network port's busy level, allowing the UCCL to dynamically sense the effective bandwidth of the RDMA network card, thereby adjusting the communication time in a timely manner and avoiding the possibility of network congestion from the source.
[0013] According to a second aspect of the embodiments of this specification, a cloud cluster device is provided, comprising a cloud cluster controller NCC, a network control protocol NCP, and a collective communication library UCCL deployed in each server, wherein the NCC is communicatively connected to the UCCL and NCP in each server, and each NCP is communicatively connected to each server. The NCC monitors the bandwidth utilization of each server through a monitoring module; When the monitoring module detects that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs of other servers other than the target server that the target server is busy, so that the UCCLs of other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
[0014] The NCC also includes: a receiving module. The receiving module is used to receive the communication topology sent by UCCL in each server and integrate it into a total communication topology. The overall communication topology is sent to the UCCL in each server via the sending module.
[0015] The overall communication topology includes some or all of the following information: GPU number, RANK number, RDMA network card number, IP address, Server number, NCP number, and port number.
[0016] Specifically, the monitoring module is used to monitor the interface occupancy rate between each NCP and its corresponding server, and to designate the server whose interface occupancy rate reaches a first threshold as the target server.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.
[0019] Figure 1This specification is a schematic diagram of the architecture of a cloud cluster device according to an exemplary embodiment.
[0020] Figure 2 This is a flowchart illustrating a traffic control method according to an exemplary embodiment of this specification.
[0021] Figure 3 This specification is a schematic diagram of the architecture of a cloud cluster device according to an exemplary embodiment. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0023] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0025] Currently, the three components of a cloud cluster switch—NCC, NCF, and NCP—correspond to the main control board, switching network board, and interface board of a traditional chassis switch, respectively.
[0026] Among them, NCC (Network Cloud Controller) is responsible for managing and controlling the operation of the entire cloud cluster switch equipment and network. NCC is responsible for the allocation and monitoring of network resources to ensure the stability and security of the entire cloud cluster switch operation.
[0027] NCF: Network Cloud Packet Forwarder, responsible for forwarding packets between NCP devices.
[0028] NCP: Network Cloud Packet-Forwarder, responsible for connecting to various servers and GPUs.
[0029] The current cloud cluster management topology and physical connections are as follows: Figure 1 As shown, NCC, NCF, and NCP are independent physical devices that integrate Comware's rich network features. They are physically connected together through the MGT management network to form a cloud cluster switch, presenting the entire data center as a single device to the outside world.
[0030] However, while cloud cluster switches address a series of challenges in network operation and maintenance, they still have some shortcomings in solving AI service communication traffic models. AIGC service traffic can be affected by user activity, event-driven processes, and content generation requests, leading to more frequent and irregular peak traffic. This places higher demands on the traffic model, requiring it to quickly adapt to and handle sudden traffic surges. While PKTC-based load balancing and DGSQ global scheduling technologies can effectively regulate and distribute traffic under stable conditions, network congestion can still occur in short periods under abnormal scenarios such as micro-bursts and link failures. In these cases, backpressure mechanisms are still needed to suppress traffic transmission from the source. Traditional PFC or FC are point-to-point local backpressure technologies; once triggered, they can spread throughout the entire network, causing network storms and other problems.
[0031] To address the aforementioned technical challenges, the inventors combined UCCL's aggregated communication mode with the dynamic scheduling capabilities of the network controller to construct a high-performance, predictable data center-level AI training traffic management solution.
[0032] like Figure 2 As shown in the embodiments of this disclosure, a method for controlling traffic is provided, the method comprising: The cloud cluster controller (NCC) in the S201 cloud cluster switch establishes communication with the collective communication library (UCCL) of each server in the server cluster; the NCC monitors the bandwidth utilization of each server. S202 When the NCC monitors that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs of other servers other than the target server that the target server is busy, so that the UCCLs of other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
[0033] To provide a full explanation of this disclosure, the applicant describes an AI computing server consisting of an upper-layer distributed AI application, a Collective Communication Library (UCCL), and an RDMA protocol. The UCCL plays a key role in distributed AI training, as it can efficiently coordinate the collaborative work of multiple GPUs (whether on a single server or across multiple servers).
[0034] In order to enable the NCC to proactively notify the UCCL network port of its busy status when it senses that the target server (such as the AI server) is busy, and to allow the UCCL to dynamically sense the effective bandwidth of the RDMA network card, thereby adjusting the communication atom time in a timely manner and avoiding the effect of network congestion from the source, in step 201, the NCC establishes communication with the UCCL in each server through the management network.
[0035] like Figure 3 As shown, NCC1 and NCC2 are in a master-slave architecture. The NCC connects to the UCCL of each server in the server cluster through the MGT management network.
[0036] Among them, UCCL (Unified Collective Communication Library) is a communication library specifically developed for multi-GPU and high-performance computing clusters. Its goal is to provide collective communication capabilities with higher performance, stronger stability, better scalability, and better resource utilization than the standard NCCL, when combined with cloud cluster switch network infrastructure. It inherits the mature design of NCCL in the field of GPU communication, while also being customized and enhanced for the characteristics of cloud cluster switch networks to accelerate distributed deep learning training and massively parallel computing tasks. UCCL connects CUDA programs to RDMA communication networks. CUDA is an abbreviation for Compute Unified Device Architecture. It is a parallel computing platform and programming model launched by NVIDIA. CUDA programs are software written using the CUDA platform and model that allows CPUs and GPUs to work together, such as AI application-layer programs similar to AI deep learning frameworks PyTorch / TensorFlow, or high-performance computing applications.
[0037] The core function of UCCL is to provide highly optimized set communication primitives (also called operators) that use standardized algorithms to solve the problem of cross-process data synchronization in parallel computing, which is crucial for distributed deep learning training. By running standardized operators, the communication capabilities provided by the network for distributed parallel computing can be measured. Common operators include set operations such as AllReduce, AlltoAll, Broadcast, AllGather, Reduce, Reduce-Scatter, and point-to-point communication (Send / Receive).
[0038] In this embodiment, the working process of UCCL can be summarized into three stages: communication initialization, topology discovery and loop establishment, and set operation execution.
[0039] Communication Initialization: This involves identifying the entities (processes) participating in the communication and establishing a communication domain. In distributed training, each GPU on the server corresponds to a process, called a rank. Each rank has a unique ID. The set of all ranks constitutes a communication domain (communicator). During initialization, the root rank (e.g., rank 0) typically generates a unique ID and passes it to all other ranks in some way (e.g., MPI broadcast). Then, each rank calls the `ucclCommInitRank` function, using this shared ID to initialize its communication domain, so that all ranks are aware of each other's existence and form a communication group.
[0040] Topology Discovery and Loop Building: UCCL automatically detects the system's hardware topology. This includes identifying GPUs, CPUs, PCIe switches, NVLink connections within nodes, and network interface cards (NICs) between nodes. UCCL organizes these hardware devices and their connections into a topology (typically represented as an XML tree). Based on this topology information, UCCL calculates efficient communication paths for aggregate operations (such as AllReduce), for example, by building rings or trees. In a ring structure, data is passed sequentially, efficiently utilizing the bandwidth of all links.
[0041] After UCCL completes the construction and ring establishment of the topology, it reports information such as GPU, RANK, and RDMA network card to NCC through the management network. NCC then establishes and maintains the overall communication topology information, including GPU, RANK, network card, IP address, server and NCP, and port number, as shown in Figure 1.
[0042] Table 1 Once the physical connection between the AI computing server cluster and the cloud cluster switch is established, the UCCL collects the communication group information and reports it to the NCC. The NCC calculates the theoretical bus bandwidth of the AlltoAll operator based on the number of servers, GPUs, and network interface card (NIC) speeds, and performs traffic planning. When sending traffic between servers, if the source IP address of the source GPU sending data is from a specific NIC, then it's generally expected that packets from that GPU will be sent out through that NIC. In this case, the IP table is used to bind packets from that source IP to that NIC for forwarding. This IP table forwarding method allows the NCC to distribute the data to the NCP, ensuring sufficient bandwidth. Simultaneously, if the link bandwidth between the NCP and NCF is insufficient, load balancing is implemented, either through port aggregation or equal-cost routing. The NCC tests the busbw of the aggregated communication by calling the UCCL communication operator to ensure that the expected values are met. If the actual value differs significantly from the theoretical value, start tuning the UCCL environment variables and adjust parameters such as the number of queue pairs used by each UCCL connection, data transfer between GPU memory and RDMA network card, cross-server node communication paths, and data transfer between processes or GPUs within the same node using shared memory, to improve communication performance and thus increase the measured value.
[0043] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0044] In step S201, the interface occupancy rate between each NCP and the corresponding server is monitored by the NCC, and the server whose interface occupancy rate reaches the first threshold is selected as the target server.
[0045] In one implementation, after the NCC determines the target server, it can send a first notification message to the UCCLs of other servers besides the target server. The first notification message carries the target address information of the target server, so that the other servers can identify that the target server corresponding to the target address information is busy after receiving the first notification message.
[0046] Then, each UCCL uses Table 1 to filter out the target traffic to be sent to the target server and adjusts to reduce the rate of the target traffic.
[0047] In another implementation, since there is interactive communication between the NCC and each UCCL, the NCC notifies the UCCL of the traffic and remaining effective bandwidth of each port. In this way, the UCCL can calculate the effective bandwidth of the RDMA network card, thereby increasing the corresponding time and reducing the data transmission rate during communication operations, thus avoiding congestion from the source.
[0048] As can be seen from the above embodiments, by establishing interactive information through the control center (NCC) of the cloud cluster switch and the UCCL communication library of the AI server, the NCC can perceive the network traffic volume from the traffic source and stabilize the network load through planning and scheduling at the traffic source. Through the cloud cluster switch's full traffic scheduling technology, we need to obtain the server GPU's traffic communication model, perform global coordination and scheduling along the entire path, and have a full-path backpressure mechanism to drive the UCCL to proactively reduce the communication atomic time when network bandwidth utilization is high, reducing packet input from the source and ensuring network congestion-free operation.
[0049] Based on the above method embodiments, this disclosure also provides a cloud cluster device, the device including: a cloud cluster controller NCC, a network control protocol NCP and a collective communication library UCCL deployed in each server, wherein the NCC is communicatively connected to the UCCL and NCP in each server respectively, and each NCP is communicatively connected to each server. The NCC monitors the bandwidth utilization of each server through a monitoring module; When the monitoring module detects that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs of other servers other than the target server that the target server is busy, so that the UCCLs of other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
[0050] The NCC also includes: a receiving module. The receiving module is used to receive the communication topology sent by UCCL in each server and integrate it into a total communication topology. The overall communication topology is sent to the UCCL in each server via the sending module.
[0051] The overall communication topology includes some or all of the following information: GPU number, RANK number, RDMA network card number, IP address, Server number, NCP number, and port number.
[0052] Specifically, the monitoring module is used to monitor the interface occupancy rate between each NCP and its corresponding server, and to designate the server whose interface occupancy rate reaches a first threshold as the target server.
[0053] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0054] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.
[0055] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.
[0056] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. A method for controlling traffic flow, characterized in that, The method includes: The cloud cluster controller (NCC) in the cloud cluster switch establishes communication with the collective communication library (UCCL) of each server in the server cluster; the NCC monitors the bandwidth utilization of each server. When the NCC monitors that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs on other servers other than the target server that the target server is busy, so that the UCCLs on other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
2. The method according to claim 1, characterized in that, The method further includes: NCC receives the communication topology sent by UCCL from each server and integrates it into the overall communication topology; The overall communication topology is sent to the UCCL in each server.
3. The method according to claim 2, characterized in that, The overall communication topology includes some or all of the following information: GPU ID, RANK ID, RDMA NIC ID, IP address, Server ID, NCP ID, and port ID.
4. The method according to claim 1, characterized in that, The methods used by NCC to monitor the bandwidth utilization of a target server include: NCC monitors the interface occupancy rate between each NCP and its corresponding server, and identifies the server whose interface occupancy rate reaches the first threshold as the target server.
5. The method according to claim 1, characterized in that, The notification that the UCCL target server is busy on servers other than the target server includes: A first notification message is sent to the UCCLs of servers other than the target server. The first notification message carries the target address information of the target server so that the other servers can identify that the target server corresponding to the target address information is busy after receiving the first notification message.
6. The method according to claim 2, characterized in that, The method of adjusting the UCCL in servers other than the target server to reduce the rate at which traffic is sent to the target server includes: This allows UCCLs in other servers to filter out target traffic destined for the target server from the overall communication topology; Adjust the rate at which the target traffic is reduced.
7. A cloud cluster device, characterized in that, The device includes: a cloud cluster controller NCC, a network control protocol NCP, and a collective communication library UCCL deployed in each server, wherein the NCC is connected to the UCCL and NCP in each server respectively, and each NCP is connected to each server. The NCC monitors the bandwidth utilization of each server through a monitoring module; When the monitoring module detects that the bandwidth utilization of the target server has reached the first threshold, it notifies the UCCLs of other servers other than the target server that the target server is busy, so that the UCCLs of other servers other than the target server adjust and reduce the rate at which they send traffic to the target server.
8. The device according to claim 7, characterized in that, The NCC also includes: a receiving module, The receiving module is used to receive the communication topology sent by UCCL in each server and integrate it into a total communication topology. The overall communication topology is sent to the UCCL in each server via the sending module.
9. The device according to claim 8, characterized in that, The overall communication topology includes some or all of the following information: GPU ID, RANK ID, RDMA NIC ID, IP address, Server ID, NCP ID, and port ID.
10. The device according to claim 7, characterized in that, The monitoring module is specifically used to monitor the interface occupancy rate between each NCP and its corresponding server, and to designate the server whose interface occupancy rate reaches a first threshold as the target server.