Network flow control method based on flow block granularity
Through a network traffic control method based on flow block granularity, hash conflicts and disorder problems in LLM training are solved, more efficient communication and load balancing are achieved, the traffic pattern of LLM is adapted and the PFC mechanism of the RDMA network is perceived, thus improving the distributed training performance of large language models.
Patent Information
- Application Number
- CN202510826648.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
In existing technologies, the ECMP mechanism causes serious hash conflicts and uneven link utilization in large language model (LLM) training, which prolongs communication time. In addition, traditional load balancing solutions fail to adapt to LLM traffic patterns and the PFC mechanism of RDMA networks, leading to out-of-order problems.
A network flow control method based on flow block granularity is adopted. By calculating the flow block granularity and the expected departure time, the egress port with the minimum expected departure time is selected for data packet transmission, dynamically adapting to the LLM traffic pattern, and perceiving the PFC mechanism to avoid disorder.
Improves the communication efficiency of LLM training, reduces the possibility of out-of-order data packets, and improves network utilization and overall training performance.
Smart Images

Figure CN120750868A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a network flow control method based on flow block granularity. Background Art
[0002] Distributed training of Large Language Models (LLMs) relies on multi-GPU parallel computing and requires large-scale collective communication coordination. However, inter-GPU communication overhead is significantly affected by the parallelization strategy and requires optimization to reduce latency. In data center networks, the ECMP mechanism splits traffic by flow across multiple paths. However, LLM training traffic, due to its low entropy (a high proportion of long flows and a small number of flows), is prone to severe hash conflicts and uneven link utilization, which in turn prolongs communication time and slows overall training efficiency.
[0003] Existing fine-grained switch-based load balancing solutions, such as CONGA and LetFlow, use flowlets (data blocks divided based on the gaps between packets within a flow) to achieve link load balancing. While this can avoid the disorder problem associated with packet-level load balancing, it has the following limitations in LLM scenarios: 1. Insufficient dynamic adaptability of flowlets: Flowlet partitioning based on fixed timeouts is difficult to adapt to different LLM traffic patterns (such as periodic bursts and long-flow dominance), resulting in rigid path selection. 2. Lack of RDMA network compatibility: Traditional solutions do not integrate the underlying PFC (priority flow control) mechanism of the RDMA network, which can easily lead to routing misjudgments and disorder problems, exacerbating congestion risks. Summary of the Invention
[0004] In view of this, an embodiment of the present application provides a network traffic control method based on flow block granularity to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of an embodiment of the present application provides a network traffic control method based on flow block granularity, the method comprising: Determine a traffic type of current traffic that a sender in a current model needs to send to a receiver, where the current traffic includes multiple data packets; Calculating a flow block granularity corresponding to the traffic type of the current traffic based on the parameters provided by the current model, and sending the flow block granularity to the switch; Determining an estimated departure time of packets queued at each egress port of the switch, and selecting an egress port with a minimum estimated departure time as a target egress port; wherein the estimated departure time is based on a queue length and a pause time of the egress port; Establishing a flow block for the current traffic according to the flow block granularity, and updating the flow block table of the switch according to the established flow block; According to the flow block table, the data packets in the flow block are sent to the receiving end through the target egress port.
[0006] Optionally, determining the traffic type of current traffic that the sending end needs to send to the receiving end in the current model includes: Determining, based on the traffic type of the current traffic, a calculation method for a flow block granularity corresponding to the traffic type of the current traffic; The calculating, based on the parameters provided by the current model, a flow block granularity corresponding to the flow type of the current flow, includes: According to the calculation method of the flow block granularity corresponding to the flow type of the current flow, the flow block granularity corresponding to the flow type of the current flow is calculated according to the parameters provided by the current model.
[0007] Optionally, before determining the expected departure time of the queued packet in each egress port of the switch, the method further includes: For a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently required to be transmitted is the first data packet in the queue pair corresponding to the transmitting end; In a case where the data packet currently to be transmitted is the first data packet in the queue pair corresponding to the transmitting end, determining the estimated departure time of the queued packet in each egress port of the switch includes: For each egress port of the switch, query the PFC status of the egress port and calculate the time during which the egress port is suspended by PFC; Calculating a sending delay of a queued packet and a processing delay of a queued packet in the egress port; determining an estimated departure time of the egress port based on the time of being paused by the PFC, a sending delay of the queued packet, and a processing delay of the queued packet; The sending of the data packets in the flow block to the receiving end through the target egress port includes: sending the first data packet to the receiving end through the target egress port; The sending delay of the queued packet is determined according to the amount of data in the queue pair and the link rate; The processing delay of the queued packets is determined according to the number of queued packets in the queue pair and the average processing delay of each queued packet; The PFC pause time is determined according to the time when the egress port resumes sending data and the current time of the routing decision.
[0008] Optionally, the method further includes: Recording a destination egress port for transmitting the first data packet; For a data packet currently to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently to be transmitted belongs to the flow block where the first data packet is located; In a case where the data packet currently to be transmitted belongs to the flow block where the first data packet is located, the data packet currently to be transmitted is sent to the receiving end through the target egress port.
[0009] Optionally, the method further includes: For a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently required to be transmitted belongs to the last data packet in the flow block where the first data packet is located; In the case that the data packet currently to be transmitted is the last data packet in the flow block where the first data packet is located, a next target egress port is re-determined, and the next target egress port is used to transmit multiple data packets in the next flow block.
[0010] Optionally, the calculating method of the flow block granularity corresponding to the flow type of the current flow according to the parameters provided by the current model includes: When the traffic type of the current traffic is PP communication traffic, calculating a flow block granularity corresponding to the traffic type of the current traffic according to the sequence length, micro-batch size, and hidden size of the current model; When the traffic type of the current traffic is DP communication traffic, the flow block granularity corresponding to the traffic type of the current traffic is calculated based on the number of parameters of the current model, the size of data parallelism and the number of segments in the segmented pipeline ring.
[0011] Optionally, when the traffic type of the current traffic is PP communication traffic, the flow block granularity The calculation formula is: ; in, is the sequence length, is the mini-batch size, is the hidden size.
[0012] Optionally, when the traffic type of the current traffic is DP communication traffic, the flow block granularity The calculation formula is: ; in, is the number of parameters of the current model, is the size of the data parallelism, is the number of stages in the segmented pipeline loop.
[0013] Optionally, the method further includes: For a data packet currently to be transmitted among a plurality of data packets included in the current traffic, determining whether a size of the data packet currently to be transmitted exceeds the flow block granularity; In a case where the size of the data packet currently to be transmitted exceeds the flow block granularity, querying the PFC state of each egress port of the switch and calculating the estimated departure time of the queued packet of each egress port; Based on the expected departure time of the queued packets of each egress port, selecting the egress port with the minimum expected departure time as the new target egress port; A new flow block is established, and according to the flow block table, the data packets in the new flow block are sent to the receiving end through the new target egress port.
[0014] A second aspect of an embodiment of the present application provides a network traffic control device based on flow block granularity, the device comprising: Beneficial effects of this application: An embodiment of the present application provides a network traffic control method based on flow block granularity, the method comprising: determining the traffic type of current traffic that a sender in a current model needs to send to a receiver, the current traffic including multiple data packets; calculating a flow block granularity corresponding to the traffic type of the current traffic based on parameters provided by the current model, and sending the flow block granularity to a switch; determining an expected departure time of queued packets in each egress port of the switch, and selecting an egress port with a minimum expected departure time as a target egress port; the expected departure time is obtained based on a queue length and a pause time of the egress port; establishing a flow block for the current traffic according to the flow block granularity, and updating a flow block table of the switch according to the established flow block; and sending the data packets in the flow block to the receiver through the target egress port according to the flow block table.
[0015] Through the technical solution of the present application, the size of the flow block granularity can be dynamically adjusted according to the traffic type of the current traffic that needs to be transmitted by the current model and the parameters provided by the current model to adapt to the traffic pattern of the current model. Furthermore, after calculating the flow block granularity, the expected departure time of each exit port of the switch is determined, and the exit port with the smallest expected departure time is selected as the target exit port. This not only improves the transmission efficiency of the flow block that currently needs to be transmitted, but also ensures that the data packets of the same flow block can be transmitted through the same exit port, reducing the possibility of out-of-order data packets in multiple data packets in the same flow block. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings that constitute a part of this application are used to provide further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute improper limitations on this application.
[0017] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a schematic diagram of hybrid parallelism in LLM training provided in the related art; Figure 2 It is a flowchart of ECMP provided in the related art; Figure 3 It is a flowchart of CONGA provided in the related art; Figure 4 It is the LetFlow flowchart provided in the related art; Figure 5 This is a flow chart of a network traffic control method based on flow block granularity provided by an embodiment of the present application; Figure 6 This is an overall flow chart of a network traffic control method based on flow block granularity provided by an embodiment of the present application; Figure 7 This is a flow chart of load balancing for the collective communication traffic of the large language model GPT-7B provided by one embodiment of the present application; Figure 8 This is a flow chart of load balancing for the collective communication traffic of a large language model LLama-65B provided by one embodiment of the present application; Figure 9 This is a schematic diagram of a framework of a network traffic control device based on flow block granularity provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0020] In related technologies, the scale of LLM parameters based on Transformer has grown exponentially, making distributed training a necessary solution. Mainstream parallelization strategies include data parallelism (DP), pipeline parallelism (PP), and tensor parallelism (TP). Figure 1 It is a schematic diagram of hybrid parallelism in LLM training provided in the related art. Table 1 lists the key parameters of hybrid parallelism in LLM training.
[0021] Table 1: Key parameters of hybrid parallelism in LLM training
[0022] refer to Figure 1 As shown in Table 1, the communication characteristics of hybrid parallelism in training of large language models are as follows: Data Parallelism (DP): Each DP group holds a complete copy of the model and regularly aggregates gradients through inter-node communication (the amount of communication is positively correlated with the number of parameters) to ensure global weight consistency; Pipeline Parallelism (PP): Model layers are distributed to multiple devices, executed in micro-batches, and peer-to-peer (P2P) communication between nodes synchronizes the computation results of each layer within the PP group. Tensor Parallelism (TP): Single-layer computations are split across multiple GPUs within a node. After self-attention and MLP computations, all-reduce synchronization of tensors (of size bsh) is required within the node. Communication is limited to within the node.
[0023] Collective communication optimization becomes key: All-reduce: used for DP gradient synchronization and TP activation synchronization. It is decomposed into reduce-scatter (spreading) and all-gather (global aggregation). The ring pipeline segmentation can reduce the overhead of large-scale parameter synchronization. All-to-all: Applied to MoE architectures (such as GPT-4), the all-to-all communication mode causes a surge in traffic within and outside nodes, requiring targeted optimization of topology-aware routing. Hybrid communication mode: The combination of different parallel strategies leads to complex communication topology (such as cross-node DP + intra-node TP), requiring coordinated optimization of bandwidth allocation and congestion control.
[0024] Challenges and Trends: As model parameters increase, all-reduce communication volume increases superlinearly. The all-to-all communication introduced by MoE places higher demands on network load balancing. There is an urgent need to design low-overhead communication protocols that combine hardware topology (such as NVLink / InfiniBand) with algorithm characteristics (such as sparse communication).
[0025] Among them, in the related art: (1) ECMP (Equal Cost Multi-Path Routing) ECMP (Equal-Cost Multi-Path) is a network routing strategy. The process is as follows Figure 2 As shown in the figure, packets can be forwarded along multiple equal-cost paths, achieving load balancing and improving network utilization. In large-scale RDMA networks, the Clos architecture based on the Fat-Tree architecture is often used to connect more servers. The fundamental concept of the Clos architecture based on the Fat-Tree architecture is to use a large number of commodity switches to construct multiple equal-cost paths between servers, thereby forming a large-scale, non-blocking network. When selecting paths for actual traffic, switches perform ECMP on the flows to achieve load balancing. ECMP typically hashes the flow's five-tuple, such as the source IP address, destination IP address, protocol, source port number, and destination port number, to select one of multiple equal-cost paths for transmission. ECMP works well in most cases. It leverages the randomness of the hash function to achieve good load balancing when there are a large number of flows.
[0026] (2) CONGA CONGA is a network-based distributed congestion-aware data center load balancing mechanism. The process is as follows: Figure 3As shown in the figure. When a data packet is sent from the edge, it first passes through the sending leaf switch. The remote congestion metric FbPathId and FbMetric from the receiving leaf switch to the sending leaf switch are queried in the Congestion-From-Leaf table. If the remote metric information of the receiving leaf switch exists, it is encapsulated into the data packet. The flowlet table is then queried to determine whether a flowlet has been established for the flow corresponding to the data packet. If the corresponding table entry exists and the flowlet has not timed out, the entry is updated, routing is performed based on the egress port recorded in the entry, the local congestion metric information CE of the egress port is measured, and the information is encapsulated into the data packet. If the flowlet has timed out or the entry does not exist, a new path with the lowest congestion is selected based on the local congestion metric information and the remote congestion metric information from the sending leaf to the receiving leaf. The corresponding information of this flowlet is recorded, the local congestion metric CE of the egress port of the new path is measured, and the information is encapsulated into the data packet. When the data packet reaches a non-leaf switch, the local congestion metric CE is updated, encapsulated into the data packet, and the corresponding egress port is selected based on the selected path. When a data packet arrives at the receiving-side leaf switch, the Congestion-To-Leaf table is updated based on the remote congestion metric information FbPathId and FbMetric from the receiving-side leaf switch to the sending-side leaf switch recorded in the data packet; and the Congestion-From-Leaf table is updated based on the local congestion metric information recorded by the CE.
[0027] (3) LetFlow LetFlow is a simple data center network load balancing mechanism that simply randomly selects paths for Flowlets and lets their elasticity naturally balance traffic on different paths. The process is as follows Figure 4 As shown in the figure. When a packet is sent from the edge, it first passes through the sending ToR switch. The flowlet table is checked to see if a flowlet has been established for the flow corresponding to the packet. If the corresponding table entry exists and the flowlet has not timed out, the entry is updated and routing is performed based on the egress port recorded in the entry. If the flowlet has timed out or the entry does not exist, a new path is randomly selected and the corresponding flowlet information is recorded. When the packet reaches the non-ToR switch, the corresponding egress port is selected based on the selected path. When the packet reaches the receiving ToR switch, it is sent to the receiving network card.
[0028] Specifically, ECMP is widely used in switches to distribute traffic on a per-flow basis. However, the traffic generated by LLM training often has a high proportion of long flows and low entropy. As a result, the opportunity to divert traffic to other available links is greatly reduced, resulting in insufficient link bandwidth utilization, severe flow hash conflicts, extended communication time, and ultimately slowing down overall training time.
[0029] CONGA and LetFlow both make rerouting decisions for each flowlet on the switch. A flowlet is a data block separated by the inter-packet interval in a flow (i.e., a new flowlet is identified when the inter-packet interval exceeds a predetermined timeout). Flowlets can achieve more even load balancing while avoiding the severe out-of-order issues that come with load balancing. Commercial switches currently support flowlet rerouting. However, while splitting flows into smaller units for load balancing has proven effective in traditional data center networks, their performance under LLM traffic has been rarely studied. Under LLM traffic, different models may require different appropriate flowlet timeout values, and timeout-based flowlets are not adaptable to traffic from different LLMs. Furthermore, the PFC mechanism in RDMA networks can further lead to out-of-order packets. If PFC is in effect, the corresponding upstream link is paused and prevented from transmitting packets. As a result, packets sent earlier may be blocked by the path where PFC is paused, and then arrive later than packets sent later in the same flow (which may have traveled over a healthy link), resulting in out-of-order packets. Neither LetFlow nor CONGA consider PFC in their rerouting decisions. For example, LetFlow randomly selects a rerouting path for each flowlet even when the outgoing link is paused. CONGA selects rerouting paths by calculating the link utilization of equal-cost paths. However, PFC's PAUSE frames reduce the link utilization of affected paths, leading to suboptimal routing decisions. Therefore, switch-based rerouting that is unaware of PFC can exacerbate out-of-order traffic. Both LetFlow and CONGA are designed primarily for traditional TCP networks and lack PFC awareness.
[0030] Based on the above defects, the technical problems to be solved by the technical solution of this application are: 1. How to consider the characteristics of different LLM training traffic and introduce the basic routing unit of flow blocks based on the characteristics of different LLM training traffic to avoid out-of-order problems and improve load balancing efficiency; 2. How to consider the congestion-awareness capability of the PFC flow control mechanism at the bottom layer of the RDMA network to avoid the problem of PFC-unaware switch-based rerouting exacerbating the disorder problem.
[0031] Therefore, this application designs a customized load balancing solution for the communication characteristics of large language models (low entropy, long and dense flows) and RDMA network architecture to balance the flowlet partitioning accuracy and data packet orderliness, and improve distributed training performance.
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] Figure 5 This is a flow chart of a network traffic control method based on flow block granularity provided by an embodiment of the present application. Figure 6 This is an overall flow chart of a network traffic control method based on flow block granularity provided by an embodiment of the present application.
[0034] The core concept of this application's technical solution lies in using size-based flow blocks, rather than timeout-based flowlets, as the rerouting unit for LLM traffic. Flow blocks are divided by size rather than by inter-packet interval. In other words, the flow block granularity (block size) is determined by the characteristics of the collective communication patterns generated by the workload layer of a large language model. This approach avoids the complex flowlet timeout selection required in traditional approaches and enables better adaptation to different LLMs.
[0035] At the application workload level, for LLM, frameworks like Megatron typically maintain model parameters and schedule collective communication for each sender-receiver pair in the training cluster. This includes parameters such as model size, configuration parameters for the parallelization strategy, and message sizes for DP and PP parallel communication. Then, given a specific set of hybrid parallelization strategy parameters and model size information, a flow block granularity is identified to enable load balancing and enforced on network switches. The network controller is responsible for distributing this flow block granularity to top-of-rack (ToR) switches.
[0036] Each ToR switch maintains a flow block table. Upon receiving flow block granularity information, the switch updates the flow block table and selects the optimal forwarding path for each flow block to minimize out-of-order packets. To accurately measure link congestion, the switch uses the pause time of each egress port and the queue length of each egress port to calculate the expected send time and select the appropriate egress port. This approach achieves load balancing for PFC-aware and LLM-aware traffic.
[0037] refer to Figure 5An embodiment of the present application provides a network traffic control method based on flow block granularity, which specifically includes steps S11 to S15: Step S11, determining the traffic type of the current traffic that the sending end in the current model needs to send to the receiving end, where the current traffic includes multiple data packets.
[0038] In a distributed training environment, cross-node communication traffic for large language models is primarily divided into two types: data parallel (DP) traffic and pipeline parallel (PP) traffic. Data parallel traffic involves synchronizing gradients between GPUs on multiple nodes, while pipeline parallel traffic involves synchronizing computational results between different layers of a large language model.
[0039] Therefore, in this embodiment, reference Figure 6 In order to optimize the efficiency of load balancing during the transmission of communication traffic across nodes in training a large language model, before the communication traffic of the current model (i.e., current traffic) is transmitted, it is first necessary to determine the traffic type of the current traffic of the current model, where the current model refers to the large language model to be trained, the current traffic refers to the multiple data packets contained in the communication traffic that need to be transmitted between GPUs of multiple different nodes in the current model, the sender refers to the GPU that needs to send the current traffic, and the receiver refers to the GPU that needs to receive the current traffic sent by the sender.
[0040] Step S12: Calculate the flow block granularity corresponding to the traffic type of the current traffic according to the parameters provided by the current model, and send the flow block granularity to the switch.
[0041] In this embodiment, after determining the traffic type of the current flow, the flow block granularity corresponding to the current traffic type is calculated based on the parameters provided by the current model. The flow block granularity is dynamically calculated based on the communication pattern of the large language model (reflected by the parameters provided by the current model), which can better adapt to different traffic types. The calculated flow block granularity is distributed to the switch (ToR switch) through the flow block granularity network controller.
[0042] Step S13, determining the expected departure time of the queued packets in each egress port of the switch, and selecting the egress port with the smallest expected departure time as the target egress port; the expected departure time is obtained based on the queue length and pause time of the egress port.
[0043] In this embodiment, after receiving the size of the flow block granularity corresponding to the current traffic, the switch needs to determine the estimated departure time of the queued packets in each egress port. The queued packets are the data packets already in the egress port and waiting to be transmitted. The estimated departure time is the waiting time required to transmit the data packets in the flow block corresponding to the current traffic.
[0044] Specifically, the estimated departure time of the egress port is determined based on the queue length and pause time of the egress port. The packet buffer queue length of the egress port can indicate the time required for the queued packet to be actually sent, while the pause time of the egress port indicates the PFC (Priority Flow Control) status of the egress port and the time it is paused by PFC.
[0045] After obtaining the estimated departure time of each egress port, by selecting the egress port with the minimum estimated departure time as the target egress port for the flow block that currently needs to be transmitted, it can ensure that the multiple data packets contained in the flow block that currently needs to be transmitted can reach the receiving end as soon as possible based on the minimum estimated departure time. The minimum estimated departure time can ensure that the current flow block is transmitted to the receiving end in the fastest time. In addition, by sensing the network congestion status, it can avoid selecting the egress port that is paused by PFC, thereby reducing the possibility of out-of-order data packets appearing during the transmission of multiple data packets in the same flow block.
[0046] Step S14: Create flow blocks for the current traffic according to the flow block granularity, and update the flow block table of the switch according to the created flow blocks.
[0047] In this embodiment, after determining the target egress port, the switch will establish a flow block for the current traffic based on the calculated flow block granularity, that is, the total size of the flow block is equal to the flow block granularity, and each flow block contains a certain number of data packets. Then, the switch updates the flow block table in the switch based on the established flow block, ensuring that data packets of the same flow block can be transmitted through the same egress port, reducing the possibility of out-of-order data packets appearing in multiple data packets in the same flow block.
[0048] Step S15: according to the flow block table, sending the data packets in the flow block to the receiving end through the target egress port.
[0049] In this embodiment, the switch sends each data packet in the flow block to the receiving end through the target egress port based on the path information in the flow block table (i.e., the target egress port), thereby ensuring that multiple data packets in the same flow block can be efficiently and orderly transmitted to the target device, reducing the number of retransmissions and improving network utilization.
[0050] Through the technical solution of the above embodiment, the size of the flow block granularity can be dynamically adjusted according to the traffic type of the current traffic that needs to be transmitted by the current model and the parameters provided by the current model. The size of the dynamic flow block granularity can better adapt to the traffic pattern of the current model. Furthermore, after calculating the flow block granularity, the expected departure time of each exit port of the switch is determined, and the exit port with the smallest expected departure time is selected as the target exit port. This not only improves the transmission efficiency of the flow block that currently needs to be transmitted, but also ensures that the data packets of the same flow block can be transmitted through the same exit port, reducing the possibility of out-of-order data packets in multiple data packets in the same flow block.
[0051] Furthermore, when selecting the target egress port, the PFC mechanism is used to avoid selecting a suspended egress port, further improving the stability and efficiency of network transmission. Ultimately, this approach significantly improves the communication efficiency of large language models in distributed training, reduces communication latency, and enhances overall training performance.
[0052] In combination with the above embodiments, the present application also provides another network traffic control method based on flow block granularity, in which: The "determining the traffic type of the current traffic that the sender in the current model needs to send to the receiver" in step S11 specifically includes: step S11-1: based on the traffic type of the current traffic, determining the calculation method of the flow block granularity corresponding to the traffic type of the current traffic.
[0053] In this embodiment, the flow block granularities of different traffic types are calculated based on different parameters provided by the current model, and the calculated flow block granularities are then sent to the switch so that the switch can perform load balancing based on these flow block granularities.
[0054] For example, for pipeline parallel (PP) traffic, the flow block granularity is related to the sequence length, mini-batch size, and hidden size of the current model; for data parallel (DP) traffic, the flow block granularity is related to the number of parameters of the current model, the data parallel size, and the number of segments in the segmented pipeline ring.
[0055] The step S12 of "calculating the flow block granularity corresponding to the traffic type of the current traffic according to the parameters provided by the current model" specifically includes: step S12-1, according to the calculation method of the flow block granularity corresponding to the traffic type of the current traffic, calculating the flow block granularity corresponding to the traffic type of the current traffic according to the parameters provided by the current model.
[0056] In this embodiment, after determining the traffic type corresponding to the current traffic that needs to be transmitted, the parameters of the current model required to calculate the flow block granularity of the traffic type are determined from the parameters provided by the current model, and the size of the flow block particle corresponding to the traffic type is calculated.
[0057] In combination with the above embodiment, the present application also provides another network traffic control method based on flow block granularity. In this method, the step S12-1 of "calculating the flow block granularity corresponding to the flow type of the current flow according to the parameters provided by the current model according to the flow block granularity calculation method corresponding to the flow type of the current flow" specifically includes steps S12-1-1 and S12-1-2: Step S12-1-1, when the traffic type of the current traffic is PP communication traffic, calculate the flow block granularity corresponding to the traffic type of the current traffic according to the sequence length, micro-batch size and hidden size of the current model.
[0058] Specifically, when the traffic type of the current traffic is PP communication traffic, the flow block granularity The calculation formula is: ; in, is the sequence length, is the mini-batch size, is the hidden size.
[0059] Step S12-1-2, when the traffic type of the current traffic is DP communication traffic, calculate the flow block granularity corresponding to the traffic type of the current traffic based on the number of parameters of the current model, the size of data parallelism and the number of segments in the segmented pipeline ring.
[0060] Specifically, when the traffic type of the current traffic is DP communication traffic, the flow block granularity The calculation formula is: ; in, is the number of parameters of the current model (i.e., the model size of the current model), is the size of the data parallelism, is the number of stages in the segmented pipeline loop.
[0061] In combination with the above embodiments, the present application further provides another network traffic control method based on flow block granularity. In this method, before executing step S13 of "determining the expected departure time of the queued packet at each egress port of the switch", step S21 is further included: Step S21, for a data packet currently required to be transmitted among multiple data packets included in the current traffic, determine in the flow block table whether the data packet currently required to be transmitted is the first data packet in the queue pair corresponding to the sending end.
[0062] In this embodiment, during the distributed training of a large language model, data packets are transmitted through queue pairs (QPs) of the RDMA network card of the GPU on the sending end. Each queue pair has a send queue (SQ) for storing data packets to be sent.
[0063] Before performing load balancing during packet transmission, it is necessary to determine whether the packet currently to be transmitted is the first packet in the queue pair. The transmission path (i.e., the target egress port) for the first packet in a flow block is selected based on the estimated departure time of each egress port. Subsequent packets in the flow block can be quickly forwarded using the existing path information of the target egress port for transmitting the first packet.
[0064] In a case where the data packet currently to be transmitted is the first data packet of the queue pair corresponding to the transmitting end, the step of "determining the estimated departure time of the queued packet in each egress port of the switch" in step S13 specifically includes steps S13-1 to S13-3: Step S13-1: For each egress port of the switch, query the PFC status of the egress port and calculate the time when the egress port is suspended by PFC; the PFC suspension time is determined based on the time when the egress port resumes sending data and the current time of the routing decision.
[0065] In this embodiment, after determining that the current data packet is the first data packet in the queue pair, it is necessary to query the PFC status of each egress port of the switch. The PFC mechanism is used to prevent network congestion. When the queue length of a certain egress port exceeds a certain threshold, the PFC mechanism will suspend data transmission at the egress port by sending a PAUSE frame. In addition, when the egress port is suspended from data transmission, it is also necessary to calculate the time the egress port is suspended by PFC to avoid selecting an egress port that has been suspended for a long time as the target egress port.
[0066] Specifically, by querying the PFC status of each egress port and calculating the time the egress port is suspended by PFC ,in, The time it takes for the egress port to resume sending data. It is determined based on the time of receiving the PAUSE frame and the pause time record preset in the PAUSE frame. is the current time.
[0067] Step S13-2: Calculate the sending delay and processing delay of the queued packets at the egress port. The sending delay of the queued packets is determined based on the data volume and link rate in the queue pair; the processing delay of the queued packets is determined based on the number of queued packets in the queue pair and the average processing delay of each queued packet.
[0068] In this embodiment, after determining the PFC status and PFC pause time for each egress port, the send delay and processing delay of queued packets at each egress port need to be calculated. The send delay of a queued packet refers to the time it takes for a packet to be sent from the queue, while the processing delay of a queued packet refers to the time a packet waits in the queue for processing.
[0069] Specifically, the calculation formula for the sending delay of the queued packet at the egress port is: ,in The amount of data in the packet buffer queue (in bytes). is the link rate.
[0070] Processing delay of queued packets at the egress port Total processing delay ,in, The number of packets in the packet buffer queue, The average processing delay for each packet.
[0071] Step S13-3: determining the estimated departure time of the egress port according to the PFC pause time, the sending delay of the queued packet, and the processing delay of the queued packet.
[0072] In this embodiment, the estimated departure time of the queued packet at each egress port is: .
[0073] The step S15 of "sending the data packets in the flow block to the receiving end through the target egress port" includes: step S15-1, sending the first data packet to the receiving end through the target egress port; In this embodiment, the first data packet is sent to the receiving end through the target egress port, thereby transmitting through the optimal path and reducing transmission delay.
[0074] In combination with the above embodiment, the present application further provides another network traffic control method based on flow block granularity, wherein the method further includes steps S31 to S33: Step S31: Record the target egress port used to transmit the first data packet.
[0075] In this embodiment, after the first data packet arrives at the switch and the path selection is completed to determine the target egress port of the first data packet, the target egress port is recorded.
[0076] Step S32: For a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, judging in the flow block table whether the data packet currently required to be transmitted belongs to the flow block where the first data packet is located.
[0077] In this embodiment, when a subsequent data packet arrives, for the data packet that currently needs to be transmitted, it is necessary to determine whether the data packet belongs to the same flow block as the first data packet, so as to determine the transmission path of the data packet.
[0078] Step S33: When the data packet currently to be transmitted belongs to the flow block where the first data packet is located, the data packet currently to be transmitted is sent to the receiving end through the target egress port.
[0079] In this embodiment, if the data packet that currently needs to be transmitted belongs to the flow block where the first data packet is located, the target egress port recorded in step S31 will be directly used to send the data packet directly to the receiving end, thereby utilizing the path information of the previously recorded target egress port, not only reducing the overhead of path selection, but also reducing the possibility of out-of-order data packets during the transmission of data packets in the same flow block.
[0080] In combination with the above embodiment, the present application further provides another network traffic control method based on flow block granularity, wherein the method further includes step S41 and step S42: Step S41, for a data packet currently required to be transmitted among multiple data packets included in the current traffic, in the flow block table, determining whether the data packet currently required to be transmitted belongs to the last data packet in the flow block where the first data packet is located.
[0081] In this embodiment, data packets are transmitted in flow blocks, each containing multiple data packets whose size is determined by the flow block granularity. To further optimize load balancing and reduce out-of-order data packets, during the transmission of each data packet, it is necessary to determine whether the data packet currently being transmitted is the last data packet in the current flow block based on the flow block table. The flow block table records the path information and current size of each flow block.
[0082] For example, assume the current stream block has a 1MB block granularity and two packets have already been transmitted, totaling 800KB. Now, a third packet (the current packet) needs to be transmitted, and its size is 200KB. The stream block table is checked to determine whether this 200KB packet is the last packet in the current stream block. Since 800KB + 200KB = 1MB, the block granularity has been reached, and therefore this packet is the last packet in the current stream block.
[0083] Step S42, when the data packet currently to be transmitted is the last data packet in the flow block where the first data packet is located, re-determine the next target egress port, and the next target egress port is used to transmit multiple data packets in the next flow block.
[0084] In this embodiment, when it is identified that the current data packet is the last data packet in the current flow block, the last data packet is still transmitted according to the target egress port corresponding to the first data packet. At the same time, it is necessary to reselect the optimal path for the next flow block, that is, to determine the next target egress port. The process of determining the next target egress port is similar to the previous one, that is, it includes querying the PFC status of each egress port, calculating the estimated departure time, and selecting the port with the smallest estimated departure time as the next target egress port, ensuring that each flow block can be transmitted through the optimal path, reducing network congestion and out-of-order data packets.
[0085] In combination with the above embodiment, the present application further provides another network traffic control method based on flow block granularity, wherein the method further includes steps S51 to S54: Step S51 : for a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, determining whether the size of the data packet currently required to be transmitted exceeds the flow block granularity.
[0086] In this embodiment, for each data packet that currently needs to be processed, it is first necessary to determine whether the relationship between the size of the data packet and the flow block granularity or the remaining size of the flow block in the flow block table meets the data packet that currently needs to be processed, in order to determine whether special processing is required for the data packet.
[0087] For example, assuming that the flow block granularity is 1MB and the current data packet size that needs to be transmitted is 1.2MB, it means that the size of the data packet exceeds the flow block granularity; assuming that the flow block granularity is 1MB, the remaining size of the flow block is 0.5MB, and the current data packet size that needs to be transmitted is 0.8MB, it means that the remaining size of the flow block does not meet the current data packet that needs to be processed.
[0088] The judgment of data packet size ensures that large data packets that require special processing can be identified, thereby preventing a single data packet from occupying too many network resources and affecting overall load balancing.
[0089] Step S52 , when the size of the data packet currently to be transmitted exceeds the flow block granularity, query the PFC state of each egress port of the switch and calculate the estimated departure time of the queued packet of each egress port.
[0090] In this embodiment, when the size of the data packet exceeds the flow block granularity or the remaining size of the current flow block does not meet the data packet that needs to be processed currently, it means that the data packet is a large data packet that cannot be transmitted based on the current flow block. Therefore, the data packet needs to be specially processed, and the PFC status of each egress port is re-queried from multiple egress ports of the switch, and the estimated departure time of the queued packet of each egress port is calculated. This process is similar to the previous one and will not be repeated here.
[0091] Step S53: Based on the expected departure time of the queued packets of each egress port, select the egress port with the minimum expected departure time as the new target egress port.
[0092] In this embodiment, for the estimated departure time of the queued packet of each egress port determined in step S52, the egress port with the smallest estimated departure time is selected as the new target egress port, and the new target egress port is used to transmit the data packet.
[0093] Step S54: Create a new flow block, and send the data packets in the new flow block to the receiving end through the new target egress port according to the flow block table.
[0094] In this embodiment, after selecting a new target egress port, the flow block size is calculated based on the model parameter information, a new flow block is established, and data packets exceeding the size limit are sent to the receiving end through the new target egress port, ensuring that the data packets can be transmitted to the receiving end efficiently and orderly.
[0095] As an example, Figure 7 This is a flowchart of load balancing the collective communication traffic of the large language model GPT-7B provided by one embodiment of the present application.
[0096] refer to Figure 7 , taking the current model as the large language model GPT-7B as an example to illustrate: Network topology: This is a fat-tree topology with K=8 nodes. This topology involves a total of 128 GPUs (4 GPUs per training node), 32 ToR switches, 32 aggregation switches, and 16 core switches. All links are 400Gbps with a propagation delay of 40ns. Congestion control: Use DCQCN as the default congestion control algorithm, and the parameters are configured as ,This configuration can provide low latency and high throughput; Workload: In this example, the workload is the model size 7B, hidden The size is 4096, the sequence length is 256, the number of stages in the segmented pipeline ring The collective communication traffic of the GPT-7B model is 1024; the parameters of the parallelization strategy are ; Workflow: First, based on the model parameter information of GPT-7B, the flow block granularity of PP communication is calculated as ; The flow block granularity of DP communication is calculated as ; Afterwards, when the data packet arrives at the ToR switch on the sending side, it determines whether the current data packet is PP communication traffic or DP communication traffic, and adopts different flow block granularity according to different traffic types; if the current data packet is the first data packet of the current traffic, or the data packets that need to be conveyed currently exceed the flow block granularity, the PFC status of each egress port is queried and the estimated departure time of the data packet at the current egress is calculated, and the egress port with the smallest estimated departure time is selected as the new target egress port to establish a new flow block; if the data packet that needs to be transmitted currently belongs to the current flow block, then it is directly forwarded to the link forwarded by the first data packet (that is, the target egress port of the current flow block).
[0097] As an example, Figure 8 This is a flowchart of load balancing the collective communication traffic of a large language model LLama-65B provided by an embodiment of the present application.
[0098] refer to Figure 8 , taking the current model, the large language model LLama-65B, as an example: Network topology: This is a fat-tree topology with K=8. This topology involves a total of 128 GPUs (4 GPUs per training node), 32 ToR switches, 32 aggregation switches, and 16 core switches. All links are 400Gbps with a propagation delay of 40ns. Congestion control: Use DCQCN as the default congestion control algorithm, and the parameters are configured as ,This configuration can provide low latency and high throughput; Workload: In this example, the workload is the model size 65B, hidden The size is 8192, the sequence length is 256, the number of stages in the segmented pipeline ring The collective communication traffic of the GPT-7B model with a capacity of 1024 is as follows; the parameters of the parallelization strategy are ; Workflow: First, based on the model parameter information of GPT-7B, the flow block granularity of PP communication is calculated as ; The flow block granularity of DP communication is calculated as ; Afterwards, when the data packet arrives at the ToR switch on the sending side, it determines whether the data packet currently to be transmitted is PP communication traffic or DP communication traffic, and adopts different flow block granularity according to different traffic types. If the data packet currently to be transmitted is the first data packet of the current traffic or the data packet currently to be transmitted has exceeded the flow block granularity, the PFC status of each egress port is re-queried and the estimated departure time of the data packet of each egress port is calculated. The egress port with the smallest estimated departure time is selected as the new target egress port, and a new flow block is established; if the data packet currently to be transmitted belongs to the current flow block, it is directly forwarded to the link forwarding the first data packet of the flow block (that is, the target egress port of the current flow block).
[0099] Under the workload of the real models GPT-7B and LLama-65B, this application uses the normalized communication time to evaluate the load balancing performance of the model. The technical solution of this application increases the completion speed of GPT-7B by 34% to 77%, and the completion speed of LLama-65B by 20% to 75%.
[0100] Overall, the technical solution of this application achieves better application performance with less communication time, significantly outperforming both flow- and flowlet-level solutions. This performance improvement over Conweave is attributed to the PFC-aware rerouting mechanism employed in this application. The number of out-of-order packets directly corresponds to the number of retransmissions in an RDMA network. Unlike LetFlow and CONGA, the technical solution of this application exhibits minimal out-of-order packets, thus avoiding bandwidth waste due to retransmissions.
[0101] In summary, the core technical solution of this application lies in proposing a flow-block granularity-based network traffic control method, specifically targeting the communication traffic of large language models (LLMs) over RDMA networks. By identifying flow blocks under different model training parallelization strategies, combined with application-layer and network-layer collaborative design, this method optimizes load balancing, avoids out-of-order packets, and perceives the underlying PFC mechanism to reduce network congestion.
[0102] Compared with related technologies, this application avoids timeouts through a Flowlet-based design, can adapt to the traffic characteristics of different large language models, and can perceive the network characteristics of PFC, thereby improving network load balancing efficiency and reducing out-of-order packet problems.
[0103] Preferably, for the technical solution of this application, possible alternatives include: 1. Load balancing based on other granularities: Using "packet" or "flow"-based granularity for load balancing can avoid the complexity of flow block granularity.
[0104] 2. Improved Flowlet-based solution: Although the traditional Flowlet timeout-based strategy has limitations, it is possible to achieve better adaptability to LLM traffic by optimizing the dynamic adjustment of the Flowlet timeout value, thereby alleviating the original out-of-order packet problem.
[0105] 3. Use dynamic congestion awareness algorithm: Use real-time network congestion status assessment and intelligent routing algorithms (such as machine learning-driven load balancing methods) to adjust routing paths in real time based on network status and traffic characteristics, avoiding the use of fixed-granularity solutions.
[0106] Through these alternatives, similar load balancing effects can be achieved without relying entirely on the granularity of the stream blocks.
[0107] Based on the same inventive concept, another embodiment of the present application further provides a network flow control device based on flow block granularity. Figure 9 This is a schematic diagram of a network flow control device based on flow block granularity provided by an embodiment of the present application, with reference to Figure 9 , the device comprises: A traffic type determination module 11 is configured to determine the traffic type of current traffic that a transmitter in a current model needs to send to a receiver, wherein the current traffic includes a plurality of data packets; A flow block granularity calculation module 12 is configured to calculate a flow block granularity corresponding to the traffic type of the current traffic based on the parameters provided by the current model, and send the flow block granularity to the switch; a target egress port determination module 13, configured to determine an estimated departure time of packets queued at each egress port of the switch, and select the egress port with the smallest estimated departure time as the target egress port; the estimated departure time is obtained based on the queue length and pause time of the egress port; A flow block establishing module 14, configured to establish a flow block for the current traffic according to the flow block granularity, and update the flow block table of the switch according to the established flow block; The data transmission module 15 is configured to send the data packets in the flow block to the receiving end through the target egress port according to the flow block table.
[0108] Optionally, the traffic type determination module 11 includes: a calculation method determining unit, configured to determine, based on the traffic type of the current traffic, a calculation method for the flow block granularity corresponding to the traffic type of the current traffic; The stream block granularity calculation module 12 includes: A flow block granularity calculation unit is used to calculate the flow block granularity corresponding to the flow type of the current flow according to the calculation method of the flow block granularity corresponding to the flow type of the current flow and according to the parameters provided by the current model.
[0109] Optionally, the device further comprises: a first judging unit configured to, before determining an estimated departure time of a queued packet in each egress port of the switch, judge, in the flow block table, for a data packet currently to be transmitted among a plurality of data packets included in the current traffic, whether the data packet currently to be transmitted is the first data packet in the queue pair corresponding to the transmitting end; In a case where the data packet currently to be transmitted is the first data packet in the queue pair corresponding to the transmitting end, the target egress port determination module 13 includes: a PFC pause time calculation unit, configured to query the PFC state of each egress port of the switch and calculate the PFC pause time of the egress port; a delay calculation unit, configured to calculate a sending delay of a queued packet and a processing delay of a queued packet in the egress port; a first estimated departure time calculation unit, configured to determine an estimated departure time of the egress port according to the time of being paused by the PFC, a sending delay of the queued packet, and a processing delay of the queued packet; The data transmission module 15 includes: a first data transmission unit, configured to send the first data packet to the receiving end through the target egress port; The sending delay of the queued packet is determined according to the amount of data in the queue pair and the link rate; The processing delay of the queued packets is determined according to the number of queued packets in the queue pair and the average processing delay of each queued packet; The PFC pause time is determined according to the time when the egress port resumes sending data and the current time of the routing decision.
[0110] Optionally, the device further comprises: a recording unit, configured to record a target egress port for transmitting the first data packet; A second judging unit is configured to judge, in the flow block table, whether a data packet currently to be transmitted among a plurality of data packets included in the current traffic belongs to the flow block where the first data packet is located; The second data transmission unit is configured to send the data packet that currently needs to be transmitted to the receiving end through the target egress port when the data packet that currently needs to be transmitted belongs to the flow block where the first data packet is located.
[0111] Optionally, the device further comprises: A third judgment unit is configured to judge, in the flow block table, whether a data packet currently to be transmitted among the multiple data packets included in the current traffic belongs to the last data packet in the flow block where the first data packet is located; The next target egress port determination unit is used to redetermine the next target egress port when the data packet currently to be transmitted belongs to the last data packet in the flow block where the first data packet is located. The next target egress port is used to transmit multiple data packets in the next flow block.
[0112] Optionally, the stream block granularity calculation unit includes: A first flow block granularity calculation unit is configured to calculate a flow block granularity corresponding to the flow type of the current flow according to the sequence length, micro-batch size, and hidden size of the current model when the flow type of the current flow is PP communication flow; The second flow block granularity calculation unit is used to calculate the flow block granularity corresponding to the traffic type of the current traffic according to the number of parameters of the current model, the size of data parallelism and the number of segments in the segmented pipeline ring when the traffic type of the current traffic is DP communication traffic.
[0113] Optionally, the device further comprises: a fourth determining unit, configured to determine, for a data packet currently to be transmitted among a plurality of data packets included in the current traffic, whether a size of the data packet currently to be transmitted exceeds the flow block granularity; a second estimated departure time calculation unit, configured to query a PFC state of each egress port of the switch and calculate an estimated departure time of a queued packet of each egress port when the size of the data packet currently to be transmitted exceeds the flow block granularity; a new target egress port determining unit, configured to select, based on the expected departure times of the queued packets of each egress port, an egress port with a minimum expected departure time as a new target egress port; The data transmission unit is used to establish a new flow block and send the data packets in the new flow block to the receiving end through the new target egress port according to the flow block table.
[0114] Based on the same inventive concept, another embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the network traffic control method based on flow block granularity as described in any of the above embodiments.
[0115] Based on the same inventive concept, another embodiment of the present application further provides a computer program product, including a computer program, which is executed by a processor to implement the network traffic control method based on flow block granularity as described in any of the above embodiments.
[0116] Based on the same inventive concept, another embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein when the program is executed by a processor, the network traffic control method based on flow block granularity as described in any of the above embodiments is implemented.
[0117] As for the device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0118] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0119] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0121] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0123] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0124] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0125] The above is a detailed introduction to a network traffic control method based on flow block granularity provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A network flow control method based on flow block granularity, characterized in that: The method comprises: Determine a traffic type of current traffic that a sender in a current model needs to send to a receiver, where the current traffic includes multiple data packets; Calculating a flow block granularity corresponding to the traffic type of the current traffic based on the parameters provided by the current model, and sending the flow block granularity to the switch; Determining an estimated departure time of packets queued at each egress port of the switch, and selecting an egress port with a minimum estimated departure time as a target egress port; wherein the estimated departure time is based on a queue length and a pause time of the egress port; Establishing a flow block for the current traffic according to the flow block granularity, and updating the flow block table of the switch according to the established flow block; According to the flow block table, the data packets in the flow block are sent to the receiving end through the target egress port.
2. The network flow control method based on flow block granularity according to claim 1 is characterized in that: The determining of the traffic type of the current traffic that the sending end needs to send to the receiving end in the current model includes: Determining, based on the traffic type of the current traffic, a calculation method for a flow block granularity corresponding to the traffic type of the current traffic; The calculating, based on the parameters provided by the current model, a flow block granularity corresponding to the flow type of the current flow, includes: According to the calculation method of the flow block granularity corresponding to the flow type of the current flow, the flow block granularity corresponding to the flow type of the current flow is calculated according to the parameters provided by the current model.
3. The network flow control method based on flow block granularity according to claim 1 is characterized in that: Before determining the expected departure time of the queued packet in each egress port of the switch, the method further comprises: For a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently required to be transmitted is the first data packet in the queue pair corresponding to the transmitting end; In a case where the data packet currently to be transmitted is the first data packet in the queue pair corresponding to the transmitting end, determining the estimated departure time of the queued packet in each egress port of the switch includes: For each egress port of the switch, query the PFC status of the egress port and calculate the time during which the egress port is suspended by PFC; Calculating a sending delay of a queued packet and a processing delay of a queued packet in the egress port; determining an estimated departure time of the egress port based on the time of being paused by the PFC, a sending delay of the queued packet, and a processing delay of the queued packet; The sending of the data packets in the flow block to the receiving end through the target egress port includes: sending the first data packet to the receiving end through the target egress port; The sending delay of the queued packet is determined according to the amount of data in the queue pair and the link rate; The processing delay of the queued packets is determined according to the number of queued packets in the queue pair and the average processing delay of each queued packet; The PFC pause time is determined according to the time when the egress port resumes sending data and the current time of the routing decision.
4. The network flow control method based on flow block granularity according to claim 3 is characterized in that: The method further comprises: Recording a destination egress port for transmitting the first data packet; For a data packet currently to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently to be transmitted belongs to the flow block where the first data packet is located; In a case where the data packet currently to be transmitted belongs to the flow block where the first data packet is located, the data packet currently to be transmitted is sent to the receiving end through the target egress port.
5. The network flow control method based on flow block granularity according to claim 3 is characterized in that: The method further comprises: For a data packet currently required to be transmitted among a plurality of data packets included in the current traffic, determining in the flow block table whether the data packet currently required to be transmitted belongs to the last data packet in the flow block where the first data packet is located; In the case that the data packet currently to be transmitted is the last data packet in the flow block where the first data packet is located, a next target egress port is re-determined, and the next target egress port is used to transmit multiple data packets in the next flow block.
6. The network flow control method based on flow block granularity according to claim 2 is characterized in that: The method of calculating the flow block granularity corresponding to the flow type of the current flow according to the parameters provided by the current model includes: When the traffic type of the current traffic is PP communication traffic, calculating a flow block granularity corresponding to the traffic type of the current traffic according to the sequence length, micro-batch size, and hidden size of the current model; When the traffic type of the current traffic is DP communication traffic, the flow block granularity corresponding to the traffic type of the current traffic is calculated based on the number of parameters of the current model, the size of data parallelism and the number of segments in the segmented pipeline ring.
7. The network traffic control method based on flow block granularity according to claim 6 is characterized in that: In the case where the traffic type of the current traffic is PP communication traffic, the flow block granularity The calculation formula is: ; in, is the sequence length, is the mini-batch size, is the hidden size.
8. The network traffic control method based on flow block granularity according to claim 6 is characterized in that: In the case where the traffic type of the current traffic is DP communication traffic, the flow block granularity The calculation formula is: ; in, is the number of parameters of the current model, is the size of the data parallelism, is the number of stages in the segmented pipeline loop.
9. The network traffic control method based on flow block granularity according to claim 1 is characterized in that: The method further comprises: For a data packet currently to be transmitted among a plurality of data packets included in the current traffic, determining whether a size of the data packet currently to be transmitted exceeds the flow block granularity; In a case where the size of the data packet currently to be transmitted exceeds the flow block granularity, querying the PFC state of each egress port of the switch and calculating the estimated departure time of the queued packet of each egress port; Based on the expected departure time of the queued packets of each egress port, selecting the egress port with the minimum expected departure time as the new target egress port; A new flow block is established, and according to the flow block table, the data packets in the new flow block are sent to the receiving end through the new target egress port.
10. A network flow control device based on flow block granularity, characterized in that: The device comprises: A traffic type determination module, configured to determine the traffic type of current traffic that a sender in a current model needs to send to a receiver, wherein the current traffic includes a plurality of data packets; a flow block granularity calculation module, configured to calculate a flow block granularity corresponding to the traffic type of the current traffic based on parameters provided by the current model, and send the flow block granularity to the switch; a target egress port determination module, configured to determine an estimated departure time of packets queued at each egress port of the switch, and select an egress port with a minimum estimated departure time as a target egress port; the estimated departure time is obtained based on a queue length and a pause time of the egress port; a flow block establishing module, configured to establish a flow block for the current traffic according to the flow block granularity, and update the flow block table of the switch according to the established flow block; The data transmission module is used to send the data packets in the flow block to the receiving end through the target egress port according to the flow block table.