Port-by-port flow control method suitable for BCube topology, medium and equipment

By adopting a port-by-port flow control method in the data center network of BCube topology, the queue allocation strategy and dynamic flow control mechanism are designed, and packet loss and deadlock problems caused by burst traffic are solved, and efficient traffic control and stable network performance are achieved.

CN119966899AActive Publication Date: 2025-05-09NANJING UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510209981.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-09
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

In the data center network of BCube topology, burst traffic can easily lead to packet loss and deadlock, affecting service performance, and the existing technology is difficult to effectively solve these problems.

Method used

The port-by-port flow control method is adopted to design queue allocation policies and dynamic flow control mechanisms on switch and host ports to achieve independent management and optimization of different types of traffic to avoid deadlocks.

Benefits of technology

It effectively reduces blocking problems caused by burst traffic, improves transmission efficiency, avoids deadlock, and ensures the stability and high performance of the data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966899A_ABST
    Figure CN119966899A_ABST
Patent Text Reader

Abstract

The invention provides a port-by-port flow control method suitable for BCube topology, a medium and equipment. The port-by-port flow control method comprises a port-by-port queue allocation strategy, a dynamic flow control mechanism and a queue grouping strategy. In the design of the switch port queues, an output port of each switch is provided with an independent queue corresponding to a downstream switch port, and the blocking problem caused by burst traffic is reduced. In the design of a host port queue, a special queue is allocated for a host according to the traffic characteristics, and queue head blockage generated when different types of traffic compete for resources is avoided. The dynamic flow control mechanism dynamically adjusts the sending rate of the upstream node queue according to the buffer state of the downstream switch to prevent the buffer area from overflowing; the queue type of the congestion port is monitored in real time, a pause or recovery signal is generated, different types of traffic are managed independently, and the transmission efficiency is improved. In the queue grouping strategy, the switch port queues are divided into different priorities, the propagation of congestion signals is reduced, cyclic dependence is avoided, and network stagnation is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer network and system structure, and relates to flow control and lossless transmission technology of data center network. Background Art

[0002] Modern data center networks often use high-bandwidth, low-latency RDMA technology. This technology achieves fast data transmission by skipping the operating system kernel and unloading the protocol stack. However, in burst traffic scenarios, RDMA is prone to packet loss, which affects lossless data transmission.

[0003] At present, the industry often uses two methods to achieve lossless networks. One is the priority-based flow control (PFC) algorithm, and the other is the credit mechanism flow control (CBFC). For example, the optimization scheme of RDMA in cloud storage proposed in the prior art improves the transmission efficiency by adjusting the hardware and flow control protocol. However, most of these methods are aimed at switch-centric tree topologies. They are not suitable for server-centric networks, such as BCube topologies. BCube networks are often used in data-intensive tasks due to their multi-path high bandwidth and modular design. However, in BCube networks, shallow buffer switches cannot handle burst traffic, and are prone to packet loss under burst traffic, affecting service performance. In addition, the pause and resume mechanism of PFC may cause circular dependencies in multi-hop paths, which will cause traffic deadlock and lead to traffic stagnation. Although traffic scheduling methods suitable for BCube have been proposed in the prior art, these existing methods only solve the problem of traffic distribution, but do not solve the problem of link layer flow control. Summary of the invention

[0004] The present invention aims at the packet loss problem and deadlock problem of BCube topology in the prior art and provides a port-by-port flow control method, medium and device suitable for BCube topology.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a port-by-port flow control method suitable for a BCube topology, wherein each host has k+1 ports, which are sequentially connected to k+1 switches distributed from layer 0 to layer k, and each switch has n ports; the port-by-port flow control method comprises:

[0007] Execute the queue allocation strategy of the switch port, and allocate queues to the switch according to the type of traffic flowing through the switch, where each switch egress port has an independent queue corresponding to the downstream switch port;

[0008] Execute the queue allocation strategy of the host port and allocate queues to the host according to traffic characteristics;

[0009] The switch detects the occurrence and release of congestion and sends out flow control messages; the host and the switch process the flow control messages and apply deadlock avoidance strategies to prevent deadlock.

[0010] Optionally, the queue allocation strategy of the switch port is as follows:

[0011] The traffic flowing through the switch is divided into two categories: the first category is the traffic whose next hop host is not the destination host and needs to be forwarded by the next hop host; the second category is the traffic that reaches the destination host in the next hop;

[0012] Corresponding to the two types of traffic flowing through the switch, the queues in the switch port are also divided into two categories: for the first type of traffic, it is placed in different first-type queues according to the egress ports through which the traffic passes in the next-hop switch, for a total of n-1 queues; for the second type of traffic, it is directly stored in a separate second-type queue.

[0013] Optionally, the queue allocation strategy of the host port is as follows:

[0014] The traffic sent from the host is classified according to the number of hops, and each category is configured with n-1 queues, corresponding to the n-1 possible output port selections for the next hop; another queue is configured as a relay queue to store the traffic that needs to be forwarded by this host, and a high-priority queue is also configured.

[0015] Optionally, the detection of congestion occurrence on the switch is specifically as follows:

[0016] For the first and second queues in the switch port, set the congestion threshold X respectively. off and

[0017] When a data packet arrives at the outbound port on the switch, port congestion detection is performed: when a data packet enters the first queue, the length of the first n-1 queues is detected to see if it exceeds the congestion threshold X. off If it exceeds, a pause message carrying the congestion queue number and congestion port number is sent to the upstream to suspend the queue of the upstream port corresponding to the congested port; when the data packet enters the second queue, it detects whether the length of the second queue exceeds the congestion threshold If it exceeds the limit, a pause message carrying the congestion queue number and the congestion port number is sent to suspend the queue of the upstream port corresponding to the congested port.

[0018] Optionally, the detection of congestion relief on the switch is specifically as follows:

[0019] For the first and second queues in the switch port, set the congestion relief threshold X respectively. onand

[0020] When a queue at the switch port is about to send a data packet, the congestion status of the switch is re-judged: when a data packet is sent from the first queue, the sum of the lengths of the first n-1 queues is determined to be less than X. on to determine whether the congestion is relieved; when the data packet is sent from the second-class queue, the length of the second-class queue is less than To determine whether the congestion is relieved.

[0021] Optionally, the host processes the flow control message as follows:

[0022] The host parses the congestion / congestion relief port number, congestion / congestion relief queue number and control message type of the downstream switch in the flow control message;

[0023] When the host receives the flow control message generated by the first type of queue on the switch, it sets the status of the local corresponding queue according to the control message type;

[0024] When the host receives the flow control message generated by the second type of queue on the switch, it sets the state of the local corresponding queue according to the control message type, and sets the state of the corresponding queue of the upstream switch.

[0025] Optionally, the switch processes the flow control message as follows:

[0026] The switch parses the congestion / congestion relief port information of the downstream switch and the control message type in the flow control message, and operates the queue corresponding to the congestion / congestion relief port in the switch according to the control message type.

[0027] Optionally, the deadlock avoidance strategy is as follows:

[0028]

[0029] In the formula, the first and second queues of the switch are recorded as class A queue and class B queue respectively; the traffic sent by the host is divided into four-hop traffic, two-hop traffic and traffic to be forwarded, which are put into class A queue, class B queue and class C queue respectively; X off and Respectively represent the congestion thresholds of class A queues and class B queues; Switch i Q A and Switch i Q B They represent the class A queue and class B queue of the outbound port of the i-th switch respectively; Host i-1 Q a and Host i-1 Q bThey represent the class a queue and class b queue of the i-1th hop host respectively; PAUSE means pause operation.

[0030] In a second aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the port-by-port flow control method suitable for BCube topology described in the first aspect.

[0031] In a third aspect, the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the port-by-port flow control method suitable for the BCube topology described in the first aspect is implemented.

[0032] The beneficial effects of the present invention are:

[0033] (1) The present invention proposes a port-by-port queue allocation strategy. In the design of the switch port queue, each switch output port has an independent queue corresponding to the downstream switch port, reducing the blocking problem caused by burst traffic. In the design of the host port queue, a special queue is allocated to the host according to the traffic characteristics, avoiding head-of-line blocking that occurs when different types of traffic compete for resources.

[0034] (2) The present invention proposes a dynamic flow control mechanism, which dynamically adjusts the sending rate of the upstream node queue according to the buffer status of the downstream switch to prevent buffer overflow; monitors the queue type of the congested port in real time, generates a pause or resume signal, and manages different types of traffic separately, thereby improving transmission efficiency.

[0035] (3) The present invention proposes a queue grouping strategy, where the switch port queues are divided into different priorities, which reduces the propagation of congestion signals, avoids circular dependencies, and prevents network stagnation. By reasonably setting the number and type of queues, the deadlock propagation can be effectively limited and network stagnation can be prevented. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is the Bcube (4, 1) network topology architecture diagram.

[0037] Figure 2 This is a schematic diagram of port queue allocation on a switch in the Bcube (4, 1) network topology architecture.

[0038] Figure 3 This is a schematic diagram of port queue allocation on the host in the Bcube (4, 1) network topology architecture.

[0039] Figure 4 This is a diagram of traffic control deadlock avoidance in the Bcube scenario.

[0040] Figure 5This is the Bcube (8, 1) network topology architecture diagram.

[0041] Figure 6 It is the cumulative distribution function (CDF) graph of the traffic size of WebSearch, FB_Hadoop, and Storage. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0043] In one embodiment, the present invention proposes a port-by-port flow control method suitable for BCube topology.

[0044] The per-port queue allocation strategy is an important part of the design. In the switch port queue design, each switch egress port has an independent queue corresponding to the downstream switch port. This design can reduce the blocking problem caused by burst traffic. In the host port queue design, a dedicated queue is assigned to the host based on the traffic characteristics, such as two-hop traffic or four-hop traffic. This can avoid head-of-line blocking that occurs when different types of traffic compete for resources.

[0045] The dynamic flow control mechanism dynamically adjusts the sending rate of the upstream node queue according to the buffer status of the downstream switch to prevent buffer overflow. The mechanism also monitors the queue type of the congested port in real time and generates a pause (PAUSE) or resume (RESUME) signal. In this way, different types of traffic can be managed separately, improving transmission efficiency.

[0046] In order to avoid deadlock problems in the network, a queue grouping strategy is adopted in the design. The switch port queues are divided into different priorities. This can reduce the propagation of congestion signals and avoid circular dependencies. In the BCube (4, 1) topology, by reasonably setting the number and type of queues, the deadlock propagation is limited to within three hops. This method can effectively prevent network stagnation.

[0047] Before explaining the queue allocation scheme of the BCube port-by-port traffic method, we first briefly introduce the construction method of the BCube network. BCube (n, k) (k ≥ 1) is a recursively defined structure, which consists of n BCube (n, k-1) and n k Each server in BCube(n, k) has k+1 ports, which are connected to k+1 switches from layer 0 to layer k in sequence. According to the above definition, the number of servers in BCube(n, k) is n. k+1 , the number of switches is (k+1)·n k These switches are distributed in layers 0 to k, and there are n switches in layer k-1.k The topology architecture of BCube (4, 1) is as follows: Figure 1 shown.

[0048] Similar to the PFC flow control algorithm commonly used in lossless RDMA technology, this embodiment also prevents the problem of buffer overflow in the downstream switch by controlling the sending of a certain queue. What is different from the PFC algorithm based on ingress port queue detection is that the flow control of this embodiment detects congestion based on the egress port queue. The method of controlling congested traffic based on egress port detection of network congestion can alleviate the head-of-line blocking problem of the ingress port flow control algorithm. The queue allocation of port-by-port flow control is mainly designed based on the fact that each switch in the BCube network topology has a fixed number of ports. The main idea is to maintain several queues for the downstream node ports in the egress port of the upstream node. These queues correspond one-to-one to the ports of the downstream nodes. When a downstream port is congested, the queue corresponding to the upstream switch stops sending to the downstream. Specifically, the queue allocation scheme of the BCube port-by-port flow control method designed in this embodiment is divided into two types: queue allocation of switch ports and queue allocation of the host end. Next, we will talk about Figure 1 Taking BCube(4,1) as an example, these two allocation methods are explained in detail.

[0049] (1) Queue allocation of switch ports

[0050] Since the adjacent nodes of the switch in the BCube network are all hosts, the traffic flowing through the switch can be divided into two categories. The first category is the traffic that needs to be forwarded by the next-hop host when the next-hop host is not the destination host, and the second category is the traffic that reaches the destination host in the next hop. Corresponding to the above two types of traffic, the queues in the switch port are also divided into two types of queues. The first type of queue stores traffic that needs to be forwarded by the next-hop host for further transmission. Since the next-hop to which this part of traffic reaches after forwarding by the next-hop host is a switch, it is placed in different queues according to the egress port that the traffic passes through on the next-hop switch. The other type stores traffic that reaches the destination in the next hop. There is only one queue of this type, and all traffic that reaches the destination in the next hop is placed in this queue.

[0051] by Figure 1 Take the topology in Figure 1 for example. Assume that there are four flows, namely F1, F2, F3 and F4. These four flows come from host No. 11 and connect to the egress port of host No. 15 through switch No. 23. The queues to which the packets of the four flows are allocated on the egress port of switch No. 23 are as follows: Figure 2As shown. After the data packets of the three flows F2, F3, and F4 arrive at host No. 15, since the destination host is not host No. 15, they need to be forwarded by host No. 15 to reach the next hop switch No. 19. Assume that all ports and queues are numbered starting from 0. Because F2 will be sent out from port No. 0 of switch No. 19, F2 enters the Q0 queue of switch No. 23; because F3 will be sent out from port No. 19, F3 enters the Q1 queue of switch No. 23; similarly, F4 is sent out from port No. 2 on switch No. 19, so the data packet of F4 is placed in the Q2 queue on switch No. 23. Different from the above three flows, after F1 is sent out from machine No. 23, the next hop is its destination host, so it will be placed in a separate queue Q3. Figure 2 A highest priority queue HQ is also drawn in the figure, which can be used to store some control messages in the network (such as ACK 1 or CNP 2 messages, etc.).

[0052] From the above examples, we can see that the queue allocation strategy of the BCube port-by-port flow control method on the switch port is to put the data packet into the queue of the corresponding next-hop switch port in the switch port according to which egress port the data packet passes through in the next-hop switch (because the next hop connected to the BCube switch is the host). If the data packet reaches the destination in the next hop, it will be placed in a separate queue. According to the number of downstream switch ports, a corresponding number of queues are allocated in the egress port of the upstream node. This queue allocation method makes the BCube port-by-port flow control use very little switch hardware resources. For the BCube (n, k) topology architecture, the switches therein have n ports, so only n-1 queues are needed in each switch port to store data packets sent to the n-1 possible egress ports of the next-hop switch, and only 1 queue is needed to store data packets that reach the destination host in the next hop. In addition, if necessary, 1 highest priority queue can be added. In summary, in the design of the port-by-port flow control method, each port of the switch in the BCube (n, k) topology only needs n+1 queues.

[0053] (2) Queue allocation of host ports

[0054] There are three types of traffic on the host in the BCube topology. The first is the four-hop traffic sent by this host, the second is the two-hop traffic sent by this host, and the third is the traffic sent by other hosts through this host and needs to be forwarded by this host. For the above three types of traffic, the port-by-port traffic control method gives different queue allocation schemes. In the BCube (n, k) topology, each switch has n ports. The first type of four-hop traffic data packets sent by this host will enter queues 0 to n-2 according to the output port number of the next-hop switch it passes through (there are n-1 possibilities). The second type of two-hop traffic data packets sent by this host will enter queues n-1 to 2n-3 according to the output port number of the next-hop switch it passes through (there are n-1 possibilities). The third type of data packets that need to be forwarded by this host will enter the "relay queue" numbered 2n-2.

[0055] by Figure 1 Take the topology in as an example, assuming that there are four-hop flows F1, F2, and F3 sent from host 15, which are forwarded by hosts 12, 13, and 14 and finally sent to their respective destination hosts 0, 1, and 2. There is also a second type of two-hop flow F4, F5, and F6 sent from host 15, which respectively pass through ports 0, 1, and 2 of switch 19 and finally reach their respective destination hosts 12, 13, and 14. The third type of flow F7 is forwarded from host 11 via host 15 and finally reaches destination host 12.

[0056] Figure 3 Taking the above traffic as an example, the queue allocation on the port of switch No. 19 connecting host No. 15 is given. First, for the first type of four-hop traffic sent from host No. 15, since F1, F2, and F3 pass through the 0, 1, and 2 output ports of the next-hop switch No. 19 respectively, F1 is placed in Q0, F2 is placed in Q1, and F3 is placed in Q2 on switch No. 15. The second type of traffic F4, F5, and F6 on switch No. 15 is two-hop traffic. Since they will be sent out from output ports 0, 1, and 2 of switch No. 19 and then reach their respective destination hosts, they are placed in Q3, Q4, and Q5 on switch No. 15 respectively. Data packets that need to be forwarded through host No. 15 will be placed in RQ (representing Relay Queue). Similarly, Figure 3 HQ is also given for storing high priority data packets.

[0057] As can be seen from the above examples, the queue allocation on the host is slightly more complicated than that on the switch. This more refined classification will be more convenient when the port-by-port flow control method is used to control the flow, making the control granularity finer and the control effect better. Even so, in the BCube (n, k) topology, the number of port queues on the host is still linearly proportional to the number of switch ports. Specifically, in the BCube (n, k) topology architecture, the traffic sent from the host can be divided into k+1 categories according to the number of hops, each of which requires n-1 queues (corresponding to the possible n-1 output port selections for the next hop) for storage, and another queue is required as a relay queue to store data packets that need to be forwarded locally, and a high priority queue is also required. In summary, in the BCube (n, k) topology, the host port requires a total of (k+1)×(n-1)+2 queues (where k is generally less than or equal to 4). Current commercial switches can generally support dozens of queues, so the queue allocation scheme designed in this embodiment can still support large BCube networks and has good scalability.

[0058] (3) Congestion Occurrence and Congestion Relief

[0059] The previous article gave the allocation method of two types of queues on the switch, so there are two sets of congestion detection and congestion relief thresholds in the port-by-port flow control method of BCube. Set the congestion threshold X for the queues (n-1, numbered 0 to n-2) in the switch egress port that store the traffic that passes through different switch egress ports. off and congestion relief threshold X on Set the congestion threshold for the second type of queue, which is the queue that stores traffic that reaches the destination host in the next hop (1 queue, numbered n-1). and decongestion threshold

[0060] When a data packet arrives at the output port of the switch, a port congestion detection is performed. The detailed logic is shown in pseudo code 1 in Table 1. Lines 4 to 18 in pseudo code 1 give the processing logic when a data packet enters the first type of queue in BCube (n, k). In this embodiment, the length of the n-1 queues and whether it exceeds the congestion threshold are used as the judgment criteria for whether the port is congested. When a data packet is received at the output port of the switch and enters the first type of queue, the length of the first type n-1 queues and whether it exceeds the congestion threshold X are detected. off, if it exceeds, a pause message (carrying the queue number and the number of the congested port) is sent to the upstream connected to other ports to suspend the queue of the upstream node port corresponding to the congested port. Similar to the operation of the first type of queue when receiving a data packet, the logic of the second type of queue that stores data packets that reach the destination in the next hop to detect whether congestion occurs when receiving a data packet is given in lines 19 to 27 of pseudo code 1. In addition, the flow control method proposed in this embodiment does not control the transmission of the highest priority. The highest priority data packets are usually small in number. If the data packet belongs to the highest priority, the highest priority queue will not be detected whether it is congested. This part of the logic is given in line 29 of pseudo code 1.

[0061] Table 1 Pseudocode 1

[0062]

[0063] When a packet is to be sent from a queue at the switch output port, the congestion status of the switch needs to be re-judged to detect whether the length of the current queue meets the congestion release condition. In the pseudo code 2 shown in Table 2, lines 3 to 15 show the sum of the lengths of n-1 queues less than the given threshold X when a packet is sent from the first queue. on To determine whether congestion is relieved. Similar to the first type of queue for congestion relief and whether to send congestion relief messages to the upstream, lines 16 to 24 in pseudo code 2 give the processing logic for the second type of queue on the switch when congestion is relieved. Similarly, the flow control logic is not processed for the data packets sent in the highest priority. The variable upstreamIsPaused is used to determine whether a control message needs to be sent upstream, and upstreamIxPaused can ensure that the same congestion or congestion relief information will not be reported to the upstream multiple times.

[0064] Table 2 Pseudocode 2

[0065]

[0066] (4) Host processing logic

[0067] Since there are two types of queues on the switch that can generate flow control messages, there are corresponding operations on the upstream host for receiving these two types of flow control messages. First, the host will parse the number of the downstream congested (or decongested) port, the number of the queue that causes congestion (or decongestion), and the type of control message (Pause or Resume) according to the content carried in the flow control message. As shown in lines 4 to 11 in pseudo code 3 shown in Table 3, when the host receives the flow control message generated by the first type of queue on the switch, the host will set the corresponding state of the queue numbered congestedQIdx in its first n-1 queues according to the type of control message (Pause or Resume). This is because in the BCube network topology, the traffic is only divided into four-hop traffic and two-hop traffic. When the host receives the control message of the first type of queue from the downstream switch, it means that the four-hop traffic sent by the local machine may cause the downstream port numbered congestedPortIdx to continue to be congested, so the local queue corresponding to the downstream congested port should be suspended.

[0068] Table 3 Pseudocode 3

[0069]

[0070] If the local machine receives a flow control message from the second-class queue of the downstream switch, it means that the queue sent by the downstream switch to the next-hop host is congested. This traffic may be two-hop traffic from this host, or it may be four-hop traffic from further upstream. Therefore, the congestedPortIdx queue (number range is (n-1) to (2n-3)) in the second-class queue of this host should be suspended accordingly, because the data packets in this queue will increase the congestion of the downstream congested port. Therefore, as shown in line 14, an offset of size n-1 needs to be added to the queue number indicated by congestedPortIdx. The above logic corresponds to lines 12 to 20 in pseudo code 3. In addition, lines 21 to 23 in pseudo code 3 show that this host needs to pass the flow control message to the switch further upstream to control the impact of the four-hop traffic sent upstream on the congested port.

[0071] By observing the host's reaction to the congestion of the two types of queues on the downstream switch, it can be found that the congestion caused by the traffic in the first type of queue on the switch (the traffic that needs to be forwarded by the next-hop host) will only suspend the queue corresponding to the upstream host of the switch; while the congestion caused by the traffic in the second type of queue on the switch (the queue storing the traffic of the next-hop host being the destination host) will not only suspend the queue corresponding to the upstream host, but also suspend the corresponding queue of the upstream switch. Since new congestion will not be detected on the host, in the flow control method designed in the embodiment, congestion will propagate at most three hops in the network.

[0072] (5) Switch processing logic

[0073] The processing of the control message received by the switch is relatively simple, which is given in the pseudo code 4 shown in Table 4. When the switch receives the flow control message, it first parses the port information of the downstream switch congested (or decongested) and the type of flow control message (pause or resume) carried in the flow control message. Then, according to the type of flow control message, the queue corresponding to the congested port in the switch is paused or resumed.

[0074] Table 4 Pseudocode 4

[0075]

[0076] (6) Deadlock avoidance algorithm

[0077] Traditional flow control algorithms, such as PFC, can cause deadlock in topology scenarios where loops exist. The control logic of the flow control algorithm is usually that the downstream node detects congestion or relieves congestion, and sends a flow control message to the upstream node to report the downstream congestion or congestion relief. After receiving the flow control message, the upstream node pauses or resumes the transmission of the corresponding port or queue according to the type of message (Pause or Resume). Under this logic, since there are loops in the BCube topology, there may be multiple flows entering the same queue, causing congestion, and then the congestion propagates hop by hop in the network, eventually forming a paused loop in the network. Each port in the paused loop is waiting for the next hop to relieve congestion. In this scenario, each node in the loop cannot send data packets, resulting in a deadlock.

[0078] The queue allocation strategy and flow control logic of this embodiment successfully destroy the loop of flow control and effectively avoid the aforementioned deadlock problem. Without considering the highest priority queue, the queue allocation strategy proposed in this embodiment can be summarized as follows: on the switch, the traffic is divided into two types: traffic that needs to be forwarded by the next-hop host and traffic whose next-hop host is the destination, and they are placed in two different queues, respectively recorded as Class B queues and Class C queues; the traffic type on the host is slightly more complicated, and is divided into three categories, four-hop traffic sent by this host, two-hop traffic sent by this host, and traffic from other hosts that needs to be forwarded by this host. The above three types of traffic are placed in three queues recorded as b, c, and d.

[0079]

[0080] The congestion propagation path after congestion is detected at the switch egress port is briefly presented in Figure 4In the formula (1) and formula (2), the situation where the upstream queue is suspended after congestion is detected on the switch. As shown in formula (1), one situation where the egress port of the i-th hop switch is congested is that the Class C queue on it is congested, and the congestion will suspend the Class C queue of the i-1th hop host and the Class B queue of the i-2th hop switch. Assuming that the Class B queue of the i-2th hop switch is also congested due to being suspended, according to formula (2), the congestion will continue to suspend the Class B queue of the i-3th hop host; so when the Class C queue of the i-th hop switch is congested, the congestion will propagate at most three hops in the network to the i-3th hop host and then be suspended. Another situation is described by formula (2), where the Class B queue of the i-th hop switch egress port is congested. In this case, the congestion will only propagate to the previous hop, that is, the i-1th hop host, and will not spread further upstream. Combined with the above analysis, it can be seen that the queue allocation strategy and congestion control logic designed in this embodiment can effectively prevent the continuous propagation of congestion, thereby preventing deadlock.

[0081] Next, the performance of the method proposed in this embodiment will be evaluated through simulation experiments. In the following experiments and analysis, the BCube port-by-port flow control method is called PPFC (Per-Port Flow Control).

[0082] The experiment selects BCube (4, 1) and BCube (8, 1) as the topology of the simulation experiment to evaluate the performance of PPFC. Figure 1 As shown in Figure 1, it consists of 16 hosts and 8 switches. In the first dimension, a 4-port switch connects four hosts to form a Pod, and in the second dimension, four 4-port switches connect four Pods. The four hosts connected by the switches in the second dimension are located in the same position of the four Pods. Similarly, BCube(8,1) is as follows Figure 5 As shown in the figure, it consists of 64 hosts and 16 switches. In the first dimension, an 8-port switch connects eight hosts to form a Pod. In the second dimension, eight 8-port switches connect eight Pods, where each switch in the second dimension connects the hosts located in the same position in the eight Pods. The link bandwidth in the above topology is set to 100Gbps, the link delay is set to 1us, and the switch buffer size is set to 5MB.

[0083] The experiment selected three common traffic scenarios in data centers: WebSearch, FB_Hadoop, and Storage. Figure 6The cumulative distribution function (CDF) graph of the three traffic sizes is given in the figure. It can be seen that the short flows (traffic less than 1KB) in FB_Hadoop are significantly more than those in WebSearch and Storage. There are more long flows at the MB level in the WebSearch traffic model. In the experiment, these three traffic models were used to generate traffic with a Poisson arrival load of 50%. In addition, incast traffic with a load of 30% was added to increase the congestion level in the network and more realistically simulate the traffic in the data center network. The scale of incast traffic is 8 hosts sending traffic to 1 host in the BCube (4, 1) network, and 32 hosts sending traffic to 1 host in the BCube (8, 1) network. The size of these flows is fixed at 1MB.

[0084] Compared with the traditional IRN algorithm, PPFC can reduce the average completion time of traffic by 16.92% to 41.65%. At the same time, the throughput is increased by 1 to 1.32 times. Compared with the GBN algorithm, PPFC performs better, with an average completion time reduction of 91.57% to 96.54% and a throughput increase of 1.74 to 20.62 times. These improvements have greatly improved the overall performance of the network. In response to the deadlock problem in the BCube network, PPFC designed a sophisticated queue allocation strategy and dynamic flow control mechanism to successfully avoid deadlock. This not only improves network stability, but also ensures the reliable operation of complex network structures.

[0085] In another embodiment, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the port-by-port flow control method suitable for BCube topology of the aforementioned embodiment.

[0086] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the port-by-port flow control method suitable for the BCube topology of the aforementioned embodiment is implemented.

[0087] In the embodiments disclosed in the present application, the computer storage medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. The computer storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of computer storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0088] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0089] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A port-by-port flow control method suitable for BCube topology, in which each host has k+1 ports, which are sequentially connected to k+1 switches distributed from layer 0 to layer k, and each switch has n ports; characterized in that: The port-by-port flow control method comprises: Execute the queue allocation strategy of the switch port, and allocate queues to the switch according to the type of traffic flowing through the switch, where each switch egress port has an independent queue corresponding to the downstream switch port; Execute the queue allocation strategy of the host port and allocate queues to the host according to traffic characteristics; The switch detects the occurrence and release of congestion and sends out flow control messages; the host and the switch process the flow control messages and apply deadlock avoidance strategies to prevent deadlock.

2. The port-by-port flow control method suitable for BCube topology according to claim 1, characterized in that: The queue allocation strategy of the switch port is as follows: The traffic flowing through the switch is divided into two categories: the first category is the traffic whose next hop host is not the destination host and needs to be forwarded by the next hop host; the second category is the traffic that reaches the destination host in the next hop; Corresponding to the two types of traffic flowing through the switch, the queues in the switch port are also divided into two categories: for the first type of traffic, it is placed in different first-type queues according to the egress ports through which the traffic passes in the next-hop switch, for a total of n-1 queues; for the second type of traffic, it is directly stored in a separate second-type queue.

3. The port-by-port flow control method suitable for BCube topology according to claim 2, characterized in that: The queue allocation strategy of the host port is as follows: The traffic sent from the host is classified according to the number of hops, and each category is configured with n-1 queues, corresponding to the n-1 possible output port selections for the next hop; another queue is configured as a relay queue to store the traffic that needs to be forwarded by this host, and a high-priority queue is also configured.

4. The port-by-port flow control method suitable for BCube topology as claimed in claim 3, characterized in that: The detection of congestion occurrence on the switch is as follows: For the first and second queues in the switch port, set the congestion threshold X respectively. off and When a data packet arrives at the outbound port on the switch, port congestion detection is performed: when a data packet enters the first queue, the length of the first n-1 queues is detected to see if it exceeds the congestion threshold X. off If it exceeds, a pause message carrying the congestion queue number and congestion port number is sent to the upstream to suspend the queue of the upstream port corresponding to the congested port; when the data packet enters the second queue, it detects whether the length of the second queue exceeds the congestion threshold If it exceeds the limit, a pause message carrying the congestion queue number and the congestion port number is sent to suspend the queue of the upstream port corresponding to the congested port.

5. The port-by-port flow control method suitable for BCube topology as claimed in claim 3, characterized in that: The detection of congestion relief on the switch is as follows: For the first and second queues in the switch port, set the congestion relief threshold X respectively. on and When a queue at the switch port is about to send a data packet, the congestion status of the switch is re-judged: when a data packet is sent from the first queue, the sum of the lengths of the first n-1 queues is determined to be less than X. on to determine whether the congestion is relieved; when the data packet is sent from the second-class queue, the length of the second-class queue is less than To determine whether the congestion is relieved.

6. The port-by-port flow control method suitable for BCube topology as claimed in claim 3, characterized in that: The host processes the flow control message as follows: The host parses the congestion / congestion relief port number, congestion / congestion relief queue number and control message type of the downstream switch in the flow control message; When the host receives the flow control message generated by the first type of queue on the switch, it sets the status of the local corresponding queue according to the control message type; When the host receives the flow control message generated by the second type of queue on the switch, it sets the state of the local corresponding queue according to the control message type, and sets the state of the corresponding queue of the upstream switch.

7. The port-by-port flow control method suitable for BCube topology as claimed in claim 3, characterized in that: The switch processes the flow control message as follows: The switch parses the congestion / congestion relief port information of the downstream switch and the control message type in the flow control message, and operates the queue corresponding to the congestion / congestion relief port in the switch according to the control message type.

8. The port-by-port flow control method suitable for BCube topology as claimed in claim 3, characterized in that: The deadlock avoidance strategy is as follows: In the formula, the first and second queues of the switch are recorded as class A queue and class B queue respectively; the traffic sent by the host is divided into four-hop traffic, two-hop traffic and traffic to be forwarded, which are put into class A queue, class B queue and class C queue respectively; X off and Respectively represent the congestion thresholds of class A queues and class B queues; Switch i Q A and Switch i Q B They represent the class A queue and class B queue of the outbound port of the i-th switch respectively; Host i-1 Q a and Host i-1 Q b They represent the class a queue and class b queue of the i-1th hop host respectively; PAUSE means pause operation.

9. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables the computer to execute the port-by-port flow control method suitable for BCube topology as described in any one of claims 1-8.

10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the port-by-port flow control method suitable for the BCube topology as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Flow isolation method for avoiding head-of-queue congestion and congestion diffusion in lossless network

    CN115134302A

  • Dynamic clearance cache management method and device for PFC (Power Factor Correction) switch

    CN116896534A

  • Fine-grained flow control method for lossless network

    CN117395207A

  • Method and device for displaying congestion notification mark and electronic equipment

    CN119052174A

  • Method and apparatus for inhibiting generation of congestion queue

    WO2023226603A1