A data center RDMA incast solution and system based on ToR switches

By perceiving the Incast situation in the ToR switch and cacheing data flow, combining the receiver rate calculation and token bucket mechanism control, the Incast problem in the RDMA network is solved, achieving fast and accurate traffic management and stable data transmission.

CN115714750BActive Publication Date: 2025-08-19BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211200973.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-08-19
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

Existing RDMA networks are prone to Incast problems at high link rates, resulting in catastrophic network crashes. Existing solutions are difficult to effectively deal with Incast problems at high link rates in RDMA networks.

Method used

The ToR switch-based method is adopted, and the Incast situation is sensed by the sending switch and cached data flow. The receiving switch calculates the reception rate and the sending speed, and uses the token bucket mechanism to control data transmission, avoid data influx, improve processing efficiency and reduce layout difficulty.

Benefits of technology

It realizes fast and accurate identification and isolation of Incast traffic, reduces switch buffer accumulation, reduces Incast's impact on non-Incast streams, and improves the processing efficiency and stability of data center networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115714750B_ABST
    Figure CN115714750B_ABST
Patent Text Reader

Abstract

The present invention provides a data center RDMA incast solution and system based on a ToR switch. The method comprises the following steps: sensing an incast situation based on a sending-end switch, and caching a data stream if the incast situation occurs; obtaining a network card receiving rate of a user receiving end based on a receiving-end switch, calculating a data volume that can be received by the user receiving end per unit time based on the network card receiving rate, calculating a link utilization rate based on the data volume that can be received by the user receiving end per unit time, calculating a data volume that can be received by the receiving-end switch per unit time based on the link utilization, calculating a receiving rate of the receiving-end switch based on the data volume that can be received by the receiving-end switch per unit time, calculating a sending rate of each sending end based on the number of sending ends connected to the sending-end switch and the receiving rate of the receiving-end switch, and wherein the sending-end switch where the incast situation occurs sends data to the receiving-end switch based on the sending rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transmission, and in particular to a data center RDMA Incast solution and system based on a ToR switch. Background Art

[0002] Cloud services, with their powerful computing resources and unlimited storage capacity, have become a crucial necessities for the development of today's information technology companies. Cloud service providers such as Amazon, Alibaba, and Google use data centers (DCs) to provide services for applications such as large-scale online learning, distributed storage, and web search. With the continuous development of upper-layer application technologies, higher performance requirements are being placed on data center infrastructure. Data center link rates have increased from 1 / 10 Gbps to 100 Gbps today. Resource decoupling requires controlling intra-network latency to 3-5 µs to ensure good program performance, while machine learning requires intra-data center tail latency of less than 50 µs. The high link rate and low latency requirements in data centers pose significant challenges to transmission control design.

[0003] The traditional TCP network protocol stack requires the intervention of the operating system kernel to copy data. At high link rates, TCP transmission will cause high CPU utilization, resulting in transmission delays of typically hundreds of microseconds. Furthermore, the maximum bandwidth supported by operating system kernel threads is only tens of Gbps. This transmission performance is no longer sufficient to meet the needs of today's data centers. Therefore, remote direct memory access (RDMA) technology has been applied to data center networks. RDMA directly reads and writes memory without the intervention of the operating systems of either communicating party. By bypassing the kernel, RDMA achieves high link rates, low CPU utilization, and microsecond-level intra-network transmission latency. The RDMA technology used in data centers is RoCE (RDMA over converged Ethernet). Its latest version, RoCEv2, uses existing traditional networks for its data link and network layers. A variety of solutions for RDMA transmission control have been proposed in academia.

[0004] However, existing RDMA networks are prone to the incast problem. In the case of shallow caches and high link rates in data center networks, the incast problem occurs when multiple servers in the network simultaneously request data transmissions from a requesting direction. This causes a large influx of data packets within the data center network in a short period of time, leading to catastrophic network failures.

[0005] Existing solutions to the incast problem usually require arranging windows between switches in adjacent layers of the RDMA network, using windows to restrict transmission layer by layer, which makes the arrangement difficult. Summary of the Invention

[0006] In view of this, an embodiment of the present invention provides a data center RDMA incast solution based on a ToR switch to eliminate or improve one or more defects in the prior art.

[0007] One aspect of the present invention provides a data center RDMA incast solution based on a ToR switch. The RDMA network includes a sending end switch connected to a sending end and a receiving end switch connected to a user receiving end. Both the sending end switch and the receiving end switch are ToR switches. The method includes the following steps:

[0008] The sending switch senses the incast situation and caches the data flow in the sending switch if an incast situation occurs.

[0009] The network card receiving rate of the user receiving end is obtained based on the receiving end switch, the amount of data that the user receiving end can receive per unit time is calculated based on the network card receiving rate, the link utilization of the receiving end switch is calculated based on the amount of data that the user receiving end can receive per unit time, the amount of data that the receiving end switch can receive per unit time is calculated based on the link utilization, the receiving rate of the receiving end switch is calculated based on the amount of data that the receiving end switch can receive per unit time, the sending rate of each sending end is calculated based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, and the sending end switch, where an incast occurs, sends data to the receiving end switch based on the sending rate.

[0010] Using the above solution, this solution only needs to deploy programs on the sending and receiving switches of the RDMA network. The sending switch uses the sending switch to detect incast situations. When an incast situation occurs, the data flow is cached to prevent a large amount of data flow from entering the RDMA network and causing data congestion. On the other hand, the receiving switch calculates the maximum receiving rate of the receiving switch based on the user receiving end situation and allocates a sending rate to each sending switch, quickly and efficiently outputting the data cached by the sending switch. This solution does not require deployment between the switches between the sending and receiving switches, which not only improves the processing efficiency of incast situations but also reduces the difficulty of deployment.

[0011] In some embodiments of the present invention, the step of sensing the incast situation based on the transmitting-end switch includes:

[0012] The sending end switch counts the amount of data sent to the same user receiving end through the sending end switch within a preset time length;

[0013] When the amount of data sent to the same user receiving end through the sending end switch within a preset time length is greater than a preset incast threshold, it is determined that an incast situation occurs.

[0014] In some embodiments of the present invention, the step of sensing an incast condition based on the transmitting-end switch further includes: upon determining that an incast condition has occurred, the transmitting-end switch continuing to transmit a first preset amount of data to the user receiving end; and in the step of caching the data stream in the transmitting-end switch if an incast condition has occurred, after the transmitting-end switch transmits the first preset amount of data to the user receiving end, caching the data stream in the transmitting-end switch.

[0015] In some embodiments of the present invention, if an incast occurs, the step of caching the data flow in the transmitting switch further includes constructing a queue for the data flow cached in the transmitting switch based on the input time of the data flow, and in the step of the transmitting switch sending data to the receiving switch based on the sending rate when the incast occurs, the transmitting switch sends the data to the receiving switch based on the queue order.

[0016] In some embodiments of the present invention, in the step of calculating the amount of data that can be received by the user receiving terminal per unit time based on the network card receiving rate, the amount of data that can be received by the user receiving terminal per unit time is calculated based on the following formula:

[0017] ;

[0018] in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the amount of data that the user receiving end can receive per unit time;

[0019] In the step of calculating the link utilization of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, the link utilization of the receiving-end switch is calculated based on the following formula:

[0020]

[0021] in, Indicates the amount of data that the user receiving end can receive per unit time. Indicates the amount of data buffered by the port connected to the current receiving switch and the user receiving end. Indicates the link utilization of the receiving switch.

[0022] In some embodiments of the present invention, in the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the link utilization, the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time is obtained, and the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization are used to calculate the amount of data that the receiving-end switch can receive per unit time in the current unit time.

[0023] In some embodiments of the present invention, in the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization, the amount of data that the receiving-end switch can receive per unit time is calculated based on the following formula:

[0024]

[0025] in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the link utilization of the receiving switch, Indicates the amount of data that the receiving switch can receive per unit time, calculated in the previous unit time. Indicates the amount of data that the receiving switch can receive in the current unit time. are all constants.

[0026] In some embodiments of the present invention, in the step of calculating the receiving rate of the receiving switch based on the amount of data that the receiving switch can receive per unit time, the receiving rate of the receiving switch is calculated based on the following formula:

[0027]

[0028] in, Indicates the receiving rate of the receiving switch. Indicates the amount of data that the receiving switch can receive in the current unit time. Indicates the length of time per unit time.

[0029] In some embodiments of the present invention, in the step of calculating the sending rate of each sending end based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, the sending rate of each sending end switch is calculated based on the following formula:

[0030]

[0031] in, Indicates the currently calculated sending rate of the sending switch. Indicates the receiving rate of the receiving switch. Indicates the number of senders connected to the sender switch currently being calculated, Indicates the total number of senders connected to the sender switch where incast occurs.

[0032] Another aspect of the present invention provides a data center RDMA Incast solution system based on a ToR switch, the system includes a sending end switch management module and a receiving end switch management module.

[0033] The sending end switch management module is used to sense the incast situation based on the sending end switch, and if the incast situation occurs, cache the data flow in the sending end switch;

[0034] The receiving-end switch management module is configured to obtain a network card receiving rate of a user receiving end based on the receiving-end switch, calculate an amount of data that can be received by the user receiving end per unit time based on the network card receiving rate, calculate a link utilization rate of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, calculate an amount of data that can be received by the receiving-end switch per unit time based on the link utilization, calculate a receiving rate of the receiving-end switch based on the amount of data that can be received by the receiving-end switch per unit time, calculate a sending rate of each sending end based on the number of sending ends connected to the sending-end switch and the receiving rate of the receiving-end switch, and wherein the sending-end switch, when an incast occurs, sends data to the receiving-end switch based on the sending rate.

[0035] Additional advantages, objects, and features of the present invention will be described in part in the following description and will become apparent to those skilled in the art after studying the following or may be learned by practice of the present invention. The objects and other advantages of the present invention may be particularly pointed out and attained in the description and drawings.

[0036] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.

[0038] Figure 1Schematic diagram of an implementation of the data center RDMA Incast solution based on ToR switches of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0040] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0041] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0042] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0043] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0044] In existing technologies, the incast problem has two main characteristics: 1. A many-to-one communication model, where multiple servers respond to the same requester's data transmission, resulting in an asymmetric data transmission pattern; 2. Response synchronization, where servers in a data center network respond to a requester's data request and begin transmitting data almost simultaneously, resulting in a large influx of traffic within the data center network within a short period of time. These characteristics of incast can have a serious impact on data center networks. In traditional TCP networks, incast can cause switch buffer overflows, resulting in a large number of packet losses and retransmissions. In RDMA networks, incast can cause head-of-line blocking, triggering priority flow control (PFC) mechanisms, and even packet loss and retransmission. The triggering of PFC mechanisms significantly impacts tail latency and can also cause pause frame storms, deadlocks, and other problems within the network. Existing technologies have proposed many TCP incast solutions based on end-to-end transmission control and ACK return mechanisms. However, due to the high link rates and changes in transmission mechanisms of RDMA networks, TCP incast solutions are ineffective in RDMA networks. An accurate and effective RDMA incast control scheme is urgently needed.

[0045] In existing technologies, RDMA incast solutions for data centers are primarily categorized into congestion control (CC) solutions and auxiliary solutions based on congestion control. In RDMA networks, the main congestion control methods currently used in the industry are as follows: 1. DCQCN uses an explicit congestion control (ECN) mechanism based on the queue length of the switch's outbound port to control the flow sending rate. 2. Using the RTT (round-trip time) gradient as a congestion feedback signal allows for more sensitive perception of network congestion and timely response to congestion. The core principles of both DCQCN and Timely are to use heuristic algorithms to adjust the sender's sending rate. This adjustment is relatively coarse-grained and cannot ensure rapid convergence of network performance when congestion changes. 3. Using in-network telemetry (INT) technology to calculate the network utilization of bottleneck links within the network as a congestion feedback signal enables more fine-grained adjustment of the sending rate and window, resulting in superior performance in terms of flow completion time (FCT), fairness, and tail latency.

[0046] Congestion control-based auxiliary solutions primarily address the characteristics of incast and incorporate the features of RDMA networks. The earliest work, DCQCN+, addresses the shortcomings of DCQCN in addressing incast. DCQCN+ maintains a flow table recording active flows at the receiving end, which responds with congestion notification packets (CNPs) based on the number of active flows. The sending end of DCQCN+ employs an improved reaction point (RP) algorithm to enhance its incast handling capabilities. Another existing approach proposes an incast control strategy based on dynamic windows in switches. The switch maintains a send window for each receiver, limiting the amount of data bursts within a fixed time interval. The upstream link updates the window value of the downstream link based on its own available bandwidth period. After the window is exhausted, excess data at the receiving end of the switch is isolated using a virtual output queue (VOQ), reducing the impact of incast on other flows and thereby improving overall network performance in incast scenarios.

[0047] However, the aforementioned traditional RDMA congestion control uses an end-to-end control scheme, requiring at least one RTT to respond to network congestion. While DCQCN+ addresses the shortcomings of DCQCN with incast, it still uses end-to-end transmission control and is similarly unsuitable for RDMA networks with high link rates. Another existing technique maintains a send window for each receiver, but to ensure the normal transmission of non-incast flows and avoid excessive switch resource usage due to frequent window updates, a large initial window value is set. This prevents incast control from achieving ideal results. As RDMA network link rates continue to increase, more and more flows will be transmitted within a single RTT. Therefore, traditional congestion control schemes are unable to respond promptly to incast issues, impacting the transmission of non-incast flows. Furthermore, most congestion control schemes use heuristic algorithms to adjust the send rate, preventing network performance from converging quickly. In summary, traditional congestion control schemes are no longer able to address the incast problem in RDMA networks.

[0048] To solve the above problems, Figure 1 As shown, the present invention proposes a data center RDMA Incast solution based on a ToR switch. The RDMA network includes a sending end switch connected to the sending end and a receiving end switch connected to the user receiving end. The steps of the method include:

[0049] In some embodiments of the present invention, both the sending-end switch and the receiving-end switch are ToR switches.

[0050] In some embodiments of the present invention, TOR (Top of Rack) access is a common cabling method from server to switch or switch to switch. A TOR switch can be an access layer switch, an aggregation layer switch, or a core layer switch. When a TOR switch is located at the top of a server cabinet, connecting the server to the core or aggregation layer switches, it is called an access layer switch; when a TOR switch is located at the top of a switch cabinet, it is called an aggregation or core layer switch.

[0051] Step S100: Based on the incast situation detected by the sending switch, if an incast situation occurs, the data flow is cached in the sending switch;

[0052] In some embodiments of the present invention, the sending-end switch is provided with a buffer space, and the data flow is buffered in the buffer space of the sending-end switch.

[0053] Step S200: Obtaining a network card receiving rate of a user receiving end based on the receiving end switch, calculating an amount of data that can be received by the user receiving end per unit time based on the network card receiving rate, calculating a link utilization rate of the receiving end switch based on the amount of data that can be received by the user receiving end per unit time, calculating an amount of data that can be received by the receiving end switch per unit time based on the link utilization, calculating a receiving rate of the receiving end switch based on the amount of data that can be received by the receiving end switch per unit time, calculating a sending rate of each sending end based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, and the sending end switch where an incast occurs sends data to the receiving end switch based on the sending rate.

[0054] In some embodiments of the present invention, in the step of calculating the sending rate of each sending end based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, a sending rate is allocated to the sending end switch based on the number of sending ends connected to the sending end switch. If the number of sending ends connected to the sending end switch is large, a higher rate is allocated; if the number of sending ends connected to the sending end switch is small, a lower rate is allocated, thereby ensuring that each sending end switch that experiences incast can output data stably.

[0055] Using the above solution, this solution only needs to deploy programs on the sending and receiving switches of the RDMA network. The sending switch is used to detect incast situations. When an incast situation occurs, the data stream is cached to prevent a large amount of data streams from entering the RDMA network and causing data congestion. On the other hand, the receiving switch calculates the maximum receiving rate of the receiving switch based on the user receiving end situation and allocates a sending rate to each sending switch, quickly and efficiently outputting the data cached by the sending switch. This solution does not require deployment between the switches between the sending and receiving switches, which not only improves the processing efficiency of incast situations, but also reduces the deployment difficulty and speeds up program processing.

[0056] In some embodiments of the present invention, the sending-end switch where an incast occurs performs rate limiting on the transmission of the incast traffic according to a token bucket mechanism.

[0057] The token bucket mechanism is a commonly used rate-limiting algorithm in switches. It generates tokens at a set rate and limits the burst size of data based on the remaining number of tokens, thereby limiting the data transmission rate. When the token transmission is completed, it is returned to the token bucket and the output task continues to be executed.

[0058] In some embodiments of the present invention, the step of sensing the incast situation based on the transmitting-end switch includes:

[0059] The sending end switch counts the amount of data sent to the same user receiving end through the sending end switch within a preset time length;

[0060] When the amount of data sent to the same user receiving end through the sending end switch within a preset time length is greater than a preset incast threshold, it is determined that an incast situation occurs.

[0061] In some embodiments of the present invention, the transmitting switch and the receiving switch each maintain a data flow table for recording internal data flows of the transmitting switch or the receiving switch, and the data flow table is cleared once every preset time interval;

[0062] In the data flow table of the sending end switch, a data sub-flow table is set for each sending end connected to the sending end switch;

[0063] In a specific implementation process, when the amount of data in the data flow table of the sending switch is greater than a preset incast threshold, it is determined that an incast situation occurs.

[0064] In some embodiments of the present invention, when the amount of data sent from the transmitting switch to the same user receiving end within a preset time period is greater than a preset incast threshold, an incast is determined to have occurred and the transmitting switch sets an incast flag. When the amount of data cached in the transmitting switch is less than or equal to the incast threshold, the incast flag is removed. After the transmitting switch removes the incast flag, the transmitting switch sends data normally.

[0065] In some embodiments of the present invention, the step of sensing an incast condition based on the transmitting-end switch further includes: upon determining that an incast condition has occurred, the transmitting-end switch continuing to transmit a first preset amount of data to the user receiving end; and in the step of caching the data stream in the transmitting-end switch if an incast condition has occurred, after the transmitting-end switch transmits the first preset amount of data to the user receiving end, caching the data stream in the transmitting-end switch.

[0066] With the above solution, if the transmitting switch immediately stops sending data after determining that an incast has occurred, which would cause the receiving switch to waste link bandwidth while waiting for the allocated rate value, this solution allows the transmitting switch to continue transmitting the first preset size of data to the user receiving end after determining that an incast has occurred, thereby avoiding wasting link bandwidth while the receiving switch waits for the allocated rate value.

[0067] In some embodiments of the present invention, if an incast occurs, the step of caching the data flow in the transmitting switch further includes constructing a queue for the data flow cached in the transmitting switch based on the input time of the data flow, and in the step of the transmitting switch sending data to the receiving switch based on the sending rate when the incast occurs, the transmitting switch sends the data to the receiving switch based on the queue order.

[0068] In some embodiments of the present invention, in the step of calculating the amount of data that can be received by the user receiving terminal per unit time based on the network card receiving rate, the amount of data that can be received by the user receiving terminal per unit time is calculated based on the following formula:

[0069] ;

[0070] in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the amount of data that the user receiving end can receive per unit time;

[0071] In the step of calculating the link utilization of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, the link utilization of the receiving-end switch is calculated based on the following formula:

[0072]

[0073] in, Indicates the amount of data that the user receiving end can receive per unit time. Indicates the amount of data buffered by the port connected to the current receiving switch and the user receiving end. Indicates the link utilization of the receiving switch.

[0074] In some embodiments of the present invention, in the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the link utilization, the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time is obtained, and the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization are used to calculate the amount of data that the receiving-end switch can receive per unit time in the current unit time.

[0075] In some embodiments of the present invention, in the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization, the amount of data that the receiving-end switch can receive per unit time is calculated based on the following formula:

[0076]

[0077] in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the link utilization of the receiving switch, Indicates the amount of data that the receiving switch can receive per unit time, calculated in the previous unit time. Indicates the amount of data that the receiving switch can receive in the current unit time. are all constants.

[0078] In some embodiments of the present invention, the calculated data volume that the receiving end switch can receive per unit time is used as the data volume that the receiving end switch can receive per unit time calculated for the previous unit time in the next calculation.

[0079] By adopting the above solution, the amount of data that the receiving end switch can receive per unit time is calculated based on the amount of data that the receiving end switch can receive per unit time calculated by the receiving end switch in the previous unit time, and the amount of data that the receiving end switch can receive per unit time in the next unit time is calculated to improve the calculation accuracy.

[0080] In some embodiments of the present invention, in the step of calculating the receiving rate of the receiving switch based on the amount of data that the receiving switch can receive per unit time, the receiving rate of the receiving switch is calculated based on the following formula:

[0081]

[0082] in, Indicates the receiving rate of the receiving switch. Indicates the amount of data that the receiving switch can receive in the current unit time. Indicates the length of time per unit time.

[0083] In some embodiments of the present invention, in the step of calculating the sending rate of each sending end based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, the sending rate of each sending end switch is calculated based on the following formula:

[0084]

[0085] in, Indicates the currently calculated sending rate of the sending switch. Indicates the receiving rate of the receiving switch. Indicates the number of senders connected to the sender switch currently being calculated, Indicates the total number of senders connected to the sender switch where incast occurs.

[0086] The above scheme is adopted. This scheme calculates the amount of data that the receiving switch can receive per unit time in the next unit time based on the amount of data that the receiving switch can receive per unit time calculated by the receiving switch in the previous unit time. The total receiving rate of the receiving switch is calculated based on the amount of data that the receiving switch can receive per unit time in the next unit time. The sending rate is then allocated to each sending switch based on the number of sending terminals connected to the currently calculated sending switch. The sending rate allocated to each sending switch can be updated every unit time, which can ensure the maximum rate transmission and solve the incast problem.

[0087] Another aspect of the present invention provides a data center RDMA Incast solution system based on a ToR switch, the system includes a sending end switch management module and a receiving end switch management module.

[0088] The sending end switch management module is used to sense the incast situation based on the sending end switch, and if the incast situation occurs, cache the data flow in the sending end switch;

[0089] The receiving-end switch management module is configured to obtain a network card receiving rate of a user receiving end based on the receiving-end switch, calculate an amount of data that can be received by the user receiving end per unit time based on the network card receiving rate, calculate a link utilization rate of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, calculate an amount of data that can be received by the receiving-end switch per unit time based on the link utilization, calculate a receiving rate of the receiving-end switch based on the amount of data that can be received by the receiving-end switch per unit time, calculate a sending rate of each sending end based on the number of sending ends connected to the sending-end switch and the receiving rate of the receiving-end switch, and wherein the sending-end switch, when an incast occurs, sends data to the receiving-end switch based on the sending rate.

[0090] Using the above solution, the system only needs to deploy programs on the sending and receiving switches of the RDMA network. The sending switch is used to detect incast situations. When an incast situation occurs, the data stream is cached to prevent a large amount of data streams from entering the RDMA network and causing data congestion. On the other hand, the receiving switch calculates the maximum receiving rate of the receiving switch based on the user receiving end situation and allocates a sending rate to each sending switch, quickly and efficiently outputting the data cached by the sending switch. This solution does not require deployment between the switches between the sending and receiving switches, which not only improves the processing efficiency of incast situations but also reduces the difficulty of deployment.

[0091] In existing technologies, the RDMA incast problem poses a significant challenge to data center networks. Traditional RDMA congestion control solutions use end-to-end control and heuristic algorithms to address network congestion. This results in a large amount of data accumulating in the switch buffer after an incast occurs, triggering the PFC mechanism, leading to head-of-line blocking, and even severe packet loss. Furthermore, the triggering of the PFC mechanism can lead to pause frame storms, deadlocks, and other problems.

[0092] The goal of RDMA incast control is to minimize the impact of incasts on non-incast traffic without restricting normal incast transmission. The key lies in fast and accurate incast identification and traffic scheduling to reduce the accumulation of incast traffic in the switch buffer. Coarse-grained incast identification and traffic scheduling using the switch's dynamic window mechanism and voice-over-queue (VOQ) control mechanism cannot quickly identify incasts, resulting in a large accumulation of incast traffic in the switch buffer at the start of an incast, severely impacting the transmission of non-incast traffic.

[0093] The solution of the present invention realizes rapid and accurate identification and isolation of incast through the sending and receiving switches, reduces the accumulation of incast traffic in the switch buffer, and uses a token bucket mechanism to control the transmission rate of incast traffic, thereby quickly converging the incast transmission rate.

[0094] The main advantages of this solution include:

[0095] 1. Rapidly and accurately identify and isolate incast traffic, reducing the accumulation of incast traffic in the switch buffer;

[0096] 2. Use a token bucket mechanism to accurately control the transmission rate of incast flows, allowing them to quickly converge to the available bandwidth of the bottleneck link, reducing the impact of incast on non-incast flow transmission;

[0097] 3. It is only deployed on the sending and receiving switches, is compatible with existing transmission control solutions, and has low hardware overhead.

[0098] An embodiment of the present invention also provides a data center RDMA incast solution device based on a ToR switch, the device including a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor being used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the device implements the steps implemented by the method described above.

[0099] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the aforementioned ToR switch-based data center RDMA incast solution. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.

[0100] It should be understood by those skilled in the art that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether to implement the system in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention. When implemented in hardware, it may be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via a data signal carried in a carrier wave.

[0101] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0102] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0103] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A data center RDMA incast solution based on ToR switches, characterized by: The RDMA network includes a sending end switch connected to a sending end and a receiving end switch connected to a user receiving end, wherein both the sending end switch and the receiving end switch are ToR switches. The method includes the following steps: The sending switch senses the incast situation and caches the data flow in the sending switch if an incast situation occurs. The network card receiving rate of the user receiving end is obtained based on the receiving end switch, the amount of data that the user receiving end can receive per unit time is calculated based on the network card receiving rate, the link utilization of the receiving end switch is calculated based on the amount of data that the user receiving end can receive per unit time, the amount of data that the receiving end switch can receive per unit time is calculated based on the link utilization, the receiving rate of the receiving end switch is calculated based on the amount of data that the receiving end switch can receive per unit time, the sending rate of each sending end is calculated based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, and the sending end switch, where an incast occurs, sends data to the receiving end switch based on the sending rate.

2. The data center RDMA incast solution based on ToR switches according to claim 1 is characterized in that: The step of sensing the incast situation based on the sending end switch includes: The sending end switch counts the amount of data sent to the same user receiving end through the sending end switch within a preset time length; When the amount of data sent to the same user receiving end through the sending end switch within a preset time length is greater than a preset incast threshold, it is determined that an incast situation occurs.

3. The data center RDMA incast solution based on a ToR switch according to claim 1 or 2, characterized in that: The step of sensing an incast condition based on the transmitting-end switch further includes, when determining that an incast condition has occurred, continuing, by the transmitting-end switch, to transmit a first preset amount of data to the user receiving end; and in the step of caching the data stream in the transmitting-end switch if an incast condition has occurred, caching the data stream in the transmitting-end switch after the transmitting-end switch has transmitted the first preset amount of data to the user receiving end.

4. The data center RDMA incast solution based on ToR switches according to claim 1 is characterized in that: The step of caching the data flow in the transmitting switch if an incast occurs further includes establishing a queue for the data flow cached in the transmitting switch based on an input time of the data flow, and in the step of the transmitting switch sending data to the receiving switch based on a sending rate when the incast occurs, the transmitting switch sends the data to the receiving switch based on a queue order.

5. The data center RDMA incast solution based on ToR switches according to claim 1 is characterized in that: In the step of calculating the amount of data that can be received by the user receiving end per unit time based on the network card receiving rate, the amount of data that can be received by the user receiving end per unit time is calculated based on the following formula: ; in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the amount of data that the user receiving end can receive per unit time; In the step of calculating the link utilization of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, the link utilization of the receiving-end switch is calculated based on the following formula: in, Indicates the amount of data that the user receiving end can receive per unit time. Indicates the amount of data buffered by the port connected to the current receiving switch and the user receiving end. Indicates the link utilization of the receiving switch.

6. The data center RDMA incast solution based on ToR switches according to claim 1, characterized in that: In the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the link utilization, the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time is obtained, and the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization are calculated.

7. The data center RDMA incast solution based on ToR switches according to claim 6 is characterized in that: In the step of calculating the amount of data that the receiving-end switch can receive per unit time based on the amount of data that the receiving-end switch can receive per unit time calculated in the previous unit time and the link utilization, the amount of data that the receiving-end switch can receive per unit time is calculated based on the following formula: in, Indicates the network card receiving rate of the user receiving end. Indicates the length of time per unit time, Indicates the link utilization of the receiving switch, Indicates the amount of data that the receiving switch can receive per unit time, calculated in the previous unit time. Indicates the amount of data that the receiving switch can receive in the current unit time. are all constants.

8. The data center RDMA incast solution based on ToR switches according to claim 1 is characterized in that: In the step of calculating the receiving rate of the receiving switch based on the amount of data that the receiving switch can receive per unit time, the receiving rate of the receiving switch is calculated based on the following formula: in, Indicates the receiving rate of the receiving switch. Indicates the amount of data that the receiving switch can receive in the current unit time. Indicates the length of time per unit time.

9. The data center RDMA incast solution based on ToR switches according to claim 1, characterized in that: In the step of calculating the sending rate of each sending end based on the number of sending ends connected to the sending end switch and the receiving rate of the receiving end switch, the sending rate of each sending end switch is calculated based on the following formula: in, Indicates the currently calculated sending rate of the sending switch. Indicates the receiving rate of the receiving switch. Indicates the number of senders connected to the sender switch currently being calculated, Indicates the total number of senders connected to the sender switch where incast occurs.

10. A data center RDMA incast solution system based on ToR switches, characterized in that: The system includes a sending end switch management module and a receiving end switch management module. The sending end switch management module is used to sense the incast situation based on the sending end switch, and if the incast situation occurs, cache the data flow in the sending end switch; The receiving-end switch management module is configured to obtain a network card receiving rate of a user receiving end based on the receiving-end switch, calculate an amount of data that can be received by the user receiving end per unit time based on the network card receiving rate, calculate a link utilization rate of the receiving-end switch based on the amount of data that can be received by the user receiving end per unit time, calculate an amount of data that can be received by the receiving-end switch per unit time based on the link utilization, calculate a receiving rate of the receiving-end switch based on the amount of data that can be received by the receiving-end switch per unit time, calculate a sending rate of each sending end based on the number of sending ends connected to the sending-end switch and the receiving rate of the receiving-end switch, and wherein the sending-end switch, when an incast occurs, sends data to the receiving-end switch based on the sending rate.

Citation Information

Patent Citations

  • TCP friendly rate control method based on change rate of handling capacity and ECN mechanism

    CN103281255A

  • Data center network congestion control method and system

    CN114760252A