Data center network congestion control method and device based on data and credit coupling
By introducing a congestion control method based on data and credit coupling in the data center network, using the ECN signal and credit rate control mechanism, the throughput decline caused by excessive response delay and credit loss is solved, efficient congestion control and flow scheduling are achieved, and the performance needs of AI applications are met.
Patent Information
- Application Number
- CN202510561461.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the prior art, the response time delay is too long, and it is difficult for the sending end to suppress the data injection rate in time, resulting in network congestion. Data transmission based on the credit mechanism has problems of data queuing instability and credit loss caused by throughput decline.
The data center network congestion control method based on data and credit coupling is adopted. By designing a credit rate control mechanism based on ECN, the credit rate is dynamically adjusted, the data and credit queues are synchronized management, resource waste is reduced, and a credit-driven flow scheduling mechanism is designed to optimize the flow completion time and deadline loss rate.
It effectively improves congestion control performance, reduces resource waste, meets the requirements of data center networks in low latency, high throughput and stream deadlines, and improves the efficiency and stability of network data transmission.
Smart Images

Figure CN120090974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital information transmission, and particularly relates to a data center network congestion control method and device based on data and credit coupling. Background Art
[0002] With the continuous development of artificial intelligence (AI) and the Internet of Things (IoT), the amount of data generated by real-time AI applications has increased sharply, and this data is usually transmitted to the data center for storage and analysis. To ensure the performance requirements of such applications, the data center must efficiently transmit data streams with low latency, high throughput, and deadline constraints.
[0003] Driven by the growing performance requirements of diverse AI applications, the link bandwidth required by the data center has rapidly increased to 100 Gbps and 400 Gbps, and this trend continues. At the same time, network traffic exhibits complex and time-varying characteristics, including both long-period continuous flows (data streams) and short-term bursty flows, which poses a huge challenge to achieving high throughput, low latency, and good stability simultaneously. Once congestion occurs, the sudden packet queuing and unpredictable packet loss will cause a sharp decline in network performance, thus seriously affecting the user experience.
[0004] Congestion control is crucial for the overall performance of data transmission. ECN (Explicit Congestion Notification) is a network protocol that allows network devices, such as routers or switches, to explicitly notify the sender or receiver of congestion by marking a header field of a packet (instead of directly dropping the packet) when congestion occurs. For congestion problems in data center networks, traditional congestion control schemes based on the try-and-backoff mechanism of ECN, RTT (Round-Trip Time), or INT (In-band Network Telemetry) feedback usually require at least one RTT measurement period to respond to congestion. Therefore, when congestion occurs, the response latency of such traditional methods is too long, and it is difficult for the sender to timely suppress the data injection rate, especially in high-speed data centers, where the problem is particularly prominent if the data flow is completed within an extremely short RTT. In recent years, receiver-driven credit-based congestion control schemes have gradually emerged, which can precisely regulate data transmission and prevent congestion from occurring at the source. In the credit-based congestion control mechanism, the receiver issues a clear credit limit according to its own reception ability, and controls the data transmission of the sender in a one-to-one credit-and-data manner, realizing fine-grained packet-level feedback and accurately reflecting the network state. However, the existing credit-based mechanisms still face problems such as unstable data queuing and throughput degradation caused by bursty data traffic and credit loss. Summary of the Invention
[0005] In view of this, to solve the problems of too long response latency in the prior art, the sender's difficulty in timely suppressing the data injection rate, and throughput degradation caused by traffic and credit loss, the present invention provides a congestion control method and device for a data center network based on data and credit coupling. The congestion control method for a data center network based on data and credit coupling can effectively improve congestion control performance, reduce resource waste, and meet the requirements of data center networks in terms of low latency, high throughput, and flow deadline.
[0006] The present invention provides a congestion control method for a data center network based on data and credit coupling, the method comprising: Obtaining the structure of the data center network, including switches, senders, and receivers; Based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, developing a credit rate control mechanism based on ECN for managing data congestion and credit congestion and performing congestion control for the entire loop of the data center network; the credit rate control mechanism based on ECN includes: Setting the switch as the congestion point location, the sender as the notification point location, and the receiver as the reaction point location; The receiving end sends credits through the reverse path of credit transmission, pulls packets from the sending end, and uses credit rate limiting to prevent data congestion on the reverse path; The same ECN feedback signal is used for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / packet, and when the credit / data queue is not congested, the ECN label is not marked for the credit / packet; According to the result of marking the ECN label, including whether the credit / packet carries the ECN label or not, a credit rate control algorithm based on ECN is used to dynamically adjust the credit rate.
[0007] Furthermore, the congestion control of the full loop of the data center network includes: Design a speculative probing mechanism to prevent bandwidth waste in the first RTT transmission of the flow during the congestion control process; the RTT is the round-trip delay of the flow, which represents the time from the sending end sending the packet to the sending end receiving the feedback; Formulate data retransmission policies and credit retransmission policies to cope with data loss and credit loss during the congestion control process; Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through credit rate limiting at the receiving end and credit selective dropping at the switch end.
[0008] In addition, the present invention also provides a data center network congestion control device based on the coupling of data and credit. The device uses the steps of the foregoing method to perform congestion control of the data center network based on the coupling of data and credit. The device includes the following modules: The first module is used to obtain the structure of the data center network, including switches, sending ends, and receiving ends; The second module is used to develop a credit rate control mechanism based on ECN for managing data congestion and credit congestion and performing congestion control of the full loop of the data center network based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework; The second module further includes: Sub-module 1 is used to set the switch as the congestion point location, the sending end as the notification point location, and the receiving end as the reaction point location; Sub-module 2 is used to make the receiving end send credits through the reverse path of credit transmission, pull packets from the sending end, and use credit rate limiting to prevent data congestion on the reverse path; Sub-module 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / packet, and when the credit / data queue is not congested, the ECN label is not marked for the credit / packet; The sub-module 4 is configured to dynamically adjust the credit rate by using an ECN-based credit rate control algorithm according to the results of marking ECN tags, including credit / data packet carrying ECN tags and not carrying ECN tags.
[0009] In summary, the present invention provides a data center network congestion control method and apparatus based on data and credit coupling. Compared with the prior art, the technical solution of the present invention has the following advantages: 1) The DCEF framework is introduced, and a new data and credit coupling congestion control scheme is proposed based on this framework. By dynamically synchronizing the credit transmission rate and the data reception rate, accurate real-time credit rate adaptation is achieved, and at the same time, one-to-one credit data transmission is enforced, effectively alleviating congestion.
[0010] 2) An ECN-based credit rate control mechanism is developed to achieve full-loop congestion control and solve the problem of credit waste.
[0011] 3) By using a credit-driven flow scheduling mechanism, the flow completion time (FCT) and deadline miss rate (DMR) of congestion control are further reduced, and high-throughput network data transmission can be achieved to meet the performance requirements of AI applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic flowchart of a data center network congestion control method based on data and credit coupling provided by the first embodiment of the present invention; Figure 2 It is a schematic diagram of bandwidth waste caused by credit loss in an embodiment of the present invention; Figure 3 It is a schematic comparison diagram of the transmission situations of flows with different traffic sizes before and after flow scheduling in an embodiment of the present invention, where are data streams with different traffic sizes, is the unit time, is the unit bandwidth, Figure 3 (a) is the transmission situation before flow scheduling, Figure 3 (b) is the transmission situation after flow scheduling; Figure 4 It is a schematic diagram of an ECN-based credit rate control framework and control process in an embodiment of the present invention, where is the threshold of the credit queue length, is the threshold of the data queue length, CP represents the congestion point, NP represents the notification point, and RP represents the reaction point; Figure 5 It is a schematic flowchart of a congestion control credit-driven flow scheduling provided in an embodiment of the present invention, where is the threshold of the credit queue length. Detailed implementation manners
[0013] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0014] In artificial intelligence (AI) application scenarios, for example, in AIoT (Artificial Intelligence of Things, a scenario where artificial intelligence technology is deeply integrated with the Internet of Things) scenarios, in order to meet the high-performance requirements of artificial intelligence (AI) applications for large amounts of data, low latency, high reliability, etc. in data center networks, the present invention proposes a congestion control solution for data centers. The solution is based on the DCEF framework, deeply explores the tight coupling mechanism between the data plane and the credit plane in congestion control design, and realizes the coordinated regulation of data and credit in a complete control loop. In addition, the present invention effectively responds to the requirements of flow deadlines by introducing a credit-driven flow scheduling mechanism.
[0015] In one embodiment, the present invention proposes a method for congestion control of a data center network based on the coupling of data and credit, as Figure 1 shown, the method includes: Obtain the structure of the data center network, including switches, senders, and receivers; Based on the principles of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, develop an ECN-based credit rate control mechanism for managing data congestion and credit congestion, and perform congestion control for the entire loop of the data center network; the ECN-based credit rate control mechanism includes: Set the switch as the congestion point location, the sender as the notification point location, and the receiver as the reaction point location; The receiver sends credit through the reverse path of credit transmission, pulls packets from the sender, and uses credit rate limiting to prevent data congestion on the reverse path; Use the same ECN feedback signal for credit congestion and data congestion; when congestion occurs in the credit / data queue, the switch marks the ECN label for the credit / packet, and when no congestion occurs in the credit / data queue, the ECN label is not marked for the credit / packet; According to the results of marking the ECN label, including whether the credit / packet carries the ECN label or not, dynamically adjust the credit rate using the ECN-based credit rate control algorithm.
[0016] Furthermore, the congestion control of the entire loop of the data center network includes: Design a speculative detection mechanism to prevent bandwidth waste in the first RTT transmission of a flow during congestion control; the RTT is the round-trip delay of the flow, representing the time from when a data packet is sent from the sender to when the sender receives feedback. Formulate a data retransmission policy and a credit retransmission policy to address data loss and credit loss during congestion control. Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limiting and switch-side credit selective dropping.
[0017] Specifically, the design goal of the present invention is to achieve a short flow completion time (FCT), high throughput, short queuing time, few packet losses, and a low deadline miss rate (DMR). This goal brings many challenges. The most difficult challenge is that when credit is discarded by the switch due to credit queue congestion, the credit will be wasted, thereby reducing the effective throughput of the flow, and it is not easy to meet the performance requirements of real-time AI applications. Therefore, based on the data and credit coupling design principle in the DCEF framework, the present invention designs a new data and credit coupling mechanism to address the above challenges.
[0018] The DCEF (Data–Credit and End-to-end Feedback) framework is an intelligent congestion control technology for modern data center networks. Its core idea is to achieve high-throughput and low-latency data center network resource management through the dynamic coupling of data and credit and the closed-loop optimization of end-to-end feedback. Compared with traditional reactive methods based on coarse-grained feedback, the participation of the credit plane can improve congestion control performance in terms of convergence and accuracy, specifically manifested as follows: 1) Credit-based congestion control performs better. In traditional round-trip time (RTT)-based congestion control, the sender considers the transmission delay of the entire path, resulting in inefficient congestion management in reducing data queuing to achieve high throughput. When the queue length exceeds a certain threshold, the ECN signal provides feedback on data congestion, but this is still coarse-grained information and cannot help the sender precisely adjust the data transmission rate. INT provides specific link load information, but the sender needs time to limit the data injection rate.
[0019] The credit mechanism provides the sender with precise packet-level network state information, enabling more effective congestion prevention. In credit-based congestion control, the receiver sends credit to the sender according to its receiving ability, and the sender sends a data packet after receiving each credit. In this way, if the credit transmission rate is well regulated, data congestion can be prevented from occurring.
[0020] 2) The full - closed - loop congestion control based on the DCEF framework helps to achieve near - optimal data transmission and improves the integrity of control by using a full control loop.
[0021] Each processing process in the data processor and the credit processor has its unique function, and the congestion control scheme should not bypass these processes. Data processing is the key to notifying the data queue of congestion, and methods that ignore this process are inefficient in managing data congestion within the network. Credit processing is crucial for credit - based approaches. It can precisely prevent data congestion at the switch, including credit rate limiting and recording credit congestion signals. Similar to controlling the credit rate at the receiving end, performing switch credit rate limiting on one side of the link can avoid reverse - path data congestion on the other side. In addition, credit congestion information is crucial for the receiving end to adjust the credit rate and avoid credit waste.
[0022] Analysis and comparison of the DCEF framework and related existing technologies show that tight data - credit cooperation improves the accuracy of control, and the mechanisms of the data plane and the credit plane should have a coupled design. The control from credit to data is the core of credit - based congestion control, directly determining the congestion control efficiency and serving as the basis for any credit - based congestion control scheme to work effectively. On the other hand, the feedback from data to credit is crucial for effectively preventing data congestion. Credit - based congestion control schemes need it to precisely adjust data transmission when unexpected congestion occurs.
[0023] The present invention obtains the key idea to promote the process of technology development from the above - mentioned analysis process, that is, high - performance congestion control requires designing a reasonable data - credit coupling mechanism, and in order to achieve network - wide congestion control, each control process should be incorporated into the control design. However, it is challenging to design a data - credit coupling mechanism to meet the flow deadline while implementing congestion control based on data - credit coupling, mainly reflected in: 1) It is necessary to alleviate the bandwidth waste caused by credit loss.
[0024] Credit - based congestion control schemes should prevent credit waste because credit waste has a great negative impact on data throughput. Since each credit processor on the transmission path performs rate limiting on credits, many credits may be discarded due to exceeding the speed limit. Once a credit is lost at the switch, the previously passed link cannot be fully utilized.
[0025] Taking a specific credit transmission process as an example to illustrate the mechanism of credit loss. For the sake of description, it is assumed here that each switch only allows one credit to pass within a given time. As Figure 2 shown in the credit transmission path, the credit transmission process involves = 3 switches and servers, showing = Transmission process of 4 credits. The server is the terminal node of the transmission, including at least a sending end as the notification point and a receiving end as the response point. Sequence numbers are assigned to the credits: credit , , where credit successively passes through the 3rd switch and the 2nd switch successfully, but credit and credit are discarded. Then, when credit is allowed to pass, credit is discarded at the 1st switch. In this case, the links between the 2nd switch and the 3rd switch and between the 1st switch and the 2nd switch are not fully utilized because credits are all discarded without corresponding credits successfully reaching the sending end.
[0026] For example, ExpressPass (Fast Credit Transmission Protocol) in the existing credit-based congestion control method is a congestion control protocol for data center networks. It dynamically adjusts credit loss through a credit feedback mechanism, but this loss rate-based algorithm cannot fundamentally reduce credit waste.
[0027] 2) A flow scheduling mechanism that meets the requirements of artificial intelligence applications needs to be designed.
[0028] The diversification of artificial intelligence applications has led to the complexity of traffic performance requirements, which are mainly divided into two types: deadline traffic and non-deadline traffic. Deadline traffic is usually related to tasks with clear completion time limits. In this type of traffic, if data transmission exceeds the deadline that users can wait for, it will lose its value, which will not only affect the user experience but also waste bandwidth. For non-deadline traffic, the completion time and throughput of the traffic become its main performance optimization goals.
[0029] The flow scheduling mechanism is the key method to improve the transmission efficiency of AI traffic with different performance requirements. Let's assign sequence numbers to the flows to be processed to obtain the set of flows , is the total number of flows. Specifically, consider four flows with different traffic sizes , which are transmitted through the bottleneck link. The specific flow attributes (traffic size and deadline) are shown in Table 1 below: Table 1 Example of flow attribute representation
[0030] Combined with Table 1, Figure 3 shows the comparison of traffic transmission situations of 4 different-sized flows (data flows) with and without flow scheduling. Among them, Denote the unit time, Denote the unit bandwidth. The relationship between the bottleneck link bandwidth and the flow transmission time is as Figure 3 (a) and Figure 3 (b) shown. When no flow scheduling is performed in Figure 3 (a), the transmission time of is , the transmission time of is the transmission time of is the transmission time of is. It can be seen that all flows will miss the deadline, and the deadline miss rate (DMR) is 100%. In contrast, after flow scheduling in Figure 3 (b), by preferentially transmitting short flows, the time taken to complete the transmission is , the time taken to complete the transmission is , the time taken to complete the transmission is , the time taken to complete the transmission is Only flow will miss the deadline. It can be seen that flow scheduling reduces the DMR to 25%. Obviously, by scheduling shorter data flows for transmission first, the DMR is reduced from 100% to 25%.
[0031] Traditional flow scheduling methods only schedule data flows in the switch queue. By adopting congestion control with data and credit coupling, the switch is usually in a near-zero queue state. When there are multiple data flows with different performance requirements at the sender or receiver that need to be sent or received simultaneously, the performance of the scheduling strategy will be greatly reduced.
[0032] Considering the above challenges, based on the principle of designing a data and credit coupling mechanism in the above DCEF framework, the present invention proposes a new credit control mechanism that utilizes all congestion feedback from two planes (data plane and credit plane) to perform credit rate control, considers ECN as the congestion signal for both planes, and integrates all control processes related to congestion regulation, including speculative probing and data retransmission, into the ECN-based credit rate control. By buffering excessive credit and using ECN-based rate control to adjust the credit rate at the receiver, credit waste is prevented; on the other hand, the present invention designs a credit-driven flow scheduling mechanism that considers indirectly adjusting the data reception order at the receiver by attempting to control the credit rate at the receiver. In this way, under the condition that the switch does not queue, traffic in the data center network that meets different performance requirements can be effectively scheduled. Specifically, it includes: 1. Develop an ECN-based credit rate control mechanism to manage data congestion and credit congestion, and achieve congestion control for the entire loop of the data center network.
[0033] In this embodiment, relying on the notifications of switches and the responses of receivers, as Figure 4 shown, the process of the ECN-based credit rate control mechanism performing congestion control involves three positions in the network, namely the congestion point (CP), the notification point (NP), and the response point (RP). The congestion point is the location where congestion actually occurs, the notification point is the location where the ECN label is marked, and the response point is the location where the ECN label is received and the credit rate is adjusted. Among them, the CP is the switch. Different from the traditional ECN-based congestion control, here the NP is the sender and the RP is the receiver. Figure 4 In the network architecture of ECN-based credit rate control, it includes a switch (CP), a sender (NP), and a receiver (RP).
[0034] The ECN-based credit rate control mechanism also includes: the receiver sends credit through the reverse path of credit transmission to pull packets from the sender, and adopts credit rate limitation to prevent data congestion in the reverse path; when congestion occurs in the credit / data queue, the switch marks the ECN label for the credit / packet, and when no congestion occurs in the credit / data queue, the switch does not mark the ECN label for the credit / packet; according to the results of marking the ECN label, including whether the credit / packet carries the ECN label or not, the credit rate is dynamically adjusted.
[0035] Next, in order to facilitate the analysis of the process of using the ECN-based credit rate control mechanism to achieve congestion control in the data center network, taking the network architecture with two switches as an example, as Figure 4 shown, the behavior performance of each position in this process is described, specifically including: 1) Behavior of CP (switch): In addition, each switch port has two queues, namely the data queue and the credit queue. Since there is a one-to-one correspondence between packets and credits, the data reception rate and the credit transmission rate of the switch port are linearly dependent. To make full use of the link bandwidth and avoid congestion, the switch limits the link bandwidth ratio of the credit rate to , for example, . The switch also needs to support symmetric routing to ensure that packets can be transmitted through the reverse path corresponding to the corresponding credit.
[0036] In the switch, the approach of the present invention mainly acts on the credit queue and the data queue. As Figure 4As shown in the first switch, the excessive credits will accumulate in the credit queue, and each credit queue maintains a fixed threshold. . If the queue length exceeds , the credit processor in the first switch will mark the credits sent by the switch as ECN tags through the credit processing program ( ). In addition, when the credit queue overflows, the credits will be discarded. It is necessary to design a credit retransmission strategy to retransmit the discarded credits by the receiving end according to the credit sequence number. To ensure the visibility of data queue congestion, an ECN tag marking mechanism for the data queue is introduced to indicate the occurrence of credit congestion; when carrying the ECN tag, it means congestion occurs, and when not carrying the ECN tag, it means congestion does not occur. When the data queue length exceeds the threshold , as Figure 4 shown in the second switch, the data processor will mark the transmitted data packets with ECN tags through the data processing program ( ). Otherwise, when the data queue overflows, the switch will discard the data packets. It is necessary to formulate a data retransmission strategy to ensure that the sending end will retransmit the data according to the data sequence number to reduce the data packet loss rate.
[0037] The present invention uses the same ECN feedback signal for credit congestion and data congestion to simplify the protocol design and reduce the amount of congestion feedback. On the one hand, whether congestion occurs in the credit queue or the data queue, the credit rate must be reduced (carrying the ECN tag). On the other hand, the credit rate can only increase when there is no congestion in both the credit queue and the data queue (not carrying the ECN tag). Therefore, only when ECN is efficient enough can the receiving end make a rate adjustment decision, that is, increase or decrease the rate.
[0038] 2) NP (sender) behavior: The sender combines two sender control programs, namely the credit self-feedback program ( ) and the credit control data program ( ) to avoid excessive congestion response. During normal transmission, the data packets will be scheduled and sent one by one together with the credits. Since both data queue congestion and data reception rate are important references for the receiving end to adjust the credit rate, the credit transmission rate already contains the feedback of data path congestion. Therefore, sending data at the credit reception rate is sufficient to effectively handle the congestion of the data queue.
[0039] When the sender receives a credit, it first records its ECN label and the credit sequence number in the corresponding predetermined data packet. In addition, the sender also assigns an incremental data sequence number to the corresponding data packet. After recording the information, the sender sends the data packet to the receiver. The credit sequence number will be used for credit retransmission, while the data sequence number will be used for data retransmission.
[0040] To prevent bandwidth waste during the first RTT transmission, the present invention needs to additionally design a speculative probing mechanism such that when a new flow arrives and needs to be transmitted, the sender directly sends probing packets equal to the bandwidth-delay product (BDP) to the receiver without performing credit control. The probing packets are different from normal data packets and are another type of data packet (or control message) designed for monitoring the network state. They carry the network metric information required for the receiver to generate and send credits, but do not carry the normal data of network users.
[0041] 3) RP (Receiver) Behavior: When the receiver obtains a credit request (the first probing packet), it starts to send credits to the sender. To prevent credit waste, the total credit amount is set according to the size of the flow. Specifically, the receiver will send as many credits as the sender needs to send data.
[0042] Let be the credit rate at which the receiver sends credits to the sender. In the initial stage, the credit rate is set to the maximum value in order to quickly reach the maximum throughput. Subsequently, is dynamically adjusted periodically according to the ECN-based credit rate control algorithm. As shown in the receiver in Figure 4 , the credit processor will perform dynamic adjustment and update of the credit rate according to the data self-feedback program ( ) and the data feedback credit program ( ) provided by the data processor.
[0043] Set the update period for adjustment to be , set the result variable for marking the ECN label within the update period, which is used to determine whether to mark the ECN label, that is, whether the credit / data packet carries the ECN label. For example, if , it means carrying the ECN label, that is, trust congestion occurs; if , it means not carrying the ECN label, that is, credit congestion does not occur; let be the link bandwidth of the network. For a normal credit control data packet, the receiver first reads the ECN label information carried in the packet. If the data packet carries the ECN label, the value within the period will be set to True, indicating that congestion occurred during this period. On the contrary, if no data packet carries an ECN label within a period, the receiving end will consider that no congestion occurred in the past period. Then, once the adjustment time arrives, the receiving end will check the value of, and adjust the credit rate based on this. If congestion occurs, the receiving end will reduce the credit rate to a slightly lower value than the current data reception rate . For example, take , where is the deceleration ratio of the credit rate; specifically depending on the minimum weight factor , is the weight factor the minimum value of, and the weight factor is used to balance the bandwidth allocation of different flows. Otherwise, if no congestion occurs, the maximum weight factor is used to dynamically adjust the weight and gradually and actively increase the credit rate.
[0044] Specifically, the ECN-based credit rate control algorithm: Step 1.1, initialize parameters: Set the update period for dynamically adjusting the credit rate ; Set the initial weight factor ; Obtain the initial ; Obtain the link bandwidth of the network ; Set the credit rate to take the maximum value , where the maximum value is assigned as: ; Set the data reception rate ; Set the set of flows ; Obtain the minimum weight factor , used to control the decrease amplitude of the credit rate; Step 1.2, for each flow , every time, repeat the following steps: According to the value of, judge whether credit congestion occurs in the current period; If credit congestion occurs, reduce the credit rate and assign a value to the credit rate: ; And use the minimum weight factor to reset the weight: ; If trust congestion does not occur, use the weight to perform a smooth assignment to the credit rate assignment: ; And dynamically adjust the weight: , where is the maximum weight factor, used to control the recovery speed; Update value; In exclude the ones that have completed execution ; Step 1.3, repeat the termination condition of Step 1.2: ; Output , weight .
[0045] The above ECN-based credit rate control algorithm is the same as the data rate adjustment algorithm in PCN, and its efficiency is high enough to converge to fairness quickly.
[0046] When using the above ECN-based credit rate control mechanism to achieve network congestion control, since credit queuing will greatly reduce the overhead of the switch buffer compared to data queuing, the advantage of credit queuing in rate control is obvious. Consider the overhead comparison of credit congestion control and data congestion control in the same network, that is, using the same link bandwidth and the same basic delay . Let the credit size be , and the packet size be ; Therefore, when the credit queue is not released, the credit rate limit is , while the data rate is . It is easy to find that the end-to-end transmission time of both credit and data is . As is well known, the reaction time determines the queue accumulation in ECN-based congestion control. Note that the accumulated data volume is equal to the link bandwidth multiplied by the time. Since the reaction times of credit congestion and data congestion are the same, and the credit rate is much lower than the data rate, the accumulation speed of the data queue is much faster than that of the credit queue.
[0047] In addition, credit queuing rarely affects data transmission delay. Data transmission delay is mainly determined by data queuing delay and propagation delay. Since propagation delay depends on network infrastructure, credit queuing has no impact on it. In addition, data queuing is determined by the credit transmission rate of the reverse path, rather than by credit queuing.
[0048] It can be seen that the present invention has obtained significant benefits by applying ECN-based rate control to credit. In this process, in order to reduce bandwidth waste and unnecessary information loss during transmission, it is also necessary to design a speculative probing mechanism and formulate a reasonable retransmission strategy for data and credit.
[0049] (2) Design a speculative probing mechanism to prevent bandwidth waste in the first RTT transmission of flows during congestion.
[0050] In the process of implementing network congestion control using the ECN-based credit rate control mechanism described above, it was found that a speculative probing mechanism needs to be designed to prevent waste of bandwidth during the first RTT transmission round-trip of a flow. Since the solution adopted in the present invention eliminates the credit reservation process, the sender sends probe packets at the line rate during the first round-trip, and is not credit-driven, while normal data packets are credit-driven.
[0051] Probe packets are not credit-controlled, which may lead to data congestion. In addition, when severe congestion occurs, they also cause packet loss. Packet loss has many impacts on network performance. Packet loss is harmful to throughput because the bandwidth occupied by the lost packets is ultimately not used for data transmission. In addition, the retransmission delay of data packets affects the flow completion time (FCT).
[0052] When packet loss occurs in the last part of a flow and the corresponding data volume is equal to the DDP (Deadline-Driven Protocol) data (the data volume threshold set by the protocol), the impact on the flow completion time (FCT) may be more severe. When a data packet is lost, the sender needs at least one round-trip delay (RTT) to make a retransmission decision. An RTT represents the time from sending the data packet to receiving feedback, which the present invention refers to as the reaction RTT. It can be assumed that when packet loss in the data flow occurs before the last DDP data, the data flow will continue to be transmitted normally during the reaction RTT. In this case, the FCT only increases by , where is the packet size, is the data transmission rate. However, when the lost data packet is located in the last DDP data of the flow, the data flow will be transmitted normally before the sender receives the feedback and retransmits the lost data. Therefore, the bandwidth will be wasted and the FCT will increase by , where is the remaining traffic size after packet loss.
[0053] To mitigate the impact of data loss on the flow completion time (FCT), the goal of the present invention is to preferentially discard normal data packets rather than the last DDP packet (the packet carrying DDP data) in the case of buffer overflow. Here, a priority queue is used to distinguish between normal data packets and the last DDP packet, and the last DDP packet is set to the highest priority to prevent its loss as much as possible.
[0054] (3) Formulate data retransmission strategies and credit retransmission strategies to address data loss and credit loss during the congestion control process.
[0055] 1) Develop a data retransmission strategy based on the continuity of data sequence numbers to address data loss caused by data queue overflow.
[0056] Under normal circumstances, if each data packet is strictly credit-driven, the data queue is usually short and limited. However, when the aforementioned speculative probing mechanism is adopted, the data queue may overflow, resulting in data packet loss. Therefore, the present invention develops a data loss recovery mechanism.
[0057] The data retransmission strategy is based on the continuity of data sequence numbers. The receiving end determines the lost data packets by calculating the difference between the data sequence numbers carried by two consecutively arriving data packets. When data loss is detected, the receiving end records the data sequence number of the lost data packet in the next credit sent. When the sending end receives the credit carrying the data sequence number, the sending end retransmits the corresponding data packet and does not send any other data packets driven by this credit.
[0058] 2) Develop a credit retransmission strategy based on the continuity of credit sequence numbers to address credit loss caused by credit queue overflow under severe credit congestion.
[0059] The credit queue can buffer thousands of credits with only dozens of kilobytes. Under normal circumstances, the size of the credit queue is sufficient for credit congestion control. However, under extreme traffic burst conditions, credits may be lost due to credit queue overflow under congestion. In addition, the total number of credits sent is the same as the total number of data packets. Therefore, when credits are accidentally lost, data transmission cannot be completed. To solve the problem of credit loss, the present invention develops a credit loss recovery strategy, that is, a credit retransmission strategy.
[0060] Each credit generated by the receiving end and transmitted to the sending end carries an incrementing credit sequence number. When the sending end receives a credit, it records the credit sequence number in the corresponding data packet sent to the destination. In this way, the sequence numbers of the non-lost credits are all fed back to the receiving end. Since the credit sequence numbers sent by the receiving end are consecutive, the receiving end can easily determine the number of lost credits by calculating the difference between the credit sequence numbers of two consecutively arriving data packets. After determining the number of credits to be retransmitted, the receiving end adds this number to the total number of credits already sent. In this way, it can be ensured that all data packets can be driven by credits in a one-to-one manner.
[0061] Note that, different from data retransmission, it is not necessary to retransmit the original lost credits because the data packets related to the lost credits can be driven by any other credits belonging to the same flow. In addition, it is not necessary to resend the credits immediately because uncontrolled credits will interfere with the credit rate control mechanism.
[0062] (4) Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limiting and switch-side credit selective discard.
[0063] Through the congestion control of the data and credit coupling mechanism, keep the low queuing of normal data packets (non-probe packets), which limits the application of traditional flow scheduling mechanisms. Traditional flow scheduling mechanisms focus on adjusting the queuing of data transmission order in the switch queue. Therefore, for the flow deadline, the present invention develops and designs a credit-driven flow scheduling mechanism to meet the deadline requirements of AI applications, including further reducing the flow completion time and deadline miss rate. This mechanism scheme achieves the target requirements through two key components, namely receiver credit rate limiting and switch-side credit selective discard. In this embodiment, taking Figure 5 as an example, illustrate the working process of the credit-driven flow scheduling mechanism, as Figure 5 shown, flow is transmitted from the first sender to the receiver, and flow is transmitted from the second sender to the receiver. The deadline of flow is less than that of flow , so the receiver gives priority to sending the credit of flow , and when the credit queue length exceeds , the first switch will selectively discard the credit of flow . The credit-driven flow scheduling mechanism specifically includes: 1) Receiver credit rate limiting: First, propose a concept called Deadline Reaching Remaining Time (DRRT), which represents the remaining time for a flow to reach the deadline, and use the set to represent the DRRT value of the flow. Let , where represents the credit rate limit of the rd flow . When the receiver receives multiple flows simultaneously, it will first meet the requirements of the shortest DRRT flow, that is, release the credit rate limit of the flow to the maximum credit rate , and mark a high-priority label for the corresponding data packet (the corresponding credit is also marked as high-priority), and the switch can recognize this label. Set as a temporary credit rate limit variable for dynamically adjusting the rate of the flow, then the credit rate limits of other data flows will be set to , so as to accelerate the transmission of the shortest DRRT flow. Note that the used to adjust the credit rate in the ECN-based credit rate control algorithm should be assigned to 。
[0064] Specifically, the following credit-driven flow scheduling algorithm is adopted to implement the credit rate limit at the receiver: Step 2.1, input parameters: the set of flows , the set of DRRT values of the flows , and the set of credit rate limits of the flows ; set the link bandwidth ratio of the credit rate ; obtain the link bandwidth of the network ; Step 2.2, initialize parameters: Set the maximum credit rate: ; Initialize the current credit rate: ; Initialize the temporary credit rate limit: ; Step 2.3, repeatedly execute the main loop step: Each time a data packet arrives, check the flow to which the data packet belongs and process the loop process: Step 2.3.1, if the data packet belongs to the current flow , , then execute in the current flow: is the minimum value in the set , then Set the credit rate limit of the th flow to the maximum credit rate: ; Update the temporary credit rate limit: ; Mark high priority; else if is not the minimum value in the set : Set the credit rate limit of the th flow to the temporary credit rate limit: ; ; Execute the end of , and remove it in ; end if If , end the loop; Step 2.3.2, if the data packet does not belong to any current flow , , process the addition of a new flow : Set the credit rate limit of the new flow to the temporary credit rate limit: ; Update the current credit rate: ; Step 2.4, the condition for ending the execution of the main loop: , output , , , high-priority flag.
[0065] 2) Credit selective discard at the switch side: The present invention sets a threshold for the credit queue of the switch , to achieve credit selective discard. Once the length of the switch's credit queue exceeds , the switch will discard the credits without high-priority tags. In this way, the credits with high priority can be transmitted faster without modifying the switch hardware. Therefore, the transmission of the shortest DRRT flow can be accelerated at the switch side, greatly improving the data throughput and facilitating high-throughput network data transmission.
[0066] The data and credit coupling mechanism adopted by the present invention designs different functional mechanisms and strategies in the data and credit coupling mechanism through dynamically binding data flow characteristics such as deadlines, priorities, real-time requirements, etc., and credit allocation strategies, and works together to achieve congestion control and efficient scheduling of the data center network.
[0067] The key innovation points designed by the present invention mainly include: 1) Introduced the DCEF framework and proposed a new data and credit coupling congestion control scheme based on this framework. By dynamically synchronizing the credit transmission rate with the data reception rate, accurate real-time credit rate adaptation is achieved, and at the same time, one-to-one credit data transmission is enforced to effectively relieve congestion; 2) Developed a credit rate control mechanism based on ECN to achieve full-loop congestion control and solve the problem of credit waste; 3) Utilized a credit-driven flow scheduling mechanism to preferentially process traffic according to the urgency of the deadline, further reducing the flow completion time (FCT) and deadline miss rate (DMR), and enabling high-throughput network data transmission to meet the performance requirements of AI applications for the data center network.
[0068] In another embodiment of the present invention, a data center network congestion control device based on data and credit coupling is protected. The device is used to perform data center network congestion control based on data and credit coupling by using the foregoing method. The device includes the following modules: The first module is used to obtain the structure of the data center network, including switches, senders, and receivers. The second module is used to develop an ECN-based credit rate control mechanism for managing data congestion and credit congestion and performing congestion control for the entire loop of the data center network based on the principles of designing a data-credit coupling mechanism in the data-credit and end-to-end feedback framework. The second module further includes: Sub-module 1 is used to set that the switch is the congestion point location, the sender is the notification point location, and the receiver is the reaction point location. Sub-module 2 is used to enable the receiver to send credits through the reverse path of credit transmission, pull data packets from the sender, and adopt credit rate limiting to prevent data congestion on the reverse path. Sub-module 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / data packet, and when the credit / data queue is not congested, it does not mark the ECN label for the credit / data packet. Sub-module 4 is used to dynamically adjust the credit rate using an ECN-based credit rate control algorithm according to the results of marking the ECN label, including whether the credit / data packet carries the ECN label or not.
[0069] In another embodiment of the present invention, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the foregoing data center network congestion control method based on data-credit coupling are implemented. This computer device can be a server. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is used to communicate with an external terminal through a network connection.
[0070] On the other hand, in an embodiment of the present invention, a storage medium is protected, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the foregoing data center network congestion control method based on data-credit coupling are implemented.
[0071] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0072] Matters not described in this invention are well-known technologies.
[0073] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0074] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A data center network congestion control method based on data and credit coupling, characterized in that: include: Obtain the structure of the data center network, including switches, transmitters, and receivers; Based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, an ECN-based credit rate control mechanism is developed to manage data congestion and credit congestion and perform congestion control on the entire loop of the data center network; the ECN-based credit rate control mechanism includes: Set the switch to the congestion point, the sender to the notification point, and the receiver to the reaction point. The receiver sends credits through the reverse path of the credit transmission, pulls data packets from the sender, and applies credit rate limiting to prevent data congestion on the reverse path; The same ECN feedback signal is used for credit congestion and data congestion. When the credit / data queue is congested, the switch marks the credit / data packet with an ECN tag. When the credit / data queue is not congested, the switch does not mark the credit / data packet with an ECN tag. According to the result of marking the ECN label, including whether the credit / data packet carries the ECN label or does not carry the ECN label, the credit rate is dynamically adjusted using an ECN-based credit rate control algorithm.
2. The data center network congestion control method based on data and credit coupling according to claim 1 is characterized in that: The data center network includes at least two switches, two sending ends and two receiving ends; The ECN is a network protocol method of explicit congestion notification, which allows a switch to explicitly notify a sender or receiver of the congestion by marking a header field of a data packet when congestion occurs.
3. The data center network congestion control method based on data and credit coupling according to claim 2 is characterized in that: The congestion control of the entire loop of the data center network includes: Design a speculative detection mechanism to prevent bandwidth waste in the first RTT transmission of the flow during congestion control; the RTT is the round-trip delay of the flow, which means the time from the sender sending a data packet to the sender receiving feedback; Formulate data retransmission strategy and credit retransmission strategy to deal with data loss and credit loss during congestion control; A credit-driven flow scheduling mechanism is designed to achieve flow scheduling control through credit rate limitation at the receiving end and selective credit discard at the switch end.
4. The data center network congestion control method based on data and credit coupling according to claim 3 is characterized in that: Let the result variable marked with ECN label be ,like , indicating that the credit / packet carries the ECN label, trust congestion occurs, if , indicating that the credit / data packet does not carry an ECN label, trust congestion does not occur; set the weight factor , used to balance bandwidth allocation among different flows; The method of dynamically adjusting the credit rate using the ECN-based credit rate control algorithm includes: Step 110, initialization parameters: set the set of flows ; Set the link bandwidth ratio of the credit rate The value of; Get the link bandwidth of the network ; Set credit rate Take the maximum value , where the maximum value is ; Set the credit rate to dynamically adjust the update cycle ; Set the initial weight factor ; Get the initial ; Get the minimum weight factor , used to control the decrease of credit rate; Step 120, each update cycle Internal, based on The value of determines whether credit congestion occurs and dynamically adjusts the credit rate: like , when credit congestion occurs, the credit rate is reduced and the credit rate is assigned: ,in, is the credit rate reduction ratio, using the minimum weight factor , reset the weight factor ; like , credit congestion does not occur, then the weight factor is used Smooth assignment of credit rate values: ; Dynamically adjust weight factors: ,in, is the maximum weight factor, used to control the recovery speed of the credit rate; renew value; exist Eliminate the execution completion ; Step 130, repeat step 120 until the termination condition is met , output .
5. The data center network congestion control method based on data and credit coupling according to claim 4 is characterized in that: The speculative detection mechanism includes: A priority queue is used to distinguish between a common data packet and a last DDP data packet, and the last DDP data packet is set to the highest priority; the DDP data packet refers to a data packet of DDP data, and the DDP data refers to a data volume threshold set by a deadline-driven protocol; When the network buffer data queue overflows, normal data packets are discarded first and DDP data packets are retained.
6. The data center network congestion control method based on data and credit coupling according to claim 5, characterized in that: The process of formulating a data retransmission strategy and a credit retransmission strategy to cope with data loss and credit loss during congestion control includes: Formulate a data retransmission strategy based on the continuity of data sequence numbers to deal with data loss caused by data queue overflow, including: the receiving end determines the lost data packet by calculating the difference between the data sequence numbers carried by two consecutive arriving data packets; when the data packet loss is detected, the receiving end records the data sequence number of the lost data packet in the next sent credit; when the sending end receives the credit carrying the data sequence number, it retransmits the lost data packet and no longer sends other data packets driven by the same credit; Formulate a credit retransmission strategy based on the continuity of credit sequence numbers to deal with credit loss caused by credit queue overflow under severe credit congestion, including: the receiving end determines the number of lost credits by calculating the difference in credit sequence numbers between two consecutive data packets; determines the number of credits that need to be retransmitted from the number of lost credits, and the receiving end adds the number of credits that need to be retransmitted to the total number of credits sent, ensuring that all data packets can be driven by credits in a one-to-one manner.
7. The data center network congestion control method based on data and credit coupling according to claim 6, characterized in that: The credit-driven flow scheduling mechanism includes: When the receiving end receives multiple flows at the same time, it will first meet the needs of the flow with the shortest DRRT, and mark a high priority label for the data packet of the flow with the shortest DRRT, and the corresponding credit is also marked with a high priority label; the DRRT represents the remaining time for the flow to reach the deadline, and the shortest DRRT flow refers to the flow whose DRRT value reaches the minimum among the multiple flows received at the same time; Adopting credit-driven flow scheduling algorithm to achieve credit rate limitation at the receiving end; Speed up the transmission of the shortest DRRT flow on the switch side through selective discard of credits on the switch side, including setting thresholds for the credit queues on the switch , when the credit queue length exceeds the threshold When a high priority tag is not tagged, the switch chooses to discard the credits that are not tagged with the high priority tag.
8. The data center network congestion control method based on data and credit coupling according to claim 7, characterized in that: The credit-driven flow scheduling algorithm includes: Step 210, input parameters: set of flows , the set of DRRT values of the flow , and the credit rate limit set for the flow ; Set the link bandwidth ratio of the credit rate The value of; Get the link bandwidth of the network ; Step 220, initialization parameters: To set the maximum credit rate: ; Initialization of current credit rate: ; Initialization of temporary credit rate limit: ; Step 230, repeat the main loop steps: Step 231, each time a data packet arrives, check the flow to which the data packet belongs; Step 232, if the data packet belongs to the current flow , , then the loop is executed in the current stream: when Is a collection The minimum value in The credit rate of a flow is limited to the maximum credit rate: ; Update temporary credit rate limit: ;Mark high priority tags; when Not a collection The minimum value in Stream The credit rate limit is a temporary credit rate limit: ; End the current flow loop, in Eliminate the execution end , jump back to step 231; If the packet does not belong to the current flow , jump out of the loop of step 232 and execute step 233; Step 233, if the data packet does not belong to the current flow , , handle the addition of new streams : Set the credit rate limit for new flows to the temporary credit rate limit: ; Update the current credit rate limit: ; Step 240, repeat step 230 until the termination condition is met , output , , , and high priority tags.
9. The data center network congestion control method based on data and credit coupling according to claim 8, characterized in that: The link bandwidth ratio is .
10. A data center network congestion control device based on data and credit coupling, characterized in that: The device uses the steps of the method according to claim 1 to perform data center network congestion control based on data and credit coupling, and the device includes the following modules: The first module is used to obtain the structure of the data center network, including switches, transmitters, and receivers; The second module is used to develop an ECN-based credit rate control mechanism based on the principle of designing a data-credit coupling mechanism in the data-credit and end-to-end feedback framework to manage data congestion and credit congestion and perform congestion control on the entire loop of the data center network. The second module also includes: Submodule 1 is used to set the switch as the congestion point, the sender as the notification point, and the receiver as the reaction point; Submodule 2 is used to enable the receiving end to send credits through the reverse path of the credit transmission, pull data packets from the sending end, and use credit rate limiting to prevent data congestion in the reverse path; Submodule 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the credit / data packet with an ECN label, and when the credit / data queue is not congested, the switch does not mark the credit / data packet with an ECN label; Submodule 4 is used to dynamically adjust the credit rate by using an ECN-based credit rate control algorithm according to the result of marking the ECN label, including whether the credit / data packet carries the ECN label or does not carry the ECN label.
Citation Information
Patent Citations
Data center network congestion control method based on credit and reaction types
CN112468405A
Method for Transmission Control Protocol window size control in Asynchronous Transfer Mode
KR1020000073790A
Technique for providing end-to-end congestion control with no feedback from a lossless network
US7035220B1