Data Center Network Congestion Control Method and Device Based on the Coupling of Data and Credit
By introducing DCEF framework and ECN feedback signals into the data center network, a congestion control method that couples data with credit is designed, which solves the problems of excessive response delay and credit loss, and realizes efficient data transmission and flow scheduling, meeting the performance needs of AI applications.
Patent Information
- Application Number
- CN202510561461.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The congestion control method of existing data center networks is too late for the response time, and it is difficult for the sending end to suppress data injection rates in a timely manner, and the throughput decline caused by traffic and credit loss. Traditional methods cannot meet the needs of low latency, high throughput and stream deadlines.
The congestion control method based on data and credit coupling is adopted, and the data and credit coupling mechanism is designed through the DCEF framework, combined with the ECN feedback signal, dynamic credit rate control and flow scheduling are realized to prevent credit waste and optimize data transmission.
It improves the response speed of congestion control, reduces resource waste, meets the low latency, high throughput and streaming deadline requirements of the data center network, and improves network performance.
Smart Images

Figure CN120090974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital information transmission, and particularly relates to a data center network congestion control method and device based on data and credit coupling. Background Art
[0002] With the continuous development of artificial intelligence (AI) and the Internet of Things (IoT), the amount of data generated by real-time AI applications has increased sharply, and this data is usually transmitted to the data center for storage and analysis. To ensure the performance requirements of such applications, the data center must efficiently transmit data streams with low latency, high throughput, and deadline constraints.
[0003] Driven by the growing performance requirements of diverse artificial intelligence applications, the link bandwidth required by the data center has rapidly increased to 100 Gbps and 400 Gbps, and this trend continues. At the same time, network traffic exhibits complex and time-varying characteristics, including both long-period continuous flows (data streams) and short-time bursty flows, which poses a huge challenge to simultaneously achieving high throughput, low latency, and good stability. Once congestion occurs, the sudden packet queuing and unpredictable packet loss will cause a sharp decline in network performance, thus seriously affecting the user experience.
[0004] Congestion control is crucial for the overall performance of data transmission. ECN (Explicit Congestion Notification) is a network protocol that allows network devices, such as routers or switches, to explicitly notify the sender or receiver of congestion by marking the header field of a data packet (instead of directly discarding the data packet) when congestion occurs. For the congestion problem in data center networks, traditional congestion control schemes based on the try-and-backoff mechanism of ECN, RTT (Round-Trip Time), or INT (In-band Network Telemetry) feedback usually require at least one RTT measurement period to respond to congestion. Therefore, when congestion occurs, the response delay of such traditional methods is too long, and it is difficult for the sender to timely suppress the data injection rate, especially in high-speed data centers. If the data stream is transmitted within an extremely short RTT, the problem is particularly prominent. In recent years, receiver-driven credit-based congestion control schemes have gradually emerged, which can precisely regulate data transmission and prevent congestion from occurring at the source. In the credit-based congestion control mechanism, the receiver issues a clear credit limit according to its own receiving ability, and controls the data transmission of the sender in a one-to-one credit and data manner, realizing fine-grained packet-level feedback and accurately reflecting the network state. However, the existing credit-based mechanisms still face problems such as unstable data queuing and throughput degradation caused by bursty data traffic and credit loss. Summary of the Invention
[0005] In view of this, in order to solve the problems of too long response delay, difficulty for the sender to timely suppress the data injection rate, and throughput degradation caused by traffic and credit loss in the prior art, the present invention provides a congestion control method and device for a data center network based on data and credit coupling. The congestion control method for a data center network based on data and credit coupling can effectively improve the congestion control performance, reduce resource waste, and meet the requirements of data center networks in terms of low latency, high throughput, and flow deadline.
[0006] The present invention provides a congestion control method for a data center network based on data and credit coupling, and the method includes:
[0007] Obtain the structure of the data center network, including switches, senders, and receivers;
[0008] Based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, develop a credit rate control mechanism based on ECN for managing data congestion and credit congestion, and perform congestion control for the entire loop of the data center network; the credit rate control mechanism based on ECN includes:
[0009] Set the switch as the congestion point location, the sender as the notification point location, and the receiver as the response point location;
[0010] The receiver sends credits through the reverse path of credit transfer, pulls data packets from the sender, and uses credit rate limiting to prevent data congestion on the reverse path;
[0011] Use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / data packet, and when the credit / data queue is not congested, it does not mark the ECN label for the credit / data packet;
[0012] According to the result of marking the ECN label, including whether the credit / data packet carries the ECN label or not, use the ECN-based credit rate control algorithm to dynamically adjust the credit rate.
[0013] Furthermore, the congestion control of the full loop of the data center network includes:
[0014] Design a speculative probing mechanism to prevent bandwidth waste in the first RTT transmission of the flow during the congestion control process; the RTT is the round-trip delay of the flow, representing the time from the sender sending the data packet to the sender receiving the feedback;
[0015] Formulate data retransmission policies and credit retransmission policies to deal with data loss and credit loss during the congestion control process;
[0016] Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limiting and switch-side credit selective discard.
[0017] In addition, the present invention also provides a data center network congestion control device based on data and credit coupling. The device uses the steps of the foregoing method to perform data center network congestion control based on data and credit coupling. The device includes the following modules:
[0018] The first module is used to obtain the structure of the data center network, including switches, senders, and receivers;
[0019] The second module is used to develop an ECN-based credit rate control mechanism based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, for managing data congestion and credit congestion, and performing congestion control for the full loop of the data center network;
[0020] The second module further includes:
[0021] Sub-module 1 is used to set the switch as the congestion point location, the sender as the notification point location, and the receiver as the response point location;
[0022] Sub-module 2 is used to enable the receiving end to send credits through the reverse path of credit transmission, pull data packets from the sending end, and adopt credit rate limiting to prevent data congestion in the reverse path;
[0023] Sub-module 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when congestion occurs in the credit / data queue, the switch marks the ECN label for the credit / data packet, and when congestion does not occur in the credit / data queue, the ECN label is not marked for the credit / data packet;
[0024] Sub-module 4 is used to dynamically adjust the credit rate by adopting an ECN-based credit rate control algorithm according to the result of marking the ECN label, including whether the credit / data packet carries the ECN label or not.
[0025] In summary, the present invention provides a data center network congestion control method and device based on data and credit coupling. Compared with the prior art, the technical solution of the present invention has the following advantages:
[0026] 1) The DCEF framework is introduced, and a new data and credit coupling congestion control scheme is proposed based on this framework. By dynamically synchronizing the credit transmission rate and the data reception rate, accurate real-time credit rate adaptation is achieved, and at the same time, one-to-one credit data transmission is enforced, effectively alleviating congestion.
[0027] 2) An ECN-based credit rate control mechanism is developed to achieve full-loop congestion control and solve the problem of credit waste.
[0028] 3) Using a credit-driven flow scheduling mechanism, the flow completion time (FCT) and deadline miss rate (DMR) of congestion control are further reduced, and high-throughput network data transmission can be achieved to meet the performance requirements of AI applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic flow chart of the data center network congestion control method based on data and credit coupling provided by the first embodiment of the present invention;
[0030] Figure 2 It is a schematic diagram of bandwidth waste caused by credit loss in an embodiment of the present invention;
[0031] Figure 3 It is a schematic diagram for comparing the transmission situations of flows with different traffic sizes in an embodiment of the present invention before and after flow scheduling. Among them, are data flows with different traffic sizes, is the unit time, is the unit bandwidth, Figure 3(a) shows the transmission situation without flow scheduling, Figure 3 (b) shows the transmission situation after flow scheduling;
[0032] Figure 4 is a schematic diagram of the ECN-based credit rate control framework and control process in an embodiment of the present invention. Among them, is the threshold of the credit queue length, is the threshold of the data queue length. CP represents the congestion point, NP represents the notification point, and RP represents the reaction point;
[0033] Figure 5 is a schematic diagram of the congestion control credit-driven flow scheduling process provided in an embodiment of the present invention. Among them, is the threshold of the credit queue length. Specific implementation manners
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] In artificial intelligence (AI) application scenarios, for example, in the AIoT (Artificial Intelligence of Things, a scenario where artificial intelligence technology is deeply integrated with the Internet of Things) scenario, to meet the high-performance requirements of artificial intelligence (AI) applications for large data volumes, low latency, high reliability, etc. in the data center network, the present invention proposes a congestion control scheme for the data center. The scheme is based on the DCEF framework, deeply explores the tight coupling mechanism between the data plane and the credit plane in congestion control design, and realizes the coordinated regulation of data and credit in a complete control loop. In addition, the present invention effectively responds to the requirements of the flow deadline by introducing a credit-driven flow scheduling mechanism.
[0036] In one embodiment, the present invention proposes a method for congestion control of a data center network based on the coupling of data and credit, as Figure 1 shown, the method includes:
[0037] Obtain the structure of the data center network, including switches, senders, and receivers;
[0038] Based on the principles of designing the data and credit coupling mechanism in the data and credit and end-to-end feedback framework, develop an ECN-based credit rate control mechanism for managing data congestion and credit congestion, and perform congestion control for the entire loop of the data center network; the ECN-based credit rate control mechanism includes:
[0039] Set the switch as the congestion point location, the sender as the notification point location, and the receiver as the reaction point location;
[0040] The receiver sends credits through the reverse path of credit transfer, pulls packets from the sender, and adopts credit rate limiting to prevent data congestion in the reverse path;
[0041] Use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / packet, and when the credit / data queue is not congested, it does not mark the ECN label for the credit / packet;
[0042] According to the results of marking the ECN label, including whether the credit / packet carries the ECN label or not, adopt an ECN-based credit rate control algorithm to dynamically adjust the credit rate.
[0043] Furthermore, the congestion control of the full loop of the data center network includes:
[0044] Design a speculative probing mechanism to prevent bandwidth waste in the first RTT transmission of the flow during the congestion control process; the RTT is the round-trip delay of the flow, representing the time from the sender sending a packet to the sender receiving the feedback;
[0045] Formulate data retransmission policies and credit retransmission policies to cope with data loss and credit loss during the congestion control process;
[0046] Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limiting and switch-side credit selective dropping.
[0047] Specifically, the design goal of the present invention is to achieve a shorter flow completion time (FCT), high throughput, short queuing time, less packet loss, and a low deadline miss rate (DMR). This goal brings many challenges. The most difficult challenge is that when the credit is discarded by the switch due to credit queue congestion, the credit will be wasted, thereby reducing the effective throughput of the flow, and it is not easy to meet the performance requirements of real-time AI applications. Therefore, based on the data and credit coupling design principle in the DCEF framework, the present invention designs a new data and credit coupling mechanism to solve the above challenges.
[0048] The DCEF (Data–Credit and End-to-end Feedback) framework is an intelligent congestion control technology for modern data center networks. Its core idea is to achieve high-throughput and low-latency data center network resource management through the dynamic coupling of data and credit, and the closed-loop optimization of end-to-end feedback. Compared with traditional reactive methods based on coarse-grained feedback, the participation of the credit plane can improve congestion control performance in terms of convergence and accuracy, specifically manifested in:
[0049] 1) Credit-based congestion control performs better. In traditional round-trip time (RTT)-based congestion control, the sender considers the transmission delay of the entire path, resulting in inefficient congestion management in reducing data queuing to achieve high throughput. When the queue length exceeds a certain threshold, the ECN signal provides feedback on data congestion, but this is still coarse-grained information and cannot help the sender precisely adjust data transmission. INT provides specific link load information, but the sender needs time to limit the data injection rate.
[0050] The credit mechanism provides the sender with precise packet-level network status information, enabling more effective congestion prevention. In credit-based congestion control, the receiver sends credits to the sender based on its receiving capacity, and the sender sends a data packet after receiving each credit. In this way, if the credit transmission rate is well regulated, data congestion can be prevented from occurring.
[0051] 2) The full closed-loop congestion control based on the DCEF framework helps achieve near-optimal data transmission and improves the integrity of control using the full control loop.
[0052] Each processing process in the data processor and the credit processor has its unique function, and the congestion control scheme should not bypass these processes. Data processing is the key to notifying data queue congestion, and methods that ignore this process are inefficient in managing data congestion within the network. Credit processing is crucial for credit-based approaches as it can precisely prevent data congestion at the switch, including credit rate limiting and credit congestion signal recording. Similar to controlling the credit rate at the receiver, performing switch credit rate limiting on one side of the link can avoid reverse path data congestion on the other side. In addition, credit congestion information is crucial for the receiver to adjust the credit rate and avoid credit waste.
[0053] Analysis and comparison with the DCEF framework and related existing technologies show that close data credit cooperation improves the accuracy of control, and the mechanisms on the data side and the credit side should have a coupled design. The control from credit to data is the core of credit-based congestion control, directly determining the congestion control efficiency and serving as the basis for the effective operation of any credit-based congestion control scheme. On the other hand, the feedback from data to credit is crucial for effectively preventing data congestion, and credit-based congestion control schemes require it to precisely adjust data transmission in case of unexpected congestion.
[0054] The present invention obtains the key idea for promoting the process of technology development from the above analysis process, that is, high-performance congestion control requires designing a reasonable data and credit coupling mechanism, and in order to achieve network-wide congestion control, each control process should be incorporated into the control design. However, it is challenging to achieve congestion control based on data and credit coupling while meeting the flow deadline by designing a data and credit coupling mechanism, which is mainly reflected in:
[0055] 1) It is necessary to alleviate the bandwidth waste caused by credit loss.
[0056] Credit-based congestion control schemes should prevent credit waste because credit waste has a great negative impact on data throughput. Since each credit processor on the transmission path rate-limits the credit, many credits may be discarded due to speeding. Once a credit is lost at the switch, the previously passed link cannot be fully utilized.
[0057] Taking a specific credit transmission process as an example to illustrate the mechanism of credit loss. For the sake of description, it is assumed here that each switch only allows one credit to pass within a given time. As Figure 2 shown in the credit transmission path, the credit transmission process involves = 3 switches and servers, showing the transmission process of = 4 credits. The server is the terminal node of the transmission, including at least a sending end as the notification point and a receiving end as the reaction point. The credits are numbered in sequence: credit , , where credit successively passes through the 3rd switch and the 2nd switch, but credit and credit are discarded. Then, when credit is allowed to pass, credit is discarded at the 1st switch. In this case, the links between the 2nd switch and the 3rd switch and between the 1st switch and the 2nd switch are not fully utilized because credits are all discarded without corresponding credits successfully reaching the sending end.
[0058] For example, ExpressPass in the existing credit-based congestion control method is a congestion control protocol for data center networks. It dynamically adjusts credit loss through a credit feedback mechanism. However, this loss rate-based algorithm cannot fundamentally reduce credit waste.
[0059] 2) A flow scheduling mechanism that meets the requirements of artificial intelligence applications needs to be designed.
[0060] The diversification of artificial intelligence applications has led to the complexity of traffic performance requirements, which are mainly divided into two types: deadline traffic and non-deadline traffic. Deadline traffic is usually related to tasks with clear completion time limits. In this type of traffic, if the data transmission exceeds the deadline that the user can wait for, it will lose its value, which not only affects the user experience but also wastes bandwidth. For non-deadline traffic, the completion time and throughput of the traffic become the main performance optimization goals.
[0061] The flow scheduling mechanism is the key method to improve the transmission efficiency of AI traffic with different performance requirements. Let's number the flows to be processed in sequence to obtain a set of flows , where is the total number of flows. Specifically, consider four flows with different sizes , which are transmitted through a bottleneck link. The specific flow attributes (flow size and deadline) are shown in Table 1 below:
[0062] Table 1 Example of flow attribute representation
[0063]
[0064] Combined with Table 1, Figure 3 shows the comparison of traffic transmission situations of 4 different-sized flows (data flows) when flow scheduling is not performed and when flow scheduling is performed. Among them, represents the unit time, represents the unit bandwidth. The relationship between the bottleneck link bandwidth and the flow transmission time is as shown in Figure 3 (a) and Figure 3 (b). When flow scheduling is not performed in Figure 3 (a), the transmission time of is , the transmission time of is , the transmission time of is , and the transmission time of is . It can be seen that all flows will miss the deadline, and the deadline miss rate (DMR) is 100%. In contrast, inFigure 3 After flow scheduling in (b), by preferentially transmitting short flows, the time to complete transmission is , the time to complete transmission is , the time to complete transmission is , the time to complete transmission is , and only flow will miss the deadline. It can be seen that flow scheduling reduces the DMR to 25%. Obviously, by scheduling shorter data flows for transmission first, the DMR is reduced from 100% to 25%.
[0065] Traditional flow scheduling methods only schedule data flows in the switch queue. By adopting congestion control with data and credit coupling, the switch is usually in a near-zero queue state. When there are multiple data flows with different performance requirements at the sending end or receiving end that need to be sent or received simultaneously, the performance of the scheduling strategy will be greatly reduced.
[0066] In view of the above challenges, based on the principle of designing a data and credit coupling mechanism in the above DCEF framework, the present invention proposes a new credit control mechanism that uses all congestion feedback from two planes (data plane and credit plane) to perform credit rate control, considers ECN as the congestion signal for both planes, and integrates all control processes related to congestion regulation, including speculative probing and data retransmission, into the ECN-based credit rate control. By buffering excessive credits and using ECN-based rate control to adjust the credit rate of the receiving end, credit waste is prevented; on the other hand, the present invention designs a credit-driven flow scheduling mechanism that considers indirectly adjusting the data reception order of the receiving end by attempting to control the credit rate of the receiving end. In this way, under the condition that the switch does not queue, traffic in the data center network with different performance requirements can be effectively scheduled. Specifically, it includes:
[0067] (1) Developing an ECN-based credit rate control mechanism for managing data congestion and credit congestion to achieve congestion control for the entire loop of the data center network.
[0068] In this embodiment, relying on the notification of the switch and the reaction of the receiving end, such as Figure 4As shown in the figure, the process of the ECN-based credit rate control mechanism performing congestion control involves three locations in the network, namely the congestion point (CP), the notification point (NP), and the reaction point (RP). The congestion point is the location where congestion actually occurs. The notification point is the location where the ECN label is marked. The reaction point is the location that receives the ECN label and adjusts the credit rate. Among them, the CP is a switch. Different from the traditional ECN-based congestion control, here the NP is the sender and the RP is the receiver. Figure 4 In it, the network architecture of the ECN-based credit rate control includes a switch (CP), a sender (NP), and a receiver (RP).
[0069] The ECN-based credit rate control mechanism also includes: the receiver sends credits through the reverse path of credit transmission to pull packets from the sender, and adopts a credit rate limit to prevent data congestion in the reverse path; when the credit / data queue is congested, the switch marks the ECN label for the credit / data packet, and when the credit / data queue is not congested, the ECN label is not marked for the credit / data packet; according to the result of marking the ECN label, including whether the credit / data packet carries the ECN label or not, the credit rate is dynamically adjusted.
[0070] Next, in order to facilitate the analysis of the process of implementing data center network congestion control using the ECN-based credit rate control mechanism, taking the network architecture with two switches as an example, as Figure 4 shown in the figure, the behavior performance of each location in this process is described, specifically including:
[0071] 1) Behavior of CP (switch):
[0072] In addition, each switch port has two queues, namely the data queue and the credit queue. Since there is a one-to-one correspondence between packets and credits, the data reception rate and the credit transmission rate at the switch port are linearly dependent. In order to make full use of the link bandwidth and avoid congestion, the switch limits the link bandwidth ratio of the credit rate to , for example, . The switch also needs to support symmetric routing to ensure that packets can be transmitted through the reverse path corresponding to the corresponding credit.
[0073] In the switch, the practice of the present invention mainly acts on the credit queue and the data queue. As Figure 4 shown in the first switch in the figure, the excessive credits will accumulate in the credit queue, and each credit queue will maintain a fixed threshold . If the queue length exceeds , the credit processor in the first switch will pass through the credit processing program ( The credit sent by the switch will be marked with an ECN label. In addition, when the credit queue overflows, the credit is discarded, and a credit retransmission strategy needs to be designed to retransmit the discarded credit by the receiving end according to the credit sequence number. To ensure the visibility of data queue congestion, an ECN label marking mechanism for the data queue is introduced to indicate the occurrence of credit congestion; when carrying an ECN label, it means congestion occurs, and when not carrying an ECN label, it means congestion does not occur. When the data queue length exceeds the threshold as Figure 4 shown in the second switch in ), the data processor will mark the ECN label for the transmitted data packets through the data processing program (
[0074] ). Otherwise, when the data queue overflows, the switch will discard the data packets, and a data retransmission strategy needs to be formulated to ensure that the sending end will retransmit the data according to the data sequence number, reducing the data packet loss rate.
[0075] 2) NP (sending end) behavior:
[0076] The sending end combines two sending end control programs, namely the credit self-feedback program ( ), and the credit control data program ( ), to avoid excessive congestion response. During normal transmission, data packets are scheduled and sent one by one together with the credit. Since both data queue congestion and data reception rate are important references for the receiving end to adjust the credit rate, the credit transmission rate already contains the feedback of data path congestion. Therefore, sending data at the credit reception rate is sufficient to effectively cope with data queue congestion.
[0077] When the sending end receives a credit, it first records its ECN label and credit sequence number in the corresponding predetermined data packet. In addition, the sending end also assigns an incremental data sequence number to the corresponding data packet. After recording the information, the sending end sends the data packet to the receiving end. The credit sequence number will be used for credit retransmission, and the data sequence number will be used for data retransmission.
[0078] To prevent bandwidth waste during the first RTT transmission, the present invention needs to design an additional speculative probing mechanism. When a new flow arrives and needs to be transmitted, the sender directly sends probing packets equal in number to the bandwidth-delay product (BDP) to the receiver without performing credit control. Probing packets are different from normal data packets and are another type of packet (or control message) designed to monitor the network state. They carry the network metric information required for the receiver to generate and send credits, but do not carry the normal data of network users.
[0079] 3) RP (Receiver) behavior:
[0080] When the receiver obtains a credit request (the first probing packet), it starts sending credits to the sender. To prevent credit waste, the total credit amount is set according to the size of the flow. Specifically, the receiver will send as much credit as the sender needs to send data.
[0081] Let be the credit rate at which the receiver sends credits to the sender. In the initial stage, the credit rate is set to the maximum value to quickly reach the maximum throughput. Subsequently, is dynamically adjusted periodically according to the ECN-based credit rate control algorithm. As shown in the receiver in Figure 4 , the credit processor will perform dynamic adjustment and update of the credit rate according to the data self-feedback program provided by the data processor ( ) and the data feedback credit program ( ).
[0082] Set the update period for adjustment to , and set the result variable that marks the ECN label within the update period, which is used to determine whether to mark the ECN label, that is, whether the credit / data packet carries the ECN label. For example, if , it means carrying the ECN label, that is, trust congestion occurs; if , it means not carrying the ECN label, that is, credit congestion does not occur; let be the link bandwidth of the network. For a normal credit control data packet, the receiver first reads the ECN label information carried in the packet. If the data packet carries the ECN label, the value within the period will be set to True , indicating that congestion has occurred within this period. On the contrary, if no data packet carries the ECN label within a period, the receiver will consider that no congestion has occurred in the past period. Then, once the adjustment time arrives, the receiver will check The value is used to adjust the credit rate accordingly. If congestion occurs, the receiver will reduce the credit rate to a slightly lower value than the current data reception rate. For example, take , where is the deceleration ratio of the credit rate; it specifically depends on the minimum weight factor , is the weight factor 's minimum value. The weight factor is used to balance the bandwidth allocation of different flows. Otherwise, if no congestion occurs, the maximum weight factor is used to dynamically adjust the weight and gradually and actively increase the credit rate.
[0083] Specifically, the ECN-based credit rate control algorithm:
[0084] Step 1.1, Initialize parameters:
[0085] Set the update period for dynamically adjusting the credit rate ; Set the initial weight factor ; Obtain the initial ; Obtain the link bandwidth of the network ;
[0086] Set the credit rate to take the maximum value , where the maximum value is assigned as: ;
[0087] Set the data reception rate ;
[0088] Set the set of flows ;
[0089] Obtain the minimum weight factor , which is used to control the decrease amplitude of the credit rate;
[0090] Step 1.2, For each flow , every time, repeat the following steps:
[0091] According to the value of , determine whether credit congestion occurs in the current period;
[0092] If credit congestion occurs, reduce the credit rate and assign a value to the credit rate: ; And use the minimum weight factor to reset the weight: ;
[0093] If trust congestion does not occur, use the weight to perform a smooth assignment of the credit rate value: ; and dynamically adjust the weight: , where is the maximum weight factor, used to control the recovery speed;
[0094] Update value;
[0095] In , eliminate the that has ended execution;
[0096] Step 1.3, repeat the termination condition of Step 1.2: ; Output , weight .
[0097] The above credit rate control algorithm based on ECN is the same as the data rate adjustment algorithm in PCN, and its efficiency is high enough to converge to fairness quickly.
[0098] When using the above credit rate control mechanism based on ECN to achieve network congestion control, since credit queuing will greatly reduce the overhead of the switch buffer compared to data queuing, the advantage of credit queuing in rate control is obvious. Considering the overhead comparison between credit congestion control and data congestion control in the same network, that is, using the same link bandwidth and the same basic delay . Let the credit size be , and the packet size be ; Therefore, when the credit queue is not released, the credit rate limit is , while the data rate is . It is easy to find that the end-to-end transmission time of both credit and data is . As is well known, the reaction time determines the queue accumulation in ECN-based congestion control. Note that the accumulated data volume is equal to the link bandwidth multiplied by the time. Since the reaction times of credit congestion and data congestion are the same, and the credit rate is much lower than the data rate, the accumulation speed of the data queue is much faster than that of the credit queue.
[0099] In addition, credit queuing rarely affects data transmission delay. Data transmission delay is mainly determined by data queuing delay and propagation delay. Since propagation delay depends on the network infrastructure, credit queuing has no impact on it. In addition, data queuing is determined by the credit transmission rate of the reverse path, rather than by credit queuing.
[0100] It can be seen that the present invention has obtained significant benefits by applying ECN-based rate control to credit. In this process, in order to reduce bandwidth waste and unnecessary information loss during transmission, a speculative probing mechanism and a reasonable retransmission strategy for data and credit also need to be designed.
[0101] (2) Design a speculative detection mechanism to prevent bandwidth waste in the first RTT transmission of a flow during congestion.
[0102] In the process of implementing network congestion control using the ECN-based credit rate control mechanism mentioned above, it is found that a speculative detection mechanism needs to be designed to prevent bandwidth waste during the first RTT transmission round-trip of a flow. Since the solution adopted in the present invention cancels the credit reservation process, the sender sends probe packets at the line rate during the first round-trip without being credit-driven, while normal data packets are credit-driven.
[0103] Probe packets are not credit-controlled, which may lead to data congestion. In addition, when severe congestion occurs, they may also cause packet loss. Packet loss has many impacts on network performance. Packet loss is harmful to throughput because the bandwidth occupied by the lost packet is ultimately not used for data transmission. In addition, the retransmission delay of data packets affects the flow completion time (FCT).
[0104] When packet loss occurs in the last part of a flow and the corresponding data volume is equal to the DDP (Deadline-Driven Protocol) data (the data volume threshold set by the protocol), the impact on the flow completion time (FCT) may be more serious. When a packet is lost, the sender needs at least one round-trip time delay (RTT) to make a retransmission decision. One RTT represents the time from sending the packet to receiving the feedback, which the present invention calls the reaction RTT. It can be assumed that when packet loss in the flow occurs before the last DDP data, the data flow will continue to be transmitted normally during the reaction RTT. In this case, the FCT only increases , where is the packet size, is the data transmission rate. However, when the lost packet is located in the last DDP data of the flow, the data flow will be transmitted normally before the sender receives the feedback and retransmits the lost data. Therefore, the bandwidth will be wasted and the FCT will increase , where is the remaining traffic size after packet loss.
[0105] To mitigate the impact of data loss on the flow completion time (FCT), the goal of the present invention is to preferentially discard normal data packets rather than the last DDP packet (the packet carrying DDP data) in the case of buffer overflow. Here, a priority queue is used to distinguish between normal data packets and the last DDP packet, and the last DDP packet is set to the highest priority to prevent its loss as much as possible.
[0106] (3) Develop data retransmission strategies and credit retransmission strategies to address data loss and credit loss during congestion control.
[0107] 1) Develop a data retransmission strategy based on the continuity of data sequence numbers to address data loss caused by data queue overflow.
[0108] Under normal circumstances, if each data packet is strictly credit-driven, the data queue is usually short and limited. However, when adopting the aforementioned speculative probing mechanism, the data queue may overflow, resulting in data packet loss. Therefore, the present invention develops a data loss recovery mechanism.
[0109] The data retransmission strategy is based on the continuity of data sequence numbers. The receiving end determines the lost data packets by calculating the difference between the data sequence numbers carried by two consecutively arriving data packets. When data loss is detected, the receiving end records the data sequence number of the lost data packet in the next credit sent. When the sending end receives a credit carrying a data sequence number, the sending end retransmits the corresponding data packet and will not send other data packets driven by this credit.
[0110] 2) Develop a credit retransmission strategy based on the continuity of credit sequence numbers to address credit loss caused by credit queue overflow under severe credit congestion.
[0111] The credit queue can buffer thousands of credits with only dozens of kilobytes. Under normal circumstances, the size of the credit queue is sufficient for credit congestion control. However, under extreme traffic burst conditions, credits may be lost due to credit queue overflow under congestion. In addition, the total number of credits sent is the same as the total number of data packets. Therefore, when credits are unexpectedly lost, data transmission cannot be completed. To solve the problem of credit loss, the present invention develops a credit loss recovery strategy, which is the credit retransmission strategy.
[0112] Each credit generated by the receiving end and transmitted to the sending end carries an incrementing credit sequence number. When the sending end receives a credit, it records the credit sequence number in the corresponding data packet sent to the destination. In this way, the sequence numbers of the non-lost credits are all fed back to the receiving end. Since the credit sequence numbers sent by the receiving end are continuous, the receiving end can easily determine the number of lost credits by calculating the difference between the credit sequence numbers of two consecutively arriving data packets. After determining the number of credits to be retransmitted, the receiving end adds this number to the total number of credits already sent. In this way, it can be ensured that all data packets can be driven by credits in a one-to-one manner.
[0113] Note that, different from data retransmission, it is not necessary to retransmit the lost original credit, because the data packets related to the lost credit can be driven by any other credit belonging to the flow. In addition, it is not necessary to retransmit the credit immediately, because uncontrolled credit will interfere with the credit rate control mechanism.
[0114] (4) Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limitation and switch-side credit selective discard.
[0115] Through the congestion control of the data-credit coupling mechanism, the low queuing of normal data packets (non-probe packets) is maintained, which limits the application of traditional flow scheduling mechanisms. Traditional flow scheduling mechanisms focus on adjusting the queuing order of data transmission in the switch queue. Therefore, for the flow deadline, the present invention develops and designs a credit-driven flow scheduling mechanism to meet the deadline requirements of AI applications, including further reducing the flow completion time and deadline miss rate. This mechanism scheme achieves the target requirements through two key components, namely receiver credit rate limitation and switch-side credit selective discard. In this embodiment, taking Figure 5 as an example, the working process of the credit-driven flow scheduling mechanism is described. As Figure 5 shown, flow is transmitted from the first sender to the receiver, and flow is transmitted from the second sender to the receiver. The deadline of flow is less than that of flow . Therefore, the receiver preferentially sends the credit of flow , and when the length of the credit queue exceeds , the first switch will selectively discard the credit of flow . The credit-driven flow scheduling mechanism specifically includes:
[0116] 1) Receiver credit rate limitation:
[0117] First, a concept called Deadline Reaching Remaining Time (DRRT) is proposed, which represents the remaining time for a flow to reach the deadline, and the DRRT value of the flow is represented by the set . Let , where represents the credit rate limitation of the th flow . When the receiver receives multiple flows simultaneously, it will first meet the requirements of the shortest DRRT flow, that is, release the credit rate limitation of the flow to the maximum credit rate , and mark a high-priority label for the corresponding data packet (the corresponding credit is also marked as high-priority), and the switch can recognize this label. Set As a temporary credit rate limit variable for dynamically adjusting the rate of a flow, the credit rate limit of other data flows will be set to , so as to accelerate the transmission of the shortest DRRT flow. Note that in the ECN-based credit rate control algorithm, the used to adjust the credit rate should be assigned to .
[0118] Specifically, the following credit-driven flow scheduling algorithm is adopted to implement the credit rate limit at the receiver:
[0119] Step 2.1, input parameters: a set of flows , a set of DRRT values of the flows , and a set of credit rate limits of the flows ; set the value of the link bandwidth ratio of the credit rate; obtain the link bandwidth of the network;
[0120] Step 2.2, initialize parameters:
[0121] Set the maximum credit rate: ;
[0122] Initialize the current credit rate: ;
[0123] Initialize the temporary credit rate limit: ;
[0124] Step 2.3, repeatedly execute the main loop steps:
[0125] Each time a data packet arrives, check the flow to which the data packet belongs and process the loop process:
[0126] Step 2.3.1, if the data packet belongs to the current flow , , then execute in the current flow:
[0127] is the minimum value in the set , then
[0128] Set the credit rate limit of the th flow to the maximum credit rate: ;
[0129] Update the temporary credit rate limit: ;
[0130] Mark high priority;
[0131] else if Not a set Minimum value in
[0132] Set the credit rate limit of the th flow to the temporary credit rate limit: ;
[0133] ;
[0134] At the end of the execution in perform culling;
[0135] end if
[0136] If end the loop;
[0137] Step 2.3.2, if the data packet does not belong to any current flow , , process the addition of a new flow :
[0138] Set the credit rate limit of the new flow to the temporary credit rate limit: ;
[0139] Update the current credit rate: ;
[0140] Step 2.4, conditions for ending the execution of the main loop: , output , , , high-priority flag.
[0141] 2) Credit selective discard at the switch end:
[0142] The present invention sets a threshold for the credit queue of the switch to achieve credit selective discard. Once the length of the switch's credit queue exceeds , the switch discards credits without high-priority tags. In this way, credits with high priority can be transmitted faster without modifying the switch hardware. Therefore, the transmission of the shortest DRRT flow can be accelerated at the switch end, greatly improving the data throughput and facilitating high-throughput network data transmission.
[0143] The data and credit coupling mechanism adopted by the present invention designs different functional mechanisms and strategies in the data and credit coupling mechanism through dynamic binding of data flow characteristics such as deadlines, priorities, real-time requirements, etc., and credit allocation strategies, and they work together to achieve congestion control and efficient scheduling of the data center network.
[0144] The key innovation points of the present invention mainly include:
[0145] 1) Introduce the DCEF framework, and propose a new data and credit coupled congestion control scheme based on this framework. By dynamically synchronizing the credit transmission rate and the data reception rate, achieve precise real-time credit rate adaptation, and at the same time enforce one-to-one credit data transmission to effectively alleviate congestion;
[0146] 2) Develop a credit rate control mechanism based on ECN to achieve full-loop congestion control and solve the problem of credit waste;
[0147] 3) Utilize a credit-driven flow scheduling mechanism to preferentially process traffic according to the deadline urgency, further reduce the flow completion time (FCT) and the deadline miss rate (DMR), and can achieve high-throughput network data transmission to meet the performance requirements of AI applications for data center networks.
[0148] In another embodiment of the present invention, a data center network congestion control device based on data and credit coupling is protected. The device is used to perform data center network congestion control based on data and credit coupling by using the foregoing method. The device includes the following modules:
[0149] The first module is used to obtain the structure of the data center network, including switches, senders, and receivers;
[0150] The second module is used to develop a credit rate control mechanism based on ECN according to the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, for managing data congestion and credit congestion and performing full-loop congestion control of the data center network;
[0151] The second module further includes:
[0152] Sub-module 1 is used to set that the switch is the congestion point location, the sender is the notification point location, and the receiver is the reaction point location;
[0153] Sub-module 2 is used to enable the receiver to send credit through the reverse path of credit transmission, pull packets from the sender, and adopt credit rate limitation to prevent data congestion in the reverse path;
[0154] Sub-module 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / packet, and when the credit / data queue is not congested, does not mark the ECN label for the credit / packet;
[0155] The sub-module 4 is configured to dynamically adjust the credit rate according to the results of marking the ECN label, including the credit / data packet carrying the ECN label and not carrying the ECN label, by using the ECN-based credit rate control algorithm.
[0156] In another embodiment of the present invention, there is provided a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the foregoing data center network congestion control method based on the coupling of data and credit are implemented. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is configured to communicate with an external terminal through a network connection.
[0157] On the other hand, in an embodiment of the present invention, there is provided a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing data center network congestion control method based on the coupling of data and credit are implemented.
[0158] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0159] Matters not described in this invention are well-known technologies.
[0160] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0161] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.
Claims
1. A data center network congestion control method based on the coupling of data and credit, characterized in that, Including: Obtain the structure of the data center network, including switches, senders, and receivers; Based on the principles of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework, develop an ECN-based credit rate control mechanism for managing data congestion and credit congestion and performing congestion control for the entire loop of the data center network; the ECN-based credit rate control mechanism includes: Set the switch as the congestion point location, the sender as the notification point location, and the receiver as the response point location; The receiver sends credit through the reverse path of credit transmission, pulls packets from the sender, and uses credit rate limiting to prevent data congestion on the reverse path; Use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / packet, and when the credit / data queue is not congested, it does not mark the ECN label for the credit / packet; According to the results of marking the ECN label, including whether the credit / packet carries the ECN label or not, dynamically adjust the credit rate using the ECN-based credit rate control algorithm; The congestion control for the entire loop of the data center network includes: Design a speculative probing mechanism to prevent bandwidth waste in the first RTT transmission of a flow during the congestion control process; the RTT is the round-trip delay of the flow, representing the time from when the sender sends a packet to when the sender receives the feedback; Formulate data retransmission policies and credit retransmission policies to handle data loss and credit loss during the congestion control process; Design a credit-driven flow scheduling mechanism to achieve flow scheduling control through receiver credit rate limiting and switch-side credit selective dropping; the credit-driven flow scheduling mechanism includes: When the receiver receives multiple flows simultaneously, it will first meet the requirements of the shortest DRRT flow and mark a high-priority label for the packets of the shortest DRRT flow, and the corresponding credit is also marked with a high-priority label; the DRRT represents the remaining time for the flow to reach the deadline, and the shortest DRRT flow refers to the flow with the smallest DRRT value among the multiple flows received simultaneously; Implement receiver credit rate limiting using the credit-driven flow scheduling algorithm; Accelerate the transmission of the shortest DRRT flow at the switch end by selectively discarding at the switch end's credit, including setting a threshold for the switch's credit queue , when the length of the credit queue exceeds the threshold , the switch selects to discard the credit without the high-priority label marked.
2. The congestion control method for a data center network based on data and credit coupling according to claim 1, characterized in that The data center network includes at least two switches, two senders, and two receivers; The ECN is an explicit congestion notification network protocol method that allows the switch to explicitly notify the sender or receiver of congestion information by marking the header field of the packet when congestion occurs.
3. The data center network congestion control method based on data and credit coupling according to claim 2, characterized in that Let the result variable for marking the ECN label be , if , it means that the credit / data packet carries the ECN label and the trust congestion occurs. If , it means that the credit / data packet does not carry the ECN label and the trust congestion does not occur; Set the weight factor , which is used to balance the bandwidth allocation of different flows; The dynamic adjustment of the credit rate using the ECN-based credit rate control algorithm includes: Step 110, initialize parameters: set the set of flows ; set the link bandwidth ratio of the credit rate value; obtain the link bandwidth of the network ; set the credit rate take the maximum value , where the maximum value is assigned to ; set the update period for dynamically adjusting the credit rate ; set the initial weight factor ; obtain the initial ; obtain the minimum weight factor , which is used to control the decline range of the credit rate; Step 120, each update period Within, according to the value, determine whether credit congestion occurs and dynamically adjust the credit rate: If , credit congestion occurs, then reduce the credit rate and assign a value to the credit rate: , where is the deceleration ratio of the credit rate, and use the minimum weight factor to reset the weight factor ; If , and credit congestion does not occur, then the weight factor is used to smoothly assign the credit rate value: ; Dynamically adjust the weight factor: , where is the maximum weight factor, which is used to control the recovery speed of the credit rate; Update Value; In exclude those that have completed execution ; Step 130, repeat Step 120 until the termination condition is met , output .
4. The method for data center network congestion control based on data and credit coupling according to claim 3, wherein The speculative probing mechanism includes: Use a priority queue to distinguish ordinary packets and the last DDP packet, and set the last DDP packet to the highest priority; the DDP packet refers to the packet of DDP data, and the DDP data refers to the data volume threshold set according to the deadline-driven protocol; In the case of network buffer data queue overflow, ordinary data packets are preferentially discarded and DDP packets are retained.
5. The data center network congestion control method based on the coupling of data and credit according to claim 4, characterized in that, The process of formulating data retransmission policies and credit retransmission policies to address data loss and credit loss during congestion control includes: Formulating a data retransmission policy based on the continuity of data sequence numbers to address data loss caused by data queue overflow, including: the receiving end determines the lost data packets by calculating the difference between the data sequence numbers carried by two consecutively arriving data packets; when detecting data packet loss, the receiving end records the data sequence number of the lost data packet in the next sent credit; when the sending end receives the credit carrying the data sequence number, it retransmits the lost data packet and no longer sends other data packets driven by the same credit. Formulating a credit retransmission policy based on the continuity of credit sequence numbers to address credit loss caused by credit queue overflow under severe credit congestion, including: the receiving end determines the number of lost credits by calculating the difference between the credit sequence numbers of two consecutively arriving data packets; determines the number of credits that need to be retransmitted from the number of lost credits, and the receiving end adds the number of credits that need to be retransmitted to the total number of sent credits to ensure that all data packets can be driven one-to-one by credits.
6. The data center network congestion control method based on data and credit coupling according to claim 5, characterized in that, The credit-driven flow scheduling algorithm includes: Step 210, input parameters: a set of flows , a set of DRRT values of the flows , and a set of credit rate limits of the flows ; set the link bandwidth ratio of the credit rate value; obtain the link bandwidth of the network ; Step 220, initializing parameters: Set the maximum credit rate: ; Initialization of the current credit rate: ; Initialization of temporary credit rate limit: ; Step 230, repeatedly executing the main loop step: Step 231, each time a data packet arrives, check the flow to which the data packet belongs. Step 232, if the data packet belongs to the current flow , , then execute a loop in the current flow: When is the minimum value in the set , set the credit rate limit of the th flow to the maximum credit rate: ; Update the temporary credit rate limit: ; Mark the high-priority label; When is not the minimum value in the set , set the credit rate limit of the th flow to the temporary credit rate limit: ; End the loop of the current stream and perform the elimination of the execution end in and jump back to step 231; If the data packet does not belong to the current flow , break out of the loop in step 232 and execute step 233; Step 233, if the data packet does not belong to the current flow , , process the addition of a new flow : Set the credit rate limit of the new flow to the temporary credit rate limit: ; Update the current credit rate limit: ; Step 240, repeat Step 230 until the termination condition is met , output , , , and high-priority tags.
7. The data center network congestion control method based on the coupling of data and credit according to claim 6, wherein The link bandwidth ratio is taken as .
8. A data center network congestion control device based on the coupling of data and credit, characterized in that, The device performs congestion control for a data center network based on the coupling of data and credit using the steps of the method according to claim 1. The device includes the following modules: The first module is used to obtain the structure of the data center network, including switches, sending ends, and receiving ends. The second module is used to develop an ECN-based credit rate control mechanism for managing data congestion and credit congestion and performing congestion control for the entire loop of the data center network based on the principle of designing a data and credit coupling mechanism in the data and credit and end-to-end feedback framework. The second module further includes: Sub-module 1 is used to set the switch as the congestion point location, the sending end as the notification point location, and the receiving end as the reaction point location. Sub-module 2 is used to enable the receiving end to send credits through the reverse path of credit transmission, pull data packets from the sending end, and adopt credit rate limiting to prevent data congestion on the reverse path. Sub-module 3 is used to use the same ECN feedback signal for credit congestion and data congestion; when the credit / data queue is congested, the switch marks the ECN label for the credit / data packet, and when the credit / data queue is not congested, it does not mark the ECN label for the credit / data packet. Sub-module 4 is used to dynamically adjust the credit rate using an ECN-based credit rate control algorithm according to the results of marking the ECN label, including whether the credit / data packet carries the ECN label or not.
Citation Information
Patent Citations
Data center network congestion control method based on credit and reaction types
CN112468405A
Technique for providing end-to-end congestion control with no feedback from a lossless network
US7035220B1
Cited By
RAN and UE driven L4S marking and processing for congestion management in an O-RAN based network architecture
US12684410B2
RAN and UE Driven L4S Marking and Processing for Congestion Management in an O-RAN Based Network Architecture
US20240334244A1