A collaborative congestion control method based on INCAST lifecycle state awareness

CN122621540BActive Publication Date: 2026-09-29无锡沐创集成电路设计有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611080181.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-29
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

除DCQCN外,TIMELY、NSCC等基于发送端独立拥塞反馈调节的方法,在大规模Incast场景下同样存在上述问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621540B_ABST
    Figure CN122621540B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data center network and RDMA network congestion control, and discloses a kind of collaborative congestion control method based on INCAST life cycle state awareness. Including: receiving end responds to data message, and the unique identification of sending end is recorded in active list;Based on the trend of the number of active sending end changes, determine the life cycle state of INCAST state machine;Based on state and available bandwidth, calculate the current effective limiting speed value;In INCAST state, generate RI information containing the limiting speed value and send to each sending end in active list. Sending end receives RI information, if it does not contain release information, enter RI control mode and reset RI timer, update target rate based on limiting speed value, if current sending rate is greater than limiting speed value, then slow down;If it contains release information and RI timer has not expired, stop RI timer, exit RI control mode, and restore the autonomous rate growth mechanism based on TC and / or BC driven.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data center network and RDMA network congestion control, and in particular to a collaborative congestion control method based on INCAST lifecycle state awareness. Background Technology

[0002] DCQCN is widely used as an end-to-end congestion control mechanism in current RoCE networks. DCQCN is a transmitter-adaptive rate control algorithm based on ECN (Explicit Congestion Notification), and its basic workflow is as follows: Figure 1 As shown, its core consists of three parts: the sending end (RP, Reaction Point), the switch (CP, Congestion Point), and the receiving end (NP, Notification Point). The sending end is responsible for dynamically adjusting the sending rate based on congestion notification messages; the switch is responsible for detecting queue congestion and marking it with an ECN; and the receiving end is responsible for returning a CNP (Congestion Notification Packet) to the sending end after receiving an ECN-marked message.

[0003] The specific working mechanism of the traditional DCQCN is as follows: (1) The transmitting end RP continuously sends an RDMA data stream to the receiving end at the current transmission rate Rc; (2) When a data packet carrying an RDMA data stream passes through a switch CP, the switch will mark the data packet carrying the RDMA data stream with ECN (Explicit Congestion Notification) after detecting that the output queue length (the number of bytes or packets of data packets waiting to be sent to the next hop link at the physical output port of the switch) exceeds the preset ECN threshold. (3) After receiving a data packet with the ECN tag, the receiving NP returns a CNP (Congestion Notification Packet) congestion notification packet to the corresponding sending end (which is a control packet and contains the network card address of the sending server). (4) After receiving the CNP, the sending end RP enters the congestion response process and performs multiplicative decrease. Its core state variables include: Rc: Current Rate; Rt: Target Rate; α: Congestion factor; g: Smoothing coefficient in α update.

[0004] In a typical implementation, the update process is as follows: Rt = Rc; Rc = Rc × (1 α / 2); α = (1 g) × α + g First, the sending end saves the current sending rate Rc as the target rate Rt; then, it multiplicatively reduces the sending rate Rc based on the current congestion factor α; at the same time, it updates α in conjunction with the currently received congestion notification messages, thereby gradually reflecting the current network congestion level.

[0005] (5) After the rate reduction is completed, the transmitting end enters the transmission rate recovery phase. Traditional DCQCN combines both TimerCounter (TC) and Byte Counter (BC) mechanisms to drive transmission rate recovery: Byte Counter: Triggers a speed-up after sending a certain number of bytes of data; Timer Counter: Triggers a speed-up after a fixed time period.

[0006] The acceleration process of DCQCN typically includes two stages: The first stage is the Fast Recovery stage. In this stage, the transmission rate Rc recovers to the historical target rate Rt at a relatively fast speed. In a typical implementation, this lasts for F iterations (e.g., F = 5), and the update method is as follows: Rc = (Rt + Rc) / 2; that is, each recovery round makes the current rate move closer to the target rate Rt by half the distance, thereby achieving faster recovery.

[0007] The second stage is the Additive Increase stage. After completing the rapid recovery, the transmitter enters a smoother linear rate-up process. Its typical update method is as follows: Rt = Rt + Rai; Rc = (Rt + Rc) / 2; where Rai is a fixed additive rate-up step size. In this stage, the target rate Rt continuously increases at a fixed step size, while the current transmission rate Rc gradually approaches the new target rate Rt, thus achieving smooth bandwidth detection.

[0008] DCQCN is essentially a sender-driven congestion control mechanism based on ECN / CNP feedback. The sender independently controls the rate of decrease and increase based on the congestion feedback returned by the network. Besides DCQCN, methods such as TIMELY and NSCC also belong to congestion control mechanisms where the sender independently adjusts the congestion feedback.

[0009] In recent years, the industry has gradually proposed receiver-driven congestion control mechanisms, such as RCCC (Receiver-Centric Congestion Control) in UEC and the grant / credit-based congestion control method in GSE.

[0010] The core idea of ​​this type of scheme is that the receiving end uniformly perceives the traffic situation of multiple sending ends and performs rate control on these multiple sending ends, rather than each sending end independently adjusting congestion. Its typical working mechanism is as follows: Figure 2 As shown. Taking RCCC as an example, its main working principle is as follows: (1) The receiving end calculates the total allocable bandwidth based on the current receiving capability. This bandwidth can be obtained through real-time measurement or by using a pre-configured fixed value. (2) The receiving end maintains an active remote list and periodically sends Credit control messages through a unified Credit timer; (3) The sending period of the Credit timer satisfies: sending period = data amount corresponding to a single Credit / total available bandwidth, where a single Credit usually corresponds to a data amount of k×MTU bytes, k is a preset constant, which is generally taken as a small value, such as 4; (4) Each time the Credit timer is triggered, the receiving end uses a round-robin method to select a sending end from the list of active remote ends and add the corresponding amount of Credit to it, usually k×MTU bytes.

[0011] (5) When the sending end sends data, it needs to consume the corresponding Credit. Data can only be sent when the accumulated Credit is sufficient. (6) When the sending end has insufficient credit, the sending end suspends sending and waits for the receiving end to authorize new credit; A smaller k value can reduce bursty traffic and improve bandwidth sharing smoothness, but it also increases the number of credit control messages and control plane overhead. Compared to traditional sender control methods based on congestion feedback, receiver-driven approaches typically achieve faster convergence, lower queue backlog, and better bandwidth fairness in large-scale incast scenarios.

[0012] Currently, existing technologies still have the following technical shortcomings: 1) Traditional sender congestion control converges slowly in large-scale incast scenarios: Traditional DCQCN is essentially a passive feedback control mechanism at the sending end based on network congestion events. That is, the sending end will only start to reduce the sending rate after the switch has generated queue backlog and triggered the ECN flag.

[0013] Therefore, in large-scale incast scenarios, the following process is likely to occur: multiple senders simultaneously send data at a high initial rate; the aggregated traffic instantly exceeds the receiver's link bandwidth; the switch output queue grows rapidly; then the switch begins to perform ECN marking; the receiver returns a CNP; and the sender gradually slows down after receiving the CNP.

[0014] Because there is a significant feedback delay throughout the control link, it can easily lead to: Sudden surges in queue length; bursts of ECN / CNP packets; increased probability of PFC (Priority Flow Control) triggering; significantly increased network latency; severe throughput jitter; unfair bandwidth allocation among multiple senders; and long incast convergence time (each sender independently competes for bandwidth, requiring multiple rounds of "deceleration-increase" oscillations before gradually reaching a fair and stable state). Besides DCQCN, methods like TIMELY and NSCC, which rely on independent congestion feedback adjustment at the sender level, also suffer from these problems in large-scale incast scenarios.

[0015] (2) Receiver-driven methods have difficulty handling internal network bottlenecks and congestion: While receiver-driven methods offer good incast performance in ideal environments with no network congestion and no oversubscription, they typically assume that "the receiver's outgoing bandwidth is the network's available bandwidth." These methods are prone to failure when real bottlenecks exist within the network. Situations where real bottlenecks exist include: ECMP hash collisions causing some traffic to concentrate on specific links; link failures reducing available paths; oversubscription in the network design; and localized hotspots within the network causing congestion.

[0016] In the above scenario, the receiver cannot accurately perceive the actual bottleneck bandwidth within the network. A practical example: the receiver's link bandwidth is 200G, so it allocates a 200G transmission allowance to the sender; however, there is actually a 100G bottleneck link in the network path. If the sender continues to transmit data at a rate of 200G, it will cause severe congestion on the bottleneck link. Therefore, the receiver-driven method lacks the ability to perceive the dynamic congestion state within the network and is difficult to adapt to complex network topologies and dynamic congestion scenarios.

[0017] (3) Receiver-driven methods incur additional control message overhead: Receiver-driven methods based on Credit / Grant typically require the periodic transmission of grant control messages. These control messages consume additional network bandwidth. For example, in a typical implementation, each Credit message allows the sender to transmit only k×MTU size data, where k is usually a small value to reduce bursty traffic. Therefore, when k is small, the receiver needs to send Credit control messages at a higher frequency, resulting in a significant increase in the number of control messages and additional control plane bandwidth overhead. Summary of the Invention

[0018] The purpose of this invention is to provide at least one cooperative congestion control method based on INCAST lifecycle state awareness, particularly an enhanced DCQCN (Data Center Quantized Congestion Notification) congestion control method for RoCE (RDMA over Converged Ethernet) network incast scenarios, belonging to the fields of high-speed Ethernet, RDMA network transmission, and data center communication technology. Unlike traditional sender-driven methods that rely solely on switch ECN tags for congestion feedback, this invention proactively senses the aggregation behavior of multiple senders at the receiver, coordinating and rate-limiting multiple senders before the incast queue is formed or the switch generates a large number of ECN tags. Unlike traditional receiver-driven methods based on Credit / Grant, this invention does not employ a "send credit" mechanism, nor does it require the sender to obtain authorization from the receiver before sending data; instead, it employs a "Receiver-assisted Rate Coordination" mechanism. Specifically: The receiver detects the current Incast lifecycle state and, based on different Incast states, uses different contention factor calculation methods and rate coordination strategies to generate the corresponding sender's suggested transmission rate, Instruction_Rate, and announces it to the corresponding sender via RI (RateInstruction) information. While retaining the dynamic congestion control capabilities of traditional DCQCN based on ECN / CNP, the sender introduces an additional rate constraint mechanism based on the current effective rate limit (Instruction_Rate), ensuring that the sender's final transmission rate is simultaneously constrained by both the actual congestion feedback within the network and the receiver's Incast coordination rate.

[0019] To address the aforementioned technical problems, at least one embodiment of this application provides a cooperative congestion control method based on INCAST lifecycle state awareness, applied at a receiving end, the method comprising: In response to receiving a data packet sent by the sender, the unique identifier information corresponding to the sender is recorded in the list of currently active senders, and the last active timestamp of the sender is updated; Based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage, the lifecycle state of the INCAST state machine is determined; the lifecycle state includes the NORMAL state and the INCAST in progress state. The contention factor is determined based on the lifecycle state, and the current effective rate limit value is determined according to the available bandwidth of the receiver and the contention factor. The current effective rate limit value is then associated and stored in the record item corresponding to the transmitter in the list of currently active transmitters. The contention factor is determined in different ways under different lifecycle states. When the lifecycle state is in the INCAST state, RI information containing the current effective speed limit value is generated, and the RI information is sent to each sender in the list of currently active senders.

[0020] To address the aforementioned technical problems, at least one embodiment of this application also provides a cooperative congestion control method based on INCAST lifecycle state awareness, applied at the transmitting end, the method comprising: In response to receiving RI information sent by the receiving end, determine whether the RI information contains RI cancellation information; If the RI information does not contain information to cancel the RI, enter the RI control mode and create or reset the RI timer to start timing. Obtain the current effective rate limit value from the RI information and update the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, control the current transmission rate of the transmitter to be no greater than the current effective rate limit value. If the RI information contains RI cancellation information and the RI timer has not timed out, control the RI timer to stop counting, exit the RI control mode, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate and Instruction_Rate is the current effective rate limit value.

[0021] To address the aforementioned technical problems, at least one embodiment of this application also provides a cooperative congestion control method based on INCAST lifecycle state awareness, the method comprising: In response to receiving a data packet sent by the sender, the receiver records the unique identifier information corresponding to the sender in the list of currently active senders and updates the last active timestamp of the sender. The receiving end determines the lifecycle state of the INCAST state machine based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage; the lifecycle state includes the NORMAL state and the INCAST in progress state. The receiving end determines the contention factor based on the lifecycle state, and determines the current effective rate limit value according to the available bandwidth of the receiving end and the contention factor, and stores the current effective rate limit value in association with the record item corresponding to the sending end in the list of currently active sending ends; wherein, the contention factor is valued differently in different lifecycle states; When the lifecycle state is in INCAST state, RI information containing the current effective rate limit value is generated at the receiving end, and the RI information is sent to each sending end in the list of currently active sending ends; In response to receiving RI information sent by the receiving end, the sending end determines whether the RI information contains RI cancellation information; If the RI information does not contain RI cancellation information, the control transmitter enters the RI control mode and creates or resets the RI timer to start timing. It obtains the current effective rate limit value from the RI information and updates the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, the control transmitter's current transmission rate is made not to exceed the current effective rate limit value. If the RI information contains RI cancellation information and the RI timer has not expired, the RI timer at the control end is stopped, the RI control mode is exited, and the autonomous rate growth mechanism driven by timer count TC and / or byte count BC is restored. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate and Instruction_Rate is the current effective rate limit value.

[0022] At least one embodiment of this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described above.

[0023] At least one embodiment of this application also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described above.

[0024] At least one embodiment of this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0025] The INCAST lifecycle state-aware collaborative congestion control method provided in this application, compared to existing technologies, allows the receiver to proactively detect the aggregation behavior of multiple senders. This enables coordinated rate limiting of multiple senders before the incast queue is formed or the switch generates large-scale ECN markings. Unlike traditional Credit / Grant-based receiver-driven methods, this invention does not employ a "send credit" mechanism, nor does it require the sender to obtain authorization from the receiver before sending data. Instead, it uses a "Receiver-assisted Rate Coordination" mechanism. Specifically, the receiver detects the current incast lifecycle state and, based on different incast states, uses different contention factor calculation methods and rate coordination strategies to generate a suggested sending rate (Instruction_Rate) for the corresponding sender. This suggested rate is then announced to the corresponding sender via RI (Rate Instruction) information. While retaining the dynamic congestion control capabilities of traditional DCQCN based on ECN / CNP, the transmitter introduces an additional rate constraint mechanism based on the current effective rate limit value (Instruction_Rate), so that the final transmission rate of the transmitter is simultaneously constrained by the actual congestion feedback within the network and the incast coordination rate of the receiver.

[0026] For methods applied to the receiving end, in some optional embodiments, the unique identification information includes at least one of the following: The QP number corresponding to the sender, the addressing information corresponding to the sender, and the last active timestamp of the sender.

[0027] For the method applied to the receiving end, in some optional embodiments, the lifecycle state of the INCAST state machine represents the development stage of the QP aggregation traffic of multiple active sending ends, the NORMAL state indicates that INCAST aggregation has not yet occurred, and the INCAST in progress state indicates that INCAST aggregation has already occurred; wherein, the INCAST in progress state includes at least one of the following: The INCAST_GROWTH status indicates that the number of active sender QPs is on a continuous increasing trend; The INCAST_STABLE status indicates that the current INCAST has been formed and the number of active sender QPs has entered a preset stable fluctuation phase; The INCAST_DECAY status indicates that the number of active sender QPs is on a continuous decreasing trend.

[0028] For methods applied to the receiving end, in some optional embodiments, the manner of sending the RI information includes: The RI information is sent via ACK, SACK, NACK, or extended control messages.

[0029] For methods applied to the receiving end, in some optional embodiments, the contention factor characterizes the number of transmitters currently participating in INCAST bandwidth contention or the degree of contention, and determining the contention factor based on the lifecycle state includes: If the lifecycle state is INCAST_GROWTH or INCAST_DECAY, then the number of active transmitters is obtained, and the number of active transmitters is used as the value of the contention factor. If the lifecycle state is INCAST_STABLE, then the value of the competition factor is determined by the following formula: N = Navg(t) = (1 β)Navg(t Tdetect)+βNactive(t) Where N is the competition factor, β is the weighting coefficient, Tdetect is the statistical period or detection time window used by the receiver for INCAST state detection, and Nactive(t) is the number of active transmitters at time t.

[0030] In some optional embodiments of the method applied to the sending end, after updating the target transmission rate of the sending end, the method further includes: If the current transmission rate is greater than the current effective rate limit, then the current effective rate limit is used as the current transmission rate, and the counting of timer count TC and / or byte count BC is paused.

[0031] In some optional embodiments of the method applied to the sending end, the method further includes: If the RI timer has expired, exit the RI control mode and set the current effective rate limit value to the maximum line speed of the transmitting link, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC.

[0032] In some optional embodiments of the method applied to the transmitting end, after obtaining the current effective rate limit value from the RI information, the method further includes: If the current transmission rate at the transmitting end is not greater than the current effective rate limit value, the counting accumulation of timer count TC and byte count BC remains active, and during the rate-up process triggered by timer count TC and / or byte count BC, the rate-up control of the transmitting end is restored to the upper limit value at the current effective rate limit value. Attached Figure Description

[0033] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0034] Figure 1 This is a schematic diagram of a traditional DCQCN congestion control process; Figure 2 This is a schematic diagram of a receiver-driven congestion control method. Figure 3 A flowchart illustrating a collaborative congestion control method based on INCAST lifecycle state awareness applied at a receiving end, provided as an embodiment of this disclosure; Figure 4 A flowchart illustrating a collaborative congestion control method based on INCAST lifecycle state awareness applied to a transmitting end, provided as an embodiment of this disclosure; Figure 5 This is a schematic diagram of a complete RI information processing flow provided in an embodiment of the present disclosure; Figure 6 A flowchart of a collaborative congestion control method based on INCAST lifecycle state awareness provided in an embodiment of this disclosure. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been presented in the various embodiments of the present invention to enable the reader to better understand the present invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and various changes and modifications based on the following embodiments.

[0036] The following explanations will first describe some of the technical terms used in the embodiments of this application and / or the prior art, so as to enable those skilled in the art to understand the technical solutions of this application. The following terms are existing technical terms in the fields of RDMA, RoCE, data center networks, and congestion control:

[0037] The following terms are technical concepts proposed, extendedly defined, or given specific implementation meanings in this invention:

[0038] Symbols / Variables:

[0039] Example 1: The embodiments of the present invention relate to a collaborative congestion control method based on INCAST lifecycle state awareness, which can be applied to the receiving end.

[0040] In this invention, the receiving end is used to count how many active sending ends are simultaneously sending data to it, calculate the fair rate, and distribute it to each sending end via RI information. The sending end, based on the received RI information, limits its own transmission rate according to instructions, while retaining the traditional DCQCN's ability to respond to internal network congestion (e.g., ECN / CNP).

[0041] Traditional ECN-based sender-driven congestion control mechanisms typically rely on the switch's output queue reaching the ECN marking threshold before triggering subsequent congestion feedback. By the time the sender actually perceives congestion, a significant queue backlog has usually already formed within the switch. In this invention, the receiver can identify incast establishment trends formed by the rapid aggregation of multiple senders in advance. This allows for earlier initiation of the sender-coordinated rate-limiting process, thereby reducing the instantaneous queue growth at the initial stage of incast.

[0042] The following is a detailed description of the implementation details of the cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0043] The cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities.

[0044] like Figure 3 As shown, the cooperative congestion control method based on INCAST lifecycle state awareness provided in this embodiment includes the following steps: Step 310: In response to receiving a data packet sent by the sender, record the unique identifier information corresponding to the sender in the list of currently active senders, and update the last active timestamp of the sender.

[0045] In some embodiments, the unique identification information includes at least one of the following: The QP number corresponding to the sender, the addressing information corresponding to the sender, and the last active timestamp of the sender.

[0046] In this embodiment, the receiver maintains a list of currently active transmitters. In an RDMA network, each local QP at the receiver corresponds to a local QP Context. Furthermore, this QP Context records: Local QP information, remote QP information connected to it (i.e., QP information of the sending network card), current connection status, traffic statistics information, and RI-related status information added by this invention.

[0047] In this embodiment, the active remote QP actually corresponds to the remote QP that is currently receiving data from the local QP. An active remote QP (i.e., an active sending QP) is a remote QP that is sending data within a preset time window. The receiving end can identify the corresponding remote QP based on its local QP Context and maintain a list of currently active sending ends (also known as a list of currently active remote QPs).

[0048] Specifically, when the unique identifier information corresponding to the sending end is recorded in the current active sending end list: if the unique identifier information corresponding to the sending end is not recorded in the current active sending end list, the unique identifier information of the sending end is added; if the unique identifier information corresponding to the sending end is already recorded in the current active sending end list, the active time information of the sending end is refreshed.

[0049] Step 320: Based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage, determine the lifecycle state of the INCAST state machine; the lifecycle state includes the NORMAL state and the INCAST in progress state. In some embodiments, the lifecycle state of the INCAST state machine represents the development stage of QP aggregation traffic from multiple active senders, the NORMAL state indicates that INCAST aggregation has not yet occurred, and the INCAST in progress state indicates that INCAST aggregation has already occurred; wherein, the INCAST in progress state includes at least one of the following: The INCAST_GROWTH status indicates that the number of active sender QPs is on a continuous increasing trend; The INCAST_STABLE status indicates that the current INCAST has been formed and the number of active sender QPs has entered a preset stable fluctuation phase; The INCAST_DECAY status indicates that the number of active sender QPs is on a continuous decreasing trend.

[0050] Optionally, the receiver maintains an Incast state machine, which includes: NORMAL state; INCAST_GROWTH state; INCAST_STABLE state; and INCAST_DECAY state.

[0051] Among them: NORMAL status indicates that no significant incast aggregation has occurred; INCAST_GROWTH status indicates that multiple remote QPs are rapidly joining the current incast; INCAST_STABLE status indicates that the current incast has been formed and the number of active remote QPs has entered a relatively stable fluctuation phase; INCAST_DECAY status indicates that the current incast is gradually receding and the number of active remote QPs is continuously decreasing.

[0052] In addition, to facilitate the description of the Incast state detection and state transition process in this invention, the following parameters and symbols are defined: Tdetect: Represents the statistical period or detection time window used by the receiver for Incast state detection.

[0053] Nactive(t): Represents the number of active remote QPs at time t.

[0054] ΔNstable: This represents the threshold used to determine whether the number of active remote QPs has entered a stable fluctuation range.

[0055] Nexit: Represents the threshold number of active remote QPs that have exited the Incast state.

[0056] Navg(t): represents the average number of active remote QPs obtained after smoothing the number of active remote QPs.

[0057] β: represents the weighting coefficient in the smoothing calculation, 0 < β ≤ 1.

[0058] It should be noted that the above parameters are merely example parameters in a preferred implementation of the present invention. In actual implementation, the above parameters can be adjusted according to network scale, link bandwidth, service model, switch buffer size, and system implementation requirements, and the present invention does not impose any limitations on them.

[0059] The transitions between each state are explained below: 1. NORMAL → INCAST_GROWTH: When the receiver detects an incast establishment trend, the state machine switches from the NORMAL state to the INCAST_GROWTH state. The incast establishment trend can be determined based on any one or more of the following conditions: the number of active remote QPs increases rapidly within a unit of time; the number of newly added active remote QPs exceeds a threshold within a unit of time; the number of active remote QPs is greater than a preset threshold; the received aggregated traffic increases rapidly; the receiver's buffer usage increases rapidly; and the receiver link utilization is close to the link limit.

[0060] 2. INCAST_GROWTH → INCAST_STABLE: When the receiver detects that: the growth rate of the number of active remote QPs decreases; the number of active remote QPs enters a stable fluctuation range; and the change in received aggregated traffic tends to stabilize, the state machine switches from the INCAST_GROWTH state to the INCAST_STABLE state.

[0061] For example: If in multiple consecutive detection cycles: |Nactive(t) Nactive(t If Tdetect)∣<Δnstable; then it can be considered that the current Incast has entered a stable phase.

[0062] 3. INCAST_STABLE → INCAST_DECAY: When the receiver detects that the incast is beginning to fade, the state machine switches from the INCAST_STABLE state to the INCAST_DECAY state. The incast fading trend can be determined based on one or more of the following conditions: the number of active remote QPs continues to decrease; a large number of remote QPs exit the active state within a unit of time; the received aggregated traffic continues to decrease; the receive link utilization decreases significantly; and the receiver buffer usage continues to decrease.

[0063] For example: If Nactive(t) occurs over multiple consecutive detection periods... <Nactive(t If Tdetect is detected, it can enter the INCAST_DECAY state.

[0064] 4. INCAST_DECAY → NORMAL: When the receiver detects that the number of active remote QPs drops below a preset threshold, the received aggregate traffic returns to normal, and the received link utilization is lower than a preset threshold, the state machine exits the Incast state and re-enters the NORMAL state.

[0065] For example, if Nactive < Nexit, then exit the current Incast state.

[0066] It should be noted that, in this embodiment, the receiver maintains the Incast life cycle state. The Incast life cycle state is used to represent the current development stage of aggregate traffic from multiple remote QPs, and different rate coordination strategies are adopted according to different stages. The several state divisions mentioned above are only a preferred implementation of the present invention. In actual implementation, the receiver can use any number of state divisions with any granularity, or use an implicit state machine to implement dynamic stage identification and coordination control, based on the change trend of the number of active remote QPs, the change trend of aggregate traffic, the change trend of cache occupancy, the change trend of link utilization, or other statistical information that can reflect changes in the Incast life cycle. The present invention does not limit this.

[0067] For example, the life cycle can be divided into only two states: NORMAL and INCAST; more life cycle stages can be further subdivided; or continuous trend functions, sliding window statistics, probability models, machine learning models and other methods can be used to identify Incast life cycle stages.

[0068] Similarly, the transition process mentioned above is only a preferred implementation of the present invention. In actual operation, due to the dynamic burstiness of data center traffic, frequent entry and exit of short flows, overlapping multi-stage communication and other characteristics, the receiver state machine can dynamically switch between different Incast states based on the change trend of the number of active remote QPs, the change trend of aggregate traffic, the change trend of receive cache, and the change trend of link utilization. Therefore, the state transition in the present invention is not limited to the fixed sequence of NORMAL → INCAST_GROWTH → INCAST_STABLE → INCAST_DECAY → NORMAL.

[0069] For example, the following state transitions may exist: INCAST_GROWTH → INCAST_DECAY; INCAST_DECAY → INCAST_GROWTH; INCAST_STABLE → INCAST_GROWTH; INCAST_GROWTH → NORMAL; INCAST_STABLE → NORMAL; or other state switching methods based on traffic change trends.

[0070] Step 330: Determine the contention factor based on the lifecycle state, and determine the current effective rate limit value according to the available bandwidth of the receiver and the contention factor, and store the current effective rate limit value in association with the record item corresponding to the transmitter in the list of currently active transmitters; wherein, the contention factor is valued differently in different lifecycle states.

[0071] In some embodiments, the contention factor characterizes the number of transmitters currently participating in INCAST bandwidth contention or the degree of contention, and determining the contention factor based on the lifecycle state includes: If the lifecycle state is INCAST_GROWTH or INCAST_DECAY, then the number of active transmitters is obtained, and the number of active transmitters is used as the value of the contention factor. If the lifecycle state is INCAST_STABLE, then the value of the competition factor is determined by the following formula: N = Navg(t) = (1 β)Navg(t Tdetect)+βNactive(t) Where N is the competition factor, β is the weighting coefficient, Tdetect is the statistical period or detection time window used by the receiver for INCAST state detection, and Nactive(t) is the number of active transmitters at time t.

[0072] Specifically, the receiving end generates an Instruction_Rate (i.e., the current effective rate limit value) based on the locally calculated fair rate Rfair, and integrates it into the RI information for transmission. The calculation formula includes: Rfair=B / N In the above formula, B represents the available receive bandwidth at the receiving end that can currently be used for Incast traffic coordination.

[0073] The available receiving bandwidth can be: the physical link bandwidth of the receiving end, or the currently remaining available bandwidth of the receiving end; The receiver may use a fixed bandwidth value pre-configured at the receiving end, a bandwidth value dynamically adjusted by the receiving end in combination with buffer usage, traffic load, link utilization, and network status, or other bandwidth parameters that can characterize the current receiving capability.

[0074] For example, in one implementation, B can directly take the link negotiation rate of the receiving end's network card, such as 100Gbps, 200Gbps, or 400Gbps.

[0075] In another implementation, the receiving end can dynamically calculate the remaining available bandwidth based on the current background traffic usage and use this remaining available bandwidth as B.

[0076] It should be noted that the specific calculation method of B in this embodiment is not a limitation. Any bandwidth estimate that can reflect the current coordinateable reception capability of the receiving end can be used as B in this invention.

[0077] In the above formula, N represents the contention factor used for fair rate calculation. The contention factor characterizes the number of remote QPs currently participating in Incast bandwidth contention or the degree of contention. N can be calculated in different ways depending on the implementation and the Incast state: In the INCAST_GROWTH state, due to the rapid increase in the number of active remote QPs, the receiver prefers to use the current real-time active remote QPs as the basis for fair rate calculation, i.e., N=Nactive(t).

[0078] This allows for a rapid reduction in the suggested transmission rate of each sender, suppressing the initial instantaneous queue growth during Incast.

[0079] In the INCAST_STABLE state, since the number of active remote QPs may fluctuate frequently due to short-circuit entry and exit, the receiver prefers to smooth the number of active remote QPs before using it for fair rate calculation.

[0080] For example, moving average, exponentially weighted average, or other low-pass filtering methods can be used: Navg(t) = (1 β)Navg(t Tdetect)+βNactive(t) This reduces jitter in the suggested transmission rate and improves the stability of the sending rate.

[0081] In the INCAST_DECAY state, as the number of active remote QPs continues to decrease, the receiver prefers to re-adopt the current real-time active remote QP number, i.e., N=Nactive(t), for fair rate calculation.

[0082] In another preferred implementation, the receiver, in the INCAST_DECAY state, can further combine the continuous detection period trend or smoothing statistical results to smoothly control the release process of the suggested transmission rate, thereby reducing the transmission rate oscillation caused by short-term traffic fluctuations.

[0083] It should be noted that using Rfair as the Instruction_Rate is only one preferred implementation method.

[0084] In other implementations, the receiver can also calculate the Instruction_Rate corresponding to different remote QPs based on: flow priority, flow type, QoS level, tenant weight, historical transmission rate, flow size; or other bandwidth allocation strategies.

[0085] In the NORMAL state, in a preferred implementation, the receiving end can notify the sending end to release the RI-based transmission rate limit by sending RI information containing a preset release flag. For example: Set the Instruction_Rate to the preset maximum rate value; Set a dedicated RI Disable flag; Alternatively, send a dedicated RI to revoke the control message.

[0086] Step 340: If the lifecycle state is in INCAST state, generate RI information containing the current effective speed limit value, and send the RI information to each sender in the list of currently active senders.

[0087] In some embodiments, the method of sending the RI information includes: The RI information is sent via ACK, SACK, NACK, or extended control messages.

[0088] Specifically, during the Incast state, the receiving end can continuously announce RI information to the corresponding remote QP via RDMA ACK / SACK / NACK or extended control messages.

[0089] The cooperative congestion control method based on INCAST lifecycle state awareness provided in this embodiment, compared with the prior art: 1. It can coordinate the transmission rate in the early stages of Incast, reducing instantaneous queue backlog.

[0090] Traditional DCQCN is a passive feedback control mechanism based on network congestion events. The sending end only receives a CNP and begins to reduce its speed after the switch has already experienced significant queue backlog and triggered an ECN flag. Therefore, in large-scale incast scenarios, multiple sending ends often inject data into the network simultaneously at high transmission rates for extended periods, easily leading to a rapid and sudden increase in switch queues.

[0091] In this invention, the receiver, by maintaining an active remote QP list and an Incast state machine, can identify Incast establishment trends in advance during the rapid aggregation phase of multiple transmitters. Before the switch generates a clear ECN marker, the receiver can calculate the fair transmission rate in advance and notify each transmitter to limit the transmission rate through RI information.

[0092] Therefore, compared with traditional DCQCN, this invention can: suppress initial traffic bursts in Incast in advance; reduce the instantaneous growth rate of switch output queues; reduce ECN tag bursts; reduce PFC trigger probability; reduce network queuing latency; and improve network stability in Incast scenarios.

[0093] 2. It can improve convergence speed and bandwidth fairness in large-scale incast scenarios.

[0094] In traditional sender-independent congestion control mechanisms, each sender typically performs deceleration and acceleration control independently based on its own received congestion feedback. The lack of unified coordination among multiple senders easily leads to: some senders occupying high bandwidth for extended periods; inconsistent deceleration timings among different senders; multiple rounds of deceleration-acceleration oscillations; and long incast convergence times.

[0095] In this invention, the receiving end uniformly maintains the number of active remote QPs currently participating in Incast and calculates the suggested transmission rate for each sender based on: Rfair = B / N. Therefore, in Incast scenarios, each sender can quickly converge to a near-fair bandwidth allocation state without experiencing prolonged contention and oscillation.

[0096] Especially in the INCAST_GROWTH state, this invention directly uses the number of real-time active remote QPs, Nactive(t), to calculate the fair rate, thereby quickly reducing the transmission rate of each sender and suppressing the initial queue growth of Incast. In the INCAST_STABLE state, this invention further uses: Navg(t) = (1 β)Navg(t The function Tdetect)+βNactive(t) smooths the number of active remote QPs, thereby reducing rate fluctuations caused by frequent short-stream entry and exit, and improving the stability of the transmission rate.

[0097] Therefore, the present invention can: improve bandwidth fairness among multiple transmitters; shorten Incast convergence time; reduce transmission rate oscillation; and improve throughput stability.

[0098] 3. Retain the ability of traditional DCQCN to adapt to real-world network congestion.

[0099] Traditional Receiver-driven solutions typically assume that the receiver's outgoing bandwidth is the same as the available network bandwidth. Therefore, when the network experiences issues such as ECMP hash collisions, localized hotspots, link failures, oversubscription, or dynamic path changes, the receiver may not accurately perceive the true bottleneck bandwidth within the network.

[0100] This invention does not replace the traditional DCQCN ECN / CNP congestion feedback mechanism, but rather adds an RI (Instruction Rate Limiting) coordination mechanism. RI is responsible for advance fair coordination in Incast scenarios, while ECN / CNP remains responsible for real congestion feedback within the network. When there is real bottleneck congestion within the network, even if the Instruction_Rate given by the receiver is high, the sender can still continue to execute the traditional DCQCN rate reduction process based on ECN / CNP.

[0101] 4. It can reduce the overhead of control messages in Receiver-driven methods.

[0102] Traditional receiver-driven methods based on Credit / Grant typically require periodically sending Credit messages, with the sender relying on Credit messages to allow data transmission. When the amount of data corresponding to a single Credit message is small, the receiver needs to send Credit control messages at a high frequency to ensure smooth traffic flow, resulting in significant control plane overhead.

[0103] In this invention, RI is not used as a transmission permission, but only as a suggested transmission rate value. The sending end does not need to wait for Credit before allowing data transmission. This invention eliminates the need to periodically send a large number of independent Credit messages and maintain a complex Credit-consuming state machine. This reduces the number of control messages, reduces control plane bandwidth usage, reduces protocol implementation complexity, and improves the effective utilization of service bandwidth.

[0104] Example 2: The embodiments of the present invention relate to a collaborative congestion control method based on INCAST lifecycle state awareness, which can be applied to the transmitting end.

[0105] In this invention, the receiving end is used to count how many active sending ends are simultaneously sending data to it, calculate the fair rate, and distribute it to each sending end via RI information. The sending end, based on the received RI information, limits its own transmission rate according to instructions, while retaining the traditional DCQCN's ability to respond to internal network congestion (e.g., ECN / CNP).

[0106] Traditional ECN-based sender-driven congestion control mechanisms typically rely on the switch's output queue reaching the ECN marking threshold before triggering subsequent congestion feedback. By the time the sender actually perceives congestion, a significant queue backlog has usually already formed within the switch. In this invention, the receiver can identify incast establishment trends formed by the rapid aggregation of multiple senders in advance. This allows for earlier initiation of the sender-coordinated rate-limiting process, thereby reducing the instantaneous queue growth at the initial stage of incast.

[0107] The following is a detailed description of the implementation details of the cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0108] The cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities.

[0109] It should be noted that the sending end in this embodiment still retains all the core mechanisms of the traditional DCQCN, specifically including: The system slowed down upon receiving congestion feedback from the CNP. α-congestion factor update; Increases triggered by timers and data transfer counters, including but not limited to Fast Recovery and Additive Increase.

[0110] In addition, at least the following technical features have been added: (1) RI information extraction mechanism at the receiving end; (2) Based on the RI transmission rate limiting mechanism, during the period when the RI is in effect, the transmission rate of the sending end must not exceed the Instruction_Rate; (3) During the period when RI is in effect, the normal rate reduction process of DCQCN based on CNP will not be affected, and the sending end can continue to respond to the internal network congestion feedback to further reduce the sending rate; (4) RI failure mechanism: When RI fails, Fast Recovery and Additive Increase are used to restore the transmission rate.

[0111] Specifically, such as Figure 4 As shown, the cooperative congestion control method based on INCAST lifecycle state awareness provided in this embodiment includes the following steps: Step 410: In response to receiving RI information sent by the receiving end, determine whether the RI information contains RI cancellation information.

[0112] After receiving the RI message, the sending end first determines whether the RI belongs to the "cancel RI" message, that is, whether the receiving end notifies the sending end that the current Incast has ended and can exit the RI control mode.

[0113] Since the RI information may contain dedicated RI cancellation control information, upon receiving RI information sent by the receiving end, it is first determined whether the RI information contains RI cancellation information. RI cancellation information is used to notify the sending end to exit RI control mode.

[0114] Specifically, when the receiving end detects the end of the incast, it sends RI information with a cancellation flag (cancellation RI information) to the sending end. Upon receiving the RI information with the cancellation flag, the sending end terminates the RI Timer early and triggers the expiration process.

[0115] The sending end extracts the RI information during the corresponding message reception and processing flow and updates the Instruction_Rate in the corresponding QP context. Simultaneously, it creates or resets the RI Timer, with the timer expiration time being RI_Expired_Time = now + RI_Effective_Time, where now is the current time and RI_Effective_Time is the locally configured default RI duration. For the complete process of receiving and processing RI information, please refer to [reference needed]. Figure 5 .

[0116] Step 420: If the RI information does not contain information to cancel the RI, enter the RI control mode and create or reset the RI timer to start timing. Obtain the current effective rate limit value from the RI information and update the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, control the current transmission rate of the transmitter to be no greater than the current effective rate limit value.

[0117] The sender must ensure that during the RI period (i.e., before the RI Timer reaches its trigger threshold), the current transmission rate does not exceed Instruction_Rate. To achieve this, the sender not only needs to limit the current transmission rate but also needs to coordinate the control of the DCQCN internal state and subsequent recovery process. Key processing includes: during the RI period, when the sender receives a CNP congestion signal, it should reduce the transmission rate according to the traditional DCQCN processing method, while ensuring that the subsequent rate increase can only reach a maximum of Instruction_Rate.

[0118] Specifically, if the received RI message does not include a release flag, the normal RI processing flow begins. First, the Instruction_Rate is parsed from the RI message and saved to the QP context or congestion control context. This is not merely a temporary storage of a rate limit value; rather, it truly integrates the RI into the DCQCN's internal state management system, including saving the Instruction_Rate, RI status, and related timeout information. This design means that the RI is no longer an external rate limiter, but rather part of the DCQCN control logic.

[0119] The sending end then creates or refreshes the RI Expire Timer, with the logic being "RI_Expired_Time = now + Effective_Time". This means that RI is a soft-state control mechanism, requiring the receiving end to refresh the RI periodically. If no new RI is received for an extended period, the sending end automatically exits RI mode and resumes normal DCQCN behavior. This prevents the sending end from permanently remaining in a rate-limited state due to lost control messages or abnormal conditions.

[0120] Next, the DCQCN target transmission rate Rt is updated. In this embodiment, the update method used is as follows: Rt = min(Rt,Instruction_Rate).

[0121] That is, by directly modifying Rt, RI is made a true part of the internal control state of DCQCN. Here, Rt is the target transmission rate, and Instruction_Rate is the current effective rate limit value.

[0122] The purpose of this mechanism is not simply to limit the current transmission rate, but to smoothly control the subsequent rate recovery process based on TC and BC mechanisms when the transmission flow has already been slowed down by CNP-based DCQCN due to network congestion, and the current rate is already lower than Instruction_Rate. By limiting Rt, it can avoid excessively fast recovery, rate oscillations, or transient bursts after RI is removed during the DCQCN recovery phase due to excessively high historical target rates, thus achieving a smoother and more stable rate recovery process.

[0123] In some embodiments, after updating the target transmission rate of the sending end, the method further includes: If the current transmission rate is greater than the current effective rate limit, then the current effective rate limit is used as the current transmission rate, and the counting of timer count TC and / or byte count BC is paused.

[0124] Specifically, after updating Rt, the process further determines whether the current sending rate Rc is still higher than Instruction_Rate. If Rc is still greater than Instruction_Rate, then "Rc = Instruction_Rate" is immediately executed, i.e., the current sending rate is limited. The purpose of this is to quickly suppress incast, thereby timely suppressing sudden incast traffic. After Rc is synchronized to Instruction_Rate, the process further disables the TC and BC counting mechanisms. Here, TC and BC correspond to the Timer Counter and Byte Counter in DCQCN, which are responsible for driving the periodic speed-up and recovery logic of DCQCN. If these mechanisms continue to run during RI, DCQCN will continuously attempt to autonomously increase its speed, thus conflicting with the RI control logic, eventually causing the sending rate to gradually increase again, triggering new queue backlogs and system oscillations. Therefore, this scheme suspends the DCQCN autonomous speed-up mechanism during RI, making receiver coordination the dominant control logic at this stage, thereby achieving a more stable congestion control effect. On the other hand, if the current Rc is already lower than or equal to Instruction_Rate, it indicates that the sending flow has already undergone a CNP-based DCQCN rate-down process due to network congestion, and the current rate has naturally decreased to within the range allowed by the receiver. In this case, the role of RI is mainly reflected in adjusting Rt to control the dynamic behavior of the subsequent recovery phase without further freezing the TC and BC mechanisms. Retaining the original recovery logic of DCQCN at this time allows for a smoother link utilization recovery by leveraging its gradual recovery characteristics, avoiding the sending rate remaining at an excessively low level for an extended period due to over-freezing of the recovery mechanism, thus balancing incast suppression effects with link utilization recovery efficiency.

[0125] In some embodiments, the method further includes: If the RI timer has expired, exit the RI control mode and set the current effective rate limit value to the maximum line speed of the transmitting link, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC.

[0126] Specifically, regarding the handling of RI Timer expiration: In a preferred implementation, when the RI Timer expires, the transmitter exits the RI active state and sets the Instruction_Rate to the maximum line rate of the transmitter link, while simultaneously restarting or restoring the TC and BC counting mechanisms. By restoring the Instruction_Rate to the maximum line rate of the link, the speed-up processing flows such as Fast Recovery and Additive Increase can continue to use a unified rate-limiting logic without the need for additional checks on whether the RI is active, thereby simplifying the transmitter's congestion control state machine and hardware implementation complexity.

[0127] In some embodiments, after obtaining the current effective speed limit value from the RI information, the method further includes: If the current transmission rate at the transmitting end is not greater than the current effective rate limit value, the counting accumulation of timer count TC and byte count BC remains active, and during the rate-up process triggered by timer count TC and / or byte count BC, the rate-up control of the transmitting end is restored to the upper limit value at the current effective rate limit value.

[0128] Specifically, modifications to the Fast Recovery and Additive Increase processes may include: During the RI period, when the current rate Rc of the transmitter is lower than or equal to the Instruction_Rate, the DCQCN can still continue to execute the speed-up processes such as Fast Recovery and Additive Increase triggered by TC and BC. However, during the speed-up process, limiting logic needs to be added to ensure that Rc and Rt do not exceed the Instruction_Rate, thereby ensuring that the transmission rate is always coordinated and controlled by the receiver during the RI period.

[0129] Step 430: If the RI information contains RI cancellation information and the RI timer has not timed out, control the RI timer to stop counting, exit the RI control mode, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC.

[0130] If the RI information includes RI cancellation information, the sending end further determines whether the local RI is still in effect, i.e., whether the local RI Timer has not yet expired. If the local RI is still in effect, the RI Timer is terminated early, the RI control mode is immediately exited, and normal DCQCN behavior is restored; if the local RI has expired, no additional processing is required, and the process ends directly.

[0131] The cooperative congestion control method based on INCAST lifecycle state awareness provided in this embodiment, compared with the prior art: 1. It can coordinate the transmission rate in the early stages of Incast, reducing instantaneous queue backlog.

[0132] Traditional DCQCN is a passive feedback control mechanism based on network congestion events. The sending end only receives a CNP and begins to reduce its speed after the switch has already experienced significant queue backlog and triggered an ECN flag. Therefore, in large-scale incast scenarios, multiple sending ends often inject data into the network simultaneously at high transmission rates for extended periods, easily leading to a rapid and sudden increase in switch queues.

[0133] In this invention, the receiver, by maintaining an active remote QP list and an Incast state machine, can identify Incast establishment trends in advance during the rapid aggregation phase of multiple transmitters. Before the switch generates a clear ECN marker, the receiver can calculate the fair transmission rate in advance and notify each transmitter to limit the transmission rate through RI information.

[0134] Therefore, compared with traditional DCQCN, this invention can: suppress initial traffic bursts in Incast in advance; reduce the instantaneous growth rate of switch output queues; reduce ECN tag bursts; reduce PFC trigger probability; reduce network queuing latency; and improve network stability in Incast scenarios.

[0135] 2. It can improve convergence speed and bandwidth fairness in large-scale incast scenarios.

[0136] In traditional sender-independent congestion control mechanisms, each sender typically performs deceleration and acceleration control independently based on its own received congestion feedback. The lack of unified coordination among multiple senders easily leads to: some senders occupying high bandwidth for extended periods; inconsistent deceleration timings among different senders; multiple rounds of deceleration-acceleration oscillations; and long incast convergence times.

[0137] In this invention, the receiving end uniformly maintains the number of active remote QPs currently participating in Incast and calculates the suggested transmission rate for each sender based on: Rfair = B / N. Therefore, in Incast scenarios, each sender can quickly converge to a near-fair bandwidth allocation state without experiencing prolonged contention and oscillation.

[0138] Especially in the INCAST_GROWTH state, this invention directly uses the number of real-time active remote QPs, Nactive(t), to calculate the fair rate, thereby quickly reducing the transmission rate of each sender and suppressing the initial queue growth of Incast. In the INCAST_STABLE state, this invention further uses: Navg(t) = (1 β)Navg(t The function Tdetect)+βNactive(t) smooths the number of active remote QPs, thereby reducing rate fluctuations caused by frequent short-stream entry and exit, and improving the stability of the transmission rate.

[0139] Therefore, the present invention can: improve bandwidth fairness among multiple transmitters; shorten Incast convergence time; reduce transmission rate oscillation; and improve throughput stability.

[0140] 3. Retain the ability of traditional DCQCN to adapt to real-world network congestion.

[0141] Traditional Receiver-driven solutions typically assume that the receiver's outgoing bandwidth is the same as the available network bandwidth. Therefore, when the network experiences issues such as ECMP hash collisions, localized hotspots, link failures, oversubscription, or dynamic path changes, the receiver may not accurately perceive the true bottleneck bandwidth within the network.

[0142] This invention does not replace the traditional DCQCN ECN / CNP congestion feedback mechanism, but rather adds an RI (Instruction Rate Limiting) coordination mechanism. RI is responsible for advance fair coordination in Incast scenarios, while ECN / CNP remains responsible for real congestion feedback within the network. When there is real bottleneck congestion within the network, even if the Instruction_Rate given by the receiver is high, the sender can still continue to execute the traditional DCQCN rate reduction process based on ECN / CNP.

[0143] 4. It can reduce the overhead of control messages in Receiver-driven methods.

[0144] Traditional receiver-driven methods based on Credit / Grant typically require periodically sending Credit messages, with the sender relying on Credit messages to allow data transmission. When the amount of data corresponding to a single Credit message is small, the receiver needs to send Credit control messages at a high frequency to ensure smooth traffic flow, resulting in significant control plane overhead.

[0145] In this invention, RI is not used as a transmission permission, but only as a suggested transmission rate value. The sending end does not need to wait for Credit before allowing data transmission. This invention eliminates the need to periodically send a large number of independent Credit messages and maintain a complex Credit-consuming state machine. This reduces the number of control messages, reduces control plane bandwidth usage, reduces protocol implementation complexity, and improves the effective utilization of service bandwidth.

[0146] Example 3: The embodiments of the present invention relate to a collaborative congestion control method based on INCAST lifecycle state awareness.

[0147] In this invention, the receiving end is used to count how many active sending ends are simultaneously sending data to it, calculate the fair rate, and distribute it to each sending end via RI information. The sending end, based on the received RI information, limits its own transmission rate according to instructions, while retaining the traditional DCQCN's ability to respond to internal network congestion (e.g., ECN / CNP).

[0148] Traditional ECN-based sender-driven congestion control mechanisms typically rely on the switch's output queue reaching the ECN marking threshold before triggering subsequent congestion feedback. By the time the sender actually perceives congestion, a significant queue backlog has usually already formed within the switch. In this invention, the receiver can identify incast establishment trends formed by the rapid aggregation of multiple senders in advance. This allows for earlier initiation of the sender-coordinated rate-limiting process, thereby reducing the instantaneous queue growth at the initial stage of incast.

[0149] The following is a detailed description of the implementation details of the cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0150] The cooperative congestion control method based on INCAST lifecycle state awareness in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities.

[0151] like Figure 6 As shown, the cooperative congestion control method based on INCAST lifecycle state awareness provided in this embodiment includes the following steps: Step 610: In response to receiving a data packet sent by the sender, the receiver records the unique identifier information corresponding to the sender in the list of currently active senders and updates the last active timestamp of the sender. Step 620: The receiving end determines the lifecycle state of the INCAST state machine based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage; the lifecycle state includes the NORMAL state and the INCAST in progress state. Step 630: The receiving end determines the contention factor based on the lifecycle state, and determines the current effective rate limit value according to the available bandwidth of the receiving end and the contention factor, and stores the current effective rate limit value in association with the record item corresponding to the sending end in the list of currently active sending ends; wherein, the contention factor is determined in different ways under different lifecycle states; Step 640: When the lifecycle state is in INCAST state, generate RI information containing the current effective rate limit value at the receiving end, and send the RI information to each sending end in the list of currently active sending ends. Step 650: In response to receiving RI information sent by the receiving end, the sending end determines whether the RI information contains RI cancellation information; Step 660: If the RI information does not contain information to cancel the RI, control the transmitter to enter the RI control mode and create or reset the RI timer to start timing. Obtain the current effective rate limit value from the RI information and update the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, control the current transmission rate of the transmitter to be no greater than the current effective rate limit value. Step 670: If the RI information contains RI cancellation information and the RI timer has not timed out, control the RI timer at the transmitting end to stop counting, exit the RI control mode, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate and Instruction_Rate is the current effective rate limit value.

[0152] The specific implementation process of each step of the method provided in this embodiment can be referred to the foregoing embodiments, and will not be repeated in this embodiment.

[0153] Finally, it should be noted that: 1. Regarding the method of carrying RI information: In this embodiment, the RI information is preferably carried through RDMA ACK / SACK / NACK messages. In other implementations, the RI information can also be implemented in the following ways: based on extended proprietary control messages, based on independent UDP control messages, based on data packet reverse path piggyback, based on RoCEv2, IPv4, or IPv6 extended header fields, or other methods capable of enabling rate coordination information transmission between the sender and receiver. This invention does not constitute a limitation.

[0154] 2. Incast lifecycle state detection: In this embodiment, the receiving end primarily performs Incast lifecycle state detection based on the number of active far QPs, aggregated traffic change trends, buffer usage change trends, and link utilization change trends. In other implementations, it can further incorporate: ECN tagging ratio, RTT change trends, receive buffer queuing latency, PFC triggering status, NIC internal queue status, or other statistical information that reflects the degree of traffic aggregation and network congestion trends to perform Incast state identification and state transition control.

[0155] 3. Regarding the fair rate calculation method: In this embodiment, the receiver preferably uses Rfair = B / N to calculate the suggested transmission rate for each sender. In other implementations, weighted fair rate calculation can also be performed based on: QoS level, flow priority, tenant weight, flow size, historical transmission rate, or other bandwidth allocation strategies.

[0156] For example: Ri = B × Wi / ΣW Where Wi represents the weight factor of the corresponding sending stream.

[0157] 4. Regarding the RI activation and deactivation mechanism: In this embodiment, the sending end removes the RI rate-limiting state through two methods: RI Expire Timer and explicit RI Release. Other implementations may also employ: automatic exit based on not receiving RI for multiple consecutive cycles, automatic exit based on partial congestion at the sending end, exit based on explicit state synchronization at the receiving end, or other timeout and state synchronization mechanisms to achieve RI lifecycle management.

[0158] 5. Methods for restoring transmission rate after RI expiration: In this embodiment, in a preferred implementation, when the RI Timer expires, the transmitter exits the RI active state and restores or restarts the TC and BC counting mechanisms, gradually restoring the transmission rate through the Fast Recovery and Additive Increase processes. In other implementations, when the RI Timer expires, the transmitter can also directly restore: Rc, Rt to the maximum line speed corresponding to the transmitter interface, or restore to the preset default transmission rate, without needing to restore through the gradual speed-up process driven by TC and BC. This method can further reduce the recovery latency after RI exit, thereby improving the link utilization recovery speed in short-term incast scenarios. This invention does not constitute a limitation.

[0159] 6. Regarding the method of RI updating the internal state of DCQCN: In this embodiment, in a preferred implementation, after receiving the RI, the sender preferably updates Rt synchronously and, if necessary, further adjusts Rc synchronously, thereby enabling the RI coordination rate to be coordinated and integrated with the congestion control state within the traditional DCQCN. In other implementations, after receiving the RI, the sender may only limit the actual transmission rate or only update Rc without modifying Rt. For example, the sender may maintain the traditional DCQCN's internal Rt unchanged and only add a transmission rate limit in the form of Rsend=min(Rdcqcn,Instruction_Rate) to the actual transmission rate. Or, only synchronously adjust the current transmission rate Rc while keeping the traditional DCQCN's target transmission rate Rt unchanged.

[0160] In the above implementation, RI can serve as: external transmission rate constraint, auxiliary coordination control information, or independent rate limiting mechanism, working in conjunction with traditional DCQCN. This invention does not constitute a limitation in this regard.

[0161] Example 4: Another embodiment of this application relates to a collaborative congestion control device based on INCAST lifecycle state awareness.

[0162] The following provides a detailed description of the implementation details of the cooperative congestion control device based on INCAST lifecycle state awareness in this embodiment. The following details are provided for ease of understanding and are not essential for implementing this solution. The cooperative congestion control device based on INCAST lifecycle state awareness provided in this embodiment includes: The receiver identification information update module is used to, in response to receiving a data packet sent by the sender, record the unique identification information corresponding to the sender in the list of currently active senders at the receiver and update the last active timestamp of the sender. The receiver lifecycle state determination module is used to determine the lifecycle state of the INCAST state machine based on the changing trend of the number of active transmitters within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage. The lifecycle state includes the NORMAL state and the INCAST in progress state. The receiver effective rate limit determination module is used to determine a contention factor based on the lifecycle state, and determine the current effective rate limit value according to the available bandwidth of the receiver and the contention factor, and store the current effective rate limit value in association with the record item corresponding to the sender in the list of currently active senders; wherein, the contention factor is determined in different ways under different lifecycle states; The receiving end RI information determination module is used to generate RI information containing the current effective speed limit value at the receiving end when the life cycle state is in the INCAST state, and send the RI information to each transmitter located in the list of currently active transmitters. The sending end determination module is used to determine whether the RI information contains RI cancellation information in response to receiving RI information sent by the receiving end. The transmitter rate limiting module is used to control the transmitter to enter the RI control mode and create or reset the RI timer to start timing when the RI information does not contain the RI cancellation information; to obtain the current effective rate limiting value from the RI information and update the target transmission rate of the transmitter; if the current transmission rate of the transmitter is greater than the current effective rate limiting value, then control the current transmission rate of the transmitter not to be greater than the current effective rate limiting value. The transmitting end rate limit removal module is used to control the transmitting end's RI timer to stop counting, exit the RI control mode, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC when the RI information contains RI removal information and the RI timer has not expired. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate and Instruction_Rate is the current effective rate limit value.

[0163] The specific implementation process of each step of the method provided in this embodiment can be referred to the foregoing embodiments, and will not be repeated in this embodiment.

[0164] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of each module in this collaborative congestion control device based on INCAST lifecycle state awareness can be referred to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0165] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0166] Example 5: Another embodiment of this application relates to an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the above embodiments.

[0167] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0168] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0169] Example 6: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0170] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0171] In some embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the methods described in the above embodiments.

[0172] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. A cooperative congestion control method based on INCAST lifecycle state awareness, characterized in that, Applied to the receiving end, the method includes: In response to receiving a data packet sent by the sender, the unique identifier information corresponding to the sender is recorded in the list of currently active senders, and the last active timestamp of the sender is updated; Based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage, the lifecycle state of the INCAST state machine is determined; the lifecycle state includes the NORMAL state and the INCAST in progress state. The contention factor is determined based on the lifecycle state, and the current effective rate limit value is determined according to the available bandwidth of the receiver and the contention factor. The current effective rate limit value is then associated and stored in the record item corresponding to the transmitter in the list of currently active transmitters. The contention factor is determined in different ways under different lifecycle states. When the lifecycle state is in INCAST state, RI information containing the current effective speed limit value is generated, and the RI information is sent to each sender in the list of currently active senders; The lifecycle state of the INCAST state machine represents the development stage of aggregated traffic from multiple active sender QPs. The NORMAL state indicates that no INCAST aggregation has occurred yet, and the INCAST in progress state indicates that INCAST aggregation has already occurred. The INCAST in progress state includes at least one of the following: INCAST_GROWTH state, indicating that the number of active sender QPs is continuously increasing; INCAST_STABLE state, indicating that INCAST has been formed and the number of active sender QPs has entered a preset stable fluctuation stage; INCAST_DECAY state, indicating that the number of active sender QPs is continuously decreasing. The contention factor characterizes the number of transmitters currently participating in INCAST bandwidth contention or the degree of contention. Determining the contention factor based on the lifecycle state includes: If the lifecycle state is INCAST_GROWTH or INCAST_DECAY, then the number of active transmitters is obtained, and the number of active transmitters is used as the value of the contention factor. If the lifecycle state is INCAST_STABLE, then the value of the competition factor is determined by the following formula: N = Navg(t) = (1 β)Navg(t Tdetect)+βNactive(t) Where N is the competition factor, β is the weighting coefficient, Tdetect is the statistical period or detection time window used by the receiver for INCAST state detection, and Nactive(t) is the number of active transmitters at time t.

2. The method according to claim 1, characterized in that, The unique identifier information includes at least one of the following: The QP number corresponding to the sender, the addressing information corresponding to the sender, and the last active timestamp of the sender.

3. The method according to claim 1, characterized in that, The methods for sending the RI information include: The RI information is sent via ACK, SACK, NACK, or extended control messages.

4. A cooperative congestion control method based on INCAST lifecycle state awareness, characterized in that, Applied to the sending end, the method includes: In response to receiving RI information sent by the receiving end, it is determined whether the RI information contains RI cancellation information; the RI information is information generated by the receiving end when the lifecycle state is INCAST and sent to each sending end in the list of currently active sending ends, and the RI information contains the current effective rate limit value. If the RI information does not contain information to cancel the RI, enter the RI control mode and create or reset the RI timer to start timing. Obtain the current effective rate limit value from the RI information and update the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, control the current transmission rate of the transmitter to be no greater than the current effective rate limit value. If the RI information contains RI cancellation information and the RI timer has not timed out, control the RI timer to stop counting, exit the RI control mode, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate, and Instruction_Rate is the current effective rate limit value; The INCAST state machine's lifecycle state represents the development stage of aggregated traffic from multiple active sender QPs. The lifecycle state includes a NORMAL state and an INCAST in progress state. The NORMAL state indicates that INCAST aggregation has not yet occurred, while the INCAST in progress state indicates that INCAST aggregation has already occurred. The INCAST in progress state includes at least one of the following: INCAST_GROWTH state, indicating that the number of active sender QPs is continuously increasing; INCAST_STABLE state, indicating that INCAST has been formed and the number of active sender QPs has entered a preset stable fluctuation phase; and INCAST_DECAY state, indicating that the number of active sender QPs is continuously decreasing.

5. The method according to claim 4, characterized in that, After updating the target transmission rate at the sending end, the method further includes: If the current transmission rate is greater than the current effective rate limit, then the current effective rate limit is used as the current transmission rate, and the counting of timer count TC and / or byte count BC is paused.

6. The method according to claim 4, characterized in that, The method further includes: If the RI timer has expired, exit the RI control mode and set the current effective rate limit value to the maximum line speed of the transmitting link, and restore the autonomous rate growth mechanism driven by timer count TC and / or byte count BC.

7. The method according to claim 4, characterized in that, After obtaining the current effective speed limit value from the RI information, the method further includes: If the current transmission rate at the transmitting end is not greater than the current effective rate limit value, the counting accumulation of timer count TC and byte count BC remains active, and during the rate-up process triggered by timer count TC and / or byte count BC, the rate-up control of the transmitting end is restored to the upper limit value at the current effective rate limit value.

8. A cooperative congestion control method based on INCAST lifecycle state awareness, characterized in that, The method includes: In response to receiving a data packet sent by the sender, the receiver records the unique identifier information corresponding to the sender in the list of currently active senders and updates the last active timestamp of the sender. The receiving end determines the lifecycle state of the INCAST state machine based on the changing trend of the number of active senders within a preset detection period, the rate of change of aggregated traffic, or the rate of change of buffer usage; the lifecycle state includes the NORMAL state and the INCAST in progress state. The receiving end determines the contention factor based on the lifecycle state, and determines the current effective rate limit value according to the available bandwidth of the receiving end and the contention factor, and stores the current effective rate limit value in association with the record item corresponding to the sending end in the list of currently active sending ends; wherein, the contention factor is valued differently in different lifecycle states; When the lifecycle state is in INCAST state, RI information containing the current effective rate limit value is generated at the receiving end, and the RI information is sent to each sending end in the list of currently active sending ends; In response to receiving RI information sent by the receiving end, the sending end determines whether the RI information contains RI cancellation information; If the RI information does not contain RI cancellation information, the control transmitter enters the RI control mode and creates or resets the RI timer to start timing. It obtains the current effective rate limit value from the RI information and updates the target transmission rate of the transmitter. If the current transmission rate of the transmitter is greater than the current effective rate limit value, the control transmitter's current transmission rate is made not to exceed the current effective rate limit value. If the RI information contains RI cancellation information and the RI timer has not expired, the RI timer at the control end is stopped, the RI control mode is exited, and the autonomous rate growth mechanism driven by timer count TC and / or byte count BC is restored. The target transmission rate of the sending end is updated based on the following formula: Rt = min(Rt,Instruction_Rate) Where Rt is the target transmission rate, and Instruction_Rate is the current effective rate limit value; The lifecycle state of the INCAST state machine represents the development stage of aggregated traffic from multiple active sender QPs. The NORMAL state indicates that no INCAST aggregation has occurred yet, and the INCAST in progress state indicates that INCAST aggregation has already occurred. The INCAST in progress state includes at least one of the following: INCAST_GROWTH state, indicating that the number of active sender QPs is continuously increasing; INCAST_STABLE state, indicating that INCAST has been formed and the number of active sender QPs has entered a preset stable fluctuation stage; INCAST_DECAY state, indicating that the number of active sender QPs is continuously decreasing. The contention factor characterizes the number of transmitters currently participating in INCAST bandwidth contention or the degree of contention. Determining the contention factor based on the lifecycle state includes: If the lifecycle state is INCAST_GROWTH or INCAST_DECAY, then the number of active transmitters is obtained, and the number of active transmitters is used as the value of the contention factor. If the lifecycle state is INCAST_STABLE, then the value of the competition factor is determined by the following formula: N = Navg(t) = (1 β)Navg(t Tdetect)+βNactive(t) Where N is the competition factor, β is the weighting coefficient, Tdetect is the statistical period or detection time window used by the receiver for INCAST state detection, and Nactive(t) is the number of active transmitters at time t.

Citation Information

Patent Citations

  • Rapid congestion feedback method based on DCTCP

    CN111464452A

  • Congestion control method and device, computer equipment and storage medium

    CN118984307A