Ethernet switch congestion control method supporting RoCEv2 lossless transmission

CN122824673APending Publication Date: 2026-09-25WUHAN CHAOQING DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611239349.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]提供一种支持RoCEv2无损传输的以太网交换机拥塞控制方法,以解决现有拥塞检测维度单一、无法及时预判队列变化以及缺少流间干扰量化、标记精度不足的问题,通过融合流-端口-时间三维特征的动态阈值映射和基于流间干扰因子矩阵的时空耦合标记调整,实现更精准的拥塞控制

Benefits of technology

动态阈值映射处理采用由历史拥塞事件样本点构成的二维曲面插值函数,以当前端口的瞬时缓冲区占用率及当前流的瞬时发送速率作为输入变量,在控制点集合形成的网格区域内执行双三次样条插值计算,生成拥塞预判索引值。该处理将历史拥塞发生时缓冲区占用率、瞬时发送速率和拥塞等级评分的对应关系融入插值曲面,使得拥塞预判不仅依赖当前单一队列深度,还能结合发送端注入速率和端口缓冲区占用率的动态组合,形成非线性的多维映射。相比固定阈值或仅依赖队列深度的线性映射方式,该映射处理能够捕捉流速率与缓冲区占用率之间的耦合变化趋势,在队列尚未达到预设警戒线时即可根据趋势提前输出更敏感的预判索引值,从而在突发流量场景下缩减拥塞标记的响应时延,减少PFC触发频次。在根据拥塞预判索引值执行显式拥塞通知标记调整时,构建流间干扰因子矩阵,矩阵元素由共享同一下行端口的两个RoCEv2流在时间维度上的发送窗口重叠概率和空间维度上的数据包交织度计算所得。将当前RoCEv2流的拥塞预判索引值乘以当前流与同端口其他所有流的流间干扰因子之和,生成本地标记概率修正值。发送窗口重叠概率通过统计两个流在序列号空间上存在重叠区间的时间点比例得到,数据包交织度则反映两个流的数据包在同一个端口输出队列中相互穿插的程度。这种加权修正使得标记概率不仅反映单一流自身引起的拥塞风险,还融入了该流与其他并行流之间的干扰强度。对于与多条流存在高度重叠和交织的流,其本地标记概率修正值相应放大,能够被优先标记以抑制对共享队列的冲击;而对于干扰程度较低的流,标记概率保持较低水平,避免不必要的速率回退。该机制实现了在共享端口上对相互干扰的流群组进行协调标记,缓解因不考虑流间关系而导致的带宽分配不公平和吞吐量抖动。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824673A_ABST
    Figure CN122824673A_ABST
Patent Text Reader

Abstract

The application discloses an Ethernet switch congestion control method supporting RoCEv2 lossless transmission and belongs to the technical field of network communication congestion control. The method collects the instantaneous sending rate, the instantaneous receiving confirmation rate and the priority flow suspension frame trigger frequency of each port based on the RoCEv2 flow identification as real-time message flow characteristic parameters; performs dynamic threshold mapping processing of the flow-port-time three-dimensional characteristics on the real-time message flow characteristic parameters, obtains the congestion prediction index value of each RoCEv2 flow at each port; performs weighted correction according to the congestion prediction index value and queries the inter-flow interference factor matrix, performs space-time coupled explicit congestion notification label adjustment, generates a local label probability correction value; and performs priority reverse pressure flow control based on the local label probability correction value, obtains a flow shaping parameter set containing a sending token bucket depth update value and a credit amount rollback step. The method can finely distinguish the inter-flow influence and realize more timely congestion response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication congestion control technology, specifically to a congestion control method for Ethernet switches that supports RoCEv2 lossless transmission. Background Technology

[0002] In Ethernet switches supporting RoCEv2 lossless transmission, to achieve low-latency, high-throughput lossless forwarding, a congestion management mechanism combining priority-based flow control and explicit congestion notification marking is typically relied upon. Existing solutions often use fixed or simple adaptive thresholds to trigger ECN marking, combined with PFC backpressure control of the sending end rate. These solutions primarily rely on macroscopic statistics such as port queue depth or average queue length for decision-making, treating all flows as homogeneous objects. The marking probability calculation lacks a fine-grained characterization of flow behavior differences and instantaneous port traffic dynamics. Existing technologies have significant shortcomings. On the one hand, the congestion detection thresholds or mapping functions are often statically configured or only linearly change with queue depth, failing to reflect the multi-dimensional characteristics of current active flows, such as sending rate, acknowledgment rate, and PFC trigger frequency, which are crucial for predicting queue change trends. This results in marking behavior lagging behind burst traffic, making effective intervention impossible before queue accumulation. On the other hand, among multiple RoCEv2 streams sharing the same downlink port, the degree of overlap of the data stream sending windows and the interleaving characteristics of data packets in the time dimension will generate complex inter-stream interference. Existing methods do not consider the mutual influence of this spatiotemporal coupling when calculating the marking probability, which makes it impossible for the marking to be accurately applied to the source streams that actually occupy buffer resources and interfere with each other, which can easily lead to over-marking or insufficient marking, thus impairing the overall throughput performance.

[0003] To address the aforementioned issues, existing limitations need to be overcome on two levels. First, an adaptive congestion prediction method needs to be established that integrates the three-dimensional characteristics of flow, port, and time. This method should utilize the correlations between multi-dimensional data from historical congestion events to accurately and instantly quantify the congestion level of the current network state. Second, a quantification mechanism for inter-flow interference needs to be introduced into the explicit congestion notification marking decision. This mechanism should incorporate the temporal and spatial interactions of flows within a shared port into the marking probability calculation, enabling the marking results to collaboratively adjust the injection rate of relevant sending ends. Summary of the Invention

[0004] This paper presents a congestion control method for Ethernet switches that supports RoCEv2 lossless transmission. This method addresses the problems of existing congestion detection methods, such as the single dimension, inability to predict queue changes in a timely manner, lack of inter-flow interference quantification, and insufficient marking accuracy. By fusing dynamic threshold mapping of flow-port-time three-dimensional features and spatiotemporal coupling marking adjustment based on the inter-flow interference factor matrix, more accurate congestion control is achieved.

[0005] To achieve the above objectives, this invention provides the following technical solution: This invention provides an Ethernet switch congestion control method supporting RoCEv2 lossless transmission. By jointly sensing and dynamically adjusting the network state from three dimensions—flow level, port level, and time level—it achieves precise control of RoCEv2 data plane congestion. This method collects real-time packet flow characteristic parameters from each port of the Ethernet switch. These real-time packet flow characteristic parameters include the instantaneous transmission rate based on each RoCEv2 flow identifier, the instantaneous receive acknowledgment rate based on each RoCEv2 flow identifier, and the priority traffic pause frame trigger frequency based on each port. Using these parameters to finely characterize flow behavior, it can comprehensively reflect the transmission rhythm, response characteristics, and backpressure intensity of each flow.

[0006] Subsequently, dynamic threshold mapping processing based on three-dimensional features of flow-port-time is performed on the real-time packet flow feature parameters to generate a congestion prediction index value for each RoCEv2 flow on each port. Unlike traditional methods that rely on fixed thresholds, this dynamic threshold mapping processing introduces a two-dimensional surface interpolation function composed of multiple historical congestion event sample points, using the instantaneous buffer occupancy rate of the current port and the instantaneous transmission rate of the current flow as input variables for this two-dimensional surface interpolation function. The historical congestion event sample points are collected and stored when the switch detects that the queue depth of any port exceeds a preset threshold, thus forming a set of control points. Each control point contains instantaneous buffer occupancy rate data, instantaneous transmission rate data, and corresponding congestion level score data. Using the instantaneous buffer occupancy rate of the current port as the input value on the first coordinate axis and the instantaneous transmission rate of the current flow as the input value on the second coordinate axis, bicubic spline interpolation calculation is performed within the grid area formed by the set of control points to obtain the congestion prediction index value. This interpolation method enables congestion prediction to automatically adapt to different load patterns and buffer utilization. In scenarios where burst traffic and stable high traffic coexist, the perception of congestion level is closer to the actual queue accumulation trend, thus providing continuous and high-resolution quantitative basis for subsequent labeling decisions.

[0007] After obtaining the congestion prediction index value, an explicit congestion notification label adjustment operation based on spatiotemporal coupling is further performed on each RoCEv2 flow according to the index value, generating a local label probability correction value for each RoCEv2 flow. Specifically, this adjustment operation queries a pre-constructed inter-flow interference factor matrix and performs a weighted correction on the congestion prediction index value. Preferably, the elements of the inter-flow interference factor matrix are calculated from the transmission window overlap probability of two RoCEv2 flows sharing the same downlink port in the time dimension and the packet interleaving degree in the spatial dimension. The transmission window overlap probability is determined as follows: the starting sequence number and window size of the transmission window of the first RoCEv2 flow at each time point are collected, as well as the starting sequence number and window size of the transmission window of the second RoCEv2 flow at each time point; for each time point, it is determined whether there is an overlapping interval between the transmission windows of the first RoCEv2 flow and the second RoCEv2 flow in the sequence number space; the ratio of the number of time points with overlapping intervals to the total number of sampling time points is used as the transmission window overlap probability. The local labeling probability correction value is generated by multiplying the congestion prediction index value of the current RoCEv2 flow by the sum of the inter-flow interference factors of the current RoCEv2 flow and all other RoCEv2 flows on the same port. By introducing the inter-flow interference factor matrix, this invention can quantify the mutual influence between adjacent flows due to temporal overlap and packet interleaving while considering the congestion risk of the flow itself. This allows the labeling probability to reflect the true competition relationship and avoids incorrectly assigning excessively high labels to isolated low-interference flows or under-labeling high-interference flows.

[0008] After obtaining the local label probability correction value, a priority-based reverse pressure flow control strategy is executed on the received RoCEv2 data packets based on this local label probability correction value, generating a traffic shaping parameter set. This traffic shaping parameter set includes the sending token bucket depth update value and credit rollback step size for each RoCEv2 flow. As a preferred embodiment of the present invention, the traffic shaping process includes: obtaining the current credit balance and the current depth value of the sending token bucket for each RoCEv2 flow; comparing the local label probability correction value with a preset probability threshold; when the local label probability correction value exceeds the probability threshold, setting the sending token bucket depth update value to the current depth value multiplied by a reduction coefficient based on the total port queue length, and setting the credit rollback step size to a positive increment based on the current flow service level. Preferably, the steps for obtaining the reduction factor based on the total port queue length are as follows: calculate the ratio of the current total port queue length to the total port buffer capacity to obtain the buffer occupancy rate; look up the basic reduction factor in a mapping table based on the buffer occupancy rate, which is a buffer occupancy rate-reduction factor mapping table pre-stored in the switch memory and contains multiple buffer occupancy rate intervals and the basic reduction factor corresponding to each interval; multiply the basic reduction factor by a decay factor based on the number of times the current RoCEv2 flow has been marked in history to generate the reduction factor based on the total port queue length. This mechanism enables the switch to not only apply forward congestion notification to the sender through ECN marking when the congestion risk is high, but also actively perform reverse pressure shaping on the downlink transmission rhythm of each flow at the switch side. The combination of token bucket tightening and credit backoff allows the flow rate to quickly converge to a level that the network can withstand, while the segmented reduction based on buffer occupancy rate avoids excessive limitation of throughput under light load.

[0009] As a further technical solution of the present invention, while implementing a reverse pressure flow control strategy based on local marker probability correction values, the change rate of the total queue depth of each port within multiple consecutive time windows is also collected. A dynamic priority weight for each port is calculated based on the change rate of the total queue depth. This dynamic priority weight is used to adjust the downlink scheduling polling order at the port level. Specifically, this involves: performing a differential operation on the change rate of the total queue depth over two consecutive time windows to generate a queue acceleration value; comparing the queue acceleration value with a preset acceleration threshold; when the queue acceleration value is greater than the acceleration threshold, setting the dynamic priority weight to a first weight value; and when the queue acceleration value is less than or equal to the acceleration threshold, setting the dynamic priority weight to a second weight value, where the first weight value is greater than the second weight value. By sensing the acceleration of queue consumption, ports that are about to enter deep congestion can be identified in advance, and their scheduling priority can be increased, thereby tilting limited scheduling opportunities towards high-risk ports and working in conjunction with flow-level shaping to suppress congestion deterioration.

[0010] This invention also provides an enhanced mechanism for globally consistent labeling. After generating the local label probability correction value for each RoCEv2 flow, the local label probability correction values ​​for all RoCEv2 flows are read and normalized to generate a normalized label probability vector. This normalized label probability vector is then input into a pre-stored port-level probability mapping table to retrieve the global forward explicit congestion notification label reference value corresponding to the current port. The port-level probability mapping table is indexed by the hash value of the normalized label probability vector. Preferably, the query process is as follows: extract the RoCEv2 flow identifiers corresponding to the largest components in the normalized label probability vector to form an active flow identifier set; perform a hash operation on the active flow identifier set to generate a hash query value; use this hash query value as an index to retrieve the port-level probability mapping table. If the search is successful, read the corresponding global forward explicit congestion notification label reference value; if the search is unsuccessful, use the average of the local label probability correction values ​​of all RoCEv2 flows in the active flow identifier set as the global forward explicit congestion notification label reference value, and write the correspondence between the hash query value and the calculated global forward explicit congestion notification label reference value into the port-level probability mapping table. Through this normalization and hash lookup method, the dispersed labeling decisions of each flow are mapped to a global reference label value for a port. This allows historical experience to be directly reused when similar patterns appear in the combination of active flows within a port, reducing computational overhead and improving the consistency of forward ECN labels in multi-flow scenarios, which is beneficial for the remote sender to perform more accurate rate adjustment.

[0011] In summary, the congestion control method of this invention integrates fine-grained flow state acquisition, three-dimensional dynamic threshold mapping, spatiotemporally coupled inter-flow interference correction, forward and reverse joint flow shaping, and scheduling weight adjustment based on port queue change rate at the switch side, forming a hierarchical lossless network congestion control system. Utilizing surface interpolation driven by historical congestion events reduces the probability of misjudgment; weighting inter-flow interference factors makes ECN marking more accurately reflect competition relationships; the synergistic effect of reverse pressure shaping and dynamic priority scheduling suppresses congestion propagation while ensuring low-latency transmission of high-priority flows; and the global marker reference value query mechanism enhances the continuity of the strategy and processing efficiency. These measures collectively guarantee the lossless transmission performance of RoCEv2 flows in an Ethernet environment.

[0012] The technical effects and advantages provided by the present invention in the above technical solution are as follows: The dynamic threshold mapping process employs a two-dimensional surface interpolation function composed of historical congestion event sample points. Using the instantaneous buffer occupancy rate of the current port and the instantaneous transmission rate of the current flow as input variables, it performs bicubic spline interpolation calculations within a grid region formed by the control point set to generate a congestion prediction index value. This process integrates the correspondence between buffer occupancy rate, instantaneous transmission rate, and congestion level score at historical congestion occurrences into the interpolation surface. This allows congestion prediction to not only rely on the current single queue depth but also combine the dynamic combination of the sender injection rate and port buffer occupancy rate, forming a non-linear, multi-dimensional mapping. Compared to fixed thresholds or linear mapping methods that only rely on queue depth, this mapping process can capture the coupled changing trend between flow rate and buffer occupancy rate. It can output a more sensitive prediction index value in advance based on the trend before the queue reaches the preset warning line, thereby reducing the response latency of congestion marking and decreasing the frequency of PFC triggering in sudden traffic scenarios. When performing explicit congestion notification marking adjustments based on the congestion prediction index value, an inter-flow interference factor matrix is ​​constructed. The matrix elements are calculated from the temporal window overlap probability and spatial packet interleaving degree of two RoCEv2 flows sharing the same downlink port. The local marking probability correction value is generated by multiplying the current RoCEv2 flow's congestion prediction index value by the sum of the inter-flow interference factors of the current flow and all other flows on the same port. The transmission window overlap probability is obtained by statistically analyzing the proportion of time points where two flows have overlapping sequence number intervals, while the packet interleaving degree reflects the degree to which packets from the two flows interweave in the same port's output queue. This weighted correction ensures that the marking probability reflects not only the congestion risk caused by a single flow itself but also the interference intensity between the flow and other parallel flows. For flows that have high overlap and interleaving with multiple flows, their local marking probability correction value is amplified, allowing them to be marked preferentially to suppress the impact on the shared queue; while for flows with low interference levels, the marking probability remains low to avoid unnecessary rate backoff. This mechanism enables coordinated marking of interfering flow groups on a shared port, mitigating unfair bandwidth allocation and throughput jitter caused by neglecting inter-flow relationships. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0014] Figure 1 This is a flowchart of a congestion control method for Ethernet switches that support RoCEv2 lossless transmission. Figure 2 This is a flowchart of the dynamic threshold mapping process; Figure 3 It is a flowchart of dynamic priority-weighted round-robin scheduling based on port queue acceleration; Figure 4 This is a distribution diagram of RoCEv2 inter-stream interference factor, transmission window overlap probability, and packet interleaving degree. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] See Figure 1 This invention provides a congestion control method for Ethernet switches supporting lossless RoCEv2 transmission. The method includes collecting real-time packet flow characteristic parameters from each port of the Ethernet switch. These parameters include the instantaneous transmission rate based on each RoCEv2 flow identifier, the instantaneous reception acknowledgment rate based on each RoCEv2 flow identifier, and the priority traffic pause frame trigger frequency based on each port. The real-time packet flow characteristic parameters are then processed using dynamic threshold mapping based on three-dimensional flow-port-time features to generate a congestion prediction index value for each RoCEv2 flow on each port. Based on the congestion prediction index value, an explicit congestion notification label adjustment operation based on spatiotemporal coupling is performed on each RoCEv2 flow to generate a local label probability correction value for each RoCEv2 flow. The explicit congestion notification label adjustment operation performs a weighted correction on the congestion prediction index value by querying a pre-constructed inter-flow interference factor matrix. Based on the local tag probability correction value, a priority-based reverse pressure flow control strategy is executed on the received RoCEv2 data packets to generate a flow shaping parameter set, which includes the sending token bucket depth update value and credit backoff step size for each RoCEv2 flow.

[0017] Example 1: See Figure 2The dynamic threshold mapping process employs a two-dimensional surface interpolation function composed of multiple historical congestion event sample points. The instantaneous buffer occupancy rate of the current port and the instantaneous transmission rate of the current flow are used as input variables for the two-dimensional surface interpolation function. Historical congestion event sample points are collected and stored by the switch when it detects that the queue depth of any port exceeds a preset threshold. The preset threshold is set to 80% of the total port buffer capacity of the switch. When the port queue depth exceeds this proportion, the switch records the port buffer occupancy rate at the trigger time, the transmission rate of each active RoCEv2 flow, and the congestion level score generated according to the switch's built-in congestion rating logic. These three types of data constitute a historical congestion event sample point. The control point set of the two-dimensional surface interpolation function is obtained. The control point set consists of the instantaneous buffer occupancy rate data, instantaneous transmission rate data, and corresponding congestion level score data corresponding to each historical congestion event. The switch divides the range of buffer occupancy rate values ​​into... The range of values ​​for the transmission rate is divided into several intervals. The system divides the network into intervals, forming a rectangular grid covering a two-dimensional plane. Each grid node corresponds to a combination of buffer occupancy and transmission rate values. For each grid node, the switch selects the sample point with the smallest spatial distance from the historical congestion event sample points, and uses the congestion level score of that sample point as the initial control point value for that grid node. If the number of historical congestion event sample points around a grid node is zero, a linear extrapolation based on the nearest neighbor sample point distance is performed on the node's location to obtain the control point value for that grid node. The buffer occupancy values, transmission rate values, and corresponding control point values ​​of all grid nodes together constitute the control point set. Using the instantaneous buffer occupancy of the current port as the input value on the first coordinate axis and the instantaneous transmission rate of the current flow as the input value on the second coordinate axis, bicubic spline interpolation is performed within the grid area formed by the control point set to generate a congestion prediction index value. The bicubic spline interpolation process is as follows: locate the surrounding input point within the grid area. A rectangular grid cell, which is composed of the buffer occupancy coordinates. and and the transmission rate coordinates and Defining, among which This represents the instantaneous buffer occupancy rate of the current port. Given the instantaneous transmission rate of the current flow; during interpolation, the control point values ​​of the grid cell and its surrounding 16 grid nodes are extracted, along with the first-order partial derivatives of each control point value along the buffer occupancy rate direction, the first-order partial derivatives along the transmission rate direction, and the cross-second-order partial derivatives. The congestion prediction index value is calculated using the following formula. :

[0018] in, and To sum the exponents, and The values ​​of are all integers from 0 to 3; and These represent the lower and upper boundary values ​​of the located rectangular grid cell on the buffer occupancy axis, respectively. and These are the lower and upper boundary values ​​of the located rectangular grid cell on the transmission rate coordinate axis, respectively; The bicubic spline interpolation coefficient matrix is ​​at the th Line 1 The elements of the column, the coefficient matrix, are obtained by solving a system of linear equations that make the function values, first-order partial derivatives, and cross-second-order partial derivatives of the interpolation surface continuous at the grid nodes, based on the control point values ​​of the 16 grid nodes, the first-order partial derivatives of the control point values ​​along the buffer occupancy direction, the first-order partial derivatives of the control point values ​​along the transmission rate direction, and the cross-second-order partial derivatives of the control point values. The switch will calculate the... This serves as the congestion prediction index for the current RoCEv2 flow on the current port.

[0019] Example 2: In practical implementation, an explicit congestion notification label adjustment operation based on spatiotemporal coupling is performed on each RoCEv2 flow according to the congestion prediction index value, generating a local label probability correction value for each RoCEv2 flow. An inter-flow interference factor matrix is ​​constructed, whose elements are calculated from the temporal overlap probability of the sending windows of two RoCEv2 flows sharing the same downlink port and the packet interleaving degree in the spatial dimension. For obtaining the sending window overlap probability, the starting sequence number and sending window size of the first RoCEv2 flow at each time point are collected, as well as the starting sequence number and sending window size of the second RoCEv2 flow at each time point. The switch records the starting and ending sequence numbers of the data packets to be sent by the first RoCEv2 flow at each time point, as the sending window interval of the first RoCEv2 flow at that time point, and simultaneously records the starting and ending sequence numbers of the data packets to be sent by the second RoCEv2 flow, as the sending window interval of the second RoCEv2 flow at that time point. For each time point, it is determined whether the sending windows of the first RoCEv2 flow and the second RoCEv2 flow have an overlapping interval in the sequence number space. The determination method is as follows: if the maximum sequence number of the sending window of the first RoCEv2 flow is greater than or equal to the minimum sequence number of the sending window of the second RoCEv2 flow, and the minimum sequence number of the sending window of the first RoCEv2 flow is less than or equal to the maximum sequence number of the sending window of the second RoCEv2 flow, then it is determined that the two sending windows have an overlapping interval. The ratio of the number of time points with overlapping intervals to the total number of sampling time points is used as the probability of overlapping sending windows. For the acquisition of packet interleaving degree, the switch monitors the dequeue sequences of the first RoCEv2 flow and the second RoCEv2 flow sharing the same downlink port during the sampling period, and records the number of alternations of the flow identifier to which the packet belongs. The packet interleaving degree is obtained by dividing the number of alternations of the flow identifier to which the packet belongs by the difference obtained by subtracting one from the total number of dequeue packets of the two RoCEv2 flows during the sampling period. The value range of the packet interleaving degree is a closed interval [0,1]. Multiplying the transmit window overlap probability by the packet interleaving degree yields the inter-flow interference factor between two RoCEv2 flows sharing the same downlink port. The index of the first RoCEv2 flow in the corresponding row of the inter-flow interference factor matrix is ​​denoted as... The index of the corresponding column in the inter-flow interference factor matrix for the second RoCEv2 flow is denoted as... The inter-flow interference factor between the first RoCEv2 flow and the second RoCEv2 flow is filled into the inter-flow interference factor matrix. Line 1 In the column position, the diagonal elements of the inter-flow interference factor matrix are set to zero. The number of rows and columns of the inter-flow interference factor matrix are equal to the total number of active RoCEv2 flows on the current port. When generating the local label probability correction value, the congestion prediction index value of the current RoCEv2 flow is obtained, and the element values ​​of all columns in the corresponding row of the current RoCEv2 flow in the inter-flow interference factor matrix are retrieved. The congestion prediction index value is multiplied by the sum of the element values ​​of all columns in the corresponding row of the current RoCEv2 flow, and the product is used as the local label probability correction value of the current RoCEv2 flow. The calculation formula is:

[0020] in, The stream identifier is The local tag probability correction value for RoCEv2 streams; The stream identifier is Congestion prediction index value for RoCEv2 streams; This indicates the total number of active RoCEv2 streams on the current port; The stream identifier is The RoCEv2 stream and stream identifier are The inter-flow interference factor between RoCEv2 flows, when hour The value is zero; The stream identifier is The sum of the inter-flow interference factors of the RoCEv2 flow and all other RoCEv2 flows on the same port.

[0021] See Figure 4 In the graph, the horizontal axis represents the packet interleaving degree between two RoCEv2 flows sharing the same downlink port, ranging from 0 to 1. The vertical axis represents the probability of overlapping transmission windows of the two RoCEv2 flows in the time dimension, ranging from 0 to 0.8. Each scatter point in the graph represents a different flow pair, and the color intensity of the scatter points corresponds to the magnitude of the inter-flow interference factor between the flow pairs. The color bar gradually changes from purple to yellow, with yellow indicating a larger inter-flow interference factor and purple indicating a smaller inter-flow interference factor.

[0022] As shown in the figure, both packet interleaving and transmission window overlap probability have a wide distribution range, with the scatter points being relatively dispersed. Furthermore, the inter-flow interference factor increases significantly with the increase of the horizontal and vertical coordinate values, exhibiting a clear positive correlation trend. When both packet interleaving and transmission window overlap probability are close to 0, the inter-flow interference factor is the lowest, and the scatter points appear in dark purple. When packet interleaving is greater than 0.6 and transmission window overlap probability is greater than 0.5, the inter-flow interference factor reaches a high value, and the scatter points tend to turn yellow, indicating that the spatiotemporal coupling effect between the two flows in this region is significant, leading to strong inter-flow interference.

[0023] Example 3: In practice, a priority-based reverse pressure flow control strategy is applied to received RoCEv2 data packets based on the local label probability correction value, generating a flow shaping parameter set. This set includes the sending token bucket depth update value and credit rollback step size for each RoCEv2 flow. The current credit balance and current sending token bucket depth value for each RoCEv2 flow are obtained. Internally, the switch maintains a credit counter and a token bucket depth register for each RoCEv2 flow. The credit counter tracks the available credit balance for the flow, and the token bucket depth register stores the maximum number of tokens the current token bucket can hold. When processing arriving data packets, the current credit balance and current sending token bucket depth value are read from the corresponding registers. The local label probability correction value is compared with a preset probability threshold. The preset probability threshold is stored in the switch configuration register. Switch administrators configure the preset probability threshold based on network topology and target congestion response sensitivity; for example, setting the preset probability threshold to 0.5. When the local label probability correction value exceeds the probability threshold, a token bucket depth update operation and a credit rollback step size setting operation are performed. The token bucket depth update value is set to the current depth value multiplied by a reduction factor based on the total port queue length. The steps to obtain the reduction factor based on the total port queue length are as follows: calculate the ratio of the current total port queue length to the total port buffer capacity to obtain the buffer occupancy rate; look up the basic reduction factor in the mapping table based on the buffer occupancy rate. The mapping table is a buffer occupancy rate-reduction factor mapping table pre-stored in the switch's memory, containing multiple buffer occupancy rate intervals and their corresponding basic reduction factors; multiply the basic reduction factor by a decay factor based on the number of times the current RoCEv2 flow has been historically marked to generate the reduction factor based on the total port queue length. The structure of the buffer occupancy rate-reduction factor mapping table is defined as follows: buffer occupancy rate interval [0, 30%) corresponds to a basic reduction factor of 1.0; interval [30%, 60%) corresponds to a basic reduction factor of 0.8; interval [60%, 80%) corresponds to a basic reduction factor of 0.6; interval [80%, 100%] corresponds to a basic reduction factor of 0.4. The basis for setting the above basic reduction factor is that as the proportion of the total port queue length to the total buffer capacity increases, the switch needs to reduce the sending token bucket depth more significantly to quickly suppress sending traffic and avoid buffer overflow. The attenuation factor based on the number of times the current RoCEv2 flow has been historically marked is calculated using the following formula: ,in, The attenuation factor has a value range of (0,1]. is the base of the natural logarithm; The attenuation constant is set to 0.2. The reason for setting it to 0.2 is to ensure that the attenuation factor drops to about 0.37 after the flow is marked 5 times, resulting in a significant reduction and enhancement effect. This represents the number of times the current RoCEv2 flow has been historically marked. This number is read from a marking counter maintained by the switch for each flow, which increments each time the switch performs an explicit congestion notification marking on the flow's data packets. The reduction factor based on the total port queue length is expressed as... , ,in, Basic reduction factor, This is the decay factor. The sent token bucket depth update value is set to the current depth value multiplied by a reduction factor based on the total queue length of the port. The credit rollback step size is set as a positive increment based on the current flow service level. The switch has a pre-configured mapping table between service levels and credit rollback step size positive increments. This table maps the RoCEv2 flow service level to a specific rollback step size value; for example, service level 0 corresponds to a rollback step size increment of 128 bytes, service level 1 corresponds to an increment of 256 bytes, and service level 2 corresponds to an increment of 512 bytes. After obtaining the service level of the current flow, the positive increment value of the credit rollback step size is determined by looking up the table, and this value is written to the credit rollback step size register, completing the credit rollback step size setting.

[0024] Example 4: See Figure 3 While implementing a reverse pressure flow control strategy that prioritizes received RoCEv2 data packets based on local tag probability correction values, the switch collects the rate of change of the total queue depth for each port within multiple consecutive time windows. The length of each time window is set to a fixed value. The switch records the initial value of the total queue depth at the beginning of each time window and the ending value at the end of each time window. The difference between the ending value and the initial value is divided by the length of the time window to obtain the rate of change of the total queue depth for that time window. Internally, the switch maintains a sliding memory for each port. This sliding memory stores the rate of change of the total queue depth for the most recent time windows in chronological order. Each time a new rate of change of the total queue depth is collected, the contents of the sliding memory are shifted forward one position, discarding the oldest rate of change and storing the latest one.

[0025] The dynamic priority weight of each port is calculated based on the rate of change of the total queue depth. The dynamic priority weight is used to adjust the scheduling polling order of the downlink at the port level.

[0026] The queue acceleration value is generated by performing a difference operation on the total queue depth change rate over two consecutive time windows. The latest total queue depth change rate in the sliding memory is recorded as the first change rate, and the previous total queue depth change rate in the sliding memory immediately preceding the first change rate is recorded as the second change rate. The queue acceleration value is calculated using the following formula:

[0027] in, This represents the queue acceleration value of the port. This represents the first rate of change, which is the rate of change of the total queue depth in the current time window; This represents the second rate of change, which is the rate of change of the total queue depth in the previous adjacent time window.

[0028] The queue acceleration value is compared with a preset acceleration threshold. The preset acceleration threshold is stored in the switch register. The switch administrator sets the preset acceleration threshold according to the network congestion sensitivity. The preset acceleration threshold is set to zero. The basis for setting it to zero is that when the queue acceleration value is positive, it indicates that the rate of change of the total queue depth is increasing and the queue accumulation speed is accelerating, requiring intervention and adjustment; when the queue acceleration value is zero or negative, it indicates that the rate of change of the total queue depth is unchanged or decreasing, and the queue status is stable or tending to be alleviated.

[0029] When the queue acceleration value exceeds the acceleration threshold, the dynamic priority weight is set to the first weight value. The first weight value is set to 2, which ensures that the corresponding port receives double the transmission opportunities relative to the base weight value in downlink round-robin scheduling, thus accelerating the clearing of queue backlog. When the queue acceleration value is less than or equal to the acceleration threshold, the dynamic priority weight is set to the second weight value. The second weight value is set to 1, which maintains the basic scheduling frequency for the corresponding port without further adjustments to the port scheduling priority. The first weight value is greater than the second weight value.

[0030] The switch downlink scheduler allocates transmission time slots to each port using a weighted round-robin method. In each polling cycle, the number of times a port is selected is proportional to its dynamic priority weight. Ports with higher dynamic priority weights receive more downlink transmission tokens in a polling cycle, thus sending backlogged packets faster in the downlink direction and reducing the total queue depth of the port.

[0031] Example 5: In practice, after generating the local label probability correction value for each RoCEv2 flow, the switch reads the local label probability correction value for each RoCEv2 flow and normalizes the local label probability correction values ​​for all RoCEv2 flows to generate a normalized label probability vector.

[0032] The normalization process is as follows: obtain the total number of active RoCEv2 streams on the current port, denoted as . and read the stream identifiers sequentially. The local tag probability correction value corresponding to the RoCEv2 stream ,in, From 1 to Integers. Calculate all. The sum of the local label probability correction values ​​is used to calculate the normalized component. The formula for calculating the normalized component is as follows:

[0033] in, This indicates that the corresponding flow identifier in the normalized labeled probability vector is... The normalized component of the RoCEv2 stream; The stream identifier is The local tag probability correction value for RoCEv2 streams; Indicates all The sum of the local label probability correction values ​​of each RoCEv2 stream; To find the sum index, iterate from 1 to... All integers. The normalized components are arranged in the order of the flow identifiers to form a normalized label probability vector.

[0034] The normalized marker probability vector is input into a pre-stored port-level probability mapping table. The reference value of the global forward explicit congestion notification marker corresponding to the current port is retrieved. The port-level probability mapping table is indexed by the hash value of the normalized marker probability vector. During the query process, the RoCEv2 flow identifiers corresponding to the largest components in the normalized marker probability vector are extracted to form an active flow identifier set. The specific number of components to be extracted is set to 3. This is based on the fact that in a standard data center RoCEv2 deployment scenario, the number of concurrent active flows on the same port is usually less than 8. Selecting the 3 largest components can cover the main congestion contributing flows while ensuring that the index space of the port-level probability mapping table is controllable. The extraction steps are as follows: all normalized components in the normalized marker probability vector are sorted in descending order. The top 3 normalized components are selected, and the RoCEv2 flow identifiers corresponding to these 3 normalized components are read. The 3 RoCEv2 flow identifiers are then arranged in ascending order according to their flow identifier values ​​to form an active flow identifier set.

[0035] A hash operation is performed on the active flow identifier set to generate a hash lookup value. The hash operation uses the CRC32 hash algorithm. The operation involves concatenating the three RoCEv2 flow identifiers from the active flow identifier set into a single byte sequence, inserting a 0xFF separator byte between adjacent flow identifiers, and inputting the concatenated byte sequence into the CRC32 calculation module to obtain a 32-bit hash lookup value. This hash lookup value is used to index the port-level probability mapping table.

[0036] The port-level probability mapping table is retrieved using the hash query value as an index. The port-level probability mapping table is stored in the switch's internal static random access memory (SRAM) and is implemented using a hash bucket structure. Each hash bucket stores a hash query value and its corresponding global forward explicit congestion notification flag (GPF) reference value. During retrieval, the calculated hash query value is matched against the hash query values ​​stored in each hash bucket of the port-level probability mapping table. If a match is found (i.e., a hash bucket stores a hash query value that matches the calculated hash query value), the corresponding GPF reference value stored in that hash bucket is read and used as the GPF reference value for the current port. If the search fails, meaning no hash query value stored in any hash bucket matches the calculated hash query value, the average of the local label probability correction values ​​of all RoCEv2 flows in the active flow identifier set is used as the global forward explicit congestion notification label reference value. The correspondence between the hash query value and the calculated global forward explicit congestion notification label reference value is written into the port-level probability mapping table. The writing location is the first empty hash bucket in the port-level probability mapping table or an alternative hash bucket determined by the open addressing method.

[0037] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A congestion control method for Ethernet switches supporting RoCEv2 lossless transmission, characterized in that, Includes the following steps: Collect real-time packet flow characteristic parameters of each port of the Ethernet switch. The real-time packet flow characteristic parameters include the instantaneous transmission rate based on each RoCEv2 flow identifier, the instantaneous reception acknowledgment rate based on each RoCEv2 flow identifier, and the priority traffic pause frame triggering frequency based on each port. The real-time message flow feature parameters are subjected to dynamic threshold mapping based on three-dimensional features of flow-port-time to generate congestion prediction index values ​​for each RoCEv2 flow on each port. Based on the congestion prediction index value, an explicit congestion notification label adjustment operation based on spatiotemporal coupling is performed on each RoCEv2 flow to generate a local label probability correction value for each RoCEv2 flow. The explicit congestion notification label adjustment operation performs a weighted correction on the congestion prediction index value by querying a pre-constructed inter-flow interference factor matrix. Based on the local tag probability correction value, a priority-based reverse pressure flow control strategy is executed on the received RoCEv2 data packets to generate a flow shaping parameter set, which includes the sending token bucket depth update value and credit backoff step size for each RoCEv2 flow.

2. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 1, characterized in that, The step of performing dynamic threshold mapping processing on the real-time packet flow feature parameters based on three-dimensional features of flow-port-time to generate a congestion prediction index value for each RoCEv2 flow on each port includes: The dynamic threshold mapping process uses a two-dimensional surface interpolation function composed of multiple historical congestion event sample points, with the instantaneous buffer occupancy rate of the current port and the instantaneous transmission rate of the current flow as the input variables of the two-dimensional surface interpolation function. The historical congestion event sample points are collected and stored by the switch when it detects that the queue depth of any port exceeds a preset threshold. Obtain the control point set of the two-dimensional surface interpolation function. The control point set consists of the instantaneous buffer occupancy rate data, instantaneous transmission rate data, and corresponding congestion level score data corresponding to each historical congestion event. Using the instantaneous buffer occupancy rate of the current port as the input value on the first coordinate axis and the instantaneous transmission rate of the current flow as the input value on the second coordinate axis, bicubic spline interpolation calculation is performed within the grid area formed by the set of control points to generate the congestion prediction index value.

3. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 1, characterized in that, The step of performing an explicit congestion notification label adjustment operation based on spatiotemporal coupling on each RoCEv2 flow according to the congestion prediction index value, and generating a local label probability correction value for each RoCEv2 flow, includes: The inter-flow interference factor matrix is ​​constructed, and the elements of the inter-flow interference factor matrix are calculated by the transmission window overlap probability of two RoCEv2 flows sharing the same downlink port in the time dimension and the packet interleaving degree in the spatial dimension. The local label probability correction value is generated by multiplying the congestion prediction index value of the current RoCEv2 flow by the sum of the inter-flow interference factors of the current RoCEv2 flow and all other RoCEv2 flows on the same port.

4. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 1, characterized in that, The reverse pressure flow control strategy, which prioritizes the received RoCEv2 data packets based on the local tag probability correction value, generates a set of flow shaping parameters, including: Get the current credit balance and the current depth of the sending token bucket for each RoCEv2 stream; The local tag probability correction value is compared with a preset probability threshold. When the local tag probability correction value exceeds the probability threshold, the sending token bucket depth update value is set to the current depth value multiplied by a reduction coefficient based on the total queue length of the port, and the credit backoff step size is set to a positive increment based on the current flow service level.

5. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 1, characterized in that, The method further includes: While implementing a reverse pressure flow control strategy that prioritizes received RoCEv2 data packets based on the local tag probability correction value, the rate of change of the total queue depth of each port within multiple consecutive time windows is collected. The dynamic priority weight of each port is calculated based on the rate of change of the total queue depth. The dynamic priority weight is used to adjust the scheduling polling order of the downlink at the port level.

6. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 5, characterized in that, The calculation of the dynamic priority weight for each port based on the rate of change of the total queue depth includes: Perform a difference operation on the total queue depth change rate of two consecutive time windows to generate queue acceleration values; The queue acceleration value is compared with a preset acceleration threshold. When the queue acceleration value is greater than the acceleration threshold, the dynamic priority weight is set to a first weight value. When the queue acceleration value is less than or equal to the acceleration threshold, the dynamic priority weight is set to a second weight value. The first weight value is greater than the second weight value.

7. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 1, characterized in that, After generating the local tag probability correction value for each RoCEv2 stream, the process further includes: Read the local label probability correction value of each RoCEv2 stream, and normalize the local label probability correction values ​​of all RoCEv2 streams to generate a normalized label probability vector. The normalized tag probability vector is input into a pre-stored port-level probability mapping table, and the reference value of the global forward explicit congestion notification tag corresponding to the current port is obtained by querying. The port-level probability mapping table is indexed by the hash value of the normalized tag probability vector.

8. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 7, characterized in that, The step of inputting the normalized marker probability vector into a pre-stored port-level probability mapping table and querying the global forward explicit congestion notification marker reference value corresponding to the current port includes: Extract the RoCEv2 stream identifiers corresponding to the largest components in the normalized label probability vector to form an active stream identifier set; Perform a hash operation on the set of active stream identifiers to generate a hash query value; The port-level probability mapping table is retrieved using the hash query value as an index. If the retrieval is successful, the corresponding global forward explicit congestion notification flag reference value is read. If the search fails, the average value of the local label probability correction values ​​of all RoCEv2 flows in the active flow identifier set is used as the global forward explicit congestion notification label reference value, and the correspondence between the hash query value and the calculated global forward explicit congestion notification label reference value is written into the port-level probability mapping table.

9. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 4, characterized in that, The steps for obtaining the reduction factor based on the total port queue length include: Calculate the ratio of the total queue length to the total capacity of the port buffer to obtain the buffer occupancy rate; The basic reduction coefficient is obtained by looking up the mapping table based on the buffer occupancy rate. The mapping table is a buffer occupancy rate-reduction coefficient mapping table pre-stored in the switch memory. The mapping table contains multiple buffer occupancy rate intervals and the basic reduction coefficient corresponding to each interval. The reduction factor based on the total port queue length is generated by multiplying the base reduction factor by a decay factor based on the number of times the current RoCEv2 stream has been marked in history.

10. The Ethernet switch congestion control method supporting RoCEv2 lossless transmission according to claim 3, characterized in that, The step of obtaining the overlap probability of the sending windows includes: Collect the starting sequence number and window size of the transmission window of the first RoCEv2 stream at each time point, and the starting sequence number and window size of the transmission window of the second RoCEv2 stream at each time point; For each time point, determine whether there is an overlapping interval between the sending window of the first RoCEv2 stream and the sending window of the second RoCEv2 stream in the sequence number space; The ratio of the number of time points where the overlapping interval exists to the total number of sampling time points is used as the probability of overlapping of the sending window.