Credit packet reuse method, and device

By maintaining active flow tables and end flow tables in intermediate devices and reusing credit packets, the problem of low credit packet utilization in long-distance data center interconnection scenarios is solved, and link utilization and transmission efficiency are improved.

WO2025148679A1PCT designated stage expired Publication Date: 2025-07-17ZTE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141802
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-12-24
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the long-distance data center interconnection scenario, the existing active congestion control methods have problems such as low credit packet utilization and low link utilization. Especially when the stream size is unknown, the waste of credit packets and the efficiency of data packet transmission is not high.

Method used

By maintaining active flow tables and end flow tables in intermediate devices, reuse credit packets of transport streams that have not ended transmission, reduce waste of credit packets, and improve the utilization rate of credit packets.

Benefits of technology

It improves link utilization in data center interconnection scenarios, reduces the waste of credit packets, and improves transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141802_17072025_PF_FP_ABST
    Figure CN2024141802_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure is a credit packet reuse method, which is applied to an intermediate device connected to at least one source sending end. The method comprises: receiving a credit packet from a destination receiving end (S401), wherein the credit packet is used for scheduling a data packet in a first transport stream corresponding to the credit packet; acquiring stream identification information of the first transport stream corresponding to the credit packet (S402); acquiring an active flow table and a terminated flow table (S403), wherein the active flow table comprises stream identification information of a transport stream, the transmission of which has not yet terminated, and the terminated flow table comprises stream identification information of a transport stream, the transmission of which has terminated; and when it is determined that the stream identification information of the first transport stream is included in the terminated flow table, reusing the credit packet on the basis of the active flow table, so as to schedule, by means of the credit packet, a data packet in the transport stream, the transmission of which has not yet terminated (S404).
Need to check novelty before this filing date? Find Prior Art

Description

Credit package reuse method and device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims priority to Chinese patent application 202410042100.4 filed with the State Intellectual Property Office of China on January 10, 2024, and the disclosure of this Chinese patent application is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates to the technical field of data center network data transmission, and in particular to a credit packet reuse method and device. Background Art

[0004] With the advancement of network technology, data centers have become critical infrastructure supporting various application services, centrally managing (storing, computing, and exchanging) data for these services. With the continuous expansion of applications and services, coupled with the need for remote disaster recovery, traditional data centers are no longer able to meet the high-performance and security requirements of applications. In the future, geographically separated data centers are becoming an inevitable development direction. These can be defined as a collection of data centers in different regions, either within the same city or in different cities. These geographically separated data centers are typically located far apart, and long-distance links are more expensive. Improving link utilization in these data center interconnection scenarios has become a pressing issue. Summary of the Invention

[0005] The present disclosure provides a credit packet reuse method, device, and computer-readable medium for improving link utilization in a data center interconnection scenario.

[0006] In the first aspect, an embodiment of the present disclosure provides a credit packet reuse method, which is applied to an intermediate device connected to at least one source sending end, and can also be applied to a structure or device set in the intermediate device, such as a chip, a chip system or a circuit system, etc., and is illustrated by taking the application of the method to the intermediate device as an example. The method includes: receiving a credit packet from the destination receiving end, the credit packet is used to schedule data packets in the first transmission stream corresponding to the credit packet, obtaining the flow identification information of the first transmission stream corresponding to the credit packet, and obtaining an active flow table and an end flow table, wherein the active flow table includes the flow identification information of the transmission stream whose transmission has not been terminated, and the end flow table includes the flow identification information of the transmission stream whose transmission has been terminated; when it is determined that the flow identification information of the first transmission stream is included in the end flow table, the credit packet is reused according to the active flow table to schedule the data packets in the transmission stream whose transmission has not been terminated through the credit packet.

[0007] In a second aspect, an embodiment of the present disclosure provides a credit package reuse device, which includes: one or more processors; a memory on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the above-mentioned credit package reuse method; and one or more I / O interfaces, connected between the processor and the memory, configured to realize information interaction between the processor and the memory.

[0008] In a third aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the above-mentioned credit packet reuse method when the computer program is executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the accompanying drawings of the embodiments of the present disclosure:

[0010] FIG1 is an architecture diagram of a data center interconnection system provided by an embodiment of the present disclosure;

[0011] FIG2 is a schematic diagram of an end-to-end active transmission based on ExpressPass provided in an embodiment of the present disclosure;

[0012] FIG3 is a schematic diagram of an end-to-end active transmission based on FlashPass provided in an embodiment of the present disclosure;

[0013] FIG4 is a flow chart of a credit packet reuse method provided by an embodiment of the present disclosure;

[0014] FIG5 is a flow chart of another credit packet reuse method provided by an embodiment of the present disclosure;

[0015] FIG6 is a flow chart of another credit packet reuse method provided by an embodiment of the present disclosure;

[0016] FIG7 is a schematic diagram of an IP packet header format provided by an embodiment of the present disclosure;

[0017] FIG8 is a block diagram of a credit packet reuse device provided by an embodiment of the present disclosure;

[0018] FIG9 is a block diagram of a computer-readable medium according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] To enable those skilled in the art to better understand the technical solution of the present disclosure, a credit packet reuse method, device, and computer-readable medium provided by an embodiment of the present disclosure are described in detail below with reference to the accompanying drawings.

[0020] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the illustrated embodiments may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that the present disclosure will be thorough and complete and will fully understand the scope of the present disclosure to those skilled in the art.

[0021] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing the detailed embodiments with reference to the accompanying drawings.

[0022] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0023] The terms used in this disclosure are only used to describe specific embodiments and are not intended to limit the disclosure. As used in this disclosure, the term "and / or" includes any and all combinations of one or more related enumerated items. As used in this disclosure, the singular forms "a" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. As used in this disclosure, the terms "comprising" and "made of" specify the presence of the features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.

[0024] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meanings as those commonly understood by those skilled in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined in this disclosure.

[0025] In this disclosure, unless otherwise specified, the following technical terms should be understood as follows:

[0026] 1) Round-trip time (RTT) refers to the time from when the source sender sends a data packet to when the source sender receives the confirmation from the destination receiver.

[0027] 2) Band-delay Product (BDP) refers to the product of bandwidth and delay, representing the number of bits in flight in the pipeline.

[0028] 3) Minimum Ethernet frame size (Minimum Transmission Unit, MinTU), the minimum frame size allowed by Ethernet, in bytes.

[0029] 4) Maximum Ethernet frame size (Maximum Transmission Unit, MTU), the maximum frame size allowed by Ethernet, in bytes.

[0030] 5) Explicit Congestion Notification (ECN), which is usually used as a congestion signal to indicate that the storage queue exceeds the threshold.

[0031] 6) Transmission stream, also known as flow, communication stream or data stream, refers to a complete transmission process between the source sender and the destination receiver, including the collection of various types of data (such as control packets and data packets) transmitted in both directions during the complete transmission process.

[0032] 7) Reactive congestion control schemes in data center interconnect scenarios are typically designed with low latency as the goal of expediting the transmission of small flows. They first send packets directly at the maximum rate, then adjust the packet sending rate or window size based on end-to-end congestion control signals to achieve congestion control. For example, Data Center Quantification Congestion Notifications (DCQCN), Swift, and High Precision Congestion Control (HPCC) use explicit congestion notifications, latency, and in-band network measurements as congestion control signals, respectively. These schemes all require a feedback delay of one round-trip (RTT) and even longer convergence times. For example, in long-distance data center interconnect scenarios, in lossless networks implemented using hop-by-hop flow control with priority-based flow control (PFC), reactive congestion control schemes require that the buffers on long-distance switch ports reach the BDP level of the long-distance link due to the need to send packets at the maximum rate. This is typically hundreds of times larger than within the data center and increases linearly with link length.

[0033] 8) Active congestion control schemes based on credits in data center interconnect scenarios typically achieve zero packet loss with smaller switch buffer requirements and offer fast convergence. Two commonly used active congestion control schemes for active inter-data center transmission are ExpressPass and FlashPass.

[0034] 9) ExpressPass is an end-to-end credit-based, latency-bound congestion control solution for data centers. Credit packets are used to control the rate and schedule data packets. The destination receiver sends a credit packet to the source sender on a per-flow basis. Credit packets schedule data packets so that a source sender only sends a data packet after receiving a credit packet from the destination receiver. Each credit packet schedules one data packet. Upon receiving a credit packet, the source sender sends a data packet if there is one, and ignores the credit packet if there is none. Different flows use different credit packets to schedule their data packet transmission. Credit packets represent the buffer size used by the destination receiver to receive data for the corresponding flow. A switch, an intermediate device between the source sender and the destination receiver, limits the credit packet rate on each link and determines the bandwidth available for data packets in the opposite direction. By limiting the credit packet rate at the switch, the system proactively controls congestion.

[0035] 10) FlashPass is another end-to-end credit-based, latency-bound congestion control solution for data centers. The source sends simulated packets to the destination on a per-flow basis. The destination triggers the sending of credit packets, which in turn schedule data packets. Each simulated packet schedules one credit packet, and each credit packet schedules one data packet. Upon receiving a credit packet, the source sends it if there is a data packet, and ignores it if there is no data packet. The switch randomly discards simulated packets and calculates the simulated packet loss rate by copying the simulated packet sequence number to the credit packet and passing it to the source. Congestion is actively controlled by adjusting the simulated packet rate.

[0036] FIG1 is a diagram of a data center interconnection system architecture provided by an embodiment of the present disclosure, and the applicable scenarios of the present disclosure are described in conjunction with FIG1. ​​In FIG1, taking the sending end as the local and the receiving end as the remote as an example, the local data center includes data center S1 to data center Sn, and the remote data center includes data center R1 to data center Rn. The source sending end of the local data center is connected to the local intermediate device S1, the local intermediate device S1 is connected to the remote intermediate device R1, and the remote intermediate device R1 is connected to the destination receiving end of the remote data center. The source sending end and the destination receiving end transmit data through the intermediate device, that is, the intermediate device forwards the data transmitted between the source sending end and the destination receiving end. It should be noted that local and remote are relative concepts in the present disclosure. For the intermediate device R1, the data center directly connected to it can also be understood as the local data center corresponding to it, and the data center indirectly connected through other intermediate devices can be understood as the remote data center corresponding to it. In the present disclosure, in order to save the storage resource usage of the intermediate device, the port of the intermediate device can be divided into a local port (LPort) and a remote port (RPort). The LPort of the intermediate device is connected to the local data center, and the RPort is connected to the remote intermediate device through a long-distance link. In the present disclosure, the source sending end and the destination receiving end may be servers in a data center, and the intermediate device may be a switch, which is not limited.

[0037] The following will describe the specific processes of ExpressPass and FlashPass in conjunction with the system architecture to which the present disclosure is applicable.

[0038] Figure 2 is a schematic diagram of an end-to-end active transmission based on ExpressPass provided by an embodiment of the present disclosure. End-to-end active transmission based on ExpressPass can be understood as a receiver-driven mode, which can be simply referred to as a request packet-credit packet-data packet mode. This mode is applicable to a typical active-active data center scenario, that is, a scenario where only two data centers are interconnected. In this scenario, the end-to-end RTT is almost the same. When a stream arrives at the source sending end or data transmission is required, a request packet of MinTU bytes is sent to the destination receiving end to schedule a credit packet. After the destination receiving end receives the request packet, it triggers the continuous generation of credit packets of MinTU bytes to be returned to the source sending end to schedule the transmission of data packets. The generation interval of the credit packet is the transmission time of one MTU. One credit packet can schedule one data packet. After receiving the credit packet, the source sending end sends a data packet of MTU bytes to the destination receiving end. The destination receiving end stops generating credit packets after receiving all the data packets corresponding to the stream. When the size of the transmission stream is unknown, the credit cannot be estimated in advance. After the stream ends, there are wasted credit packets that last for the RTT. The credit packet loss rate is calculated by copying the credit packet sequence number to the corresponding data packet and transmitting it to the destination receiving end. The destination receiving end adjusts the credit packet sending rate through feedback control to keep the loss rate within the preset range. The local port and remote port of the intermediate device respectively maintain two queues: the credit queue and the data packet queue, and limit the credit queue so that the credit packet forwarding interval of the intermediate device credit queue is the transmission time of one MTU. At the same time, a smaller threshold value N is set for the credit queue. th Credit packets exceeding this threshold are discarded fairly. Strong congestion avoidance is achieved by allowing reverse credit packets to be discarded. To ensure that reverse credit packets and forward data packets follow the same path and that credit packet planning is effective, intermediate devices use symmetric hashing with deterministic equal-cost multi-path (ECMP) forwarding. Symmetric routing does not affect the performance of long-distance data center interconnections.

[0039] Figure 3 is a schematic diagram of an end-to-end active transmission based on FlashPass provided by an embodiment of the present disclosure. End-to-end active transmission based on FlashPass can be understood as a sender-driven mode, which can be simply referred to as an analog packet-credit packet-data packet mode. This mode is suitable for a multi-data center cluster interconnected structure where different long-distance links in the structure bring about large RTT differences. In this mode, it is necessary to determine the maximum RTT between data centers, which is recorded as RTT maxAt time T, the source transmitter has a flow to send and begins to continuously send simulated packets of MinTU bytes to the destination receiver. The interval between simulated packets is the transmission time of one MTU. The destination receiver returns a credit packet of MinTU bytes for each simulated packet it receives. The source transmitter triggers the sending of a data packet of MTU bytes for each credit packet it receives. When the size of the transmission flow is unknown, credit cannot be estimated in advance. After the flow ends, there are wasted credit packets lasting for the RTT. The local port and remote port of the intermediate device respectively maintain two queues: the simulated packet queue and the data packet queue. The forwarding interval of the simulated packet queue is the transmission time of one MTU. At the same time, a small threshold value N is set for the simulated packet queue. th , the simulation packets exceeding the threshold are discarded fairly, and the credit packets are almost never discarded. The source sender limits the number of packets to T+RTT. max This forces synchronization of all forward traffic flows, preventing packet accumulation due to RTT variations on long-distance links. Intermediary devices calculate the simulated packet loss rate by copying the simulated packet sequence number into a credit packet and transmitting it to the source sender. They then adjust the simulated packet rate to keep the simulated packet loss rate within a preset range. While source-driven active congestion control in the prior art does not require symmetric routing, the method of the present embodiment does, and can be achieved through ECMP.

[0040] In some related technologies, in long-distance data center interconnection scenarios, active transmission can adopt reactive congestion control and credit-based active congestion control. Since reactive congestion control has obvious disadvantages, it is no longer applicable to long-distance data center interconnection scenarios. The embodiments of the present disclosure mainly relate to credit-based active congestion control, such as ExpressPass and FlashPass mentioned above. The disadvantages of active congestion control are also obvious. The flow requires a basic delay of RTT. If data packets are also transmitted within the first RTT, it is inevitable that unscheduled data packets (which can be understood as data packets that are not scheduled by credit packets and are directly sent) will collide with scheduled data packets (data packets triggered by credit packets), resulting in the selective discard of data packets. It is necessary to introduce a data packet retransmission mechanism, which reduces efficiency. In addition, long-distance links between data centers have high latency and load traffic has large and small flow characteristics (various types of flows). In the more general case where the flow size is unknown, it is impossible to estimate in advance whether the credit packet supply of the destination receiver is sufficient. The destination receiver can only continue to send credit packets until it receives all the data packets of the flow. After the flow ends, there are wasted credit packets lasting for the RTT. These credit packets will not trigger data packet transmission, which will result in wasted credit packets that do not schedule data packets on long-distance links, reducing link utilization.

[0041] In some related technologies, for unknown-sized flows, by assuming their size, the destination receiving end estimates in advance whether the generated credit packets are sufficient to meet the data transmission needs, and reduces the credit packet overhead by assigning priorities to the credit packets and selectively discarding the credit packets based on the priorities. However, such assumptions about the unknown flow size are too radical and impractical. In addition, due to the phenomenon of large and small flows between data centers, there are a large number of small flows whose flow sizes are smaller than the long-distance BDP. The existence of this small flow causes the credit packet allocation mechanism for active transmission to have a low credit packet utilization rate. Assuming that the bandwidth-delay product of the long-distance link is BDP bytes and the flow size is f size Bytes (obtained after transmission is completed), the number of flows sharing the long-distance link is N, and it is assumed that fair credit packets or simulated packets are dropped to obtain fair bandwidth shares, and the credit packet utilization is u c :

[0042] The additional overhead of active transmission is:

[0043] In an ideal situation (each credit packet dispatches a data packet of MTU size), the theoretical link utilization is:

[0044] It was found that when the flow size is smaller than the BDP, a large additional overhead occurs, deviating from the original design intention of active transmission to achieve low queues with low overhead. In the long-distance data center interconnection scenario, the BDP increases rapidly due to long delays, resulting in large additional overhead for active transmission and low utilization of the effective link for data transmission.

[0045] In light of this, embodiments of the present disclosure provide a credit packet reuse method, device, and computer-readable medium, primarily for use in scenarios where active transmission employs active congestion control in long-distance data center interconnects. These methods are applicable to the ExpressPass and FlashPass scenarios described above, as well as the system architecture of Figure 1. These methods are described in detail below with reference to the accompanying drawings.

[0046] Figure 4 is a flow chart of a credit packet reuse method provided by an embodiment of the present disclosure. The method is applied to an intermediate device connected to at least one source sending end. The intermediate device can be a switch or other device with data processing and forwarding functions. The method includes steps S401 to S404.

[0047] In step S401, a credit packet is received from a destination receiving end, where the credit packet is used to schedule data packets in a first transport stream corresponding to the credit packet.

[0048] The embodiment of the present disclosure does not limit the number of credit packets.

[0049] In some embodiments, to conserve storage resources on the intermediate device, the ports of the intermediate device are divided into local ports (LPorts) and remote ports (RPorts). The local ports of the intermediate device are connected to the source sending end located in the local data center, while the remote ports are connected to the remote intermediate device via a long-distance link. The remote intermediate device is then connected to the destination receiving end. In step S401, a credit packet can be received from the destination receiving end via the remote port of the intermediate device.

[0050] In step S402, the stream identification information of the first transport stream corresponding to the credit packet is obtained.

[0051] It can be understood that different transport streams correspond to different credit packets, and different transport streams can be distinguished by stream identification information.

[0052] In some embodiments, the flow identification information may include at least a source Internet Protocol (IP) address and a destination IP address. In addition, the flow identification information may also include a source port number and a destination port number. In this disclosure, the source IP address may be understood as the IP address of the source sending end, the source port number refers to the port number of the source sending end, the destination IP address may be understood as the IP address of the destination receiving end, and the destination port number refers to the port number of the destination receiving end.

[0053] In some embodiments, the stream identification information of the first transport stream is included in a credit packet, and the corresponding stream identification information of the first transport stream can be obtained from the credit packet.

[0054] In some embodiments, the storage unit of the intermediate device stores the credit packet and the flow identification information of the transmission flow corresponding to the credit packet, and the flow identification information of the first transmission flow corresponding to the received credit packet can be obtained from the storage unit of the intermediate device.

[0055] In step S403, the active flow table and the end flow table are obtained.

[0056] The Active Flow Table (AFT) includes flow identification information of transmission flows that have not yet completed transmission, and the Complete Flow Table (CFT) includes flow identification information of transmission flows that have completed transmission.

[0057] In step S404, when it is determined that the flow identification information of the first transport flow is included in the end flow table, the credit packets are reused according to the active flow table to schedule data packets in the transport flow whose transmission has not ended through the credit packets.

[0058] Through the above method, a credit packet is received from the destination receiving end, and the flow identification information of the first transmission flow corresponding to the credit packet, as well as the active flow table and the end flow table are obtained. If it is determined that the flow identification information of the first transmission flow is included in the end flow table, it can be determined that the first transmission flow has ended transmission, and then the active flow table can be used to reuse the credit packets that no longer schedule data packets in the first transmission flow, so as to schedule the data packets in the transmission flow that have not ended transmission in the active flow table through the credit packets. This can reduce invalid credit packets (also called wasted credit packets) in the transmission link, improve the utilization rate of credit packets, and thus improve the link utilization rate.

[0059] In some embodiments, the end flow table may also include expiration time information, which is used to determine the expiration time of the flow identification information of the transmission flow that has ended transmission. In this case, the intermediate device may also determine the expiration time based on the expiration time information and, after the expiration time is reached, delete the flow identification information of the transmission flow that has ended transmission corresponding to the expiration time information from the end flow table. This allows for timely deletion of invalid information in the end flow table, improving storage space utilization.

[0060] In some embodiments, the expiration time information includes a timer with a preset duration or an expiration timestamp. In some embodiments, the preset duration can be less than or equal to the RTT. The RTT can be obtained through measurement. In long-distance data center interconnection scenarios, the latency ranges from hundreds of microseconds to tens of milliseconds. The basic time unit of the timer can be set to 1 microsecond.

[0061] If the expiration time information includes a timer with a preset duration, the timer can be set to a timeout timer with the preset duration, with the timer's initial time being the preset duration. When the timer reaches zero, the flow identification information of the corresponding transport flow whose transmission has ended is determined to be invalid. This timer entry with the preset duration can be used to determine the time range for reusing wasted credit packets. This design, using relative time to determine the expiration time, saves storage space compared to using absolute time to determine the expiration time. Furthermore, using a timer can prevent excessive accumulation of entries in the end flow table and prevent misjudgment.

[0062] When the expiration time information includes an expiration timestamp, the expiration timestamp may be an absolute expiration time, upon which the corresponding stream identification information of the completed transmission stream is determined to be invalid. For example, the expiration timestamp may be an absolute time stamp on the intermediate device.

[0063] In some embodiments, before determining whether the flow identification information of the first transmission flow is included in the end flow table, the method further includes querying whether the end flow table contains the flow identification information of the first transmission flow, and if not, directly forwarding the credit packet. The query method is not limited. For example, a Bloom filter can be used for fast query. The advantages of using a Bloom filter are high space efficiency and low query time.

[0064] In some embodiments, reusing credit packets based on an active flow table includes: determining updated flow identification information from the active flow table based on a preset rule, updating the flow identification information of a first transmission flow corresponding to the credit packet to the updated flow identification information, and forwarding the credit packet based on the updated flow identification information to schedule data packets in the transmission flow corresponding to the updated flow identification information using the credit packet. In this way, credit packets can be used to schedule data packets in transmission flows that have not yet completed transmission, thereby avoiding credit packet waste, improving credit packet utilization, and thereby improving link utilization.

[0065] In some embodiments, the flow identification information includes the source IP address of the source sending end and the destination IP address of the destination receiving end of the transmission flow, the preset rule includes a path maximum coverage rule, and the path maximum coverage rule includes being the same as the first source IP address and / or the first destination IP address included in the flow identification information of the first transmission flow. In this case, according to the preset rule, determining the updated flow identification information from the active flow table may include: comparing the first source IP address and / or the first destination IP address with the source IP address and / or the destination IP address included in the flow identification information of the transmission flow of unfinished transmission included in the active flow table; and determining the flow identification information of the transmission flow of unfinished transmission in the active flow table that meets the path maximum coverage rule as the updated flow identification information. Reusing credit packets through the path maximum coverage rule can minimize the impact on the network data packet queue and minimize the impact on the path planned by active congestion control.

[0066] In some embodiments, the preset rules further include a random determination rule. The random determination rule can be understood as randomly determining at least one updated flow identification information from the flow identification information of the transmission flow of the unfinished transmission included in the active flow table, so that the wasted credit packets can be randomly and evenly re-matched to the flow identification information in the active flow table.

[0067] In some embodiments, the updated flow identification information may be determined from the active flow table based on the maximum path coverage rule. If there is no matching item in the maximum path coverage rule, the updated flow identification information may be determined from the active flow table using a random determination rule.

[0068] In some embodiments, when multiple updated flow identification information are determined from the active flow table according to preset rules, the multiple updated flow identification information can be used to update the original flow identification information of the credit packet, which can be understood as evenly matching the credit packet to the multiple updated flow identification information.

[0069] In some embodiments, before obtaining the active flow table and ending the flow table, it also includes: receiving a credit scheduling package from the source sending end, the credit scheduling package is used to schedule the credit package, obtaining the flow identification information of the transmission flow corresponding to the credit scheduling package, wherein the credit scheduling package and the credit package it schedules correspond to the same transmission flow, and adding the flow identification information of the transmission flow corresponding to the credit scheduling package to the active flow table to generate or update the active flow table.

[0070] In some embodiments, the credit scheduling packet may be received from the source sender via a local port of the intermediate device.

[0071] In some embodiments, the credit scheduling packet may be a credit request packet and / or a simulation packet.

[0072] In some embodiments, before adding the flow identification information of the transmission flow corresponding to the credit scheduling package to the active flow table, the active flow table is queried to see whether the flow identification information of the transmission flow corresponding to the credit scheduling package is already included. If the flow identification information of the transmission flow corresponding to the credit scheduling package is included, the credit scheduling package is directly forwarded. If the flow identification information of the transmission flow corresponding to the credit scheduling package is not included, the flow identification information of the transmission flow corresponding to the credit scheduling package is added to the active flow table and the credit scheduling package is forwarded. Forwarding the credit scheduling package can use methods known in the art. The query method is not limited. For example, a Bloom filter can be used for fast query.

[0073] In some embodiments, before obtaining the active flow table and the end flow table, the method further includes: receiving a data packet from a source sending end, the data packet including flow end identification information, the flow end identification information being used to determine whether the transmission flow corresponding to the data packet has ended transmission, obtaining the flow identification information of the transmission flow corresponding to the data packet, and when it is determined based on the flow end identification information that the transmission flow corresponding to the data packet has ended transmission, deleting the flow identification information of the transmission flow corresponding to the data packet included in the active flow table to update the active flow table, and adding the flow identification information of the transmission flow corresponding to the data packet to the end flow table to generate or update the end flow table. In this embodiment, expiration time information corresponding to the flow identification information of the transmission flow corresponding to the data packet can also be added to the end flow table. Determining the transmission status of the transmission flow by carrying the flow end identification information in the data packet is more timely than using the timeout mechanism of the state machine to determine the transmission status of the transmission flow.

[0074] It should be noted that the stream end identification information is used to determine whether the transmission stream corresponding to the data packet has ended transmission. It can also be understood that the stream end identification information is used to indicate or determine whether the data packet is the last data packet of the corresponding transmission stream.

[0075] In some embodiments, the data packet may be received from the source sender via a local port of the intermediate device.

[0076] In some embodiments, the flow end identification information is at least one bit in the IP header of the data packet. Unused bits in the IP header can be reused, or bits can be added to the IP header.

[0077] In some embodiments, the end-of-stream identifier is the ECN bit in the IP packet header. This minimizes impact on the IP protocol stack. There are no restrictions on the value of the ECN field used to indicate or determine whether the transport stream corresponding to a data packet has ended. In one possible design, ECN = 00 indicates that the transport stream corresponding to the data packet has ended or that the data packet is the last data packet in the transport stream. ECN ≠ 00 indicates that the transport stream corresponding to the data packet has not ended or that the data packet is not the last data packet in the transport stream.

[0078] In order to enable those skilled in the art to more clearly understand the technical solutions provided by the embodiments of the present disclosure, the technical solutions provided by the embodiments of the present disclosure are further described below through specific examples.

[0079] Figure 5 is a flow chart of another credit packet reuse method provided by an embodiment of the present disclosure. In this embodiment, the intermediate device is a switch, and the switch's ports include local ports (LPorts) and remote ports (RPorts). The switch's storage area maintains two data structures related to transmission flow information: an AFT and a CFT. The AFT includes flow identification information (source IP address, destination IP address, source port, destination port) for transmission flows that have not yet completed transmission, and the CFT includes flow identification information and corresponding timers (source IP address, destination IP address, source port, destination port, and RTT) for transmission flows that have completed transmission. The AFT and CFT stored in the switch only include flow identification information for transmission flows connected to a source sender located in a local data center. The switch uses symmetric hashing for ECMP forwarding. Flow end identification information in data packets is implemented using ECN, where ECN = 00 indicates that the transmission flow corresponding to the data packet has completed transmission. The method includes the following steps.

[0080] Switch LPort processing process:

[0081] In step S501, the switch determines whether the packet received by the LPort is a request packet or a simulated packet. If not, step S502 is executed. If it is a request packet or a simulated packet, the AFT is maintained. Maintaining the AFT may include: obtaining the flow identification information of the transport flow corresponding to the request packet or simulated packet, and using a Bloom filter to query whether the flow identification information is included in the AFT. If the flow identification information is not included in the AFT, the flow identification information is added to the AFT and the request packet or simulated packet is forwarded. If the flow identification information is included in the AFT, the request packet or simulated packet is directly forwarded.

[0082] In step S502, the switch determines whether the packet received by the LPort is a data packet. If it is not a data packet, the switch directly forwards it; if it is a data packet, the switch executes step S503.

[0083] In step S503, the switch determines whether the transport flow corresponding to the data packet has completed transmission based on the value of the ECN field in the data packet header. If ECN = 00, the transport flow corresponding to the data packet has completed transmission. The switch then maintains the CFT, updates the AFT, and forwards the data packet. Maintaining the CFT may include adding the flow identification information of the transport flow corresponding to the data packet to the CFT and setting a timer with a corresponding duration equal to the RTT. After the timer expires, the corresponding flow identification information in the CFT is deleted. Updating the AFT may include querying and deleting the flow identification information of the transport flow corresponding to the data packet contained in the AFT. If ECN ≠ 00, the transport flow corresponding to the data packet has not completed transmission, and the data packet is forwarded.

[0084] Switch Rport processing process:

[0085] In step S504, the switch determines whether the packet received by the RPort is a credit packet. If it is not a credit packet, the switch directly forwards it; if it is a credit packet, the switch executes step S505.

[0086] In step S505, the switch obtains the flow identification information (source IP address 1, destination IP address 1, source port number 1, destination port number 1) of the first transmission flow corresponding to the credit packet, and uses the Bloom filter to query whether the CFT contains the flow identification information of the first transmission flow based on the flow identification information of the first transmission flow. If the CFT does not contain the flow identification information of the first transmission flow, the credit packet is forwarded; if the CFT contains the flow identification information of the first transmission flow, S506 is executed.

[0087] At S506, the switch reuses the credit packet according to the AFT to schedule data packets in the transmission flow whose transmission has not yet been completed using the credit packet. The credit packet reuse includes:

[0088] a) Determining updated flow identification information from the AFT based on a maximum path coverage rule and a random determination rule may include: prioritizing matching entries in the AFT whose source IP address and destination IP address match source IP address 1 and destination IP address 1 of the first transmission flow; if no matching entry exists, matching entries in the AFT whose source IP address matches source IP address 1 of the first transmission flow or matching entries in the AFT whose destination IP address matches destination IP address 1 of the first transmission flow; if no matching entry exists, randomly and uniformly matching entries in the AFT and determining the flow identification information corresponding to the matched entry in the AFT as the updated flow identification information. For fairness, when there are multiple matching entries, the flow identification information corresponding to the multiple matched entries is determined as the updated flow identification information.

[0089] b) The switch updates the flow identification information corresponding to the credit packet to the updated flow identification information matched from the AFT, and forwards the credit packet to the corresponding source sending end via the source port identified by the updated flow identification information based on the updated flow identification information, so as to schedule the data packets in the transmission flow identified by the updated flow identification information through the credit packet, thereby reusing the part of the credit packets that were originally wasted.

[0090] Figure 6 is a flow chart of another credit packet reuse method provided by an embodiment of the present disclosure. In this embodiment, the intermediate device executing the disclosed method is a switch (which can be understood as a local switch). The switch's ports include local ports (LPorts) and remote ports (RPorts). The switch's storage area maintains two data structures related to transmission flow information: an AFT and a CFT. The AFT includes flow identification information for transmission flows that have not yet completed transmission (source sender IP address, destination receiver IP address, source sender port, destination receiver port), and the CFT includes flow identification information for transmission flows that have completed transmission and a corresponding timer with an RTT duration (source sender IP address, destination receiver IP address, source sender port, destination receiver port, and RTT duration). The AFT and CFT stored in the switch only include flow identification information for transmission flows connected to the source sender located in the local data center. The switch uses symmetric hashing for ECMP forwarding.

[0091] As shown in Figure 6, a switch connects four source senders, S1, S2, S3, and S4, to remote switches and destination receivers via long-distance links. The destination receivers include R1, R2, R3, and R4. Source sender S1 sends traffic flow F1 to destination receiver R1 through the switch, and also sends traffic flow F2 to destination receiver R2 through the switch. Source sender S2 sends traffic flow F3 to destination receiver R2 through the switch. Source sender S3 sends traffic flow F4 to destination receiver R3 through the switch. Source sender S4 sends traffic flow F5 to destination receiver R4 through the switch. The following describes AFT maintenance, CFT maintenance, and credit packet reuse, using Figure 6.

[0092] (1) AFT maintenance process

[0093] The switch maintains the AFT data structure shown in Table 1 based on the request packet or simulation packet received through Lport. The data structure includes the flow identification information of the transmission flow, the source sender IP address occupies 32 bits, the destination receiver IP address occupies 32 bits, the source sender port occupies 16 bits, and the destination receiver port occupies 16 bits.

[0094] Table 1 AFT data structure

[0095] The following describes the AFT process using this data structure, with reference to Figure 6. The switch receives a request packet or simulation packet (corresponding to transmission flows F1 to F5, respectively) via Lport, and obtains the flow identification information for transmission flows F1 to F5 corresponding to these request packets or simulation packets. Assuming that the flow identification information for F1 is (S1IP, R1IP, S1 port 1, R1 port), the flow identification information for F2 is (S1IP, R2IP, S1 port 2, R2 port), the flow identification information for F3 is (S2IP, R2IP, S2 port, R2 port), the flow identification information for F4 is (S3IP, R3IP, S3 port, R3 port), and the flow identification information for F5 is (S4IP, R4IP, S4 port, R4 port), the Bloom filter is used to query whether a corresponding entry exists in the AFT table. If not, the AFT table entry shown in Table 2 is established based on the obtained flow identification information.

[0096] Table 2 AFT list

[0097] (2) CFT maintenance process

[0098] A data packet is received from the source sender via Lport. The data packet includes flow end identification information. The switch maintains a CFT data structure as shown in Table 3 based on the received data packet. The data structure includes flow identification information of the transmission flow and the corresponding timer. The source sender IP address occupies 32 bits, the destination receiver IP address occupies 32 bits, the source sender port occupies 16 bits, the destination receiver port occupies 16 bits, and the timer occupies 16 bits. In this embodiment, the flow end identification information is implemented through ECN. Referring to Figure 7, an embodiment of the present disclosure provides a schematic diagram of an IP packet header format. The seventh and eighth bits of the 8-bit Type of Service (TOS) field in the IP packet header are used as ECN flags in data center networks. Data center switches now widely support the ECN function. Since actively transmitted data packets do not use the ECN function, this embodiment can use the ECN bits to implement flow end identification information. ECN = 00 indicates that the transmission flow corresponding to the data packet has ended transmission, and ECN ≠ 00 indicates that the transmission flow corresponding to the data packet has not ended transmission.

[0099] Table 3 CFT data structure

[0100] After the switch receives a data packet from the source sender through Lport, it obtains the flow identification information of the transmission flow corresponding to the data packet and checks the flow end identification information in the packet header of the data packet. When the transmission flow F1 from the local data center ends, the ECN in the packet header of the last data packet corresponding to F1 is 00. When the switch detects that the ECN of the data packet is 00, it queries the AFT shown in Table 2 and deletes the flow identification information of F1 in the AFT (S1IP, R1IP, S1 port 1, R1 port) to obtain the AFT list shown in Table 4, and adds the flow identification information of F1 and the corresponding timer with the RTT duration to the CFT table shown in Table 3 to obtain the CFT list shown in Table 5. When the timer times out, the table entry in the CFT is deleted.

[0101] Table 4 AFT list

[0102] Table 5 CFT list

[0103] (3) Credit package reuse process

[0104] The switch receives the credit packet corresponding to F1 from the destination receiving end through RPort, obtains the flow identification information of F1 corresponding to the credit packet (S1IP, R1IP, S1 port 1, R1 port), queries CFT, and determines that the flow identification information of F1 is included in the CFT. It can be understood that the flow identification information of F1 matches the table entry in the CFT that has not timed out, and the credit packet is reused according to AFT. Reusing the credit packet according to AFT may include, first, determining the updated flow identification information that matches the flow identification information of F1 from the AFT shown in Table 4 according to the path maximum coverage rule. At this time, it can be seen that there is no table entry in AFT that is completely consistent with the source sender IP address and the destination receiver IP address of F1, but there is a table entry that is consistent with the source sender IP address of F1. This table entry corresponds to the flow identification information of the transmission flow F2, so the flow identification information of F2 is determined as the updated flow identification information. Further, the flow identification information of F1 corresponding to the credit packet is updated to the flow identification information of F2, and the credit packet is forwarded according to the flow identification information of F2, so that the data packet in F2 is scheduled through the credit packet.

[0105] Figure 8 provides a credit package reuse device for an embodiment of the present disclosure, which includes: one or more processors 801; a memory 802, on which one or more programs are stored, and when the one or more programs are executed by one or more processors 801, the one or more processors 801 implement the above-mentioned credit package reuse method; and one or more I / O interfaces 803, connected between the processor 801 and the memory 802, and configured to implement information interaction between the processor 801 and the memory 802.

[0106] The processor 801 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 802 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 803 is connected between the processor 801 and the memory 802, and can realize information exchange between the processor 801 and the memory 802, including but not limited to a data bus (Bus), etc.

[0107] In some embodiments, the processor 801 , the memory 802 , and the I / O interface 803 are connected to each other via a bus 804 , and further connected to other components of the computing device.

[0108] FIG9 shows a computer-readable medium provided by an embodiment of the present disclosure, on which a computer program is stored. When the program is executed by a processor, the above-mentioned credit packet reuse method is implemented.

[0109] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units in the apparatus disclosed above may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0110] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be performed by several physical components in cooperation.

[0111] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; compact disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cassettes, tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0112] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for reusing credit packets, which is applied to an intermediate device connected to at least one source sender, and the method includes: Receiving a credit packet from a destination receiver, where the credit packet is used to schedule data packets in a first transmission stream corresponding to the credit packet; Obtaining flow identification information of the first transmission stream corresponding to the credit packet; Obtaining an active flow table and an ended flow table, where the active flow table includes flow identification information of transmission streams that have not ended transmission, and the ended flow table includes flow identification information of transmission streams that have ended transmission; And When it is determined that the flow identification information of the first transmission stream is included in the ended flow table, reusing the credit packet according to the active flow table to schedule data packets in the transmission stream that has not ended transmission through the credit packet.

2. The method according to claim 1, wherein The ended flow table further includes expiration time information, and the expiration time information is used to determine the expiration time of the flow identification information of the transmission stream that has ended transmission; The method further includes: Determining the expiration time according to the expiration time information; and After reaching the expiration time, deleting the flow identification information of the transmission stream that has ended transmission corresponding to the expiration time information in the ended flow table.

3. The method according to claim 2, wherein The expiration time information includes a timer with a preset duration or an expiration timestamp.

4. The method according to claim 1, wherein Reusing the credit packet according to the active flow table includes: Determining updated flow identification information from the active flow table according to a preset rule; Updating the flow identification information of the first transmission stream corresponding to the credit packet to the updated flow identification information; and Forwarding the credit packet according to the updated flow identification information to schedule data packets in the transmission stream corresponding to the updated flow identification information through the credit packet.

5. The method according to claim 4, wherein The flow identification information includes the source Internet Protocol (IP) address of the source sender of the transmission stream corresponding to the flow identification information and the destination IP address of the destination receiver corresponding to the flow identification information, and the preset rule includes a path maximum coverage rule, and the path maximum coverage rule includes being the same as the first source IP address and / or the first destination IP address included in the flow identification information of the first transmission stream; Determining updated flow identification information from the active flow table according to a preset rule includes: Comparing the first source IP address and / or the first destination IP address with the source IP address and / or the destination IP address included in the flow identification information of the transmission streams that have not ended transmission included in the active flow table; And Determining the flow identification information of the transmission stream that has not ended transmission in the active flow table and satisfies the path maximum coverage rule as the updated flow identification information.

6. The method according to claim 1, wherein, Before obtaining the active flow table and the ended flow table, the method further includes: Receiving a credit scheduling packet from a source sender, where the credit scheduling packet is used to schedule credit packets; Obtaining flow identification information of the transmission stream corresponding to the credit scheduling packet, where the credit scheduling packet corresponds to the same transmission stream as the credit packet it schedules; and Adding the flow identification information of the transmission stream corresponding to the credit scheduling packet to the active flow table to generate or update the active flow table.

7. The method according to claim 1, wherein, Before obtaining the active flow table and the ended flow table, the method further includes: Receive a data packet from a source sender, where the data packet includes stream end identification information for determining whether the transmission of the transmission stream corresponding to the data packet has ended; Obtain stream identification information of the transmission stream corresponding to the data packet; and When it is determined according to the stream end identification information that the transmission stream corresponding to the data packet has ended transmission, delete the stream identification information of the transmission stream corresponding to the data packet included in the active flow table to update the active flow table, and add the stream identification information of the transmission stream corresponding to the data packet to the ended flow table to generate or update the ended flow table.

8. The method according to claim 7, wherein The stream end identification information is at least one bit in the IP packet header of the data packet.

9. The method according to claim 8, wherein The stream end identification information is the ECN bit in the IP packet header.

10. A credit packet reuse device, comprising: One or more processors; A memory storing one or more programs which, when executed by the one or more processors, cause the one or more processors to implement the method according to any one of claims 1 to 9; And One or more I / O interfaces connected between the processor and the memory and configured to enable information interaction between the processor and the memory.

11. A computer-readable medium having stored thereon a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for providing credit-based flow control

    CN104320350A

  • A method and apparatus for sending a stop signal of a credit packet

    CN109039900A

  • Credit packet-based active transmission method for data center

    CN112437019A

  • Credit adjustment circuit and method, task scheduling circuit and method, circuit and medium

    CN114896045A

  • Positive feedback ethernet link flow control for promoting lossless ethernet

    US20140204742A1