Congestion handling method, network device and storage medium

By obtaining cache space parameters in the data center network and adjusting ECN parameters, the problem of low-priority services is solved, and the network transmission efficiency and performance of high-priority services is improved.

CN113810309BActive Publication Date: 2025-05-06ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010547269.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-16
Publication Date
2025-05-06
Estimated Expiration
2040-06-16

AI Technical Summary

Technical Problem

In data center networks, the congestion control mechanism of low-priority services is low, resulting in packet loss and network performance degradation when network equipment exits are congested.

Method used

By obtaining the cache space parameters of network devices, the cache space allocated to the high-priority message queue is increased, and the explicit congestion notification ECN parameters are adjusted according to the queue delay, so as to improve the flexibility of congestion control and network transmission efficiency.

Benefits of technology

It effectively avoids packet loss of high-priority services, improves network transmission efficiency, and ensures network transmission performance of high-priority services, especially in the case of network congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113810309B_ABST
    Figure CN113810309B_ABST
Patent Text Reader

Abstract

The present invention discloses a congestion processing method, a network device and a storage medium. The congestion processing method obtains a cache space parameter, increases the cache space allocated to a first message queue with a higher priority according to the cache space parameter, thereby avoiding the loss of messages in the first message queue, realizing dynamic adjustment, and improving the flexibility of congestion control. On this basis, the queuing delay of the first message queue after the cache space is increased is obtained, and the explicit congestion notification ECN parameter is adjusted according to the queuing delay to avoid the degradation of delay performance caused by the increase of the cache space, thereby improving the network transmission efficiency of high-priority services in the case of network congestion and ensuring the network transmission performance of high-priority services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a congestion processing method, network equipment and storage medium. Background Art

[0002] In existing data center networks, different types of services are usually mapped to different priorities. Generally speaking, the congestion control mechanism for low-priority services (such as TCP services) is less efficient. However, network devices (such as switches) have limited cache space except for the fixed static cache space and header cache space. Since the congestion adjustment mechanism for low-priority service queues is less efficient, when the network device outlet is congested, high-priority services are prone to packet loss, resulting in reduced network performance. Summary of the invention

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] The embodiment of the present invention provides a congestion processing method, a network device and a storage medium, which can improve the network transmission efficiency of high-priority services under the condition of network congestion.

[0005] In a first aspect, an embodiment of the present invention provides a congestion processing method, which is applied to a network device, wherein the network device forwards at least a first message queue and a second message queue, and the priority of the first message queue is higher than the priority of the second message queue. The method includes:

[0006] Acquire a cache space parameter of the network device, and increase the cache space allocated to the first message queue according to the cache space parameter;

[0007] The queuing delay of the first message queue after the buffer space is increased is obtained, and an explicit congestion notification (ECN) parameter is adjusted according to the queuing delay, wherein the ECN parameter is used to trigger congestion control.

[0008] In a second aspect, an embodiment of the present invention further provides a network device, comprising at least one processor and a memory for communicating with the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the congestion handling method as described in the first aspect.

[0009] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the congestion handling method described in the first aspect.

[0010] The embodiment of the present invention includes: obtaining a cache space parameter, increasing the cache space allocated to the first message queue according to the cache space parameter, obtaining the queuing delay of the first message queue after the cache space is increased, and adjusting the explicit congestion notification ECN parameter according to the queuing delay. By obtaining the cache space parameter, increasing the cache space allocated to the first message queue with a higher priority according to the cache space parameter, thereby avoiding the loss of messages in the first message queue, realizing dynamic adjustment, and improving the flexibility of congestion control. On this basis, the queuing delay of the first message queue after the cache space is increased is obtained, and the explicit congestion notification ECN parameter is adjusted according to the queuing delay to avoid the degradation of delay performance caused by the increase of the cache space, thereby improving the network transmission efficiency of high-priority services in the case of network congestion and ensuring the network transmission performance of high-priority services.

[0011] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.

[0013] Figure 1 It is a schematic diagram of the transmission architecture of a switch in a mixed operation scenario of a data center network provided by an embodiment of the present invention;

[0014] Figure 2 is a flow chart of a congestion handling method provided by an embodiment of the present invention;

[0015] Figure 3 is a schematic diagram of a cache message queue provided by an embodiment of the present invention;

[0016] Figure 4 It is a flow chart of increasing the cache space allocated to the first message queue according to the cache space parameter provided by an embodiment of the present invention;

[0017] Figure 5 is a flow chart of obtaining the queuing delay of the first message queue after the cache space is increased, provided by an embodiment of the present invention;

[0018] Figure 6 It is a flow chart of adjusting ECN parameters according to queuing delay provided by an embodiment of the present invention;

[0019] Figure 7It is a flowchart of adjusting ECN parameters according to the size relationship between the first queuing delay and the second queuing delay provided by an embodiment of the present invention;

[0020] Figure 8 is a flow chart of a congestion handling method provided by another embodiment of the present invention;

[0021] Fig. 9 is a schematic diagram of cache space allocation provided by an embodiment of the present invention;

[0022] Fig.10 It is a structural diagram of a network device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] It should be understood that in the description of the embodiments of the present invention, the meaning of multiple (or multiple) is more than two, greater than, less than, and exceeding are understood to exclude the number itself, and above, below, and within are understood to include the number itself. If there is a description of "first", "second", etc., it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0025] The embodiments of the present invention provide a congestion processing method, a network device and a storage medium, which are highly flexible and can improve the network transmission efficiency of high-priority services.

[0026] In existing data center networks, different types of services are usually mapped to different priorities. Generally speaking, the congestion control mechanism for low-priority services (such as TCP services) is less efficient. However, network devices (such as switches) have limited cache space except for the fixed static cache space and header cache space. Since the congestion adjustment mechanism for low-priority service queues is less efficient, when the network device outlet is congested, high-priority services are prone to packet loss, resulting in reduced network performance.

[0027] An embodiment of the present invention provides a congestion handling method, which is applied to a network device, which may be a switch, a router, etc. This embodiment is described using a switch as an example, and the cache space described in the embodiment of the present invention refers to a shared cache space of the switch.

[0028] A data center is a globally collaborative network of specific devices used to transmit, accelerate, display, calculate, and store data information on the Internet infrastructure. The data center network can transmit a wide variety of data types. The present invention uses a mixed scenario of TCP (Transmission Control Protocol) and RDMA (Remote Direct Memory Access) to illustrate. The priority of the TCP message queue is lower than that of the RDMA message queue. Figure 1 , is a schematic diagram of the transmission architecture of a switch in a mixed-running scenario of a data center network provided by an embodiment of the present invention, wherein A1 to A n There are multiple message queues entering the switch at port A, and then transmitted from the switch's B exit to the next node.

[0029] When the data traffic transmitted in the data center increases, congestion may easily occur at the B exit of the switch. Since TCP controls congestion through congestion window adjustment and packet retransmission, the efficiency is low, and the shared memory space of the switch is limited, it is easy for the TCP message queue to occupy a large amount of the switch's shared cache, which in turn restricts the RDMA message queue from entering the cache space. As a result, when the B exit is congested, the RDMA message queue quickly reaches the cache cutoff position, which is prone to packet loss.

[0030] Based on this, refer to Figure 2 The congestion handling method provided in the embodiment of the present application includes but is not limited to the following steps 201 to 203:

[0031] Step 201: Obtain cache space parameters;

[0032] Among them, in step 201, the cache space parameters include a first cache space parameter corresponding to the first message queue and a second cache space parameter corresponding to the second message queue, and the priority of the first message queue is higher than the priority of the second message queue; exemplarily, the first message queue is an RDMA message queue, and the second message queue can be a TCP message queue; those skilled in the art can understand that the first message queue and the second message queue can also be other types of message queues.

[0033] Step 202: increasing the buffer space allocated to the first message queue according to the first buffer space parameter and the second buffer space parameter;

[0034] In step 202, the buffer space allocated to the first message queue is increased by allocating most of the remaining buffer space to the first message queue, or allocating part of the remaining buffer space to the first message queue, depending on actual conditions.

[0035] In one embodiment, the first buffer space parameter may be a first queue depth of the first message queue, and the second buffer space parameter may be a second queue depth of the second message queue. The first queue depth and the second queue depth may be used to determine the buffer space occupancy of the first message queue and the second message queue.

[0036] Reference Figure 3 ,Take the second queue depth as an example, when the second queue depth reaches the cache cutoff position, ,it proves that the traffic of TCP packets is large and congestion is ,prone to occur.

[0037] Step 203: Obtain the queuing delay of the first message queue after the buffer space is increased, and adjust the explicit congestion notification ECN parameters according to the queuing delay.

[0038] Among them, in step 203, the ECN parameter is used to control whether to trigger congestion control. By obtaining the queuing delay of the first message queue after the cache space is increased in step 202, the explicit congestion notification ECN parameter is adjusted according to the queuing delay to avoid the delay performance degradation caused by the increase of the cache space.

[0039] In the above steps 201 to 203, by acquiring the cache space parameter, the cache space allocated to the first message queue is increased according to the cache space parameter, the queuing delay of the first message queue after the cache space is increased is acquired, and the explicit congestion notification ECN parameter is adjusted according to the queuing delay. By acquiring the cache space parameter, the cache space allocated to the first message queue with a higher priority is increased according to the cache space parameter, thereby avoiding the loss of messages in the first message queue, realizing dynamic adjustment, and improving the flexibility of congestion control. On this basis, the queuing delay of the first message queue after the cache space is increased is acquired, and the explicit congestion notification ECN parameter is adjusted according to the queuing delay to avoid the degradation of delay performance caused by the increase of the cache space, thereby improving the network transmission efficiency of high-priority services in the case of network congestion and ensuring the network transmission performance of high-priority services.

[0040] Reference Figure 4 In one embodiment, in the above step 202, increasing the buffer space allocated to the first message queue according to the first buffer space parameter and the second buffer space parameter may specifically include the following steps 401 to 402:

[0041] Step 401: Compare the first queue depth and the second queue depth;

[0042] Among them, in step 401, the traffic statistics sampling technology of the switch can be used to send the sampled traffic to the CPU, and the message queues where the traffic of different priorities is located and the queue depth of each message queue can be identified according to the second and third layer header fields, and then the queue depths of different message queues can be compared.

[0043] Step 402: Determine whether the first queue depth is less than the second queue depth. If the first queue depth is less than the second queue depth, jump to step 403;

[0044] Step 403: Increase the buffer space allocated to the first message queue.

[0045] Among them, in step 402 to step 403, when the depth of the first queue is less than the depth of the second queue, that is, the number of messages in the second message queue is large, there is a risk of packet loss in the first message queue, so the cache space allocated to the first message queue is increased, thereby avoiding the loss of messages in the first message queue, realizing dynamic adjustment, and improving the flexibility of congestion control.

[0046] In one embodiment, the cache space parameters may include not only the first queue depth and the second queue depth, but also the cache utilization of the switch. When the cache utilization of the switch is too high, the cache space may be exhausted at any time, and packet loss may occur easily. On this basis, in the above step 401, the cache utilization may be used as a judgment basis, that is, when the first queue depth is less than the second queue depth and the cache utilization exceeds the first threshold, the cache space allocated to the first message queue is increased. By combining the first queue depth, the second queue depth and the cache utilization as the basis for increasing the cache space allocated to the first message queue, the rationality of the cache space adjustment may be improved. Exemplarily, when the first queue depth is less than the second queue depth, but the cache utilization of the switch is not large, for example, about 40%, the cache space allocated to the first message queue may not be increased at this time, so as to ensure that both the first message queue and the second message queue can be efficiently transmitted, and the overall performance of the network is ensured.

[0047] It is understandable that the first threshold may be preset according to actual conditions, for example, may be set to 90%.

[0048] In one embodiment, in the above step 202, the buffer space allocated to the first message queue is increased according to the first buffer space parameter and the second buffer space parameter, which may be specifically:

[0049] The cache cutoff position of the first message queue is increased according to the first cache space parameter and the second cache space parameter. Figure 3, increasing the cache cutoff bit of the first message queue can increase the maximum value of the first queue depth, so that the switch can cache more messages in the first message queue, thereby avoiding the loss of messages in the first message queue.

[0050] Among them, raising the cache cutoff of the first message queue may be to allocate all the remaining cache space to the first message queue. Exemplarily, when the depth of the first queue is less than the depth of the second queue and the cache utilization of the switch exceeds 90%, the remaining 10% of the cache space is allocated to the first message queue. At this time, the messages of the second message queue are no longer cached, that is, the messages of the second message queue are directly forwarded. Alternatively, most of the cache space in the remaining cache space is allocated to the first message queue. Exemplarily, when the depth of the first queue is less than the depth of the second queue and the cache utilization of the switch exceeds 90%, the remaining 8% of the cache space is allocated to the first message queue, and the remaining 2% is allocated to the second message queue.

[0051] In one embodiment, before adjusting the buffer space of the first message queue, a pre-judgment condition may be added to improve the rationality of the adjustment. Specifically, the second threshold, the third threshold, and the fourth threshold are first pre-set, wherein the second threshold is the bandwidth utilization when congestion occurs, the third threshold is the transmission delay of the switch, and the fourth threshold is the PFC (Priority-based Flow Control) packet sending rate.

[0052] Before adjusting the cache space of the first message queue, the first bandwidth utilization is obtained. When the first bandwidth utilization exceeds the second threshold, it can be determined that the current bandwidth utilization is too high and the congestion of the switch may be aggravated. Therefore, the cache space parameter of the switch is obtained to adjust the cache space of the first message queue and the second message queue. Exemplarily, the second threshold can be set to 98%.

[0053] The first transmission delay of the switch is obtained. When the first transmission delay exceeds the third threshold, it can be determined that the current transmission delay of the switch is too high and the congestion of the switch may be aggravated. Therefore, the buffer space parameter of the switch is obtained to adjust the buffer space of the first message queue and the second message queue. Exemplarily, the third threshold can be set to 50 microseconds.

[0054] The packet sending rate of the priority-based flow control PFC is obtained. When the packet sending rate exceeds the fourth threshold, it is easy to trigger too many PFCs and increase the risk of deadlock and packet loss. Therefore, the cache space parameters of the switch are obtained to adjust the cache space of the first message queue and the second message queue. Exemplarily, the third threshold can be set to 10 per second.

[0055] It can be understood that the above-mentioned pre-judgment process based on the second threshold, the third threshold and the fourth threshold can be set one by one or all of them, depending on the specific network requirements.

[0056] Reference Figure 5 In one embodiment, in the above step 203, obtaining the queuing delay of the first message queue after the buffer space is increased may specifically include the following steps 501 to 502:

[0057] Step 501: Obtaining a queue length of a first message queue and a transmission rate of the first message queue within a unit time;

[0058] In step 501, the unit time can be freely set according to the actual situation. Figure 3 , which can be implemented by using a timestamp tag. For example, the unit time can be 10 microseconds. The queue length of the first message queue can reflect the number of messages in the first message queue per unit time, which can be read using the switch's built-in function. The transmission rate of the first message queue can be the instantaneous rate of the tail of the queue leaving the queue, where the instantaneous rate can be obtained by dividing the total number of messages cached per unit time by the unit time.

[0059] Step 502: Obtain the queuing delay of the first message queue according to the queue length and the transmission rate.

[0060] In step 502, the queue delay of the first message queue is obtained by dividing the queue length of the first message queue by the transmission rate.

[0061] In one embodiment, the queuing delay of the first message queue may be averaged to obtain an average queuing delay, and the average queuing delay may be used as a basis for judgment, which is conducive to improving the accuracy of judgment.

[0062] Reference Figure 6 In one embodiment, in the above step 203, adjusting the ECN parameter according to the queuing delay may specifically include the following steps 601 to 602:

[0063] Step 601: obtaining a current first queuing delay of a first message queue and an initial second queuing delay of the first message queue;

[0064] Wherein, in step 601, the current first queuing delay of the first message queue and the initial second queuing delay of the first message queue are continuously obtained. When the cache space of the first message queue is not increased during the first acquisition, the queuing delay of the first message queue is the second queuing delay, and after the cache space of the first message queue is increased, the queuing delay of the first message queue is the first queuing delay; when the next acquisition is performed, the first queuing delay obtained in the previous round is the second queuing delay of the current round, and the queuing delay of the first message queue re-acquired in the current round is the new first queuing delay, and so on. Of course, the first queuing delay when first acquired can also be the queuing delay of the first message queue after the cache space of the first message queue is increased, depending on the timing of acquisition. In short, the first queuing delay is the queuing delay obtained in the current round, and the second queuing delay is the queuing delay obtained in the previous round.

[0065] Step 602: Adjust the ECN parameter according to the magnitude relationship between the first queuing delay and the second queuing delay.

[0066] In step 602, when the first queuing delay is greater than the second queuing delay, it means that congestion is aggravated; when the first queuing delay is less than the second queuing delay, it means that congestion is relieved, so the ECN parameters are adjusted according to the congestion situation.

[0067] Reference Figure 7 In one embodiment, in the above step 602, adjusting the ECN parameter according to the magnitude relationship between the first queuing delay and the second queuing delay may specifically include the following steps 701 to 703:

[0068] Step 701: when the first queuing delay is greater than the second queuing delay, obtaining the difference between the first queuing delay and the second queuing delay;

[0069] Wherein, in step 701, T1 represents the first queuing delay, T2 represents the second queuing delay, and the above difference β can be expressed as:

[0070] β = T1-T2;

[0071] When β is less than 0, it means that congestion is relieved; when β is greater than 0, it means that congestion is aggravated.

[0072] Step 702: Obtaining a threshold adjustment coefficient according to the difference;

[0073] In step 702, an adjustable parameter α is introduced, and a threshold adjustment coefficient F is obtained according to the difference, that is:

[0074] F new =(1-αβ)*F old , where F new is the threshold adjustment coefficient of the current round, F oldis the threshold adjustment coefficient of the previous round, where 0<α<1, and is used to fine-tune β.

[0075] Step 703: Use the threshold adjustment coefficient to reduce the ECN threshold value and / or reduce the ECN marking probability.

[0076] In step 703, the ECN parameter K is adjusted using the threshold adjustment coefficient, that is, K new =K old *F new , where K new is the ECN threshold parameter of the current round, K old is the ECN threshold parameter of the previous round, wherein, when the buffer space of the first message queue is not adjusted, the initial ECN threshold parameter can be obtained by using a DC-QCN (Date Center Quantized Congestion Notification) algorithm.

[0077] In one embodiment, the ECN parameters may include an ECN threshold value and an ECN marking probability. Therefore, in the above step 703, the ECN threshold value may be reduced by using a threshold adjustment coefficient, or the ECN marking probability may be reduced by using a threshold adjustment coefficient. Reducing the ECN threshold value may facilitate timely triggering of ECN marking, thereby performing congestion control, and reducing the ECN marking probability may ensure the throughput of data packets with large traffic. It is understandable that the above-mentioned reduction of the ECN threshold value and the reduction of the ECN marking probability may be performed either or both, depending on the specific network adjustment requirements.

[0078] In one embodiment, after adjusting the ECN parameters, a verification step may be performed, which may be:

[0079] The second bandwidth utilization is obtained. When the second bandwidth utilization is lower than the fifth threshold, the cache space allocated to the first message queue and the second message queue is restored to an initial state, wherein the second bandwidth utilization is the bandwidth utilization of the switch after the ECN parameters are adjusted. Correspondingly, the fifth threshold may be 70%.

[0080] A second transmission delay is obtained. When the second transmission delay is lower than a sixth threshold, the cache space allocated to the first message queue and the second message queue is restored to an initial state, wherein the second transmission delay is the transmission delay of the switch after the ECN parameters are adjusted. Correspondingly, the sixth threshold may be 40 microseconds.

[0081] Among them, the cache space allocated to the first message queue and the second message queue is restored to the initial state, that is, the cache space allocation of the switch when increasing the cache space of the first message queue and adjusting the ECN parameters in the above embodiment is not performed. The initial allocation method depends on the specific network requirements and is not listed here one by one.

[0082] The cache space is restored to its initial allocation state, which can ensure that messages of various priorities can be effectively transmitted.

[0083] It is understandable that the above judgment process based on the fifth threshold and the sixth threshold can be set either one or all of them, depending on specific network requirements.

[0084] The congestion handling method of the present application is described in detail below with a practical example.

[0085] Reference Figure 8 The embodiment of the present invention further provides a congestion handling method, including but not limited to the following steps 801 to 810:

[0086] Step 801: Determine whether the cache space of the switch is occupied. If the cache space of the switch is occupied, jump to step 802, otherwise end the process;

[0087] Step 802: Obtain the bandwidth utilization and transmission delay preset by the switch;

[0088] Step 803: Determine whether the bandwidth utilization exceeds a preset value. If the bandwidth utilization exceeds the preset value, jump to step 805; otherwise, jump to step 804.

[0089] Step 804: determine whether the transmission delay exceeds the preset value. If the transmission delay exceeds the preset value, jump to step 805, otherwise end the process;

[0090] Step 805: Obtain the queue depths of the message queues of different priorities and the buffer utilization of the switch;

[0091] Step 806: Determine whether the queue depth of the low priority message queue is greater than the queue depth of the high priority message queue, and whether the cache utilization of the switch exceeds the preset value. If the queue depth of the low priority message queue is greater than the queue depth of the high priority message queue and the cache utilization of the switch exceeds the preset value, jump to step 807, otherwise end the process;

[0092] Step 807: Allocate all remaining buffer space of the switch to the high-priority message queue, and directly forward the low-priority message queue;

[0093] Step 808: By means of timestamp marking, the number of buffers and queue length of the high priority message queue in a unit time are counted to obtain the queuing delay of the high priority message queue, determine the threshold adjustment coefficient, and dynamically adjust the ECN parameter;

[0094] Step 809: Obtain the bandwidth utilization and transmission delay of the switch after adjusting the ECN parameters;

[0095] Step 810: if the bandwidth utilization of the switch after adjusting the ECN parameters is lower than the preset value, jump to step 811; if the bandwidth utilization of the switch after adjusting the ECN parameters is higher than the preset value, jump to step 801;

[0096] Step 811: If the transmission delay of the switch is lower than the preset value after adjusting the ECN parameters, the process ends; otherwise, jump to step 801.

[0097] In the above steps 801 to 811, whether the switch is congested is first determined by determining whether the cache space of the switch is occupied. If the cache space of the switch is not occupied, that is, the switch is not congested, no processing is required at this time. When the switch is congested, it is determined whether the bandwidth utilization and transmission delay of the switch exceed the preset values. If they do not exceed the preset values, it means that the network is in good condition and no processing is required. When the bandwidth utilization and transmission delay of the switch exceed the preset values, it means that the network congestion is serious. By obtaining the queue depths of message queues of different priorities and the cache utilization of the switch, the cache space allocated to the message queue with higher priority is increased according to the size relationship between the queue depths of message queues of different priorities and the cache utilization of the switch, thereby avoiding the loss of messages in the high-priority message queue, realizing dynamic adjustment, and improving the flexibility of congestion control. On this basis, the queuing delay of the high-priority message queue after the cache space is increased is obtained, and the threshold adjustment coefficient is determined according to the queuing delay. The ECN parameters are dynamically adjusted to avoid the degradation of delay performance caused by the increase of the cache space, thereby improving the network transmission efficiency of high-priority services in the case of network congestion and ensuring the network transmission performance of high-priority services. After adjusting the ECN parameters, by obtaining the bandwidth utilization and transmission delay of the switch again, it is determined whether it is necessary to re-execute the steps of increasing the cache space of the high-priority message queue and adjusting the ECN parameters until the bandwidth utilization and transmission delay of the switch meet the requirements, that is, the above steps 801 to 811 are executed in a loop.

[0098] For example, the high-priority message queue is an RDMA message queue, and the low-priority message queue is a TCP message queue. After all the remaining cache space of the switch is allocated to the high-priority message queue, the cache space of the switch is allocated as follows: Fig. 9 shown.

[0099] It should also be understood that the various implementations provided in the embodiments of the present invention can be combined arbitrarily to achieve different technical effects.

[0100] Fig.10 The embodiment of the present invention provides a network device 1000 provided by the embodiment of the present invention. The network device 1000 includes: a memory 1001, a processor 1002, and a computer program stored in the memory 1001 and executable on the processor 1002. When the computer program is executed, it is used to execute the above congestion processing method.

[0101] The processor 1002 and the memory 1001 may be connected via a bus or other means.

[0102] The memory 1001 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs and non-transitory computer executable programs, such as the congestion handling method described in the embodiment of the present invention. The processor 1002 implements the above congestion handling method by running the non-transitory software programs and instructions stored in the memory 1001.

[0103] The memory 1001 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store and execute the above-mentioned congestion handling method. In addition, the memory 1001 may include a high-speed random access memory 1001, and may also include a non-volatile memory 1001, such as at least one disk memory 1001 piece, a flash memory device or other non-volatile solid-state memory 1001 piece. In some embodiments, the memory 1001 may optionally include a memory 1001 remotely arranged relative to the processor 1002, and these remote memories 1001 may be connected to the network device 1000 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0104] The non-transient software programs and instructions required to implement the above congestion handling method are stored in the memory 1001. When executed by one or more processors 1002, the above congestion handling method is executed, for example, Figure 4 Steps 401 to 402 of the method, Figure 6 In steps 601 to 602 of the method, Figure 7 In steps 701 to 703 of the method, Figure 8 Method steps 801 to 811.

[0105] The embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the above-mentioned congestion handling method.

[0106] In one embodiment, the computer-readable storage medium stores computer-executable instructions, which are executed by one or more control processors 1002, for example, by a processor 1002 in the network device 1000, so that the one or more processors 1002 can execute the congestion handling method, for example, Figure 4 Steps 401 to 402 of the method, Figure 6 In steps 601 to 602 of the method, Figure 7 In steps 701 to 703 of the method, Figure 8 Method steps 801 to 811.

[0107] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0108] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory 1001 technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically include computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0109] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions under the shared conditions without violating the spirit of the present invention. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A congestion handling method, comprising: Acquire a cache space parameter, wherein the cache space parameter includes a first cache space parameter corresponding to the first message queue and a second cache space parameter corresponding to the second message queue, the priority of the first message queue is higher than the priority of the second message queue, the cache space parameter is used to characterize the cache space occupancy, and the cache space is a shared cache space of the switch; Determining, according to the first cache space parameter and the second cache space parameter, that the cache space occupancy of the first message queue is lower than the cache space occupancy of the second message queue, and increasing the cache space allocated to the first message queue; Acquire a queuing delay of the first message queue after the cache space is increased, and adjust an explicit congestion notification (ECN) parameter according to the queuing delay, wherein the ECN parameter is used to trigger congestion control; The increasing the buffer space allocated to the first message queue includes one of the following: Allocating all remaining buffer space to the first message queue; Part of the remaining buffer space is allocated to the first message queue.

2. The congestion handling method according to claim 1, characterized in that: The first cache space parameter includes a first queue depth of the first message queue; The second buffer space parameter includes a second queue depth of the second message queue.

3. The congestion handling method according to claim 2, characterized in that: The determining, according to the first cache space parameter and the second cache space parameter, that the cache space occupancy of the first message queue is lower than the cache space occupancy of the second message queue, and increasing the cache space allocated to the first message queue includes: Comparing the first queue depth and the second queue depth; When the first queue depth is less than the second queue depth, the buffer space allocated to the first message queue is increased.

4. The congestion handling method according to claim 3, characterized in that: The cache space parameter also includes cache utilization, and when the first queue depth is less than the second queue depth, increasing the cache space allocated to the first message queue includes: When the first queue depth is less than the second queue depth and the cache utilization exceeds a first threshold, the cache space allocated to the first message queue is increased.

5. The congestion handling method according to claim 3 or 4, characterized in that: The increasing the cache space allocated to the first message queue includes: Raise the cache cutoff bit of the first message queue.

6. The congestion handling method according to claim 1, characterized in that: The obtaining of the queuing delay of the first message queue after the cache space is increased includes: Obtaining a queue length of the first message queue and a transmission rate of the first message queue within a unit time; The queuing delay of the first message queue is obtained according to the queue length and the transmission rate.

7. The congestion handling method according to claim 1, characterized in that: The adjusting of ECN parameters according to the queuing delay includes: Obtaining a current first queuing delay of the first message queue and an initial second queuing delay of the first message queue; The ECN parameter is adjusted according to the magnitude relationship between the first queuing delay and the second queuing delay.

8. The congestion handling method according to claim 7, characterized in that: The adjusting the ECN parameter according to the magnitude relationship between the first queuing delay and the second queuing delay includes: When the first queuing delay is greater than the second queuing delay, obtaining a difference between the first queuing delay and the second queuing delay; Obtaining a threshold adjustment coefficient according to the difference; The threshold adjustment coefficient is used to reduce the ECN threshold value and / or reduce the ECN marking probability.

9. The congestion handling method according to claim 1, characterized in that: The obtaining of cache space parameters includes at least one of the following: Acquire a first bandwidth utilization, and when the first bandwidth utilization exceeds a second threshold, acquire a cache space parameter; Acquire a first transmission delay, and when the first transmission delay exceeds a third threshold, acquire a cache space parameter; A packet sending rate of priority-based flow control PFC is obtained, and when the packet sending rate exceeds a fourth threshold, a buffer space parameter is obtained.

10. The congestion handling method according to claim 1, characterized in that: The method further comprises at least one of the following: acquiring a second bandwidth utilization, and when the second bandwidth utilization is lower than a fifth threshold, restoring the buffer space allocated to the first message queue and the second message queue to an initial state; A second transmission delay is obtained, and when the second transmission delay is lower than a sixth threshold, the buffer space allocated to the first message queue and the second message queue is restored to an initial state.

11. A network device, characterized in that: It includes at least one processor and a memory for communicating with the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the congestion handling method as described in any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the congestion handling method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Self-adaptive ECN (Explicit Congestion Notification) marking method and device in data center

    CN106789701A

  • Signal route conversion device based on priority strategy and automobile

    CN210380905U