Network interface device for user-defined congestion control
By implementing a user-defined congestion control algorithm on the network interface device and dynamically adjusting packet forwarding parameters, the problem of network congestion management is solved, and network performance and stability are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
- Filing Date
- 2025-10-21
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to effectively manage and anticipate network congestion, especially in shared network environments where applications have varying bandwidth requirements, leading to reduced data transmission speeds and packet loss.
By configuring the network interface device to execute a user-defined congestion control algorithm, the first and second processing layers work together to dynamically adjust packet forwarding parameters to cope with network congestion.
It enables flexible congestion control based on specific network environment needs, improves network performance and robustness, adapts to changes in service characteristics, and reduces waiting time and packet loss.
Smart Images

Figure CN121967330A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to network interface devices for user-defined congestion control. Background Technology
[0002] When traffic exceeds the capacity of the network environment, network congestion typically occurs, leading to reduced data transmission speeds, increased latency, and potential packet loss. In practice, controlling network congestion can be challenging due to various factors. First, the dynamic and unpredictable nature of network traffic in different network environments makes it difficult to effectively anticipate and manage congestion. Furthermore, diverse applications with their own bandwidth requirements often share the network environment. For example, bandwidth-intensive applications such as video streaming, online gaming, and distributed training of artificial intelligence (AI) models can cause traffic surges and exacerbate network congestion. Therefore, implementing congestion control to improve network performance is desirable. Summary of the Invention
[0003] According to examples of this disclosure, a network interface device can be configured to perform congestion control based on any suitable user-defined congestion control algorithm. In one aspect, examples of this disclosure provide a network interface device (see...). Figure 1 103), comprising: a processor; and a non-transitory computer-readable medium having instructions stored thereon, which, when executed by the processor, cause the processor to implement a first processing layer and a second processing layer to perform congestion control (see [reference]). Figure 1 (110 to 120 in the middle).
[0004] In one example, the first processing layer may receive an event notification from the second processing layer. In response to determining that congestion control is needed based on the event notification, the first processing layer may determine an adjustment to packet forwarding parameters by applying a congestion control algorithm. The congestion control algorithm may be one of several congestion control algorithms programmable to be applied by the first processing layer. The first processing layer may generate and send an instruction to the second processing layer to perform the adjustment. Based on the instruction, the second processing layer may adjust the packet forwarding parameters from a first value to a second value, specifically by configuring components of the network interface device (e.g., a hardware scheduler) to control packet forwarding to the physical network based on the second value. See also Figure 1 Between 140 and 160.
[0005] Another aspect may include a non-transitory computer-readable storage medium comprising a set of instructions that, in response to execution by a processor of the network interface device, cause the processor to implement a first processing layer and a second processing layer to perform congestion control according to an example of the present disclosure. Yet another aspect may include a method for a network interface device including a first processing layer and a second processing layer to perform congestion control. Still another aspect may include a computer system comprising a network interface device according to an example of the present disclosure. Attached Figure Description
[0006] Figure 1 This is a schematic diagram illustrating an example network interface device used to perform user-defined congestion control in a network environment.
[0007] Figure 2 This is a flowchart of an example process for a network interface device to perform user-defined congestion control in a network environment.
[0008] Figure 3 This is a flowchart illustrating a detailed example process for a network interface device to perform user-defined congestion control in a network environment.
[0009] Figure 4A It is a diagrammatic explanation. Figure 1 A schematic diagram of the first example programming of the first processing layer in the diagram.
[0010] Figure 4B It is a diagrammatic explanation. Figure 1 A schematic diagram of the second example programming of the first processing layer in the diagram.
[0011] Figure 5 This is a diagram illustrating how the first processing layer applies user-defined rules to determine whether to send a probe packet.
[0012] Figure 6 This is a schematic diagram illustrating an example rate-based congestion control algorithm that can be applied by programming the first processing layer.
[0013] Figure 7 This is a schematic diagram illustrating an example window-based congestion control algorithm that can be applied by programming the first processing layer.
[0014] Figure 8 This is a schematic diagram illustrating an example of a user-defined congestion control algorithm that can be applied by programming the first processing layer.
[0015] Figure 9 This is a schematic diagram illustrating an example distributed training environment in which network interface devices can be deployed to perform congestion control.
[0016] Figure 10This is a diagram illustrating an example software-defined networking (SDN) environment. Detailed Implementation
[0017] In the following embodiments, reference is made to the accompanying drawings, which form a part of this document. In the drawings, similar symbols generally identify similar components unless the context otherwise requires. The illustrative embodiments described in the embodiments, drawings, and claims are not intended to be limiting. Other embodiments and changes may be utilized without departing from the spirit or scope of the subject matter presented herein. It should be readily understood that the aspects of this disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are expressly covered herein.
[0018] Although the terms “first” and “second” are used to describe various elements, these elements should not be limited by these terms. These terms are used to distinguish one element from another. For example, the first element can be called the second element, and vice versa. As used herein, the phrase “at least one of…” following a list of items (where each item is separated by the terms “and” or “or”) modifies the entire list, not each member of the list (i.e., each item). The phrase “at least one of…” does not require a selection from at least one of every listed item; rather, the phrase allows for the meaning of at least one of any item and / or at least one of any combination of items. By example, the phrases “at least one of A, B, and C” or “at least one of A, B, or C” each refer to only A, only B, or only C; and / or any combination of A, B, and C. In instances where a selection is intended to be made from "at least one of each of A, B and C" or, alternatively, "at least one of A, at least one of B and at least one of C", it will be explicitly described as such.
[0019] Example network interface device
[0020] Figure 1This is a schematic diagram illustrating an example network interface device 103 used to perform user-defined congestion control in network environment 100. Here, network environment 100 may include a first computer system 101 capable of communicating with second computer systems 105 to 106 via physical network 104. Computer system 101 may implement any suitable application 102 (for simplicity, one application is shown). The term "application" can generally refer to any suitable software capable of running on computer system 101. Application 102 may be configured to perform any suitable task or function as a standalone application or as part of a larger software suite. For example, application 102 may be implemented by worker nodes in a distributed environment (see...). Figure 9 This is implemented in the training of artificial intelligence (AI) models, etc.
[0021] To transmit data via physical network 104, computer system 101 may use network interface device 103 to send and receive packets. As used herein, the term "network interface device" generally refers to any suitable device configured to interface with or connect to a physical network in order to receive data from and transmit data to the physical network. Network interface device 103 may include any software, firmware, and / or hardware components that enable computer system 101 to exchange data with physical network 104. The term "physical network" generally refers to a network of interconnected physical devices. Physical devices may include physical servers, physical routers, physical switches, any combination thereof, etc.
[0022] The network interface device 103 may be a standalone component (e.g., a card inserted into a slot within the computer system 101) or integrated with another component of the computer system 101 (e.g., a motherboard). Figure 1 In the examples provided, network interface device 103 may be referred to as a physical network interface controller (NIC). Depending on the network environment, network interface device 103 may be referred to as a "network adapter," "network interface card," "network interface unit," "Ethernet card," etc. Various examples will be described using NIC 103 in the following text.
[0023] In practice, it has been observed that no single congestion control algorithm can achieve optimal performance across all types of network environments 100. Service characteristics and congestion conditions often vary from one network environment to another, making it challenging to effectively respond to and manage congestion. To improve congestion control and network performance, examples of this disclosure can be implemented to facilitate the programming of user-defined congestion control algorithms on the NIC 103. This capability allows users (e.g., network administrators) to develop and fine-tune their own congestion control algorithms on the NIC 103 to meet the specific needs of network environment 100.
[0024] exist Figure 1 In the examples described herein, NIC 103 may include multiple layers, such as a first processing layer 110, a second processing layer 120, and a hardware layer 130. As used herein, the term "layer" generally refers to one or more components configured to provide a set of functions or capabilities within NIC 103. For example, the first processing layer 110 and the second processing layer 120 may be implemented using software, firmware, hardware, or any combination thereof. The term "software" generally refers to a program, process, or instruction that enables NIC 103 to perform the examples of this disclosure. The term "firmware" generally refers to a type of software that may, for example, be embedded in the hardware components of NIC 103.
[0025] According to examples of this disclosure, the first processing layer 110 is programmable or configurable to apply one of a plurality of user-defined congestion control algorithms, such as congestion control algorithm 111 (denoted as "A1"). As a congestion control measure, the user-defined congestion control algorithm 111 may specify user-defined formulas for determining adjustments to packet forwarding parameters. The first processing layer 110 may also include user-defined state machine logic 112 (more commonly referred to as "user-defined logic"), which specifies rules for determining actions (e.g., state transitions) based on inputs (e.g., event notifications). For example, the user-defined state machine logic 112 may determine whether to immediately send telemetry packets (e.g., probe packets) for measurement purposes, or trigger a delay action to send the probe packets at a later time. Depending on the implementation, the user-defined state machine logic 112 may also determine how to handle other events, such as congestion events, session events, etc. The first processing layer 110 may further include any other modules or components, such as an initialization / configuration handler 113, a session event handler 114, etc.
[0026] As used herein, the term "congestion control" generally refers to a method for controlling the amount of data (e.g., packets) flowing through a network. For example, congestion control may be performed to reduce the number of packets transmitted via physical network 104. The term "congestion control algorithm" generally refers to steps or operations that may be performed to manage congestion. The term "user-defined state machine logic" or "user-defined logic" generally refers to one or more rules used to determine whether to perform an action (e.g., transition from one state to another) based on input (e.g., event notification). The term "packet" generally refers to a group of bits that can be transmitted together and may take another form, such as a "frame," "message," "segment," etc. The term "traffic" or "flow" generally refers to multiple packets. Packets may be data / control packets, etc.
[0027] The term "user-defined" generally refers to functionality specified or programmed by a user rather than pre-configured or provided by a manufacturer or provider. The term "user" generally refers to any suitable entity capable of programming the first processing layer 110, such as a human user (e.g., a network administrator, device customer), software application, AI agent, etc. The term "programmable to apply" generally refers to the first processing layer 110 being configured (e.g., using instructions executable by processor 131) to run or execute congestion control algorithm 111 and / or state machine logic 112.
[0028] The second processing layer 120 may represent a framework for implementing multiple congestion control algorithms configured to support programmable applications of the first processing layer 110. For example, the second processing layer 120 may be configured to provide various support functions to allow any (compatible) first processing layer 110 to utilize hardware layer 130 for congestion control. Figure 1 In the example, the second processing layer 120 may include a telemetry module 121 for providing detection generation and processing functions, an event loop module 122 for providing event detection and processing functions, a parameter adjustment engine 123 for providing parameter adjustment functions, and a data storage area (not shown) for storing session context information.
[0029] In practice, the first processing layer 110 may be referred to as a user-defined congestion control (UDCC) procedure, and the second processing layer 120 may be referred to as a UDCC framework. Here, the first processing layer 110 may represent a programmable component residing on top of the UDCC framework provided by the second processing layer 120. As will be further described below, the first processing layer 110 may control how events reported by the second processing layer 120 are processed, and how packet forwarding parameters may be adjusted. To facilitate inter-layer communication, the first processing layer 110 and the second processing layer 120 may be configured to have a shared interface (e.g., API) and / or a shared event data structure.
[0030] Hardware layer 130 may include any suitable physical or hardware components, such as processor 131, memory / storage device 132 for storing program code or instructions executable by processor 131 to implement layers 110 to 120, hardware scheduler 133, hardware queues 134 to 135, etc. Processor 131 may include an embedded central processing unit (eCPU), etc. Hardware layer 130 may include any hardwired circuit system, such as one or more application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), etc. Here, the term "hardware scheduler" may generally refer to a hardware component configured to manage the timing and / or order of packet transmission. For example, hardware scheduler 133 may operate at the hardware layer to control how packets are queued and transmitted via physical network 104. The hardware queue may include a transmit (TX) queue 134 for storing outgoing (i.e., transmitted) packets for transmission by NIC 103 and a receive (RX) queue 135 for storing incoming (i.e., received) packets received by NIC 103. Processing and hardware layers 110 to 130 will be further described below.
[0031] User-defined congestion control
[0032] Figure 2 This is a flowchart of an example process 200 for performing congestion control using network interface device 103. Example process 200 may include one or more operations, functions, or actions illustrated by one or more blocks (e.g., 210 to 260). Depending on the implementation, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or eliminated. Examples of this disclosure can be implemented using any suitable "network interface device" (e.g., NIC 103) capable of interfacing with physical network 104, etc.
[0033] exist Figure 2 At positions 210 to 220, the first processing layer 110 can receive an event notification that identifies an event detected by the second processing layer 120 (see [link]). Figure 1 140 in the first example). Figures 6 to 7 In the description, an event notification can indicate that a telemetry packet (e.g., a probe response) has been received for metric measurement. In practice, metric measurement can be performed to implement a congestion control algorithm based on in-band telemetry (INT). In the second example (which will use...), Figure 8In the description, the event notification may indicate that the second processing layer 120 has detected at least one of the following congestion events: a retransmission timeout (RTO) event, a sequence error negative acknowledgment (NAK) event, and a congestion notification point (CNP) event. For example, a CNP event may be sent by the destination (e.g., the second computer system 105 / 106) to signal congestion to the source (e.g., the first computer system 101) for implementation of a congestion control algorithm based on explicit congestion notification (ECN). See also Figure 2 221 to 222.
[0034] exist Figure 2 At positions 230 to 240, in response to an event-based notification indicating the need for congestion control, the first processing layer 110 may execute a user-defined congestion control algorithm 111 to determine an adjustment of the packet forwarding parameter (P) from a first value (v1) to a second value (v2). For example, using... Figure 1 As explained, the user-defined congestion control algorithm 111 can be one of a plurality of congestion control algorithms that can be programmed to be applied to the first processing layer 110.
[0035] As used herein, the phrase “determines congestion control is needed” at box 230 should be interpreted broadly as including the first processing layer 110 performing the determination based on any suitable information (e.g., metric information, events, instructions, control signals, etc.) at least specified by or derived from event notifications. In the first example (see...) Figures 6 to 7 In the example, based on an event notification indicating that a probe response has been received, the first processing layer 110 can execute block 230 based on metric information determined based on the probe response (i.e., derived from the probe response). In the second example (see...) Figure 8 In this context, based on an event notification indicating that a congestion event has been detected, the first processing layer 110 may execute block 230 based on the congestion event (e.g., a form of instruction / control signal to perform congestion control). Any additional and / or alternative methods may be implemented for block 230.
[0036] exist Figure 2 At position 250, the first processing layer 110 can generate and send to the second processing layer 120 the adjustment to the packet forwarding parameters (see [reference]). Figure 1 The instructions (150) in the second processing layer. As used herein, the term "instruction" can generally refer to an indication that specifies the action to be performed. Any suitable form of instruction may be used, such as invoking an application programming interface (API) call supported by the second processing layer 120.
[0037] exist Figure 2At point 260, the second processing layer 120 can adjust P from a first value (v1) to a second value (v2) based on instructions. In practice, block 260 may involve the interaction between the second processing layer 120 and the hardware layer 130 (see [link to documentation]). Figure 1 The hardware layer 130 (v2) is configured to control packet forwarding to the physical network 104 based on a second value (v2). The hardware layer 130, containing any suitable components, may be referred to as a hardware engine, hardware pipeline, etc. In practice, the “component” may be a hardware scheduler 133 configured to control packet forwarding to the physical network 104. The term “control” generally refers to managing the allocation of resources and / or timing associated with packet forwarding. For example, the hardware scheduler 133 may manage the timing and / or order of packet transmission, organize packets into queues, perform any combination thereof, etc., based on a specific transmission rate and / or congestion window size. Here, the term “configuration” generally refers to sending instructions or control signals to the component. Although the hardware scheduler 133 is used for demonstration, it should be understood that the scheduler may be implemented using hardware, firmware, software, or any combination thereof.
[0038] As used herein, the term "packet forwarding parameters" can generally refer to any suitable settings used to control the process of receiving and / or transmitting packets. In an example (which will use...), Figure 4A and Figure 6 In the description, the user-defined congestion control algorithm 111 can be a rate-based congestion control algorithm, in which case P can be the transmission (TX) rate associated with the hardware scheduler 133. For example, the TX rate can be reduced from v1 to v2 to reduce the amount of data transmitted to the physical network 104. See also Figure 2 241 in the middle.
[0039] In another example (which will be used) Figure 4B and Figure 7 In the description, the user-defined congestion control algorithm 111 can be a window-based congestion control algorithm, in which case P can be the size of the congestion window (“CWND”) associated with the hardware scheduler 133. Here, P = the congestion window size can be reduced to limit the number of incomplete (unacknowledged) packets that can be transmitted to the physical network 104 within a given time period. See also Figure 2 242. Any additional and / or alternative packet forwarding parameters can be adjusted.
[0040] Using examples from this disclosure, the first processing layer 110 can be programmed to execute any suitable user-defined congestion control algorithm 111 tailored to a specific network environment. In this way, the first processing layer 110 can be programmed to determine adjustments to any suitable parameters using any user-defined formula. The second processing layer 120 can be configured as an intermediary between the first processing layer 110 and the hardware layer 130 to perform parameter adjustments, in particular, to support different congestion control algorithms. This flexibility in customizing and adjusting the congestion control algorithm 111 provides several benefits. For example, it allows network administrators to improve network performance based on the specific service patterns and network conditions of the network environment 100. This adaptability also enables responses to changes in network demands and service characteristics over time, thereby enhancing the overall reliability or robustness of the network environment 100.
[0041] Furthermore, using the examples of this disclosure, the first processing layer 110 can be programmed to apply any user-defined state machine logic 112. For example, using... Figure 1 As described, user-defined state machine logic 112 can specify user-defined rules for determining whether to perform an action (e.g., sending a probe packet to measure metric information) based on input (e.g., event notification). The term "metric information" generally refers to any suitable measurable quantity that provides insight into the performance, health, or state of a network. Example metric information may include round-trip time (RTT), latency, throughput, packet loss, jitter, bandwidth utilization, error rate, etc. The term "probe packet" generally refers to a packet sent to measure metric information. (The term will be used...) Figures 3 to 10 Discuss various examples.
[0042] Example programming for the first processing layer
[0043] Figure 3 This is a flowchart of an example detailed process 300 for network interface device 103 to perform congestion control in network environment 100. Example process 300 may include one or more operations, functions, or actions illustrated by one or more blocks (e.g., 310 to 390). Depending on the implementation, various blocks may be combined into fewer blocks, divided into additional blocks, and / or eliminated.
[0044] exist Figure 3At point 310, the first processing layer 110 can be programmed to apply a user-defined congestion control algorithm 111 and a user-defined state machine logic 112. For example, the user-defined congestion control algorithm 111 can determine how to adjust packet forwarding parameters. The user-defined state machine logic 112 can specify one or more rules for determining whether telemetry packets (e.g., probe packets) are needed. The configuration of the first processing layer 110 can be initiated using a second processing layer 120 (e.g., a configuration acquisition / setting module 311 used to interact with the initialization / configuration handler 113 of the first processing layer 110).
[0045] Any suitable programming language can be used to implement the instructions or program code associated with algorithm 111 and / or state machine logic 112. Once generated, one or more firmware images implementing the first processing layer 110 and the second processing layer 120 can be loaded onto the programmable NIC 103. In practice, the term "firmware image" can generally refer to a file (e.g., a binary file) containing the low-level software required to control the hardware of the NIC. The firmware image can provide the processor 131 on the NIC 103 with the necessary instructions to implement the processing layers 110 / 120 according to the examples of this disclosure. Depending on the implementation to be performed, the first processing layer 110 and the second processing layer 120 may use multiple firmware images (see [link to relevant documentation]). Figure 4A Implemented via B) or a single firmware image (not shown).
[0046] Figure 4A The first example is shown in the figure, which is an illustration. Figure 1 A schematic diagram of a first example programming of the first processing layer 110 (see 400) is shown. Here, the first firmware image 410 may contain instructions 411 executable by the processor 131 to implement the first processing layer 110, particularly the first algorithm 111 (labeled "A1") and the first logic 112 (labeled "L1"). The second firmware image 420 may contain instructions 421 for implementing the second processing layer 120 to provide various support functions for the first processing layer 110. During the programming process, the first firmware image 410 and the second firmware image 420 may then be loaded onto the NIC 103 using any suitable firmware update method (see 422).
[0047] Figure 4B The second example is shown in the figure, which is an illustration. Figure 1A schematic diagram of a second example programming of the first processing layer 110 (see 401) is shown. Here, the third firmware image 430 may contain instructions 431 that can be executed by the processor 131 to implement the first processing layer 110, in particular the second algorithm 451 (labeled "A2") and the second logic 452 (labeled "L2"). The fourth firmware image 440 may further contain instructions 441 for implementing the second processing layer 120. Firmware images 430 to 440 may be loaded onto the NIC 103 using any suitable firmware update method (see 442). Note that the instructions 441 in the fourth firmware image 440 may be the same as the instructions 421 in the second firmware image 420. Furthermore, the user-defined state machine logic 112 / 452 may be the same as that of algorithm 111 / 451 (such as...). Figure 4A (As shown in B) Separate software components or parts of algorithm 111 / 451. When a user (e.g., a network administrator) wishes to update user-defined algorithm 111 / 451 and / or logic 112 / 452, the firmware image 410 / 430 associated with the first processing layer 110 can be updated accordingly without modifying the second processing layer 120.
[0048] Depending on the implementation plan, A1 111 can be a rate-based congestion control algorithm, which detects congestion based on metric information and adjusts the parameter = TX rate to control congestion. Any suitable rate-based congestion control algorithm can be used. One example is TIMELY, a congestion control algorithm that relies on RTT information to adjust the TX rate. TIMELY is explained in R. Mittal et al., “TIMELY: RTT-based Congestion Control for the Datacenter” (Proceedings of the 2015 ACM Special Interest Group on Data Communications (SIGCOMM '15), London, UK, 2015, pp. 537-550), which is incorporated herein by reference.
[0049] Compared to A1 111, A2 451 can be a congestion control algorithm based on different rates, a window-based congestion control algorithm, or any other algorithm. For example, SWIFT is a congestion control algorithm that relies on RTT information to adjust the congestion window size, with the goal of maintaining packet delay at approximately the target delay. SWIFT is explained in “Swift: Delay is Simple and Effective for Congestion Control in the Datacenter” by G. Kumar et al. (Proceedings of the ACM Special Interest Group on Data Communications '20, Virtual Event, USA, 2020, pp. 1–15), which is incorporated herein by reference.
[0050] Using a rate-based congestion control algorithm, the TX rate of a source (e.g., computer system 101) can be shaped according to the desired rate. This allows the hardware scheduler 131 to transmit packets at a specific rate (e.g., a constant fixed rate), similar to the leaky bucket method. In contrast, a window-based congestion control algorithm can use a congestion window to limit the number of incomplete (e.g., unacknowledged) packets that the source can transmit within a given time period, which can lead to bursty packet transmissions. Window-based algorithms may require the hardware scheduler 133 to transmit packets based on a specific number of tokens, etc. (The last sentence appears to be incomplete and possibly refers to a different algorithm.) Figures 6 to 7 Let's explain the example congestion control algorithm.
[0051] The first processing layer 110 can be programmed to apply any additional or alternative congestion control algorithms. One example is High Precision Congestion Control Enhanced (HPCC+), an advanced congestion control mechanism designed for high-speed, large-scale networks. It utilizes INT to collect more accurate real-time link load information, enabling precise flow rate adjustment. By leveraging this detailed telemetry data, HPCC+ can quickly converge to optimal bandwidth utilization while avoiding congestion and maintaining near-zero queues in the network, which is crucial for achieving ultra-low latency. This approach allows HPCC+ to provide predictable delivery performance, making it highly effective for applications requiring high throughput and low latency, such as data center networks.
[0052] Using examples from this disclosure, the second processing layer 120 may support different congestion control algorithms (e.g., A1 111 and A2 452) and state machine logic (e.g., L1 112 and L2 452). Examples from this disclosure should be compared to conventional hardware-based methods that rely on hardware logic to perform telemetry and rate adjustment based on static formulas. The parameters used in the static formulas may be configurable, but the actual formulas themselves are typically immutable.
[0053] Event-driven architecture
[0054] The first processing layer 110 and the second processing layer 120 can be configured to implement an event-driven architecture to handle various events related to congestion control. This architecture allows the first processing layer 110 to be decoupled from the underlying hardware layer 130. (See again...) Figure 3 Detectable example events include session events (see 320), service events (see 330), telemetry events (see 350 to 360), congestion events (see 362), etc.
[0055] The event loop module 122 of the second processing layer 120 can be configured to monitor various events. Events can be detected based on hardware interrupts, hardware status registers, firmware events, hardware counter polling, queues, etc. Example hardware counters may include Classification and Forwarding Architecture (CFA) flow counters for managing traffic flows, Remote Direct Memory Access (RDMA) (RoCE) counters via converged Ethernet, etc. In practice, RoCE is a network protocol used in data centers to implement RDMA over Ethernet networks. Depending on the implementation, the second processing layer 120 may send event notifications to the first processing layer 110 via API calls, queuing mechanisms, or a combination thereof. Each event notification may contain timestamp information associated with the detected event. Various events are described below.
[0056] (a) Session events
[0057] exist Figure 3 At point 320, the session event may include session creation and deletion events. Here, the term "session" or "network session" may refer to a connection between two endpoints. For example, in response to detecting that application 102 has created a session, the second processing layer 120 may generate and send an event notification identifying the session creation event to the session event handler 114. Based on the event notification, the first processing layer 110 may store session context information, such as session state, tuple information associated with the session, etc. The tuple information (e.g., source / destination address information, source / destination port number, protocol) can be used for probe packet generation.
[0058] In response to detecting that application 102 has deleted a session, the second processing layer 120 may generate and send an event notification identifying the session deletion event to the session event handler 114. The session event may be stored in a data storage area (not shown) maintained by the second processing layer 120. Based on the event notification, the first processing layer 110 may delete any session context information associated with the session.
[0059] (b) Business Events
[0060] exist Figure 3 At positions 330 to 331, the second processing layer 120 can generate and send an event notification identifying a service event to the first processing layer 110. The term "service event" can generally refer to any suitable event related to packet forwarding. An example service event is the TX event (shown in...). Figure 3 The TX event (not shown) specifies the amount of data (e.g., cumulative byte count) transmitted by the NIC 103 since the last TX event notification. Another example is the ACK RX event (not shown), which specifies the amount of data (e.g., in bytes) acknowledged by the receiver since the last event notification.
[0061] Depending on the implementation, the second processing layer 120 may monitor the hardware layer 130 (including the TX queue 134) to determine whether the amount of data transmitted / acknowledged exceeds a configurable threshold. If so (i.e., exceeds the threshold), the event loop module 122 of the second processing layer 120 may generate and send a service event notification to the first processing layer 110. Any additional and / or alternative service events may be monitored.
[0062] (c) Telemetry events
[0063] Will use Figure 5 Example description boxes 330 to 350 in the figure illustrate a schematic diagram of the first processing layer 110 applying user-defined rules to determine whether to send a probe packet. Figure 5 At point 510, the first processing layer 110 can receive an event notification that identifies a TX event (described above) associated with the session. In response, user-defined state machine logic 112 can be applied to determine whether the session requires a probe packet. See also Figure 3 340 in the middle.
[0064] Figure 5 The following section demonstrates some example rules. Note that one or more rules can be applied. In the first example, user-defined state machine logic 112 extracts a certain amount of cumulative byte count from TX event notification 510 and applies the first rule (see [link]). Figure 5 The "R1" in the table determines whether the cumulative byte count is greater than or equal to a first threshold (T1). If it is (i.e., R1 is satisfied), then a probe packet is required. Otherwise, no probe packet is required.
[0065] In the second example, user-defined state machine logic 112 may extract timestamp information associated with the TX event notification 510 and apply a second rule (see “R2”) to determine whether the timestamp information is greater than or equal to a second threshold (T2). For example, T2 may be a user-defined threshold specifying the time elapsed since the last probe packet was sent. If so (i.e., R2 is satisfied), then a probe packet is determined to be needed. Otherwise, a probe packet is not needed.
[0066] In the third example, user-defined state machine logic 112 can determine the number of TX events received within a pre-configured time period based on TX event notification 510 and other prior notifications. In this case, user-defined state machine logic 112 can apply a third rule (see “R3”) to determine whether the number of TX events is greater than or equal to a third threshold (T3). If so (i.e., R3 is satisfied), then a probe packet is required. Otherwise, no probe packet is required. Any additional and / or alternative rules can be defined and applied.
[0067] exist Figure 5 At point 530, in response to determining that one or more rules are met, user-defined state machine logic 112 can determine that a probe packet is needed. In this case, user-defined state machine logic 112 can generate and send a request to the second processing layer 120 to request that a probe packet be sent to the destination associated with the session (i.e., the probe packet responder). Figure 5 (The label is "Request: Probe"). Otherwise, no probe packet is requested. Requests can be sent using any suitable method, such as the first processing layer 110 invoking an API call that causes the second processing layer 120 to generate and send a probe packet.
[0068] exist Figure 5 At position 540, based on a request, the telemetry module 121 of the second processing layer 120 can generate and send a probe packet to the destination. The probe packet 540 can be placed in the TX queue 134 before being forwarded to the physical network 104. Figure 5 In the example, probe packet 540 can be sent to measure metrics such as RTT. In practice, RTT typically refers to the time (in milliseconds) it takes for a packet to travel from a source (e.g., first computer system 101) to a destination (e.g., second computer system 105 / 106) and back, providing insights into network latency and performance.
[0069] exist Figure 5 At point 550, once probe packet 540 has been sent, telemetry module 121 can generate an event notification to first processing layer 110 to report that probe packet 540 was sent at TX time = t1. TX time can be reported immediately after probe packet 540 is sent to the line to exclude any waiting time that may be included in RTT calculation at TX queue 134.
[0070] First example: Rate-based congestion control algorithm
[0071] Will use Figure 6 explain Figure 3 The box in the middle is between 360 and 390. Figure 6This is a schematic diagram illustrating a first example (see 700) of a user-defined congestion control algorithm 111 that can be programmable to be applied in the first processing layer 110. In this example, A1 111 may be a rate-based congestion control algorithm.
[0072] exist Figure 6 At positions 610 to 620, telemetry module 121 can receive probe responses via RX queue 135. Any suitable format can be used for probe packet 540 and probe response 610, such as Management Datagram (MAD) format, In-Band Flow Analyzer (IFA) format, etc. For example, MAD is defined by the InfiniBand architecture, a high-performance networking standard for high-performance computing (HPC) environments, data centers, and enterprise networks. In practice, MAD-based network probe packets (e.g., 256-byte messages) can be used to collect metrics about physical network 104 by exchanging probe packets between an initiator (e.g., computer system 101) and a responder (e.g., computer systems 105 / 106). In another example, IFA allows for the insertion and collection of predefined and customized telemetry information (i.e., metadata) on a hop-by-hop basis. Metadata containing timestamp information inserted by the responder can be used for RTT calculation.
[0073] exist Figure 6 At point 630, in response to receiving a probe response via RX queue 135, telemetry module 121 can generate and send an event notification (labeled "Event: PROBE_RES") to first processing layer 110. Event notification 630 can indicate that a telemetry event has occurred, specifically the receipt of probe response 610. Furthermore, event notification 630 can specify (t2, t3, t4) for first processing layer 110 to perform RTT calculation. Here, t2 = RX time of probe packet 540 at the responder, t3 = TX time of probe response 610 at the responder, and t4 = RX time of probe response 610 at the initiator. See also Figure 3 360 and 361 (are).
[0074] exist Figure 6 At points 640 to 641, in response to determining that congestion control is needed, the first processing layer 110 may apply A1 111 to determine an adjustment to the parameter = TX rate associated with the TX queue 134. For example (see 640), A1 111 may be performed to determine the metric information = RTT based on (t1, t2, t3, t4) discussed above, for example, by applying the formula RTT = (t4 - t1) – (t3 - t2), etc. Additionally (see 641), the first processing layer 110 may determine whether congestion control is needed, for example, by comparing the calculated RTT or derived value with a user-defined threshold.
[0075] For example, the TIMELY algorithm (discussed above) can be applied to monitor RTT to infer network congestion levels. This is in response to determining that the RTT is less than a user-defined low threshold (T). low This can be adjusted to increase the TX rate. In response to determining that the RTT is greater than a user-defined high threshold (T...),... high Furthermore, congestion control is required, and adjustments can be calculated to reduce the TX rate. Additionally, a delay gradient value representing the derivative of the queue with respect to time can be calculated based on the current RTT and previous RTT calculations. In response to determining that the delay gradient value is less than or equal to 0 (i.e., a negative gradient value indicates a positive decrease in RTT), adjustments can be calculated to increase the TX rate to utilize available bandwidth more efficiently. Otherwise (i.e., a positive gradient value indicates a positive increase in RTT and congestion control is required), adjustments can be calculated to decrease the TX rate to reduce the load on physical network 104. Any suitable user-defined formula can be used to calculate a specific adjustment from the first value (v1) to the second value (v2). See also Figure 3 370 to 371.
[0076] exist Figure 6 At position 650, the first processing layer 110 can generate and send an instruction (labeled "INSTR") to the second processing layer 120, causing the parameter adjustment engine 123 to perform an adjustment. Figure 6 At positions 660 to 670, the parameter adjustment engine 123 configures the hardware layer 130, specifically the hardware scheduler 133, to update the TX rate from v1 to v2. Depending on the implementation, this configuration can be performed using any suitable hardware-readable instructions, control signals, etc. See also Figure 3 380 to 390.
[0077] Second example: Window-based congestion control algorithm
[0078] Figure 7 This is a schematic diagram illustrating a second example (see 700) of a user-defined congestion control algorithm 451 that is programmable to be applied in the first processing layer 110. (The following text will be used...) Figure 4B Explanation of the second congestion control algorithm (A2) 451 and the second state machine logic (L2) 452. Figure 7 In this example, A2 451 could be a window-based congestion control algorithm to adjust packet forwarding parameters that are the size of the congestion window (CWND).
[0079] exist Figure 7In steps 710 to 730, in response to receiving a probe response via RX queue 135, telemetry module 121 can generate and send an event notification (labeled "Event: PROBE_RES") to first processing layer 110. Event notification 730 can indicate that probe response 710 has been received. Event notification 730 can specify (t2, t3, t4) for first processing layer 110 to perform RTT calculations. Similar to... Figure 6 In the example, t2 = RX time of probe packet 540 at the responder, t3 = TX time of probe response 710 at the responder, and t4 = RX time of probe response 710 at the initiator. See also Figure 3 360 and 361 (are).
[0080] exist Figure 7 At points 740 to 741, in response to determining that congestion control is needed, the first processing layer 110 may apply A2 451 to determine the adjustment of parameter = CWND. For example (see 740), the metric information = RTT can be calculated based on (t1, t2, t3, t4), such as RTT = (t4 - t1) - (t3 - t2), etc. Additionally (see 741), based on RTT, the first processing layer 110 may determine whether congestion control is needed, for example, by comparing the RTT or derived value with a user-defined threshold.
[0081] For example, the SWIFT algorithm (discussed above) can be applied to detect congestion by monitoring RTT. In response to determining that the measured RTT is greater than a threshold (i.e., the target RTT), the first processing layer 110 can apply the SWIFT algorithm to determine that congestion control is needed. In this case, an adjustment to reduce CWND from a first value (v1) to a second value (v2), for example, by multiplication, can be determined. Conversely, in response to determining that the measured RTT is less than or equal to a threshold (i.e., the target RTT), the first processing layer 110 can apply the SWIFT algorithm to determine that congestion control is not needed. In this case, an adjustment to increase CWND additively can be determined. This method is called Additive Increase Multiplicative Decrease (AIMD).
[0082] exist Figure 7 At position 750, the first processing layer 110 can generate and send an instruction (labeled "INSTR") to the second processing layer 120, causing the parameter adjustment engine 123 to perform an adjustment. Figure 7 At positions 760 to 770, parameter tuning engine 123 can configure hardware layer 130, specifically hardware scheduler 133, to update CWND from v1 to v2. For example, v2 < v1 is used as a congestion control measure to reduce the amount of data forwarded to physical network 104. See also Figure 3 380 to 390.
[0083] Examples of this disclosure can utilize the ability of hardware layer 130 to send and receive probe packets from processor 131 (e.g., eCPU) and its ability to adjust packet forwarding parameters. As a software solution, the logic and mathematical calculations behind a particular congestion control algorithm can be variable. This allows the congestion control algorithm to respond to congestion by adjusting calculations using different telemetry schemes and different parameters.
[0084] Congestion control granularity
[0085] Depending on the implementation plan, any suitable congestion control granularity can be implemented using congestion control algorithm 111 / 451, such as by destination, by QP (queue pair), by path, etc. A queue pair contains a send queue and a receive queue to manage communication between two endpoints in a network session. By destination congestion control granularity refers to the ability to manage congestion for one or more QPs heading towards the same destination (e.g., a destination Internet Protocol (IP) address). By QP congestion control granularity refers to the ability to manage congestion individually for each QP (regardless of the destination). In this mode, a session can be created when each QP is created.
[0086] Path-based congestion control granularity refers to the ability to manage congestion individually for each QP (regardless of destination). Regarding scale and granularity, path-based congestion control granularity is a compromise between destination-based and QP-based configuration. With path-based granularity, session creation is based not only on the destination IP address but also on tuple information associated with the path. In practice, path-associated tuple information may include the source IP address, destination IP address, source port number, destination port number, and protocol information. In this mode, if multiple paths exist from the source node to the destination node, each path can be probed independently, and the rate can be adjusted.
[0087] Third example: Congestion event
[0088] Refer again Figure 3 At 362, the event loop module 122 of the second processing layer 120 can generate and send event notifications related to congestion events to the first processing layer 110. Here, the term "congestion event" can generally refer to any event triggered by the detection of congestion (e.g., based on any suitable congestion condition or performance problem). For example, when a session experiences congestion, the event loop module 122 can generate one of the following congestion events: RTO event, sequence error NAK event, CNP event, etc.
[0089] Figure 8The diagram illustrates examples of user-defined congestion control algorithms 801 (see 800) that can be programmably applied to the first processing layer 110. In practice, an RTO event associated with a session can be generated in response to the event loop module 122 determining that an RTO condition has been met. This can occur whenever a packet destined for the destination is dropped, in which case an ACK response is expected but not received before a timeout is triggered. RTO conditions can be configured at any suitable granularity, such as for QP connections. See [link to documentation]. Figure 8 830 in the middle.
[0090] A sequence error negative acknowledgment (NAK) event associated with the session can be generated in response to event loop module 122 determining that an out-of-order condition has been met. This can occur at any time when a packet destined for the destination is dropped, in which case the received packet is out of sequence (e.g., the sequence number does not match the expected number). This causes the destination to send an error-indicating packet to the source (i.e., computer system 101). The sequence error NAK event can be detected for any QP belonging to the session. See also Figure 8 810 to 830.
[0091] A CNP event can be generated in response to event loop module 122 detecting that a CNP packet has been received. A CNP event indicates that the path connecting computer system 101 and the destination is experiencing congestion. This can occur when intermediate network devices (e.g., switches) along the path have tagged some packets with an ECN field. When an ECN-tagged packet is received, the destination can use a CNP packet to indicate congestion to respond to computer system 101. CNP events can be detected for any QP belonging to a session. See also Figure 8 810 to 830.
[0092] exist Figure 3 At point 371, in response to receiving an event notification indicating a congestion event, the first processing layer 110 can determine an adjustment to the packet forwarding parameters from a first value (v1) to a second value (v2). Using the TX rate as an example, rate reduction can be performed based on RTO events and / or sequence error NAK events. Rate reduction can also be performed based on CNP events (e.g., when a CNP packet is received after sending probe packet 540). Any suitable user-defined formula can be used to calculate the adjustment. See also Figure 8 840 to 870.
[0093] Example network environment for AI applications
[0094] The examples disclosed herein can be implemented in any suitable network environment, such as a network environment used to support any AI application. One example is shown below. Figure 9The figure described herein is a schematic diagram illustrating an example distributed training environment 900 in which a network interface device 103 can be deployed to perform congestion control. Here, the term "distributed training environment" generally refers to a network environment in which the workload associated with training a model can be distributed across multiple worker nodes. In practice, distributed training can be performed to improve speed (i.e., training time), scalability (e.g., easier handling of large datasets and complex models), and efficiency (e.g., better utilization of computational resources) during training. Although Figure 9 Not shown in the document, but the examples disclosed herein can be implemented to support inference using AI models.
[0095] exist Figure 9 In the example, a cluster of multiple (N) worker nodes 911 to 91N can be deployed in a distributed training environment 900 to perform distributed training. For example, a first worker node 911 running on computer system 101 can be configured to train model 921 based on dataset 931. A second worker node 912 can be configured to train model 922 based on dataset 932. Similarly, the Nth worker node 91N can be configured to train model 92N based on dataset 93N. As used herein, the term "worker node" generally refers to a computing resource capable of performing tasks related to model training. In practice, worker nodes 911 to 91N may be equipped with one or more accelerators to accelerate the computation of training tasks, such as graphics processing units (GPUs), tensor processing units (TPUs), etc. "Worker node" may also be referred to as "computing node," "training node," "processing node," "computing resource," "GPU node" (where a GPU is provided), etc. In another example, training can be performed by any suitable software and / or hardware component of computer system 101.
[0096] In practice, the distributed training environment 900 can implement any suitable parallelism strategy to extend training across multiple worker nodes, such as data parallelism, model parallelism, or a combination of both (i.e., hybrid parallelism). For example, using data parallelism, worker nodes 911 to 91N can each train copies or duplicates of the same model (see 921 to 92N) using different datasets 931 to 93N. In this way, large datasets can be divided into small chunks 931 to 93N, allowing each chunk to be processed independently by different worker nodes. In another example, using model parallelism, the model can be split into multiple parts (also 921 to 92N), each of which is trained using a different worker node. This is particularly useful when the model is too large to fit into the memory of a single node. Using hybrid parallelism, a combination of data parallelism and model parallelism can be implemented to take advantage of both.
[0097] The term "model" generally refers to a mathematical representation or algorithm that can be trained in a distributed training environment to make predictions or decisions based on input data. Figure 9 In the examples, AI models (see 921 to 92N) can be trained in a distributed manner, such as machine learning (ML) models, deep learning models, etc. Broadly speaking, deep learning is a subset of machine learning, where multiple layers of neural networks are used for feature extraction and pattern analysis and / or classification. The term "depth" in deep learning typically refers to the number of layers in the neural network. For example, a deep learning model can have dozens or even hundreds of layers compared to a shallow learning model. This allows deep learning models to extract more complex and nuanced features from the input data, resulting in more accurate output data. Although described using AI models, it should be understood that non-AI models, such as linear regression models, decision trees, random forests, etc., can be trained.
[0098] During training, worker nodes 911 / 912 / 91N process dataset 931 / 932 / 93N to generate model information associated with model 921 / 922 / 92N. Here, the term "model information" can generally refer to any suitable information generated by the worker nodes during model training. For example, model information may include gradient coordinate values (also called "gradients" and "gradient vectors") or parameters associated with model 921 / 922 / 92N. In practice, a gradient can represent the direction and rate of change of the model's parameters (e.g., weights) relative to a loss function. Thus, the gradient indicates the degree of deviation of the model's predictions from actual values, guiding the learning process to minimize error. Using data parallelism, each worker node 911 / 912 / 91N can compute model information based on its dataset 931 / 932 / 93N (e.g., one or more chunks of a larger dataset).
[0099] In practice, distributed training of AI models requires the transmission of large amounts of data via physical network 104, for example, for data synchronization during training. The examples disclosed herein can be implemented to facilitate the customization of congestion control for the unique needs and characteristics (e.g., low latency, high bandwidth, etc.) of supporting distributed training environments 900. Figure 9 In the example, computer system 101 may include network interface device 103 for performing congestion control on a first worker node 911. Congestion control algorithms 111 / 451 and state machine logic 112 / 452 may be defined to support more efficient data transfer between the first worker node 911 and another worker node 912 / 91N, thereby reducing training time and improving overall system performance.
[0100] Software-defined networking (SDN) environment
[0101] Depending on the implementation plan, computer system 101 may be a host deployed in a software-defined networking (SDN) environment (such as a public or private cloud environment). It will use... Figure 10 The diagram illustrates an example SDN environment 1000 in which congestion control can be implemented. In this example, the SDN environment 1000 may contain any suitable number of hosts, such as host-A 1010A and host-B 1010B. In practice, Figure 9 Worker nodes 911 to 912 can be implemented using virtualized computing instances in the form of virtual machines (VMs), containers, etc.
[0102] Hosts 1010A / 1010B may include suitable hardware 1012A / 1012B and virtualization software (e.g., monitor-A 1014A, monitor-B 1014B) to support various VMs. For example, host-A 1010A may support VM1 1031 and VM2 1032, while VM3 1033 and VM4 1034 are supported by host-B 1010B. Hardware 1012A / 1012B includes suitable physical components such as a central processing unit (CPU) or processor 1020A / 1020B, memory 1022A / 1022B, physical network interface controller (PNIC) 1024A / 1024B, storage disk 1026A / 1026B, GPU 1028A / 1028B, etc.
[0103] The monitoring programs 1014A / 1014B maintain the mapping between the underlying hardware 1012A / 1012B and the virtual resources allocated to the corresponding VMs. Virtual resources are allocated to the corresponding VMs 1031 through 1034 to support the guest operating system and applications; see 1041 through 1044, 1051 through 1054. For example, virtual resources may include virtual CPUs, guest physical memory, virtual disks, virtual network interface controllers (VNICs), etc. Hardware resources can be emulated using a virtual machine monitor (VMM). For example, in... Figure 10 In this diagram, VNICs 1061 to 1064 are virtual network adapters for VMs 1031 to 1034, and are emulated by corresponding VMMs (not shown) instantiated by their respective monitors on the corresponding hosts-A 1010A and B 1010B. A VMM can be considered part of the corresponding VM, or alternatively, separate from the VM. Although a one-to-one relationship is shown, a VM can be associated with multiple VNICs (each with its own network address).
[0104] Although the examples in this disclosure involve VMs, it should be understood that a “virtual machine” running on a host is only one example of a “virtualized computing instance” or “workload.” A virtualized computing instance may represent an addressable data compute node (DCN) or an isolated user-space instance. In practice, any suitable technology may be used to provide isolated user-space instances, not just hardware virtualization. Other virtualized computing instances may include containers (e.g., running within a VM or on top of a host operating system without requiring a monitor or separate operating system, or implemented as operating system-level virtualization), virtual private servers, client computers, etc. This container technology is available from Docker Inc. and other companies. A VM may also be a complete computing environment containing a virtual equivalent of the hardware and software components of a physical computing system. Depending on the implementation, the examples in this disclosure may also utilize any suitable serverless computing technology. One example is Function as a Service (FaaS), which allows developers to execute code (e.g., in response to events) without having to manage the underlying cloud infrastructure. Another example is serverless GPU (also known as Accelerator as a Service), which allows developers to access powerful GPU resources for their applications.
[0105] The term "monitor" typically refers to a software layer or component that supports the execution of multiple virtualized compute instances, including system-level software within the guest VM that supports namespace containers (such as Docker). Monitors 1014A through B can each implement any suitable virtualization technology, such as VMware ESX® or ESXi. TM (Available from VMware LLC), kernel-based virtual machines (KVM), etc. The term "packet" can generally refer to a group of bits that can be transmitted together, and can also take other forms such as "frame," "message," "segment," etc. The terms "service" or "flow" can generally refer to multiple packets. In the Open Systems Interconnection (OSI) model, the term "Layer-2" can generally refer to the data link layer or the Media Access Control (MAC) layer; "Layer-3" can generally refer to the network or Internet Protocol (IP) layer; and "Layer-4" can generally refer to the transport layer (e.g., using Transmission Control Protocol (TCP), User Datagram Protocol (UDP), etc.), although the concepts described herein can be used with other network models.
[0106] SDN controller 1070 and SDN manager 1072 are example network management entities in SDN environment 100. An example of an SDN controller is the VMware NSX® (available from VMware LLC) NSX controller component that operates on a central control plane. SDN controller 1070 can be a component of a controller cluster (not shown for simplicity) that can be configured using SDN manager 1072. Network management entities 1070 / 1072 can be implemented using physical machines, VMs, or both. To send or receive control information, a local control plane (LCP) agent (not shown) on hosts 1010A / 1010B can interact with SDN controller 1070 via control plane channels 1001 / 1002.
[0107] By virtualizing network services in the SDN environment 100, logical networks (also known as overlay networks or logical overlay networks) can be deployed, modified, stored, deleted, and restored programmatically without requiring reconfiguration of the underlying physical hardware architecture. Monitoring programs 1014A / 1014B implement virtual switches 1015A / 1015B and logical distributed router (DR) instances 1017A / 1017B to handle outgoing packets from VMs 1031 to 1034 and incoming packets destined for those VMs. In the SDN environment 100, logical switches and logical DRs can be implemented in a distributed manner and can span multiple hosts.
[0108] For example, logical switches (LS) can be deployed to provide logical layer-10 connectivity (i.e., overlay network) to VMs 1031 through 1034. The logical switches can be implemented jointly by virtual switches 1015A through B, and represented internally at each virtual switch 1015A through B using forwarding tables 1016A through B. Forwarding tables 1016A through B can each contain entries for the respective logical switches that are being implemented. Furthermore, logical DRs providing logical layer-3 connectivity can be implemented jointly by DR instances 1017A through B, and represented internally at each DR instance 1017A through B using routing tables (not shown). Each routing table can contain entries for the respective logical DRs that are being implemented.
[0109] Packets can be received from or sent to each VM via their associated logical ports. For example, logical switch ports 1065 through 1068 (labeled "LSP1" through "LSP4") are associated with corresponding VMs 1031 through 1034. Here, the terms "logical port" or "logical switch port" generally refer to the port on the logical switch to which the virtualized compute instance is connected. "Logical switch" generally refers to the software-defined networking (SDN) architecture implemented by virtual switches 1015A through B, while "virtual switch" generally refers to a software switch or software implementation of a physical switch. In practice, there is typically a one-to-one mapping between logical ports on a logical switch and virtual ports on virtual switches 1015A / 1015B. However, in some scenarios, this mapping can change, for example, when logical ports are mapped to different virtual ports on different virtual switches after the migration of the corresponding virtualized compute instance (e.g., when the source and destination hosts do not have a distributed virtual switch spanning them).
[0110] Logical overlay networks can be formed using any suitable tunneling protocol, such as Virtual Extensible Local Area Network (VXLAN), Stateless Transport Tunneling (STT), Geneve Network Virtualization Encapsulation (GENEVE), Geneve Routing Encapsulation (GRE), etc. For example, VXLAN is a Layer-2 overlay scheme on a Layer-3 network, using tunnel encapsulation to allow Layer-2 segments to span multiple host extensions that can reside on different physical networks. Monitors 1014A / 1014B can implement Virtual Tunnel Endpoints (VTEPs) 1019A / 1019B to encapsulate and decapsulate packets with an outer header (also called a tunnel header) that identifies the relevant logical overlay network (e.g., VNI). Hosts 1010A to B can maintain data plane connectivity with each other via physical network 1005 to facilitate east-west communication among VMs 1031 to 1034.
[0111] Computer System
[0112] The above examples can be implemented by hardware (including hardware logic circuitry), software or firmware, or a combination thereof. The above examples can be implemented by any suitable computing device, computer system, etc. A computer system may include a processor, memory units, and a physical NIC that can communicate with each other via a communication bus, etc. A computer system may include a non-transitory computer-readable medium on which instructions or program code are stored, which, when executed by a processor, cause the processor to perform the processes described herein with reference to the figures.
[0113] The technologies introduced above can be implemented in dedicated hardwired circuit systems, software and / or firmware combined programmable circuit systems, or any combination thereof. Dedicated hardwired circuit systems can take the form of, for example, one or more ASICs, PLDs, FPGAs, etc. The term 'processor' should be interpreted broadly to include processing units, ASICs, logic units, or programmable gate arrays, etc.
[0114] The foregoing embodiments have been described by using block diagrams, flowcharts, and / or examples to illustrate various embodiments of the apparatus and / or processes. As long as these block diagrams, flowcharts, and / or examples contain one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within these block diagrams, flowcharts, or examples can be implemented individually and / or collectively by a wide range of hardware, software, firmware, or any combination thereof.
[0115] Those skilled in the art will recognize that some aspects of the embodiments disclosed herein can be implemented, in whole or in part, as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computing systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as virtually any combination thereof, and that designing circuit systems and / or writing code for software and / or firmware in accordance with this disclosure will be well within the skill of those skilled in the art.
[0116] Software and / or firmware used to implement the techniques introduced herein may be stored on a non-transitory computer-readable storage medium and may be executed by one or more general-purpose or special-purpose programmable microprocessors. As used herein, "computer-readable storage medium" includes any means of providing (i.e., storing and / or transmitting) information in a form accessible by a machine (e.g., a computer, network device, personal digital assistant (PDA), mobile device, manufacturing tool, any device having one or more processors, etc.). Computer-readable storage medium includes recordable / non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk or optical storage media, flash memory devices, etc.).
[0117] The diagrams are merely illustrative examples, and the units or processes shown in the diagrams are not necessarily essential for implementing this disclosure. Those skilled in the art will understand that the units in the illustrated apparatus may be arranged as described in the illustrated apparatus, or alternatively located in one or more apparatuses different from those in the illustrated apparatus. The units in the described examples may be combined into a module or further divided into multiple sub-units.
Claims
1. A network interface device, comprising: processor; and A non-transitory computer-readable medium having instructions stored thereon, which, when executed by the processor, cause the processor to implement a first processing layer and a second processing layer to perform the following operations: The first processing layer receives the event notification from the second processing layer; In response to the first processing layer determining that congestion control is needed based on the event notification, The first processing layer determines the adjustment of the packet forwarding parameters from a first value to a second value by applying a congestion control algorithm, wherein the congestion control algorithm is one of a plurality of congestion control algorithms that the first processing layer can programmably apply; The instruction to perform the adjustment is generated by the first processing layer and sent to the second processing layer; and Based on the instructions, the second processing layer configures the components of the network interface device to control packet forwarding to the physical network based on the second value of the packet forwarding parameters.
2. The network interface apparatus of claim 1, wherein the instruction for determining that congestion control is required causes the processor to: The first processing layer determines the metric information associated with packet forwarding based on the event notification indicating that the second processing layer has received a probe response; and The first processing layer determines that congestion control is needed based on the metric information.
3. The network interface device according to claim 2, wherein the instructions further cause the processor to: Before receiving the event notification, the first processing layer determines whether a probe packet is needed based on one or more user-defined rules that can be programmably applied by the first processing layer; and In response to determining that a probe packet is needed, the first processing layer generates and sends a request to the second processing layer to send the probe packet to a receiver capable of sending the probe response.
4. The network interface apparatus of claim 1, wherein the instruction for determining the adjustment causes the processor to: The adjustment of the packet forwarding parameters in the form of transmission rate is determined by the first processing layer, wherein the congestion control algorithm is a rate-based congestion control algorithm.
5. The network interface apparatus of claim 1, wherein the instruction for determining the adjustment causes the processor to: The adjustment of the packet forwarding parameters, which is in the form of a congestion window size, is determined by the first processing layer, wherein the congestion control algorithm is a window-based congestion control algorithm.
6. The network interface apparatus of claim 1, wherein the instruction for determining that congestion control is required causes the processor to: The first processing layer determines that congestion control is needed based on the event notification indicating that the second processing layer has detected at least one of the following events: Retransmission Timeout (RTO) event, Sequence Error Negative Acknowledgment (NAK) event, and Congestion Notification Point (CNP) event.
7. The network interface apparatus of claim 1, wherein the instructions for configuring the component cause the processor to: The component is configured as a hardware scheduler by the second processing layer, which acts as an intermediary between the first processing layer and the hardware scheduler.
8. A non-transitory computer-readable storage medium comprising a set of instructions, the set of instructions being responsive to execution by a processor of a network interface device to cause the processor to implement a first processing layer and a second processing layer to perform a congestion control method, wherein the method includes: The first processing layer receives the event notification from the second processing layer; In response to the first processing layer determining that congestion control is needed based on the event notification, The first processing layer determines the adjustment of the packet forwarding parameters from a first value to a second value by applying a congestion control algorithm, wherein the congestion control algorithm is one of a plurality of congestion control algorithms that the first processing layer can programmably apply; The instruction to perform the adjustment is generated by the first processing layer and sent to the second processing layer; and Based on the instructions, the second processing layer configures the components of the network interface device to control packet forwarding to the physical network based on the second value of the packet forwarding parameters.
9. The non-transitory computer-readable storage medium of claim 8, wherein determining the need for congestion control comprises: The first processing layer determines the metric information associated with packet forwarding based on the event notification indicating that the second processing layer has received a probe response; and The first processing layer determines that congestion control is needed based on the metric information.
10. The non-transitory computer-readable storage medium of claim 9, wherein the method further comprises: Before receiving the event notification, the first processing layer determines whether a probe packet is needed based on one or more user-defined rules that can be programmably applied by the first processing layer; and In response to determining that a probe packet is needed, the first processing layer generates and sends a request to the second processing layer to send the probe packet to a receiver capable of sending the probe response.
11. The non-transitory computer-readable storage medium of claim 8, wherein determining the adjustment comprises: The adjustment of the packet forwarding parameters in the form of transmission rate is determined by the first processing layer, wherein the congestion control algorithm is a rate-based congestion control algorithm.
12. The non-transitory computer-readable storage medium of claim 8, wherein determining the adjustment comprises: The adjustment of the packet forwarding parameters, which is in the form of a congestion window size, is determined by the first processing layer, wherein the congestion control algorithm is a window-based congestion control algorithm.
13. The non-transitory computer-readable storage medium of claim 8, wherein determining the need for congestion control comprises: The first processing layer determines that congestion control is needed based on the event notification indicating that the second processing layer has detected at least one of the following events: Retransmission Timeout (RTO) event, Sequence Error Negative Acknowledgment (NAK) event, and Congestion Notification Point (CNP) event.
14. The non-transitory computer-readable storage medium of claim 8, wherein configuring the component comprises: The component is configured as a hardware scheduler by the second processing layer, which acts as an intermediary between the first processing layer and the hardware scheduler.
15. A method for performing congestion control on a network interface device, wherein the network interface device includes a first processing layer and a second processing layer, and the method includes: The first processing layer receives the event notification from the second processing layer; In response to the first processing layer determining that congestion control is needed based on the event notification, The first processing layer determines the adjustment of the packet forwarding parameters from a first value to a second value by applying a congestion control algorithm, wherein the congestion control algorithm is one of a plurality of congestion control algorithms that the first processing layer can programmably apply; The instruction to perform the adjustment is generated by the first processing layer and sent to the second processing layer; and Based on the instructions, the second processing layer configures the components of the network interface device to control packet forwarding to the physical network based on the second value of the packet forwarding parameters.
16. The method of claim 15, wherein determining the need for congestion control comprises: The first processing layer determines the metric information associated with packet forwarding based on the event notification indicating that the second processing layer has received a probe response; and The first processing layer determines that congestion control is needed based on the metric information.
17. The method of claim 16, wherein the method further comprises: Before receiving the event notification, the first processing layer determines whether a probe packet is needed based on one or more user-defined rules that can be programmably applied by the first processing layer; and In response to determining that a probe packet is needed, the first processing layer generates and sends a request to the second processing layer to send the probe packet to a receiver capable of sending the probe response.
18. The method of claim 15, wherein determining the adjustment comprises one of the following: The adjustment of the packet forwarding parameters, in the form of transmission rate, is determined by the first processing layer, wherein the congestion control algorithm is a rate-based congestion control algorithm; and The adjustment of the packet forwarding parameters, which is in the form of a congestion window size, is determined by the first processing layer, wherein the congestion control algorithm is a window-based congestion control algorithm.
19. The method of claim 15, wherein determining the need for congestion control comprises: The first processing layer determines that congestion control is needed based on the event notification indicating that the second processing layer has detected at least one of the following events: Retransmission Timeout (RTO) event, Sequence Error Negative Acknowledgment (NAK) event, and Congestion Notification Point (CNP) event.
20. The method of claim 15, wherein configuring the component comprises: The component is configured as a hardware scheduler by the second processing layer, which acts as an intermediary between the first processing layer and the hardware scheduler.