Packet buffering techniques

By using SoC circuits in the switch to select the memory for storing packets and using overflow queue buffering packets, the problems of packet loss and delay increase in the data center network are solved, and the stability and throughput of the network are improved.

CN119996359APending Publication Date: 2025-05-13INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411365546.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-09-29
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

There may be problems with packet loss and increased latency in data center networks, mainly due to switch buffer overflow caused by burst traffic and incoming traffic, and some CC protocols based on orderly packet delivery, and the receiver may discard out-of-order packets.

Method used

Using circuits in a switch system-on-chip (SoC) to select between the first memory and a plurality of second memory devices to store the packets based on the reception of packets and the level of the first queue. The circuit stores packets in selected memory and uses overflow queues to buffer packets, reducing packet loss and delay.

Benefits of technology

By reducing packet loss and buffer overflow, the possibility of increased delay at the switch or network interface device is reduced, and network stability and throughput are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996359A_ABST
    Figure CN119996359A_ABST
Patent Text Reader

Abstract

The invention relates to packet buffering techniques. Examples described herein relate to a switch. In some examples, the switch includes circuitry configured to make a selection between a first memory and a second memory device among a plurality of second memory devices to store a packet based on reception of the packet and a level of a first queue, store the packet in the first memory based on the selection of the first memory, and store the packet in the second memory based on the level of the first queue. And store the packet into a selected second memory device based on a selection of the second memory device among the plurality of second memory devices. In some examples, the packet is associated with an ingress port and an egress port, and the selected second memory device is associated with a third port that is different from the ingress port and the egress port associated with the packet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to packet buffering techniques. Background Art

[0002] Core characteristics of data center networks include high throughput, low latency, and network stability. However, even with fine-tuned congestion control (CC) protocols, packet loss may still occur in the network due to bursty traffic and / or large amounts of in-cast traffic causing switch buffers to overflow. In-cast traffic can be observed at the Top-of-Rack (ToR) switches. In addition, some CC protocols are based on in-order packet delivery, and out-of-order packets may be discarded by the receiver. Such discards may result in increased packet reception latency due to retransmissions. Summary of the invention

[0003] According to an embodiment of the present disclosure, a device is provided, comprising: a switch system on chip (SoC), comprising a circuit, wherein the circuit is used to: based on reception of a packet and a level of a first queue, make a selection between a first memory and a second memory device among a plurality of second memory devices to store the packet, based on the selection of the first memory, store the packet in the first memory, and based on the selection of the second memory device among the plurality of second memory devices, store the packet in a selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the selected second memory device is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

[0004] According to an embodiment of the present disclosure, a method is provided, comprising: in a switch: based on reception of a packet and a level of a first queue, making a selection between a first memory and a second memory device among a plurality of second memory devices to store the packet, based on the selection of the first memory, storing the packet in the first queue in the first memory, and based on the selection of the second memory device among the plurality of second memory devices, storing the packet in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

[0005] According to an embodiment of the present disclosure, at least one non-transitory computer-readable medium is provided, the medium including instructions stored thereon, and if the instructions are executed by one or more circuits, the one or more circuits are used in the switch to: based on reception of a packet and a level of a first queue, select between a first memory and a second memory device among a plurality of second memory devices to store the packet; based on the selection of the first memory, store the packet in the first queue in the first memory; and based on the selection of the second memory device among the plurality of second memory devices, store the packet in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 An example system is shown.

[0007] Figure 2 An example system is depicted.

[0008] Figures 3A-3D An example of the operation is depicted.

[0009] Figures 3E-1 to 3E-3 An example of the operation is depicted.

[0010] Figure 4 An example process is depicted.

[0011] Figure 5 An example network interface device is depicted.

[0012] Figure 6A-6B An example network interface device is depicted.

[0013] Figure 7 An example system is depicted. DETAILED DESCRIPTION

[0014] In order to reduce the possibility of packet loss, and reduce the possibility of increasing delay due to buffer overflow at the switch or network interface device (e.g., due to the burst of received packets), various examples of switches or network interface devices can utilize supplementary memory to store packets. For example, when the target queue assigned for the received packet is full or reaches or exceeds the packet or byte occupancy level, the last hop switch (e.g., top of rack (TOR or ToR) switch) before the destination receiver can forward or copy the received packet to the overflow queue in the supplementary memory. The overflow queue can be assigned to the stream and / or other streams of the received packet, the same or different inlet port associated with the inlet port of the target queue, and / or the same or different outlet port associated with the outlet port of the target queue. The supplementary memory may include one or more of the following: on-chip memory in the switch, internal auxiliary memory in the switch, pooled memory connected via a device interface or a network interface, memory connected by a device interface, host memory, and / or memory hopped from the switch by one or more network interface devices. For example, if the host memory is utilized as a supplementary memory, the memory of the host connected to the switch can be utilized as a supplementary memory.

[0015] In some cases, buffering packets in the supplemental memory may introduce a delay to the delivery of packets at the receiver. This delay may come from the time to copy the packet header and / or payload data to the supplemental memory through the device interface or other interface, and the time to transmit the packet from the supplemental memory out of the egress port of the switch. The switch may cause the packet to be discarded and retransmitted by the sender, based on the time to retransmit the discarded packet being less than or equal to the delay associated with storing the packet in the selected supplemental memory and the switch transmitting the packet out. In some cases, the switch may cause subsequent received packets associated with the packets stored in the overflow queue to be stored in the overflow queue even if the occupancy level of the target queue has not reached full or even empty. The switch may cause the packet to be transmitted from the target queue based on a configured wait time, which may be based on a predicted burst duration.

[0016] The endpoint receiver may utilize a reordering buffer to store and reorder received packets and wait for a flush period before providing data of in-order packets to a packet processing protocol (e.g., Transmission Control Protocol (TCP) and other protocols). Packet ordering may be based at least on a packet sequence number (e.g., a TCP packet sequence number) or other value.

[0017] Figure 1An example system is shown. One or more of the sender network interface devices 116-0 to 116-A (where A is an integer) may send packets to one or more of the receiver network interface devices 126-0 to 126-B (where B is an integer) via the ToR switch 110, the spine switch 100, and the ToR switch 120 (e.g., the last hop switch). In some examples, the ToR switch 120 may utilize a queue monitoring circuit 122 to monitor queue utilization of one or more queues associated with an egress port. When the utilization of a queue exceeds a dynamically defined threshold, the queue monitoring circuit 122 may determine whether to transfer the packet traffic to an overflow buffer 126 in a supplemental memory or a memory of the ToR switch 120, or to discard the packet traffic. For example, the supplemental memory may be part of the host system 124. In some examples, if the destination is the receiver 126-B, the memory of the server 130-0 may be used as the supplemental memory.

[0018] The overflow buffer 126 may include a mix of packets of one or different flows and be associated with one or more ingress ports or different egress ports of the ToR switch 120. After buffering one or more packets for a configured amount of time, the ToR switch 120 may cause the packet traffic to be routed from the overflow buffer 126 to an egress port of the ToR switch 120 for transmission to one or more of the receivers 126-0 through 126-B. By utilizing the overflow buffer 126, packet loss at the ToR switch 110 may be reduced and the overall flow completion time (FCT) may be shortened. Although the examples are described with respect to the ToR switch 120, such examples may also be applied to the operations of the ToR switch 110 (e.g., the queue monitor 112 and the overflow buffer 116) for sending packets to one or more of the transmitters 116-0 through 116-B.

[0019] Transferring the packet to the overflow buffer 126 may introduce a delay in the packet's traversal to the receiver network interface device. The ToR switch 120 may copy or transfer the received packet to the overflow buffer 126 before forwarding it based on: (a) the baseline latency between the sender and the ToR switch 120 (e.g., latency in an unloaded structure) and / or (b) when the latency of the packet transfer is less than the latency introduced by the packet retransmission. For example, the latency of the packet retransmission may be based on the number of network interface device hops traversed from the sender to the ToR switch 120, where such hop traversal count may be recorded in the header of the packet. When the baseline latency level between the sender and the ToR switch 120 (with or without packet retransmission) exceeds a certain level, transferring the packet to the overflow buffer 126 may result in a latency below the baseline latency level and may reduce the packet traversal latency from the sender to the ToR switch 120. When the delay level caused by transferring packets to overflow buffer 126 and reordering at the receiver exceeds the delay introduced by packet retransmission (eg, baseline delay), discarding packets and causing packet retransmission may result in a reduction in packet reception delay.

[0020] If the overflow buffer 126 is used by packets of other flows, for example, the supplemental memory is connected to a host that sends and receives packets of the flow, the supplemental memory may become congested. Therefore, the queue monitor 122 can load balance between the supplemental memory devices when selecting which supplemental memory to use or whether to use the supplemental memory. Examples of ways to select the supplemental memory to be used include, but are not limited to, selecting the supplemental memory with the lowest utilization. In some examples, the selected supplemental memory may be used for an integer N consecutive packets that belong to the same or different flows as the packets buffered in the overflow buffer 126 in an attempt to provide ordered packet delivery to the receiver. In the case of a supplemental memory that also stores packets for independent traffic (e.g., packets from the host 124 or the server 130-0), the utilization of the supplemental memory by packets of the independent traffic and the transferred packets from the ToR switch 120 can be considered to determine the utilization level of the supplemental memory.

[0021] The reflection or transfer delay can be based on the amount of time that the packet is stored in the overflow buffer 126. The overflow buffer 126 can buffer packets that would otherwise be discarded due to overflow, for example, during a traffic burst, which may be of short duration. Therefore, in some cases, packets transferred to the overflow buffer 126 may not be forwarded to the next hop (e.g., one or more of the receivers 126-0 and 126-1) until the burst subsides. If the ToR switch 120 can observe the egress utilization of the next hop, the ToR switch 120 (e.g., the queue monitor 122) can implement a dynamic wait time so that the ToR switch 120 can cause the packets buffered in the overflow buffer 126 to be retrieved and forwarded after the wait time. The wait time can be based on historical measurements of the burst duration (or after a predefined wait time).

[0022] In some examples, determining whether to place a packet in the overflow buffer 126 or to discard the packet can be based on the priority of the packet's flow, a per-packet priority, a service level agreement (SLA), a service level objective (SLO), a priority of the packet flow, or other factors. For example, a higher priority packet can be stored in the overflow buffer 126 in priority over a lower priority packet, while the lower priority packet can be discarded. In some examples, it can be determined when the ToR switch 120 transmits a packet from the overflow buffer 126 based on the priority of the packet's flow, a per-packet priority, an SLA, an SLO, a priority of the packet flow, or other factors. For example, a higher priority packet can be transmitted from the overflow buffer 126 in priority over a lower priority packet.

[0023] In the presence of diverted packets and reflection (diversion) delays, the receiver may receive out-of-order packets. Packet reordering 128-0 to 128-B can reorder packets based on the sequence number provided in the packet header, and after the flush duration, provide the ordered packets to the protocol stack in the host (e.g., one or more of the servers 130-0 to 130-B) for protocol processing (e.g., processing of the packet header by the operating system (OS)). The flush duration can be greater than or equal to the delay introduced by buffering packets in the overflow buffer 126, thereby increasing the probability that the packet protocol layer receives ordered packets when the reordering buffer is flushed. However, this waiting adds delay, so the flush duration can be shortened to reduce delay.

[0024] The flush duration and reflection delay may be low so as not to add too much latency, which may result in delayed congestion notifications being returned to the sender and introduce delayed reactions and slowdowns, thus potentially increasing packet losses. Small message flows that can utilize a small amount of bandwidth and may be latency sensitive may bypass supplemental buffering and queue in buffer 126 without significantly degrading application performance. Such flows may be identified as sub-maximum transmission unit (MTU) packets based on header information.

[0025] Protocol processing performed by one or more of the servers 130-0 to 130-B may apply media access control (MAC) layer processing to the packet using MAC context information (including using driver data structures, driver statistics structures, etc.). The MAC context information may be pre-fetched into a cache of a core performing MAC layer processing. Protocol processing may apply IP layer processing of an Internet Protocol (IP) header and determine whether the packet is pushed to TCP layer processing or forwarded to another device. The IP layer may extract information from the packet to check the IPv4 context (e.g., action, forwarding, upstream host). The IP context information pre-fetched into the cache may be used to determine the next stage of the received packet. TCP layer processing may include checking the TCP header, determining TCP compliance to check whether the sequence number is as expected. The TCP context information may be loaded into a cache of a core performing TCP layer processing to process the packet. The TCP context information may include one or more of the following: sequence number, congestion window, packets to be processed, out-of-order queue information, etc. For example, the TCP context information may be loaded into a cache of a core performing TCP layer processing.

[0026] Figure 2 For example, network interface 206 may transmit packet 208 to network interface 240 via switch 220 and possibly one or more other switches or routers at the request of process 204 executed by circuitry of node 202. Figure 7 At least one example of node 202 is depicted. Switch 220 may be positioned as a ToR switch and positioned as the last hop before an endpoint receiver network interface device 240 or other switch.

[0027] The switch 220 can utilize the packet processing circuit 222 to parse and process the packet header based on the match-action operation. In some examples, the packet processing circuit 222 can perform the operation of the queue selector 223. The queue selector 223 can select a buffer from the buffer 226 allocated in the memory 224 or one or more of the buffers 230-0 to 230-X allocated in the respective memory devices 228-0 to 228-X to store the packet received from the network interface 206, where X is an integer. For example, the selection of one or more of the buffers 226 or buffers 230-0 and / or buffers 230-X can be based on various criteria described herein. In some examples, one or more of the buffers 230-0 to 230-X can be allocated in the memory 224 (e.g., on the same chip as the packet processing circuit 222) and / or in one or more of the memories 228-0 to 228-X.

[0028] In some examples, the packet processing circuit 222 may be included in a system on chip (SoC). In some examples, the packet processing circuit 222 may be implemented as a packet processing pipeline that performs a match-action based on a configuration. For example, it may be determined in which memory or buffer the packet is stored based on a match-action operation. In some examples, the memory 224 may be connected to the SoC of the switch 220, wherein the SoC is connected to one or more ingress ports and one or more egress ports. In some examples, the memory devices 228-0 to 228-X may be connected to the switch SoC via at least one device interface.

[0029] To transfer a packet in packets 208 to one or more of buffers 230-0 to 230-X, packet processing circuitry 222 may select another egress port for the packet instead of the assigned egress port. In some examples, another egress port may provide communication to a selected supplemental memory device or device interface (e.g., Compute Express Link (CXL) or Peripheral Component Interconnect Express (PCIe)) to provide communication with the assigned supplemental memory device. If the packet is not transferred to the selected supplemental memory, packet processing circuitry 222 may forward the packet to the assigned egress port for transmission to network interface 240.

[0030] In some examples, the packet processing circuit 222 may select the supplemental memory with the lowest queue utilization to achieve load-balanced traffic transfer according to the following process.

[0031]

[0032] Line 2 determines that the sender-receiver pair is not in the same ToR. Line 3 checks if the current packet size is equal to or greater than one Maximum Transmission Unit (MTU). If the packet is small, for example, there is only one flow of sub-MTU packets, this check bypasses the supplemental memory. Line 4 determines if the sender-receiver pair is not in the same ToR. Line 3 checks if the current packet size is equal to or greater than one Maximum Transmission Unit (MTU). If the packet is small, for example, there is only one flow of sub-MTU packets, then this check bypasses the supplemental memory. Line 4 divert ) checks if the current packet cannot fit into the memory, and line 5 checks if the packet has already been transferred to avoid transferring the same packet twice. In line 6, if the four conditions in lines 2, 3, 4, and 5 are true, find the port with the lowest queue occupancy of the ToR other than the original destination egress port. min In line 7, if port min Not the actual receiver port packet.port egress , you can set the packet's egress port to port in line 8 min , and the packet is marked as transferred in line 9. Thus, the packet can be transferred to the selected supplemental storage, or the packet is forwarded to the intended receiver. If any of the conditions in lines 2, 3, 4, or 5 are not true, and the packet has been marked as transferred in line 10, the state of the packet is reset to not transferred in line 11.

[0033] At the supplemental storage, the buffered packets may be routed through one or more first in first out (FIFO) queues (e.g., one or more of buffers 230-0 to 230-X). If the supplemental storage stores packets for flows initiated by a host server, switch 220 may perform arbitration to determine whether to transmit packets initiated by a host server (not shown) or packets buffered due to overflow. In some examples, arbitration may prioritize the transmission of packets from the host server over the diverted packets.

[0034] The switch 220 may transmit the packet 232 to the network interface 240 based on the packets stored in the buffer 226 or one or more of the buffers 230-0 and / or the buffers 230-X. In some examples, the packet processing circuit 222 may reorder the packets stored in the buffer 226 and one or more of the buffers 230-0 to 230-X based on the packet sequence number specified in the header field of the packet, and transmit the packets of the same flow to the network interface 240 in the sequence number order. In some examples, the packet processing circuit 222 may cause the packets stored in the buffer 226 and one or more of the buffers 230-0 to 230-X to be transmitted to the network interface 240 without considering the packet ordering according to the sequence number. For example, the packet processing circuit 222 may cause the transmission of the packets stored in the buffers 230-0 to 230-X to take precedence over the packets in the buffer 226 to reduce the packet delivery delay from the buffers 230-0 to 230-X to the network interface 240.

[0035] In some examples, the network interface 240 can perform packet reordering 242 to reorder received packets 232 from the switch 230 after the flush duration and before copying the received packets for processing by the node 250. For example, the node 250 can process the packets using a driver or an operating system. The flush duration of the receiver's reorder buffer can be greater than or equal to the delay introduced by buffering packets in the supplemental memory, thereby increasing the probability that the packet protocol layer receives in-order packets when the reorder buffer is flushed.

[0036] A packet may be used herein to refer to a collection of bits in various formats that can be sent over a network, such as an Ethernet frame, an IP packet, a TCP segment, a UDP datagram, etc. In addition, as used in this document, references to L2, L3, L4, and L7 layers (or Layer 2, Layer 3, Layer 4, and Layer 7) refer to the second data link layer, the third network layer, the fourth transport layer, and the seventh application layer, respectively.

[0037] A flow can be a sequence of packets transmitted between two endpoints, generally representing a single session using a known protocol. Therefore, a flow can be identified by a set of defined tuples or header field values, and for routing purposes, a flow is identified by two tuples (e.g., source address and destination address) that identify the endpoints. For content-based services (e.g., load balancers, firewalls, intrusion detection systems, etc.), flows can be distinguished with finer granularity by using N-tuples (e.g., source address, destination address, IP protocol, transport layer source port and destination port). Packets in a flow are expected to have the same set of tuples in the packet header. A packet flow can be identified by a combination of tuples (e.g., Ethernet type field, source and / or destination IP address, source and / or destination user datagram protocol (UDP) port, source / destination TCP port, or any other header field) and a unique source and destination queue pair (queue pair, QP) number or identifier.

[0038] Reference to flows may alternatively or additionally refer to tunnels (e.g., Multiprotocol Label Switching (MPLS) Label Distribution Protocol (LDP), Segment Routing over IPv6 dataplane (SRv6), VXLAN tunnel traffic, GENEVE tunnel traffic, network slicing based on virtual local area network (VLAN), the techniques described by Mudigonda, Jayaram et al. in "Spain: Cots data-center ethernet for multipathing over arbitrary topologies" (NSDI. Vol. 10, 2010) (hereinafter referred to as "SPAIN"), etc.

[0039] Figures 3A-3D Describes the sequence of operations. Figure 3A An example of transmission of packet 208 by network interface 206 via switch 220 to network interface 240 is depicted. Figure 3B An example is depicted where one or more of packets 208 are stored in buffer 226 for forwarding to an egress port. Additionally, queue selector 223 may transfer one or more of packets 208 to one or more of buffers 230-0 to 230-X based on the level of buffer 226 and factors such as whether another packet of the same flow is stored in one or more of buffers 230-0 to 230-X. Figure 3C An example of forwarding packets from buffer 226 and / or one or more of buffers 230 - 0 through 230 -X to network interface 240 is depicted. Figure 3D An example of packet reordering of received packets 232 is depicted.

[0040] Figures 3E-1 to 3E-3 An example of operation is depicted. In some examples, buffers iq1, iq2, iq3, eq can be allocated in the memory and / or supplemental memory of the switch. buffer 、iq buffer and / or eq receiver One or more of the buffers iq1 to iq3 may be associated with one or more inlet ports, and the buffers eq receiver Can be associated with one or more egress ports. Figure 3E-1 An example is shown in which the egress queue packet occupancy level exceeds the overflow level, and packet 5 is transferred to the overflow egress queue (eq buffer ),like Figure 3E-2 As shown. Arbitration for outgoing packets from queues iq1-iq3 can be based on flow priority, round-robin, weighted round-robin, or other factors. In some examples, the egress buffer eq buffer Can be associated with a port or device interface connected to a buffer node (e.g., supplemental memory in a host or other network interface device). A buffer node may include a host with supplemental memory connected to a switch via a device interface or network port. Examples of buffer nodes include a host with supplemental memory, a network interface device with supplemental memory, a memory pool, or a memory device connected to a device interface.

[0041] Figure 3E-3 An example of egressing packets from an overflow egress queue is depicted. Egressing packets from an overflow queue in a supplemental storage may be based on one or more of: prioritizing packets initiated for transmission by a host server associated with the supplemental storage, first-in, first-out (FIFO), prioritizing diverted packet traffic, egress based on packet priority, or otherwise. Note that while only a single overflow queue is shown, multiple overflow queues may be used, with two or more overflow queues associated with different priority levels, or with at least one overflow queue assigned to diverted traffic and at least one other overflow queue associated with traffic initiated by a host server.

[0042] In some examples, the transferred packets can be in the slave buffer eq receiver Before being transmitted, it is allocated in the exit buffer eq buffer and is allocated to the inlet buffer iq buffer In some examples, the egress buffer eq bufferand / or inlet buffer iq buffer can be allocated in the memory and / or supplemental memory of the switch. In some examples, the buffer eq buffer andiq buffer Can be associated with one or more network interface ports and / or device interfaces connected to the buffer node. In some examples, the transferred packet 5 is not associated with the ingress queue (iq buffer ) is associated with eq buffer Transfer group 5 in is directly assigned to eq receiver , without putting it into iq buffer middle.

[0043] Figure 4 An example process is depicted. The process may be performed by a switch. At 402, based on receiving a packet and according to a configuration, the received packet may be stored in a first buffer or a second buffer, or may be discarded. For example, the first buffer may be allocated in a memory of the switch, and the second buffer may be allocated in a memory of the switch and / or a supplemental memory coupled to the switch, such as in a memory pool, a memory device, or a host system one or more hops away from the switch. For example, the packet may be stored in the first buffer based on a configuration that specifies that the level of the first buffer is equal to or below a configured threshold level. For example, the packet may be stored in the second buffer based on one or more of the following: a previous packet of the same flow is stored in the second buffer, or the level of the first buffer is above a configured threshold level. For example, a determination to discard a packet may be based on a configuration that specifies that the packet is discarded instead of being stored in the second buffer, for example, if the sender-to-switch delay (e.g., duration or number of hops) of the packet retransmission is less than or equal to the delay introduced by storing the packet in the second buffer and reordering the packet.

[0044] Based on storing the packet in the first buffer, the packet may be stored in the first buffer at 404. Based on the scheduled transmission of the packet, the packet may be transmitted to a receiver through the selected egress port.

[0045] Based on storing the packet in the second buffer, at 410, the packet may be stored in the second buffer. Based on the expiration of a timer specified in the configuration, the packet may be forwarded from the second buffer to the receiver through the egress port of the switch. For example, the timer may be based on a predicted duration of a burst of packets received at the switch. In some examples, the second buffer may be allocated in a supplemental memory connected to the switch via a device interface or other connection. The switch may load balance between available supplemental memory devices to select a supplemental memory for storing the packet.

[0046] In some examples, based on the level of the first buffer being equal to or above the level, the packet may be dropped at 420. In some examples, the switch may send a negative acknowledgement to the sender of the packet to cause a retransmission of the packet, or the switch may send a pause request to reduce the transmission rate of packets for the flow of the dropped packet.

[0047] Figure 5 An example network interface device or packet processing device is depicted. In some examples, the circuitry of the network interface device may buffer packets in a supplemental memory, as described herein. In some examples, the packet processing device 500 may be implemented as a network interface controller, a network interface card, a host fabric interface (HFI), or a host bus adapter (HBA), and such examples may be interchangeable. The packet processing device 500 may be coupled to one or more servers using a bus, PCIe, CXL, or a double data rate (DDR). The packet processing device 500 may be embodied as part of a system on chip (SoC) including one or more processors, or may be included in a multi-chip package that also includes one or more processors.

[0048] Some examples of the packet processing device 500 are part of an infrastructure processing unit (IPU) or a data processing unit (DPU), or are utilized by an IPU or DPU. An xPU may refer to at least an IPU, a DPU, a GPU, a GPGPU, or other processing unit (e.g., an accelerator device). An IPU or DPU may include a network interface having one or more programmable or fixed-function processors to perform load transfer of operations that would otherwise be performed by a CPU. An IPU or DPU may include one or more memory devices. In some examples, an IPU or DPU may perform virtual switch operations, manage storage transactions (e.g., compression, cryptography, virtualization), and manage operations performed on other IPUs, DPUs, servers, or devices.

[0049] The network interface 500 may include a transceiver 502, a processor 504, a transmit queue 506, a receive queue 508, a memory 510, and a host interface 512, and a DMA engine 552. The transceiver 502 may be capable of receiving and sending packets that conform to an appropriate protocol, such as Ethernet described in IEEE 802.3, although other protocols may be used. The transceiver 502 may receive and send packets from and to the network via a network medium (not depicted). The transceiver 502 may include a PHY circuit 514 and a media access control (MAC) circuit 516. The PHY circuit 514 may include encoding and decoding circuits (not shown) to encode and decode data packets according to applicable physical layer specifications or standards. The MAC circuit 516 may be configured to assemble the data to be sent into packets that include destination and source addresses as well as network control information and error detection hash values.

[0050] Processor 504 and / or system on chip (SoC) 550 may include one or more of the following: a processor, a core, a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other programmable hardware device that allows programming of network interface 500. For example, a "smart network interface" may utilize processor 504 to provide packet processing capabilities in the network interface.

[0051] The processor 504 and / or the system on chip 550 may include one or more packet processing pipelines that may be configured to perform match-actions on received packets to identify packet processing rules and next hops using information stored in a ternary content-addressable memory (TCAM) table or an exact match table in some embodiments. For example, a match-action table or circuit may be used, through which a hash of a portion of a packet is used as an index to find an entry. The packet processing pipeline may perform one or more of the following: packet parsing (parser), exact match-action (e.g., a small exact match (SEM) engine or a large exact match (LEM)), wildcard match-action (WCM), longest prefix match block (LPM), hash block (e.g., receive side scaling (RSS)), packet modifier (modifier), or traffic manager (e.g., transmission rate metering or shaping). For example, the packet processing pipeline may implement an access control list (ACL) or drop packets due to queue overflow.

[0052] The configuration of the operation of the processor 504 and / or the system-on-chip 550 (including its data plane) can be programmed based on one or more of the following: a protocol-independent packet processor (P4), software for open networking in the cloud (SONiC), Network Programming Language (NPL), DOCA TM , Infrastructure Programmer Development Kit (IPDK), etc.

[0053] As described herein, processor 504, system on chip 550, or other circuitry may be configured to allocate packets for storage in a buffer in a memory of network interface 750 or another device, or to drop packets and determine when to transmit packets.

[0054] The packet distributor 524 can use the timeslot allocation or RSS described herein to provide distribution of received packets so that they can be processed by multiple CPUs or cores. When the packet distributor 524 uses RSS, the packet distributor 524 can calculate a hash or make another determination based on the content of the received packets to determine which CPU or core is to process the packet.

[0055] Interrupt coalescing 522 may perform interrupt throttling, whereby the network interface interrupt coalescing 522 waits for multiple packets to arrive, or waits for a timeout to expire, before generating an interrupt to the host system to process the received packet(s). Receive Segment Coalescing (RSC) may be performed by the network interface 500, whereby portions of an incoming packet are combined into segments of the packet. The network interface 500 may provide the coalesced packet to the application.

[0056] The direct memory access (DMA) engine 552 can copy packet headers, packet payloads and / or descriptors directly from host memory to the network interface, or vice versa, rather than copying the packets to an intermediate buffer at the host and then using another copy operation from the intermediate buffer to the destination buffer.

[0057] Memory 510 may be any type of volatile or non-volatile memory device and may store any queue or instruction for programming network interface 500. Transmit queue 506 may include data or references to data for transmission by the network interface. Receive queue 508 may include data or references to data received by the network interface from the network. Descriptor queue 520 may include descriptors that reference data or packets in transmit queue 506 or receive queue 508. Host interface 512 may provide an interface with a host device (not depicted). For example, host interface 512 may be compatible with PCI, PCI Express, PCI-x, Serial ATA, and / or USB compatible interfaces (although other interconnect standards may be used).

[0058] Fig. 6A An example system is depicted. Host 600 may include a processor, memory devices, device interfaces, and other circuits, such as reference Figure 2 , Figure 5 and / or Figure 6BThose described in one or more of . The processor of the host 600 can execute software, such as processes (e.g., applications, microservices, virtual machines (VMs), micro VMs, containers, processes, threads, or other virtualized execution environments), operating systems (OS), and device drivers. The OS or device driver can configure the network interface device or packet processing device 610 to communicate with a software defined networking (SDN) controller 650 via a network using one or more control planes to configure the operation of one or more control planes. The host 600 can be coupled to the network interface device 610 via a host or device interface 644.

[0059] The network interface device 610 may include multiple computing complexes, such as an Acceleration Compute Complex (ACC) 620 and a Management Compute Complex (MCC) 630, as well as packet processing circuits 640 and network interface technology for communicating with other devices via a network. The ACC 620 may be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or at least a combination of a plurality of processors. Figure 6B and / or Figure 7 Similarly, the MCC 630 may be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or at least a combination of the above. Figure 2 , Figure 5 and / or Figure 6B In some examples, the ACC 620 and MCC 630 can be implemented as separate cores in a CPU, different cores in different CPUs, different processors in the same integrated circuit, or different processors in different integrated circuits. In some examples, the circuitry and software of the network interface device 610 can be configured to determine whether to transfer a packet to a buffer in supplemental memory, or to discard a packet, and when to transmit a packet, as described herein.

[0060] The network interface device 610 may be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or at least a plurality of processors. Figure 2 , Figure 5and / or Figure 6B The packet processing pipeline circuitry 640 may process packets as directed or configured by one or more control planes executed by the plurality of compute complexes. For example, the processing pipeline circuitry 640 may be configured to determine whether to store a packet in a buffer in a memory of the network interface 750, or in another device, or to discard a packet, and when to transmit a packet, as described herein. In some examples, the ACC 620 and the MCC 630 may execute respective control planes 622 and 632.

[0061] The SDN controller 650 can upgrade or reconfigure the software (e.g., control plane 622 and / or control plane 632) executed on the ACC 620 by the content of the packets received via the packet processing device 610. In some examples, the ACC 620 can execute a control plane operating system (OS) (e.g., Linux) and / or a control plane application 622 (e.g., a user space or kernel module) used by the SDN controller 650 to configure the operation of the packet processing pipeline 640. The control plane application 622 can include a Generic Flow Table (GFT), ESXi, NSX, Kubernetes control plane software, application software for managing cryptographic configuration, a Programming Protocol-independent Packet Processor (P4) runtime daemon, a target-specific daemon, a Container Storage Interface (CSI) agent, or a remote direct memory access (RDMA) configuration agent.

[0062] In some examples, the SDN controller 650 can communicate with the ACC 620 using a remote procedure call (RPC), such as Google remote procedure call (gRPC) or other services, and the ACC 620 can convert the request to a target-specific protocol buffer (protobuf) request to the MCC 630. gRPC is a remote procedure call solution based on data packets sent between a client and a server. Although gRPC is an example, other communication schemes may also be used, such as but not limited to Java remote method calls, Modula-3, RPyC, distributed Ruby, Erlang, Elixir, action message format, remote function call, open network computing RPC, JSON-RPC, and the like.

[0063] In some examples, the SDN controller 650 can provide packet processing rules for execution by the ACC 620. For example, the ACC 620 can program table rules (e.g., header field matches and corresponding actions) applied by the packet processing pipeline circuit 640 based on changes in policy and changes in VMs, containers, microservices, applications, or other processes. The ACC 620 can be configured to provide network policies as flow cache rules into a table to configure the operation of the packet processing pipeline 640. For example, the control plane application 622 executed by the ACC can configure the rule table applied by the packet processing pipeline circuit 640 with rules based on packet type and content to define traffic destinations. The ACC 620 can program table rules (e.g., matches-actions) into a memory accessible to the packet processing pipeline circuit 640 based on changes in policy and changes in VMs.

[0064] For example, the ACC 620 can implement a virtual switch, such as a vSwitch or Open vSwitch (OVS), Stratum, or Vector Packet Processing (VPP), which provides communication between virtual machines executed by the host 600 or with other devices connected to the network. For example, the ACC 620 can configure the packet processing pipeline circuit 640 regarding which VM receives traffic and what kind of traffic the VM can transmit. For example, the packet processing pipeline circuit 640 can implement a virtual switch, such as a vSwitch or Open vSwitch, which provides communication between virtual machines executed by the host 600 and the packet processing device 610.

[0065] The MCC 630 may execute a host management control plane, a global resource manager, and perform hardware register configuration. The control plane 632 executed by the MCC 630 may perform provisioning and configuration of the packet processing circuit 640. For example, a VM executing on the host 600 may utilize the packet processing device 610 to receive or send packet traffic. The MCC 630 may execute startup, power, management, and manageability software (SW) or firmware (FW) code to start and initialize the packet processing device 610, manage device power consumption, provide connectivity with a management controller (e.g., a baseboard management controller (BMC)), and other operations.

[0066] One or both control planes of ACC 620 and MCC 630 may define traffic routing table contents and network topology, which are applied by packet processing circuitry 640 to select a path for a packet to go to the next hop in the network or to a destination network connection device. For example, a VM executing on host 600 may utilize packet processing device 610 to receive or send packet traffic.

[0067] The ACC 620 may execute a control plane driver to communicate with the MCC 630. At least to provide a configuration and provisioning interface between the control planes 622 and 632, the communication interface 625 may provide control plane to control plane communication. The control plane 632 may perform gatekeeper operations for configuration of shared resources. For example, via the communication interface 625, the ACC control plane 622 may communicate with the control plane 632 to perform one or more of the following: determine hardware capabilities, access data plane configurations, reserve hardware resources and configurations, communicate between the ACC and MCC via interrupts or polling, subscribe to receive hardware events, perform indirect hardware register reads and writes for debuggability, flash and physical layer interface (phy) configuration, or perform system provisioning for different deployments of network interface devices, such as storage nodes, tenant hosting nodes, microservices backends, compute nodes, or others.

[0068] The communication interface 625 may be utilized by negotiation protocols and configuration protocols running between the ACC control plane 622 and the MCC control plane 632. The communication interface 625 may include a general mailbox for different operations performed by the packet processing circuit 640. Examples of operations of the packet processing circuit 640 include: issuance of non-volatile memory express (NVMe) reads or writes, non-volatile memory express over fabrics (NVMe-oF), and non-volatile memory over fabrics (NVMe-oF). TM ) issuance of reads or writes, a lookaside cryptoEngine (LCE) (e.g., compression or decompression), an address translation engine (ATE) (e.g., an input output memory management unit (IOMMU) to provide virtual to physical address translation), encryption or decryption, configured as a storage node, configured as a tenant hosting node, configured as a compute node, providing multiple different types of services between different Peripheral Component Interconnect Express (PCIe) endpoints, or others.

[0069] The communication interface 625 may include one or more mailboxes that can be accessed as registers or memory addresses. For communications from the control plane 622 to the control plane 632, the communications may be written to the one or more mailboxes by the control plane driver 624. For communications from the control plane 632 to the control plane 622, the communications may be written to the one or more mailboxes. The communications written to the mailboxes may include descriptors, including message opcodes, message errors, message parameters, and other information. The communications written to the mailboxes may include defined format messages that convey data.

[0070] The communication interface 625 may provide communication based on writing or reading to specific memory addresses (e.g., dynamic random access memory (DRAM)), registers, and other mailboxes that are written and read to pass commands and data. In order to provide secure communication between the control planes 622 and 632, the registers and memory addresses (and memory address translations) used for communication can only be written or read by the control planes 622 and 632 or cloud service provider (CSP) software executing on the ACC 620 and device vendor software, embedded software, or firmware executing on the MCC 630. The communication interface 625 may support communication between multiple different computing complexes, such as from the host 600 to the MCC 630, the host 600 to the ACC 620, the MCC 630 to the ACC 620, the baseboard management controller (BMC) to the MCC 630, the BMC to the ACC 620, or the BMC to the host 600.

[0071] The packet processing circuit 640 may be implemented using one or more of the following: an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a processor executing software, or other circuits. The control plane 622 and / or 632 may configure the packet processing pipeline circuit 640 or other processor to perform operations related to NVMe, NVMe-oF read or write, backup cryptographic engine (LCE), address translation engine (ATE), local area network (LAN), compression / decompression, encryption / decryption, or other accelerated operations.

[0072] Various message formats may be used to configure the ACC 620 or the MCC 630. In some examples, a P4 program may be compiled and provided to the MCC 630 to configure the packet processing circuit 640. The following is a JSON configuration file that may be transferred from the ACC 620 to the MCC 630 to obtain the capabilities of the packet processing circuit 640 and / or other circuits in the packet processing device 610. More specifically, the file may be used to specify a number of transmit queues, a number of receive queues, a number of supported traffic classes (TCs), a number of available interrupt vectors, a number of available virtual ports and port types, the size of the allocated memory, supported parser profiles, exact match table profiles, packet mirroring profiles, and the like.

[0073] Figure 6B An example network interface device system is depicted. Various examples of packet processing devices or network interface devices 610 may be utilized Figure 2 , Figure 5 and / or Fig. 6A Components of a system. In some examples, the network interface device 610 may be configured to determine whether to store the packet in a buffer in the memory 684 or in an overflow buffer in a different attached memory device, or to discard the packet, and when to transmit the packet, as described herein. In some examples, the packet processing device or network interface device may refer to one or more of the following: a network interface controller (NIC), a NIC enabling remote direct memory access (RDMA), a SmartNIC, a router, a switch, a forwarding element, an infrastructure processing unit (IPU), or a data processing unit (DPU). The network subsystem 660 may be communicatively coupled to the computing complex 680. The device interface 662 may provide an interface for communicating with the host. Various examples of the device interface 662 may utilize protocols based on peripheral component interconnect express (PCIe), compute express link (CXL) or other protocols, and virtual device interfaces, such as virtual device interfaces.

[0074] The interface 664 can initiate and terminate at least offloaded remote direct memory access (RDMA) operations, non-volatile memory express (NVMe) read or write operations, and LAN operations. The packet processing pipeline 666 can perform packet processing (e.g., packet header and / or packet payload) based on the configuration and support quality of service (QoS) and telemetry reporting. The inline processor 668 can perform offloaded encryption or decryption of packet communications (e.g., Internet Protocol Security (IPSec) or other). The traffic shaper 670 can schedule the transmission of communications. The network interface 672 can provide at least an interface to an Ethernet network through media access control (MAC) and serializer / deserializer (Serdes) operations.

[0075] Core 682 may be configured to perform infrastructure operations such as a storage initiator, a Transport Layer Security (TLS) proxy, a virtual switch (e.g., a vSwitch), or other operations. Memory 684 may store applications and data to be executed or processed. Offload circuitry 686 may perform at least cryptographic and compression operations for use by a host or compute complex 680. Offload circuitry 686 may include one or more graphics processing units (GPUs) that may access memory 684. Management complex 688 may perform secure boot, lifecycle management, and management of network subsystem 660 and / or compute complex 680.

[0076] Figure 7A system is depicted. In some examples, the circuitry of system 700 can determine whether to store a packet in a buffer in network interface device 750 or in a buffer in supplemental memory, or to discard a packet, and when to transmit a packet, as described herein. System 700 includes a processor 710 that provides processing, operational management, and execution of instructions for system 700. Processor 710 may include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), XPU, processing core, or other processing hardware used to provide processing for system 700, or a combination of processors. The XPU may include one or more of the following: a CPU, a graphics processing unit (GPU), a general purpose GPU (GPGPU), and / or other processing units (e.g., an accelerator or a programmable or fixed function FPGA). The processor 710 controls the overall operation of the system 700 and may be or may include one or more programmable general or special purpose microprocessors, digital signal processors (DSPs), programmable controllers, application specific integrated circuits (ASICs), programmable logic devices (PLDs), etc., or a combination of such devices.

[0077] In one example, the system 700 includes an interface 712 coupled to the processor 710, which may represent a higher speed interface or a high throughput interface for system components that require a higher bandwidth connection, such as the memory subsystem 720 or the graphics interface component 740, or the accelerator 742. The interface 712 represents an interface circuit, which may be a separate component or may be integrated onto the processor die. If present, the graphics interface 740 interfaces with the graphics component for providing a visual display to a user of the system 700. In one example, the graphics interface 740 generates a display based on data stored in the memory 730 or based on operations performed by the processor 710, or both. In one example, the graphics interface 740 generates a display based on data stored in the memory 730 or based on operations performed by the processor 710, or both.

[0078] The accelerator 742 may be a programmable or fixed-function load transfer engine that may be accessed or used by the processor 710. For example, one of the accelerators 742 may provide data compression (DC) capabilities, cryptographic services (e.g., public key encryption (PKE)), cryptography, hashing / authentication capabilities, decryption, or other capabilities or services. In some cases, the accelerator 742 may be integrated into a CPU socket (e.g., a connector of a motherboard or circuit board that includes a CPU and provides an electrical interface with the CPU). For example, the accelerator 742 may include a single-core or multi-core processor, a graphics processing unit, a logic execution unit, a single-level or multi-level cache, a functional unit that can be used to independently execute a program or thread, an application specific integrated circuit (ASIC), a neural network processor (NNP), programmable control logic, and a programmable processing element, such as a field programmable gate array (FPGA). The accelerator 742 may provide multiple neural networks, CPUs, processor cores, general purpose graphics processing units, or graphics processing units for use by artificial intelligence (AI) or machine learning (ML) models. For example, the AI ​​model may use or include any one or combination of the following: a reinforcement learning scheme, a Q learning scheme, a deep Q learning, or an asynchronous advantage actor-evaluator (A3C), a combination neural network, a recursive combination neural network, or other AI or ML models. Multiple neural networks, processor cores, or graphics processing units may be used by the AI ​​or ML model to perform learning and / or reasoning operations.

[0079] The memory subsystem 720 represents the main memory of the system 700 and provides storage for the code to be executed by the processor 710 or the data values ​​to be used when executing the routine. The memory subsystem 720 may include one or more memory devices 730, such as read-only memory (ROM), flash memory, one or more varieties of random access memory (RAM) (such as DRAM), or other memory devices, or a combination of such devices. The memory 730 stores and hosts (among other things) an operating system (OS) 732 to provide a software platform for the execution of instructions in the system 700. In addition, applications 734 can be executed from the memory 730 on the software platform of the OS 732. Applications 734 represent programs with their own operating logic to perform the execution of one or more functions. Processes 736 represent agents or routines that provide auxiliary functions to the OS 732 or one or more applications 734 or a combination thereof. The OS 732, applications 734, and processes 736 provide software logic to provide functions for the system 700. In one example, the memory subsystem 720 includes a memory controller 722, which is a memory controller for generating and issuing commands to the memory 730. It will be appreciated that the memory controller 722 can be a physical part of the processor 710 or a physical part of the interface 712. For example, the memory controller 722 can be an integrated memory controller that is integrated onto a circuit with the processor 710.

[0080] Application 734 and / or process 736 may alternatively or additionally refer to a virtual machine (VM), container, microservice, processor or other software. Various examples described herein may execute an application consisting of microservices, wherein the microservices run in their own processes and communicate using protocols (e.g., application program interface (API), hypertext transfer protocol (HTTP) resource API, message service, remote procedure call (RPC), or Google RPC (gRPC)). Microservices may communicate with each other using a service grid and be executed in one or more data centers or edge networks. Microservices may be independently deployed using centralized management of these services. Management systems may be written in different programming languages ​​and use different data storage technologies. Microservices may be characterized by one or more of the following: polyglot programming (e.g., code written in multiple languages ​​to capture additional functionality and efficiency that a single language cannot provide), or lightweight container or virtual machine deployment, and decentralized continuous microservice delivery.

[0081] In some examples, OS 732 may be Server or personal computer, VMware vSphere, openSUSE, RHEL, CentOS, Debian, Ubuntu, or any other operating system. The OS and drivers can be found on Texas etc., and executes on processors sold or designed by the company.

[0082] In some examples, OS 732, a system administrator, and / or a coordinator can configure network interface 750 to determine whether to store packets in a buffer in a memory of network interface 750, in another device, or to discard packets, and when to transmit packets, as described herein.

[0083] Although not specifically illustrated, it will be understood that system 700 may include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, an interface bus, or other. A bus or other signal line may couple components together communicatively or electrically, or simultaneously couple components communicatively and electrically. A bus may include a physical communication line, a point-to-point connection, a bridge, an adapter, a controller, or other circuits or a combination of these. A bus may include, for example, one or more of the following: a system bus, a peripheral component interconnect (PCI) bus, a Hyper Transport or an industry standard architecture (ISA) bus, a small computer system interface (SCSI) bus, a universal serial bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus (Firewire).

[0084] In one example, system 700 includes interface 714, which can be coupled to interface 712. In one example, interface 714 represents interface circuit, which can include independent components and integrated circuits. In one example, multiple user interface components or peripheral components or both are coupled to interface 714. Network interface 750 provides system 700 with the ability to communicate with remote devices (e.g., servers or other computing devices) through one or more networks. Network interface 750 may include Ethernet adapter, wireless interconnection component, cellular network interconnection component, USB (universal serial bus) or other wired or wireless standard or exclusive interface. Network interface 750 can transmit data to equipment or remote devices in the same data center or rack, which can include sending data stored in memory. Network interface 750 can receive data from remote devices, which can include storing received data in memory. In some examples, the packet processing device or network interface device 750 may refer to one or more of the following: a network interface controller (NIC), a remote direct memory access (RDMA) enabled NIC, a SmartNIC, a router, a switch, a forwarding element, an infrastructure processing unit (IPU), or a data processing unit (DPU). Figure 4 , Figure 5 , Fig. 6A , Figure 6B and / or Figure 7 An example IPU or DPU is described.

[0085] In one example, system 700 includes one or more input / output (I / O) interfaces 760. I / O interfaces 760 may include one or more interface components through which a user interacts with system 700. Peripheral interfaces 770 may include any hardware interface not specifically mentioned above. Peripherals generally refer to devices that are dependently connected to system 700.

[0086] In one example, the system 700 includes a storage subsystem 780 to store data in a non-volatile manner. In one example, in certain system implementations, at least some components of the storage device 780 may overlap with components of the memory subsystem 720. The storage subsystem 780 includes (one or more) storage devices 784, which may be or may include any conventional media for storing large amounts of data in a non-volatile manner, such as one or more magnetic, solid-state, or optical-based disks, or a combination of these. The storage device 784 saves code or instructions and data 786 in a persistent state (e.g., the value is retained despite power interruptions to the system 700). The storage device 784 may be generally considered to be "memory," although the memory 730 is generally an execution or operating memory to provide instructions to the processor 710. The storage device 784 is non-volatile, while the memory 730 may include volatile memory (e.g., if power to the system 700 is interrupted, the value or state of the data is indeterminate). In one example, storage subsystem 780 includes controller 782 to interface with storage device 784. In one example, controller 782 is a physical part of interface 714 or processor 710, or may include circuitry or logic in both processor 710 and interface 714.

[0087] Volatile memory may include memory whose state (and therefore data stored therein) is indeterminate if power to the device is interrupted. Non-volatile memory (NVM) devices may include memory whose state is deterministic even if power to the device is interrupted.

[0088] In some examples, system 700 may be implemented using an interconnected computing platform of processors, memory, storage, network interfaces, and other components. High-speed interconnects can be used, such as Ethernet (IEEE 802.3), remote direct memory access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connection (QUIC), RDMA over Converged Ethernet (RoCE), Peripheral Component Interconnect express (PCIe), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omni-Path, Compute Express Link (CXL), HyperTransport, High-Speed ​​Fabric, NVLink, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Infinity Fabric ( Fabric, IF), Cache Coherent Interconnect for Accelerator (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and variants thereof. Data may be copied or stored to a virtualized storage node, or accessed using a protocol such as NVMe over Fabric (NVMe-oF) or NVMe (e.g., a non-volatile memory express (NVMe) device may operate in a manner consistent with the Non-Volatile Memory Express (NVMe) Specification Revision 1.3c, released on May 24, 2018 (the "NVMe Specification"), or a derivative or variant thereof).

[0089] Communications between devices may occur using a network that provides die-to-die communications; chip-to-chip communications; board-to-board communications; and / or package-to-package communications.

[0090] In one example, the system 700 may be implemented using an interconnected computing platform of processors, memory, storage, network interfaces, and other components. A high-speed interconnect such as PCIe, Ethernet, or optical interconnect (or a combination of these) may be used.

[0091] Examples herein can be implemented in various types of computing and networking devices, such as switches, routers, racks and blade servers, such as those employed in data center and / or server farm environments. Servers used in data centers and server farms include array server configurations, such as rack-based servers or blade servers. These servers are interconnected in communications via various network configurations, such as dividing groups of servers into local area networks (LANs), with appropriate switching and routing facilities between LANs to form private intranets. For example, cloud hosting facilities typically employ large data centers with numerous servers. The blade includes a separate computing platform that is configured to perform server-type functions, i.e., a "server on a card." Therefore, the blade includes components common to traditional servers, including a main printed circuit board (motherboard) that provides internal wiring (e.g., a bus) for coupling appropriate integrated circuits (ICs) and other components mounted on the board.

[0092] Various examples may be implemented using hardware elements, software elements, or a combination of the two. In some examples, hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, ASICs, PLDs, DSPs, FPGAs, memory cells, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some examples, software elements may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, APIs, instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination of these. Determining whether an example is implemented using hardware elements and / or software elements may vary according to any number of factors as desired by a given implementation, such as desired computing rates, power levels, heat tolerance, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints. A processor may be a hardware state machine, digital control logic, a central processing unit, or one or more combinations of any hardware, firmware, and / or software elements.

[0093] Some examples may be implemented using or as an article of manufacture or at least one computer-readable medium. The computer-readable medium may include a non-transitory storage medium to store logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and the like. In some examples, logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, APIs, instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination of these.

[0094] According to some examples, a computer-readable medium may include a non-transitory storage medium to store or maintain instructions that, when executed by a machine, a computing device, or a system, cause the machine, the computing device, or the system to perform methods and / or operations according to the described examples. The instructions may include any appropriate type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. The instructions may be implemented according to a predetermined computer language, manner, or syntax to instruct a machine, a computing device, or a system to perform a specific function. The instructions may be implemented using any appropriate high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.

[0095] One or more aspects of at least one example may be implemented by representative instructions stored on at least one machine-readable medium representing various logic within the processor, which when read by a machine, computing device or system causes the machine, computing device or system to fabricate logic to perform the techniques described herein. Such representations, known as "IP cores", may be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to load into fabrication machines that actually fabricate the logic or processor.

[0096] The appearance of the phrase "an example" or "an example" does not necessarily refer to the same example or embodiment. Any aspect described herein may be combined with any other aspect or similar aspect described herein, whether or not these aspects are described with respect to the same figure or element. The division, omission or inclusion of block functions depicted in the figures does not infer that the hardware components, circuits, software and / or elements used to implement these functions will necessarily be divided, omitted or included in the embodiments.

[0097] Some examples may be described using the expressions "coupled" and "connected" and their derivatives. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" may also refer to two or more elements that are not in direct contact, but still cooperate or interact.

[0098] The terms "first", "second" and the like do not represent any order, quantity or importance in this article, but are used to distinguish one element from another element. The term "one" herein does not represent a limitation on quantity, but represents the existence of at least one mentioned item. The term "assertion" used when referring to a signal herein refers to a state of the signal, in which the signal is valid, and the state can be achieved by applying any logic level (whether logic 0 or logic 1) (for example, low effective or high effective) to the signal. The term "subsequently" or "afterwards" may refer to immediately following or following after some or some other events. According to alternative embodiments, other operation sequences may also be performed. In addition, depending on the specific application, additional operations may be added or removed. Any combination of variations may be used, and those of ordinary skill in the art who benefit from the present disclosure will understand its many variations, modifications and alternative embodiments.

[0099] Unless specifically stated otherwise, disjunctive language such as the phrase "at least one of X, Y, or Z" is understood within the context to be generally used to state that an item, term, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that certain embodiments require the presence of at least one X, at least one Y, or at least one Z. Furthermore, unless specifically stated otherwise, conjunctive language such as the phrase "at least one of X, Y, and Z" should also be understood to refer to X, Y, Z, or any combination thereof, including "X, Y, and / or Z."

[0100] Illustrative examples of the devices, systems, and methods disclosed herein are provided below. Embodiments of the devices, systems, and methods may include any one or more of the examples described below, as well as any combination thereof.

[0101] Example 1 includes an apparatus comprising: a switch system on chip (SoC), comprising circuitry for: based on reception of a packet and a level of a first queue, selecting between a first memory and a second memory device among a plurality of second memory devices to store the packet, based on the selection of the first memory, storing the packet in the first memory, and based on the selection of the second memory device among the plurality of second memory devices, storing the packet in a selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the selected second memory device is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

[0102] Example 2 includes one or more examples wherein: the circuit causes the packet to be discarded based on a retransmission of the packet and a delay associated with storing the packet in the selected second memory device.

[0103] Example 3 includes one or more examples wherein: after storing the packet in the selected second memory device, the circuitry causes subsequent received packets of a flow associated with the packet to be stored in the selected second memory device.

[0104] Example 4 includes one or more examples, wherein: the circuit causes the packet to be transmitted from the selected second memory device based on a configured latency.

[0105] Example 5 includes one or more examples wherein the selected second memory device is to store packets of a plurality of different packet flows.

[0106] Example 6 includes one or more examples wherein the circuit includes a packet processing pipeline.

[0107] Example 7 includes one or more examples wherein the switch SoC is located in a last-hop switch before an endpoint receiver of the packet.

[0108] Example 8 includes one or more examples and includes at least one ingress port coupled to the switch SoC; at least one egress port coupled to the switch SoC; and one or more device interfaces coupled to the switch SoC.

[0109] Example 9 includes one or more examples and includes a method comprising: in a switch: based on reception of a packet and a level of a first queue, making a selection between a first memory and a second memory device among a plurality of second memory devices to store the packet, based on the selection of the first memory, storing the packet in the first queue in the first memory, and based on the selection of the second memory device among the plurality of second memory devices, storing the packet in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

[0110] Example 10 includes one or more examples and includes discarding the packet based on a time to retransmit the packet and a delay associated with storing the packet in the selected second memory device.

[0111] Example 11 includes one or more examples and includes, after storing the packet in the second queue, storing subsequently received packets of a flow associated with the packet in the second queue.

[0112] Example 12 includes one or more examples, and includes egressing the packet from the second queue based on a configured wait time.

[0113] Example 13 includes one or more examples wherein the second queue stores packets of a plurality of different packet flows.

[0114] Example 14 includes one or more examples wherein the switch is located in a last hop switch before an endpoint receiver of the packet.

[0115] Example 15 includes one or more examples and includes at least one non-transitory computer-readable medium, the medium including instructions stored thereon, which, if executed by one or more circuits, cause the one or more circuits in a switch to: based on reception of a packet and a level of a first queue, select between a first memory and a second memory device among a plurality of second memory devices to store the packet, based on the selection of the first memory, store the packet in the first queue in the first memory, and based on the selection of the second memory device among the plurality of second memory devices, store the packet in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

[0116] Example 16 includes one or more examples and includes instructions stored thereon that, if executed by one or more circuits, cause the one or more circuits to, in the switch, discard the packet based on a time to retransmit the packet and a delay associated with storing the packet in the selected second memory device.

[0117] Example 17 includes one or more examples and includes instructions stored thereon, which if executed by one or more circuits, cause the one or more circuits in the switch to: after storing the packet in the second queue, store subsequently received packets of the flow associated with the packet in the second queue.

[0118] Example 18 includes one or more examples and includes instructions stored thereon that, if executed by one or more circuits, cause the one or more circuits to, in the switch, egress the packet from the second queue based on a configured wait time.

[0119] Example 19 includes one or more examples wherein the second queue stores packets of a plurality of different packet flows.

[0120] Example 20 includes one or more examples and includes instructions stored thereon, which, if executed by one or more circuits, cause the one or more circuits to be used in the switch to: select the second memory device to store the packet based on a priority of the packet, a priority of a flow of the packet, or a service level agreement (SLA).

Claims

1. A device comprising: A switch system on chip (SoC) comprising a circuit for: selecting between a first memory and a second memory device among a plurality of second memory devices to store the packet based on receipt of the packet and a level of the first queue, Based on the selection of the first memory, storing the packet in the first memory, and Based on a selection of a second memory device among the plurality of second memory devices, the group is stored in the selected second memory device, wherein the group is associated with an ingress port and an egress port, and the selected second memory device is associated with a third port, the third port being different from the ingress port and the egress port associated with the group.

2. The device of claim 1, wherein: The circuitry causes the packet to be discarded based on a retransmission of the packet and a delay associated with storing the packet in the selected second memory device.

3. The device of claim 1, wherein: The circuitry causes the packet to be transferred out of the selected second memory device based on a configured latency.

4. The device according to claim 1, wherein: The selected second memory device is used to store packets of a plurality of different packet flows.

5. The device according to claim 1, wherein: The circuit includes a packet processing pipeline.

6. The device according to claim 1, wherein: The switch SoC is located in the last hop switch before the endpoint receiver of the packet.

7. The device of claim 1, comprising: at least one ingress port coupled to the switch SoC; at least one egress port coupled to the switch SoC; as well as One or more device interfaces coupled to the switch SoC.

8. The device according to any one of claims 1 to 7, wherein: After storing the packet in the selected second memory device, the circuitry causes storage of subsequently received packets of a flow associated with the packet in the selected second memory device.

9. A method comprising: In the switch: selecting between a first memory and a second memory device among a plurality of second memory devices to store the packet based on receipt of the packet and a level of the first queue, Based on the selection of the first memory, storing the packet in the first queue in the first memory, and Based on a selection of a second memory device among the plurality of second memory devices, the packet is stored in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

10. The method of claim 9, comprising: The packet is discarded based on a time to retransmit the packet and a delay associated with storing the packet in the selected second memory device.

11. The method of claim 9, comprising: After storing the packet in the second queue, subsequently received packets of the flow associated with the packet are stored in the second queue.

12. The method of claim 9, wherein: The second queue stores packets of a plurality of different packet flows.

13. The method of claim 9, wherein: The switch is located in the last hop switch before the endpoint receiver of the packet.

14. The method according to any one of claims 9 to 13, comprising: The packet is egressed from the second queue based on a configured wait time.

15. At least one non-transitory computer-readable medium, the medium comprising instructions stored thereon, the instructions, if executed by one or more circuits, causing the one or more circuits in a switch to: selecting between a first memory and a second memory device among a plurality of second memory devices to store the packet based on receipt of the packet and a level of the first queue, Based on the selection of the first memory, storing the packet in the first queue in the first memory, and Based on a selection of a second memory device among the plurality of second memory devices, the packet is stored in a second queue of the selected second memory device, wherein the packet is associated with an ingress port and an egress port, and the second queue is associated with a third port, the third port being different from the ingress port and the egress port associated with the packet.

16. The computer-readable medium of claim 15, comprising instructions stored thereon that, if executed by one or more circuits, cause the one or more circuits in the switch to: The packet is discarded based on a time to retransmit the packet and a delay associated with storing the packet in the selected second memory device.

17. The computer-readable medium of claim 15, comprising instructions stored thereon that, if executed by one or more circuits, cause the one or more circuits in the switch to: After storing the packet in the second queue, subsequently received packets of the flow associated with the packet are stored in the second queue.

18. The computer-readable medium of claim 15, comprising instructions stored thereon that, if executed by one or more circuits, cause the one or more circuits in the switch to: The packet is egressed from the second queue based on a configured wait time.

19. The computer-readable medium of claim 15, wherein: The second queue stores packets of a plurality of different packet flows.

20. The computer-readable medium of any one of claims 15-19, comprising instructions stored thereon, which if executed by one or more circuits, cause the one or more circuits in the switch to: The second memory device is selected to store the packet based on a priority of the packet, a priority of a flow of the packet, or a service level agreement (SLA).