Time-controlled switching

By synchronizing network devices and switches with time stamps, packet delivery is ordered, addressing data center congestion and packet loss issues, resulting in improved latency and efficiency.

DE102025106530A1Pending Publication Date: 2025-10-02INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025106530
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-02-20
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Data centers experience congestion and packet loss due to excessive data traffic, leading to slower data transfer and disordered packet reception, which is exacerbated by packet spraying techniques.

Method used

Implement time synchronization across network interface devices and switches to insert transmit time stamps into packets, enabling switches to order packets based on timestamps for delivery, thereby reducing out-of-order delivery and mitigating incast scenarios.

Benefits of technology

This approach reduces packet delivery latency and improves data flow efficiency by ensuring packets are received in the correct order, lowering the likelihood of congestion and enhancing overall network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The examples described herein relate to a network interface device including an interface to a port and circuitry. The circuitry may be configured to: receive a first packet including a timestamp associated with a previous or original transmission of the first packet by a sender network interface device; enqueue an entry for the first packet; and dequeue the entry based at least in part on the timestamp.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] A data center may contain multiple computing platforms connected via network interface devices and switches. Excessive traffic on the data center network can cause congestion and packet loss, resulting in slower traffic. To avoid congested queues of network devices, a network interface device can transmit packets from a data flow through different output ports and different paths using a technique called packet spraying. However, packet spraying can result in out-of-order reception of packets, which can cause packet reordering and delays in delivering packets to a recipient. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows an example system. Fig. Figure 2 shows an example of a system and an operation. Fig. 3A and Fig. 3B shows examples of the order of outputs from different input queues. Fig. Figure 4 shows an example process. Fig. Figure 5 shows an example of a network interface device. Fig. 6A-6C show example switches. Fig. Figure 7 shows an example system. DETAILED DESCRIPTION

[0002] To at least attempt to reduce the out-of-order delivery of packets to an endpoint network interface device, various examples can time-synchronize network interface devices transmitting packets to a switch and configure the switch to output packets in timestamp order, from oldest timestamp to newest timestamp. An endpoint sender or switch can insert a timestamp into a packet before transmitting it to indicate the time of packet transmission. The sender end stations and the switches can participate in the time synchronization described here. When the packet arrives (enters) at the switch, the switch can place the packet into an output queue in time order so that the packets leaving the switch are time-ordered.The switch can perform timestamp-based load balancing of packets received from different input ports and assigned to different queues. The switch can perform packet-by-packet egress port scheduling based on timestamp values ​​for one or more flows instead of, or in addition to, flow-based scheduling.

[0003] A switch that dispatches traffic based on timestamp order can reduce the likelihood of packets arriving out of order at the receiver. In some cases, incast scenarios (e.g., all-to-one or many-to-one traffic patterns) can be mitigated, resulting in lower latency and faster receipt of packets from a flow. Tail latency is a tail of the probability distribution of the packet retrieval latency (also called the latency profile) of the switch fabric. Tail latency refers to the worst-case latencies that occur with a very low probability.

[0004] Fig. 1 shows an example switch. Servers 100-0 through 100-N, where N is an integer, may configure one or more of the respective network interface devices 102-0 through 102-N to insert a transmission timestamp into the packets before transmitting the packets to switch 110, or to prevent one or more of the respective network interface devices 102-0 through 102-N from inserting a transmission timestamp into the packets before transmitting the packets to switch 110. Various examples of servers 100-0 through 100-N are described at least with reference to Fig. 7 described.

[0005] In some examples, switch 110 may include one or more of the following elements: a network interface controller (NIC), an RDMA-capable NIC, a SmartNIC, a router, a switch, a virtual switch, a forwarding element, an infrastructure processing unit (IPU), a data processing unit (DPU), or an edge processing unit (EPU). An edge processing unit (EPU) may be a network interface that uses processors and accelerators (e.g., digital signal processors (DSPs), signal processors or wireless-specific accelerators for virtualized radio access networks (vRANs), cryptographic operations, compression / decompression, etc.). A virtual switch may enable communication between virtual machines on the same server or between different servers.

[0006] In some examples, the network interface devices 102-0 through 102-N may be time synchronized based on time synchronization technologies such as IEEE 1588 Precision Time Protocol (PTP), Peripheral Component Interconnect Express (PCIe) Precision Time Measurement (PTM), a Global Positioning System (GPS) signal, satellites (e.g., Global Navigation Satellite Systems (GNSS)), or other technologies.

[0007] In some examples, Media Access Control (MAC) processors and / or Physical Layer Interfaces (PHYs) of network interface devices 102-0 through 102-N may be configured to insert a transmission timestamp into a packet prior to transmission to switch 110. A network interface device may insert a timestamp into a header field of a packet. Timestamp fields may be inserted into network and / or transport protocol headers transmitted end-to-end, such as an Internet Protocol (IP) header field, an Ethernet header field, or other header fields.

[0008] Network interface devices 102-0 through 102-N can transmit packets of the same or different data flow over different paths using one or more output ports. For example, network interfaces 102-0 through 102-N can transmit packets based on multipath packet transmission technologies such as Internet Engineering Task Force (IETF) Request for Comments (RFC) 6824, "TCP Extensions for Multipath Operation with Multiple Addresses" (2020), or others.

[0009] Switch 110 can be configured as a TOR (Top of Rack), EOR (End of Row), or MOR (Middle of Row) switch. Switch 110 can be positioned as a leaf or spine switch. Switch 110 can receive packets from network interface devices 102-0 through 102-N through one or more downstream ports (DPs) 112-0 through 112-D, where D is an integer. Switch 110 can store received packets in queues 132-0 through 132-Q, where Q is an integer. For example, different queues among queues 132-0 through 132-Q can be associated with different DPs, so that packets from a particular DP are stored in a particular queue.

[0010] At the input, packet processing circuitry 120 can separate received packets with timestamp values ​​from packets without timestamp values, so that a queue of queues 132-0 through 132-Q does not store packets with timestamp values ​​and packets without timestamp values. However, packet processing circuitry 120 can store packets with timestamp values ​​and packets without timestamp values ​​in a queue of queues 132-0 through 132-Q, so that the queue of queues 132-0 through 132-Q stores packets with timestamp values ​​and packets without timestamp values.

[0011] A packet can refer to various formatted collections of bits that can be sent over a network, such as Ethernet frames, Internet Protocol (IP) packets, Transmission Control Protocol (TCP) segments, User Datagram Protocol (UDP) datagrams, Real-time Transport Protocol (RTP) segments, etc. A packet can be associated with a flow. A flow can consist of one or more packets transmitted between two endpoints. A flow can be identified by a set of defined tuples, such as two tuples that identify the endpoints (e.g., source and destination addresses). For some services, flows can be identified at a finer granularity by using five or more tuples (e.g., source address, destination address, IP protocol, transport layer source port, and destination port).

[0012] Based on the arrival of a packet at a DP, the circuitry of switch 110 can queue the packet per port so that packets from a DP are time-ordered for that port. In some cases, packets from a DP are stored in a queue with a specific priority level. For example, one or more of queues 132-0 through 132-Q can store packets from a specific DP and for one or more priority levels.

[0013] An orchestrator or data center administrator may configure packet processing circuitry 120 to output packets from queues 132-0 through 132-Q in timestamp order to prioritize outputting the oldest timestamp value associated with packets in queues 132-0 through 132-Q. Packet processing circuitry 120 may output packets according to timestamp order for traffic flowing from servers 100-0 through 100-N through one or more of ports UP 122-0 through 122-U upstream to a spine layer or switch fabric, where U is an integer. Similarly, for received downstream packets, the packet processing circuit 120 may output packets according to the timestamp order for traffic flowing to one or more of the network interface devices 110-0 through 110-N via one or more of the ports DP 112-0 through 112-D from the fabric to the server direction.

[0014] For example, to determine which packet from queues 132-0 through 132-Q should be output from a port under UP 122-0 through 122-U, packet processing circuitry 120 may select an oldest packet from input queues 132-0 through 132-Q per port. In some cases, packets arriving at a switch port DP may be time-ordered within that port because these packets were transmitted according to the time order.

[0015] Packet processing pipeline 120 may output packets based on timestamp order and on a priority basis, such that packets of higher priority levels are output before lower priority packets according to timestamp order. Packet processing pipeline 120 may output packets based on timestamp order unless another packet with a more recent timestamp has a higher priority level. Packet processing pipeline 120 may output packets such that a higher priority packet among packets with the same timestamp values ​​or timestamp values ​​within a configured timestamp value difference is output before lower priority packets. Packet processing pipeline 120 may output packets with the same timestamp values ​​or timestamp values ​​within a range simultaneously from multiple independent ports.For example, given a packet P1 with a timestamp value of T2 and a high priority level and a packet P2 with a timestamp value of T1 and a low priority level, where T1 is older than T2, the packet processing pipeline 120 may output the packet P1 before P2 if T2 is within a configured timestamp difference from T1.

[0016] If the upstream ports UP of switch 110 are connected to another switch in a fabric, the incoming packets may also be time-ordered. To determine an oldest packet, packet processing circuitry 120 may select a packet with the oldest timestamp (e.g., with a lower timestamp value) at the head of one or more input queues connected to a switch port. If a number of queues or entries in a queue is at or above a configured number (e.g., 64 or 128), packet processing circuitry 120 may select an oldest timestamp from the configured number of queues and / or entries in a queue, even if the oldest timestamp is not the head of the queue.However, packet processing circuitry 120 may schedule packets for output based on receipt of the packets, rather than waiting to schedule output based on reaching a number of entries in a queue.

[0017] Note that timestamp values ​​can be rolled over (reset to zero), and packet processing circuitry 120 can track the last timestamp received from an input port. If a newly received timestamp value is lower than the last received one, packet processing circuitry 120 can detect timestamp rollover and consider the timestamp rollover to determine which timestamp is the oldest for scheduling a packet for output. For example, based on detecting a timestamp, packet processing circuitry 120 can select a packet with a higher value than the older timestamp and schedule that selected packet for output.

[0018] Some examples can increase network utilization by reducing congestion caused by packet spraying. Some examples can reduce congestion caused by selecting egress ports based on flows (e.g., Equal-Cost Multi-Path Routing (ECMP) hashing).

[0019] For example, switch metrics can be inserted into a packet and transmitted with the packet or sent to the endpoint sender or a leaf switch. Switch metrics can include the time a packet spends in a switch hop, and such switch metrics can be used to choose a different network path for packets to reduce the time packets spend in a switch hop. In some examples, switch metrics can be conveyed in metadata of in-band telemetry schemes, such as those described in the following documents: "In-band Network Telemetry (INT) Dataplane Specification, v2.0", P4.org Applications Working Group (February 2020); IETF draft-lapukhov-dataplane-probe-01, "Data-plane probe for in-band telemetry collection" (2016); and IETF draft-ietf-ippm-ioam-data-09, “In-situ Operations, Administration, and Maintenance (IOAM)” (March 8, 2020).In-situ Operations, Administration, and Maintenance (IOAM) records operational and telemetry data within the packet as the packet traverses a path between two points in the network. IOAM discusses the data fields and associated data types for in-situ OAM. In-situ OAM data fields can be encapsulated in a variety of protocols such as NSH, Segment Routing, Geneve, IPv6 (via an extension header), or IPv4.

[0020] In some examples where a time gap between received timestamp values ​​is greater than a configured value, packet processing circuitry 120 may indicate to an endpoint that reordering should occur and / or proactively send a negative acknowledgment (NACK) to a sender to indicate packet loss and prompt retransmission of a packet with such a timestamp gap.

[0021] In some examples, a sorting layer in a network interface device or host software may perform timestamp-based egress of packets.

[0022] Packet processing circuitry 120 may select from UP 122-0 through 122-U to output packets in timestamp order from queues 132-0 through 132-Q based on one or more of the following criteria: lowest load, round-trip, or weighted round-trip. Packet processing circuitry 120 may prioritize the output of packets with timestamps over packets without timestamps. For example, packets without timestamps may be forwarded on a best-effort basis. Packet processing circuitry 120 may schedule the simultaneous output of packets with the same timestamp value over multiple independent ports.

[0023] In some cases where the received packets P1 and P2 are associated with the same flow, were received at the same input port, and P1 has a larger byte size than P2, P2 may arrive at the destination earlier than P1 if smaller packets traverse the network faster and P1 and P2 arrive at the destination out of order. In some cases, packet processing circuitry 120 may schedule the output of P2 to exit at a different output port than P1 after P1 completes exit to reduce the likelihood of P1 and P2 arriving at an endpoint out of order.

[0024] In an all-to-all pattern, the starting point for initiating all-to-all communication could have a skew map that accounts for skews in transmission timing from the start of an all-to-all communication task on a particular node. For example, if four nodes N1, N2, N3, and N4 are communicating, this mapping could specify that if timestamp t is the start time of a communication at a node, the transmission start time of N1 to N2, N3 to N2, and N4 to N2 are delayed relative to each other so that their arrival at N2 is staggered. This can reduce the likelihood of exceeding queue thresholds, which can lead to congestion. The Linux® Earliest Txtime First (ETF) feature can be used to specify these times.

[0025] In some examples, switch 110 or an endpoint receiver network interface device may perform packet reordering based on transmission timestamps. Reordering methods may be applied at an end station (e.g., Amazon Scalable Reliable Datagram (SRD), Data Plane Development Kit (DPDK) reordering technologies, or others).

[0026] In some examples, packet processing circuitry 120 may be implemented as one or more field programmable gate arrays (FPGAs), accelerators, application specific integrated circuits (ASICs), processors, or other circuitry.

[0027] The switch 110 may be implemented as a switch chip (e.g., a switch system on a chip (SoC), one or more tiles, a multi-chip module, and / or a package). In some examples, a switch may be implemented with a physical package containing one or more chips or tiles. In some examples, a physical package may include discrete chips and tiles connected together by a network or other interconnect, as well as an interface and heat dissipation. A chip may include semiconductor devices forming one or more processing devices or other circuits. A tile may include semiconductor devices forming one or more processing devices or other circuits. A physical package may include, for example, one or more chips, a plastic or ceramic package for the chips, and conductive contacts connected to a circuit board.

[0028] Fig. Figure 2 shows an example system and operation. In this example, network interface devices N1, N2, and N3 send packets to three different destination network interface devices N4, N5, and N6 via four ToR switches (T1, T2, T3, and T4) and two spine switches (S1 and S2). Network interface devices N1-N6 can be time-synchronized as described here.

[0029] In some examples, a leaf node switch (e.g., one or more of T1-T4) may insert a transmission timestamp into a packet at the time packets are received and leverage time synchronization with other endpoint sender NICs (N1 through N6) that might also insert timestamps.

[0030] In this example, packets from different flows traverse switch S2 due to packet spray, but packets from a flow destined for network interface device N6 also traverse switch S1. The numbers indicate the order of the timestamps, with a lower number representing an older timestamp value. If packet 9 arrives at switch T3 before packet 8 because packet 8 is older than packet 9, the arbiter in switch T3 outputs packet 8 before outputting packet 9. An arbiter in switch T3 selects the oldest packet received from various input ports for egress and blocks the egress of packets from other input ports until other input ports receive the oldest packet at the top of their queue.

[0031] Fig. Figure 3A shows an example of the order of outputs from different input queues. For two ports of a switch that are active for input, a queue can be associated with an input port and store entries that specify a packet number and the timestamp for the corresponding packet's output. In some examples, the order of output scheduling might look like this: At (1), packet 1 (Pkt1) received at port 1 (Port1) can be output with a timestamp value t=1. At (2), packet 1 (Pkt1) received at port 2 (Port2) can be output with a timestamp value t=1. Note that Pkt1 received at port 1 and Pkt1 received at port 2 can be output from different output ports at the same time or with time overlap.In this example, Pkt1 received on port 1 may have a higher priority than Pkt1 received on port 2, and Pkt1 received on port 1 may be output before Pkt1 received on port 2 even though it has the same timestamp value as Pkt1 received on port 1.

[0032] In (3), packet 2 (Pkt2) received at port 1 (Port1) can be output with a timestamp value of t=2. In (4), packet 3 (Pkt3) received at port 1 (Port1) can be output with a timestamp value of t=3. In (5), packet 2 (Pkt2) received at port 2 (Port2) with a timestamp value of t=3 can be output to different output ports. Note that Pkt3 received at port 1 and Pkt2 received at port 2 can be output from different output ports at the same time or with a time overlap. In this example, Pkt3 received at port 1 can have a higher priority than Pkt2 received at port 2, and Pkt3 received at port 1 can be output before Pkt2 received at port 2, even though it has the same timestamp value as Pkt3 received at port 1.

[0033] In (6), packet 3 (Pkt3) received at port 2 (Port2) can be output with a timestamp value of t=4. In (7), packet 4 (Pkt4) received at port 2 (Port2) can be output with a timestamp value of t=6. In (8), packet 4 (Pkt4) received at port 1 (Port1) can be output with a timestamp value of t=7.

[0034] In (9), packet 5 (Pkt5) received at port 1 (Port1) can be output with a timestamp value of t=8. In (10), packet 5 (Pkt5) received at port 2 (Port2) can be output with a timestamp value of t=8. Note that Pkt5 received at port 1 and Pkt5 received at port 2 can be output from different output ports simultaneously or with a time overlap. In this example, Pkt5 received at port 1 can have a higher priority than Pkt5 received at port 2, and Pkt5 received at port 1 can be output before Pkt5 received at port 2, even though it has the same timestamp value as Pkt5 received at port 1.

[0035] Fig. Figure 3B illustrates an example of the order of packet input and output. Packets may be received at one or more input ports in the order (1) Pt1 with timestamp value t=3, (2) Pt2 with timestamp value t=1, (3) Pt3 with timestamp value t=2, and (4) Pt4 with timestamp value t=4. Based on the timestamp order in the manner described here, a network interface device may output packets from one or more ports in the following order: (1) Pt2 with timestamp value t=1, (2) Pt3 with timestamp value t=2, (3) Pt1 with timestamp value t=3, and (4) Pt4 with timestamp value t=4.

[0036] Based on the order of timestamps and priority, where packet Pkt1 has a higher priority than packet Pkt3 but a newer timestamp, and packet Pkt1 has the same or lower priority than packet Pkt2, a network interface device can output packets from one or more ports in the following order: (1) Pkt2 with timestamp value t=1, (2) Pkt1 with timestamp value t=3, (3) Pkt3 with timestamp value t=2, and (4) Pkt4 with timestamp value t=4.

[0037] Fig. 4 illustrates an example process. The process may be performed by a timestamp arbiter circuit, such as packet processing circuitry in a switch or other network interface device. At 402, a determination may be made as to whether a packet is available for egress. In some examples, a packet may be available for egress if the packet arrives at a particular input port. In some examples, a packet may be available for egress if the packet arrives at a particular input port and is subject to cut-through mode, where the incoming packet is moved to the front of a queue. In some examples, a packet may be available for egress if the packet arrives at a particular input port and has the highest priority. The input queue per port may be associated with packets with egress timestamps received at a particular input port.The input queue per port can be assigned a specific priority level (e.g., high, medium, low) of packets with output timestamps received at a specific input port. If a packet is available for output, the process can continue to 404. If a packet is not available for output, the process can be repeated to 402.

[0038] At 404, it can be determined whether the packet is at the head of a per-port input queue. For example, a per-port input queue head might be a packet that is next in line to be output based on first-in-first-out ordering. Based on the fact that the packet is next in the queue, the process can continue with 406. Based on the fact that the packet is not next in line to be output based on first-in-first-out ordering, the process can continue with 420.

[0039] At 406, it can be determined whether the packet is associated with an oldest timestamp among other packets in other input queues per port, or whether it is subject to a cut-through. Based on whether the packet is associated with an oldest timestamp among other packets in other input queues per port, has the highest priority, or is subject to a cut-through, the process can continue with 408. Since the packet is not associated with an oldest timestamp among other packets in other input queues per port, is not among the highest priority packets, and is not subject to a cut-through, the process can continue with 404.

[0040] At 408, the packet may be removed from the queue and not associated with the respective port's ingress queue; it may be marked to be scheduled for egress next. For example, the packet may be associated with an available, valid egress port for egress. A valid egress port may be available for use by a packet routing scheme. The process may return to 404.

[0041] At 420, the packet can be placed in a per-port inbound queue or remain in the per-port inbound queue. The process can return to 404.

[0042] Fig. 5 illustrates an example of a network interface device. In some examples, the processors and / or FPGAs 530 may be configured to insert timestamps into packets prior to transmission or to receive (and possibly reorder) time-ordered packets, as described herein. The network interface device may be configured as an endpoint receiver that has multiple homes and receives time-ordered packets, possibly reordering packets and merging packet contents before copying packet contents to a host. Some examples of network interfaces 500 are part of, or used by, an infrastructure processing unit (IPU) or data processing unit (DPU). An xPU may refer to at least one of an IPU, DPU, graphics processing unit (GPU), general-purpose GPU (GPGPU), or other processing units (e.g., accelerators).An IPU or DPU may include a network interface with one or more programmable pipelines or fixed function processors to offload operations that could have been performed by a CPU. The IPU or DPU may include one or more memory devices. In some examples, the IPU or DPU may perform virtual switching, manage memory transactions (e.g., compression, cryptography, virtualization), and manage operations running on other IPUs, DPUs, servers, or devices.

[0043] Network interface 500 may include a transceiver 502, processors 504, a transmit queue 506, a receive queue 508, memory 510, an interface 512, and a DMA module 514. Transceiver 502 may receive and transmit packets in accordance with applicable protocols, such as Ethernet as described in IEEE 802.3, although other protocols may also be used. Transceiver 502 may receive and transmit packets to and from a network over a network medium (not shown). Transceiver 502 may include PHY circuitry 504 and Media Access Control (MAC) circuitry 505. PHY circuitry 504 may include encoding and decoding circuitry (not shown) to encode and decode data packets according to applicable physical layer specifications or standards.MAC circuitry 505 can be configured to perform MAC address filtering on received packets, process MAC headers of received packets by verifying data integrity, remove preambles and padding, and prepare packet contents for processing by higher layers. MAC circuitry 516 can be configured to assemble the data to be transmitted into packets containing destination and source addresses along with network control information and error detection hashes.

[0044] As described here, PHY 504 and / or MAC 505 can be configured to insert timestamps into packets before transmission or to output packets according to the timestamp values ​​and priority level as described here.

[0045] Processors 530 may be one or more of the following combinations: a processor, a core, a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other programmable hardware device that enables programming of network interface 500. For example, an "intelligent network interface" or SmartNIC may provide packet processing functions in the network interface using processors 530.

[0046] Processors 530 may include a programmable processing pipeline or offload circuitry programmable with P4, Software for Open Networking in the Cloud (SONiC), Broadcom® Network Programming Language (NPL), NVIDIA® CUDA®, NVIDIA® DOCA™, Data Plane Development Kit (DPDK), OpenDataPlane (ODP), Infrastructure Programmer Development Kit (IPDK), eBPF, x86-compatible executable binaries, or other executable binaries. A programmable processing pipeline may include one or more Match Action Units (MAUs) configured based on a programmable pipeline language instruction set. Processors, FPGAs, other specialized processors, controllers, devices, and / or circuitry may be used for packet processing or packet modification. Ternary content-addressable memory (TCAM) can be used for parallel matching or lookup operations on packet header contents.

[0047] Packet arbiter 524 may divide received packets for processing by multiple CPUs or cores using Receive Side Scaling (RSS). When packet arbiter 524 uses RSS, packet arbiter 524 may compute a hash or make another determination based on the contents of a received packet to determine which CPU or core should process a packet.

[0048] Interrupt Coalesce 522 can perform interrupt moderation, where Interrupt Coalesce 522 waits until multiple packets arrive or a timeout expires before generating an interrupt to the host system to process the received packets. Receive Segment Coalescing (RSC) can be performed by network interface 500, where portions of the incoming packets are combined into a single aggregate packet. Network interface 500 makes this aggregate packet available to an application.

[0049] The DMA engine 514 may copy a packet header, packet payload, and / or descriptor directly from host memory to the network interface or vice versa, rather than copying the packet to an intermediate buffer on the host and then performing another copy operation from the intermediate buffer to the destination buffer.

[0050] In some examples, processors 530 may be configured to perform packet reordering to reorder packets according to timestamp values ​​and / or packet sequence numbers before copying the packets to a host system via interface 512.

[0051] Memory 510 may be volatile and / or non-volatile memory and may store any queues or instructions used to program network interface 500. The transmit traffic manager may schedule the transmission of packets from transmit queue 506. Transmit queue 506 may contain data or references to data to be transmitted over the network interface. Receive queue 508 may contain data or references to data received from a network over the network interface. Descriptor queues 520 may contain descriptors that reference data or packets in transmit queue 506 or receive queue 508. Interface 512 may be an interface to the host device (not shown).For example, the interface 512 may be compatible with or at least partially based on PCI, PCIe, PCI-x, Serial ATA and / or USB (although other connection standards may also be used), or even proprietary variants thereof.

[0052] Fig. 6A illustrates an example of a network interface. The host 600 may include processors, memory devices, device interfaces, and other circuitry as described herein. The processors of the host 600 may execute software such as applications (e.g., microservices, virtual machines (VMs), microVMs, containers, processes, threads, or other virtualized execution environments), operating systems, and device drivers. An operating system or device driver may configure the network interface device or packet processing device 610 to use one or more control planes to communicate with the software-defined networking (SDN) controller 650 over a network and to configure the operation of the one or more control planes.

[0053] Packet processing device 610 may include multiple compute complexes, such as an acceleration compute complex (ACC) 620 and a management compute complex (MCC) 630, as well as packet processing circuitry 640 and network interface technologies for communicating with other devices over a network. ACC 620 may be implemented as one or more of the following: a microprocessor, processor, accelerator, field programmable gate array (FPGA), application-specific integrated circuit (ASIC), or circuitry at least described herein. Similarly, MCC 630 may be implemented as one or more of the following: a microprocessor, processor, accelerator, field programmable gate array (FPGA), application-specific integrated circuit (ASIC), or circuitry described herein.In some examples, ACC 620 and MCC 630 may be implemented as separate cores in a CPU, as different cores in different CPUs, as different processors in the same integrated circuit, or as different processors in different integrated circuits.

[0054] Packet processing device 610 may be implemented as one or more of the following: a microprocessor, processor, accelerator, FPGA (Field Programmable Gate Array), application-specific integrated circuit (ASIC), or circuits described herein. Packet processing pipeline circuitry 640 may process packets according to the instructions or configuration of one or more control planes executed by multiple compute complexes. In some examples, ACC 620 and MCC 630 may execute respective control planes 622 and 632.

[0055] Packet processing device 610, ACC 620, and / or MCC 630 may be configured to insert timestamps into packets prior to transmission or output packets according to the timestamp values ​​and priority level, as described herein.

[0056] SDN controller 642 may update or reconfigure software executing on ACC 620 (e.g., control plane 622 and / or control plane 632) through the contents of packets received via packet processing device 610. In some examples, ACC 620 may execute a control plane operating system (OS) (e.g., Linux) and / or a control plane application 622 (e.g., user-space or kernel modules) used by SDN controller 642 to configure the operation of packet processing pipeline 640. The control plane application 622 may include generic flow tables (GFT), ESXi, NSX, and Kubernetes control plane software, application software for managing crypto configurations, a programming protocol-independent packet processor (P4) runtime daemon, a target-specific daemon, Container Storage Interface (CSI) agents, or Remote Direct Memory Access (RDMA) configuration agents.

[0057] In some examples, the SDN controller 642 can communicate with ACC 620 using a remote procedure call (RPC) such as Google Remote Procedure Call (gRPC) or another service, and ACC 620 can convert the request into a target-specific protocol buffer request (protobuf) to MCC 630. gRPC is a remote procedure call solution based on data packets sent between a client and a server. Although gRPC is one example, other communication schemes can also be used, such as Java Remote Method Invocation, Modula-3, RPyC, Distributed Ruby, Erlang, Elixir, Action Message Format, Remote Function Call, Open Network Computing RPC, JSON-RPC, and so on.

[0058] In some examples, the SDN controller 642 may provide packet processing rules for the performance of ACC 620. For example, ACC 620 may program table rules (e.g., header field matching and corresponding action) that are applied by the packet processing pipeline circuitry 640 based on policy changes and changes in VMs, containers, microservices, applications, or other processes. ACC 620 may be configured to provide network policies as flow cache rules in a table to configure the operation of the packet processing pipeline 640. For example, the control plane application 622 executed by ACC may configure rule tables that are applied by the packet processing pipeline circuitry 640 with rules to define a traffic destination based on packet type and content. ACC 620 may program table rules (e.g.,Match-Action) into memory that the packet processing pipeline circuit 640 can access based on changes in policy and changes in VMs.

[0059] For example, ACC 620 may implement a virtual switch such as vSwitch or Open vSwitch (OVS), Stratum, or Vector Packet Processing (VPP), which enables communication between virtual machines running on host 600 or with other devices connected to a network. For example, ACC 620 may configure packet processing pipeline circuitry 640 to determine which VM should receive traffic and what type of traffic a VM can transmit. For example, packet processing pipeline circuitry 640 may implement a virtual switch such as vSwitch or Open vSwitch, which enables communication between the virtual machines running on host 600 and packet processing device 610.

[0060] MCC 630 may execute a host management control plane and a global resource manager and configure hardware registers. Control plane 632, executed by MCC 630, may provision and configure packet processing circuitry 640. For example, a VM running on host 600 may use packet processing device 610 to receive or send packet traffic. MCC 630 may execute boot, power, management, and manageability software (SW) or firmware (FW) code to start and initialize packet processing device 610, manage device power consumption, provide connectivity to the baseboard management controller (BMC), and perform other operations.

[0061] One or both control planes of ACC 620 and MCC 630 can define the contents of the traffic routing table and the network topology applied by packet processing circuitry 640 to select a path for a packet in a network to a next hop or to a destination device connected to the network. For example, a VM executing on host 600 can use packet processing device 610 to receive or send packet traffic.

[0062] ACC 620 may execute control drivers to communicate with MCC 630. At a minimum, to provide a configuration and provisioning interface between control planes 622 and 632, communication interface 625 may enable control plane-to-control plane communication. Control plane 632 may perform a gatekeeper operation to configure shared resources.For example, via the communication interface 625, the ACC control plane 622 can communicate with the control plane 632 to achieve one or more of the following goals: determining hardware capabilities, accessing the data plane configuration, reserving hardware resources and configuration, communicating between ACC and MCC via interrupts or polling, subscribing to receive hardware events, performing indirect reading and writing of hardware registers for debugging, configuring flash and physical interfaces (PHY), or performing system provisioning for various network interface device deployments, such as: storage nodes, tenant hosting nodes, microservices backend, compute nodes, or others.

[0063] The communication interface 625 may be used by a negotiation and configuration protocol running between the ACC control plane 622 and the MCC control plane 632. The communication interface 625 may include a general-purpose mailbox for various operations performed by the packet processing circuit 640. Examples of operations of the packet processing circuit 640 include issuing non-volatile memory Express (NVMe) reads or writes, issuing non-volatile memory Express over Fabrics (NVMe-oF™) reads or writes, Lookaside Crypto Engine (LCE) (e.g., compression or decompression), Address Translation Engine (ATE) (e.g.,Input-Output Memory Management Unit (IOMMU) for virtual to physical address translation), encryption or decryption, configuration as a storage node, configuration as a tenant hosting node, configuration as a compute node, provisioning several different types of services between different Peripheral Component Interconnect Express (PCIe) endpoints, or others.

[0064] The communication interface 625 may include one or more mailboxes accessible as registers or memory addresses. For communication from the control plane 622 to the control plane 632, the communication from the control drivers 624 may be written to the one or more mailboxes. For communication from the control plane 632 to the control plane 622, the communication may be written to one or more mailboxes. Messages written to mailboxes may include descriptors containing the message opcode, message errors, message parameters, and other information. Messages written to mailboxes may include defined format messages that transfer data.

[0065] Communications interface 625 may enable communication based on writing or reading to specific memory addresses (e.g., dynamic random access memory (DRAM)), registers, or other mailboxes to which instructions and data are written and read. To ensure secure communication between control planes 622 and 632, registers and memory addresses (and memory address translations) for communication may only be available for writing or reading by control planes 622 and 632 or the cloud service provider (CSP) software executing on ACC 620 and the device manufacturer's software, embedded software, or firmware executing on MCC 630. Communications interface 625 may support communication between multiple different computing complexes, e.g.,between Host 600 and MCC 630, Host 600 and ACC 620, MCC 630 and ACC 620 or others.

[0066] Packet processing circuitry 640 may be implemented using one or more of the following: application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), processors executing software, or other circuitry. Control plane 622 and / or 632 may configure packet processing pipeline circuitry 640 or other processors to perform operations related to NVMe, NVMe-oF reads or writes, Lookaside Crypto Engine (LCE), Address Translation Engine (ATE), Local Area Network (LAN), compression / decompression, encryption / decryption, or other accelerated operations.

[0067] Various message formats can be used to configure the ACC 620 or MCC 630. In some examples, a P4 program can be compiled and provided to the MCC 630 to configure the packet processing circuitry 640 to output packets according to the timestamp values ​​and priority level, as described herein.

[0068] Fig. 6B illustrates an example switch. Various examples may be used in or with the switch to insert timestamps into packets prior to transmission or to output packets according to the timestamp values ​​and priority level as described herein. The switch 654 may forward packets or frames of any format or specification from any port 652-0 through 652-X to any of the ports 656-0 through 656-Y (or vice versa). One or more of the ports 652-0 through 652-X may be connected to a network of one or more interconnected devices. Similarly, any of the ports 656-0 through 656-Y may be connected to a network of one or more interconnected devices.

[0069] In some examples, switch fabric 660 may provide for forwarding packets from one or more input ports for processing prior to output from switch 654. Switch fabric 660 may be implemented as one or more multi-hop topologies, with example topologies including torus, butterflies, buffered multi-stage, etc., or shared memory switch fabric (SMSF), among other implementations. SMSF may be any switch fabric connected to input and output ports in the switch, with the input subsystems writing (storing) packet segments to the fabric's memory, while the output subsystems reading (fetching) packet segments from the fabric's memory.

[0070] Memory 658 may be configured to store packets received at the ports before they are output from one or more ports. Packet processing pipelines 662 may include incoming and outgoing packet processing circuitry to process incoming packets and output packets, respectively.

[0071] In some examples, packet processing pipelines 662 may include a parser, an input pipeline, a buffer, a scheduler, and output pipelines. In some examples, packet processing pipelines 662 may perform traffic manager 663 operations.

[0072] Packet processing pipelines 662 may determine which port to transmit packets or frames to based on a table that maps packet characteristics to a corresponding output port. Packet processing pipelines 662 may be configured to perform matching actions on received packets to identify packet processing rules and next hops using information stored in ternary content-addressable memory (TCAM) tables or, in some examples, exact-match tables. For example, match-action tables or circuits may be used, which use a hash of a portion of a packet as an index to find an entry (e.g., a forwarding decision based on the contents of a packet header). Packet processing pipelines 662 may implement access control lists (ACLs) or packet loss due to queue overflow.Packet processing pipelines 662 can be configured to insert timestamps into packets prior to transmission or to output packets according to the timestamp values ​​and priority level, as described herein. The operation of packet processing pipelines 662, including their data plane, can be configured using P4, C, Python, Broadcom Network Programming Language (NPL), or x86-compatible executable binaries or other executable binaries. Processors 666 and FPGAs 668 can be used for packet processing or modification.

[0073] Traffic manager 663 can perform hierarchical scheduling, rate adjustment, and packet transmission measurement from one or more packet queues. Traffic manager 663 can perform congestion management such as flow control, Congestion Message Generation and Reception (CNM), Priority Flow Control (PFC), and others. Network interface circuitry and software described herein can be used by switch 650, including a MAC and SerDes.

[0074] Fig. 6C shows an example of a switch. Various examples can be used in or with the switch to output packets according to the timestamp values ​​and priority level, as described herein. The switch 680 can include a network interface 682 that can provide a consistent Ethernet interface. The network interface 682 can support 25 GbE, 50 GbE, 100 GbE, 200 GbE, and 400 GbE Ethernet interfaces. The cryptographic circuitry 684 can perform at least Media Access Control Security (MACsec) or Internet Protocol Security (IPSec) decryption for received packets or encryption for packets to be transmitted.

[0075] Different circuits can perform one or more of the following tasks: service measurement, packet counting, operations, administration and management (OAM), protection module, instrumentation and telemetry, and clock synchronization (e.g., based on IEEE 1588).

[0076] Database 686 may store a device's profile to configure the operation of switch 680. Memory 688 may include high bandwidth memory (HBM) for packet buffering. Packet processor 690 may perform one or more of the following functions: next hop decision related to packet forwarding, packet counting, access list operations, bridging, routing, Multiprotocol Label Switching (MPLS), Virtual Private LAN Service (VPLS), L2VPNs, L3VPNs, OAM, Data Center Tunneling Encapsulations (e.g., VXLAN and NV-GRE), or others. Packet processor 690 may include one or more FPGAs. Buffer 694 may store one or more packets. Traffic Manager (TM) 692 may provide per-subscriber bandwidth guarantees in accordance with service level agreements (SLAs) and provide hierarchical quality of service (QoS).The fabric interface 696 may include a serializer / de-serializer (SerDes) and provide an interface to a switch fabric.

[0077] The functions of the components of switches of the example devices described herein can be combined, and components of the switches described herein can be incorporated into other example switches of the example devices described herein. For example, the components of the example switches described herein can be implemented in a switch system on a chip (SoC) that includes at least one interface to other circuitry in a switch system. A switch SoC can be coupled to other devices in a switch system, such as input or output ports, memory devices, or host interface circuitry.

[0078] Fig.7 illustrates a system. In some examples, the circuitry of system 700 may configure network interface device 750 to insert timestamps into packets prior to transmission or to output packets according to the timestamp values ​​and priority level, as described herein. System 700 includes a processor 710 that provides processing, operational management, and instruction execution for system 700. Processor 710 may include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), XPU, processing core, or other processing hardware to provide processing for system 700, or a combination of processors. An XPU may include one or more of the following: a CPU, a graphics processing unit (GPU), a general-purpose GPU (GPGPU), and / or other processing units (e.g., accelerators or FPGAs with programmable or fixed functions).The processor 710 controls the overall operation of the system 700 and may be or include one or more programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or the like, or a combination of such devices.

[0079] In one example, system 700 includes an interface 712 connected to processor 710, which may provide a higher-speed or high-throughput interface for system components requiring higher-bandwidth connections, such as memory subsystem 720 or graphics interface components 740, or accelerators 742. Interface 712 represents interface circuitry that may be a standalone component or integrated into a processor chip. If present, graphics interface 740 provides an interface to graphics components to provide a visual representation to the user of system 700. In one example, graphics interface 740 generates a display based on the data stored in memory 730 or based on operations performed by processor 710, or both.In one example, the graphics interface 740 generates a display based on the data stored in the memory 730 or based on the operations performed by the processor 710, or both.

[0080] Accelerators 742 may be a programmable or fixed function offload engine that a processor 710 can access or utilize. For example, an accelerator among accelerators 742 may provide data compression (DC) capabilities, cryptographic services such as public key encryption (PKE), ciphering, hashing / authentication capabilities, decryption, or other capabilities or services. In some cases, accelerators 742 may be integrated into a CPU socket (e.g., into a port on a motherboard or circuit board containing a CPU and providing an electrical interface with the CPU).Accelerators 742 may include, for example, a single- or multi-core processor, a graphics processing unit, a logical execution unit with a single- or multi-level cache, functional units for independent execution of programs or threads, application-specific integrated circuits (ASICs), neural network processors (NNPs), programmable control logic, and programmable processing elements such as field-programmable gate arrays (FPGAs). Accelerators 742 may provide multiple neural networks, CPUs, processor cores, general-purpose graphics processing units, or graphics processing units for use by artificial intelligence (AI) or machine learning (ML) models.For example, the AI ​​model may use or include any one or a combination of the following: a reinforcement learning scheme, a Q-learning scheme, deep Q-learning, or asynchronous advantage actor-critic (A3C), a combinatorial neural network, a recurrent combinatorial neural network, or another AI or ML model. Multiple neural networks, processor cores, or graphics processing units may be made available to AI or ML models to perform learning and / or inference operations.

[0081] The memory subsystem 720 represents the main memory of the system 700 and provides storage space for the code to be executed by the processor 710, or for data values ​​used in the execution of a routine. The memory subsystem 720 may include one or more memory devices 730, such as read-only memory (ROM), flash memory, one or more types of random access memory (RAM) such as DRAM, or other memory devices, or a combination of such devices. The memory 730 stores and hosts, among other things, the operating system (OS) 732 to provide a software platform for executing instructions in the system 700. In addition, applications 734 may be executed on the software platform of the operating system 732 from the memory 730. The applications 734 are programs that have their own operating logic to perform one or more functions.Processes 736 are agents or routines that provide support functions for the operating system 732 or one or more applications 734, or a combination thereof. The operating system 732, the applications 734, and the processes 736 provide software logic to provide functions for the system 700. In one example, the memory subsystem 720 includes a memory controller 722 that generates and issues instructions for the memory 730. It should be understood that the memory controller 722 may be a physical part of the processor 710 or a physical part of the interface 712. The memory controller 722 may, for example, be an integrated memory controller integrated into circuitry with the processor 710.

[0082] Applications 734 and / or processes 736 may instead or additionally refer to a virtual machine (VM), a container, a microservice, a processor, or other software. Various examples described herein may execute an application composed of microservices, where a microservice runs in its own process and communicates via protocols (e.g., application program interface (API), Hypertext Transfer Protocol (HTTP) Resource API, messaging service, Remote Procedure Calls (RPC), or Google RPC (gRPC)). Microservices may communicate with each other via a service mesh and may run in one or more data centers or edge networks. Microservices may be deployed independently of each other, with management of these services performed centrally. The management system may be written in various programming languages ​​and use various data storage technologies.A microservice can be characterized by one or more of the following features: polyglot programming (e.g., code written in multiple languages ​​to provide additional functionality and efficiency not available in a single language), or lightweight deployment in containers or virtual machines and decentralized continuous delivery of microservices.

[0083] In some examples, the 732 operating system may be Linux®, FreeBSD, Windows® Server or Personal Computer, FreeBSD®, Android®, macOS®, iOS®, VMware vSphere, openSUSE, RHEL, CentOS, Debian, Ubuntu, or another operating system. The operating system and driver may run on a processor sold or developed by Intel®, ARM®, AMD®, Qualcomm®, IBM®, Nvidia®, Broadcom®, and Texas Instruments®, among others.

[0084] In some examples, the operating system 732, a system administrator, and / or an orchestrator may enable or disable the network interface 750 to insert timestamps into packets prior to transmission or to output packets according to the timestamp values ​​and priority level, as described herein.

[0085] Although not specifically shown, system 700 may include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, interface buses, or others. Buses or other signal lines may communicatively or electrically couple components together, or may couple components both communicatively and electrically. Buses may include physical communication lines, point-to-point connections, bridges, adapters, controllers, or other circuitry, or a combination thereof. Buses include, for example, one or more system buses, a Peripheral Component Interconnect (PCI) bus, a HyperTransport (ISA) bus, an Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB) bus, or an Institute of Electrical and Electronics Engineers (IEEE) 1394 (Firewire) bus.

[0086] In one example, system 700 includes interface 714, which is connectable to interface 712. In one example, interface 714 represents interface circuitry that may include discrete components and integrated circuits. In one example, multiple user interface components or peripheral components, or both, are connected to interface 714. Network interface 750 provides system 700 with the ability to communicate with remote devices (e.g., servers or other computing devices) over one or more networks. Network interface 750 may include an Ethernet adapter, wireless interconnect components, cellular network interconnect components, Universal Serial Bus (USB), or other wired or wireless standards-based or proprietary interfaces.The network interface 750 may transmit data to a device located in the same data center or rack, or to a remote device, which may include sending data stored in memory. The network interface 750 may receive data from a remote device, which may include storing the received data in memory. In some examples, the packet processing device or network interface device 750 may refer to one or more of the following devices: a network interface controller (NIC), a remote direct memory access (RDMA)-capable NIC, a SmartNIC, a router, a switch, a forwarding element, an infrastructure processing unit (IPU), or a data processing unit (DPU). An example IPU or DPU is described herein.

[0087] In one example, system 700 includes one or more input / output (I / O) interfaces 760. I / O interface 760 may include one or more interface components through which a user interacts with system 700. Peripheral interface 770 may include any hardware interface not explicitly mentioned above. Peripherals generally refer to devices that connect dependently on system 700.

[0088] In one example, system 700 includes a memory subsystem 780 for storing data in non-volatile form. In one example, in certain system implementations, at least certain components of memory 780 may overlap with components of memory subsystem 720. Memory subsystem 780 includes storage device(s) 784, which may be or include any conventional medium for storing large amounts of data in non-volatile form, such as one or more magnetic, solid-state, or optical-based disks, or a combination thereof. Memory 784 contains code or instructions and data 786 in a persistent state (i.e., the value persists despite interruption of power to system 700). Memory 784 may be generally considered "memory," although memory 730 is typically the execution or operational memory that provides instructions to processor 710.While memory 784 is non-volatile, memory 730 may include volatile memory (i.e., the value or state of the data is indeterminate if power to system 700 is interrupted). In one example, memory subsystem 780 includes a controller 782 as an interface to memory 784. In one example, controller 782 is a physical part of interface 714 or processor 710, or may include circuitry or logic in both processor 710 and interface 714.

[0089] Volatile memory may include memory whose state (and thus the data stored therein) is indeterminate when the device's power supply is interrupted. Non-volatile memory (NVM) may include memory whose state is retained even when the device's power supply is interrupted.

[0090] In some examples, system 700 may be implemented using interconnected computing platforms with processors, memories, storage units, network interfaces, and other components. High-speed connections may be used, such as Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Peripheral Component Interconnect express (PCIe), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omni-Path, Compute Express Link (CXL), HyperTransport, High-Speed ​​Fabric, NVLink, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, OpenCAPI, Gen-Z, Infinity Fabric (IF), Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and variants thereof. Data can be copied or stored on virtualized storage nodes, or accessed using a protocol such as NVMe over Fabrics (NVMe-oF) or NVMe (e.g.a Non-Volatile Memory Express (NVMe) device may operate in compliance with the Non-Volatile Memory Express (NVMe) Specification, Revision 1.3c, published on May 24, 2018 (“NVMe Specification”) or derivatives or variations thereof).

[0091] In one example, system 700 may be implemented using interconnected computing platforms with processors, memory, storage media, network interfaces, and other components. High-speed interconnects such as PCIe, Ethernet, or optical interconnects (or a combination thereof) may be used.

[0092] These examples can be implemented in various types of computing and networking equipment, such as switches, routers, racks, and blade servers, as used in data centers and / or server farms. The servers used in data centers and server farms consist of arranged server configurations such as rack-based servers or blade servers. These servers are interconnected through various networking facilities, such as dividing groups of servers into local area networks (LANs) with appropriate switching and routing facilities between the LANs to form a private intranet. For example, cloud hosting facilities typically utilize large data centers with numerous servers. A blade consists of a separate computing platform configured to perform server-like functions—i.e., a "server on a card."Accordingly, a blade contains components common to traditional servers, including a motherboard that has internal wiring (e.g., buses) for coupling appropriate integrated circuits (ICs) and other components mounted on the board.

[0093] Various examples may be implemented using hardware elements, software elements, or a combination of both. In some examples, hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, ASICs, PLDs, DSPs, FPGAs, memory units, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some examples, software elements may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.The decision whether to implement an example using hardware elements and / or software elements may vary depending on any number of factors, such as the desired computational performance, power consumption, thermal tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints as desired for a particular implementation. A processor may be one or more combinations of a hardware state machine, digital control logic, a central processing unit, or any of the hardware, firmware, and / or software elements.

[0094] Some examples may be implemented with or as an article of manufacture or with at least one computer-readable medium. A computer-readable medium may include a non-transferable storage medium for storing logic. In some examples, the non-transitory storage medium may comprise one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on.In some examples, the logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, API, instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.

[0095] According to some examples, a computer-readable medium may include a non-transitory storage medium for storing or maintaining instructions that, when executed by a machine, computing device, or system, cause the machine, computing device, or system to perform methods and / or operations consistent with the described examples. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. The instructions may be implemented according to a predefined computer language, manner, or syntax to instruct a machine, computing device, or system to perform a particular function. The instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.

[0096] One or more aspects of at least one example may be implemented by representative instructions stored on at least one machine-readable medium representing various logic within the processor that, when read by a machine, computing device, or system, cause the machine, computing device, or system to generate logic to perform the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and provided to various customers or manufacturing facilities for loading into the manufacturing machines that actually manufacture the logic or processor.

[0097] The phrases "an example" or "an example" do not necessarily all refer to the same example or embodiment. Any aspect described herein may be combined with any other aspect described herein or with any similar aspect, regardless of whether the aspects are described with respect to the same figure or element. The partitioning, omission, or inclusion of the block functions illustrated in the accompanying figures does not imply that the hardware components, circuits, software, and / or elements for implementing those functions have necessarily been partitioned, omitted, or included in embodiments.

[0098] Some examples can be described using the terms "coupled" and "connected," and their derivatives. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact. However, the term "coupled" can also mean that two or more elements are not in direct contact but still cooperate or interact.

[0099] The terms "first," "second," and the like herein do not denote order, quantity, or importance, but are used to distinguish one element from another. The terms "a" and "an" herein do not imply a limitation of quantity, but denote the presence of at least one of the mentioned elements. The term "asserted," as used herein with reference to a signal, denotes a state of the signal in which the signal is active, which can be achieved by applying any logic level, either logic 0 or logic 1, to the signal. The terms "follow" or "after" can refer to immediately following or following one or more other events. In alternative embodiments, other sequences of operations may be performed. Furthermore, additional operations may be added or removed depending on the application.Any combination of changes may be used, and one skilled in the art, having the benefit of this disclosure, will understand the many variations, modifications, and alternative embodiments.

[0100] Disjunctive expressions such as “at least one of X, Y, or Z,” unless explicitly stated otherwise, are generally understood in context to mean that an item, term, etc., can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Such disjunctive wording is not intended to, and should not, imply that at least one of X, at least one of Y, or at least one of Z must be present in particular embodiments. Furthermore, conjunctive expressions such as “at least one of X, Y, and Z,” unless explicitly stated otherwise, should also be understood to mean X, Y, Z, or any combination thereof, including “X, Y, and / or Z.”

[0101] The following are illustrative examples of the devices, systems, and methods disclosed herein. An embodiment of the devices, systems, and methods may include any one or more, and any combination, of the examples described below.

[0102] Example 1 includes one or more examples and includes an apparatus comprising: a network interface device comprising: an interface to a port; and circuitry configured to: receive a first packet comprising a timestamp associated with a previous or original transmission of the first packet by a sender network interface device; enqueue an entry for the first packet; and dequeue the entry based at least in part on the timestamp.

[0103] Example 2 includes one or more examples where the circuit is to output the first packet associated with the queue and a second packet associated with a second queue based on timestamps associated with the first packet and the second packet, wherein the first packet and the second packet are received by the network interface device from a plurality of network interface devices.

[0104] Example 3 includes one or more examples where the first packet and the second packet are connected to the same flow.

[0105] Example 4 includes one or more examples where the network interface device includes a switch chip.

[0106] Example 5 includes one or more examples where the circuit outputs a second packet before the first packet based on a priority level of the second packet.

[0107] Example 6 includes one or more examples where the circuit outputs the packet and the second packet over multiple ports using a multipath protocol.

[0108] Example 7 includes one or more examples where the network interface device is to receive the packet and the second packet in multiple input ports, and where the multiple network interface devices are time-synchronized.

[0109] Example 8 includes one or more examples, wherein the network interface device is to receive the packet and the second packet in a single input port from the plurality of network interface devices, and wherein the plurality of network interface devices are time-synchronized.

[0110] Example 9 includes one or more examples where the circuit is to store packets with associated timestamps in a first queue and packets without associated timestamps in a second queue.

[0111] Example 10 includes one or more examples and includes at least one input port and at least one output port, wherein the at least one input port and the at least one output port are connected to the interface.

[0112] Example 11 includes one or more examples and includes at least one non-transitory computer-readable medium comprising instructions stored thereon that, when executed by a switch circuit, cause the switch circuit to output a packet associated with one queue and a second packet associated with a second queue based on timestamps associated with the packet and the second packet, wherein the packet and the second packet are received by a plurality of network interface devices, and wherein the timestamps associated with the packet and the second packet indicate transmission of the packet and the second packet.

[0113] Example 12 includes one or more examples and includes instructions stored thereon that, when executed by a switch circuit, cause the switch circuit to output the packet associated with the queue and a second packet associated with a second queue based on timestamps associated with the packet and the second packet, wherein the packet and the second packet are received by the switch circuit from a plurality of network interface devices.

[0114] Example 13 includes one or more examples where the packet and the second packet are connected to the same flow.

[0115] Example 14 includes one or more examples and includes instructions stored thereon that, when executed by a switch circuit, cause the switch circuit to output a second packet before the first packet based on a priority level of the second packet.

[0116] Example 15 includes one or more examples and includes instructions stored thereon that, when executed by a switch circuit, cause the switch circuit to output the packet and the second packet over multiple ports using a multipath protocol.

[0117] Example 16 includes one or more examples and includes a method comprising: a switch performing: outputting a packet associated with one queue and a second packet associated with a second queue based on timestamps associated with the packet and the second packet, wherein the packet and the second packet are received by the switch from a plurality of network interface devices.

[0118] Example 17 includes one or more examples and includes outputting a second packet before the packet based on a priority level of the second packet.

[0119] Example 18 includes one or more examples and includes outputting the packet and the second packet over multiple ports using a multipath protocol.

[0120] Example 19 includes one or more examples and includes receiving the packet and the second packet in a plurality of input ports, wherein the plurality of network interface devices are time synchronized.

[0121] Example 20 includes one or more examples and includes storing packets with associated timestamps in a first queue and storing packets without associated timestamps in a second queue.

[0122] Example 21 includes one or more examples and includes a network interface device having a physical layer interface (PHY) or media access control (MAC) circuit that inserts a timestamp value into a packet.

[0123] Example 22 includes one or more examples and includes a switch that inserts a timestamp value into a packet.

[0124] Example 23 includes the time synchronization of network interface devices that insert timestamps into packets based on one or more of the following protocols: IEEE 1588 Precision Time Protocol (PTP), Peripheral Component Interconnect Express (PCIe) Precision Time Measurement (PTM), a Global Positioning System (GPS) signal, a satellite signal (e.g., Global Navigation Satellite Systems (GNSS)), or other technologies.

Claims

[1] Device comprising: a network interface device comprising: an interface to a port; and a circuit designed for the following: receiving a first packet containing a timestamp associated with a previous or original transmission of the first packet by a sender network interface device; Putting an entry for the first packet into a queue; and Remove the entry from the queue based at least in part on the timestamp. [2] The apparatus of claim 1, wherein the circuit outputs the first packet associated with the queue and outputs a second packet associated with a second queue based on timestamps associated with the first packet and the second packet, the first packet and the second packet being received by the network interface device from a plurality of network interface devices. [3] The apparatus of claim 2, wherein the first packet and the second packet are connected to the same flow. [4] The apparatus of claim 1, wherein the network interface device comprises a switch chip. [5] The apparatus of claim 2, wherein the circuit outputs the second packet before the first packet based on a priority level of the second packet. [6] The apparatus of claim 2, wherein the circuit is to output the first packet and the second packet over multiple ports using a multipath protocol. [7] The apparatus of any one of claims 2 to 6, wherein the network interface device is to receive the first packet and the second packet at a plurality of input ports, and wherein the plurality of network interface devices are time-synchronized. [8] The apparatus of any one of claims 2 to 6, wherein the network interface device is to receive the first packet and the second packet via a single input port from the plurality of network interface devices, and wherein the plurality of network interface devices are time-synchronized. [9] Apparatus according to any one of claims 1 to 8, wherein the circuit is arranged to store packets with associated timestamps in a first queue and to store packets without associated timestamps in a second queue. [10] The device of claim 1, comprising at least one input port and at least one output port, wherein the at least one input port and the at least one output port are connected to the interface. [11] At least one non-transitory computer-readable medium having stored thereon instructions that, when executed by a switch circuit, cause the switch circuit to: Outputting a first packet associated with a first queue and outputting a second packet associated with a second queue based on timestamps associated with the first packet and the second packet, wherein: the first packet and the second packet are received by multiple network interface devices, the multiple network interface devices are synchronized in time, and the timestamps associated with the first packet and the second packet indicate the transmission of the first packet and the second packet. [12] At least one non-transitory computer-readable medium according to claim 11, comprising instructions stored thereon that, when executed by a switch circuit, cause the switch circuit to: Outputting the first packet associated with the first queue and the second packet associated with the second queue based on timestamps associated with the first packet and the second packet. [13] At least one non-transitory computer-readable medium according to claim 11, wherein the first packet and the second packet are associated with the same flow. [14] At least one non-transitory computer-readable medium according to any one of claims 11 to 13, comprising instructions stored thereon which, when executed by a switch circuit, cause the switch circuit to: Output the second packet before the first packet based on a priority level of the second packet. [15] At least one non-transitory computer-readable medium according to any one of claims 11 to 14, comprising instructions stored thereon which, when executed by a switch circuit, cause the switch circuit to: Output the first packet and the second packet over multiple ports using a multipath protocol. [16] Procedure comprising: a switch that performs the following: Outputting a first packet associated with a first queue and a second packet associated with a second queue based on timestamps associated with the first packet and the second packet, wherein the first packet and the second packet are received by the switch from a plurality of network interface devices, and wherein the plurality of network interface devices are time-synchronized. [17] A method according to claim 16, comprising: Output the second packet before the first packet based on a priority level of the second packet. [18] A method according to claim 16, comprising: Receiving the first and second packets on multiple input ports. [19] A method according to any one of claims 16-18, comprising: Output the first packet and the second packet over multiple ports using a multipath protocol. [20] A method according to any one of claims 16-19, comprising: Storing packets with associated timestamps in the first queue and Storing packets without timestamps in the second queue.