Computer network system, computer networking method, and computer readable medium
By introducing a processing unit into a network interface card (NIC), using the data path on the NIC to spy on the data traffic generated and received by the host, identifying the timestamps of forward and reverse packet flows, and calculating the delay between servers, solving the problem of large-scale consumption of host computing resources in the prior art, and improving the accuracy and reliability of delay measurements.
Patent Information
- Application Number
- CN202510419805.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-14
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-16
AI Technical Summary
When measuring the delay between servers in a data center, the prior art relies on the ping requests/replies generated by the host, resulting in a large amount of computing resources of the host CPU and memory, affecting processing efficiency.
By introducing a processing unit into a network interface card (NIC), the data path on the NIC is used to spy on the data traffic generated and received by the host, identify the timestamps of forward and reverse packet flows, calculate the delay between servers, and reduce the dependence on the host computing resources.
Improves the accuracy and reliability of latency measurements, reduces the load on the host CPU and memory, and improves the overall network processing efficiency.
Smart Images

Figure CN120017558A_ABST
Abstract
Description
[0001] This application is a divisional application with a filing date of June 30, 2022, application number 202210761437.1, and invention name “Determining Delay Using a Network Interface Card with a Processing Unit”, and all contents are incorporated herein by reference.
[0002] This application claims the benefit of U.S. Patent Application No. 17 / 806,865, filed on June 14, 2022, which claims the benefit of Indian Provisional Patent Application No. 202141029411, filed on June 30, 2021. The entire contents of each of these applications are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to computer networks. Background Art
[0004] In a typical cloud data center environment, there are a large number of interconnected servers that provide computing and / or storage capabilities to run various applications. For example, a data center may include facilities that host applications and services for users (i.e., customers of the data center). For example, a data center may host all infrastructure equipment, such as networking and storage systems, redundant power supplies, and environmental controls. In a typical data center, clusters of storage servers and application servers (computing nodes) are interconnected via a high-speed switching fabric provided by one or more layers of physical network switches and routers. More complex data centers provide user support equipment located in various physical hosting facilities for infrastructures spread across the globe.
[0005] The connection between the server and the switching fabric occurs on a hardware module called a network interface card (NIC). A traditional NIC includes an application-specific integrated circuit (ASIC) that performs packet forwarding, which includes some basic Layer 2 / Layer 3 (L2 / L3) functions. In a traditional NIC, packet processing, supervision, and other advanced functions known as the "data path" are performed by the host central processing unit (CPU) (i.e., the CPU of the server that includes the NIC). Therefore, applications running on the server and data path processing share CPU resources in the server. For example, in a 4-core x86 server, one of the cores can be reserved for the data path, leaving 3 cores (or 75% of the CPU) for applications and the host operating system.
[0006] Some NIC vendors have begun to include additional processing units in the NIC itself to offload at least some of the data path processing from the host CPU to the NIC. The processing unit in the NIC may be, for example, a multi-core ARM processor with some programmable hardware acceleration provided by a data processing unit (DPU), a field programmable gate array (FPGA), and / or an ASIC. NICs that include such enhanced processing capabilities are often referred to as SmartNICs. Summary of the invention
[0007] Generally, techniques are described for an edge service platform that utilizes processing units of SmartNICs to enhance the processing and networking capabilities of a network of servers that include SmartNICs. The functionality provided by the edge service platform may include: orchestration of NICs; application programming interface (API)-driven service deployment on NICs; NIC addition, removal, and replacement; monitoring of services and other resources on NICs; and managing connections between various services running on NICs.
[0008] A network may include multiple servers, wherein packets are propagated throughout the network between one or more pairs of the multiple servers. The amount of time it takes for a packet to travel back and forth between two of the multiple servers defines the latency between the two servers. It may be beneficial to determine the latency between one or more pairs of servers so that a controller can monitor the performance of the network. Existing latency measurement and monitoring techniques include Pingmesh, a program that uses a mesh visualization to determine and display the latency between any two servers in a data center. Pingmesh creates such a graph by having each server send a periodic "ping" (i.e., an Internet Control Message Protocol (ICMP) echo request) to each other server, which then immediately responds with an ICMP echo reply. Pingmesh can optionally use Transmission Control Protocol (TCP) and Hypertext Transfer Protocol (HTTP) pings. Pingmesh collects latency information based on the ping request / reply round-trip time and uses that information to generate a mesh visualization.
[0009] As further described in detail herein, the processing unit of the NIC may execute an agent that determines the delay between the device hosting the NIC and another device. When determining the delay, the processing unit of the NIC uses little or no computing resources of the host device, such as the host CPU and memory, relative to the prior art (such as Pingmesh that relies on host-generated ping requests / replies). In some examples, the agent executed on the NIC processing unit may determine the delay by snooping the data or control traffic generated and received by the host device and passing through the data path on the NIC. For example, the agent may detect the forward packets of the forward packet flow and the reverse packets of the reverse packet flow, generate and obtain the timestamp information of the forward packets and the reverse packets, and calculate the round trip time according to the timestamp information to determine the delay between the source device and the destination device of the forward and reverse packet flows. The agent may perform a similar process for the address resolution protocol (ARP) request / reply. In this way, the agent may determine the delay without the agent or host separately generating and sending a ping request or response to another device for the purpose of delay measurement. The agent may instead snoop the existing data traffic exchanged between the devices. The proxy may also perform a similar process for ICMP echo request / reply message pairs. In addition, because the proxy is executed on the NIC, the timestamps of at least some of the packet streams, ARP request / reply packets, and ICMP echo request / reply packets are not affected by delays within the host caused by the kernel network stack, DMA transfers or memory copies between NIC memory and host memory, interrupts and polling, process context switches, and other non-deterministic timing that typically affects delays in packet processing. Therefore, these techniques can improve the accuracy and reliability of timestamp and round-trip time calculations.
[0010] In some examples, a system includes: a network interface card (NIC) of a first computing device, wherein the NIC includes: a set of interfaces configured to receive one or more packets and to send one or more packets; and a processing unit configured to: identify information indicating a forward packet; calculate a delay between the first computing device and a second computing device based on a first time corresponding to the forward packet and a second time corresponding to a reverse packet associated with the forward packet, wherein the second computing device includes a destination for the forward packet and a source for the reverse packet; and output information indicating the delay between the first computing device and the second computing device.
[0011] In some examples, a method includes: identifying, by a processing unit of a network interface card (NIC) of a first computing device, information indicating a forward packet, wherein the NIC includes a set of interfaces, and wherein the set of interfaces is configured to receive one or more packets and to send one or more packets; calculating, by the processing unit, a delay between the first computing device and a second computing device based on a first time corresponding to the forward packet and a second time corresponding to a reverse packet associated with the forward packet, wherein the second computing device includes a destination of the forward packet and a source of the reverse packet; and outputting, by the processing unit, information indicating the delay between the first computing device and the second computing device.
[0012] In some examples, a non-transitory computer-readable medium includes instructions for causing a processing unit of a network interface card (NIC) of a first computing device to: identify information indicating a forward packet; calculate a delay between the first computing device and a second computing device based on a first time corresponding to the forward packet and a second time corresponding to a reverse packet associated with the forward packet, wherein the second computing device includes a destination for the forward packet and a source for the reverse packet; and output information indicating the delay between the first computing device and the second computing device.
[0013] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram illustrating an example network system with data centers, in accordance with one or more techniques of this disclosure.
[0015] Figure 2 is a block diagram illustrating an example computing device using a network interface card with a separate processing unit to execute services managed by an edge service platform, in accordance with one or more techniques of this disclosure.
[0016] Figure 3 is a conceptual diagram illustrating a data center with servers, each server including a network interface card with a separate processing unit controlled by an edge service platform, in accordance with one or more techniques of this disclosure.
[0017] Figure 4 is a block diagram illustrating an example computing device using a network interface card with a separate processing unit to execute services managed by an edge service platform, in accordance with one or more techniques of this disclosure.
[0018] Figure 5 is a diagram showing a flow of packets including one or more techniques according to the present disclosure Figure 1 A block diagram of components of an exemplary network system.
[0019] Figure 6 is a flow diagram illustrating a first example operation for determining a delay between two devices in accordance with one or more techniques of this disclosure.
[0020] Figure 7 is a flow diagram illustrating a second example operation for determining a delay between two devices in accordance with one or more techniques of this disclosure.
[0021] Like reference characters refer to like elements throughout the specification and drawings. DETAILED DESCRIPTION
[0022] Figure 1 is a block diagram illustrating an exemplary network system 8 having a data center 10 in accordance with one or more techniques of the present disclosure. In general, the data center 10 provides an operating environment for applications and services at a customer site 11 (illustrated as "customer 11"), which has one or more customer networks coupled to the data center through a service provider network 7. For example, the data center 10 may host infrastructure equipment such as networking and storage systems, redundant power supplies, and environmental controls. The service provider network 7 is coupled to a public network 4, which may represent one or more networks managed by other providers and, thus, may form part of a large-scale public network infrastructure (e.g., the Internet). The public network 4 may represent, for example, a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a layer 3 virtual private network (VPN), an Internet Protocol (IP) intranet operated by a service provider operating the service provider network 7, an enterprise IP network, or some combination thereof.
[0023] Although customer sites 11 and public networks 4 are primarily illustrated and described as edge networks of service provider network 7, in some examples, one or more customer sites 11 and public networks 4 may be tenant networks within data center 10 or another data center. For example, data center 10 may host multiple tenants (customers), each tenant being associated with one or more virtual private networks (VPNs), each of which may implement one customer site 11.
[0024] The service provider network 7 provides packet-based connectivity to connected customer sites 11, data centers 10, and public networks 4. The service provider network 7 may represent a network owned and operated by a service provider to interconnect multiple networks. The service provider network 7 may implement multi-protocol label switching (MPLS) forwarding, and in this case may be referred to as an MPLS network or MPLS backbone. In some cases, the service provider network 7 represents multiple interconnected autonomous systems, such as the Internet, that provide services from one or more service providers.
[0025] In some examples, data center 10 may represent one of many geographically distributed network data centers. Figure 1 As shown in the example of , data center 10 can be a facility that provides network services to customers. Customers of a service provider can be collective entities, such as businesses, governments, or individuals. For example, a network data center may host web services for several businesses and end users. Other exemplary services may include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. Although shown as a separate edge network of a service provider network 7, elements of data center 10 (such as one or more physical network functions (PNFs) or virtualized network functions (VNFs)) may be included in the core of the service provider network 7.
[0026] In this example, data center 10 includes storage and / or computing servers interconnected via a switching fabric 14 provided by one or more layers of physical network switches and routers, wherein servers 12A-12X (referred to herein as "servers 12") are depicted coupled to top-of-rack switches 16A-16N. Servers 12 may also be referred to herein as "hosts" or "host devices." Although Figure 1 Only the server coupled to TOR switch 16A is shown in detail, but the data center 10 may include many additional servers coupled to other TOR switches 16 of the data center 10.
[0027] The switch fabric 14 in the illustrated example includes interconnected top-of-rack (TOR) (or other "leaf") switches 16A-16N (collectively, "TOR switches 16") that are coupled to a distribution layer of chassis (or "spine" or "core") switches 18A-18M (collectively, "chassis switches 18"). Although not shown, the data center 10 may also include, for example, one or more non-edge switches, routers, hubs, gateways, security devices such as firewalls, intrusion detection and / or intrusion prevention devices, servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices.
[0028] In this example, the TOR switch 16 and the chassis switch 18 provide redundant (multi-homed) connections to the IP structure 20 and the service provider network 7 for the server 12. The chassis switch 18 aggregates traffic flows and provides connections between the TOR switches 16. The TOR switch 16 can be a network device that provides Layer 2 (MAC) and / or Layer 3 (e.g., IP) routing and / or switching functions. The TOR switch 16 and the chassis switch 18 can each include one or more processors and memories, and can execute one or more software processes. The chassis switch 18 is coupled to the IP structure 20, and the IP structure can perform Layer 3 routing to route network traffic between the data center 10 and the customer site 11 through the service provider network 7. The switching architecture of the data center 10 is only an example. For example, other switching architectures can have more or fewer switching layers.
[0029] The term "packet flow", "traffic flow", or simply "flow" refers to a group of packets that originate from a specific source device or endpoint and are sent to a specific destination device or endpoint. A single packet flow can be identified, for example, by a 5-tuple: <source network address, destination network address, source port, destination port, protocol>. The 5-tuple generally identifies the packet flow to which the received packets correspond. An n-tuple refers to any n items extracted from a 5-tuple. For example, a 2-tuple of a packet can refer to a combination of <source network address, destination network address> or <source network address, source port> of the packet.
[0030] Each server 12 may be a computing node, an application server, a storage server, or other type of server. For example, each server 12 may represent a computing device, such as an x86 processor-based server, configured to operate according to the techniques described herein. The server 12 may provide a network function virtualization infrastructure (NFVI) for the NFV architecture.
[0031] The server 12 hosts one or more virtual network endpoints 23 (in Figure 1 These virtual networks are shown as "EP" 23 in the figure, and run on a physical network represented here by IP fabric 20 and switch fabric 14. Although described primarily with respect to data center based switch networks, other physical networks such as service provider network 7 can serve as the basis for one or more virtual networks.
[0032] Servers 12 each include at least one network interface card (NIC) of NICs 13A-13X (collectively "NICs 13"), each NIC including at least one port through which packets are exchanged to send and receive packets over the communication link. For example, server 12A includes NIC 13A.
[0033] In some examples, each NIC 13 provides one or more virtual hardware components for virtualized input / output (I / O). The virtual hardware component for I / O can be a virtualization of a physical NIC 13 ("physical function"). For example, in single root I / O virtualization (SR-IOV) described in the Peripheral Component Interface Special Interest Group SR-IOV specification, the PCIe physical function of a network interface card (or "network adapter") is virtualized to present one or more virtual network interface cards as "virtual functions" for use by corresponding endpoints executed on the server 12. In this way, virtual network endpoints can share the same PCIe physical hardware resources, and the virtual function is an example of a virtual hardware component. As another example, one or more servers 12 can implement Virtio, a paravirtualization framework that can be used for example in a Linux operating system, which provides simulated NIC functions as a type of virtual hardware component. As another example, one or more servers 12 can implement an open vSwitch to perform distributed virtual multilayer switching between one or more virtual NICs (vNICs) of a hosted virtual machine, where such a vNIC can also represent a type of virtual hardware component. In some cases, the virtual hardware component is a virtual I / O (e.g., NIC) component. In some cases, the virtual hardware component is an SR-IOV virtual function, and a direct process user space access based on a data plane development kit (DPDK) can be provided to SR-IOV.
[0034] In some examples, including Figure 1 In the example shown, one or more NICs 13 may include multiple ports. The NICs 13 may be interconnected via the ports and communication links of the NICs 13 to form a NIC fabric 23 having a NIC fabric topology. The NIC fabric 23 is a collection of NICs 13 connected to at least one other NIC 13.
[0035] The NICs 13 each include a processing unit to offload various aspects of the data path. The processing unit in the NIC may be, for example, a multi-core ARM processor with hardware acceleration provided by a data processing unit (DPU), a field programmable gate array (FPGA), and / or an ASIC. The NIC 13 may alternatively be referred to as a SmartNIC or a GeniusNIC.
[0036] According to various aspects of the techniques described in this disclosure, the edge service platform utilizes the processing unit 25 of the NIC 13 to enhance the processing and networking functions of the switch fabric 14 and / or the server 12 including the NIC 13.
[0037] The edge service controller 28 manages the operation of the edge service platform within the NIC 13 by orchestrating services to be executed by the processing unit 25; API-driven service deployment on the NIC 13; adding, deleting, and replacing NICs 13 within the edge service platform; monitoring services and other resources on the NIC 13; and managing connections between various services running on the NIC 13.
[0038] The edge service controller 28 may transmit information describing the services available on the NIC 13, the topology of the NIC fabric 13, or other information about the edge service platform to an orchestration system (not shown) or a network controller 24. Exemplary orchestration systems include OpenStack, VMWARE's vCenter, or Microsoft's System Center. Exemplary network controllers 24 include controllers of Juniper Networks or Tungsten Fabric. Additional information regarding the controller 24 in conjunction with other devices of the data center 10 or other software defined network operations is found in International Application No. PCT / US2013 / 044378, filed on June 5, 2013, entitled “PHYSICAL PATH DETERMINATION FOR VIRTUAL NETWORK PACKETFLOWS,” and in U.S. Patent Application No. 14 / 226,509, filed on March 26, 2014, entitled “Tunneled Packet Aggregation for Virtual Networks,” each of which is incorporated herein by reference as if fully set forth herein.
[0039] In some examples, a NIC of a first computing device (e.g., NIC 13A of server 12A), wherein the NIC includes a set of interfaces configured to receive one or more packets and send one or more packets. A forward packet may represent a packet sent from server 12A to another computing device, and a reverse packet may represent a packet received by server 12A from another computing device in response to a forward packet. Thus, NIC 13A may send and receive packets, and NIC 13A may process packets to determine whether a packet represents a forward packet or a reverse packet. In some examples, NIC 13A includes a processing unit 25A configured to identify information indicating a forward packet received by the set of interfaces. Processing unit 25A may calculate a delay between server 12A and another computing device (e.g., server 12X) based on a first time corresponding to the forward packet and a second time corresponding to a reverse packet associated with the forward packet, wherein server 12X includes a destination for the forward packet and a source for the reverse packet. The delay between server 12A and server 12X may represent the amount of time it takes for a packet to propagate from server 12A to server 12X or from server 12X to server 12A. That is, the delay between server 12A and server 12X may represent half of the amount of time it takes for a packet to travel back and forth between server 12A and server 12X. In some examples, processing unit 25A may output information indicating the delay between the first computing device and the second computing device.
[0040] It may be beneficial for the processing unit 25 of the NIC 13 to analyze one or more forward packets and one or more reverse packets to determine the delay between the servers in the data center 10. That is, by analyzing the one or more packets processed by the packets of the NIC 13 to determine the delay value, the processing unit 25 can determine the delay value while consuming a smaller amount of network resources compared to a system that does not use the computing resources of the NIC to determine the delay value. In addition, the processing unit 25 effectively uses the resources of the NIC 13 to determine one or more delays based on packets that exist for one or more purposes other than determining the delay. In other words, the NIC 13 may not send ping packets solely for the purpose of determining the delay. The NIC 13 analyzes packets that have other purposes, thereby reducing the amount of network resources consumed compared to a system that determines the delay by sending ping packets.
[0041] In some examples, the edge service controller 28 is configured to receive information indicating a delay between the server 12A and the server 12X from the NIC 13A. The edge service controller 28 may update a delay table to include a delay between the server 12A and the server 12X. The delay table maintained by the edge service controller 28 may indicate a plurality of delays, each of the plurality of delays corresponding to a respective server pair of the server 12. For example, the delay table may include a delay between the server 12A and the server 12B, a delay between the server 12A and the server 12C, a delay between the server 12B and the server 12C, and the like. Whenever the edge service controller 28 receives a delay value from one NIC 13A, the edge service controller 28 may maintain a delay table to indicate the received delay value. In some examples, the edge service controller 28 may generate a Pingmesh graph based on the delay table. The edge service controller 28 may output the Pingmesh graph to a user interface so that an administrator can view the health of the network.
[0042] In addition to determining the delay between server 12A and server 12X, processing unit 25A may determine the delay between server 12A and one or more other servers of server 12. For example, processing unit 12A may identify information indicating a forward packet received by NIC 13A. NIC 13A may determine that the source device of the forward packet is server 12A, and the destination device of the forward packet is server 12C. Processing unit 25A may be configured to calculate the delay between server 12A and server 12C. Server 12C includes a destination for the forward packet and a source of a reverse packet corresponding to the forward packet. Processing unit 25A may output information indicating the delay between server 12A and server 12C. The processing unit of the NIC may determine the delay between the host server of the NIC and one or more other servers of the data center 10. For example, processing unit 25A of NIC 13A may determine the delay between server 12A and server 12B, the delay between server 12A and server 12C, the delay between server 12A and server 12D, the delay between server 12A and server 12X, and the delay between server 12A and one or more other computing devices configured to receive forward packets and output reverse packets. Additionally or alternatively, processing unit 25B of NIC 13B may determine the delay between server 12B and server 12A, the delay between server 12B and server 12C, the delay between server 12B and server 12D, the delay between server 12B and server 12X, and the delay between server 12B and one or more other computing devices configured to receive forward packets and output reverse packets. Processing units 25C-25X may determine the delay between their respective host servers and other servers within or outside of data center 10.
[0043] In some examples, NIC 13A may receive a packet. The source device of the forward packet may be server 12A, and the destination device of the forward packet may be server 12X. Processing unit 25A may be configured to identify a source Internet Protocol (IP) address and a destination IP address in a header of the packet. Based on the source IP address and the destination IP address, processing unit 25A may determine that the packet is a forward packet originating from server 12A and destined for server 12X. Therefore, when NIC 13A sends a forward packet to server 12X and when NIC 13A receives a reverse packet from server 12X in response to server 12X receiving the forward packet, processing unit 25A may be configured to determine the delay between server 12A and server 12X.
[0044] In some examples, processing unit 25A may be configured to determine the delay between server 12A and server 12X only when server 12X immediately sends a reverse packet in response to receiving a forward packet from server 12A. When server 12X does not immediately send a reverse packet, processing unit 25A may not be configured to determine the delay between server 12A and server 12X based on the time when the reverse packet arrives at server 12A because the reverse packet is delayed. Certain types of packets are configured to elicit an immediate reverse packet from a destination device. For example, a transmission control protocol (TCP) packet with any one or more of a synchronization (SYN) TCP packet flag, an urgent (URG) TCP packet flag, and a push (PSH) TCP packet flag may elicit an immediate reverse packet. Additionally or alternatively, a packet sent according to one or both of an Internet Control Message Protocol (ICMP) and an Address Resolution Protocol (ARP) may elicit an immediate reverse packet. In any case, it may be beneficial for processing unit 25A to identify the type of forward packet and determine whether the type represents a packet type that triggers an immediate reverse packet.
[0045] Processing unit 25A can create a flow structure in response to identifying a forward packet with a packet type that leads to an immediate response packet. The flow structure can represent information indicating the reverse packet that processing unit 25A expects to receive in response to outputting the forward packet. For example, when processing unit 25A identifies the source IP address and destination IP address corresponding to the forward packet, processing unit 25A can create a flow structure to indicate the expected source IP address of the reverse packet and the expected destination IP address of the reverse packet. The expected source IP address of the reverse packet can represent the destination IP address of the forward packet, and the expected destination IP address of the reverse packet can represent the source IP address of the forward packet. Additionally or alternatively, processing unit 25A can create a timestamp corresponding to the time when the forward packet leaves server 12A and goes to a destination device (e.g., server 12X).
[0046] NIC 13A may output a forward packet to server 12X. In response to outputting the forward packet to server 12X, NIC 13A may receive a reverse packet. Processing unit 25A is configured to identify information indicating a reverse packet received by NIC 13A. Processing unit 25A may determine that the reverse packet represents a packet received by NIC 13A in response to outputting the forward packet based on a flow structure generated by processing unit 25A in response to receiving the forward packet. For example, processing unit 25A may identify a source IP address of the reverse packet and a destination IP address of the reverse packet. When the source IP address of the reverse packet and the destination IP address of the reverse packet match an expected source IP address and an expected destination IP address in the flow structure, processing unit 25A may determine that the reverse packet represents a packet received by NIC 13A in response to outputting the forward packet.
[0047] Processing unit 25A may determine the delay between server 12A and server 12X based on the round trip time of the packets between server 12A and server 12X. For example, processing unit 25A may identify the time corresponding to the arrival of the reverse packet at NIC 13A. Since the timestamp generated by processing unit 25A corresponds to the time when the forward packet leaves server 12A, processing unit 25A may calculate the round trip time of the packets between server 12A and server 12X by subtracting the timestamp from the time corresponding to the arrival of the reverse packet at NIC 13A. In some examples, the delay may represent half of the round trip time of the packets. Processing unit 25A may output information indicating the delay between server 12A and server 12X to edge service controller 28.
[0048] In some examples, edge service controller 28 configures processing unit 25A to, for example, begin measuring latency between server 12A and other servers 12 or between server 12A and specific other servers 12. Edge service controller 28 may configure processing unit 25A to stop measuring latency, or parameterize an algorithm executed by processing unit 25A to calculate latency.
[0049] Figure 2 is a block diagram illustrating an example computing device using a network interface card with a separate processing unit to execute services managed by an edge service platform, in accordance with one or more techniques of this disclosure. Figure 2 The computing device 200 may represent a real or virtual server and may represent Figure 1 24. In this example, the computing device 200 includes a bus 242 that couples hardware components of the computing device 200 hardware environment. The bus 242 couples a network interface card (NIC) 230 that supports SR-IOV, a storage disk 246, and a microprocessor 210. In some cases, the front-end bus can couple the microprocessor 210 and the memory device 244. In some examples, the bus 242 can couple the memory device 244, the microprocessor 210, and the NIC 230. The bus 242 can represent a peripheral component interface (PCI) express (PCIe) bus. In some examples, a direct memory access (DMA) controller can control DMA transfers between components coupled to the bus 242. In some examples, components coupled to the bus 242 control DMA transfers between components coupled to the bus 242.
[0050] Microprocessor 210 may include one or more processors, each processor including independent execution units ("processing cores") to execute instructions in accordance with an instruction set architecture. The execution units may be implemented as separate integrated circuits (ICs) or may be combined in one or more multi-core processors (or "many-core" processors), each of which is implemented using a single IC (i.e., a chip multiprocessor).
[0051] Disk 246 represents a computer-readable storage medium, which includes volatile and / or nonvolatile, removable and / or non-removable media implemented in any method or technology for storing information such as processor-readable instructions, data structures, program modules or other data. Computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), EEPROM, flash memory, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by microprocessor 210.
[0052] The main memory 244 includes one or more computer-readable storage media, which may include random access memory (RAM), such as various forms of dynamic RAM (DRAM), for example, DDR2 / DDR3 SDRAM or static RAM (SRAM), flash memory or any other form of fixed or removable storage medium that can be used to carry or store desired program code and program data in the form of instructions or data structures and can be accessed by the computer. The main memory 144 provides a physical address space consisting of addressable memory locations.
[0053] The network interface card (NIC) 230 includes one or more interfaces 232 configured to exchange packets using the links of the underlying physical network. The interface 232 may include a port interface card having one or more network ports. The NIC 230 also includes, for example, an on-card memory 227 for storing packet data. Direct memory access transmissions between the NIC 230 and other devices coupled to the bus 242 may read from / write to the memory 227.
[0054] The memory 244 , NIC 230 , storage disk 246 , and microprocessor 210 provide an operating environment for executing a software stack of a hypervisor 214 and one or more virtual machines 228 managed by the hypervisor 214 .
[0055] Typically, a virtual machine provides a virtualized / guest operating system for executing applications in an isolated virtual environment. Because the virtual machine is virtualized from the physical hardware of the host server, the executing application is isolated from both the hardware of the host and other virtual machines.
[0056] An alternative to virtual machines is virtualized containers, such as those provided by the open source DOCKER container application. Like virtual machines, each container is virtualized and can remain isolated from the host and other containers. However, unlike virtual machines, each container can omit a separate operating system and only provide an application suite and application-specific libraries. Containers are executed by the host as isolated user space instances and can share an operating system and common libraries with other containers executing on the host. Therefore, containers may require less processing power, storage, and network resources than virtual machines. As used herein, containers may also be referred to as virtualization engines, virtual private servers, silos, or jails. In some cases, the technology described herein involves containers and virtual machines or other virtualization components.
[0057] Although shown and described with respect to a virtual machine Figure 2 , but other operating environments such as one or more containers (e.g., Docker containers) can also implement virtual network endpoints. For example, containers can be deployed using Kubernetes pods. The operating system kernel ( Figure 2 ) can execute in kernel space and can include, for example, Linux, Berkeley Software Distribution (BSD), another Unix variant kernel, or a Windows server operating system kernel available from Microsoft.
[0058] The computing device 200 executes a hypervisor 214 to manage virtual machines 228. Exemplary hypervisors include Kernel-based Virtual Machine (KVM) for the Linux kernel, Xen, ESXi available from VMWARE, Windows Hyper-V available from Microsoft, and other open source and proprietary hypervisors. The hypervisor 214 may represent a virtual machine manager (VMM).
[0059] The virtual machine 228 can host one or more applications, such as virtual network function instances. In some examples, the virtual machine 228 can host one or more VNF instances, wherein each VNF instance is configured to apply a network function to a packet.
[0060] The hypervisor 214 includes a physical driver 225 to use the physical functions provided by the network interface card 230. In some cases, the network interface card 230 can also implement SR-IOV to enable sharing of physical network functions (I / O) between virtual machines 228. Each port of the NIC 230 can be associated with a different physical function. The shared virtual device, also known as a virtual function, provides dedicated resources so that each virtual machine 228 (and the corresponding client operating system) can access the dedicated resources of the NIC 230, so that for each virtual machine 228, the NIC 230 behaves as a dedicated NIC. The virtual function can represent a lightweight PCIe function that shares physical resources with physical functions and other virtual functions. According to the SR-IOV standard, the NIC 230 can have thousands of virtual functions available, but for I / O intensive applications, the number of configured virtual functions is typically much less.
[0061] The virtual machine 228 includes a corresponding virtual NIC 229 that is directly presented in the guest operating system of the virtual machine 228, thereby providing direct communication between the NIC 230 and the virtual machine 228 via the bus 242 using the virtual functions assigned to the virtual machine. This can reduce the hypervisor 214 overhead involved in software-based VIRTIO and / or vSwitch implementations, where the hypervisor 214 memory address space of the memory 244 stores packet data and the packet data copied from the NIC 230 to the hypervisor 214 memory address space and from the hypervisor 214 memory address space to the virtual machine 228 memory address space consumes cycles of the microprocessor 210.
[0062] The NIC 230 may also include a hardware-based Ethernet bridge 234 to perform Layer 2 forwarding between the virtual functions and the physical functions of the NIC 230. The bridge 234 thus provides hardware acceleration of packet forwarding between virtual machines 228 via the bus 242 and between the hypervisor 214 and any virtual machines 228 accessing the physical functions via the physical driver 225.
[0063] Computing device 200 may be coupled to a physical network switch fabric that includes an overlay network that extends the switch fabric from physical switches to software or "virtual" routers of physical servers coupled to the switch fabric, including virtual router 220. A virtual router may be a physical server (e.g., Figure 1A process or thread or component thereof executed by a server 12) that dynamically creates and manages one or more virtual networks that can be used for communication between virtual network endpoints. In one example, the virtual router implements each virtual network using an overlay network that provides the ability to separate the virtual address of the endpoint from the physical address (e.g., IP address) of the server on which the endpoint executes. Each virtual network can use its own addressing and security scheme and can be considered orthogonal to the physical network and its addressing scheme. Various techniques can be used to transmit packets within and across virtual networks through the physical network. At least some of the functions of the virtual router can be performed as a service 233.
[0064] exist Figure 2 In the exemplary computing device 200 of , the virtual router 220 executes within the hypervisor 214 using the physical functions of the I / O, but the virtual router 220 may execute within the processing unit 25 of the hypervisor, the host operating system, the host application, a virtual machine 228 and / or the NIC 230.
[0065] Typically, each virtual machine 228 may be assigned a virtual address for use within a corresponding virtual network, wherein each virtual network may be associated with a different virtual subnet provided by virtual router 220. Virtual machines 228 may be assigned their own virtual layer 3 (L3) IP addresses, e.g., for sending and receiving communications, but may not be aware of the IP address of the computing device 200 on which the virtual machine executes. In this manner, a "virtual address" is an address of an application program that is different from the logical address of the underlying physical computer system (e.g., computing device 200).
[0066] In one implementation, the computing device 200 includes a virtual network (VN) agent (not shown) that controls the coverage of the virtual network of the computing device 200 and coordinates the routing of data packets within the computing device 200. Typically, the VN agent communicates with the virtual network controllers of multiple virtual networks, which generate commands to control the routing of packets. The VN agent can operate as a proxy for control plane messages between the virtual machine 228 and the virtual network controller (such as the controller 24). For example, a virtual machine can request to send a message via the VN agent using its virtual address, and the VN agent can send the message in turn and request to receive a response to the message for the virtual address of the virtual machine that initiated the first message. In some cases, the virtual machine 228 can cause a procedure or function call provided by the application programming interface of the VN agent, and the VN agent can also handle the encapsulation of messages, including addressing.
[0067] In one example, a network packet (e.g., a layer 3 (L3) IP packet or a layer 2 (L2) Ethernet packet) generated or consumed by an instance of an application executed by a virtual machine 228 within a virtual network domain can be encapsulated in another packet (e.g., another IP or Ethernet packet) transmitted by a physical network. The packets transmitted in the virtual network may be referred to as "internal packets" herein, while the physical network packets may be referred to as "external packets" or "tunnel packets" herein. The virtual router 220 may perform encapsulation and / or decapsulation of virtual network packets within physical network packets. This function is referred to as a tunnel herein and can be used to create one or more overlay networks. In addition to IPinIP, other example tunnel protocols that can be used include IP on Generic Routing Encapsulation (GRE), VxLAN, Multiprotocol Label Switching (MPLS) on GRE, MPLS on User Datagram Protocol (UDP), etc.
[0068] As described above, a virtual network controller may provide a logically centralized controller to facilitate the operation of one or more virtual networks. The virtual network controller may, for example, maintain a routing information base, e.g., one or more routing tables storing routing information for a physical network and one or more overlay networks. The virtual router 220 of the hypervisor 214 implements a network forwarding table (NFT) 222A-222N for the N virtual networks for which the virtual router 220 operates as a tunnel endpoint. Typically, each NFT 222 stores forwarding information for the corresponding virtual network and identifies where data packets will be forwarded and whether the packets will be encapsulated in a tunnel protocol, such as using a tunnel header, which may include one or more headers of different layers of the virtual network protocol stack. Each NFT 222 may be an NFT for a different routing instance (not shown) implemented by the virtual router 220.
[0069] According to the technology described in the present disclosure, the edge service platform utilizes the processing unit 25 of the NIC 230 to enhance the processing and networking capabilities of the computing device 200. The processing unit 25 includes a processing circuit 231 to execute services orchestrated by the edge service controller 28. The processing circuit 231 can represent any combination of processing cores, ASICs, FPGAs, and other integrated circuits and programmable hardware. In one example, the processing circuit may include a system on a chip (SoC) having, for example, one or more cores, a network interface for high-speed packet processing, one or more acceleration engines for specialized functions (e.g., security / cryptography, machine learning, storage), programmable logic, integrated circuits, etc. Such a SoC may be referred to as a data processing unit (DPU).
[0070] In the exemplary NIC 230, the processing unit 25 executes an operating system kernel 237 and a user space 241 for services. The kernel may be a Linux kernel, a Unix or BSD kernel, a real-time OS kernel, or other kernel for managing the hardware resources of the processing unit 25 and managing the user space 241.
[0071] Services 233 may include network, security, storage, data processing, co-processing, machine learning, or other services. Processing unit 25 may execute services 233 and edge service platform (ESP) agent 236 as processes and / or execute services 233 and edge service platform (ESP) agent 236 within a virtual execution element such as a container or a virtual machine. As described elsewhere herein, services 233 may enhance the processing power of a host processor (e.g., microprocessor 210) by, for example, enabling computing device 200 to offload packet processing, security, or other operations that would be performed by the host processor.
[0072] Processing unit 25 executes edge service platform (ESP) agent 236 to exchange data and control data with an edge service controller of the edge service platform. Although shown in user space 241, ESP agent 236 may be a kernel module 237 in some cases.
[0073] As an example, ESP agent 236 may collect and send to the ESP controller telemetry data generated by services 233, the telemetry data describing traffic in the network, computing device 200 or network resource availability, resource availability (such as memory or core utilization) of resources of processing unit 25. As another example, ESP agent 236 may receive from the ESP controller service code for executing any service 233, service configuration for configuring any service 233, packets for injection into the network, or other data.
[0074] The edge service controller 28 manages the operation of the processing unit 25 by, for example, orchestrating and configuring services 233 performed by the processing unit 25; deploying services 233; adding, deleting, and replacing NICs 230 within the edge service platform; monitoring services 233 and other resources on the NIC 230; and managing connections between various services 233 running on the NIC 230. Exemplary resources on the NIC 230 include memory 227 and processing circuitry 231.
[0075] Processing unit 25 may execute delay agent 238 to determine a delay between computing device 200 and one or more other computing devices. In one example, computing device 200 may represent server 12A, and processing unit 25 may execute delay agent 238 to determine a delay between server 12A and server 12X, for example, based on receiving a forward packet indicating server 12A as a source device and server 12X as a destination device. In some examples, delay agent 238 may be part of ESP agent 236. For example, NIC 230 may receive the forward packet from one or more components of computing device 200 (e.g., microprocessor 210). For example, the packet may be generated by an application running on computing device 200 and executed by microprocessor 210.
[0076] The packet may be propagated to the interface 232 via the Ethernet bridge 234. The interface 232 may be configured to receive one or more packets and to send one or more packets. Thus, the NIC 230 is configured to receive one or more packets from components of the computing device 200 via the bus 242 and to receive one or more packets from other computing devices via the interface 232. Additionally or alternatively, the NIC 230 is configured to send one or more packets to components of the computing device 200 via the bus 242 and to send one or more packets to other computing devices via the interface 232.
[0077] When the NIC 230 receives a packet via the interface 232 or bus 242, the processing unit 25 may perceive the packet from the wire; or the Ethernet bridge 234 may be configured with a filter that matches packets used to determine the delay between the computing device 200 and another computing device and switches these packets to the processing unit 25 for further processing; or the processing unit 25 may (in some cases) include the Ethernet bridge 234 and apply processing packets for determining delay, as described below and elsewhere in this disclosure.
[0078] Processing unit 25 may execute delay agent 238 to identify information corresponding to the packet and analyze the information. For example, delay agent 238 may determine whether the packet is used to determine the delay between computing device 200 and another computing device. In some examples, when processing unit 25 receives information corresponding to a packet arriving at NIC 230, processing unit 25 may execute delay agent 238 to run an algorithm to determine a delay value. The following exemplary computer code may represent an algorithm for processing packet information to determine a delay value:
[0079] For each group P,
[0080]
[0081] As seen in the exemplary computer code, the delay agent 238 can determine whether a packet arriving at the NIC 230 is according to TCP, ICMP, or ARP (e.g., the "if protocolType(P) == TCP", "elseif protocolType(P) == ICMP", and "elseif protocolType(P) == ARP" lines in the exemplary computer code). The delay agent 238 can determine whether the packet is according to TCP, ICMP, or ARP because some packets sent according to TCP, ICMP, and ARP may cause an immediate reverse packet from the destination device. For example, when the NIC 230 sends a forward packet according to ICMP to the destination device, when the destination device receives the forward packet, the destination device may immediately send a reverse packet back to the NIC 230. Therefore, packets sent according to TCP, ICMP, and ARP can be used to determine the delay between two computing devices because the source device can determine the round-trip time based on the time it takes for the forward packet to be sent from the source and the time it takes for the reverse packet to return to the source. When the delay proxy 238 determines that the packet is not according to TCP, ICMP, or ARP, the algorithm can end and the delay proxy 238 can apply the algorithm to the next packet that arrives at the NIC 230.
[0082] When delay agent 238 determines that the packet arriving at NIC 230 is sent according to TCP, delay agent 238 may proceed to determine whether the packet represents a forward packet. For example, delay agent 238 may determine the source IP address and the destination IP address corresponding to the packet. When the source IP address corresponds to computing device 200 and the destination device corresponds to another computing device, the packet represents a forward packet.
[0083] When the delay agent 238 determines that the packet arriving at the NIC 230 is a forward packet sent according to TCP, the delay agent 238 can determine whether the packet includes at least one of a set of packet flags (e.g., the "if isSet(P->flags, URG||SYN||PSH)" line in the exemplary computer code). As seen in the exemplary computer code, the set of TCP packet flags can include a synchronization (SYN) TCP packet flag, an urgent (URG) TCP packet flag, and a push (PSH) TCP packet flag. The SYN packet flag, the URG packet flag, and the PSH packet flag can indicate a TCP packet that triggers an immediate reverse packet from a destination device. One or more TCP packets that do not include a SYN packet flag, a URG packet flag, or a PSH packet flag may not trigger an immediate reverse packet from a destination device. Therefore, it may be beneficial for the delay agent 238 to determine the delay based on a TCP packet that includes any one or more of the SYN packet flag, the URG packet flag, and the PSH packet flag.
[0084] The processing unit 25 may execute the delay agent 238 to create a flow structure based on identifying the forward packet sent according to the TCP protocol (e.g., the line "reverseFlow = createFlow (P->dip, P->sip, P->dport, P->sport)" in the exemplary computer code). The flow structure may include an expected source IP address and an expected destination IP address, which correspond to the reverse packets expected to arrive at the NIC 230 in response to the forward packets arriving at the destination device. In some examples, the expected source IP address of the reverse packet is the destination IP address of the forward packet, and the expected destination IP address of the reverse packet is the source IP address of the forward packet, because the forward packet and the reverse packet complete a "round trip" between a pair of devices.
[0085] Additionally or alternatively, the processing unit 25 can create a timestamp based on identifying the forward packet sent according to the TCP protocol (e.g., the line "reverseFlow->timeStamp=getTime()" in the exemplary computer code). The NIC 230 can output the forward packet to the destination device. The processing unit 25 can create a timestamp to indicate the approximate time when the forward packet leaves the NIC 230 for the destination device.
[0086] When the NIC 230 receives a reverse packet based on the destination device receiving the forward packet, the processing unit 25 may identify information corresponding to the reverse packet. For example, the processing unit 25 may execute the delay agent 238 to identify the source IP address and the destination IP address indicated by the reverse packet. The processing unit 25 may determine, based on the flow structure created for the forward packet, that the reverse packet represents a packet sent by the destination device in response to receiving the forward packet (e.g., the "forwardFlow = getFlow (P->sip, P->dip, P->sport, P->dport)" and "if valid (forwardFlow)" lines in the exemplary computer code). Based on determining that the reverse packet corresponds to the forward packet, the NIC 230 may execute the delay agent 238 to determine the round-trip time between the computing device 200 and the destination device of the forward packet (e.g., the "rtt = getTime () - forwardFlow->timeStamp" line in the exemplary computer code). To determine the round trip time, the delay agent 238 may subtract the timestamp corresponding to the time when the forward packet left the NIC 230 (e.g., "forwardFlow->timestamp") from the current time when the reverse packet arrived at the NIC 230 (e.g., "getTime()"). Since the forward packet immediately causes the destination device to send a reverse packet, the round trip time indicates the delay between the computing device 200 and the destination device. The delay agent 238 may then update the delay between the computing device 200 and the destination device.
[0087] When delay proxy 238 determines that a packet arriving at NIC 230 is a forward packet sent according to ICMP or ARP, delay proxy 238 may perform a process similar to that described for TCP packets, except that in the example of ICMP and ARP packets, delay proxy 238 may not check the packet flags.
[0088] In some examples, the NIC may use the rate at which TCP sequence numbers move in an elephant flow to calculate the throughput between two nodes (e.g., between two servers 12). The delay agent 238 may execute an algorithm to track one or more packet flows and track the throughput corresponding to the one or more packet flows.
[0089] In some examples, a media access control (MAC) address table on the NIC identifies whether the NIC is active. If the NIC is not communicating with any other node, the MAC table entry will time out after 3 minutes. The ESP agent 236 and / or the delay agent 238 can use this timeout event to maintain the reachability state of the ESP agent 236.
[0090] In some examples, edge service controller 28 may configure delay proxy 238 via ESP proxy 236 with a list of endpoints of interest. Such endpoints may be IP addresses of one or more other computing devices. In this case, delay proxy 238 may apply only the delay determination techniques described herein to calculate the delay between computing device 200 and those computing devices in the list of endpoints of interest. Optionally, Ethernet bridge 234 filters may be configured to switch packets to processing unit 25 that have packet header information identifying such packets as being associated with the list of endpoints of interest.
[0091] In some examples, the n-tuple (P->sip, P->dip, P->sport, P->dport) packet information used in the above algorithm represents packet information in an internal packet header of an overlay / virtual network. That is, the endpoints are virtual network endpoints, such as virtual machines or pods, and the latency information is calculated for packets exchanged between virtual network endpoints rather than between servers (or in addition thereto).
[0092] The processing unit 25 can measure the delay between two nodes of the network (e.g., any two of the servers 12) by tracking packets passing through the NIC 13. In some examples, the processing unit 25 may determine the delay between the two devices based on TCP packets with one or more of a set of packet flags. Before the destination device sends a reverse packet, the destination device may process some TCP packets (e.g., TCP packets without any of a set of packet flags). The destination device may immediately send a response to a TCP packet with a URG, PSH, or SYN flag. For example, when the server 12A sends a forward TCP packet with a URG flag, a PSH flag, or a SYN flag to the server 12D, the server 12D may immediately send a reverse packet upon receiving the forward packet. Additionally or alternatively, the destination device may immediately send a reverse packet upon receiving an ICMP and ARP request. In this way, the processing unit 25 of the NIC 230 may record the time elapsed from the server computing device 200 sending the forward packet to the computing device 200 receiving the reverse packet from the destination device. The elapsed time may represent the delay between the computing device 200 and the destination device. By determining the delay between computing device 200 and the destination device without sending a probe packet from computing device 200 to the destination device, processing unit 25 can collect information for the Pingmesh graph while consuming a smaller amount of network resources than a system that sends ping packets to determine the delay. Processing unit 25 can timestamp TCP packets with one or more labels (e.g., URG, PSH, or SYN) to calculate the delay. In some examples, a processing unit (e.g., processing unit 25) can execute an algorithm for determining the delay between two servers of server 12.
[0093] Figure 3 is a conceptual diagram illustrating a data center with servers, each server including a network interface card with a separate processing unit controlled by an edge service platform, according to one or more techniques of the present disclosure. Computing node racks 307A-307N (collectively referred to as "computing node racks 307") may correspond to Figure 1 The server 12, and the switches 308A-308N (collectively referred to as "switches 308") may correspond to Figure 1 The agent 302 or orchestrator 304 represents software executed by a processing unit (referred to herein as a data processing unit or DPU) and receives configuration information of the processing unit and sends telemetry and other information including the NIC of the processing unit to the orchestrator 304. In some examples, the agent 302 includes Figure 2 ESP agent 236 and delay agent 238. In some examples, agent 302 includes a JESP agent. Network services 312, L4-L7 services 314, telemetry services 316, and Linux and software development kit (SDK) services 318 may represent examples of services 233. Orchestrator 304 may represent Figure 1 In some examples, proxy 302 may send one or more calculated delays to orchestrator 304. Orchestrator 304 may maintain a delay table to include delays received from proxy 302 and agents executed by other devices. Orchestrator 304 may generate a Pingmesh graph based on the delay table, wherein the Pingmesh graph indicates Figure 1 The health status of the network system 8.
[0094] The network automation platform 306 connects and manages the network devices and the orchestrator 304, and the network automation platform 306 can utilize the edge service platform through the network devices and the orchestrator. The network automation platform 306 can, for example, deploy network device configurations, manage the network, extract telemetry data, and analyze and provide indications of network status.
[0095] Figure 4 400 is a block diagram illustrating an exemplary computing device that uses a network interface card with a separate processing unit to execute services managed by an edge service platform according to the techniques described herein. Although a virtual machine is shown in this example, other instances of computing device 400 may also or alternatively run containers, local processes, or other endpoints of packet flows. Different types of vSwitches may be used, such as an open vSwitch or a virtual router (e.g., a track cloud). Other types of interfaces between endpoints and NICs, such as tap interfaces, veth pair interfaces, etc., are also contemplated.
[0096] Figure 5 is a diagram showing a flow of packets including one or more techniques according to the present disclosure Figure 1 A block diagram of components of an exemplary network system 8 is shown. Figure 5 As shown, a first forward packet is propagated from server 12A to server 12X via connection 502A, connection 502B, connection 502C, and connection 502D. A first reverse packet is propagated from server 12X to server 12A via connection 504A, connection 504B, connection 504C, and connection 504D. A second forward packet is propagated from server 12A to server 12B via connection 506A and connection 506B. A second reverse packet is propagated from server 12B to server 12A via connection 508A and connection 508B. Processing unit 25A may be configured to determine a delay between server 12A and server 12X based on the first forward packet and the first reverse packet, and processing unit 25A may be configured to determine a delay between server 12A and server 12B based on the second forward packet and the second reverse packet. Thus, processing unit 25A may be configured to determine the delay between two servers based on packets propagating through switch fabric 14 , and processing unit 25A may be configured to determine the delay between two servers based on packets propagating between endpoints 23 without propagating through switch fabric 14 .
[0097] Figure 6 is a flowchart illustrating a first exemplary operation for determining a delay between two devices according to one or more techniques of the present disclosure. Figure 1 Network system 8 and Figure 2 The computing device 200 is described Figure 6 .However, Figure 6 The techniques may be performed by different components of network system 8 and computing device 200, or by additional or alternative devices.
[0098] NIC 230 may receive a forward packet (602). In some examples, NIC 230 may receive a forward packet via bus 242, because a forward packet represents a packet originating from computing device 200 and destined for a real or virtual destination device. The application data of the packet may originate from a process executing on computing device 200, such as an application, a service (in user space 245), a kernel module of a host OS in kernel space 243. Processing unit 25 may identify information corresponding to the forward packet (604) and verify the forward packet based on the information (606). In some examples, processing unit 25 may verify the forward packet by identifying the source IP address and the destination IP address indicated by the packet. In order to verify the forward packet, processing unit 25 may confirm that the packet represents a forward packet. Processing unit 25 may determine the packet type of the forward packet (608). For example, in order to determine the packet type, the processing unit may determine whether the packet is sent according to TCP, ICMP, or ARP. If the packet is sent according to TCP, processing unit 25 may determine whether the packet includes at least one of a set of packet headers. Processing unit 25 may create a flow structure corresponding to the forward packet (610). In some examples, in response to the destination device receiving the forward packet, the flow structure indicates information corresponding to a reverse packet that is expected to arrive at NIC 230. Processing unit 25 may create a timestamp corresponding to the time when the forward packet leaves NIC 230 for the destination device (612).
[0099] NIC 230 may send a forward packet to a destination device (614). In some examples, NIC 230 is an example of NIC 13A of server 12A, and the destination device represents server 12B, but this is not required. NIC 230 may represent any NIC 13, and the destination device may represent any server 12 that does not host NIC 230. Server 12B may receive the forward packet (616). Server 12B may process the forward packet (618). In some examples, server 12B may recognize that the packet represents a forward packet from server 12A. Server 12B may send a reverse packet to server 12A (620). In some examples, server 12B immediately sends a reverse packet upon detecting the arrival of the forward packet.
[0100] NIC 230 may receive the reverse packet (622). Processing unit 25 may identify the reverse packet information (624) and validate the reverse packet (626). In some examples, processing unit 25 may validate the reverse packet by determining that the source IP address and the destination IP address of the reverse packet match the expected source IP address and the expected destination IP address of the flow structure (step 610). Processing unit 25 may identify a current time corresponding to the time when NIC 230 received the reverse packet (628). Processing unit 25 may calculate a delay between server 12A and server 12B based on the timestamp and the current time (630). Processing unit 25 may output information indicating the delay (632).
[0101] Figure 7 is a flow chart illustrating a second exemplary operation for determining a delay between two devices according to one or more techniques of the present disclosure. Figure 1 Network system 8 and Figure 2 The computing device 200 is described Figure 7 .However, Figure 7 The techniques may be performed by different components of network system 8 and computing device 200, or by additional or alternative devices.
[0102] In some examples, processing unit 25 of NIC 230 of computing device 200 may identify information indicating a forward packet (702). In some examples, the information includes a source device of the forward packet and a destination device of the forward packet. In some examples, the information includes a protocol of the packet. The information may inform processing unit 25 whether the forward packet will cause the destination device to immediately send a reverse packet upon receiving the forward packet, thereby allowing processing unit 25 to calculate a delay between the first computing device and the second computing device. Based on determining that the information indicates that the forward packet will cause the destination device to immediately send a reverse packet, processing unit 25 may determine that a delay may be calculated based on the forward packet and the reverse packet.
[0103] Processing unit 25 may calculate a delay between the first computing device and the second computing device based on a first time corresponding to the forward packet and a second time corresponding to the reverse packet associated with the forward packet (704). In some examples, processing unit 25 is configured to calculate the delay based on a time difference between the first time and the second time. Processing unit 25 may output information indicating the delay between the first computing device and the second computing device (706).
[0104] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. The various features described as modules, units, or components may be implemented together in an integrated logic device, or individually as separate but interoperable logic devices or other hardware devices. In some cases, the various features of the electronic circuit may be implemented as one or more integrated circuit devices, such as an integrated circuit chip or chipset.
[0105] If implemented in hardware, the present disclosure may be directed to a device, such as a processor or an integrated circuit device, such as an integrated circuit chip or chipset. Alternatively or additionally, if implemented in software or firmware, the technology may be implemented at least in part by a computer-readable data storage medium including instructions that, when executed, cause a processor to perform one or more of the above methods. For example, a computer-readable data storage medium may store such instructions that are executed by a processor.
[0106] The computer-readable medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include computer data storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. In some examples, the article of manufacture may include one or more computer-readable storage media.
[0107] In some examples, computer-readable storage media may include non-transitory media. The term "non-transitory" may indicate that the storage medium is not contained in a carrier wave or propagating signal. In some examples, non-transitory storage media may store data that may change over time (e.g., in RAM or cache).
[0108] The code or instructions may be software and / or firmware executed by a processing circuit that includes one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described in the present disclosure may be provided within a software module or a hardware module.
Claims
1. A computer network system, comprising: A network interface card NIC of a first computing device, wherein the NIC comprises: a set of interfaces; and The processing unit is configured as follows: identifying information indicating a generated forward packet to be output by the NIC; determining, based on the information indicative of the forward packet, whether the forward packet is configured to elicit a reverse packet from the second computing device immediately upon receipt of the forward packet by the second computing device; When the forward packet is configured to elicit the reverse packet from the second computing device immediately upon receipt of the forward packet by the second computing device, calculating a delay between the first computing device and the second computing device based on a first time corresponding to the forward packet and a second time corresponding to the reverse packet associated with the forward packet, wherein the second computing device includes a destination of the forward packet and a source of the reverse packet; and Information indicative of the delay between the first computing device and the second computing device is output.
2. The computer network system according to claim 1, wherein: In order to determine whether the forward packet is configured to elicit the reverse packet from the second computing device immediately when the second computing device receives the forward packet, the processing unit is configured to determine whether the forward packet represents at least one packet type from a set of packet types that are configured within the processing unit to be useful for determining delay.
3. A computer network system according to claim 2, when the forward packet is configured to immediately elicit the reverse packet from the second computing device when the second computing device receives the forward packet, the difference between the first time and the second time represents the round-trip time between the first computing device and the second computing device.
4. The computer network system according to claim 2, wherein: In order to determine whether the forward packet represents the at least one packet type from the set of packet types, the processing unit is configured to: It is determined whether the forward packet includes at least one Transmission Control Protocol TCP packet flag from a set of TCP packet flags, wherein the set of TCP packet flags includes a Synchronization SYN TCP packet flag, an Urgent URG TCP packet flag, and a Push PSH TCP packet flag.
5. The computer network system according to claim 2, wherein: In order to determine whether the forward packet represents the at least one packet type from the group of packet types, the processing unit is configured to determine whether the forward packet is an Internet Control Message Protocol ICMP packet or an Address Resolution Protocol ARP packet.
6. The computer network system according to any one of claims 1 to 5, further comprising: a plurality of computing devices, including the first computing device and the second computing device; as well as A controller, wherein the controller is configured to: receiving information indicative of the delay between the first computing device and the second computing device; and A delay table is updated to include the delay between the first computing device and the second computing device, wherein the delay table indicates a plurality of delays, each delay in the plurality of delays corresponding to a pair of computing devices in the plurality of computing devices.
7. The computer network system according to any one of claims 1 to 5, wherein: The forward packet represents a first forward packet, wherein the reverse packet represents a first reverse packet, and wherein the processing unit is configured to: identifying information indicating a generated second forward packet to be output by the NIC; determining, based on the information indicating the second forward packet, whether the second forward packet is configured to elicit a second reverse packet from the third computing device immediately upon receipt of the second forward packet by the third computing device; when the second forward packet is configured to elicit the second reverse packet from the third computing device immediately when the third computing device receives the second forward packet, calculating a delay between the first computing device and the third computing device based on a third time corresponding to the second forward packet and a fourth time corresponding to the second reverse packet associated with the second forward packet, wherein the third computing device includes a destination of the second forward packet and a source of the second reverse packet; and Information indicative of the delay between the first computing device and the third computing device is output.
8. The computer network system according to any one of claims 1 to 5, wherein: In order to identify the information indicating the forward packet, the processing unit is configured to: identifying a source Internet Protocol (IP) address and a destination IP address in a header of a packet received at the NIC; and Based on the source IP address and the destination IP address, it is determined that the packet represents the forward packet.
9. The computer network system according to any one of claims 1 to 5, wherein: Based on identifying the information, the processing unit is configured to: creating a flow structure indicating information corresponding to the reverse packet associated with the forward packet; and A timestamp is created corresponding to a time when the forward packet leaves the first computing device for the second computing device, wherein the first time is based on the time.
10. The computer network system according to claim 9, wherein: The NIC is configured to output the forward packet via the set of interfaces and to the second computing device, and wherein the processing unit is further configured to: identifying information indicative of packets received by the set of interfaces; determining, based on the flow structure, that the packet represents the reverse packet associated with the forward packet; identifying a time corresponding to arrival of the reverse packet at the set of interfaces; and The delay between the first computing device and the second computing device is determined based on a timestamp corresponding to a time when the forward packet leaves the set of interfaces and a time corresponding to a time when the reverse packet arrives at the set of interfaces.
11. A computer networking method, comprising: identifying, by a processing unit of a network interface card (NIC) of a first computing device, information indicating a generated forward packet to be output by the NIC, wherein the NIC includes a set of interfaces; determining, by the processing unit, based on the information indicating the forward packet, whether the forward packet is configured to elicit a reverse packet from the second computing device immediately when the second computing device receives the forward packet; When the forward packet is configured to elicit the reverse packet from the second computing device immediately when the second computing device receives the forward packet, the processing unit calculates a delay between the first computing device and the second computing device based on a first time corresponding to the forward packet and a second time corresponding to the reverse packet associated with the forward packet, wherein the second computing device includes a destination of the forward packet and a source of the reverse packet; and Information indicative of the delay between the first computing device and the second computing device is output by the processing unit.
12. The computer networking method according to claim 11, wherein: Determining whether the forward packet is configured to elicit the reverse packet from the second computing device immediately when the second computing device receives the forward packet includes: determining whether the forward packet represents at least one packet type from a set of packet types configured within the processing unit to be useful for determining delay.
13. The computer networking method according to claim 12, wherein: Determining that the forward packet represents the at least one packet type from the set of packet types comprises: The processing unit determines that the forward packet includes at least one Transmission Control Protocol TCP packet flag from a set of TCP packet flags, wherein the set of TCP packet flags includes a synchronization SYN TCP packet flag, an urgent URG TCP packet flag, and a push PSH TCP packet flag.
14. The computer networking method according to claim 12, wherein: Determining whether the forward packet represents the at least one packet type from the set of packet types includes determining, by the processing unit, whether the forward packet is an Internet Control Message Protocol (ICMP) packet or an Address Resolution Protocol (ARP) packet.
15. The computer networking method according to any one of claims 11 to 14, further comprising: receiving, by a controller, information indicating the delay between the first computing device and the second computing device, wherein the plurality of computing devices includes the first computing device and the second computing device; and A delay table is updated by the controller to include the delay between the first computing device and the second computing device, wherein the delay table indicates a plurality of delays, each delay in the plurality of delays corresponding to a pair of computing devices in the plurality of computing devices.
16. The computer networking method according to any one of claims 11 to 14, wherein: The forward packet represents a first forward packet, wherein the reverse packet represents a first reverse packet, and wherein the method further comprises: identifying, by the processing unit, information indicating a generated second forward packet to be output by the NIC; determining, by the processing unit, based on the information indicating the second forward packet, whether the second forward packet is configured to elicit a second reverse packet from the third computing device immediately when the third computing device receives the second forward packet; When the second forward packet is configured to elicit the second reverse packet from the third computing device immediately when the third computing device receives the second forward packet, calculating, by the processing unit, a delay between the first computing device and the third computing device based on a third time corresponding to the second forward packet and a fourth time corresponding to the second reverse packet associated with the second forward packet, wherein the third computing device includes a destination of the second forward packet and a source of the second reverse packet; and Information indicative of the delay between the first computing device and the third computing device is output by the processing unit.
17. The computer networking method according to any one of claims 11 to 14, wherein: Identifying the information indicative of the forward packet comprises: identifying, by the processing unit, a source Internet Protocol (IP) address and a destination IP address in a header of a packet received at the NIC; and The processing unit determines, based on the source IP address and the destination IP address, that the packet represents the forward packet.
18. The computer networking method according to any one of claims 11 to 14, wherein: Based on identifying the information indicative of the forward packet, the method further comprises: creating, by the processing unit, a flow structure indicating information corresponding to the reverse packet associated with the forward packet; and A timestamp is created, by the processing unit, corresponding to a time at which the forward packet leaves the first computing device for the second computing device, wherein the first time is based on the time.
19. The computer networking method according to claim 18, further comprising: outputting, by the NIC, the forward packet to the second computing device via the set of interfaces; identifying, by the processing unit, information indicative of packets received by the set of interfaces; determining, by the processing unit based on the flow structure, that the packet represents the reverse packet associated with the forward packet; identifying, by the processing unit, a time corresponding to arrival of the reverse packet at the set of interfaces; as well as The delay between the first computing device and the second computing device is determined by the processing unit based on a timestamp corresponding to a time when the forward packet leaves the set of interfaces and a time corresponding to a time when the reverse packet arrives at the set of interfaces.
20. A non-transitory computer-readable medium encoded with instructions for causing one or more programmable processors to be configured as a computer network system according to any one of claims 1 to 10 or to be configured to perform a computer networking method according to any one of claims 11 to 19.
Citation Information
Patent Citations
Tunneled packet aggregation for virtual networks
US9571394B1