Switch-managed resource allocation and software enforcement
A switch with integrated CPU and accelerator capabilities offloads network processing, addressing the strain on server resources by managing virtualization environments and orchestrating server resources, thereby enhancing data center efficiency and compliance with SLAs.
Patent Information
- Application Number
- JP2022568889
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-18
- Filing Date
- 2020-12-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2040-12-11
AI Technical Summary
Increasing network speeds in data centers to handle growing east-west traffic volumes strains server processor resources, necessitating more efficient packet processing and resource allocation to meet service level agreements (SLAs) without overburdening CPUs.
Implementing a switch with integrated CPU and accelerator capabilities to offload network processing, including protocol termination, decryption, and memory transactions, while managing virtualization environments and orchestrating server resources to free up CPU cycles for value-added services.
Reduces server processor utilization by offloading network processing to the switch, enabling timely compliance with SLAs and reducing data center total cost of ownership through efficient packet processing and resource management.
Smart Images

Figure 0007729016000001 
Figure 0007729016000002 
Figure 0007729016000003
Abstract
Description
[Technical Field]
[0001] [Priority Claim] This application claims priority under 35 U.S.C. 365(c) to U.S. Application No. 16 / 905,761, filed June 18, 2020, entitled "SWITCH-MANAGED RESOURCE ALLOCATION AND SOFTWARE EXECUTION," which is incorporated herein in its entirety. [Background technology]
[0002] In the context of cloud computing, cloud service providers (CSPs) offer various services to other businesses or individuals for use, such as infrastructure as a service (IaaS), software as a service (SaaS), or platform as a service (PaaS). Hardware infrastructure, including compute, memory, storage, accelerators, networking, etc., runs and supports the software stacks provided by the CSP and its customers.
[0003] CSPs may have experience with complex networking environments where packets are parsed, decapsulated, decrypted, and sent to the appropriate virtual machine (VM). In some cases, packet flows are balanced and metered to fulfill service level agreement (SLA) requirements. In some cases, network processing occurs on servers in data centers. However, increasing packet volume and the increasing volume and complexity of packet processing activities are placing increased strain on servers. While central processing units (CPUs) or other server processor resources are used for packet processing, the CPU and other processor resources could be used for other services that can be billed or generate higher revenue than packet processing. The impact of this problem increases significantly when using high bitrate network devices, such as 100 Gbps and faster networks. [Brief explanation of the drawings]
[0004] [Figure 1A] 1 illustrates an exemplary switch system. [Figure 1B] 1 illustrates an exemplary switch system. [Figure 1C] 1 illustrates an exemplary switch system. [Figure 1D] 1 illustrates an exemplary switch system.
[0005] [Figure 2A] 1 illustrates an exemplary overview of a system for managing resources in a rack.
[0006] [Figure 2B] 1 shows an exemplary overview of various management hierarchies.
[0007] [Figure 3] 1 illustrates an exemplary system in which a switch can respond to memory access requests.
[0008] [Figure 4A] Here is an example of a Memcached server running on a server and on a switch.
[0009] [Figure 4B] 1 shows the Ethernet packet flow for a single request.
[0010] [Figure 5A] 1 illustrates an exemplary system in which packets may terminate at a switch. [Figure 5B] 1 illustrates an exemplary system in which packets may terminate at a switch. [Figure 5C] 1 illustrates an exemplary system in which packets may terminate at a switch.
[0011] [Figure 6]1 illustrates an example of a switch that runs an orchestration control plane to manage which devices run virtualized execution environments.
[0012] [Figure 7A] 1 illustrates an example of migration of a virtualization execution environment from a server to another server.
[0013] [Figure 7B] 1 shows an example of migration of a virtualized execution environment.
[0014] [Figure 8A] 1 illustrates an exemplary process. [Figure 8B] 1 illustrates an exemplary process. [Figure 8C] 1 illustrates an exemplary process.
[0015] [Figure 9] Shows the system.
[0016] [Figure 10] Show the environment.
[0017] [Figure 11] 1 illustrates an exemplary network element. DETAILED DESCRIPTION OF THE INVENTION
[0018] Within a data center, north-south traffic may include packets flowing in and out of the data center, while east-west traffic may include packets flowing between nodes (e.g., racks of servers) within the data center. North-south traffic may be considered a product delivered to customers, while east-west traffic may be considered overhead. East-west traffic volume is growing at a significantly higher rate than north-south traffic, and processing east-west traffic flows in a timely manner to comply with applicable SLAs while reducing the data center's total cost of ownership (TCO) is a growing challenge within data centers.
[0019] Increasing network speeds within data centers (e.g., 100 Gbps Ethernet and above) to provide faster traffic rates within the data center is a way to address traffic growth. However, increasing network speeds can involve even more packet processing activity, which uses processor resources that could otherwise be used for other tasks.
[0020] Some solutions reduce CPU utilization and accelerate packet processing by offloading tasks to network controller hardware, including dedicated hardware, which is limited to current workloads and may not have the flexibility to accommodate different workloads or packet processing activities in the future.
[0021] Some solutions attempt to reduce packet processing overhead through protocol simplification, but still use significant CPU utilization to perform packet processing.
[0022] System Overview: Various embodiments provide a switch that attempts to reduce server processor utilization and reduce or control the growth of east-west traffic in a data center while providing sufficiently fast packet processing. Various embodiments provide a switch with infrastructure offload capabilities that comprehensively include one or more CPUs or other accelerator devices. Various embodiments provide a switch with specific packet processing network interface card (NIC) capabilities to enable the switch to perform packet processing or network termination, freeing up server CPUs to perform other tasks. The switch may include or have access to server-class processors, switch blocks, accelerators, offload engines, ternary content addressable memories (TCAMs), and packet processing pipelines. The packet processing pipelines may be programmable using P4 or other programming languages. The switch may be connected to one or more CPUs or host servers using various connections. For example, direct attach copper (DAC), fiber optic cable, or other cables may be used to connect the switch to one or more CPUs, compute hosts, or servers, including servers in a rack. In some examples, the length of the connection may be less than 6 feet (approximately 1.8 meters) to reduce the bit error rate (BER). It should be noted that references to switches may refer to multiple connected switches or distributed switches, and that a rack may include multiple switches that logically divide the rack into two half racks or into pods (e.g., one or more racks).
[0023] Various embodiments of the rack switch may be configured to perform one or more of: (1) telemetry aggregation over high-speed connections, such as packet transmission rates, response latencies, cache misses, and virtualization execution environment requests; (2) orchestration of server resources connected to the switch based at least on the telemetry; (3) orchestration of virtual execution environments running on various servers based at least on the telemetry; (4) network termination and protocol processing; (5) completion of memory transactions by retrieving data associated with the memory transaction and providing the data to a requestor or forwarding the memory transaction to a target from which the data associated with the memory transaction can be retrieved; (6) caching data for access by one or more servers in the rack or group of racks; (7) management of Memcached resources in the switch; (8) execution of one or more virtualization execution environments to perform packet processing (e.g., header processing according to applicable protocols); (9) management of virtualization execution environment execution in the switch or servers or both for load balancing or redundancy; or (10) migration of virtualization execution environments between the switch and servers or from server to server. Thus, improvements to rack switch operation can free up server CPU cycles for use for billable or value-added services.
[0024] Various embodiments may terminate network processing at the switch instead of the server. For example, the switch may perform protocol termination, decryption, decapsulation, acknowledgment (ACK), integrity checking, and network-related tasks may be performed by the switch rather than handled by the server. The switch may include dedicated offload engines for known protocols or computations and may be extensible or programmable via software or field-programmable gates (FPGAs) to handle new or vendor-specific protocols to flexibly support future needs.
[0025] Network termination at the switch can reduce or eliminate the transfer of data for processing by multiple VEEs potentially on different servers, or even different racks, for service function chain processing. The switch can perform the network processing and, after processing, provide the resulting data to a destination server in the rack.
[0026] In some examples, a switch can manage memory input / output (I / O) requests by directing I / O requests to a target device instead of the server determining the target device and directing the I / O request to the server to transmit the I / O request to another server or target device. The server may include a memory pool, a storage pool or server, a compute server, or may provide other resources. Various embodiments may be used in scenarios where Server 1 issues an I / O request to access memory, Server 2 accesses near memory, and Server 3 accesses far memory (e.g., two-level memory (2LM), memory pooling, or thin memory provisioning). For example, a switch may receive a request from Server 1 requesting a read or write to memory intended for System 2. The switch may be configured to identify that the memory address referenced by the request is within memory associated with Server 3, and the switch can forward the request to Server 3 instead of sending the request to Server 2, which may transmit the request to Server 3. As such, the switch can reduce the time it takes to complete a memory transaction. In some instances, the switch may perform caching of data on the same rack to reduce east-west traffic for subsequent requests for the data.
[0027] Note that the switch may notify Server 2 that an access to Server 3's memory has occurred so that Server 2 and Server 3 may maintain coherency or consistency of data associated with the memory addresses. If Server 2 writes a cache line or posts a dirty (modified) cache line, a coherency protocol and / or a producer-consumer model may be used to maintain consistency of the data stored on Server 2 and Server 3.
[0028] In some examples, the switch can perform orchestration, hypervisor functions, and manage service chaining functions. The switch can orchestrate the processor and memory resources and execution of virtual execution environments (VEEs) across a rack of servers to present the rack's aggregated resources as a single composite server. For example, the switch can allocate use of compute threads, memory threads, and accelerator threads for execution by one or more VEEs.
[0029] In some examples, the switch may be located at the top of the rack (TOR) or middle of the rack (MOR) relative to the connected servers to reduce the length of the connection between the switch and the servers. For example, with a switch located at the TOR (e.g., furthest from the floor of the rack), the servers connect to the switch so that the copper cables from the servers to the rack switch are contained within the rack. The switch can link the rack to the data center network using fiber optic cables that run from the rack to the aggregation area. For an MOR switch location, the switch is located toward the center of the rack between the bottom and top of the rack. Other rack locations for the switch, such as at the end of the row (EOR), can be used.
[0030] 1A illustrates an exemplary switch system. Switch 100 may include or have access to switch circuit 102 communicatively coupled to port circuits 104-0 through 104-N. Port circuits 104-0 through 104-N can receive packets and provide the packets to switch circuit 102. If port circuits 104-0 through 104-N are Ethernet-enabled, port circuits 104-0 through 104-N may include a physical layer interface (PHY) (e.g., a physical medium attachment (PMA) sublayer, a physical medium dependent (PMD), a forward error correction (FEC), and a physical coding sublayer (PCS)), a media access control (MAC) encoding or decoding, and a reconciliation sublayer (RS). An optical / electrical signal interface can provide electrical signals to the network ports. Modules can be constructed using standard mechanical and electrical form factors, such as small form-factor pluggable (SFP), quad small form-factor pluggable (QSFP), quad small form-factor pluggable double density (QSFP-DD), micro QSFP, or OSFP (octa small format pluggable) interfaces, or other form factors, as described in Annex 136C of IEEE Standard 802.3cd-2018 and references therein.
[0031] Packet may be used herein to refer to various formatted collections of bits that may be transmitted across a network, e.g., an Ethernet frame, an IP packet, a TCP segment, a UDP datagram, etc. Also, as used in this document, references to the L2, L3, L4, and L7 layers (or layers 2, 3, 4, and 7) may refer to the second data link layer, the third network layer, the fourth transport layer, and the seventh application layer, respectively, of the OSI (Open Systems Interconnection) layer model.
[0032] A flow may be a sequence of packets transferred between two endpoints, which generally represents a single session using a known protocol. Thus, a flow may be identified by a defined set of N tuples; for routing purposes, a flow may be identified by tuples that identify endpoints, e.g., source and destination addresses. For content-based services (e.g., load balancers, firewalls, command detection systems, etc.), a flow may be identified with higher granularity by using five or more tuples (e.g., source address, destination address, IP protocol, transport layer source port, and destination port). Packets within a flow are expected to have the same set of tuples in their packet headers. A flow may be unicast, multicast, anycast, or broadcast.
[0033] The switch circuitry 102 may provide connectivity to, from, and between multiple servers, and may perform one or more of traffic aggregation and action table matching for routing, tunneling, buffering, VxLAN routing, Network Virtualization using Generic Routing Encapsulation (NVGRE), Generic Network Virtualization Encapsulation (Geneve) (e.g., a currently draft Internet Engineering Task Force (IETF) standard), and access control lists (ACLs) to allow or restrict packets from proceeding.
[0034] Processors 108-0 through 108-M may be coupled to switch circuit 102 via respective interfaces 106-0 through 106-M. Interfaces 106-0 through 106-M may provide low-latency, high-bandwidth memory-based interfaces, such as Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), memory interfaces (e.g., any type of double data rate (DDRx), CXL.io, CXL.cache, or CXL.mem), and / or network connections (e.g., Ethernet or InfiniBand). In cases where a memory interface is used, the switch may be identified as a memory address.
[0035] One or more of the processor modules 108-0 through 108-M may represent a server including a CPU, random access memory (RAM), persistent or non-volatile storage, and accelerators, and a processor module may be one or more servers in a rack. For example, the processor modules 108-0 through 108-M may represent multiple separate physical servers communicatively coupled to the switch 100 using connections. A physical server may be separate from another physical server by providing different physical CPU devices, random access memory (RAM) devices, persistent or non-volatile storage devices, or accelerator devices. However, separate physical servers may include devices with the same performance specifications. As used herein, a server may refer to a physical server or a composite server that aggregates resources from one or more separate physical servers.
[0036] Processor modules 108-0 through 108-M and processor 112-0 or 112-1 may include one or more cores and system agent circuitry. A core may be an execution core or computational engine capable of executing instructions. A core may have access to its own cache and read-only memory (ROM), or multiple cores may share a cache or ROM. Cores may be homogeneous (e.g., same processing function) and / or heterogeneous (e.g., different processing functions). Core frequency or power consumption may be adjustable. Any type of inter-processor communication technique may be used, such as, but not limited to, messages, inter-processor interrupts (IPIs), and inter-processor communications. Cores may be connected in any type of fashion, such as, but not limited to, a bus, a ring, or a mesh. Cores may be coupled via an interconnect to a system agent (uncore).
[0037] The system agent may include a shared cache, which may include any type of cache (e.g., level 1, level 2, or last level cache (LLC)). The system agent may include one or more of a memory controller, a shared cache, a cache coherency manager, an arithmetic logic unit, a floating point unit, a core or processor interconnect, or a bus or link controller. The system agent or uncore may provide one or more of a direct memory access (DMA) engine connection, a non-cache coherent master connection, data cache coherency and cache request coordination between cores, or an Advanced Microcontroller Bus Architecture (AMBA) function. The system agent or uncore may manage the receive and transmit priorities and clock speeds of the fabric and memory controller.
[0038] The cores may be communicatively connected using a high-speed interconnect compatible with, but not limited to, the Intel Quick Path Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, or Compute Express Link (CXL). The number of core tiles is not limited to this example and may be any number, such as 4 and 8.
[0039] As described in more detail herein, the orchestration control plane, Memcached server, and one or more virtualization execution environments (VEEs) may execute on one or more of processor modules 108-0 through 108-M, or on processor 112-0 or 112-1.
[0040] VEE may include at least a virtual machine or container. A virtual machine (VM) may be software that runs an operating system and one or more applications. A VM may be defined by a specification, configuration files, virtual disk files, nonvolatile random access memory (NVRAM) settings files, and log files and is backed by the physical resources of a host computing platform. A VM may be an OS or application environment installed on software that mimics dedicated hardware. End users have the same experience on a virtual machine as they would on dedicated hardware. Dedicated software called a hypervisor fully emulates the CPU, memory, hard disk, network, and other hardware resources of a PC client or server, allowing virtual machines to share resources. A hypervisor may emulate multiple virtual hardware platforms isolated from each other, allowing virtual machines to run Linux® and Windows® server operating systems on the same underlying physical host.
[0041] A container can be a software package of applications, configurations, and dependencies that ensures that an application runs reliably in one computing environment versus another. A container can share an operating system installed on a server platform and run as an isolated process. A container can be a software package that includes everything needed to run the software, including system tools, libraries, and settings.
[0042] Various embodiments provide driver software for various operating systems (e.g., VMWare®, Linux®, Windows® Server, FreeBSD, Android®, MacOS®, iOS®, or any other operating system) for an application or VEE to access switch 100. In some examples, the driver may present the switch as a peripheral device. In some examples, the driver may present the switch as a network interface controller or network interface card. For example, the driver may provide VEE with the ability to configure or access the switch as a PCIe endpoint. In some examples, a virtual function driver, such as an adaptive virtual function (AVF), may be used to access the switch. One example of an AVF is described in at least the “Intel® Ethernet® Adaptive Virtual Function Specification,” Revision 1.0 (2018). In some examples, VEE can interact with the driver to turn on or off any of the features in the switch described herein.
[0043] Device drivers (e.g., NDIS-Windows, NetDev-Linux, etc.) running on processor modules 108-0 through 108-M can couple to switch 100 and expose the functionality of switch 100 to a host operating system (OS) or any OS running in VEE. Applications or VEE can configure or access switch 100 using SIOV, SR-IOV, MR-IOV, or PCIe transactions. By incorporating PCIe endpoints as interfaces to switch 100, switch 100 can be enumerated on any of processor modules 108-0 through 108-M as a PCIe Ethernet device or a CXL device as a locally attached Ethernet device. For example, switch 100 can be presented as a physical function (PF) to any server (e.g., any of processor modules 108-0 through 108-M). When switch 100 resources (e.g., memory, accelerators, network, CPU) are assigned to a server, the resources will appear logically to the server as if they were attached via a high-speed link (e.g., CXL or PCIe). A server may access resources (e.g., memory or accelerators) as hot-plugged resources, or alternatively, these resources may appear as pooled resources currently available to the server.
[0044] In some examples, the processor modules 108-0 through 108-M and the switch 100 may support the use of single root I / O virtualization (SR-IOV). The PCI-SIG Single Root IO Virtualization and Sharing Specification v1.1 and its predecessor and successor versions describe the use of a single PCIe physical device under a single root port that appears as multiple separate physical devices to a hypervisor or guest operating system. SR-IOV uses physical functions (PFs) and virtual functions (VFs) to manage the overall functionality of an SR-IOV device. A PF may be a PCIe function that can configure and manage SR-IOV functionality. For example, a PF may configure or control a PCIe device, and the PF may have the ability to move data in and out of the PCIe device. For example, in the case of the switch 100, a PF is a PCIe function of the switch 100 that supports SR-IOV. The PF includes functionality to configure and manage the SR-IOV functionality of the switch 100, such as enabling virtualization and managing PCIe VFs. A VF is associated with a PCIe PF on switch 100, and a VF represents a virtualized instance of switch 100. A VF may have its own PCIe configuration space but may share one or more physical resources on switch 100, such as external network ports, with the PF and other PFs or other VFs. In other examples, the reverse relationship may be used, where any server (e.g., processor modules 108-0 through 108-M) is represented as a PF, and a VEE running on switch 100 utilizes a VF to configure or access any server.
[0045] In some examples, the platform 1900 and the NIC 1950 can communicate using Multi-Root I / O Virtualization (MR-IOV). The Multiple Root I / O Virtualization (MR-IOV) and Sharing Specification Revision 1.0 (May 12, 2008) from the PCI Special Interest Group (SIG) is a specification for sharing PCI Express (PCIe) devices between multiple computers.
[0046] In some examples, the processor modules 108-0 through 108-M and the switch 100 may support the use of Intel® Scalable I / O Virtualization (SIOV). For example, the processor modules 108-0 through 108-M may access the switch 100 as SIOV-enabled devices, or the switch 100 may access the processor modules 108-0 through 108-M as SIOV-enabled devices. An SIOV-enabled device may be configured to group its resources into multiple isolated assignable device interfaces (ADIs). Direct memory access (DMA) transfers to and from each ADI are tagged with a unique process address space identifier (PASID) number. The switch 100, processor modules 108-0 through 108-M, network controllers, storage controllers, graphics processing units, and other hardware accelerators may utilize SIOV across many virtualized execution environments. Unlike SR-IOV's coarse-grained device partitioning approach to creating multiple VFs on a PF, SIOV allows software to flexibly compose virtual devices with hardware assistance for fine-grained device sharing. Performance-critical operations on the configured virtual devices are mapped directly to the underlying device hardware, while non-critical operations are emulated through device-specific synthesis software on the host. The technical specification for SIOV is the Intel® Scalable I / O Virtualization Technical Specification, Revision 1.0 (June 2018).
[0047] Multi-tenant security may be employed in which switch 100 is granted access to some or all server resources within a rack. Access to any server by switch 100 may require the use of encryption keys, checksums, or other integrity checks. Any server may use access control lists (ACLs) to ensure that communications from switch 100 are permitted, but may filter out (e.g., drop) communications from other sources.
[0048] Examples of packet transmission using the switch 100 are described below. In some examples, the switch 100 acts as a network proxy for a VEE running on a server. The VEE running on the switch 100 can form packets for transmission using the switch 100's network connections according to any applicable communication protocol (e.g., a standardized protocol or a proprietary protocol). In some examples, the switch 100 can originate packet transmissions where a workload or VEE running on a core is within or accessible by the switch 100. The switch 100 can access a connected internal core in a manner similar to accessing any other externally connected host. One or more hosts may be located within the same chassis as the switch 100. In some examples where a VEE or service runs on the switch 100's CPU, such a VEE can originate packets for transmission. For example, if VEE runs a Memcached server on the CPU of switch 100, switch 100 may generate packets for transmission to respond to any request for data, or in the case of a cache miss, to query another server or system for the data, retrieve the data, and update its cache.
[0049] 1B illustrates an exemplary system. Switch system 130 may include or have access to switch circuit 132 communicatively coupled to port circuits 134-0 through 134-N. Port circuits 134-0 through 134-N may receive packets and provide packets to switch circuit 132. Port circuits 134-0 through 134-N may be similar to any of port circuits 104-0 through 104-N. Interfaces 136-0 through 136-M may provide communication with respective processor modules 138-0 through 138-M. As described in more detail herein, an orchestration control plane, a Memcached server, or one or more virtualized execution environments (VEEs) running any applications (e.g., web servers, databases, Memcached servers) may execute on one or more of processor modules 138-0 through 138-M. Processor modules 138-0 to 138-M may be similar to respective processor modules 108-0 to 108-M.
[0050] 1C illustrates an exemplary system. Switch system 140 may include or have access to switch circuit 142 communicatively coupled to port circuits 144-0 through 144-4. Port circuits 144-0 through 144-4 may receive packets and provide packets to switch circuit 142. Port circuits 144-0 through 144-N may be similar to any of port circuits 104-0 through 104-N. Interfaces 146-0 through 146-1 may provide communication with respective processor modules 148-0 through 148-1. As described in more detail herein, an orchestration control plane, a Memcached server, or one or more virtualized execution environments (VEEs) running any applications (e.g., web servers, databases, Memcached servers) may execute on processor module 147-0 or 147-1, or on one or more of processor modules 148-0 through 148-1. Processor modules 148-0 to 148-1 may be similar to any of processor modules 108-0 to 108-M.
[0051] 1D shows an exemplary system. In this example, aggregation switch 150 is coupled to multiple switches in different racks. A rack may include switch 152 coupled to servers 154-0 through 154-N. Another rack may include switch 156 coupled to servers 158-0 through 158-N. One or more of the switches may operate according to embodiments described herein. A core switch or other access point may connect aggregation switch 150 to the Internet for packet transmission and reception at another data center.
[0052] Note that the depiction of servers relative to switches is not intended to represent physical placement, as TOR, MOR, or any other switch position (e.g., end of line (EOR)) can be used for the servers.
[0053] The embodiments described herein are not limited to data center operations, but may be applied to operations across multiple data centers, enterprise networks, on-premise, or hybrid data centers.
[0054] Because network processing can be moved to the switch, any type of configuration that requires power cycling (e.g., after an NVM update or a firmware update (e.g., a Basic Input / Output System (BIOS), Generic Extensible Firmware Interface (UEFI), or boot loader update)) can be performed in isolation, avoiding the need to power cycle the entire switch and affecting all servers in the rack connected to the switch.
[0055] Dual Control Planes 2A shows an example overview of a system for managing resources in a rack. Various embodiments provide a switch 200 having an orchestration control plane 202 that can manage the control planes in one or more servers 210-0 through 210-N connected to the switch 200. The orchestration control plane 202 can receive SLA information 206 for one or more VEEs (e.g., 214-0-0 through 214-0-P or 214-N-0 through 214-NP), telemetry information 204 from the servers in the rack, such as resource utilization, measured device throughput (e.g., memory read or write completion time), available memory or storage bandwidth, or resource needs of the servers connected to the switch or, more broadly, in the rack. By using telemetry information 204 to affect compliance with the VEE's SLA, orchestration control plane 202 can actively control, moderate, or pause the network bandwidth allocated to the server (e.g., the data transmission rate from switch 200 to the server or from the server to switch 200), thereby moderating the rate of communications sent from or received by the VEE running on the server.
[0056] In some examples, orchestration control plane 202 can allocate one or more of compute resources, network bandwidth (e.g., between switch 200 and another switch (e.g., an aggregation switch or a switch in another rack)), and memory or storage bandwidth to any server's hypervisor (e.g., 212-0 through 212-N). For example, switch 200 can proactively manage data transmission or reception bandwidth to any VEE in the rack prior to receipt of any flow control message, but can also manage data transmission bandwidth from any VEE upon receipt of a flow control message (e.g., XON / XOFF or Ethernet PAUSE) to reduce or pause transmission of a flow. Orchestration control plane 202 can monitor the activity of all servers 210-0 through 210-N in its rack based at least on telemetry data and can manage hypervisors 212-0 through 212-N to control VEE traffic generation. For example, if congestion is detected, switch 200 may perform flow control to pause packet transmitters from either the local VEE or a remote sender. In other cases, hypervisors 212-0 through 212-N may compete for resources from orchestration control plane 202 to allocate managed VEEs, but such a scheme may not result in under-allocation of resources to some VEEs.
[0057] For example, to allocate or de-allocate resources, orchestration control plane 202 can configure a hypervisor (e.g., 212-0 or 212-N) associated with a server running one or more VEEs. For example, servers 210-0 through 210-N can run respective hypervisor control planes 212-0 through 212-N to manage the data plane for the VEEs running on the server. For a server, the hypervisor control plane (e.g., 212-0 through 212-N) can track the SLA requirements of the VEEs running on that server and manage those requirements within allocated compute resources, network bandwidth, and memory or storage bandwidth. Similarly, the VEEs can manage contention between flows within their granted resources.
[0058] Orchestration control plane 202 may be given privileges within switch 200 and servers 210-0 through 210-N to configure resource allocation to at least the servers. Orchestration control plane 202 may be protected from untrusted VEEs that may compromise the servers. Orchestration control plane 202 may monitor the VFs of the VEEs or the PFs of the servers in the NICs and shut them down if malicious activity is detected.
[0059] An example of layered configurability by orchestration control plane 202 of hypervisor control plane 212 is described below. A server's hypervisor control plane 212 (e.g., any of hypervisor control planes 212-0 through 212-N) may decide whether to configure the resources provided to the VEE and the operation of the VEE in response to receiving a physical host configuration request from an administrator, for example, orchestration control plane 202, such as as a result of updates to policies associated with the tenant on which the VEE runs.
[0060] Configurations from the orchestration control plane 202 can be classified as trusted or untrusted. The server's hypervisor control plane 212 can allow any trusted configuration to be enforced for the VEE. In some examples, bandwidth allocations, VEE transition initiation or termination, and resource allocations made by the orchestration control plane 202 can be classified as trusted. The hypervisor 212 can restrict untrusted configurations to execute certain configurations, but not certain hardware access / configuration operations that exceed the level of trust. For example, untrusted configurations cannot issue device resets, change link configurations, write to sensitive / device-wide registers, update device firmware, etc. By separating configurations into trusted and untrusted, the hypervisor 212 can neutralize a potential attack surface by filtering out untrusted requests. Additionally, the hypervisor 212 can exhibit different capabilities for each of its different VEEs, thus allowing the host / provider to isolate tenants as needed.
[0061] 2B shows an exemplary overview of the various management hierarchies. In representation 250, as previously described, the orchestration control plane issues trusted configurations to the hypervisor control plane of the server. Some or all commands or configurations from the orchestration control plane sent to the hypervisor control plane may be considered trusted. The hypervisor control plane provides configuration for the VEEs managed by the hypervisor.
[0062] In representation 260, the switch controls the server as if the server represented a physical function (PF) and the associated virtual functions (VF-0 through VF-N) represented a VEE. When SR-IOV is used, a bare metal server (e.g., a single-tenant server) or an OS hypervisor corresponds to a PF, and the VEE accesses the PF using their corresponding VFs.
[0063] In representation 270, the orchestration control plane manages the hypervisor control plane. Indirectly, the orchestration control plane can manage the server data planes DP-0 to DP-N to control allocated resources, allocated network bandwidth (e.g., transmit or receive), and migration or termination of any VEEs.
[0064] Memory Transactions 3 illustrates an exemplary system in which a switch can respond to memory access requests. A requester device, or a VEE running in server 310 or on server 301, can request data stored in server 312. Switch 300 can receive and process the memory access request and determine a destination server or device (e.g., IP address or MAC address) to provide the memory access request for completion (e.g., read or write) in memory pool 332. Instead of providing the memory access request to server 312, which will transmit the request to memory pool 332, switch 300 can forward the request to memory pool 332.
[0065] In some examples, the switch 300 can access a mapping table 302 that shows a mapping of a memory address associated with a memory access request to a device's physical address (e.g., a destination IP address or MAC address). In some examples, the switch 300 is responsible for translating the target device's address and virtual address (provided in the memory access request) to a physical address. In some examples, the switch 300 can request a memory access (e.g., a read or write) on behalf of a requester of the memory access at the target device.
[0066] In some examples, the switch 300 can directly access the memory pool 332 to retrieve data for read operations or to write data. For example, if the server 310 requests data from the server 312, but the data is stored in the memory pool 332, the switch 300 retrieves the requested data from the memory pool 332 (or another server) and provides the data to the server 310, potentially storing the data in the memory 304 or the server 312. The switch 300 can fetch data from the memory pool 332 (or another device, server, or storage pool) by issuing a data read request to the switch 320 to retrieve the data. The memory pool 332 can be located in the same data center as the switch 300 or outside the data center. The switch 300 can store the fetched data in the memory 304 (or the server 312) to enable multiple read / write transactions with low latency by servers in the same rack as the switch 300. A high-speed connection can provide data from the memory 304 to the server 310, or vice versa. When CXL.mem is used to transfer data from server 310 to memory 304 and from memory 304 to server 310, applicable protocol rules may be followed. Switch 300 may update data from memory pool 332 if the data from memory 304 is modified.
[0067] Therefore, to significantly mitigate the latency penalty associated with retrieving data for processing by the VEE, a two-level memory (2LM) architecture can be implemented to copy the data to local memory accessible over a fast connection.
[0068] If the memory access request is a read request and the data is stored by a server or device connected to another switch (e.g., switch 320) and in another rack, switch 300 can respond to the memory request by forwarding the request to the target device that stores the data. For example, switch 300 can use packet processing 306 to change the destination IP or MAC address of the packet that carried the memory access request to the destination IP or MAC address of the target device, or encapsulate the request in another packet but maintain the destination IP or MAC address of the received memory access request.
[0069] Providing thin memory allows for less memory on a compute node and for building a memory pool that is shared by multiple compute nodes. Shared memory can be dynamically allocated and deallocated to compute nodes, with allocation set at page or cache line granularity. In aggregate, the memory allocated to all compute nodes and the memory in the shared pool may be less than the amount of memory allocated to the compute node. For example, if a thin memory provision is used for server 310, data may be stored in memory on the same rack as server 310 and potentially in a remote memory pool 332.
[0070] For a memory access request from server 310 that is a write operation, if the target device is not on the rack of switch 300, switch 300 queues the write, reports the write operation as completed to server 310 (e.g., VEE), and can then update memory pool 332 (e.g., flush posted writes) as memory bandwidth permits or as required by memory ordering and cache coherency requirements.
[0071] In some examples, switch 300 can handle memory accesses to regions of memory having corresponding addresses and, in the case of writes, the corresponding data to be written. Switch 300 can read data from or store data in memory pool 332 using remote direct memory access (e.g., InfiniBand, iWARP, RoCE, and RoCE v2), NVMe over Fabrics (NVMe-oF), or NVMe. For example, NVMe-oF, and its predecessors, successors, and proprietary variants, are described at least in NVM Express Base Specification Revision 1.4 (2019). NVMe, and its predecessors, successors, and proprietary variants, are described, for example, in NVM Express™ Base Specification Revision 1.3c (2018). If the data is stored by a server or device (e.g., memory pool 332) connected to another switch (e.g., switch 320), switch 300 can retrieve the data or write the data as if the data were stored on a server in the same rack as server 310.
[0072] In addition to the cache or memory space on each server, the switch 300 may also contribute aggregated cache space. Smart cache allocation may place data in the memory of the server that accesses it. Data that has been thrashed (e.g., accessed and modified by several servers) may be placed in the memory 304 of the switch 300 or server 312, where it can be accessed using minimal connections or Ethernet link traversals.
[0073] Memcached Example Memcached may provide a distributed memory caching system within a data center or across multiple data centers. For example, Memcached may provide a distributed database to speed up applications by alleviating database load. In some examples, dedicated servers may be used as Memcached servers to consolidate resources across servers (e.g., over Ethernet) and cache frequently accessed data to speed up access to that data. In various embodiments, a switch may manage data stored as part of a Memcached object, data, or string storage in at least some memory resources in the servers connected to the switch.
[0074] FIG. 4A shows an example of Memcached servers running on a server (system 400) and on a switch (system 450). The use of Memcached allows for faster delivery of frequently requested data by using a hash lookup instead of a database (or any other complex) query, although database queries may be used in any embodiment. The first request for data may be relatively slow because it results in the retrieval of the data. Future requests for the same data may be faster because the data is stored and can be served from the data server. In system 400, the requester may be a different rack in a row of a data center, a client / server on a different row within the data center, or an external request from outside the data center. The request may be received at aggregation switch 402 and provided to switch 404 using an Ethernet link. Switch 404 may in turn provide the request using an Ethernet link to Memcached server 408 running on server 406-0, which in turn provides the request for data to server 406-1. Even though data server 406-1 is in the same rack as Memcached server 406-0, there are multiple Ethernet connections within the same rack to provide the desired data. The Ethernet connections may contribute to east-west traffic within the data center.
[0075] In system 450, a request may be received at aggregation switch 402 and provided to switch 452 using an Ethernet link. Switch 452 runs Memcached server 408 using one or more processors and determines the server device that stores the requested data. If the data is stored in the same rack to which switch 452 provides connectivity (e.g., using PCIe, CXL, DDRx), the request may be provided to server 460-1 and does not contribute to east-west traffic. If the requester is in the same rack (e.g., server 460-N), the request may be processed internally to switch 454 and does not travel over Ethernet to be fulfilled because switch 454 is the network endpoint. In the case of a cache miss (e.g., the data is not stored in server 460-1), in some scenarios, the data may be retrieved from another server (e.g., 460-0) over a connection.
[0076] For example, switch 452 may run Memcached in VEE running on the switch, aggregating resources across racks over high speed connections into a virtual pool of combined cache and memory.
[0077] Additionally, with switch 452 handling NIC endpoint operations, all requests can be automatically routed through Memcached server 408 running in a VEE running on switch 452, eliminating the need for client requesters to maintain a list of Memcached servers. The Memcached server VEE can automatically update its cache (e.g., shown as data in server 460-1) based on how it is configured to improve data locality to the requester and further reduce latency.
[0078] Figure 4B shows the Ethernet packet flow for a single request. Each arrow represents the traversal of an Ethernet link and its contribution to east-west or north-south traffic. For system 400, in the case of a cache miss, this results in the traversal of a total of 10 Ethernet links (or other formats) in the event of a cache miss, where the data is not available at the data server. The requester sends the request to the aggregation switch, which provides the request to the switch, which in turn provides the request to the Memcached server. The Memcached server provides the request through the switch to be sent to the data server. The data server responds through the switch to the Memcached server by indicating that the data does not exist. The Memcached server receives the cache miss response, which causes the Memcached server to update its cache with the data so that the next request for that data does not result in a cache miss. The Memcached server provides the data to the requester, even in the case of a cache miss.
[0079] If the Memcached server is in a different rack in the data center than the rack that stores the data, for a request to be fulfilled, the request travels to the different rack and the response is provided to the Memcached server. However, the switch may issue an Ethernet request to the rack that stores the data. In some examples, the switch may bypass the Memcached server and request the data directly from the data source.
[0080] For system 450, a requester provides a request to the switch through an aggregation switch, which accesses the Memcached server and data within its rack via a connection (e.g., PCIe, CXL, DDRx), and provides response data for the requester to the requester via the aggregation switch. In this example, traversal of four Ethernet links occurs. By providing the Memcached service at the switch, network access to databases on other racks can be reduced, and further, by performing Memcached data location lookups at the switch, east-west traffic within the rack can be reduced. In some cases, if the data is cached in the switch's memory (e.g., memory 304) or in a server in the rack, the switch can directly provide the requested data in response to the request. In the event of a cache miss, less Ethernet communication is performed by system 450 because a server in the same rack can access the data through switch 452 (FIG. 4A) using a high-speed connection (e.g., PCIe, CXL, DDR, etc.) to obtain the data to be cached.
[0081] Network termination at the switch FIG. 5A illustrates an exemplary system in which packets may terminate at a switch. The packets may be received by the switch 502, for example, from an aggregation switch. The packets may be Ethernet-compatible and may use any type of transport layer (e.g., Transmission Control Protocol (TCP), Data Center TCP (DCTCP), User Datagram Protocol (UDP), Quick User Datagram Protocol Internet Connection (QUIC)). Various embodiments of the switch 502 may execute one or more VEEs (e.g., 504 or 506) to terminate packets by performing network protocol activities. For example, the VEEs 504 or 506 may perform one or more of the following network protocol processing or termination for the switch 502: segmentation, reassembly, acknowledgement (ACK), negative acknowledgement (NACK), packet retransmission identification and request, congestion management (e.g., transmitter flow control), Secure Sockets Layer (SSL) or Transport Layer Security (TLS) termination for HTTP and TCP. When a memory page is entered (e.g., at the socket layer), the page may be copied to a destination server on the rack using a high-speed connection and corresponding protocol (e.g., CXL.mem) for access by the bare-metal host or VEE. In some examples, the protocol processing VEE 504 or 506 may perform network service chain features, such as firewall, network address translation (NAT), intrusion protection, decryption, evolved packet core (EPC), encryption, packet filtering based on virtual local area network (VLAN) tags, encapsulation, etc.
[0082] For example, switch 502 can execute protocol processing VEE 504 and 506 when the switch's processor utilization is low. Additionally or alternatively, protocol processing VEE may execute on the computational resources of one or more servers in a rack. Switch 502 may include a packet buffer or have access to it via a high-speed connection for receiving or transmitting packets.
[0083] In some examples, the VEE 504 or 506 can perform packet protocol termination or network termination for at least some received packets at the switch 502. For example, the VEE 504 or 506 can perform packet processing at any of Layers 2-4 of the Open Systems Interconnection Model (OSI Model) (e.g., the Data Link Layer, the Network Layer, or the Transport Layer (e.g., TCP, UDP, QUIC)). Additionally or alternatively, the VEE 504 or 506 can perform packet processing at any of Layers 5-7 of the OSI Model (e.g., the Session Layer, the Presentation Layer, or the Application Layer).
[0084] In some examples, the VEE 504 or 506 may provide tunnel endpoints by performing tunnel initiation or termination by providing encapsulation or decapsulation of technologies such as, but not limited to, Virtual Extensible LAN (VXLAN) or Network Virtualization using Generic Routing Encapsulation (NVGRE).
[0085] In some examples, the VEE 504 or 506 or any device (e.g., programmable or fixed function) in the switch 502 may perform one or more of large receive offload (LRO), large send / segmentation offload (LSO), TCP segmentation offload (TSO), transport layer security (TLS) offload, receive side scaling (RSS) to allocate queues or cores to handle payload, dedicated queue allocation, or another layer protocol processing.
[0086] LRO may refer to a switch 502 (e.g., a VEE 504 or 506 or a fixed or programmable device) that reassembles incoming network packets, forwards the packet contents (e.g., payload) into larger contents, and forwards the resulting larger contents but fewer packets for access by a host system or VEE. LSO may refer to the switch 502 (e.g., VEE 504 or 506) or server 510-0 or 510-1 (e.g., VEE 514-0 or 514-1) creating a multi-packet buffer and providing the contents of the buffer to the switch 502 (e.g., VEE 504 or 506 or a fixed or programmable device) to be divided into separate packets for transmission. TSO may allow the switch 502 or server 510-0 or 510-1 to construct larger TCP messages (or other transport layer) (e.g., 64 KB long), and the switch 502 (e.g., VEE 504 or 506 or a fixed or programmable device) segment the messages into smaller data packets for transmission.
[0087] TLS is defined at least in The Transport Layer Security (TLS) Protocol version 1.3, RFC8446 (August 2018). TLS offloading may refer to the offloading of encryption or decryption of content to the switch 502 (e.g., a VEE 504 or 506 or a fixed or programmable device) in accordance with TLS. The switch 502 may receive data for encryption from a server 510-0 or 510-1 (e.g., a VEE 514-0 or 514-1) or a VEE 504 or 506 and perform encryption of the data in one or more packets before transmission of the encrypted data. The switch 502 may receive packets and decrypt the contents of the packets before forwarding the decrypted data to the server 510-0 or 510-1 for access by the VEE 514-0 or 514-1 or a VEE 504 or 506. In some examples, any type of encryption or decryption may be performed by the switch 502, such as, but not limited to, Secure Sockets Layer (SSL).
[0088] RSS may refer to the switch 502 (e.g., the VEE 504 or 506 or a fixed or programmable device) calculating a hash or making another decision based on the contents of a received packet to determine and select which CPU or core will process the payload from the received packet. Other manners of distributing payloads to cores may be performed. In some examples, the switch 502 (e.g., the VEE 504 or 506 or a fixed or programmable device) may perform RSS to select a non-uniform memory access (NUMA) node having a core and memory pair to identify the NUMA node to store and process the payload from the received packet. In some examples, the switch 502 (e.g., the VEE 504 or 506 or a fixed or programmable device) may perform RSS to select a core on the switch 502 or a server to store and process the payload from the received packet. In some examples, the switch 502 may perform RSS to assign one or more cores to perform packet processing (on the switch 502 or a server).
[0089] In some examples, the switch 502 can allocate dedicated queues in memory to applications or VEEs according to application device queues (ADQs) or similar techniques. The use of ADQs can dedicate queues to applications or VEEs, allowing them to be accessed exclusively by the applications or VEEs. ADQs can prevent network traffic contention, which occurs when different applications or VEEs attempt to access the same queue, causing locking or contention and making packet availability performance (e.g., latency) unpredictable. ADQs also provide quality of service (QoS) control of dedicated application traffic queues for received packets or packets to be transmitted. For example, using ADQs, the switch 502 can allocate packet payload content to one or more queues, which are mapped for access by software such as an application or VEE. In some examples, the switch 502 can utilize ADQs to dedicate one or more queues for packet header processing operations.
[0090] 5C illustrates an exemplary method of NUMA node, CPU, or server selection by switch 502 (e.g., VEE 504 or 506 or a fixed or programmable device). For example, resource selector 572 performs a hash calculation (e.g., a hash calculation on a packet flow identifier) on the header of a received packet to determine an indirection table stored in switch 502 that maps to a queue (e.g., from queue 576), which in turn maps to a NUMA node, CPU, or server. Resource mapping 574 may include the indirection table and the mapping to the queue, as well as an indicator of which connection (e.g., a CXL link, a PCIe connection, or a DDR interface) to use to copy the header and / or payload of the received packet to memory (or cache) associated with the selected NUMA node, CPU, or server. In some cases, resource selector 572 performs RSS to select a NUMA node, CPU, or server. For example, resource selector 572 may select CPU 1 in NUMA node 0 on server 580-1 to process the header and / or payload of a received packet. A NUMA node on a server may have its own connection to switch 570 to allow writing to memory within the server without traversing the UPI bus. The VEE may run on one or more cores or CPUs, and the VEE may process the received payload.
[0091] Referring again to FIG. 5A , to perform packet protocol processing, the VEE 504 or 506 may execute processes based on the Data Plane Development Kit (DPDK), Storage Performance Development Kit (SPDK), Open Data Plane, Network Functions Virtualization (NFV), Software-Defined Networking (SDN), Evolved Packet Core (EPC), or 5G Network Slicing. Some example implementations of NFV are described in the European Telecommunications Standards Institute (ETSI) specifications or the Open Source NFV Management and Orchestration (MANO) from the ETSI Open Source Mano (OSM) group. A virtual network function (VNF) may include a service chain or sequence of virtualized tasks running on general-purpose configurable hardware, such as a firewall, Domain Name System (DNS), cache, or Network Address Translation (NAT), and may run in the VEE. VNFs may be linked together as a service chain. In some examples, the EPC is at least a 3GPP-specific core architecture for Long Term Evolution (LTE) access. 5G network slicing can provide multiplexing of virtualized, independent logical networks over the same physical network infrastructure.
[0092] In some examples, any protocol processing, protocol termination, network termination, or offload operations may be performed by a programmable or fixed function device in the switch 502 instead of, or in addition to, using a VEE running in the switch 502.
[0093] In some examples, processing packets at the switch 502 may enable faster packet disposition (e.g., forward or drop) decisions compared to when the packet disposition decisions were made at the server. Additionally, when packets are dropped, bandwidth utilization of the connection between the server and the switch may be conserved. If a packet is identified as being associated with malicious activity (e.g., a DDoS attack), the packet is dropped to protect the server from potential exposure to the malicious activity.
[0094] The VEEs 504 and 506 running on the computational resources of the switch 502 complete the network processing, and the resulting data is transferred to a data buffer for the VEE 514-0 or 514-1 via DMA, RDMA, PCIe, or CXL.mem, regardless of the network protocol used to transmit the packets. In other words, the VEEs 504 and 506 running on the computational resources of the switch 502 may act as proxy VEEs for the respective VEEs 514-0 or 514-1 running on the respective servers 510-0 and 510-1. For example, the VEEs 504 or 506 may perform protocol stack processing. The VEEs running on the switch 502 (e.g., VEEs 504 or 506) may provide socket buffer entries and buffered data for the hosts (e.g., 512-0 or 512-1).
[0095] Based at least on successful protocol layer processing and the absence of any deny conditions in the ACL, the payload from the packet may be copied to a memory buffer (e.g., 512-0 or 512-1) at the destination server (e.g., 510-0 or 510-1). For example, the VEEs 504 and 506 may cause the packet payload to be copied to a buffer associated with the VEE processing the packet payload (e.g., VEEs 514-0 and 514-1) for performance of direct memory access (DMA) or RDMA operations. A descriptor may be a data structure provided to the switch 500 by the orchestrator or the VEEs 514-0 and 514-1 to identify an area of memory or cache available for receiving the packet. In some examples, the VEEs 504 and 506 may complete a receive descriptor indicating the destination location of the packet payload in a buffer at the destination server (e.g., 510-0 or 510-1) and copy the completed receive descriptor for access by the VEE processing the packet payload.
[0096] In some examples, the switch 502 may run a VEE for each of the VEEs running on the servers in its rack or an optimized subset. In some examples, the subset of VEEs running on the switch may correspond to VEEs running on servers with low latency requirements, are primarily network-intensive, or other criteria.
[0097] In some examples, switch 502 is connected to servers 510-0 and 510-1 using connections that allow switch 502 to access all CPUs, memory, and storage in the rack. An orchestration layer can manage VEE on some or all of switch 502 and resource allocation to any server in the rack.
[0098] The VEEs 514-0 and 514-1 running on the respective servers 510-0 and 510-1 may select a mode for notifying data availability, such as polling mode, busy poll, or interrupt. Polling mode may involve the VEE polling for new packets by actively sampling the status of a buffer to determine if a new packet has arrived. Busy polling may allow socket layer code to poll the receive queue and disable network interrupts. An interrupt causes a running process to save its state and execute the process associated with the interrupt (e.g., processing the packet or data).
[0099] Instead of operating in polling mode for packet processing, the servers 510-0 or 510-1 in the rack can receive interrupts. Interrupts can be issued to the servers by the switch 502 for higher-level transactions, rather than for each packet. For example, if the VEE 514-0 or 514-1 operates a database, an interrupt can be provided to the VEE 514-0 or 514-1 by the VEE 504 or 506 when a record update is complete, even if the record update uses many packets. For example, if the VEE 514-0 or 514-1 operates a web server, an interrupt can be provided to the VEE 514-0 or 514-1 by the VEE 504 or 506 after receiving a complete form, even if one or more packets provide the form. Polling of received packets or data can be used in any case.
[0100] FIG. 5B illustrates an example of a VEE configuration on a server and a switch. In this example, VEE 552 executes on switch 550 and performs protocol processing or packet protocol termination on packets having payloads to be processed by VEE 562 executing on server 550. VEE 552 may execute on one or more cores on switch 550. For example, VEE 552 may process packet headers of packets utilizing TCP / IP or other protocols or combinations of protocols. VEE 552 may write the payload of processed packets to socket buffer 566 in server 560 via socket interface 554-socket interface 564 and high-speed connection 555 (e.g., PCIe, CXL, DDRx (x is an integer)). Socket buffer 566 may be represented as a memory address. An application (e.g., running on server 560 executing VEE 562) may access socket buffer 566 to use or process data. The VEE 552 can provide TCP Offload Engine (TOE) operation without requiring any protocol stack modifications (eg, TCP Chimney).
[0101] In some examples, network termination occurs at the VEE 552 of the switch 550, and the server 560 does not receive any packet headers in the socket buffer 566. For example, the VEE 552 of the switch 550 may perform protocol processing for Ethernet, IP, and transport layer (e.g., TCP, UDP, QUIC) headers, and such headers are not provided to the server 560.
[0102] Some applications have their own headers or markers, and switch 550 may forward or copy those headers or markers to socket buffer 566 in addition to the payload data. Thus, VEE 562 may access data in socket buffer 566 regardless of the protocol used to transmit the data (e.g., Ethernet, Asynchronous Transfer Mode (ATM), Synchronous Optical Networking (SONET), Synchronous Digital Hierarchy (SDH), Token Ring, etc.).
[0103] In some examples, VEEs 552 and 562 may be related as a network service chain (NSC) or service function chain (SFC), whereby VEE 552 passes data to VEE 562 in a trusted environment, or at least by sharing memory space. Network service VEE 552 may be chained to application service VEE 562, and VEEs 552 and 562 may have shared memory buffers for passing Layer 7 data.
[0104] Telemetry Aggregation In a data center, device (e.g., compute or memory) utilization and performance, as well as software performance, can be measured to assess server utilization and whether adjustments to resources or software should be made. Examples of telemetry data include device temperature readings, application monitoring, network utilization, disk space utilization, memory consumption, CPU utilization, fan speeds, and specific telemetry streams from VEE applications running on the server. For example, telemetry data can include processor or core utilization statistics, device and partition input / output statistics, memory utilization information, storage utilization information, bus or interconnect utilization information, counters for processor hardware registers that count instructions executed, cache misses incurred, predicted branch misses, and other hardware events, or performance monitoring events. For workload requests being executed or completed, one or more of the following may be collected: telemetry data, such as, but not limited to, Top-down Micro-Architecture Method (TMAM), execution of Unix® System Activity Reporter (SAR) commands, and output from the Emon command monitoring tool, which can profile application and system performance. However, additional information may be collected, such as output from various monitoring tools, including, but not limited to, output from the Linux perf command, Intel PMU toolkit, Iostat, VTune Amplifier, or use of monCli or other Intel Benchmark Install and Test Tool (Intel® BITT) tools. Other telemetry data may be monitored, such as, but not limited to, power consumption and inter-process communication. Various telemetry techniques may be used, such as those described with respect to collected daemons.
[0105] As VEEs in the data center transmit telemetry data to a central orchestrator, bandwidth requirements can be enormous, and east-west traffic can be overwhelmed by telemetry data. In some cases, key performance indicators (KPIs) are provided by the server, and if one of these KPIs indicates a problem, the server sends a more robust set of telemetry to allow for more detailed investigation.
[0106] In some embodiments, when high-speed connections are used between the server and the switch, much more information can be passed from the server to the switch without burdening the network with east-west traffic. The switch can collect a minimum set of telemetry (e.g., KPIs) from the server without burdening the network with excessive east-west traffic overhead. However, in some examples, the server may send KPIs to the switch unless more data or history is required, such as in the case of an error. An orchestrator running on the switch (e.g., orchestration control plane 202 of FIG. 2A ) can use the expanded telemetry data (e.g., telemetry 204 of FIG. 2A ) to determine the available capacity of each of the servers on its rack and can provide improved multi-server job placement to maximize performance by taking into account telemetry from multiple servers.
[0107] Running and Migrating VEE 6 shows an example of a switch that runs an orchestration control plane to manage which devices run VEEs. An orchestration control plane 604 running on switch 602 may monitor the performance of one or more VEEs in terms of compliance with applicable SLAs, and if a VEE is not complying with or is close to non-compliance with an SLA requirement (e.g., application availability (e.g., 99.999% on workdays and 99.9% on evenings or weekends), maximum acceptable response time to a query or other call, requirements for the actual physical location of stored data, or encryption or security requirements), the orchestration control plane 604 can instantiate one or more new VEEs to balance the workload among the VEEs. As the workload subsides, excess VEEs are torn down or deactivated, freeing up resources to be allocated to another VEE (or the same VEE at a later time) for use when the load reaches capacity. For example, the workload may include at least any type of activity, such as protocol processing and network termination for a packet or Memcached server, a database, or a web server, etc. For example, VEE 606 may perform protocol processing, and if the workload increases, multiple instances of VEE 606 may be instantiated on switch 602.
[0108] In some examples, the orchestration control plane 604 running on the switch 602 may determine whether to migrate any VEE running on the switch 602 or a server to execution on another server. For example, migration may depend on a shutdown or restart of the switch 602 on which the VEE is running, which may cause the VEE to run on the server. For example, VEE migration may depend on a shutdown or restart of the server on which the VEE is running, which may cause the VEE to run on the switch 602 or another server.
[0109] In some examples, the orchestration control plane 604 can determine whether to run the VEE on a particular processor or migrate the VEE between switches 602 or between any of the servers 608-0 through 608-N. The VEE 606 or VEE 610 may migrate from a server to a switch, from a switch to a server, or from a server to another server, as needed. For example, the VEE 606 may run on the switch 602 for a short period in connection with the server being rebooted, and the VEE may be migrated back to the rebooted server or another server.
[0110] In some examples, the switch 602 may run a virtual switch (vSwitch) that enables communication between VEE running on the switch 602 or any servers connected to the switch 602. Virtual switches may include Microsoft Hyper-V, Open vSwitch, VMware vSwitches, and the like.
[0111] The switch 602 may support S-IOV, SR-IOV, or MR-IOV for its VEE. In this example, the VEE running on the switch 602 utilizes resources in one or more servers via S-IOV, SR-IOV, or MR-IOV. S-IOV, SR-IOV, or MR-IOV may allow connectivity or bus sharing across VEEs. In some examples, when the VEE running on the switch 602 acts as a network termination proxy VEE, one or more corresponding VEEs in the rack and in the switch 602 run on one or more servers. The VEE running on the switch 602 can process packets, and the VEE running on the server or cores on the switch 602 can run applications (e.g., databases and web servers). The use of S-IOV, SR-IOV, or MR-IOV (or other schemes) may allow server resources to be configured so that physically distributed servers appear logically as one system, but with tasks divided so that network processing occurs on the switch 602.
[0112] As previously mentioned, switch 602 may use a high-speed connection to at least some of the resources on one or more servers 608-0 through 608-N in the rack, thereby providing access to resources from any of the servers in the rack to VEE 606 running on switch 602. Orchestration control plane 604 can efficiently allocate VEE to resources and is not limited to being able to run on a single server, but also on switch 602 and servers 608-0 through 608-N. This feature allows potentially constrained resources, such as accelerators, to be optimally allocated.
[0113] 7A shows an example of VEE migration from one server to another. For example, live VEE migration (e.g., Microsoft® HyperV or VMware® vSphere) can be performed to migrate an active VEE. At (1), the VEE is transmitted to a TOR switch. At (2), the VEE is transmitted through a data center core network, and at (3), the VEE is transmitted to a TOR switch in another rack. At (4), the VEE is transmitted to a server, where it can begin running in another hardware environment.
[0114] FIG. 7B shows an example of VEE migration. In this example, VEE may run on a switch using the resources of the switch and connected servers in the rack. At (1), VEE is transmitted from the switch to the core network. At (2), VEE is transmitted to another switch for execution. The other switch may use the resources of the switch and connected servers in the rack. In another example, as in the example of FIG. 7A, the destination of the VEE may be a server. Thus, by running VEE on a switch with expanded server resources, there are fewer steps in VEE migration, and VEE can begin execution sooner in the scenario of FIG. 7B than in the scenario of FIG. 7A.
[0115] FIG. 8A illustrates an exemplary process. The process may be performed by a switch with an enhanced processor according to various embodiments. At 802, the switch may be configured to execute an orchestration control plane. For example, the orchestration control plane may manage the compute, memory, and software resources of the switch and one or more servers connected to the switch that are in the same rack as the switch. The servers may run a hypervisor that controls the execution of a virtualized execution environment and may or may not allow configuration by the orchestration control plane. For example, a connection may be used to provide communication between the switch and the servers. The orchestration control plane may receive telemetry from servers in the rack via the connection without the telemetry contributing to east-west traffic within the data center. Various examples of connections are described herein.
[0116] At 804, the switch may be configured to execute a virtualization execution environment to perform protocol processing for at least one virtualization execution environment executing on the server. Various examples of protocol processing are described herein. In some examples, the switch may perform network termination of received packets and provide data from the received packets to a memory buffer of the server or the switch. However, the virtualization execution environment may perform any type of operation related to or unrelated to packet or protocol processing. For example, the virtualization execution environment may run a Memcached server or retrieve data from a memory device in another rack or outside the data center, or from a web server or database.
[0117] At 806, the orchestration control plane may determine whether to change the allocation of resources to the virtualized execution environments. For example, based on whether applicable SLAs for the virtualized execution environments or the flow of packets processed by the virtualized execution environments are met or not, the orchestration control plane may determine whether to change the allocation of resources to the virtualized execution environments. For scenarios in which the SLAs are not met or are deemed likely to be violated, at 808 the orchestration control plane may add additional computing, networking, or memory resources for use by the virtualized execution environments or may instantiate one or more additional virtualized execution environments to assist in processing. In some examples, the virtualized execution environments may be migrated from switches to servers to improve resource availability.
[0118] In scenarios where the SLA is met, the process returns to 806. Note that in some cases, when packet processing activity is low or idle, the orchestration control plane may deallocate computational resources available to a virtualized execution environment. In some examples, when the SLA is met, a virtualized execution environment may be migrated from a switch to a server to provide resources for another virtualized execution environment to utilize.
[0119] FIG. 8B illustrates an exemplary process. The process may be performed by a processor-enhanced switch according to various embodiments. At 820, a virtualization execution environment executing on the switch may perform packet processing of the received packet. Packet processing may include one or more of header parsing, flow identification, segmentation, reassembly, acknowledgement (ACK), negative acknowledgement (NACK), packet retransmission identification and request, congestion management (e.g., transmitter flow control), checksum verification, decryption, encryption, or secure tunneling (e.g., Transport Layer Security (TLS) or Secure Sockets Layer (SSL)), or other operations. For example, the packet and protocol processing virtualization execution environment may perform polling, busy polling, or rely on interrupts to detect newly received packets received in a packet buffer from one or more ports. Based on the detection of the newly received packet, the virtualization execution environment processes the received packet.
[0120] At 822, a virtualization execution environment executing on the switch may determine whether to make the data from the packet available or discard it. For example, if the packet is subject to a deny status in an access control list (ACL), the packet may be discarded. If it is determined that the data is to be provided to the next virtualization execution environment, the process may proceed to 824. If it is determined that the packet is to be discarded, the process may proceed to 826, where the packet is discarded.
[0121] At 824, the virtualization execution environment may notify the virtualization execution environment executing on the server that the data is available and provide the data for access by the virtualization execution environment executing on the server. The virtualization execution environment executing on the switch may cause the data to be copied to a buffer accessible to the virtualization execution environment executing on the server. For example, direct memory access (DMA), RDMA, or other direct copy schemes may be used to copy the data to the buffer. In other examples, the data is made available to the virtualization execution environment executing on the switch for processing.
[0122] 8C illustrates an exemplary process that may be performed by a processor-enhanced switch according to various embodiments. At 830, the switch may be configured to execute a virtualization execution environment to retrieve data from or copy data to devices in the same or a different rack as the switch.
[0123] At 832, the virtualized execution environment may be configured to include information of a destination device associated with the memory address. For example, the information may indicate a translation of a destination device or server (e.g., an IP address or MAC address) corresponding to the memory address in the memory transaction. For example, in the case of a read memory transaction, the device or server may store data corresponding to the memory address, and data may be read from the memory address at the device or server. For example, in the case of a write memory transaction, the device or server may receive and store data corresponding to a address for the write transaction.
[0124] At 834, the switch may receive a memory access request from a server in the same rack. At 836, a virtualization execution environment executing on the switch may manage the memory access request. In some examples, the performance of 836 may include performance of 838, where the virtualization execution environment executing on the switch may forward the memory access request to a destination server. In some examples, if the memory access request is sent to a server but the server does not store the requested data, the switch may redirect the memory access request to a destination server that stores the requested data instead of sending the memory access request to a server that would in turn send the request to the destination server.
[0125] In some examples, the capabilities of 836 may include capabilities of 840, where a virtualized execution environment executing on the switch may execute a memory access request. If the memory access request is a write command, the virtualized execution environment may write data to a memory address corresponding to the memory access request in a device in the same or a different rack. If the memory access request is a read command, the virtualized execution environment may copy data from a memory address corresponding to the memory access request in a device in the same or a different rack. For example, remote direct memory access may be used to write or read the data.
[0126] In the case of a read request, the switch may cache the data locally for access by the servers connected to the switch. When the orchestration control plane manages the memory resources of the switch and the servers, the retrieved data may be stored in a memory device of the switch or any server so that any virtualization execution environment running on any server in the rack can access or modify the data. For example, a memory device accessible to the switch and the servers in the rack can access the data as near memory. If the data is updated, the switch may write the updated data to the memory device to store the data.
[0127] For example, block 840 may be performed in a scenario where a switch runs a Memcached server and data is stored on a server in the same rack as the switch. The Memcached server running on the switch may respond to memory access requests corresponding to cache misses by retrieving data from another server and storing the retrieved data in a cache in the memory or storage of the rack.
[0128] 9 illustrates a system that utilizes a switch to manage resources within the system and may implement other embodiments described herein. System 900 includes a processor 910 that provides processing, operational management, and instruction execution for system 900. Processor 910 may include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), processing core, or other processing hardware or combination of processors to provide processing for system 900. Processor 910 controls the overall operation of system 900 and may be or include one or more programmable general-purpose or programmable special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or the like, or a combination of such devices.
[0129] In one example, system 900 includes an interface 912 coupled to processor 910, which may represent a high-speed or high-throughput interface for system components requiring a higher-bandwidth connection, such as memory subsystem 920 or graphics interface component 940, or accelerator 942. Interface 912 represents interface circuitry that may be a standalone component or integrated on a processor die. If present, graphics interface 940 interfaces to a graphics component for providing a visual display to a user of system 900. In one example, graphics interface 940 may drive a high-definition (HD) display that provides output to the user. High resolution may refer to a display having a pixel density of approximately 100 PPI (pixels per inch) or greater and may include formats such as Full HD (e.g., 1080p), Retina display, 4K (ultra-high definition or UHD), or others. In one example, the display may include a touchscreen display. In one example, graphics interface 940 generates a display based on data stored in memory 930, based on operations performed by processor 910, or both. In one example, graphics interface 940 generates a display based on data stored in memory 930, or based on operations performed by processor 910, or both.
[0130] Accelerators 942 may be programmable or fixed-function offload engines that can be accessed or used by processor 910. For example, one accelerator in accelerators 942 may provide cryptographic services such as compression (DC) functions, public key encryption (PKE), ciphers, hash / authentication functions, decryption, or other functions or services. In some embodiments, additionally or alternatively, one accelerator in accelerators 942 provides a field selection controller function described herein. In some cases, various devices (e.g., a connector to a motherboard or circuit board that contains the CPU and provides an electrical interface with the CPU) accelerator 942 may be integrated into or connected to the CPU. For example, the accelerator 942 may include programmable processing elements such as a single-core processor or a multi-core processor, a graphics processing unit, a single-level cache of a logical execution unit or a multi-level cache of a logical execution unit, a functional unit usable for independently executing programs or threads, an application-specific integrated circuit (ASIC), a neural network processor (NNP), programmable control logic, and a field-programmable gate array (FPGA). The accelerator 942 may provide multiple neural networks, CPUs, processor cores, general-purpose graphics processing units, or graphics processing units for use by artificial intelligence (AI) or machine learning (ML) models. For example, the AI models may use or include any or a combination of reinforcement learning schemes, Q-learning schemes, deep Q-learning, or Asynchronous Advantage Actor-Critic (A3C), combinatorial neural networks, recurrent combinatorial neural networks, or other AI or ML models. Multiple neural networks, processor cores, or graphics processing units may be made available for use by the AI or ML models.
[0131] Memory subsystem 920 represents the main memory of system 900 and provides storage for data values used to execute code or routines executed by processor 910. Memory subsystem 920 may include one or more memory devices 930, such as one or more varieties of random access memory (RAM), such as read-only memory (ROM), flash memory, DRAM, or other memory devices, or a combination of such devices. Memory 930 stores and hosts, among other things, an operating system (OS) 932 for providing a software platform for executing instructions within system 900. Additionally, applications 934 may execute on the software platform of OS 932 from memory 930. Applications 934 represent programs having their own operating logic for performing one or more functions. Processes 936 represent agents or routines that provide auxiliary functionality to OS 932, one or more applications 934, or a combination thereof. OS 932, applications 934, and processes 936 provide the software logic for providing functionality for system 900. In one example, memory subsystem 920 includes memory controller 922, which is a memory controller for generating and issuing commands to memory 930. It will be understood that memory controller 922 may be a physical part of processor 910 or a physical part of interface 912. For example, memory controller 922 may be an integrated memory controller integrated into circuitry with processor 910.
[0132] Although not specifically shown, it will be understood that system 900 may include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, an interface bus, or others. A bus or other signal lines may communicatively or electrically couple components to each other or communicatively and electrically couple components. A bus may include a physical communication line, a point-to-point connection, a bridge, an adapter, a controller, or other circuitry, or a combination thereof. A bus may include, for example, one or more of a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) Standard 1394 bus (Firewire).
[0133] In one example, system 900 includes an interface 914 that may be coupled to interface 912. In one example, interface 914 represents an interface circuit that may include standalone components and integrated circuits. In one example, multiple user interface and / or peripheral components are coupled to interface 914. Network interface 950 provides system 900 with the ability to communicate with remote devices (e.g., servers or other computing devices) over one or more networks. Network interface 950 may include an Ethernet adapter, a wireless interconnection component, a cellular network interconnection component, a Universal Serial Bus (USB), or other wired or wireless standards-based or proprietary interface. Network interface 950 may transmit data to devices in the same data center or rack or to remote devices, and may include transmitting data stored in memory. Network interface 950 may receive data from remote devices, and the remote devices may include storing the received data in memory. Various embodiments may be used in conjunction with network interface 950, processor 910, and memory subsystem 920.
[0134] In one example, system 900 includes one or more input / output (I / O) interfaces 960. I / O interface 960 may include one or more interface components through which a user interacts with system 900 (e.g., voice, alphanumeric, haptic / touch, or other interfaces). Peripheral interface 970 may include any hardware interface not specifically mentioned above. In general, a peripheral refers to a device that is dependently connected to system 900. A dependent connection is one in which system 900 provides a software and / or hardware platform on which operations are performed and with which a user interacts.
[0135] In one example, system 900 includes a storage subsystem 980 for storing data in a nonvolatile manner. In one example, in some system implementations, at least certain components of storage 980 may overlap with components of memory subsystem 920. Storage subsystem 980 includes storage device 984, which may be or include any conventional medium for storing large amounts of data in a nonvolatile manner, such as one or more magnetic, solid-state, or optical-based disks, or a combination thereof. Storage 984 holds code or instructions and data 986 in a persistent state (e.g., values are retained despite interruption of power to system 900). Memory 930 is typically an execution or operating memory that provides instructions to processor 910, although storage 984 may generally be considered “memory.” While storage 984 is nonvolatile, memory 930 may include volatile memory (e.g., the value or state of data is indeterminate if power is interrupted to system 900). In one example, storage subsystem 980 includes a controller 982 for interfacing with storage 984. In one example, controller 982 may be a physical part of interface 914 or processor 910, or may include circuitry or logic in both processor 910 and interface 914.
[0136] Volatile memory is memory whose state (and therefore the data stored therein) is indeterminate when power to the device is interrupted. Dynamic volatile memory requires the data stored in the device to be refreshed in order to maintain its state. An example of dynamic volatile memory includes DRAM (Dynamic Random Access Memory) or some variant such as Synchronous DRAM (SDRAM). Another example of volatile memory includes cache or static random access memory (SRAM). As described herein, the memory subsystem may be compatible with many memory technologies, such as DDR3 (Double Data Rate Version 3, first released by JEDEC (Semiconductor Engineering Association) on June 27, 2007). DDR4 (DDR Version 4, initial specification published by JEDEC in September 2012), DDR4E (DDR Version 4), LPDDR3 (Low Power DDR Version 3, JESD209-3B, published by JEDEC in August 2013), LPDDR4 (LPDDR Version 4, JESD209-4, first published by JEDEC in August 2014), WIO2 (Wide Input / Output Version 2, JESD229-2, first published by JEDEC in October 2014), HBM (High Bandwidth Memory, JESD325, first published by JEDEC in October 2013), LPDDR5 (currently under review by JEDEC), HBM2 (HBM Version 2), currently under review by JEDEC, etc., or any other combination of memory technologies, and technologies based on derivatives or extensions of such specifications. For example, DDR or DDRx may refer to any version of DDR, where x is an integer.
[0137] A non-volatile memory device (NVM) is a memory whose state is determined even when power to the device is interrupted. In one embodiment, an NVM device may include a block-addressable memory device, such as NAND technology, or more specifically, multi-threshold level NAND flash memory (e.g., single-level cell ("SLC"), multi-level cell ("MLC"), quad-level cell ("QLC"), tri-level cell ("TLC"), or some other NAND). NVM devices can include byte-addressable write-in-place 3D cross-point memory devices such as single-level or multilevel phase change memory (PCM) or switched phase change memory (PCMS) or other byte-addressable write-in-place NVM devices (also referred to as persistent memory), NVM devices using chalcogenide-based phase change materials (e.g., chalcogenide glasses), resistive memories including metal oxide-based, oxygen vacancy-based, and conductive bridge random access memories (CB-RAM), nanowire memories, ferroelectric random access memories (FeRAM, FRAM®), magnetoresistive random access memories (MRAM) incorporating memristor technology, spin-transfer torque (STT) MRAM, spintronic magnetic junction memory-based devices, magnetic tunnel junction (MTJ)-based devices, DW (domain wall) and SOT (spin orbit transfer)-based devices, thyristor-based memory devices, or any combination of the above or other memories.
[0138] A power source (not shown) provides power to the components of system 900. More specifically, the power source typically interfaces with one or more power supplies within system 900 to provide power to the components of system 900. In one example, the power supply includes an AC-DC (alternating current to direct current) adapter that plugs into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source, such as an external AC-DC converter. In one example, the power source or power supply includes wireless charging hardware for charging via proximity to a charging field. In one example, the power source can include an internal battery, an AC supply, a motion-based power supply, a solar power supply, or a fuel cell power source.
[0139] In one example, system 900 may be implemented using interconnected compute sleds of processors, memory, storage, network interfaces, and other components. High-speed interconnects such as PCIe, Ethernet, or optical interconnects (or combinations thereof) may be used.
[0140] In one example, system 900 may be implemented using compute threads interconnected with processors, memory, storage, network interfaces, and other components. High-speed interconnects such as Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWarp), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Quick User Datagram Protocol Internet Connection (QUIC), RDMA over Converged Ethernet (RoCE), Peripheral Component Interconnect Express (PCIe), Intel Quick Path Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omni-Path, Compute Express Link (CXL), HyperTransport, High-Speed Fabric, NVLink, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and variations thereof may be used.
[0141] Embodiments herein may be implemented in various types of computing, smartphone, tablet, personal computer, and networking equipment, such as switches, routers, rack, and blade servers, for example, used in data center and / or server farm environments. Servers used in data centers and server farms include array server configurations, such as rack-based servers or blade servers. These servers are interconnected to communicate via various network provisioning, for example, partitioning sets of servers into local area networks (LANs), which have appropriate switching and routing capabilities between the LANs to form private intranets. For example, cloud hosting functions may typically utilize large data centers with numerous servers. A blade comprises a separate computing platform configured to perform server-type functions, i.e., a "server on a card." Thus, each blade includes components common to traditional servers, including a main printed circuit board (mainboard) that provides internal wiring (e.g., buses) for coupling appropriate integrated circuits (ICs) with other components mounted on the board.
[0142] FIG. 10 illustrates an environment 1000 including multiple computing racks 1002, each including a top-of-rack (ToR) switch 1004, a pod manager 1006, and multiple pooled system drawers. Switch embodiments herein can be used to manage device resources, virtual execution environment operation, and data locality for a VEE (e.g., storing data within the same rack that runs the VEE). In general, a pooled system drawer may include a pooled compute drawer and a pooled storage drawer. Optionally, a pooled system drawer may also include a pooled memory drawer and a pooled input / output (I / O) drawer. In the illustrated embodiment, the pooled system drawer includes an Intel® XEON® pooled computer drawer 1008, an Intel® ATOM pooled compute drawer 1010, a pooled storage drawer 1012, a pooled memory drawer 1014, and a pooled I / O drawer 1016. Each of the pooled system drawers is connected to the ToR switch 1004 via a high-speed link 1018, such as a 40 Gigabit per second (Gb / s) or 100 Gb / s Ethernet link or a 100+ Gb / s Silicon Photonics (SiPh) optical link.
[0143] As shown by their connections to network 1020, multiple ones of the computing racks 1002 may be interconnected via their ToR switches 1004 (e.g., to a pod-level switch or a data center switch). In some embodiments, groups of computing racks 1002 are managed as separate pods via pod managers 1006. In one embodiment, a single pod manager is used to manage all racks in a pod. Alternatively, a distributed pod manager may be used for pod management operations.
[0144] The environment 1000 further includes a management interface 1022 that is used to manage various aspects of the environment, including managing rack configurations, the corresponding parameters of which are stored as rack configuration data 1024.
[0145] FIG. 11 illustrates an exemplary network element that may be used by embodiments of the switch herein. Various embodiments of the switch may perform any of the operations of the network interface 1100. In some examples, the network interface 110 may be implemented as a network interface controller, a network interface card, a host fabric interface (HFI), or a host bus adapter (HBA). The network interface 1100 may be coupled to one or more servers using a bus, PCIe, CXL, or DDRx. In some examples, the network interface 1100 may be embodied as part of a system-on-chip (SoC) that includes one or more processors, or may be included in a multi-chip package that also includes one or more processors.
[0146] The network interface 1100 may include a transceiver 1102, a processor 1104, a transmit queue 1106, a receive queue 1108, a memory 1110, a bus interface 1112, and a DMA engine 1152. The transceiver 1102 is capable of receiving and transmitting packets conforming to an applicable protocol, such as Ethernet as described in IEEE 802.3, although other protocols may be used. The transceiver 1102 can receive packets from a network and transmit packets to the network via a network medium (not shown). The transceiver 1102 may include a PHY circuit 1114 and a media access control (MAC) circuit 1116. The PHY circuit 1114 may include encoding and decoding circuitry (not shown) for encoding and decoding data packets in accordance with applicable physical layer specifications or standards. The MAC circuit 1116 may be configured to assemble data to be transmitted into packets that include destination and source addresses along with network control information and error detection hash values. The processor 1104 may be any combination of processors, cores, graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or other programmable hardware devices that enable programming of the network interface 1100. For example, the processor 1104 may provide identification of resources used to execute a workload and generate a bitstream for execution on selected resources. For example, a "smart network interface" may use the processor 1104 to provide packet processing capabilities at the network interface.
[0147] The packet allocator 1124 may provide distribution of received packets for processing by multiple CPUs or cores using time slot allocation or RSS as described herein. If the packet allocator 1124 uses RSS, the packet allocator 1124 may calculate a hash or make another determination based on the contents of the received packets to determine which CPU or core should process the packets.
[0148] Interrupt coalescing 1122 can perform interrupt moderation, whereby the network interface's interrupt coalescing 1122 waits until multiple packets arrive or until a timeout expires before generating an interrupt to the host system to process the received packets. Receive segment coalescing (RSC) can be performed by the network interface 1100, whereby some of the incoming packets can be combined into segments of packets. The network interface 1100 provides the fused packets to the application.
[0149] Instead of copying the packet to an intermediate buffer at the host and then using another copy operation from the intermediate buffer to the destination buffer, the direct memory access (DMA) engine 1152 may copy the packet header, packet payload, and / or descriptors directly from host memory to the network interface, or vice versa. In some examples, the DMA engine 1152 may perform writes to data in any cache, such as by using data direct I / O (DDIO).
[0150] The memory 1110 may be any type of volatile or non-volatile memory device and may store any queues or instructions used to program the network interface 1100. The transmit queue 1106 may contain data or references to data for transmission by the network interface. The receive queue 1108 may contain data or references to data received from the network by the network interface. The descriptor queue 1120 may contain descriptors that reference data or packets in the transmit queue 1106 or receive queue 1108. The bus interface 1112 may provide an interface to a host device (not shown). For example, the bus interface 1112 may be compatible with PCI, PCI Express, PCI-x, PHY Interface for PCI Express (PIPE), Serial ATA, and / or USB-compatible interfaces (although other interconnect standards may be used).
[0151] In some examples, the network interfaces and other embodiments described herein may be used in connection with base stations (e.g., 3G, 4G, 5G, etc.), macro base stations (e.g., 5G networks), pico stations (e.g., IEEE 802.11-enabled access points), nano stations (e.g., for point-to-multipoint (PtMP) applications), on-premises data centers, off-premises data centers, edge network elements, fog network elements, and / or hybrid data centers (e.g., data centers that use virtualization, cloud, and software-defined networking to deliver application workloads across physical data centers and distributed multi-cloud environments).
[0152] Various examples may be implemented using hardware elements, software elements, or a combination of both. In some examples, hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, ASICs, PLDs, DSPs, FPGAs, memory units, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. In some examples, software elements may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The decision whether to implement an example using hardware and / or software elements may depend on various factors, such as desired computation rate, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints desired for a given implementation. A processor may be a hardware state machine, digital control logic, a central processing unit, or any combination of one or more hardware, firmware, and / or software elements.
[0153] Some examples may be implemented using an article of manufacture or at least one computer-readable medium. The computer-readable medium may include a non-transitory storage medium for storing logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. In some examples, the logic may include various software elements such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.
[0154] According to some examples, a computer-readable medium may include a non-transitory storage medium that stores or maintains instructions that, when executed by a machine, computing device, or system, cause the machine, computing device, or system to perform methods and / or operations according to the described examples. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. The instructions may be implemented according to a predefined computer language, method, or syntax to instruct a machine, computing device, or system to perform a particular function. The instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0155] One or more aspects of at least one example may be implemented by representative instructions stored on at least one machine-readable medium that represent various logic within a processor, which, when read by a machine, computing device, or system, causes the machine, computing device, or system to manufacture logic that performs the techniques described herein. Such representations, known as "IP cores," may be stored on tangible machine-readable media and supplied to various customers or manufacturing facilities that load them into manufacturing machines that actually create the logic or processor.
[0156] Appearances of the phrase "one example" or "example" do not necessarily all refer to the same example or embodiment. Any aspect described herein may be combined with any other or similar aspect described herein, whether or not these aspects are described with reference to the same drawing or element. The division, omission, or inclusion of block functions described in the accompanying drawings does not imply that hardware components, circuits, software, and / or elements for implementing those functions are necessarily divided, omitted, or included in the embodiments.
[0157] Some examples may be described using the terms "coupled" or "connected," along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.
[0158] As used herein, the terms “first,” “second,” etc., do not denote any order, quantity, or importance, but rather are used to distinguish one element from another. As used herein, the terms “a” and “an” do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced items. As used herein, the term “asserted” as used with respect to a signal refers to a state of the signal where the signal is active, which may be achieved by applying a logic level of either logic 0 or logic 1 to the signal. The terms “follow” or “after” may refer to something that immediately follows or follows some other event. Other sequences of steps may also be performed by alternative embodiments. Furthermore, additional steps may be added or deleted depending on the specific application. Any combination of changes may be used, and one of ordinary skill in the art, having the benefit of this disclosure, will recognize many variations, modifications, and alternative embodiments of the present disclosure.
[0159] Disjunctive language, e.g., "at least one of X, Y, or Z," is understood within its commonly used context to indicate that an item, term, etc., can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z, respectively, be present. Conjunctive language, such as "at least one of X, Y, and Z," should also be understood to mean any combination including X, Y, Z, or X, Y, and / or Z, unless specifically specified to the contrary. Illustrative examples of the devices, systems, and methods disclosed herein are provided below. An embodiment of the device, system, and method may include any one or more of the examples described below, and any combination thereof.
[0160] Example 1 includes a method comprising a switch device for a rack of two or more physical servers, the switch device coupled to the two or more physical servers, the switch device performing packet protocol processing termination of received packets and providing payload data from the received packets, not including the headers of the received packets, to a destination buffer of a destination physical server in the rack.
[0161] Example 2 includes any example in which the switch device comprises at least one central processing unit, and the at least one central processing unit performs packet processing operations on the received packets.
[0162] Example 3 includes any example in which a physical server executes at least one virtualized execution environment (VEE), and the at least one central processing unit executes the VEE for packet processing of packets containing data accessed by the physical server executing the VEE.
[0163] Example 4 includes any example in which the switch device stores a mapping of memory addresses and corresponding destination devices, and in which the switch device executes memory transactions based on receiving the memory transactions from physical servers in the rack.
[0164] Example 5 includes any example in which the switch device performing the memory transaction includes, in the case of a read request, the switch device obtaining data from a physical server connected to the rack or another device in a different rack based on the mapping, and storing the data in a memory managed by the switch device.
[0165] Example 6 includes any example in which the switch device stores a mapping of memory addresses and corresponding destination devices, and based on receiving a memory transaction from a physical server in the rack, transmits the memory transaction to a destination server in another rack based on a memory address associated with the memory transaction according to the mapping, receives a response to the memory transaction, and stores the response in a memory of the rack.
[0166] Example 7 includes any example where the switch device comprises at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers that are part of the rack, and the control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of: assigning execution of a virtualized execution environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on at least one central processing unit of the switch device; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a physical server in the rack for access by a VEE executing on a physical server in the rack.
[0167] Example 8 includes any example where the switch device includes at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers that are part of the rack, and the control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack and selectively terminates or migrates a VEE to execution on another physical server in the rack or on the switch device.
[0168] Example 9 includes any example embodiment and includes the apparatus comprising: a switch including at least one processor, wherein the at least one processor performs packet termination processing on received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server.
[0169] Example 10 includes any example in which the at least one processor executes a virtualized execution environment (VEE), and the VEE performs the packet termination processing.
[0170] Example 11 includes any example in which, based on receiving a memory transaction from a physical server through the connection, the at least one processor executes the memory transaction based on a mapping of a memory address to a corresponding destination device.
[0171] Example 12 includes any example where, to perform the memory transaction, the at least one processor, in the case of a read request, retrieves data from a physical server connected to the at least one processor or another device in a different rack through the connection and stores the data in memory managed by the at least one processor.
[0172] Example 13 includes any example in which, based on receiving a memory transaction from a physical server in a rack associated with the switch, the at least one processor performs transmission of the memory transaction to the destination server based on a memory address associated with the memory transaction associated with a destination server in another rack according to a mapping of memory addresses and corresponding destination devices, the at least one processor accesses a response to the memory transaction, and the at least one processor causes the response to be stored in a memory of the rack.
[0173] Example 14 includes any example where the at least one processor executes a control plane of one or more physical servers that are part of a rack associated with the switch, and the control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of: assigning execution of a virtualized execution environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on the at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory in the rack for access by a VEE running on a physical server in the rack.
[0174] Example 15 includes any example in which the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch, and the control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack and selectively terminates or migrates a VEE to execution on another physical server in the rack or on at least one processor that is part of the switch.
[0175] Example 16 includes any example in which the connection is compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR).
[0176] Example 17, including any example, includes at least one non-transitory computer-readable medium having stored thereon instructions that, when executed by a switch, cause the switch to execute a control plane at the switch to collect telemetry data from one or more physical servers, and, based on the telemetry data, perform one or more of: assigning execution of a Virtualization Execution Environment (VEE) to a physical server in a rack including the switch; migrating a VEE from a physical server in the rack to execution on the at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a server in the rack for access by a VEE running on a physical server in the rack.
[0177] Example 18 includes any example comprising stored instructions that, when executed by a switch, cause the switch to store a mapping of memory addresses and corresponding destination devices, and based on receiving a memory transaction from a physical server over a connection and based on the mapping of memory addresses and corresponding destination devices, cause the switch to retrieve data from a physical server or another device in a different rack connected to the switch over the connection and store the data in a memory managed by the switch.
[0178] Example 19 includes any example comprising stored instructions that, when executed by a switch, cause the switch to store a mapping of memory addresses and corresponding destination devices; based on receiving a memory transaction from a server in a rack associated with the switch, the switch transmits the memory transaction to a destination server in another rack based on a memory address associated with the memory transaction according to the mapping; the switch receives a response to the memory transaction; and the switch stores the response in a memory of the rack.
[0179] Example 20 includes any example in which the connections between the switch and one or more physical servers in the rack are compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR).
[0180] Example 21 includes any example embodiment, including a network device comprising: circuitry for performing network protocol termination of received packets; at least one Ethernet port; and a plurality of connections connected to different physical servers in a rack, wherein the circuitry for performing network protocol termination of received packets provides payloads of received packets without associated headers to the physical server.
[0181] [Other possible items] [Item 1] 1. A method comprising: a switch device for a rack of two or more physical servers, the switch device coupled to the two or more physical servers, the switch device performing packet protocol processing termination of received packets and providing payload data from the received packets, not including the headers of the received packets, to a destination buffer of a destination physical server in the rack. [Item 2] Item 10. The method of item 1, wherein the switch device comprises at least one central processing unit, and the at least one central processing unit performs packet processing operations on the received packets. [Item 3] A physical server runs at least one Virtualization Execution Environment (VEE), the at least one central processing unit executing a VEE for packet processing of packets containing data accessed by the physical server executing the at least one VEE; The method described in item 2. [Item 4] the switch device stores a mapping of memory addresses to corresponding destination devices; upon receiving a memory transaction from a physical server in the rack, the switch device executes the memory transaction. The method according to item 1. [Item 5] The switch device performing the memory transaction, In the case of a read request, the switch device obtains data from a physical server connected to the rack or another device in a different rack based on the mapping, and stores the data in a memory managed by the switch device. The method according to item 4. [Item 6] the switch device stores a mapping of memory addresses to corresponding destination devices; Upon receiving memory transactions from physical servers in the rack, transmitting the memory transaction to a destination server in another rack based on a memory address associated with the memory transaction, the memory transaction being associated with the destination server in the other rack according to the mapping; receiving a response to the memory transaction; storing the response in a memory of the rack; The method according to item 1. [Item 7] the switch device comprises at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers associated with the rack; the control plane collects telemetry data from the one or more physical servers, and based on the telemetry data, performs one or more of the following: assigning execution of a virtualized execution environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on at least one central processing unit of the switch device; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a physical server in the rack for access by a VEE executing on a physical server in the rack; The method according to item 1. [Item 8] the switch device includes at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers that are part of the rack; The control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack, and selectively terminates or migrates a VEE to execution on another physical server in the rack or on the switch device. The method according to item 1. [Item 9] a switch including at least one processor that performs packet termination for received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server; An apparatus comprising: [Item 10] 10. The apparatus of claim 9, wherein the at least one processor executes a virtualization execution environment (VEE), and the VEE performs the packet termination processing. [Item 11] and upon receiving a memory transaction from a physical server over the connection, the at least one processor executes the memory transaction based on a mapping of a memory address to a corresponding destination device. Item 9. The device according to item 9. [Item 12] To perform the memory transaction, the at least one processor: In the case of a read request, data is obtained from a physical server connected to the at least one processor or another device in a different rack through the connection, and the data is stored in a memory managed by the at least one processor. Item 12. The device according to item 11. [Item 13] upon receiving a memory transaction from a physical server in a rack associated with the switch; the at least one processor performs a transmission of the memory transaction to the destination server based on a memory address associated with the memory transaction associated with a destination server in another rack according to the mapping of memory addresses to corresponding destination devices; the at least one processor accessing a response to the memory transaction; and The at least one processor causes the response to be stored in a memory of the rack. Item 13. The device according to item 12. [Item 14] the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch; The control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of the following: assigning execution of a virtualized execution environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on the at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of the rack of servers for access by a VEE running on a physical server in the rack. Item 9. The device according to item 9. [Item 15] the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch; The control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack and selectively terminates or migrates a VEE to execution on another physical server in the rack or on at least one processor that is part of the switch. Item 9. The device according to item 9. [Item 16] 10. The apparatus of claim 9, wherein the connection is compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR). [Item 17] At least one non-transitory computer-readable medium storing instructions that, when executed by a switch, cause the switch to: a control plane running on the switch to collect telemetry data from one or more physical servers and, based on the telemetry data, perform one or more of: assigning execution of a virtualized execution environment (VEE) to a physical server in a rack that includes the switch; migrating a VEE from a physical server in the rack to execution on the at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a server in the rack for access by a VEE running on a physical server in the rack. At least one non-transitory computer-readable medium. [Item 18] instructions stored on the switch, the instructions causing the switch to: storing a mapping of memory addresses to corresponding destination devices; Based on receiving a memory transaction from a physical server through the connection and based on mapping a memory address to a corresponding destination device, retrieve data from a physical server or another device in a different rack connected to the switch through the connection, and store the data in a memory managed by the switch. Item 18. At least one non-transitory computer-readable medium according to item 17. [Item 19] instructions stored on the switch, the instructions causing the switch to: storing a mapping of memory addresses to corresponding destination devices; upon receiving a memory transaction from a server in a rack associated with the switch; transmitting the memory transaction to a destination server based on a memory address associated with the memory transaction associated with a destination server in another rack according to the mapping; receiving a response to the memory transaction; storing the response in the memory of the rack; Item 18. At least one non-transitory computer-readable medium according to item 17. [Item 20] The connections between said switch and one or more physical servers in said rack are compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR). Item 18. At least one non-transitory computer-readable medium according to item 17. [Item 21] 1. A network device, comprising: circuitry that performs network protocol termination of received packets; At least one Ethernet port and and a plurality of connections connected to different physical servers within the rack, wherein the circuitry performing network protocol termination of the received packets provides the payload of the received packets, without the associated header, to the physical servers. Network devices.
Claims
1. 1. A method implemented using a packaged integrated circuit, the packaged integrated circuit being configurable for use in switching operations in association with at least one network, a plurality of graphics processing units (GPUs), a plurality of Compute Express Link (CXL.mem) memory devices, and a plurality of central processing units (CPUs), the packaged integrated circuit comprising: an interface circuit and a switch circuit, the interface circuit communicatively coupled to the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; the switch circuit is a switch circuit for a rack of two or more physical servers, the two or more physical servers including the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; The method includes using the switch circuitry to implement switching operations associated with respective data communication processes, the switching operations being performed via the interface circuitry associated with the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; wherein the switching operation includes performing packet protocol processing termination of a received packet by the switch circuit to provide payload data from the received packet, not including a header of the received packet, to a destination buffer of a destination physical server in the rack, the switch circuit being coupled to the two or more physical servers; The plurality of CXL.mem memory devices are configured in a pooled configuration; the switch circuitry performs, at least in part, the respective data communication operations associated with the at least one network in accordance with a Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) protocol; the switch circuitry performs at least in part the respective data communication operations associated with the plurality of GPUs and the plurality of CPUs according to a Peripheral Component Interconnect Express (PCIe) protocol; the switch circuitry performs, at least in part, the respective data communication operations associated with the plurality of CXL.mem memory devices in accordance with a CXL protocol; the switch circuitry implements the switching operations and / or the respective data communication operations related to aggregations of compute resources and / or accelerator resources and / or compositions of compute resources and / or accelerator resources; the switching operations and / or the respective communication processes are at least partially software programmable; the switch circuitry at least partially performs the respective data communication operations associated with the plurality of CXL.mem memory devices associated with memory page data transfers; method.
2. 2. The method of claim 1, wherein the switch circuitry comprises at least one central processing unit, the at least one central processing unit performing packet processing operations on the received packets.
3. The physical server runs at least one Virtualization Execution Environment (VEE); the at least one central processing unit executing a VEE for packet processing of packets containing data accessed by the physical server executing the at least one VEE; The method of claim 2.
4. the switch circuit stores a mapping of memory addresses to corresponding destination CXL.mem memory devices; the switch circuit executes processing for the request of the memory transaction or transfers the request of the memory transaction based on reception of a request of the memory transaction included in each of the data communication processes from the physical servers in the rack; 4. The method according to any one of claims 1 to 3.
5. the switch circuit performing the processing for the request of the memory transaction, In the case of a read request, the switch circuit retrieves data from a physical server connected to the rack or another device in a different rack based on the mapping, and stores the data in a memory managed by the switch circuit. The method of claim 4.
6. the switch circuit stores a mapping of memory addresses to corresponding destination CXL.mem memory devices; Upon receiving a request for a memory transaction included in each of the data communication processes from a physical server in the rack, transmitting the request for the memory transaction to a destination server in another rack based on a memory address associated with the request for the memory transaction, the request being associated with the destination server in the other rack according to the mapping; receiving a response to the request for the memory transaction; storing the response in a memory of the rack; 6. The method according to any one of claims 1 to 5.
7. the switch circuitry comprises at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers associated with the rack; the control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of the following: assigning execution of a Virtualization Execution Environment (VEE) to a physical server of the rack; migrating a VEE from a physical server of the rack to execution on at least one central processing unit of the switch circuit; migrating a VEE from a physical server of the rack to execution on another physical server of the rack; or allocating memory of a physical server of the rack for access by a VEE executing on the physical server of the rack; 7. The method according to any one of claims 1 to 6.
8. the switch circuitry includes at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers that are part of the rack; The control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack, and selectively terminates or migrates a VEE to execution on another physical server in the rack or on the switch circuitry.
8. The method according to any one of claims 1 to 7.
9. performing packet protocol processing termination of received packets by a switch device for two or more racks of physical servers to provide payload data from the received packets, not including headers of the received packets, to a destination buffer of a destination physical server in the rack, the switch device being coupled to the two or more physical servers; the switch device stores a mapping of memory addresses to corresponding destination devices; Upon receiving a request for a memory transaction from a physical server in the rack, the switch device performs processing on the request for the memory transaction or forwards the request for the memory transaction. method.
10. performing packet protocol processing termination of received packets by a switch device for two or more racks of physical servers to provide payload data from the received packets, not including headers of the received packets, to a destination buffer of a destination physical server in the rack, the switch device being coupled to the two or more physical servers; the switch device stores a mapping of memory addresses to corresponding destination devices; upon receiving a request for a memory transaction from a physical server in the rack; transmitting the request for the memory transaction to a destination server in another rack based on a memory address associated with the request for the memory transaction, the request being associated with the destination server in the other rack according to the mapping; receiving a response to the request for the memory transaction; storing the response in a memory of the rack; method.
11. performing packet protocol processing termination of received packets by a switch device for two or more racks of physical servers to provide payload data from the received packets, not including headers of the received packets, to a destination buffer of a destination physical server in the rack, the switch device being coupled to the two or more physical servers; the switch device comprises at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers associated with the rack; the control plane collects telemetry data from the one or more physical servers, and based on the telemetry data, performs one or more of the following: assigning execution of a Virtualization Execution Environment (VEE) to a physical server of the rack; migrating a VEE from a physical server of the rack to execution on at least one central processing unit of the switch device; migrating a VEE from a physical server of the rack to execution on another physical server of the rack; or allocating memory of a physical server of the rack for access by a VEE executing on the physical server of the rack; method.
12. performing packet protocol processing termination of received packets by a switch device for two or more racks of physical servers to provide payload data from the received packets, not including headers of the received packets, to a destination buffer of a destination physical server in the rack, the switch device being coupled to the two or more physical servers; the switch device includes at least one central processing unit, the at least one central processing unit executing a control plane for one or more physical servers that are part of the rack; The control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack, and selectively terminates or migrates a VEE to execution on another physical server in the rack or on the switch device. method.
13. 1. An apparatus comprising: The apparatus is configurable to be used for switching operations relating to at least one network, a plurality of graphics processing units (GPUs), a plurality of Compute Express Link (CXL) .mem memory devices, and a plurality of central processing units (CPUs); The device comprises: an interface circuit communicatively coupled to the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; and a switch circuit for a rack in which two or more physical servers are mounted, the switch circuit including at least one processor and implementing the switching operations related to each data communication process; the at least one processor performs packet termination for the received packet and copies payload data from the received packet, not including an associated received packet header, over the connection to a destination buffer of a destination physical server; wherein the switching operations are performed via the interface circuitry associated with the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; the two or more physical servers include the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; where: The plurality of CXL.mem memory devices are configured in a pooled configuration; the switch circuitry performs, at least in part, the respective data communication operations associated with the at least one network in accordance with a Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) protocol; the switch circuitry performs at least in part the respective data communication operations associated with the plurality of GPUs and the plurality of CPUs according to a Peripheral Component Interconnect Express (PCIe) protocol; the switch circuitry performs, at least in part, the respective data communication operations associated with the plurality of CXL.mem memory devices in accordance with a CXL protocol; the switch circuitry implements the switching operations and / or the respective data communication operations related to aggregations of compute resources and / or accelerator resources and / or compositions of compute resources and / or accelerator resources; the switching operations and / or the respective communication processes are at least partially software programmable; the switch circuitry at least partially performs the respective data communication operations associated with the plurality of CXL.mem memory devices associated with memory page data transfers; Device.
14. 14. The apparatus of claim 13, wherein the at least one processor executes a Virtualization Execution Environment (VEE), the VEE performing the packet termination processing.
15. Upon receiving a request for a memory transaction included in the respective data communication operation from a physical server through the connection, the at least one processor performs processing on the request for the memory transaction based on a mapping of a memory address to a corresponding destination CXL.mem memory device or forwards the request for the memory transaction.
15. Apparatus according to claim 13 or 14.
16. To perform the processing for the request of the memory transaction, the at least one processor: For a read request, obtain data from a physical server connected to the at least one processor or another CXL.mem memory device in a different rack through the connection, and store the data in memory managed by the at least one processor.
16. The apparatus of claim 15.
17. upon receiving the request for the memory transaction included in each of the data communication processes from a physical server in a rack associated with the switch circuit, the at least one processor performs a transmission of the request for the memory transaction to the destination server based on a memory address associated with the request for the memory transaction associated with a destination server in another rack according to the mapping of memory addresses to corresponding destination CXL.mem memory devices; the at least one processor accesses a response to the request for the memory transaction; and The at least one processor causes the response to be stored in a memory of the rack.
17. The apparatus of claim 16.
18. the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch circuit; The control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of the following: assigning execution of a Virtualized Execution Environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on at least one central processing unit of the switch circuit; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a server in the rack for access by a VEE running on a physical server in the rack.
18. Apparatus according to any one of claims 13 to 17.
19. the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch circuit; The control plane distributes execution of a virtualized execution environment (VEE) among one or more physical servers in the rack and selectively terminates or migrates a VEE to execution on another physical server in the rack or on at least one processor that is part of the switch circuitry.
19. Apparatus according to any one of claims 13 to 18.
20. 20. The apparatus of claim 13, wherein the connection is compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR).
21. a switch including at least one processor that performs packet termination for received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server; Equipped with The apparatus, wherein the at least one processor executes a Virtualization Execution Environment (VEE), the VEE performing the packet termination process.
22. a switch including at least one processor that performs packet termination for received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server; Equipped with and wherein, based on receiving a request for a memory transaction from a physical server through the connection, the at least one processor performs processing on the request for the memory transaction based on mapping a memory address to a corresponding destination device or forwards the request for the memory transaction.
23. a switch including at least one processor that performs packet termination for received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server; Equipped with the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch; the control plane collects telemetry data from the one or more physical servers and, based on the telemetry data, performs one or more of the following: assigning execution of a Virtualization Execution Environment (VEE) to a physical server in the rack; migrating a VEE from a physical server in the rack to execution on at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a server in the rack for access by a VEE running on a physical server in the rack.
24. a switch including at least one processor that performs packet termination for received packets and copies payload data from the received packets, not including associated received packet headers, over a connection to a destination buffer of a destination physical server; Equipped with the at least one processor executes a control plane for one or more physical servers that are part of a rack associated with the switch; The apparatus, wherein the control plane distributes execution of a Virtualized Execution Environment (VEE) among one or more physical servers in the rack and selectively terminates or migrates a VEE to execution on another physical server in the rack or on at least one processor that is part of the switch.
25. A computer program comprising: A switch installed in a rack containing two or more physical servers executing a control plane in the switch to collect telemetry data from one or more physical servers, and based on the telemetry data, perform one or more of the following: assigning execution of a Virtualization Execution Environment (VEE) to a physical server in a rack containing the switch; migrating a VEE from a physical server in the rack to execution on at least one central processing unit of the switch; migrating a VEE from a physical server in the rack to execution on another physical server in the rack; or allocating memory of a server in the rack for access by a VEE running on a physical server in the rack; The computer program further comprises: The switch storing a mapping of memory addresses to corresponding destination devices; based on receiving a request for a memory transaction from a physical server through the connection and based on said mapping of memory addresses to corresponding destination devices, obtaining data from a physical server or another device in a different rack connected to said switch through said connection, and storing said data in a memory managed by said switch; Run Computer program.
26. The switch storing a mapping of memory addresses to corresponding destination devices; upon receiving the request for the memory transaction from a server in a rack associated with the switch; transmitting the request for the memory transaction to the destination server based on a memory address associated with the request for the memory transaction associated with a destination server in another rack according to the mapping; receiving a response to the request for the memory transaction; storing the response in a memory of the rack; 26. A computer program product according to claim 25, which causes the computer to execute the following:
27. The connections between the switch and one or more physical servers in the rack are compatible with one or more of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or any type of Double Data Rate (DDR).
27. A computer program according to claim 25 or 26.
28. 28. A computer readable storage medium storing a computer program according to any one of claims 25 to 27.
29. 1. A network device, comprising: circuitry that performs network protocol termination of received packets; At least one Ethernet port; a plurality of connections connected to different physical servers in the rack, wherein the circuitry performing network protocol termination of received packets provides payloads of received packets without associated headers to the physical servers; the circuitry stores a mapping of memory addresses to corresponding destination devices, and upon receiving a request for a memory transaction from a physical server in the rack, the circuitry performs processing on the request for the memory transaction or forwards the request for the memory transaction. Network devices.
30. 1. A packaged integrated circuit comprising: The packaged integrated circuit is configurable for use in switching operations associated with at least one network, a plurality of graphics processing units (GPUs), a plurality of Compute Express Link (CXL) .mem memory devices, and a plurality of central processing units (CPUs), the packaged integrated circuit comprising: an interface circuit communicatively coupled to the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; and switch circuitry that implements the switching operations associated with each data communication transaction; the switching operation is performed via the interface circuitry in association with the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; where: The plurality of CXL.mem memory devices are configured in a pooled configuration; the switch circuitry performs, at least in part, the respective data communication operations associated with the at least one network in accordance with a Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) protocol; the switch circuitry performs at least in part the respective data communication operations associated with the plurality of GPUs and the plurality of CPUs according to a Peripheral Component Interconnect Express (PCIe) protocol; the switch circuitry performs, at least in part, the respective data communication operations associated with the plurality of CXL.mem memory devices in accordance with a CXL protocol; the switch circuitry implements the switching operations and / or the respective data communication operations related to aggregations of compute resources and / or accelerator resources and / or compositions of compute resources and / or accelerator resources; the switching operations and / or the respective communication processes are at least partially software programmable; the switch circuitry at least partially performs the respective data communication operations associated with the plurality of CXL.mem memory devices associated with memory page data transfers; Packaged Integrated Circuits.
31. the packaged integrated circuit comprises a system-on-chip; 31. The packaged integrated circuit of claim 30.
32. the packaged integrated circuit implements control plane / management processes associated with the switching operations and / or the respective data communication processes; 32. The packaged integrated circuit of claim 31.
33. the packaged integrated circuit implements congestion control and / or load balancing associated with the switching operations and / or the respective data communication processes; 33. The packaged integrated circuit of claim 32.
34. the plurality of GPUs are configurable to implement operations related to artificial intelligence and / or machine learning models; 34. The packaged integrated circuit of claim 33.
35. the packaged integrated circuit is included in a multi-switch network; 35. The packaged integrated circuit of claim 34.
36. the packaged integrated circuit comprises an application specific integrated circuit; 36. The packaged integrated circuit of claim 35.
37. The packaged integrated circuit is provided in a server system.
37. The packaged integrated circuit of any one of claims 30 to 36.
38. The server system is one of a plurality of server systems included in a data center system; the data center system is communicatively coupled to the at least one network; The plurality of server systems comprises the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs.
38. The packaged integrated circuit of claim 37.
39. The server system is mounted in a rack containing two or more physical servers.
39. The packaged integrated circuit of claim 38.
40. 1. A method implemented using a packaged integrated circuit, the packaged integrated circuit being configurable for use in switching operations in association with at least one network, a plurality of graphics processing units (GPUs), a plurality of Compute Express Link (CXL.mem) memory devices, and a plurality of central processing units (CPUs), the packaged integrated circuit comprising an interface circuit and a switch circuit, the interface circuit communicatively coupled to the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs, the method comprising: using the switch circuitry to implement the switching operations associated with respective data communication processes, the switching operations being performed via the interface circuitry associated with the at least one network, the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs; where: The plurality of CXL.mem memory devices are configured in a pooled configuration; the switch circuitry performs, at least in part, the respective data communication operations associated with the at least one network in accordance with a Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) protocol; the switch circuitry performs at least in part the respective data communication operations associated with the plurality of GPUs and the plurality of CPUs according to a Peripheral Component Interconnect Express (PCIe) protocol; the switch circuitry performs, at least in part, the respective data communication operations associated with the plurality of CXL.mem memory devices in accordance with a CXL protocol; the switch circuitry implements the switching operations and / or the respective data communication operations related to aggregations of compute resources and / or accelerator resources and / or compositions of compute resources and / or accelerator resources; the switching operations and / or the respective communication processes are at least partially software programmable; the switch circuitry at least partially performs the respective data communication operations associated with the plurality of CXL.mem memory devices associated with memory page data transfers; method.
41. the packaged integrated circuit comprises a system-on-chip; 41. The method of claim 40.
42. the packaged integrated circuit implements control plane / management processes associated with the switching operations and / or the respective data communication processes; 42. The method of claim 41.
43. the packaged integrated circuit implements congestion control and / or load balancing associated with the switching operations and / or the respective data communication processes; 43. The method of claim 42.
44. the plurality of GPUs are configurable to implement operations related to artificial intelligence and / or machine learning models; 44. The method of claim 43.
45. the packaged integrated circuit is included in a multi-switch network; 45. The method of claim 44.
46. the packaged integrated circuit comprises an application specific integrated circuit; 46. The method of claim 45.
47. The packaged integrated circuit is provided in a server system.
41. The method of claim 40.
48. The server system is one of a plurality of server systems included in a data center system; the data center system is communicatively coupled to the at least one network; The plurality of server systems comprises the plurality of GPUs, the plurality of CXL.mem memory devices, and the plurality of CPUs.
48. The method of claim 47.
49. The server system is mounted in a rack containing two or more physical servers.
49. The method of claim 48.
50. At least one machine-readable storage medium storing instructions for execution by at least one machine, the instructions, when executed by the at least one machine, result in performance of the method of any one of claims 40 to 49. At least one machine-readable storage medium.
Citation Information
Patent Citations
Virtual access router
JP2009027755A
Agile Data Center Network Architecture
JP2012528552A
Physical interface to virtual interface fault propagation
US10263832B1
Processing packet data using an offload engine in a service provider environment
US10412002B1
Reference Architecture For Improved Scalability Of Virtual Data Center Resources
US20130114607A1