Server offload card with SOC and FPGA
By using the SoC and FPGA on the offload card on the uninstall card in the cloud server, the problem of excessive CPU resource utilization in traditional cloud servers is solved, and computing power and server efficiency are improved.
Patent Information
- Application Number
- CN202080037234.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-03
- Filing Date
- 2020-04-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-04-16
AI Technical Summary
Because the virtual machine monitor occupies CPU resources, the cloud platform's overall computing power facing customers has decreased.
The uninstall card, including SoC and FPGA, is adopted to uninstall the virtual machine monitor function, thereby reducing the processing burden of the CPU complex and improving server efficiency.
By uninstalling the virtual machine monitor function, the use of CPU resources is reduced, the computing power of the cloud platform to customers is improved, and the efficiency and security of the server is enhanced.
Smart Images

Figure CN113841120B_ABST
Abstract
Description
Background Art
[0001] Cloud platforms such as Microsoft Azure and Amazon AWS run on a large number of physical servers (referred to herein as cloud servers) distributed across geographically dispersed data centers. A large portion of these cloud servers implement a virtualization software layer called a hypervisor that allows for the hosting of virtual machines (VMs). Among other things, this enables the Infrastructure as a Service (IaaS) scenario in which customers of the cloud platform can purchase and use VMs to execute their application workloads.
[0002] Traditionally, in each cloud server implementing a hypervisor, a certain percentage of the CPU (Central Processing Unit) cores of the cloud server are reserved for use by the hypervisor. While this reservation ensures that the hypervisor has sufficient computing resources to perform its functions, it also reduces the number of CPU cores available for use by, for example, customer VMs. On a large scale, this can result in a significant overall reduction in the cloud platform's customer-facing computing power. Summary of the Invention
[0003] A physical server having an offloading system is disclosed, the offloading system including a SoC (System on Chip) and an FPGA (Field Programmable Gate Array). One possible embodiment of the offloading system is on a card. According to a set of embodiments, the SoC can be configured to offload one or more hypervisor functions suitable for execution in software from the CPU complex of the server, and the FPGA can be configured to offload one or more hypervisor functions suitable for execution in hardware from the CPU complex. Brief Description of the Drawings
[0004] Figure 1 A physical server topology according to certain embodiments is depicted, which includes an offloading card having a SoC and an FPGA.
[0005] Figure 2 Depicts according to certain embodiments for Figure 1 the architecture of the offloading card.
[0006] Figure 3 A JTAG (Joint Test Action Group) multiplexer implementation according to certain embodiments is depicted.
[0007] Figure 4 An example network processing flow according to certain embodiments is depicted. Detailed Description
[0008] In the following description, for purposes of explanation, numerous examples and details are set forth to provide an understanding of various embodiments. However, it will be apparent to those skilled in the art that some embodiments may be practiced without these details, or may be practiced with modifications or their equivalents.
[0009] 1. Overview
[0010] Embodiments of the present disclosure are directed to the design of a physical server employing an offload card that includes a SoC (System on Chip) and an FPGA (Field Programmable Gate Array). In various embodiments, the SoC and the FPGA may run hypervisor functions that are typically performed by the CPU complex of the server, thereby offloading the processing burden for these functions from the CPU complex. For example, the SoC of the offload card may run hypervisor functions that require the flexibility of a general-purpose processor or benefit from the flexibility of a general-purpose processor (e.g., networking and storage control plane functions), while the FPGA of the offload card may run hypervisor functions that are suitable for implementation / acceleration in hardware (e.g., networking and storage data plane functions).
[0011] With this general architecture, it is possible to move most, if not all, of the hypervisor processing from the CPU complex of the server to the offload card, which advantageously allows the CPU complex to focus on running tenant (e.g., customer) VM workloads. In the case where the hypervisor completely exits the CPU complex, tenant code may potentially run on the CPU complex in a "bare metal" manner (i.e., without any intermediate hypervisor virtualization layer).
[0012] In addition, since the execution of the hypervisor code / logic on the offload card is physically isolated from the execution of the tenant code on the CPU complex, this solution protects the hypervisor from side-channel attacks that may attempt to use the tenant code as an attack vector.
[0013] Furthermore, by employing an FPGA to accelerate certain hypervisor functions that lend themselves to hardware implementation, the offload card can improve the efficiency of the server while maintaining architectural flexibility. For example, if needed, the FPGA can be reprogrammed from accelerating one type / category of functions (e.g., networking) to accelerating another type / category of functions (e.g., storage). This is not possible for logic-based hard accelerators such as ASICs (Application Specific Integrated Circuits).
[0014] The foregoing and other aspects of the present disclosure will be described in further detail in the following sections.
[0015] 2. Server Topology
[0016] Figure 1 FIG. is a simplified block diagram showing a high-level topology of a physical server 100 according to certain embodiments of the present disclosure. In a set of embodiments, the physical server 100 may be a cloud server deployed as part of the infrastructure of a cloud platform. In these embodiments, the physical server 100 may be mounted in a server rack within a data center operated by a cloud platform provider. In some other embodiments, the physical server 100 may be deployed in other scenarios and / or via other form factors, such as being deployed in a pre-built enterprise IT environment in the form of, for example, a stand-alone server.
[0017] As pointed out in the background section, cloud servers typically implement a virtual machine monitor for virtualization, which allows cloud platforms to provide services such as IaaS (Infrastructure as a Service). However, due to the use of some platform resources (including CPU cores) for virtual machine monitor (also known as "hosting") purposes, traditional cloud servers cannot expose their full CPU capabilities to VMs, thus reducing the efficiency of the platform.
[0018] To address this and other issues, the physical server 100 includes a novel offload card 102, which includes a SoC 104 and an FPGA 106. In the illustrated embodiment, the offload card 102 is implemented as a PCIe (Peripheral Component Interconnect Express)-based expansion card, and thus interfaces with the motherboard of the physical server 100 via a standard PCIe x16 3.0 edge connector interface 108. In some other embodiments, the offload card 102 may be implemented using any other type of peripheral interface.
[0019] As shown, the SoC 104 has its own RAM (Random Access Memory) 110 and flash memory 112, and is communicatively coupled to the FPGA 106 at least via interfaces (PCIe interface 114 and Ethernet interface 116) within the offload card 102. Additionally, the SoC 104 is communicatively coupled to the baseboard management controller (BMC) 118 of the physical server 100 via an I2C interface 108 and several other channels (such as USB and COM).
[0020] The FPGA 106 also has its own RAM 120 and flash memory 122, and is communicatively coupled to the CPU complex 124 of the physical server 100 via a PCIe edge connector interface 108. This CPU complex includes the main CPU cores of the physical server 100 and associated RAM modules. Additionally, the FPGA 106 includes two external Ethernet interfaces, one of which is connected to an external network 126 (via, for example, a TOR (Top of Rack) switch or some other network device), and the other is connected to a NIC (Network Interface Card / Controller) 128 within the physical server 100.
[0021] Generally, Figure 1 the topology shown enables some or all of the hypervisor functions that typically run on the CPU complex 124 of the physical server 100 to instead run on the SoC 104 and FPGA 106 of the offload card 102, and thus be offloaded to the SOC 104 and FPGA 106 of the offload card 102. For example, hypervisor functions that benefit from the flexibility of general-purpose processors (or are simply too complex / dynamic to be implemented in hardware) can run on the SoC 104 incorporating one or more general-purpose processing cores. Examples of such functions include SDN (software-defined networking) control plane functions, which require complex routing calculations and need to be updated relatively frequently to support new protocols and features.
[0022] On the other hand, hypervisor functions suitable for hardware acceleration can be implemented via the logic blocks on the FPGA 106. Examples of such functions include SDN data plane functions (involving forwarding network data traffic based on control plane decisions) and storage data plane functions (such as data replication, deduplication, etc.).
[0023] This solution has many advantages compared to traditional server designs. First, by offloading some of the hosting processing responsibilities of the CPU complex 124, the amount of platform resources (including the CPU cores in the CPU complex 124) used by the hypervisor can be reduced, which in turn increases the platform capabilities available to the VMs (also known as "guests"). This is particularly beneficial in public cloud platforms, where every improvement in server efficiency has a significant impact at scale. In some embodiments, the hypervisor can completely exit the CPU complex 124 and be moved to the offload card 102, in which case the CPU complex 124 can run a minimal hypervisor that processes only issues that can run on the CPU complex itself, such as accessing certain registers, or no hypervisor at all, and the remaining computing power of the CPU complex 124 can be dedicated to guest workloads.
[0024] Second, by implementing both the SoC 104 (which processes non-hardware-accelerated functions) and the FPGA 106 (which processes hardware-accelerated functions) on the offload card 102 and tightly coupling the two, the hypervisor code running on the SoC 104 can more easily interact with the logic implemented in the FPGA 106, and vice versa. There could be alternative embodiments that include hardware accelerators only on the offload card 102, but these embodiments would require data flows to properly coordinate the activities of the hardware accelerators with the server's main CPU. Additionally, these alternative embodiments may not support "bare metal" platforms and may not be able to offload a large amount of work.
[0025] Third, since the host code running on the offload card 102 is physically isolated from the guest code running on the CPU complex 124, it is more difficult for malicious entities to attack the hypervisor via the VM. This is particularly important in view of certain side-channel vulnerabilities recently discovered in modern CPU architectures. Although these known vulnerabilities can be patched, other similar vulnerabilities may be discovered in the future.
[0026] Fourth, by using an FPGA instead of an ASIC for hardware acceleration, the offload card 102 can be easily reused for different usage scenarios or the same usage scenario improved by reprogramming the FPGA, and the FPGA logic can be updated if needed. This is advantageous in large-scale deployments where it may not be desirable to pull and replace a large number of cards already in use in the field.
[0027] It should be understood that Figure 1 the specific topology of the physical server 100 shown is exemplary and various modifications are possible. For example, although the SoC 104 and the FPGA 106 are shown as being implemented on an expansion card (i.e., the offload card 102) that interfaces with the motherboard of the physical server via a peripheral (e.g., PCIe) interface, in some embodiments, alternative offload architectures can be used. In a particular embodiment, one or more of the SoC 104 and / or the FPGA 106 can be implemented directly on the server motherboard.
[0028] As another example, although the NIC 128 is described as a stand-alone component, in some embodiments, the functionality of the NIC 128 can be incorporated into Figure 1 one or more of the other components shown, such as in the FPGA 106. Those of ordinary skill in the art will recognize other variations, modifications, and alternatives.
[0029] 3. Offload Card Architecture
[0030] Figure 2 is schematic diagram 200, which presents additional details regarding the architecture of the offload card 102 in accordance with certain embodiments. Figure 1 The various aspects of this architecture will be discussed in turn below.
[0031] 3.1 SoC
[0032] The SoC 104 can be implemented using any of several existing system-on-chip designs as follows: existing system-on-chips including one or more general-purpose processing cores, interfaces for memory, storage, and peripherals, and a NIC. In a particular embodiment, the SoC 104 can incorporate general-purpose processing cores based on the ARM microprocessor architecture.
[0033] As shown, SoC 104 is communicatively coupled to FPGA 106 via three separate interfaces, which will be discussed in Section 3.2 below. Additionally, the SoC (1) is connected to one or more DRAM (Dynamic Random Access Memory) modules 202 corresponding to Figure 1 RAM 110 via memory interface 204, (2) is connected to an eMMC (Embedded Multimedia Card) device 206 corresponding to Figure 1 flash memory 112 via storage interface 208, (3) is connected to BIOS flash component 210 via an SPI (Serial Peripheral Interface) interface 212 and an intermediate security chip 214, and (4) is connected to several I2C (Inter-Integrated Circuit) devices, such as EEPROM 216, a hot-swap controller 218, and a temperature sensor 220, via an I2C bus 222 (which is also connected to FPGA 106 and PCIe edge connector interface 108).
[0034] Regarding (1), SoC 104 can use the DRAM module(s) 202 as its working memory for running program code, including hypervisor code unloaded from the CPU complex 124 of the physical server 100. The specific number and capacity of the DRAM module(s) 202 and the specifications of the memory interface 204 can vary according to the implementation. In a particular embodiment, the DRAM module(s) 202 can include 8GB (gigabytes) of DDR4 DRAM, which is organized as a single memory bank of 1024M (megabits) x 64 bits + ECC (Error Correction Code), and the memory interface 204 can be configured as a single DDR4-2400 memory channel.
[0035] Regarding (2), SoC 104 can use eMMC device 204 as a non-transitory storage medium for storing and booting program code to be executed on the SoC, including hypervisor code unloaded from the CPU complex 124, and for storing the FPGA configuration image to be applied to FPGA 106.
[0036] Regarding (3), the BIOS flash component 210 can store the system firmware for SoC 104, and the security chip 214 can ensure that the system firmware is not intentionally or accidentally modified or damaged by an attacker, etc.
[0037] Regarding (4), the I2C devices 216, 218, and 220 can provide various management information about the offload card 102 to the BMC 118. This information can include information such as operating temperature data, manufacturing information, and power consumption data.
[0038] In addition to the above, the SoC 104 includes USB (Universal Serial Bus), COM, and JTAG (Joint Test Action Group) interfaces 224, 225, and 226 respectively connected to external headers 228, 230, and 232, which can be used to connect the SoC 104 to the BMC 118 or external devices for debugging or management. There is also a power throttling signal 234 that can be sent by the BMC 118 to the SoC 104 via the PCIe edge connector interface 108.
[0039] 3.2 Interface between SoC and FPGA
[0040] As previously described, the SoC 104 is communicatively coupled to the FPGA 106 through Figure 2 three internal chip-to-chip interfaces (PCIe interface 236, Ethernet interface 238, and JTAG interface 240). In various embodiments, the PCIe interface 236 provides control and data transfer / switching capabilities. For control capabilities, the system-on-chip 104 can use the PCIe interface 236 (or alternatively, the JTAG interface) to manage and update the FPGA 106. For example, the SoC 104 can verify the FPGA configuration image transferred from the RAM 110 to the FPGA 106 and can use this interface to update the image on the FPGA or in the flash memory 122 of the FPGA. For data capabilities, the PCIe interface 236 can enable program code running on the SoC 104 to send data to and receive data from the FPGA 104. This is useful for, for example, virtual machine monitor code that has been written to exchange data via PCIe, as such code can be ported with relatively few changes for execution on the SoC 104 (or implementation on the FPGA 106). In a particular embodiment, the PCIe interface 236 can have 8 PCI 3.0 lanes (i.e., corresponding to a PCI 3.0 8x interface). In some other embodiments, any other number (e.g., 4, 12, 16, etc.) of PCI lanes can be supported.
[0041] The Ethernet interface 238 allows the SoC 104 and the FPGA 106 to exchange data in the form of network packets. For example, this is useful for hypervisor code that has been written to exchange data via network packets, as such code can be ported with relatively few changes for execution on the SoC 104 (or implementation on the FPGA 106). For example, consider a scenario where network flow-based forwarding is implemented in hardware on the FPGA 106 and the network control plane for determining the routing of network flows is implemented in software on the SoC 104. In this case, flow table exceptions and rules can be transmitted between the FPGA 106 and the SoC 104 in the form of network packets. In a particular embodiment, the Ethernet interface 238 may support 25G (gigabit) Ethernet.
[0042] The JTAG interface 240 provides a way for the SoC 104 to communicate with the FPGA 106 for low-level testing (e.g., debugging) and programming purposes. In some embodiments, a JTAG multiplexer may be inserted in the JTAG path between the SoC 104 and the FPGA 106, allowing an external programmer device to connect to the device interface 240 via the external header 232. In these embodiments, the "current" signal from the external programmer device switches the signal path of the JTAG interface 240 from the SoC 104 to the device. This is very useful for initial offload card startup when loading the initial bitstream and for FPGA application development when the SoC-to-FPGA JTAG path is not ready. Figure 3 An example diagram 300 of this architecture with a JTAG multiplexer 302 is depicted in accordance with certain embodiments.
[0043] 3.3 FPGA
[0044] The FPGA 106 can be implemented using any of several existing FPGA chips. In a particular embodiment, the FPGA 106 can be implemented using an existing FPGA chip that supports a specific minimum number of programmable logic elements (e.g., 1000K elements) and a specific minimum transceiver / FPGA architecture speed grade (e.g., grade 2). As Figure 2 shown, the FPGA 106 is communicatively coupled to the I2C bus 222 and the SoC 104 via the above-described interfaces 236 - interface 240. Additionally, the FPGA 106 (1) is connected to the PCIe edge connector interface 108 via an internal PCIe interface 242, (2) is connected to one or more DRAM modules 244 corresponding to Figure 1 the RAM 120 via a memory interface 246, and (3) is connected to corresponding to Figure 1The QSPI (Quad-Serial Peripheral Interface) flash memory module 248 of the flash memory 122, and (4) are respectively connected to two network transceiver modules 250 and 252 via the Ethernet interface 254 and the Ethernet interface 256.
[0045] Regarding (1), the internal PCIe interface 242 enables the FPGA 106 to communicate with the CPU complex 124 and other PCIe devices installed in the physical server 100 (including, for example, the NIC 128). In a specific embodiment, the PCIe interface 242 can be a PCIe 3.0x16 interface.
[0046] Regarding (2), when executing the logic programmed into the device (including the hypervisor logic unloaded from the CPU complex 124), the FPGA 106 can use the (one or more) DRAM modules 244 as its working memory. The specific number and capacity of the (one or more) DRAM modules 244 and the specifications of the memory interface 246 can vary according to the implementation. In a specific embodiment, the (one or more) DRAM modules 202 can include 8GB (gigabytes) of DDR4 DRAM, which is organized into two 4GB memory banks of 512m×64 bits + ECC, and the memory interface 246 can be configured as a dual DDR4-2400 memory channel.
[0047] Regarding (3), the QSPI flash memory module 248 can store one or more FPGA configuration images, which the FPGA 106 can load when powered on to configure itself to perform its specified functions. In a specific embodiment, the QSPI flash memory module 248 can store at least three independent images, which will be described in Article 3.4 below. In addition to the configuration from the flash memory, the FPGA 106 can also support configuration via an external JTAG programmer device, JTAG commands sent by the SoC 104 through the JTAG interface 240, CvP (Configuration via Protocol) through PCIe, and partial reconfiguration through PCIe.
[0048] Regarding (4), the network transceiver module 250 enables the FPGA 106 to receive incoming network traffic from the external network 126 and send outgoing network traffic to the external network 126. Additionally, the network transceiver module 252 enables the FPGA 106 to exchange network traffic with the NIC 128. This is useful in scenarios where the FPGA 106 implements network plane functions, as the FPGA 106 can receive outgoing network packets from the NIC 128 via module 252, process / transform them appropriately, and send them to the external network 126 via module 250. Conversely, the FPGA 106 can receive incoming network packets from the external network 126 via module 250, process / transform them appropriately, and send them to the NIC 128 via module 252 (at which point they can be forwarded to the correct destination VM). An example network data flow that utilizes the FPGA 106 for network data plane acceleration in this manner will be discussed in Section 4 below. In a particular embodiment, the network transceiver module 250 and the network transceiver module 252 can be QSFP28 optical modules, and the Ethernet interfaces 254 and 256 can support 100G Ethernet.
[0049] 3.4 FPGA Flash Configuration Details
[0050] In a set of embodiments, the QSPI flash module 248 can store at least three independent configuration images for the FPGA 106: a golden image, a fail-safe image, and a user application image. The golden image is factory-tested at initial manufacturing and includes the normal expected functionality for the FPGA 106. The fail-safe image is programmed at the factory and is not overwritten after manufacturing. In various embodiments, this fail-safe image contains the minimum set of functions required by the offload card 102 upon power-up, and the network interfaces of the FPGA 106 are forced into a bypass mode, in which all traffic is passed directly between the interfaces without any intermediate processing by the FPGA. Finally, the user application image is an image defined by the user / customer.
[0051] Upon power-up of the offload card 102, by default, the golden image will be loaded from the QSPI flash module 248 and applied to the FPGA 106 to configure its fabric. If any errors occur during the power-up process (or if problems are detected during server runtime), the card can be rebooted to load the fail-safe image instead of the golden image.
[0052] 4. Example Network Processing Workflow
[0053] By considering the foregoing offload card architecture, Figure 4FIG. 400 is a flow chart depicting an exemplary network processing workflow that may be implemented by physical server 100 according to certain embodiments. Flow chart 400 assumes that the FPGA 106 of offload card 106 is configured to maintain a flow table that includes network flows determined by the network control plane running on SoC 104, and to forward data packets according to the flow table.
[0054] Starting at block 402, the NIC 128 of physical server 100 may present a SR-IOV (Single Root I / O Virtualization) interface to a VM running on the server 100. This SR-IOV interface (referred to as a virtual function) enables the VM to communicate directly with the NIC 128 without involving the virtual machine monitor.
[0055] At block 404, the VM may create a data payload for a network packet to be transmitted to a remote destination and may notify the NIC 128 of this. In response, the NIC 128 may read the data payload from the VM's guest memory space (block 406), assemble the data payload into one or more network packets that have headers identifying, among other things, the IP address of the VM and the IP address of the intended destination (block 408), and output the network packets from an egress port of its network transceiver module 252 that is connected to the FPGA 106 (block 410).
[0056] At blocks 412 and 414, the FPGA 106 may receive the network packets and apply its network data plane logic to perform a lookup of the 5-tuple (source IP address, source port, destination IP address, destination port, protocol) of the network packets in the flow table. If a matching entry is found in the table (block 416), the FPGA 106 may identify the next-hop destination for the network packet in the entry (block 418), update the header of the packet (block 420), and send the packet from the network transceiver module 250 to the external network 126 (block 422), thus ending the workflow.
[0057] On the other hand, if no matching entry is found in the table at block 416 (indicating that this is the first packet in the flow), the FPGA 106 may send the network packet to the SoC 104 via the internal Ethernet interface 238 (block 424). The network control plane component running on the SoC 104 may then calculate the next-hop destination for the packet and add a new entry for the network flow of the packet to the flow table of the FPGA via the interface 238 (block 426). Using the new entry, the FPGA 106 may perform blocks 420 and 422, and the workflow may end.
[0058] The foregoing description illustrates various embodiments of the present disclosure and examples of how aspects of these embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments, and are presented to illustrate the flexibility and advantages of the present disclosure as defined by the following claims. For example, although specific process flows and steps of certain embodiments have been described, it will be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described flows and steps. Steps described as sequential may be performed in parallel, the order of the steps may be varied, and steps may be modified, combined, added, or omitted. As another example, although certain embodiments have been described using a specific combination of hardware and software, it should be recognized that other combinations of hardware and software are possible, and specific operations described as being implemented in software may also be implemented in hardware, and vice versa.
[0059] Accordingly, the detailed description and the drawings are to be regarded as illustrative rather than restrictive. Other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be used without departing from the spirit and scope of the present disclosure as described in the following claims.
Claims
1. A server, comprising: A central processing unit (CPU) complex ; And An offload card, the offload card comprising: A system-on-chip (SoC); And A field-programmable gate array (FPGA), the FPGA being external to the SoC and coupled to the SoC, Wherein the CPU complex is configured to execute one or more virtual machines (VMs), wherein the SoC is configured to execute one or more first functions of a virtual machine monitor associated with the one or more VMs in software, and Wherein the FPGA is configured to execute one or more second functions of a virtual machine monitor associated with the one or more VMs in hardware.
2. The server according to claim 1, wherein the SoC and the FPGA are communicatively coupled to each other via a peripheral component interconnect express (PCIe) interface inside the offload card and via an Ethernet interface inside the offload card.
3. The server according to claim 2, wherein the SoC and the FPGA are further communicatively coupled to each other via a joint test action group (JTAG) interface inside the offload card.
4. The server according to claim 1, wherein the offload card is inserted into a motherboard of the server via a PCIe edge connector interface.
5. The server according to claim 4, wherein the SoC is communicatively coupled to a baseboard management controller (BMC) of the server via the PCIe edge connector interface.
6. The server according to claim 4, wherein the FPGA is communicatively coupled to the CPU complex via the PCIe edge connector interface.
7. The server according to claim 1, wherein the SoC is communicatively coupled to one or more volatile memory modules residing on the offload card, the one or more volatile memory modules serving as working memory, and the SoC is capable of executing the one or more first functions from the working memory.
8. The server according to claim 1, wherein the SoC is communicatively coupled to a flash memory module residing on the offload card, and the flash memory module stores program code for the one or more first functions.
9. The server according to claim 1, wherein the FPGA is communicatively coupled to one or more volatile memory modules residing on the offload card, the one or more volatile memory modules serving as working memory for the FPGA when executing the one or more second functions.
10. The server according to claim 1, wherein the FPGA is communicatively coupled to a flash memory module residing on the offload card, and the flash memory module stores at least one configuration image for configuring the FPGA to execute the one or more second functions.
11. The server according to claim 10, wherein the flash memory module stores a first configuration image and a second configuration image, the first configuration image corresponding to a normal operation configuration for the FPGA, and the second configuration image corresponding to a fail-safe operation configuration for the FPGA.
12. The server according to claim 11, wherein the first configuration image is default applied to the FPGA when the offload card is powered on.
13. The server according to claim 12, wherein the second configuration image is applied to the FPGA in the case of an error occurring when applying the first configuration image.
14. The server according to claim 1, wherein the FPGA includes a first external network interface communicatively coupled to a top-of-rack TOR network switch and a second external network interface communicatively coupled to a network interface card NIC of the server.
15. The server according to claim 1, wherein the SoC is communicatively coupled to a basic input / output BIOS flash component residing on the offload card via a security chip, and the security chip is configured to verify the integrity of the firmware stored on the BIOS flash component.
16. The server according to claim 1, wherein the one or more first functions include a network control plane function or a storage control plane function.
17. The server according to claim 1, wherein the one or more second functions include a network data plane function or a storage data plane function.
18. A server, comprising: A central processing unit CPU complex configured to execute one or more virtual machines VM ; And An offload card, the offload card comprising: A first means for executing in software one or more first functions of a virtual machine monitor associated with the one or more VMs; And A second means for executing in hardware one or more second functions of the virtual machine monitor associated with the one or more VMs, wherein the second means is external to and coupled to the first means.
19. A method for processing a flow table, comprising: Receiving, by a field programmable gate array FPGA residing on an offload card of a server, a network packet from a network interface card NIC of the server, wherein the network packet is received via an Ethernet interface interconnecting the FPGA and the NIC; Performing, by the FPGA in hardware, a lookup in a flow table based on a header of the network packet; When determining that no matching entry for the header is found in the flow table, forwarding, by the FPGA, the network packet to a system-on-chip SoC residing on the offload card, wherein the network packet is forwarded via an Ethernet interface interconnecting the FPGA and the SoC; Calculating, by the SoC in software, a next-hop destination for the network packet; And Updating, by the SoC in software, the flow table with a new flow entry including the next-hop destination.
20. The method according to claim 19, the method further comprising, when determining that a matching entry is found in the flow table: Updating, by the FPGA, the network packet based on the matching entry; and Transmitting, by the FPGA, the network packet to an external network via an external network interface of the FPGA.
Citation Information
Patent Citations
Network Interface Controller with Integrated Network Flow Processing
US20160232019A1
Self-morphing server platforms
US20180189081A1