Non-Volatile Storage Express over Fabric (NVMeOF) using a volume management device
Rack Scale Design with NVMeOF and NVMe device emulation optimizes resource allocation in data centers, addressing inefficiencies by managing and virtualizing storage resources, thereby reducing costs and improving resource utilization.
Patent Information
- Application Number
- DE102018004046
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-05-18
- Filing Date
- 2018-05-18
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2038-05-18
AI Technical Summary
Current enterprise/cloud computing systems face inefficiencies in resource allocation and management, leading to higher costs and lower return on investment due to the inefficient use of rack resources in data centers, particularly in managing volatile and non-volatile memory within compute nodes.
Implementing Rack Scale Design (RSD) architecture with Non-Volatile Memory Express over Fabric (NVMeOF) protocol, utilizing NVMe storage devices via a low latency fabric, and employing NVMe device emulation to manage and virtualize storage resources across a data center, allowing dynamic assembly of resources based on workload-specific demands.
Enhances resource management efficiency, reduces overall costs, and improves return on investment by optimizing the use of compute, storage, and network resources in data centers, enabling seamless access to remote storage devices as local NVMe drives without software modifications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
STATE OF THE ART
[0001] The availability and use of "cloud" computing has grown exponentially in recent years. Under a traditional computing approach, users run software applications on their own computers and / or access software services located on local servers (such as servers operated by a business organization). In contrast, with cloud computing, the computing and storage resources are "in the cloud," meaning they are physically hosted in a remote facility accessed over a computer network, such as the Internet. Computing and storage resources hosted by a cloud operator can be accessed through "services," commonly referred to as cloud-based services, web services, or simply services.
[0002] Cloud-based services are typically hosted by a data center, which features a physical arrangement of servers that comprise a cloud or a specific portion of a cloud. Data centers generally employ a physical hierarchy of compute, network, and shared storage resources to support scaling workload requirements. Fig. 1 shows a portion of an exemplary physical hierarchy in a data center 100 having an L number of pods 102 and an M number of racks 104, each having slots for an N number of bays 106. Each bay 106 may, in turn, have multiple sleds 108. For clarity, each pod 102, rack 104, and bay 106 is labeled with a corresponding identifier, such as Pod 1, Rack 2, Bay 1B, etc. Bays may also be referred to as trays, and sleds may also have various forms, such as modules or nodes. In addition to the tray and sled configuration, racks may be provided using enclosures in which various forms of servers are installed, such as blade server enclosures and server blades. While the term "slot" is used herein, those skilled in the art will recognize that trays and enclosures are generally analogous terms.
[0003] At the top of each rack 104, a respective top-of-rack (ToR) switch 110 is shown, also labeled with a ToR switch number. Generally, ToR switches 110 refer to both ToR switches and any other switching device that supports switching between racks 104. It is common practice to refer to these switches as ToR switches, regardless of whether they are physically located at the top of a rack or not (although they generally are).
[0004] Each pod 102 further includes a pod switch 112 to which the pod's ToR switches 110 are coupled. Pod switches 112 are, in turn, coupled to a data center (DC) switch 114. The data center switches may be hierarchically located at the top of the data center switch, or there may be one or more additional levels not shown. For clarity, the hierarchies described herein are physical hierarchies that utilize physical LANs. In practice, it is common to deploy virtual LANs that utilize underlying physical LAN switching devices.
[0005] Cloud-hosted services are generally categorized into Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), and Infrastructure-as-a-Service (IaaS). SaaS services, also commonly called web services and cloud application services, provide access to services running on data center servers through a network connection and client-side interface, such as a web browser. Common examples of SaaS services include email web services (e.g., Google Gmail, Microsoft Hotmail, Yahoo mail), Microsoft Office 365, Salesforce.com, and Google Docs. PaaS, also known as cloud platform services, are used for applications and other deployments, providing cloud components for software. Examples of PaaS include Amazon Web Services (AWS), Elastic Beanstalk, Windows Azure, and Google App Engine.
[0006] IaaS are services for accessing, monitoring, and managing remote data center infrastructure, such as (virtualized or bare-metal) compute, storage, networks, and network services (e.g., firewalls). Instead of purchasing and operating their own physical hardware, users can purchase IaaS on a consumption-based basis. For example, AWS and Windows Azure each offer the use of Amazon and Microsoft data center resources on a resource allocation / consumption basis. Amazon Elastic Compute Cloud (EC2) is a core part of AWS.
[0007] IaaS usage for a given customer typically involves an allocation of data center resources. For example, a typical AWS user may request the use of one of 24 different EC2 instances, ranging from a t2.nano instance with 0.5 gigabytes (GB) of memory, 1 core / variable cores / compute units, and no instance storage to an hs 1.8xlarge with 117 GB of memory, 16 / 35 cores / compute units, and 48,000 GB of instance storage. Each allocated EC2 instance consumes specific physical data center resources (e.g., compute, storage). At the same time, data center racks can support a variety of different configurations. To maximize resource allocation, the IaaS operator must track which resources are available in which rack.
[0008] Document US 9,294,567 B2 discloses systems and methods for enabling access to expandable storage devices over a network as local storage via NVMe controllers. Document US 2016 / 0259568 A1 discloses methods and apparatus for storing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The foregoing aspects and many of the attendant advantages of this invention will become more readily apparent as the latter becomes better understood by reference to the following detailed description, taken in conjunction with the accompanying drawings, in which like reference characters designate like parts throughout the several views unless otherwise specified: Fig. 1 is a schematic representation of a conventional physical rack configuration in a data center, Fig. 2 is a schematic diagram of a Rack Scale Design (RSD) configuration in a data center according to one embodiment, Fig. 3 is a block diagram of an RSD management architecture according to one embodiment, Fig. Figure 4 is a schematic diagram showing further details of an RSD rack implementing Pooled System Management Engines (PSMEs), Fig. Figure 5 is a schematic diagram providing an overview of a remote storage access mechanism where servers are enabled to access NVMe storage devices over a low-latency fabric, Fig. Figure 5a is a schematic diagram showing an alternative implementation of the scheme of Fig. 5, where NVMe emulation is implemented to enable low-latency access to non-NVMe storage devices over the fabric, Fig. 6 is a schematic diagram of an implementation environment including multiple CPU sleds and memory slots communicatively coupled via a fabric, according to one embodiment, Fig. 6a is a schematic diagram showing an implementation environment employing an Ethernet fabric, according to one embodiment, Fig. 7 is a schematic diagram illustrating an access scheme of virtualized storage in which remote storage devices on a target are exposed to an operating system on a launcher as local storage devices, with an NVMe function implemented in a network interface controller (NIC), Fig. 7a is a schematic diagram illustrating an access scheme of virtualized storage in which remote storage devices on a target are exposed to an operating system on a launcher as local storage devices, with an NVMe function implemented in a fabric controller, Fig. Figure 7b is a schematic diagram showing a modification of the embodiment of Fig. 7 illustrates accessing an assessment of non-NVMe storage devices on the target using NVMe emulation, Fig. Figure 7c is a schematic diagram illustrating an access scheme of a virtualized storage in which remote storage devices on a target are exposed to an operating system on a launcher as local storage devices, with an NVMe function implemented in a field-programmable gate array (FPGA), Fig. Figure 8 illustrates an architecture and an associated process flow diagram in which virtualization of local and remote NVMe storage devices is implemented by extensions with a processor in combination with firmware and software drivers, Fig. Figure 8a is a schematic diagram showing the architecture of Fig. 8, which provides a launcher that accesses remote storage devices on a target, and further illustrates a virtualized I / O hierarchy exposed to an operating system running on the launcher, Fig. 9 is a flowchart illustrating operations performed during initialization of a compute node, according to one embodiment, Fig. 10 is a schematic diagram illustrating a firmware-based extension used to support a storage virtualization layer that abstracts physical storage devices and exposes them as virtual devices to software running on the compute node, the software comprising a hypervisor and multiple virtual machines, Fig. Figure 10a is a schematic diagram showing an extension of the embodiment of Fig. 10, which uses a software virtualization extension as part of the hypervisor, Fig. Figure 11 is a schematic diagram showing an extension of the embodiment of Fig. 10, which uses an operating system virtualization layer and multiple containers, Fig. Figure 11a is a schematic diagram showing an extension of the embodiment of Fig. 11a, which uses a software virtualization extension as part of the operating system virtualization layer, and Fig. 12 is a flowchart illustrating operations and logic for implementing a virtualized memory access scheme that may be used with one or more of the embodiments set forth herein. DETAILED DESCRIPTION
[0010] Embodiments of non-volatile storage express over fabric (NVMeOF) using volume management device (VMD) schemes and associated methods, systems, and software are described herein. In the following description, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention. However, those skilled in the art will recognize that the invention may be practiced without one or more of the specific details, or with different methods, components, materials, etc. In other instances, well-known structures, materials, or acts are not shown or described in detail to avoid obscuring aspects of the invention.
[0011] References in this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, the phrase "in one embodiment" in various places throughout this specification does not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0012] For clarity, individual components in the figures herein may also be referred to by their identifiers in the figures instead of a specific reference numeral. In addition, reference numerals that designate a particular type of component (as opposed to a specific component) may be shown with a reference numeral followed by "(Type)" meaning "typical." It is understood that the configuration of these components is typical of similar components that may exist but are not shown in the figures for simplicity and clarity, or other similar components that are not designated by separate reference numerals. Conversely, "(Type)" should not be construed to mean that the component, element, etc., is typically used to designate a disclosed function, implementation, purpose, etc.
[0013] Recently, INTEL® Corporation introduced a new rack architecture called Rack Scale Design (RSD) (formerly Rack Scale Architecture). Rack Scale Design is a logical architecture that partitions compute, storage, and network resources and introduces the ability to pool these resources for more efficient resource utilization. It simplifies resource management and provides the ability to dynamically aggregate resources based on workload-specific demand.
[0014] RSD uses compute, fabric, storage, and management modules that work together to enable a selectable configuration of a wide range of virtual systems. The design uses four basic pillars that can be configured based on user needs. These include 1) a Pod Manager (PODM) for managing multiple racks, comprising firmware and software application programming interfaces (APIs) that enable resource and policy management and expose the underlying hardware and the orchestration layer above it through a standard interface; 2) a system pool of compute, network, and storage resources that can be selectively composed based on workload requirements; 3) pod-wide storage built on connected storage that leverages storage algorithms to support a range of uses;deployed as resources with multiple racks or storage hardware and compute nodes with local storage, and 4) a configurable network fabric for hardware, interconnects with cables and backplanes, and for management software to support a range of cost-effective network topologies, including current top-of-rack switch designs and distributed switches within the platforms.
[0015] An example RSD environment 200 is shown in Fig. 2. RSD environment 200 includes a plurality of compute racks 202, each including a top-of-rack (ToR) switch 204, a pod manager 206, and a plurality of system pool bays. Generally, system pool bays may include compute pool bays and storage pool bays. Optionally, system pool bays may also include storage pool bays and input / output (I / O) pool bays. In the illustrated embodiment, the system pool bays include an INTEL® XEON® compute pool bay 208 and an INTEL® ATOM™ compute pool bay 210, a storage pool bay 212, a storage pool bay 214, and an I / O pool bay 216. Each of the system pool bays is connected to ToR switch 204 via a high-speed link 218, such as a 40 Gigabit / second (Gb / s) or 100 Gb / s Ethernet link or a 100+ Gb / s silicon photonics (SiPh) optical link.In one embodiment, high-speed link 218 comprises an 800 Gb / s SiPh optical link.
[0016] Multiple compute racks 200 may be interconnected via their ToR switches 204 (e.g., a pod-level switch or a data center switch), as represented by connections to a network 220. In some embodiments, groups of compute racks 202 are managed as separate pods via pod managers 206. In one embodiment, a single pod manager is used to manage all racks in the pod. Alternatively, distributed pod managers may be used for pod management operations.
[0017] RSD environment 200 further includes a management interface 222 used to manage various aspects of the RSD environment. This includes managing the rack configuration with corresponding parameters stored as rack configuration data 224.
[0018] Fig. 3 shows one embodiment of an RSD management architecture 300. The RSD management architecture includes multiple software and firmware components configured in a layered architecture including an orchestration layer 302, an RSD pod management foundation API (application programming interface), a pod manager 306, and an RSD management firmware API 308. The lower layer of an RSD management architecture includes a compute platform management component 310, a storage management component 312, a rack management component 314, and a network switch management component 316.
[0019] Compute platform management component 310 performs operations associated with compute bays and includes a system pool, a management system, node management, switch configuration, and a boot service. Storage management component 312 is configured to support operation management of storage pool bays. Rack management component 314 is configured to manage rack temperature and power subsystems. Network switch management component includes a distributed switch manager.
[0020] INTEL® Rack Scale Design is designed to shift the focus from the platform architecture of individual servers to a converged infrastructure consisting of compute, network, and storage, as outlined above and in Fig. 2. Resource management is performed at the rack and pod levels. The focus on rack-level resource management also requires management of rack-level environments, such as power and cooling zones, as well as providing a rack-level chain of trust for relative location information. This role is fulfilled by a Rack Management Module (RMM) together with a sub-rack unit manager (the plug-in units in RSD terminology), called the Pooled System Management Engine (PSME). The management elements of the RSD, RMM, and the PSMEs are connected to a private network that is not accessible from outside the rack, as shown in Fig. 4 and is explained below.
[0021] Fig. 4 shows an embodiment of a rack configuration 400 employing rack management and rack configuration components that communicate over a private rack management network. The rack management and rack configuration components include an RMM 402 communicatively coupled to a rack management switch 404 via a connection 406. A respective PSME 408 is associated with each of five system pool bays 410. Each PSME 408 is connected to a rack management switch 404 via a connection 412. The rack management switch is also connected to a POD manager 206. In the illustrated embodiment, each system pool bay 1 and 2 includes multiple compute nodes 500, while system pool bays 3, 4, and 5 each include multiple memory resources 414, multiple storage resources 415, and multiple IO acceleration resources 416.
[0022] In a data center environment, such as a RSD, data center management software is capable of assembling a single instance or compute node from various rack resources to meet user performance requirements. Generally, allocating resources to meet performance results in inefficient use of rack resources, resulting in higher total cost of ownership (TCO) and lower return on investment (ROI).
[0023] Current enterprise / cloud computing systems have volatile data storage, such as DRAM (Dynamic Random Access Memory), and non-volatile memory-class data storage, such as 3D Crosspoint technology (3D XPOINT™) DIMMs (Dual In-line Memory Modules), located locally within the compute node. Other data storage types can also be used.
[0024] Non-volatile data storage devices are storage media that do not require energy to maintain the data state stored by the medium. Non-limiting examples of non-volatile data storage devices may include any of the following or a combination thereof: solid-state storage devices (such as planar or 3D NAND flash memory or NOR flash memory), 3D crosspoint memory, storage devices that use chalcogenide phase-change material (e.g., chacogenide glass), byte-addressable non-volatile memory devices, ferroelectric data storage devices, silicon oxide nitride oxide silicon (SONOS) data storage devices, polymer data storage devices (e.g.,Ferroelectric polymer memory, ferroelectric transistor random access memory (Fe-TRAM), ovonic memory, nanowire memory, electrically erasable programmable read-only memory (EEPROM), other various types of non-volatile random access memories (RAMs), and magnetic memory. In some embodiments, 3D crosspoint memory may be a stackable, fewer-transistor crosspoint architecture in which memory cells are located at the intersection of word lines and bit lines and are individually addressable, and in which bit storage is based on a change in bulk resistance.In certain embodiments, a memory module with non-volatile memory may conform to one or more standards promulgated by the Joint Electron Device Engineering Council (JEDEC), such as JESD218, JESD219, JESD220-1, JESD223B, JESD223-1, or other suitable standards (the JEDEC standards cited herein are available at www.jedec.org).
[0025] Volatile memories are storage media that require power to maintain the data state stored in the medium. Examples of volatile memories can include various types of random-access memory (RAM), such as dynamic random-access memory (DRAM) or static random-access memory (SRAM). One specific type of DRAM that can be used in a memory module is synchronous dynamic random-access memory (SDRAM). In certain embodiments, the DRAM of the memory modules conforms to a standard promulgated by JEDEC, such as JESD79F for Double Data Rate (DDR) SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, or JESD79-4A for DDR4 SDRAM (these standards are available at www.jedec.org).These standards (and similar standards) may be referred to as DDR-based standards, and communication interfaces of storage devices that implement these standards may be referred to as DDR-based interfaces.
[0026] The rack-scale system uses these compute nodes and storage nodes (non-volatile memory, SATA, and NVM-Express (NVMe) storage devices, etc.) to assemble a system based on user needs. In embodiments described below, NVMe storage devices in storage nodes (also known as storage pool bays) are accessed by compute platforms over a low-latency fabric using an NVMe-Over-Fabric (NVMe-OF) protocol. In conjunction with accessing NVMe storage devices over the fabric, NVMe devices are designed to appear to an operating system as locally attached devices (i.e., attached to the compute platforms) that can be accessed by existing software, such as applications, without requiring any software modification.As described in detail below, the foregoing functionality is implemented via a hardware-based scheme, a primarily software-based scheme, and a combined hardware and firmware scheme.
[0027] An overview that presents NVMe storage aspects of the concept is provided in Fig. 5. Under the mechanism, each of a plurality of compute nodes, such as servers 500, is capable of accessing NVMe storage devices 502-1 ... 502-6 in a storage pool shelf 704 via a fabric 506 and a fabric switch 508. Storage pool shelf 704 further includes a storage distributor 510 with a fabric connector 512 and an NVMe driver interface 514. A PSME 516 is also coupled to storage distributor 510; in general, the PSME may either be integrated into a storage pool shelf (as shown) or reside external to the storage pool shelf.
[0028] In one embodiment, the low-latency fabric comprises an INTEL® Omni-Path fabric employing the INTEL® Omni-Path Architecture (OPA). OPA employs a host fabric interface (HFI) at each fabric endpoint and a fabric switch used to forward fabric packets along links between fabric endpoints. In another embodiment, the low-latency fabric comprises an Ethernet network employing high-speed Ethernet links (e.g., 100 Gigabits per second) and Ethernet switches. For clarity, fabric switches are shown, which correspond to Omni-Path fabric switches or Ethernet switches, depending on the implementation context. In addition to Omni-Path and Ethernet fabrics, other existing and future network technologies may be used in a manner similar to that described and illustrated herein.
[0029] In addition to using NVMe storage devices in storage pool devices, embodiments may use non-NVMe storage devices in combination with NVMe device emulation. One embodiment of this approach is described in Fig. 5a, which is similar to the embodiment shown in Fig. 5, except for the following differences. Storage pool bay 504 has been replaced by a storage pool bay 504a, which includes storage devices 503-1 ... 503-6. In general, various types of storage devices may be used, including, but not limited to, solid-state drives (SSDs), magnetic disk drives, and optical storage devices. Storage pool bay 504a further includes a storage manifold 510a, which includes a storage device interface 514a configured to interface with storage devices 503-1 ... 503-6.
[0030] According to one aspect of a non-NVMe storage device scheme, software and / or firmware providing NVMe emulation is implemented either at the storage distributor 510a, in servers 500, or a combination of the two (i.e., part of the NVMe emulation is performed by software and / or firmware in a server 500, and another part is performed by software and / or firmware in storage distributor 510a). In an emulated NVMe storage device implementation, NVMe emulation enables existing NVMe software, such as NVMe device drivers running on a server, to access non-NVMe storage devices via fabric 506 without requiring any changes to the NVMe device drivers.
[0031] An exemplary NVMeOF storage architecture 600 corresponding to one embodiment of an implementation in an RSD environment is shown in Fig. 6. NVMeOF storage architecture 600 includes multiple CPU sleds 602-1 - 602-M, also labeled as sled 1, sled 2, ... sled M. Each CPU sled 602 includes one or more compute nodes 604 having one or more CPUs and memory and an HFI 606. Each CPU sled 602 is connected to a PSME 610 via a high-speed interconnect 612, such as a high-speed Ethernet connection or a SiPh optical interconnect.
[0032] Each HFI 606 is connected to an OPA fabric, which includes multiple fabric interconnects 614 and a fabric switch 616. The OPA fabric enables high-speed, low-latency communication between compute nodes 604 and a pair of storage pool bays 618-1 and 618-2. Each storage pool bay 618-1 and 618-2 includes an HFI 620 and a storage distributor 622 coupled to multiple NVMe storage devices 624 via a PCIe interconnect hierarchy. A storage pool bay may also include a CPU 626. In one embodiment, a storage distributor is implemented as a system-on-chip including a CPU (not shown). Each storage pool shelf 618-1 and 618-2 is coupled to a respective PSME 628 and 630, while PSMEs 610, 628 and 630 are connected to a POD Manager 632.
[0033] Fig. Figure 6a shows an embodiment with NVMeOF storage architecture 600a that uses a high-speed Ethernet fabric instead of the OPA fabric used in Fig. 6. In general, aspects of NVMeOF storage architectures 600 and 600a are similar, except for the differences corresponding to components with an 'a' appended to their reference numerals. These components include CPU sleds 602a-1 - 602a-M, which now have network interface controllers (NICs) 606a instead of HFIs 606, an Ethernet switch 616a instead of fabric switch 616, and storage pool slots 618a-1 and 618a-2, which now have NICs 620a instead of HFIs 620.
[0034] In addition to the fabrics shown, other fabric types can also be used, including InfiniBand and Gen-Z fabrics. Generally, the primary differences would be replacing the illustrated fabric controllers with InfiniBand Host Controller Adaptors (HCAs) for InfiniBand fabrics and Gen-Z interfaces with the fabric controllers for Gen-Z fabrics. Example network / fabric controller implementation of NVMeOF using VMD
[0035] Fig. Figure 7 illustrates one embodiment of a hardware-based implementation of NVMeOF using VMD. In the implementation, a compute node 700 is connected in communication with a storage node 702 via a fabric having fabric interconnects 704 and 706 and a fabric switch 708. As shown, compute node 700 is a "starter," while storage node 702 is a "target."
[0036] Compute node 700 includes a processor 710 having a core portion 712 including multiple processor cores 714 and a PCIe root controller 716. For simplicity, additional details of the processor 710, caches, interconnects, memory controller(s), and PCIe interfaces are not shown. PCIe root controller 716 is coupled to a PCIe-based NVMe drive interface 718 and a PCIe-based network interface controller 720.
[0037] Storage node 702 is generally representative of an aggregated storage device in the system architecture, such as a storage pool shelf under RSD, and includes a network interface 722, a storage distributor 724, and an NVMe drive interface 726 coupled to N NVMe storage devices, including NVMe storage devices 728, 730, and 732. Storage node 702 is also illustrated as including an optional processor 734 having one or more cores 736.
[0038] In general, a given compute node can utilize local and / or remote storage resources. As used herein, a "local" storage resource is a storage device directly coupled to a compute node, such as via a storage device interface. In contrast, a "remote" storage device corresponds to a storage device accessed via a fabric, such as storage devices in storage pool bays. In the illustrated embodiment, compute node 700 is coupled to two local NVMe storage devices 738 and 740, which are coupled to NVMe drive interface 718. These NVMe storage devices are also labeled NVMe L1 (local 1) and NVMe L2 (local 2).
[0039] Compute node 700 is also configured to access two virtual NVMe storage devices 742 and 744, labeled NVMe V1 (virtual 1) and NVMe V2 (virtual 2). For the purposes of running software on compute node 700, virtual NVMe storage devices 742 and 744 are local NVMe drives directly connected to the compute node. In reality, virtual NVMe storage devices 742 and 744 are physically implemented NVMe storage devices 728 and 730, respectively.
[0040] In the embodiment of Fig. 7, the virtualization of NVMe storage devices is implemented primarily in hardware as part of an embedded function in NIC 720. As illustrated, NIC 720 includes two embedded functions: a network function 746 and an NVMe function 748. Network function 746 includes embedded logic and functionality implemented by a conventional NIC to support network or fabric operations. For example, in an implementation where NIC is an Ethernet NIC, network function 746 includes hardware-based facilities for handling the physical (PHY) and media access channel (MAC) layers for the applicable Ethernet protocol.
[0041] The NVMe function corresponds to added functionality (to the NIC) to support virtualization of NVMe storage devices. For illustrative purposes, virtual NVMe storage devices 742 and 744 are shown connected to NVMe function 748; in reality, there are no NVMe storage devices connected to NVMe function 748. At the same time, the NVMe function exposes the presence of such NVMe storage devices to an operating system as part of the enumerated PCIe device hierarchy.
[0042] Specifically, during system boot, the PCIe device and interconnect hierarchy is enumerated as part of conventional PCIe operations. Each PCIe device in the system is enumerated, and the information corresponding to the enumerated PCIe devices is exposed to the operating system running on the compute node. In the embodiment of Fig. 7, the software includes an operating system (OS) 750 and a NIC + NVMe driver 752. The NIC + NVMe driver is implemented as a software interface between the NIC 720 and the OS 750. During the PCIe enumeration process, the NVMe + NVMe driver 752 (vs. the OS 750) exposes the NIC 720 as two separate devices: 1) a NIC, and 2) an NVMe interface connected to two local NVMe storage devices.
[0043] A section 751 of the enumerated PCIe device hierarchy is shown at the bottom left of Fig. 7 and includes PCIe root controller 716, NVMe drive interface 718 connected to local NVMe drives L1 and L2, NIC 754, and NVMe drive interface 756 connected to (virtualized) local NVMe drives L3 and L4. Therefore, with respect to the operating system 750, compute node 700 is directly coupled to four local NVMe drives L1, L2, L3, and L4.
[0044] Fig. Figure 7a shows an embodiment of a hardware-based implementation of NVMeOF using VMD that utilizes an Omni-Path fabric. As indicated by components sharing the same reference numerals, the embodiments of Fig. 7 and Fig. 7a, with the primary difference that in Fig. 7 an Ethernet fabric is used and that in Fig. 7a an Omni-Path fabric is used.
[0045] Consequently, in the embodiment of Fig. 7a, a launcher comprising a compute node 700a is communicatively coupled to a target comprising a storage node 702a via an Omni-Path fabric having interconnects 705 and 707 and a fabric switch 709. NIC 720 has been replaced by a fabric controller 721, and network interface 722 has been replaced by a fabric controller 723. As further illustrated, fabric controller 721 has a network function 746a and an NVMe function 748a and is coupled to HFI 758. Fabric controller 723 is coupled to HFI 760. As shown in the lower left portion of Fig. 7a, the software components include operating system 750, an HFI + NVMe driver 753, and an enumerated PCIe hierarchy section 751a, which includes PCIe root controller 716, NVMe drive interface 718 connected to local NVMe drives L1 and L2, an HFI / Fabric controller 762, and an NVMe drive interface 756 connected to (virtualized) local NVMe drives L3 and L4. As in the embodiment of Fig. 7 is directly coupled to four local NVMe drives L1, L2, L3 and L4 with respect to the operating system 750 of the compute node 700.
[0046] Fig. 7b shows an extension to the embodiment of Fig. 7, which supports the use of non-NVMe storage devices through NVMe emulation. As illustrated, a storage node 702b has a storage device interface 726b coupled to (or integrated with) a storage distributor 724 and coupled to N storage devices, illustrated as storage devices 729, 731, and 733. NVMe emulation is implemented by NVMe emulation block 764. As described above with reference to Fig. As set forth in Section 5a, NVMe emulation may be implemented using software and / or firmware in a storage pool shelf or storage node, software and / or firmware in a compute node or server, or using a combination of both.
[0047] In addition to implementing the NVMe function in a NIC or fabric controller, in one embodiment, the NVMe function is implemented in a field-programmable gate array. In general, the FPGA can be implemented either on the processor system on a chip (SoC) (e.g., an embedded IP block) or as a separate component on the platform. In one embodiment, the FPGA appears (to the software and / or components on the platform) as one or more physically attached storage devices, such as one or more SSDs. For example, the FPGA can be "attached" to an I / O interface, such as, but not limited to, a PCIe interface. In one embodiment of a PCIe implementation, the FPGA provides enumeration information during platform boot indicating that it is a storage device controller connected to one or more storage devices.In reality, however, the FPGA is not connected to any storage device, but has storage virtualization logic to implement NVMe functionality to access remote storage devices in remote storage nodes accessed through the fabric.
[0048] Fig. Figure 7c shows a compute node 700c having an FPGA 749 implementing an NVMe function 748c. As shown, the FPGA 749 exposes virtual NVMe storage devices 742 and 744 to the platform, while the actual physical storage devices are located in storage nodes 702. In the illustrated embodiment, the FPGA is connected to PCIe root controller 716 (note that the FPGA may be connected to a root connector (not shown), which in turn is connected to the PCIe root controller). In one embodiment, the connection between the PCIe root controller (or a PCIe root connector) and FPGA 749 is a PCIe connection 751. Other connection types between processor 710 and FPGA 749 may be implemented, including a Universal Path Interconnect (UPI) connection, an Intel® Accelerator connection, and a Gen-Z connection.Optionally, there may be a connection between FPGA 749 and NIC 720c, as represented by optional connection 753.
[0049] In general, FPGAs can be programmed at the factory, or they can be programmed during system boot or runtime using known techniques for programming FPGAs and similar programmable logic devices. For example, in one approach, a bitstream is downloaded to the FPGA, which contains instructions for programming the gates in the FPGA. Example processor-based implementation of NVMeOF using VMD
[0050] According to aspects of some embodiments, volume management device (VMD) components are used to virtualize memory access. VMD is a new technology used to improve PCIe management. VMD maps the entire PCIe tree to its own address space, which is controlled by a VMD driver and enabled in the BIOS (at a x4 root-connector PCIe granularity in one embodiment). The operating system enumerates the VMD devices, and the OS enumeration for the attached child devices ends there. Control of the device domain passes to a VMD device driver, which is independent of child devices. A VMD driver sets up the domain (enumerating child devices) and clears the FastPath path for child devices. A VMD driver can load additional child device drivers that are VMD-aware of the respective child devices.
[0051] Interrupts from subordinate devices are detected by the operating system as VMD interrupts and are first forwarded to the VMD driver's interrupt service routine (ISR). A VMD driver can reissue the signal to the ISR of the corresponding subordinate device driver.
[0052] An architecture and associated process diagram 800 for one embodiment of a VMD implementation is shown in Fig. 8. The architecture is divided into a software section, shown in the upper part of the diagram, and a hardware section, shown in the lower part of the diagram. The software section includes a standard NIC driver 802 and an NVMe VMD driver 804. The NVMe VMD driver includes a NIC driver 806, an NVMe driver 808, and a VMD driver 810. NVMe VMD driver 804 may also include optional performance enhancements 811.
[0053] The hardware portion of diagram 800 represents a "non-core" 812, which generally includes the portion of a processor SoC that is not part of the processor core, including the processor cores and processor caches (e.g., Level 1 and Level 2 caches). Non-core 812 includes a VMD block 814, a bridge 816, and three 4-line (x4) PCIe root ports (RPs) 818, 820, and 822. Non-core 812 further includes a PCIe root complex, which is not shown to avoid confusion.
[0054] The hardware section further includes a NIC 824 and an NVMe solid-state drive (SSD) 826. NIC 824 includes two functions 828 and 830, labeled "Func0" and "Func1," respectively. NIC 824 further includes an Ethernet port connected to an Ethernet fabric 834 via an Ethernet connection 836. The hardware section also includes platform firmware (FW) that includes an NVMe / VMD Unified Extensible Firmware Interface (UEFI) driver 838.
[0055] One embodiment of a process for configuring a compute node (e.g., server) operating as a starter and employing the architecture of diagram 800 proceeds as follows, with processes represented by circled numbers 1-4. In a first process, an 8-line PCIe root port (x8) is automatically branched into two 4-line root ports. As shown, the x4 root port 818 is connected to Func0, while the x4 root port 820 is connected to Func1 of NIC 824. This enables Func1 to be under VMD 814 and Func0 to utilize standard NIC driver 820, thereby enabling NIC 820 to support standard NIC operations.
[0056] In a second operation, NVMe-VMD-UEFI driver 838 responds to a boot discovery. NVMe-VMD-UEFI driver 838 is an extension of the platform's firmware and is loaded as part of the UEFI firmware deployed to the platform during boot processes. During this operation, NVMe-VMD-UEFI driver 838 stores the boot MAC address in flash memory (not shown). NVMe-VMD-UEFI driver 838 also exposes VMD 814 to the operating system.
[0057] After the UEFI firmware loads, operating system boot processes are initiated, which include loading software-based drivers. As shown in process '3', the OS sees the VID / DID (virtual identifier, device identifier) of VMD 814 and loads NVMe VMD driver 804. The NVMe VMD driver abstracts the inner workings of NVMe virtualization from the software users (i.e., software components used to access an NVMe storage device). This abstraction is performed in a manner that requires no changes to existing user code. Upon loading, NVMe VMD driver 804 enumerates the VMD domain (for VMD 814). It locates NIC 824, bridge 816, and NVMe SSD 826 and loads the respective software drivers (not shown). NVMe VMD driver 804 can also load optional performance enhancements 811 (if supported).
[0058] Fig. Figure 8a shows a system having the components of architecture 800 implemented in a launcher 840 communicatively coupled to a target comprising a storage node 702 via a fabric having a fabric switch 709. As with previous embodiments, launcher 840 can access both local NVMe SSD 826 and remote NVMe drives while exposing the remote NVMe drives as local storage devices. For example, as shown by I / O hierarchy 841, the operating system "sees" a PCI root controller 842 to which NIC 824 and VMD 814 are coupled. As further shown, each SSD 826 and the two remote NVMe drives 1 and 2 appear to the operating system as local NVMe drives 844 and 846.
[0059] Fig. 9 shows a flowchart 900 illustrating operations performed during compute node initialization according to one embodiment. In a block 902, the compute node hardware is initialized. During this process, the compute node's firmware is loaded and a communication channel is configured that does not require an operating system to be running on the compute node. This allows the memory resource configuration information to be communicated to the compute node.
[0060] In a block 904, the compute node resources are assembled, including an allocation of remote storage devices. For example, in one process flow, an IaaS customer requests the use of a compute node with specific resources, and the compute node is assembled by the POD Manager. In a block 906, the POD Manager communicates resource configuration information to the PSME attached to the compute node (or otherwise communicatively coupled to the compute pool shelf for the compute node). The PSME then communicates the resource configuration information to the firmware running on the compute node hardware via the communication channel established in block 904.
[0061] In a block 904, the firmware configures a launcher-target device mapping. This contains information for mapping remote storage devices as local storage devices for the compute node. In a block 910, the OS is booted. Subsequently, the remote storage devices mapped to the compute node are exposed to the OS as local NVMe devices.
[0062] It is noted that the process flow in flowchart 900 is exemplary and not limiting. For example, in another embodiment, the launcher-target configuration mapping is maintained by a software-based driver rather than a firmware device. For example, in architecture 800, the NVMe VMD driver 804 maintains the launcher-target mapping information. Example firmware storage virtualization layer implementation of NVMeOF using VMD
[0063] According to aspects of some embodiments, a firmware-based extension is used to support a storage virtualization layer that abstracts physical storage devices and exposes them as virtual devices to running software on the compute node. One embodiment of this scheme is described in Fig. 10, which has a compute node 1000 including a starter coupled to storage node 1002 via a fabric having fabric interconnects 704 and 706 and a fabric switch 708. As indicated by like-numbered components in Fig. 7 and Fig. As shown in Figure 10, compute node 100 includes a processor 710 having a core 712 with multiple cores 714 and a PCIe root controller 716 to which a PCIe-based NVMe drive interface 718 is coupled. As further shown, a pair of local NVMe drives 738 and 740 are coupled to NVMe drive interface 718.
[0064] Compute node 1000 further includes platform firmware with a storage virtualization extension 1004. This is a firmware-based component that implements an abstraction layer between local and remote physical storage devices by virtualizing the storage devices, exposing them to software running on the compute node, such as an operating system, as local NVMe storage devices. As further shown, the platform firmware with a storage virtualization extension 1004 includes a starter-target configuration map 1006 that includes information for mapping virtual disks to their physical counterparts.
[0065] Storage node 1002 is similar to storage node 702 and includes a network interface 722, a storage distributor 724, and an NVMe drive interface 726 to which multiple NVMe storage devices are coupled, as illustrated by NVMe drives 1008, 1010, and 1012. As before, storage node 1002 may further include an optional processor 734. Storage node 1002 also includes a starter-target configuration map 1014.
[0066] As explained above, the platform firmware with storage virtualization extension virtualizes the physical storage resources and exposes them as local NVMe storage devices to run software on the platform. In the embodiment of Fig. 10, the software includes a Type 1 hypervisor 1016. Type 1 hypervisors, also known as "bare-metal" hypervisors, run directly on platform hardware without an operating system and are used to abstract the platform's physical resources. As further shown, hypervisor 1016 hosts multiple virtual machines, designated as VM1...VM N Each of the VM1 and VM N operates with a respective operating system having an NVMe driver 1018 that is part of the operating system kernel (K) and runs one or more applications in the user space (U) of the operating system.
[0067] In the embodiment of Fig. 10, virtual machines are assembled instead of physical compute nodes. As before, this is enabled in one embodiment by the POD Manager and a PSME communicatively coupled to the compute node. Furthermore, during initialization of the compute node hardware, a communication channel is configured that allows the PSME to convey configuration information to platform firmware with storage virtualization extension 1004. This includes an allocation of local and remote storage resources to the virtual machines VM1 ... VM N .
[0068] As in the Fig. As shown in Figure 10, local NVMe drive 738 and remote NVMe drive 1008 were mapped to VM1 as local virtual NVMe drives 1020 and 1022. Meanwhile, local NVMe drive 740 and remote NVMe drives 1010 and 1012 were mapped to VM Nas virtual NVMe disks 1024, 1026, and 1028. In conjunction with the allocation of physical storage resources to virtual machines, each starter-target configuration map 1006 and starter-target configuration map 1014 are updated with mappings between the virtual NVMe disks and remote NVMe disks accessed through the target (storage node 1002).
[0069] The primary use of starter-target configuration mapping 1006 is to map storage access requests from the NVMe drivers 1018 to remote NVMe drives. Internally, the storage access requests are encapsulated in fabric packets that are sent to the target across the fabric. For an Ethernet fabric, these requests are encapsulated in Ethernet packets that are transmitted across the Ethernet fabric via Ethernet frames. For other fabric types, storage access requests are encapsulated in the fabric packet / frame applicable to the fabric.
[0070] The mapping information in starter-target configuration mapping 1014 is used for a different purpose at the target. The target uses the starter-target mapping information to ensure that each of the target's storage devices can only be accessed by the VM(s) to which each storage device has been mapped. For example, in the configuration shown in Fig. 10, if VM1 attempts to access remote NVMe drive 1010, the target would check the mapping information in launcher-to-target configuration mapping 1014, determine that VM1 is not a permitted launcher for remote NVMe drive 1010, and reject the access request.
[0071] Fig. Figure 10a shows an embodiment similar to that shown in Fig. 10 and set forth above, except that the functionality of platform firmware with a storage virtualization extension 1004 is now implemented as a software storage virtualization extension 1004a in a hypervisor 1106 that is part of the platform software rather than the platform firmware. In a manner similar to that described above, a starter-target configuration mapping 1006 is maintained by software storage virtualization extension 1004a. As before, with respect to operating systems 1...N, virtual NVMe drives 1022, 1026, and 1028 are exposed to local storage devices.
[0072] In addition to supporting virtual machines using the platform firmware with storage virtualization extension scheme, container-based virtual machines are also supported. An example implementation architecture is shown in Fig. 11. As shown, all hardware components (including firmware) are the same. The primary differences between the embodiments of Fig. 10 and Fig. 11 are that hypervisor 1016 has been replaced with an OS virtualization layer 1100 and VM1 ... VM N with container1 ... container N were replaced. In the illustrated embodiment, the OS virtualization layer 1100 is shown running on the platform hardware in a manner similar to a Type 1 hypervisor. In an alternative scheme, an operating system (not shown) may run on the platform hardware, with the OS virtualization layer running on top of the operating system. In this case, the memory abstraction is between platform firmware with memory virtualization extension 1004 and the operating system, rather than between the platform firmware with memory virtualization extension 1004 and the OS virtualization layer.
[0073] Fig. 1a shows an extension of the scheme in Fig. 11, wherein the functionality of platform firmware with storage virtualization extension 1104 is implemented as a software storage virtualization extension 1004b implemented in an OS virtualization layer 1100a that is part of the platform software.
[0074] Fig. 12 shows a flowchart 1200 illustrating operations and logic for implementing a virtualized storage access scheme that may be used with one or more of the embodiments outlined above. In a block 1202, a read or write storage access request is received from an operating system NVMe driver. In a block 1204, the request is received from the storage virtualization device on the platform hardware or software. For example, in the embodiment of Fig. 7, Fig. 7a, Fig. 7b and Fig. 10 and Fig. 11 the device receiving the request is a firmware-based device, while in the embodiments of Fig. 8 and Fig. 8a, Fig. 10a and Fig. 11a the device is a software-based device. In the embodiment of Fig. 7c, the memory virtualization facility is implemented in FPGA 749.
[0075] In a block 1206, a storage device search is performed to determine whether the storage device to be accessed is a local device or a remote device. In one embodiment, the starter-target mapping includes mapping information for local storage devices as well as remote storage devices, and this mapping information is used to determine the storage device and its location. As shown in a decision block 1208, a determination is made as to whether the storage device is local or remote. If it is a local device, the logic proceeds to a block 1210 where the local storage device is accessed. If the access request is a write request, the data supplied with the request is written to the storage device. If the access request is a read request, the data is read from the storage device and, in turn, passed back to the NVMe driver to process the request.
[0076] If the search in block 1206 determines that the device is a remote device, the logic proceeds to a block 1212 where the target is identified. The access request is then encapsulated in a fabric packet (or multiple fabric packets, if applicable), and the fabric packet(s) is / are sent from the initiator across the fabric to the identified target. In a block 1214, the request is received by the target. The target decapsulates the access request and searches for the initiator in its initiator-target mapping to determine whether the initiator is permitted to access the storage device identified in the request, as represented by a decision block 1216. If the answer is NO, the logic proceeds to an end block 1218 where access is denied. Optionally, information identifying the request as denied may be passed back to the requestor (not shown).
[0077] If access is granted, the logic proceeds to a decision block 1220 where a determination is made as to whether the access request is a read or write request. If it is a write request, the logic proceeds to block 1222 where the data is written to the applicable storage device on the target. If the response to decision block 1220 is a read request, the data is read from the applicable storage device and passed back to the launcher in one or more fabric packets sent over the fabric, as shown in block 1224. The read data in the received fabric packet(s) is then decapsulated in a block 1226, and the read data is passed back to the OS NVMe driver to process the read request.
[0078] With respect to accessing virtual NVMe storage devices 742 and 744, the embodiment of Fig.7c in a manner similar to that illustrated in flowchart 1200, with the following difference. Software running on the compute node issues a storage device read or write request through a storage device driver or the like to access data logically stored on one of NVMe virtual storage devices 742 or 744. The read or write request is passed to FPGA 749 via PCIe root controller 716. In response to receiving the request, NVMe function 748c in FPGA 749 generates a network message identifying the network address of storage mode 702 as the destination and includes the read or write request to be sent across the fabric via NIC 720c. In one embodiment, the network message is routed to compute node 700c via the PCIe interconnect fabric. Alternatively, the network message is sent to NIC 720c via interconnect 751.At this point, the operations of blocks 1212, 1214, 1216, 1218, 1220, 1222, 1224, and 1226 are performed in a manner similar to that described above and illustrated in flowchart 1200.
[0079] Further aspects of the subject matter described herein are set out in the following numbered clauses: 1. A method implemented by a compute node communicatively coupled to a remote storage node via a fabric, the remote storage node having a plurality of remote storage devices accessed by the compute node via the fabric, the method comprising: Determining one or more remote storage devices that have been assigned to the compute node, and Exposing each of the one or more remote storage devices to running software on the compute node as a local non-volatile memory express (NVMe) storage device. 2. The method of clause 1, further comprising sending data between the compute node and the remote storage node using an NVMe-Over-Fabric (NVMe-OF) protocol. 3. The method of clause 1 or 2, wherein each of the one or more remote storage devices exposed to the operating system as a local NVMe storage device is a remote NVMe storage device. 4. The method of any preceding clause, wherein the compute node comprises a local NVMe storage device accessed via an NVMe driver configured to access local NVMe storage devices, and the software comprises an operating system running on the compute node, further comprising enabling the operating system to access the one or more remote storage devices via the NVMe driver. 5. The method of any preceding clause, wherein the software comprises an operating system running on physical compute node hardware, and wherein each of the one or more remote storage devices is exposed to the operating system as a local NVMe storage device through use of a hardware-based component in the physical compute node hardware. 6. The method of clause 5, wherein the hardware-based component comprises a network interface controller or a fabric controller. 7. The method of any preceding clause, wherein the software comprises an operating system running on physical compute node hardware, and wherein exposing each of the one or more remote storage devices to running software on the compute node as a local NVMe storage device is implemented by a combination of one or more hardware-based components in the physical compute node hardware and one or more software-based components. 8. The method of clause 7, wherein the one or more software-based components comprise an NVMe volume management device (VMD) driver. 9. The method of clause 8, wherein the NVMe VMD driver comprises a network interface controller (NIC) driver, an NVMe driver, and a VMD driver. 10. The method of clause 8, wherein the compute node comprises a processor having a VMD component, and wherein the one or more hardware-based components comprise the VMD component. 11. Procedure under any of the preceding clauses, which further comprises: Virtualizing a physical storage access infrastructure that has both local and remote storage resources with platform firmware that has a storage virtualization extension, and Exposing, via one of platform firmware, a field-programmable gate array (FPGA), and a software-based hypervisor or operating system virtualization layer, each of the one or more remote storage devices to executing software on the compute node as a local NVMe storage device. 12. A compute node configured to be implemented in a data center environment having a plurality of bays interconnected via a fabric, the plurality of bays comprising a storage pool bay having a plurality of remote storage devices and communicatively coupled via the fabric to a compute node bay in which the compute node is configured to be installed, the compute node comprising: a processor, a memory coupled to the processor, and a fabric controller or network interface controller (NIC) that is operatively coupled to the processor and configured to access the fabric, wherein, when the compute node is installed and operating in the compute node bay, the compute node is configured to expose each of one or more of the plurality of remote storage devices to executing software on the compute node as a local non-volatile memory express (NVMe) storage device. 13. The compute node according to clause 12, wherein the compute node is configured to communicate with the storage pool shelf using an NVMe-Over-Fabric (NVMe-OF) protocol. 14. Compute nodes according to clause 12 or 13, wherein each of the one or more remote storage devices exposed to the software as a local NVMe storage device is a remote NVMe storage device. 15. Compute node according to any of clauses 12-14, wherein the software comprises an operating system having an NVMe driver configured to access local NVMe storage devices, further comprising: at least one local NVMe storage device, wherein the compute node is configured to expose the at least one local NVMe storage device and each of the one or more remote storage devices as multiple local NVMe storage devices and to deploy the NVMe driver to access each of the at least one local NVMe storage device and the one or more remote storage devices. 16. Compute nodes according to any of clauses 12-14, wherein the fabric controller or NIC comprises embedded hardware configured to expose the one or more remote storage devices to the software as local NVMe storage devices. 17. The compute node of any of clauses 12-14, wherein the compute node comprises a combination of one or more hardware-based components and one or more software-based components configured to expose the one or more remote storage devices to the operating system as a local NVMe storage device upon booting of an operating system. 18. Compute nodes according to clause 17, wherein the one or more software-based components comprise an NVMe volume management device (VMD) driver. 19. Compute nodes according to clause 18, wherein the NVMe VMD driver comprises a network interface controller (NIC) driver, an NVMe driver, and a VMD driver. 20. The compute node of clause 18, wherein the compute node comprises a processor having a VMD component, and wherein the one or more hardware-based components comprise the VMD component. 21. Compute nodes according to clause 12, wherein the compute node further comprises: Platform firmware that has a storage virtualization extension, and Software that includes a hypervisor or operating system virtualization layer, and wherein the compute node is further configured to virtualize a physical storage access infrastructure that includes the one or more remote storage devices, and Expose the one or more remote storage devices to the hypervisor or operating system virtualization layer as local NVMe storage devices. 22. The compute node of any of clauses 12-14, further comprising a field-programmable gate array (FPGA) configured to expose the one or more remote storage devices to the software as local NVMe storage devices. 23. Compute nodes according to clause 12, wherein the compute node further comprises: Software that includes a hypervisor or operating system virtualization layer with storage virtualization extension that is configured to virtualize a physical storage access infrastructure that includes the one or more remote storage devices, and Expose each of the one or more remote storage devices to an operating system in a virtual machine or container hosted by the hypervisor or operating system virtualization layer as local NVMe storage devices. 24. System comprising: a plurality of bays interconnected via a fabric, the plurality of bays comprising a storage pool bay having a plurality of remote storage devices and communicatively coupled via the fabric to a compute node bay having a plurality of compute nodes, at least one of the plurality of compute nodes comprising: a processor, a memory coupled to the processor, and a fabric controller or network interface controller (NIC) that is operatively coupled to the processor and configured to access the fabric, wherein, when the compute node is operating, the compute node is configured to expose each of the one or more of the plurality of remote storage devices to executing software on the compute node as a local non-volatile memory express (NVMe) storage device. 25. Scheme under Clause 24, which further includes: a computing device management device that is communicatively coupled to the POD manager and computing node, wherein the system is further configured to receive a request from an Infrastructure-as-a-Service (IaaS) customer requesting the use of computing and storage resources available through the system, create a compute node to use one or more remote storage devices as local NVMe drivers, and Communicate configuration information to the compute node via the compute shelf manager, specifying use of the one or more remote storage devices as local NVMe drives. 26. The system of clause 24 or 25, wherein the compute node is configured to access the one or more remote storage devices using an NVMe-Over-Fabric (NVMe-OF) protocol. 27. The system of any of claims 24-26, wherein the compute node is a launcher and the storage pool slot is a target, and wherein the system is further configured to: Implement starter-target mapping information that maps remote storage devices used by the compute node to be accessed by the storage pool shelf, and use the starter-target mapping information to enable access to the remote storage devices. 28. The system of any of clauses 24-27, wherein the compute node is configured to communicate with the storage pool shelf using an NVMe-Over-Fabric (NVMeOF) protocol. 29. The system of any of clauses 24-28, wherein each of the one or more remote storage devices exposed to the software as local NVMe storage devices is a remote NVMe storage device. 30. The system of any of clauses 24-28, wherein the software comprises an operating system having an NVMe driver configured to access local NVMe storage devices, and a compute node further comprising: at least one local NVMe storage device, wherein the compute node is configured to expose the at least one local NVMe storage device and each of the one or more remote storage devices as multiple local NVMe storage devices and to deploy the NVMe driver to access each of the at least one local NVMe storage device and the one or more remote storage devices. 31. The system of any of clauses 24-28, wherein the fabric controller or NIC comprises embedded hardware configured to expose the one or more remote storage devices to the software as local NVMe storage devices. 32. The system of any of clauses 24-28, wherein the compute node comprises a combination of one or more hardware-based components and one or more software-based components configured to expose each of the one or more remote storage devices to the operating system as local NVMe storage devices upon booting an operating system. 33. A system according to clause 32, wherein the one or more software-based components comprise an NVMe volume management device (VMD) driver. 34. The system of clause 33, wherein the NVMe VMD driver comprises a network interface controller (NIC) driver, an NVMe driver, and a VMD driver. 35. The system of clause 33, wherein the compute node comprises a processor having a VMD component, and wherein the one or more hardware-based components comprise the VMD component. 36. The system of any of clauses 24-29, wherein the compute node further comprises: Platform firmware that has a storage virtualization extension, and Software that includes a hypervisor or operating system virtualization layer, and wherein the compute node is further configured to virtualize a physical storage access infrastructure that includes the one or more remote storage devices, and Expose the one or more remote storage devices to the hypervisor or operating system virtualization layer as local NVMe storage devices. 37. The system of any of clauses 24-29, further comprising a field-programmable gate array (FPGA) configured to expose the one or more remote storage devices to the software as local NVMe storage devices. 38. The system of any of clauses 24-29, wherein the compute node further comprises: Software that includes a hypervisor or operating system virtualization layer with storage virtualization extension that is configured to virtualize a physical storage access infrastructure that includes the one or more remote storage devices, and expose each of the one or more remote storage devices to an operating system in a virtual machine or container hosted by the hypervisor or operating system virtualization layer as a local NVMe storage device. 39. A non-transitory machine-readable medium having instructions stored thereon that is configured to be executed by a processor in a compute node in a data center environment having a plurality of bays interconnected by a fabric, the plurality of bays comprising a storage pool bay having a plurality of remote storage devices and communicatively coupled by the fabric to a compute node bay in which the compute node is installed, the compute node having a fabric controller or network interface controller (NIC) operably coupled to the processor and configured to access the fabric, the instructions configured to, when executed by the processor, subject each of the one or more remote storage devices to executing software on the compute node as local non-volatile memory express (NVMe) storage devices. 40. Non-transitory machine-readable medium as defined in Clause 39, where the instructions are firmware instructions. 41. Non-transitory machine-readable medium as defined in Clause 40, wherein the execution of software on the compute node to which the one or more remote storage devices are exposed as local NVMe storage devices comprises a hypervisor. 42. Non-transitory machine-readable medium as defined in Clause 40, wherein the software executing on the compute node to which the one or more remote storage devices are exposed as local NVMe storage devices comprises an operating system. 43. Non-transitory machine-readable medium as defined in Clause 40, wherein executing software on the compute node to which the one or more remote storage devices are exposed as local NVMe storage devices comprises an operating system virtualization layer. 44. Non-transitory machine-readable medium as defined in Clause 39, where the instructions are software instructions that include a hypervisor. 45. Non-transitory machine-readable medium as defined in Clause 39, where the instructions are software instructions that include an operating system virtualization layer. 46. Non-transitory machine-readable medium as defined in Clause 39, wherein each of the one or more remote storage devices exposed to the Software as local NVMe storage devices is a remote NVMe storage device.
[0080] Although some embodiments have been described with respect to particular implementations, other implementations are possible according to some embodiments. Furthermore, the arrangement and / or order of elements or features illustrated in the drawings and / or described herein need not be arranged in the particular manner illustrated and described. Many other arrangements are possible according to some embodiments.
[0081] In any system shown in a figure, the elements may, in some cases, have the same reference number or a different reference number to indicate that the illustrated elements may be different and / or similar. However, an element may be flexible enough to have different implementations and function with some or all of the systems shown and described herein. The various elements shown in the figures may be the same or different. What is referred to as the first element and what is called the second element is arbitrary.
[0082] Throughout the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms are not to be construed as synonyms. In certain embodiments, "connected" may be used to indicate that two or more elements are in direct physical or electrical contact with each other. "Coupled" may mean that two or more elements are in direct physical or electrical contact. However, "coupled" may also mean that two or more elements are not in direct contact with each other, but rather cooperate or interact with each other.
[0083] An embodiment is one implementation or example of the inventions. References in the specification to "one embodiment," "some embodiments," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments is included in at least some embodiments, but not necessarily all embodiments of the inventions. The various uses of "one embodiment" or "one embodiment" do not necessarily all refer to the same embodiments.
[0084] Not all components, features, structures, properties, etc., described and illustrated herein need not be included in a particular embodiment or embodiments. The specification states that a component, feature, or property "may" or "could" be included; for example, the particular component, feature, structure, or property need not be included. If the specification or claim refers to "a" element, that does not mean that only one of the element is present. If the specification or claims mention "an additional" element, that does not preclude the presence of more than one of the additional elements.
[0085] Letters such as 'M' and 'N' in the foregoing detailed description and the drawings are used to denote an integer, and the use of a particular letter is not limited to specific embodiments. Furthermore, the same letter may be used in separate claims to denote separate integers, or different letters may be used. Furthermore, the use of a particular letter in the detailed description may or may not match the letter in a claim pertaining to the same subject matter in the detailed description.
[0086] As stated above, various aspects of the embodiments herein may be enabled by corresponding software and / or firmware components and applications, such as software and / or firmware executed by an embedded processor or the like. Therefore, embodiments of this invention may be used as or support software programs, software modules, firmware, and / or distributed software executing in or supporting some type of processor, processor core, or embedded logic or virtual machine running on or otherwise implemented or realized by a processor or core, or within a computer-readable or machine-readable non-transitory storage medium. A computer-readable or machine-readable non-transitory storage medium includes any mechanism for storing or transmitting information in a readable form by a machine (e.g.,a computer). For example, a computer-readable or machine-readable non-transitory storage medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form accessible by a computer or computing machine (e.g., computing device, electronic system, etc.), such as writable / non-writable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.). The content may be directly executable ("object" or "executable" form), source code, or other code ("delta" or "patch" code). A computer-readable or machine-readable non-transitory storage medium may also include memory or a database from which content may be downloaded.The computer-readable or machine-readable non-transitory storage medium may also comprise a device or product having content stored thereon at a time or point of delivery. Therefore, delivering a device with stored content or offering content for download via a communications medium may be considered providing an article of manufacture comprising a computer-readable or machine-readable non-transitory storage medium with such content as described herein.
[0087] Various components, referred to above as processes, servers, or tools, described herein may be a means for performing the described functions. The operations and functions performed by various components described herein may be implemented by executing software on a processing element, via embedded hardware or the like, or any combination of hardware and software. Such components may be implemented as software modules, hardware modules, special-purpose hardware (e.g., application-specific hardware, ASICs, DSPs, etc.), embedded controllers, hard-wired circuits, hardware logic, etc. Software content (e.g., data, instructions, configuration information, etc.)) may be provided via an article of manufacture comprising a computer-readable or machine-readable non-transitory storage medium that provides content that represents executable instructions. The content may cause a computer to perform various functions / operations described herein.
[0088] As used herein, a list of items joined by the phrase "at least one of" may mean any combination of the listed terms. For example, the phrase "at least one of A, B, or C" may mean A; B; C; A and B; A and C; B and C; or A, B, and C.
Claims
[1] A method (900) implemented by a compute node (602) communicatively coupled to a remote storage node (618) via a fabric (614), the remote storage node having a plurality of remote storage devices (624) accessed by the compute node via the fabric, the method comprising: Assembling compute node resources, including associating one or more of the plurality of remote storage devices with the compute node prior to starting an operating system on the compute node (904), and Exposing each of the one or more remote storage devices associated with the compute node (904) to software running on the compute node as a local non-volatile memory express, NVMe, storage device (912). [2] The method of claim 1, further comprising sending data between the compute node and the remote storage node using an NVMe-Over-Fabric, NVMe-OF, protocol. [3] The method of claim 1 or 2, wherein each of the one or more remote storage devices exposed to the operating system as a local NVMe storage device is a remote NVMe storage device. [4] The method of any preceding claim, wherein the compute node comprises a local NVMe storage device accessed via an NVMe driver configured to access local NVMe storage devices, and the software comprises an operating system running on the compute node, further comprising enabling the operating system to access the one or more remote storage devices via the NVMe driver. [5] The method of any preceding claim, wherein the software comprises an operating system running on physical compute node hardware, and wherein each of the one or more remote storage devices is exposed to the operating system as a local NVMe storage device by use of a hardware-based component in the physical compute node hardware. [6] The method of claim 5, wherein the hardware-based component comprises a network interface controller or a fabric controller. [7] The method of any preceding claim, wherein the software comprises an operating system running on physical compute node hardware, and wherein exposing each of the one or more remote storage devices to running software on the compute node as a local NVMe storage device is implemented by a combination of one or more hardware-based components in the physical compute node hardware and one or more software-based components. [8] The method of claim 7, wherein the one or more software-based components comprise an NVMe volume management device, VMD, driver. [9] The method of claim 8, wherein the NVMe VMD driver comprises a network interface controller, NIC, driver, an NVMe driver, and a VMD driver. [10] The method of claim 8, wherein the compute node comprises a processor having a VMD component, and wherein the one or more hardware-based components comprise the VMD component. [11] A method according to any one of the preceding claims, further comprising: Virtualizing a physical storage access infrastructure that includes both local and remote storage resources with platform firmware that includes a storage virtualization extension, and Exposing, via one of platform firmware, a field-programmable gate array, FPGA, and a software-based hypervisor or operating system virtualization layer, each of the one or more remote storage devices to executing software on the compute node as a local NVMe storage device. [12] A compute node (210) configured to be implemented in a data center environment (202) having a plurality of bays interconnected via a fabric (218), the plurality of bays comprising a storage pool bay (212) having a plurality of remote storage devices and communicatively coupled via the fabric to a compute node bay (210) in which the compute node is configured to be installed, the compute node comprising: a processor, a memory coupled to the processor, and a fabric controller or network interface controller, NIC, that is operatively coupled to the processor and configured to access the fabric, wherein, when the compute node is installed and operating in the compute node bay, the compute node is configured to Assemble compute node resources, including associating one or more of the plurality of remote storage devices with the compute node prior to starting an operating system on the compute node (904), expose each of one or more of the plurality of remote storage devices associated with the compute node (904) to software running on the compute node as a local non-volatile memory express, NVMe, storage device. [13] The compute node of claim 12, wherein the compute node is configured to communicate with the storage pool shelf using an NVMe-Over-Fabric, NVMe-OF, protocol. [14] The compute node of claim 12 or 13, wherein each of the one or more remote storage devices exposed to the software as a local NVMe storage device is a remote NVMe storage device. [15] The compute node of any of claims 12-14, wherein the software comprises an operating system having an NVMe driver configured to access local NVMe storage devices, further comprising: at least one local NVMe storage device, wherein the compute node is configured to expose the at least one local NVMe storage device and each of the one or more remote storage devices as multiple local NVMe storage devices and to deploy the NVMe driver to access each of the at least one local NVMe storage device and the one or more remote storage devices. [16] The compute node of any of claims 12-14, wherein the fabric controller or NIC comprises embedded hardware configured to expose the one or more remote storage devices to the software as local NVMe storage devices. [17] The compute node of any of claims 12-14, wherein the compute node comprises a combination of one or more hardware-based components and one or more software-based components configured to expose the one or more remote storage devices to the operating system as local NVMe storage devices upon booting an operating system. [18] The compute node of claim 17, wherein the one or more software-based components comprise an NVMe volume management device, VMD, driver. [19] The compute node of claim 18, wherein the NVMe VMD driver comprises a network interface controller, NIC, driver, an NVMe driver, and a VMD driver. [20] The compute node of claim 18, wherein the compute node comprises a processor having a VMD component, and wherein the one or more hardware-based components comprise the VMD component. [21] The compute node of claim 12, wherein the compute node further comprises: Platform firmware that has a storage virtualization extension, and Software that includes a hypervisor or operating system virtualization layer, and wherein the compute node is further configured to virtualize a physical storage access infrastructure that includes the one or more remote storage devices, and Expose the one or more remote storage devices to the hypervisor or operating system virtualization layer as local NVMe storage devices. [22] System (200) comprising: a plurality of shelves interconnected via a fabric (218), the plurality of shelves comprising a storage pool shelf (212) having a plurality of remote storage devices and communicatively coupled via the fabric to a compute node shelf (210) comprising a plurality of compute nodes, at least one of the plurality of compute nodes comprising: a processor, a memory coupled to the processor, and a fabric controller or network interface controller, NIC, operatively coupled to the processor and configured to access the fabric, wherein, when the compute node is operating, the compute node is configured to Assemble compute node resources, including associating one or more of the plurality of remote storage devices with the compute node prior to starting an operating system on the compute node (904), expose each of one or more of the plurality of remote storage devices associated with the compute node (904) to software running on the compute node as a local non-volatile memory express, NVMe, storage device. [23] The system of claim 22, further comprising: a compute module management device communicatively coupled to the pod manager and the compute node, wherein the system is further configured to receive a request from an Infrastructure-as-a-Service (IaaS) customer requesting the use of computing and storage resources available through the system, create a compute node to use one or more remote storage devices as local NVMe drives, and Communicate configuration information to the compute node via the compute bay manager, specifying use of the one or more remote storage devices as local NVMe drives. [24] The system of claim 22 or 23, wherein the compute node is configured to access the one or more remote storage devices using an NVMe-Over-Fabric, NVMe-OF, protocol. [25] The system of any of claims 22-24, wherein the compute node is a launcher and the storage pool slot is a target, and wherein the system is further configured to: Implement starter-target mapping information that maps remote storage devices used by the compute node to be accessed by the storage pool shelf, and use the starter-target mapping information to enable access to the remote storage devices.
Citation Information
Patent Citations
Method and apparatus for storing data
US20160259568A1
Systems and methods for enabling access to extensible storage devices over a network as local storage via NVME controller
US9294567B2