Multi-rental SSD configuration
By identifying tenant identifiers and assigning performance level generation parameters to multi-tenant SSD configurations, the problem of inconsistent tenant performance management in multi-tenant SSDs is solved, thereby improving tenant service quality and resource utilization.
Patent Information
- Application Number
- CN202480026582.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2024-04-17
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multi-tenant SSD configurations struggle to achieve differentiated management and optimization of performance levels among tenants when sharing computing resources, leading to increased latency, access inconsistencies, and low resource utilization.
By identifying tenant identifiers, assigning different performance levels, generating corresponding performance parameters, and sending configuration messages to storage devices, customized management of multi-tenant SSDs can be achieved. This includes identifying virtual functions and reclaimed cell handles. The configuration messages contain performance parameters and tenant identifiers, supporting multi-tenant SSD access control and load balancing.
It improved the quality of service for tenants, reduced SSD latency, improved access and sharing efficiency among multiple tenants, extended the lifespan and reliability of SSDs, enhanced application performance and infrastructure efficiency, and enabled customized SSD operations.
Smart Images

Figure CN120981791A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 460,300, filed April 18, 2023, U.S. Provisional Patent Application Serial No. 63 / 550,033, filed February 5, 2024, and U.S. Patent Application Serial No. 18 / 634,873, filed April 12, 2024, which are incorporated herein by reference for all purposes. Technical Field
[0003] This disclosure generally relates to memory systems, and more specifically to multitenancy SSD configurations. Background Technology
[0004] This background section is intended to provide context only, and the disclosure of any concepts in this section does not constitute an admission that the concepts are prior art.
[0005] Cloud computing delivers computing services over the internet. These services include servers, storage, databases, networking, software, analytics, and intelligence. Multitenancy occurs when several different cloud customers are accessing the same computing resources, such as when several different companies are storing data on the same physical server.
[0006] The information disclosed in this background section is only intended to enhance the understanding of the background of this disclosure, and therefore may contain information that does not constitute prior art. Summary of the Invention
[0007] In various embodiments, the systems and methods described herein include systems, methods, and apparatus for multi-tenant SSD configuration. In some aspects, the technology described herein relates to a multi-tenant method comprising: identifying an identifier of a first tenant of a storage device; assigning a first performance level to the first tenant; generating a first performance parameter based on the first performance level; and sending a configuration message to the storage device including the first performance parameter and the identifier of the first tenant.
[0008] In some aspects, the techniques described herein relate to a method in which the identifier of the first tenant includes the virtual function of the storage device and the recycling unit handle of the storage device.
[0009] In some respects, the techniques described herein relate to a method in which the identifier of the first tenant includes at least one of the following: physical function of the storage device, port of the storage device, stream of the storage device, zone of the storage device, logical block address range of the storage device, non-volatile memory (NVM) controller of the storage device, commit queue, or scalable input / output virtualization.
[0010] In some respects, the techniques described herein relate to a method in which a storage device identifier group, a command type associated with a command generated by a first tenant, or a command identifier associated with a command generated by a first tenant.
[0011] In some respects, the techniques described herein relate to a method in which configuration messages include one or more tenant identifier fields.
[0012] In some respects, the technology described herein relates to a method in which a configuration message includes one or more performance parameter fields based on a first performance level, the one or more performance parameter fields including at least one field for the maximum allowed input / output operations per second (IOPS) for the first tenant on the storage device or for the reservation level of IOPS for the first tenant on the storage device.
[0013] In some respects, the technology described herein relates to a method in which a configuration message includes one or more performance parameter fields based on a first performance level, the one or more performance parameter fields including at least one field for the maximum available communication bandwidth between the first tenant and the storage device or for a reservation level of the communication bandwidth between the first tenant and the storage device.
[0014] In some respects, the technology described herein relates to a method in which: a configuration message includes a performance parameter field for a tenant load of a first tenant, the performance parameter field indicating a performance level of a request based on a first performance level, and the tenant load of the first tenant includes at least one of a queue depth (QD) for commands generated by the first tenant or a command length associated with commands generated by the first tenant.
[0015] In some aspects, the techniques described herein relate to a method in which a configuration message includes one or more performance parameter fields for at least one of the following: a first tenant’s proportional bandwidth win rate relative to at least one other tenant’s bandwidth win rate, a first tenant’s proportional IOPS win rate relative to at least one other tenant’s IOPS win rate, a maximum level of change in access to the storage device over a period of time, an access consistency level between the first tenant and the storage device, or a maximum allowed access latency between the first tenant and the storage device.
[0016] In some respects, the techniques described herein relate to a method that also includes: identifying an identifier for a second tenant; and assigning a second performance level to the second tenant that is different from the first performance level.
[0017] In some respects, the techniques described herein relate to a method that also includes generating a second performance parameter based on a second performance level, wherein: the second performance parameter is different from the first performance parameter, and the configuration message includes the second performance parameter and an identifier of the second tenant.
[0018] In some respects, the techniques described herein relate to a method in which the format of configuration messages is based on a fast-blocking format of non-volatile memory.
[0019] In some respects, the techniques described herein relate to a method in which the storage device includes a solid-state drive.
[0020] In some aspects, the technology described herein relates to a device comprising: at least one memory; and at least one processor coupled to the at least one memory, configured to: identify an identifier of a first tenant of the storage device; assign a first performance level to the first tenant; generate a first performance parameter based on the first performance level; and send a configuration message to the storage device including the first performance parameter and the identifier of the first tenant.
[0021] In some respects, the technology described herein relates to a device that also includes: an identifier for identifying a second tenant; and assigning a second performance level to the second tenant that is different from the first performance level.
[0022] In some aspects, the technology described herein relates to a device in which at least one processor is configured to generate a second performance parameter based on a second performance level, wherein: the second performance parameter is different from a first performance parameter, and a configuration message includes the second performance parameter and an identifier of a second tenant.
[0023] In some aspects, the technology described herein relates to a device in which the identifier of the first tenant includes the virtual function of the storage device and the recycling unit handle of the storage device.
[0024] In some respects, the techniques described herein relate to a non-transitory computer-readable medium storing code including instructions executable by a processor of a device to: identify an identifier of a first tenant of the storage device; assign a first performance level to the first tenant; generate a first performance parameter based on the first performance level; and send a configuration message to the storage device including the first performance parameter and the identifier of the first tenant.
[0025] In some respects, the techniques described herein relate to a non-transitory computer-readable medium, wherein the code includes further instructions executable by a processor to cause a device to: identify an identifier of a second tenant; and assign a second performance level to the second tenant that is different from the first performance level.
[0026] In some respects, the techniques described herein relate to a non-transitory computer-readable medium, wherein the code includes further instructions executable by a processor to cause a device to: generate a second performance parameter based on a second performance level, wherein the second performance parameter is different from a first performance parameter, and a configuration message includes the second performance parameter and an identifier of a second tenant.
[0027] A computer-readable medium is disclosed. The computer-readable medium may store instructions that, when executed by a computer, cause the computer to perform operations that are substantially the same as or similar to those described herein. Similarly, non-transitory computer-readable media, apparatuses, and systems for performing operations substantially the same as or similar to those described herein are further disclosed.
[0028] The systems and methods described in this paper offer numerous advantages and benefits. For example, the multi-tenant systems and methods described herein improve the quality of service for host tenants, thereby providing improved service consistency within defined service variation constraints. Multi-tenant systems and methods can reduce SSD latency for one or more tenants, thereby improving the sharing of SSD access among multiple tenants. Multi-tenant systems and methods enable balanced and customized SSD operations, thereby increasing SSD lifespan and reliability. The multi-tenant systems and methods described herein provide improved application performance and infrastructure efficiency, which improves overall data center system efficiency and utilization. Multi-tenant systems and methods provide improved SSD access control and customization, enabling SSDs to adapt to varying tenant loads across multiple tenants. Therefore, the systems and methods described herein allow hosts to specify or request relative settings for various tenants with increased specificity. Attached Figure Description
[0029] The foregoing and other aspects of this system and method will be better understood when this application is read in conjunction with the following accompanying drawings, in which like reference numerals denote similar or identical elements. Furthermore, the drawings provided herein are for illustrative purposes only; other embodiments, which may not be explicitly shown, are not excluded from the scope of this disclosure.
[0030] These and other features and advantages of this disclosure will be appreciated and understood by referring to the specification, claims and drawings, wherein:
[0031] Figure 1 An example system according to one or more embodiments as described herein is shown.
[0032] Figure 2 The following are illustrated according to one or more embodiments as described herein. Figure 1 The details of the system.
[0033] Figure 3An example system according to one or more embodiments as described herein is shown.
[0034] Figure 4 An example system according to one or more embodiments as described herein is shown.
[0035] Figure 5 A grouping format according to one or more embodiments as described herein is shown.
[0036] Figure 6-11 An example system according to one or more embodiments as described herein is shown.
[0037] Figure 12 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0038] Figure 13 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0039] Figure 14 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0040] Figure 15 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0041] Figure 16 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0042] Figure 17 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0043] Figure 18 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0044] Figure 19 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0045] While this system and method may have various modifications and alternatives, specific embodiments thereof are illustrated by way of example in the accompanying drawings and will be described herein. The drawings may not be to scale. However, it should be understood that the drawings and their detailed description are not intended to limit the system and method to the specific forms disclosed, but rather are intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the system and method as defined by the appended claims. Detailed Implementation
[0046] Details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims.
[0047] Various embodiments of this disclosure will now be described more fully below with reference to the accompanying drawings, which illustrate some, but not all, of the embodiments. In fact, this disclosure may be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure may meet applicable legal requirements. Unless otherwise stated, the term “or” is used herein in a meaning that is alternative and combined. The terms “illustrative” and “example” are used as examples without an indication of quality level. The same reference numerals always denote the same elements. Arrows in each figure depict bidirectional data flow and / or bidirectional data flow capability. The terms “path,” “path,” and “route” are used interchangeably herein.
[0048] Embodiments of this disclosure can be implemented in various ways, including as a computer program product comprising an article of manufacture. A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program components, scripts, source code, program code, object code, bytecode, compiled code, interpreted code, machine code, executable instructions, etc. (also referred to herein as executable instructions, instructions for execution, computer program product, program code, and / or similar terms used interchangeably herein). Such a non-transitory computer-readable storage medium includes all computer-readable media (including volatile and non-volatile media).
[0049] In one embodiment, a non-volatile computer-readable storage medium may include a floppy disk, flexible disk, hard disk, solid-state storage (SSS) (e.g., a solid-state drive (SSD)), solid-state card (SSC), solid-state module (SSM), enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium. Non-volatile computer-readable storage media may include punched cards, paper tape, optical marking sheets (or any other physical medium having a pattern of holes or other optically identifiable markings), compact disc read-only memory (CD-ROM), rewritable compact disc (CD-RW), digital versatile disc (DVD), Blu-ray disc (BD), or any other non-transitory optical medium. Such non-volatile computer-readable storage media may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., serial, NAND, NOR, etc.), multimedia memory card (MMC), secure digital (SD) memory card, smart media card, compact flash (CF) card, memory stick, etc. In addition, non-volatile computer-readable storage media may include conductive bridged random access memory (CBRAM), phase change random access memory (PRAM), ferroelectric random access memory (FeRAM), non-volatile random access memory (NVRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), silicon-oxide-nitride-oxide-silicon memory (SONOS), floating junction gate random access memory (FJG RAM), Millipede memory, racetrack memory, etc.
[0050] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), high-bandwidth memory (HBM), fast page mode dynamic random access memory (FPM DRAM), extended data output dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type 2 synchronous dynamic random access memory (DDR2 SDRAM), double data rate type 3 synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), dual transistor RAM (TTRAM), thyristor RAM (T-RAM), zero capacitor (Z-RAM), Rambus through-hole memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, etc. It should be understood that, where embodiments are described as using computer-readable storage media, other types of computer-readable storage media may be used in place of the computer-readable storage media described herein, or other types of computer-readable storage media may be used in addition to the computer-readable storage media described herein.
[0051] High-bandwidth memory (HBM) can be a type of computer memory that uses 3D stacking technology to provide high bandwidth and low power consumption. HBM can be used in high-performance computing applications that require high data speeds. An HBM stack can contain up to several DRAM modules (e.g., eight DRAM modules), each connected via two channels.
[0052] As should be understood, the various embodiments of this disclosure can be implemented as methods, apparatus, systems, computing devices, computing entities, etc. Thus, embodiments of this disclosure can take the form of apparatuses, systems, computing devices, computing entities, etc., that execute instructions stored on a computer-readable storage medium to perform certain steps or operations. Therefore, embodiments of this disclosure can take the form of entirely hardware embodiments performing certain steps or operations, entirely computer program product embodiments, and / or embodiments including a combination of computer program products and hardware.
[0053] This document describes embodiments of the present disclosure with reference to block diagrams and flowcharts. Therefore, it should be understood that each block of the block diagrams and flowcharts can be implemented as a computer program product, a complete hardware embodiment, a combination of hardware and computer program products, and / or as instructions, operations, steps, and interchangeable similar terms (e.g., executable instructions, instructions for execution, program code, etc.) for operation on a computer-readable storage medium, in the form of an apparatus, system, computing device, computing entity, etc. For example, code retrieval, loading, and execution can be performed sequentially, such that one instruction is retrieved, loaded, and executed at a time. In some example embodiments, retrieval, loading, and / or execution can be performed in parallel, such that multiple instructions are retrieved, loaded, and / or executed together. Therefore, such embodiments can produce machines specifically configured to perform the steps or operations specified in the block diagrams and flowcharts. Thus, the block diagrams and flowcharts support various combinations of embodiments for performing specified instructions, operations, or steps.
[0054] The following description is presented to enable those skilled in the art to make and use the subject matter disclosed herein and incorporate it into the context of a particular application. While specific examples are given below, other and further examples may be devised without departing from its basic scope.
[0055] Various modifications and multiple uses in different applications will be apparent to those skilled in the art, and the general principles defined herein can be applied to a wide range of embodiments. Therefore, the subject matter disclosed herein is not intended to be limited to the presented embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0056] The provided description sets forth numerous specific details to provide a more thorough understanding of the subject matter disclosed herein. However, it will be apparent to those skilled in the art that the subject matter disclosed herein can be practiced without being limited to these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the subject matter disclosed herein.
[0057] Unless otherwise expressly stated, all features disclosed in this specification (e.g., any appended claims, abstract, and drawings) may be replaced by alternative features for the same, equivalent, or similar purposes. Therefore, unless otherwise expressly stated, each disclosed feature is merely an example of a general series of equivalent or similar features.
[0058] This document describes various features with reference to the accompanying drawings. It should be noted that the drawings are intended only to facilitate the description of the features. The various features described are not intended as an exhaustive description of the subject matter disclosed herein or as a limitation on the scope of the subject matter disclosed herein. Furthermore, the examples shown do not need to possess all the aspects or advantages shown. Aspects or advantages described in connection with an example are not necessarily limited to that example and can be practiced in any other example, even if not so shown or so explicitly described.
[0059] Furthermore, any element in the claims that does not expressly state "means" for performing the specified function or "step" for performing a particular function shall not be construed as a "means" or "step" as specified in paragraph 6 of Section 112 of 35 USC. In particular, the use of "step" or "action" in the claims herein is not intended to invoke the provisions of paragraph 6 of 35 U.S.C. 112.
[0060] It should be noted that, if used, the labels left, right, front, back, top, bottom, forward, reverse, clockwise, and counterclockwise are for convenience only and are not intended to suggest any particular fixed direction. Rather, the labels are used to reflect the relative position and / or orientation between the various parts of an object.
[0061] Any data processing may include data buffering, alignment of incoming data from multiple communication channels, forward error correction (“FEC”), and / or others. For example, data may first be received by an analog front end (AFE) that prepares the incoming data for digital processing. The digital portion of the transceiver (e.g., a DSP) may provide skew management, equalization, reflection cancellation, and / or other functions. It should be understood that the processes described herein can provide numerous benefits, including power and cost savings.
[0062] Furthermore, the terms "system," "component," "module," "interface," and "model" are generally intended to refer to computer-related entities, hardware, combinations of hardware and software, software, or software in operation. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer. For illustration, both an application running on a controller and the controller itself can be components. One or more components may reside within a process and / or a thread of execution, and components may be centralized on a single computer and / or distributed across two or more computers.
[0063] Unless otherwise explicitly stated, each value and range may be interpreted as approximate, as the words "about" or "approximately" precede the value or range. Signals and their corresponding nodes or ports may be represented by the same name and are interchangeable here for the purpose of reference.
[0064] While embodiments may have been described with respect to circuit functionality, embodiments of the subject matter disclosed herein are not limited. Possible implementations may be embodied in a single integrated circuit, a multi-chip module, a single card, a system-on-a-chip, or a multi-card circuit package. It will be apparent to those skilled in the art that various embodiments can also be implemented as part of a larger system. Such embodiments can be used in conjunction with, for example, digital signal processors, microcontrollers, field-programmable gate arrays, application-specific integrated circuits, or general-purpose computers.
[0065] It will be apparent to those skilled in the art that the various functions of circuit elements can also be implemented as processing blocks in a software program. Such software can be used in, for example, digital signal processors, microcontrollers, or general-purpose computers. Such software can be embodied in the form of program code contained in a tangible medium, such as a magnetic recording medium, optical recording medium, solid-state memory, floppy disk, CD-ROM, hard disk drive, or any other non-transitory machine-readable storage medium, which, when loaded into and executed by the program code, becomes a means for practicing the subject matter disclosed herein. When implemented on a general-purpose processor, the program code segments are combined with the processor to provide a unique device that operates similarly to a particular logic circuit. The described embodiments can also be embodied in the form of bit streams or sequences of other signal values transmitted electrically or optically through a medium using the methods and / or apparatus described herein, magnetic field variations stored in a magnetic recording medium, etc.
[0066] In some examples, a solid-state drive (SSD) is a storage device used in computers to store data on solid-state flash memory (e.g., NAND flash memory). NAND flash memory is a non-volatile storage technology that stores data without requiring power. NAND flash memory can also be referred to as a memory chip. Flash memory cards and SSDs use multiple NAND flash memory chips to store data. In data management, "hot" data is data that is frequently accessed and / or in high demand, while "cold" data includes data that is infrequently accessed and / or infrequently demanded (e.g., settings and forgotten data). Hot data can include data that is regularly demanded, data that is in transit or periodically transited, and / or data that is stored for a relatively short period of time. Data hotness can be the relative degree of frequency of accessing or requesting data.
[0067] Hot data can be stored on media designed for fast access with multiple connections and high performance. Cold data can be stored on media with relatively slow access times. Cold storage is ideal for data that needs to be retained for long periods and / or is unlikely to change (e.g., historical data, compliance data, and legal documents). Therefore, hot storage can be used for data that is needed quickly or accessed frequently, while cold storage can be used for data that is rarely needed. Cold storage solutions typically have robust data security features, including encryption, access control, and redundancy. Cold cloud storage is relatively cheaper than warm or hot storage, but it has a higher cost per operation than other types of cloud storage.
[0068] SSDs work alongside a computer's memory (Random Access Memory (RAM)) and processor to access and use data. This includes files such as operating systems, programs, documents, games, images, and media. SSDs are permanent or non-volatile storage devices, meaning they retain stored data even when the computer is powered off. SSDs can be used as secondary storage within a computer's storage tier.
[0069] In an SSD, a page can be the smallest unit, while a block can be the smallest unit of access. A page can be 4 kilobytes (KB) in size. A page, consisting of several storage units, is the smallest unit of an SSD. Several pages on an SSD can be grouped together as a block. A block is the smallest unit of access on an SSD (e.g., read, write, erase, etc.). In some examples, 128 pages can be combined into a single block, which comprises 512 KB. A block can be called an erase unit. The size of a block or erase unit determines the granularity of garbage collection (GC) on the SSD (e.g., at the SSD software level). A Logical Block Address (LBA) is a standard used to specify the address for read and write commands on an SSD. Most SSDs report their LBA size as 512 bytes, even if they physically use larger blocks. These blocks are typically 4 kiB, 8 kiB, or sometimes larger.
[0070] Unlike hard disk drives (HDDs), SSDs and other NAND flash memory storage do not overwrite existing data. Instead, SSDs can undergo programming / erasing cycles. SSD garbage collection (GC) is an automated process that improves the write performance of an SSD. The goal of garbage collection is to periodically optimize the drive so that it operates efficiently and maintains performance throughout its lifespan. Using SSD garbage collection, the SSD (e.g., the SSD's storage controller or storage processing unit) searches for pages that have been marked as obsolete (e.g., outdated, obsolete, or no longer accurate data). The SSD copies the data still in use to a new block and then deletes all data from the old block. The SSD marks the old data as invalid and writes the new data to a new physical location.
[0071] Compute storage can include storage device architectures that allow data processing at the storage device level. Compute storage adds computing resources (e.g., processing units) to storage devices. It is also known as in-situ processing or in-storage computing. Compute storage devices have processors that can run specific computational functions directly within the storage hardware. This allows for the ability to perform selected computational tasks within or near the storage device (rather than the central processing unit of a server or computer). Therefore, compute storage reduces the amount of data that needs to be moved between the storage plane and the compute plane. Compute storage adds computation to storage in a way that drives efficiency and enables enhanced complementary functionality. Compute storage architectures improve application performance and infrastructure efficiency.
[0072] Peripheral Component Interconnect Fast (PCIe) can include an interface for high-speed data transfer between electronic components in a computer system. PCIe can be used to connect expansion cards to the motherboard, such as graphics cards, network cards, storage devices (e.g., SSDs), storage controllers, memory devices, memory controllers, processors, etc. In some examples, PCIe slots can connect the computer motherboard to peripheral components (e.g., PCIe x1, PCIe x4, PCIe x8, PCIe x16). PCIe can be forward- and / or backward-compatible. For example, a PCIe 3.0 card can be placed in a PCIe 4.0 slot, but the PCIe 3.0 card can be limited to the lower speed of PCIe 3.0.
[0073] In some examples, Non-Volatile Memory Fast (NVMe) is a data transfer protocol that can be configured to connect SSD storage to a server and / or processor using a PCIe bus. NVMe was created to improve the speed and performance of computer systems. An NVMe controller may include a logical device interface specification that allows access to a computer's non-volatile storage media. NVMe controllers are optimized for high-performance random read / write operations. In some cases, an NVMe controller can perform flash management operations on the on-chip SSD while consuming negligible host processing and memory resources. NVMe can perform parallel input / output (I / O) operations with multi-core processors to facilitate high throughput. NVMe controllers can map I / O and responses to shared memory in the host computer via a PCIe interface. In some cases, NVMe controllers can communicate directly with the host central processing unit (CPU).
[0074] In some examples, PCIe can use functions to enable individual access to its resources. These functions can include physical functions (e.g., PCIe physical functions) and / or virtual functions (e.g., PCIe virtual functions). In some cases, a PCIe device can be partitioned into multiple physical functions. In some examples, a Single Root I / O Virtualization (SR-IOV) interface is an extension of PCIe. SR-IOV can configure a physical device to behave as multiple separate physical devices (e.g., for a hypervisor, for a guest operating system, etc.). In some cases, SR-IOV allows a device (e.g., a network adapter) to decouple access to its resources among various PCIe hardware functions. These functions can include physical functions (e.g., PCIe physical functions) and / or virtual functions (e.g., PCIe virtual functions). In some examples, SR-IOV can enable a PF0 and one or more VFs (e.g., where VFs and PFs serve similar functions). In some cases, reconfiguration can provide various combinations of PFs and VFs.
[0075] In some scenarios, a PCIe Physical Function (PF) encompasses the primary functionality of a device. A PF can announce the device's SR-IOV capabilities. A PF is a full-featured PCIe function that can be managed, discovered, and manipulated like any other PCIe device. A PF can be configured to behave independently. In some cases, a PF can be associated with a parent or controlling device (e.g., in a hardware virtualization environment). In some examples, a PCIe Virtual Function (VF) includes lightweight PCIe functionality on a network adapter that supports SR-IOV. A VF can share one or more physical resources of a device with a PF and / or other VFs on the device. For example, a VF can share one or more physical resources (e.g., memory, network ports, etc.) with a PF and other VFs on the device. Unlike a PF, a VF can only be configured to behave independently. A VF can be associated with a sub-partition in a virtualization environment. In some examples, SR-IOV can be used for networking latency-sensitive or CPU-intensive virtual machines. It can also enable the sharing of a GPU across multiple users or VMs, but provides a similar performance level to discrete processors.
[0076] In some examples, a device may provide one or more Functional Functions (PFs) to a host (e.g., PFs only). In some cases, one or more PFs may be assigned as parents (e.g., given a control characteristic), enabling this PF to manage other PFs. Alternatively, other PFs (e.g., non-parent PFs) may not be granted the privilege of managing themselves or other PFs. In some cases, virtual function and / or physical function namespaces may appear as separate SSDs to connected hosts.
[0077] Write amplification (WA) is a phenomenon that occurs when more data is written to the storage medium than expected. This can happen in flash memory and solid-state drives (SSDs). WA occurs when the host computer writes a logical amount of data that differs from the amount of physical data written. In other words, WA occurs when the actual amount of physical data written differs from the amount of logical data written by the host computer. WA can be caused by a disconnect between the device and the host. The host may not have sufficient information to understand the physical layout of the device or to know the data that is frequently used together. WA can negatively impact storage performance and durability, and can also shorten the lifespan of the device.
[0078] In some examples, the write amplification factor (WAF) is a multiplier applied to the data during a write operation. WAF is the factor by which written data is amplified. WAF is calculated by dividing the amount of data written to the flash media by the amount of data written by the host. An ideal SSD has a WAF of 1.0 (e.g., WAF=1). A WAF of 1 indicates no write amplification. SSDs can use wear leveling to distribute writes evenly across the drive, which can lead to write amplification. SSDs can use garbage collection to reclaim unused space, which can also result in write amplification.
[0079] Some methods of SSD data placement can lead to write amplification, which can be caused by storing different types of data (e.g., hot data, cold data) in the same NAND block (e.g., the same erase unit). For example, at time 0, block A includes pages a, b, c, d, and e:
[0080] Block A at time 0: [(a)(b)(c)(d)(e)]; a / c = hot data; b / d / e = cold data.
[0081] Pages a and c contain hot data, while pages b, d, and e contain cold data. As a result, the "hot" data on pages a and c is likely to be updated, while the "cold" data on pages b, d, and e is likely to remain unchanged for a given period of time.
[0082] At time 1, block A is selected as a garbage collection (GC) candidate. For example, at time 1, the versions of pages a and c in block A are data exhausted, and the data in pages a and c is invalidated due to updates (e.g., updates to pages a and c):
[0083] Block A at time 1: [(a (b)(c) (d)(e)];a / c= Invalid; b / d / e = cold data;
[0084] Therefore, based on the fact that pages a and c in block A are outdated data (e.g., invalid or stale data), block A is selected as a GC candidate. Because flash memory blocks cannot be updated in-place (e.g., updating pages a and c in block A), at time 1+ (e.g., some time after time 1), the updated data for pages a and c is written to another block (e.g., block C).
[0085] Block C at time 1+: [(a)(c)( )( )( )]; a new version of a / c.
[0086] However, pages b, d, and e of block A remain valid (e.g., fresh data). Therefore, at time 1+ (e.g., sometime after time 1, before or after writing a and c to block C), pages b, d, and e are written to a new block (e.g., written to block B, causing write amplification) before block A can be erased:
[0087] Block A at time 1+: [(a (b)(c) (d)(e)];a / c= Invalid; b / d / e are retained in block A;
[0088] Block B at time 1+: [(b)(d)(e)( )( )]; copied to b / d / e of block B.
[0089] Therefore, the data movement from block A to block B results in write amplification because pages b, d, and e are now written twice in two separate blocks of the NAND.
[0090] Flexible Data Placement (FDP) is a feature of the NVMe specification designed to improve performance by reducing write amplification. FDP reduces write amplification (WA) when multiple applications write, modify, and read data on the same device. FDP gives host servers more control over where data resides within the SSD. FDP achieves this by enabling the host to prompt the device when a write request occurs. For example, the host can provide prompts in a write command to indicate where to place data via a virtual handle or pointer. Use cases for FDP are similar to those for other NVMe features such as Streams and Partition Namespaces (ZNS). ZNS divides the logical address space into fixed-size regions. ZNS devices partition functionality between the device controller and host software. Streams can include or be associated with descriptors called Stream Granularity (SGS). SGS can be used in a manner similar to how RU sizes can be used. In some examples, stream numbers and RUH IDs are used interchangeably.
[0091] Durability groups can include groups of one or more reclaim groups. In some examples, durability groups can include separate storage pools for wear leveling purposes. Each durability group can have its own dedicated spare block pool, and drives report separate wear statistics for each durability group. In some examples, NVMe durability group management allows media to be configured as durability groups and NVM collections. Durability groups can enable granular access to SSDs.
[0092] In SSDs, a Reclaim Unit (RU) is a unit of NVM storage. Data from applications is written to an RU, which is also referred to as one or more blocks (e.g., 128 pages or 512KB blocks). The host system tells the SSD where to place the data. For example, data from an application can be written to an application-specific area of the SSD, written to a so-called reclaim unit. An RU can include one or more physical NAND blocks within a reclaim group. Note that without FDP, data from different applications is written across all blocks. A reclaim unit can correspond to a physical memory cell and / or a logical memory cell. The SSD can be allowed to select which RU is filled at any time, and the SSD can select which physical NAND blocks constitute each RU.
[0093] A Reclaim Unit Handle (RUH) can be a resource in an SSD that manages and buffers logical blocks for writing to RUs. Each RUH can identify a different RU used for writing user data and can select a new, unique RU after the current RU is filled. A Reclaim Group (RG) can be a group of two or more RUs. In some cases, a namespace can access one or more RUHs. In some implementations, placement identifiers can be used to indirectly identify RUHs. In NVMe technology, an SSD namespace can be a set of logical block addresses (LBAs) accessible to host software. In some cases, namespaces divide an NVMe SSD into logically separate and individually addressable storage spaces, where each namespace can have its own I / O queue.
[0094] Note that the driver can freely choose which RUs the RUH uses at any time. In some cases, the RUH can be considered a pointer to a specific RU at any given time. A RUH can point to a RU within an RG. In some cases, within an RG, the RUH can have a RU that it points to and is filled at a given time. RUs can be changed and selected by the driver at any time during RU filling, and the new RU selection can be used. In some cases, an RG can be considered a physical boundary. For example, a die can be an RG, each die can have one RG, all dies on a channel can have one RG, and so on. RUs can be of the same size. Several erase blocks (EBs) can be grouped together to form RUs. In some cases, the EBs supporting the RUs can change at any time. A RUH can be a pointer. A pointer can identify a RU within each RG. An RG / RUH pair can individually identify a RU currently being filled with data. Therefore, each tenant can have one RG, each tenant can have one RUH, and / or each tenant can have one RG / RUH pair. In some cases, the RUs for each tenant may not be configurable because the RUs may not be addressable by the host.
[0095] In some examples, a virtual machine (VM) can be a virtualization or emulation of a computer system. VMs can be based on a computer architecture and provide the functionality of a physical computer. Their implementation can involve dedicated hardware, software, or a combination of both. In some cases, VMs can differ and can be organized by their functionality, as shown here:
[0096] Virtual machines (VMs) can be software-based computers that act as physical computers. A VM can be referred to as a guest machine. VMs can be created by borrowing resources from a physical host computer or a remote server. One or more virtual "guest" machines run on a physical "host" machine.
[0097] A hypervisor (also known as a virtual machine monitor (VMM) or virtualizer) can comprise a class of computer software, firmware, and / or hardware that creates and runs virtual machines. The term hypervisor can be a variant of "supervisor," a term that can be used for the kernel of an operating system: a hypervisor can be thought of as a supervisor of a supervisor, where management is used as a stronger variant of supervision from a "supervisor." The computer on which the hypervisor runs one or more virtual machines can be called the host machine, and each virtual machine can be called a guest machine. The hypervisor provides a virtual operating platform to the guest operating system and manages the operation of the guest operating system. Unlike an emulator, the guest runs most instructions on native hardware. Multiple instances of various operating systems can share virtualized hardware resources: for example, Linux, Windows, and macOS instances can all run on a single physical x86 machine. This contrasts with operating system-level virtualization, where all instances (often called containers) must share a single kernel, although the guest operating systems can differ in user space, such as different Linux distributions with the same kernel.
[0098] In some cases, hypervisors allow a single host computer to support multiple virtual machines (VMs) by sharing resources such as storage and processing power. Hypervisors do this by allocating compute, storage, and networking resources from the host server according to the needs of each VM. In some cases, hypervisors virtualize the compute and hardware resources of computers and servers, enabling cloud computing. In others, hypervisors isolate the hypervisor operating system and resources from the VMs, making it possible to create and manage those VMs. The VMs may be unaware that their access to the hardware is virtualized, emulated, or protected from other users of the same hardware.
[0099] In some examples, multitenancy involves an architectural design that allows multiple users to access a single application or system. Multitenancy can include an architecture where a software application or system serves multiple tenants or customers on shared infrastructure (e.g., a multitenant architecture). In cloud computing, multitenancy can refer to allowing multiple customers to use one or more shared SSDs. Multitenancy can create isolated environments within a single physical infrastructure, such as virtual machines, servers, cloud platforms, etc. Instances (tenants) can be logically isolated but physically integrated (e.g., using the same storage devices, memory devices, and / or processing units, etc.).
[0100] In some examples, multitenancy can include isolated tenancy, where each tenant's data and compute resources can remain separate. In some cases, multitenancy can include shared tenancy, where all customer data can be stored on servers, storage devices, and / or databases that are still shared during hosting. Additionally or alternatively, multitenancy can include hybrid tenancy, which combines two or more multitenancy types. Tenant isolation can include persistent isolation (e.g., tenants with persistent isolation are permanently isolated from other tenants, or as long as the host or hypervisor specifies a tenant as persistently isolated) and initial isolation (e.g., tenants with initial isolation are temporarily isolated from at least one other tenant, such as when a new tenant is added, a tenant is initially configured, etc.).
[0101] In some cases, RocksDB is a high-performance embedded database for key-value data (e.g., a persistent key-value store for fast storage). RocksDB can be optimized to take advantage of multi-core processors and efficiently utilize fast storage, such as SSDs, for input / output (I / O)-constrained workloads. RocksDB can be based on a log structure merge tree (LSM tree) data structure. RocksDB and other LSM tree databases can be examples of system applications that can differentiate between levels of transformation from hot to cold data. In some cases, CacheIB can be a caching engine for web-scale services. CacheIB can include libraries that provide thread-safe application programming interfaces (APIs) for building caching services with high throughput and low overhead. CacheIB can allow services to customize and scale highly concurrent caches. These caches can be identified as having different data hotness levels. It may be desirable to place the caching layer on the drive to mirror the cache management structure in the host software.
[0102] In some examples, multitenancy can include shared software instances. In some cases, multitenant facilities (e.g., hosts, hypervisors, servers, shared computing resources, cloud resources, etc.) can store metadata about each tenant and use that metadata to modify the software instance at runtime to suit each tenant's needs. In some cases, tenants can be isolated from each other via licenses. Even if tenants can share the same software instance, each tenant can use and experience the software differently.
[0103] Multitenant clusters can be shared by multiple users and / or workloads (e.g., tenants). Operators of multitenant clusters can isolate tenants from each other to minimize the potential damage that compromised or malicious tenants might cause to the cluster and other tenants. Furthermore, cluster resources can be allocated fairly among tenants. Multitenant clusters can include several advantages over multiple single-tenant clusters, such as reduced administrative overhead, reduced resource fragmentation, and no need to wait for new tenants to create the cluster.
[0104] Multitenant architectures can be configured with one or more types of SSDs. Local SSDs include SSDs physically attached to the servers hosting VM instances. Local SSDs can provide high input / output operations per second (IOPS) and low latency. In some cases, local SSDs can be configured to provide temporary storage. Partitioned storage SSDs can improve the overall efficiency and utilization of the data center system, enabling multitenancy. Provisioning high-capacity SSDs among tenants can provide several benefits, such as improved performance and efficient resource utilization.
[0105] Tenant isolation in multi-tenant architecture. One of the biggest considerations in designing a multi-tenant architecture is the level of isolation required for each tenant. Isolation can mean different things: having a single shared infrastructure, having separate instances of applications, and having a separate database for each tenant.
[0106] Asynchronous Transfer Mode (ATM) can encompass telecommunications standards for digital transmission of various types of services. ATM is designed for integrated telecommunications networks. It can handle both traditional high-throughput data services and real-time, low-latency content such as telephone (voice) and video. ATM can provide functionality that utilizes the features of both circuit-switched and packet-switched networks by using asynchronous time-division multiplexing.
[0107] In the Open Systems Interconnection (OSI) reference model, the basic transmission unit can be called a frame. In ATM, these frames can be of a fixed length (e.g., 53 octets) called cells. ATM can use a connection-oriented model, where a virtual circuit must be established between two endpoints before data exchange begins. These virtual circuits can be permanent (e.g., a dedicated connection pre-configured by a service provider) or switched (e.g., established on a per-call basis using signaling and disconnected when the call can be terminated).
[0108] The ATM network reference model can be approximately mapped to the three lowest layers of the OSI model: the physical layer, the data link layer, and the network layer. ATM can include core protocols used in the backbone of the public switched telephone network and the Synchronous Optical Network and Synchronous Digital Architecture (SDA) of the Integrated Services Digital Network (ISDN) (e.g., Synchronous Optical Network (SONET), Synchronous Digital Architecture (SDH)).
[0109] In computer networking, network services can include applications running at or above the network application layer, which provide data storage, manipulation, presentation, communication, or other capabilities. These services can typically be implemented using a client-server or peer-to-peer architecture based on application layer network protocols.
[0110] In some cases, each service may be provided by a server component running on one or more computers (typically a dedicated server computer that provides multiple services) and accessed over a network by a client component running on other devices. However, both the client and server components may run on the same machine. Both the client and server will typically have a user interface and sometimes other hardware associated with them.
[0111] In computer network programming, the application layer may include an abstraction layer reserved for communication protocols and methods designed for process-to-process communication across IP networks. Application layer protocols use underlying transport layer protocols to establish host-to-host connections for network services.
[0112] When a network service (e.g., an application) uses a broadband network (e.g., an ATM network) to deliver services, the network service can inform the network about the type of service to be transmitted and / or the performance requirements of that service. The application provides this information to the network in the form of a service contract. In addition to determining the service type based on service rate, the ATM Forum can also define QoS and service parameters to measure ATM service quality. QoS and service parameters can be negotiated between the ATM user and the ATM network or between two ATM networks before an ATM connection is established. These negotiated parameters can form a service contract.
[0113] When an application can request a connection, it can indicate to the network the requested type of service (e.g., service level, relatively high data rate service, relatively low data rate service, etc.), parameters for each data stream in both directions, and / or the requested quality of service (QoS) parameters for each direction. These parameters can form at least a portion of a communication descriptor (e.g., a service descriptor, an access descriptor) used for the connection.
[0114] These service categories provide a method for associating service characteristics and QoS requirements with network behavior. Service categories can be characterized as real-time or non-real-time. Real-time service categories can include at least one of constant bit rate (CBR) and / or real-time variable bit rate (rt-VBR). Non-real-time service categories can include unspecified bit rate (UBR), adaptive bit rate (ABR), and / or non-real-time variable bit rate (nrt-VBR).
[0115] In communications, traffic policing can include the process of monitoring network traffic to ensure compliance with traffic contracts and taking steps to enforce those contracts. For example, NAND flash memory may have available bandwidth or access capabilities. Therefore, an SSD controller (e.g., firmware and / or hardware) can implement a traffic policing scheme, allowing the SSD to maintain traffic contracts. Traffic sources aware of the traffic contracts can apply traffic shaping to ensure their output remains within the contract, thus preventing it from being dropped. Depending on the management policy and the characteristics of the excess traffic, traffic exceeding the traffic contract can be immediately dropped, throttled, blocked and later released, marked as incompatible, or left as is. When incoming traffic exceeds the contract, receivers of policing traffic will observe packet loss distributed throughout the time period. In some cases, receivers may observe reduced performance (e.g., throttled bandwidth, delayed I / O completion, and / or reduced / slowed completion rates). If the source does not limit its sending rate (e.g., through a feedback mechanism), this will continue, and to the receiver it may appear as if a link error or some other interruption is causing random packet loss (e.g., as if the SSD is performing poorly). In some examples, the SSD can be tested in synthetic tests. For example, a host might maintain a constant queue of Y commands of size X (e.g., a feedback mechanism on a host to keep the SSD load constant during testing). If a censorship event occurs, the SSD's performance may be lower during the event (e.g., lower bandwidth and / or longer latency per command during runtime). This throttling can stress the host because the SSD slows down or stops retrieving commands from the submission queue, and the host doesn't need to replenish the SQ as much. While components downstream of the censor in the network may introduce jitter, received traffic that has already undergone censorship en route will generally be contractually compliant. Utilizing reliable protocols, such as TCP instead of UDP, dropped packets will not be acknowledged by the receiver and will therefore be retransmitted by the transmitter, generating more traffic. In some cases, the SSD may execute slower because it cannot drop commands. In real-world applications, rather than synthetic applications, the host can submit the same amount of work to the SSD, but the work may take longer depending on the load.
[0116] Service policing in ATM networks can be referred to as usage / network parameter control. The network can also drop non-compliant services (e.g., using priority control). References to both service policing and service shaping in ATM (given by the ATM Forum and ITU-T) can include the Common Cell Rate Algorithm (GCRA). GCRA can include scheduling algorithms used in ATM networks. In some cases, GCRA measures cell timing on Virtual Channels (VCs) and Virtual Paths (VPs) against bandwidth and jitter limits. GCRA can include a leaky bucket algorithm version. In some cases, GCRA works by dividing time into cells of a certain size (e.g., cells per second). When a request is made, a new cell can be created. When the number of requests in a cell exceeds the limit, all subsequent requests can be blocked until the next cell. For example, when the rate limit is 10,000 requests / hour, GCRA can ensure that users are not allowed to make all 10,000 requests within a relatively short time period.
[0117] A leaky bucket algorithm is an algorithm based on the analogy of how a bucket with a constant leakage rate will overflow if the average rate at which water is poured in exceeds the rate at which the bucket leaks, or if more water than the bucket's capacity is poured in at once. The leaky bucket algorithm can be used to determine whether a sequence of discrete events conforms to defined limits on its average and peak rates or frequency; for example, to restrict actions associated with these events to these rates or to delay them until they do conform. The leaky bucket algorithm can be used to check conformance individually or to limit to an average rate (e.g., to remove any variation from the average). In packet-switched computer networks and telecommunications networks, the leaky bucket algorithm can be used for traffic policing, traffic shaping, and / or data transmission scheduling in the form of packets to define limits on bandwidth and burstiness (e.g., a measure of variation in traffic flow).
[0118] One version of the leaky bucket algorithm, the Universal Cell Rate algorithm, can be used with ATM networks in UPCs and NPCs at user network interfaces, inter-network interfaces, or network-to-network interfaces to protect the network from excessive traffic levels on connections routed through it. The Universal Cell Rate algorithm, or an equivalent algorithm, can also be used by network interface cards to shape transmissions onto ATM networks.
[0119] The token bucket algorithm can be compared to one of two versions of the leaky bucket algorithm. This comparable version of the leaky bucket algorithm can be described on the relevant Wikipedia page as a leaky bucket algorithm as a meter. This can be a mirror image of the token bucket algorithm, where consistent groups add fluid (equivalent to the tokens removed by consistent groups in the token bucket algorithm) to a finite-capacity bucket, and then the fluid is drained from the finite-capacity bucket at a constant rate, equivalent to adding tokens at a fixed rate.
[0120] In some cases, GCRA, leaky bucket, and token bucket algorithms are reference algorithms. They can be used in network standards because they are well-known and well-described. However, in implementations, other tokenization, tracking, or arbitration schemes can be implemented to mimic these basic algorithms. Other implementations can achieve hardware reduction, increased responsiveness, additional features beyond the basic characteristics, and / or other advantages in driver implementations. In some cases, delivering host requests within the framework of GCRA understanding has become an industry standard. The transformation of these parameters can be used for tokenization schemes implemented in drivers (e.g., SSDs). For example, communication setups for GCRA can implement methods for associating service characteristics and QoS requirements with network behavior, where service categories can be characterized as real-time or non-real-time (e.g., CBR, rt-VBR, UBR, ABR, nrt-VBR, etc.).
[0121] In some examples, statistical multiplexing is based on techniques that require dynamic allocation of time slots. Statistical multiplexing can be a core technology behind the ATM-based Broadband Integrated Services Digital Network (SSD) concept. In communication networks, statistical multiplexing can be performed by a switching system that combines data packets from multiple input lines and forwards them to multiple outputs. In some cases, statistical multiplexing models can involve bursty and correlated input processes. A burst can mean that several cells arrive within a given time period (e.g., relatively simultaneously). In the context of an SSD, a burst can mean that a group of commands is submitted to the SQ within a given time period. These commands can be requests to perform a certain amount of work on the SSD (read, write, deallocate, other commands). The length of the command or the number of LBAs affected by the command can vary. Therefore, a burst can mean a large group of commands submitted together, a small number of commands with a large number of affected LBAs, or any combination thereof.
[0122] Data packets may include a Cyclic Redundancy Check (CRC) field. In some examples, a CRC field is a mathematical technique for detecting errors in transmitted data. CRC fields can be used in digital networks and storage devices to prevent common types of errors on communication channels.
[0123] Data packets may include a Frame Check Sequence (FCS) field. The FCS field may include a 2-byte or 4-byte field used to detect errors in the frame during transmission across the network. The FCS can be added to the frame before transmission, and the destination can calculate the new FCS code. The destination can then compare the calculated code with the FCS bits of the received frame. When the FCS matches, the transmission can be considered successful. When the FCS does not match, the frame can be discarded, and a retransmission of the frame can be requested.
[0124] GCRA can be likened to two leaky buckets occurring simultaneously. One can be the bandwidth limit and tokenization, and the other can be the number of limits and tokenizations (input / output operations). In this way, read commands can have one GCRA applied to them, and the driver can then satisfy independent bandwidth and IOPS targets or a combination of both when command sizes are mixed. Similarly, other GCRAs can be implemented simultaneously on other command types (e.g., write, deallocation, etc.). This allows the host to specify the behavior of each GCRA for each command type, and the driver can implement the limits for each command independently. One GCRA can be used for read IOPS and bandwidth, one GCRA can be used for write IOPS and bandwidth, and one GCRA can be used for other commands.
[0125] In some approaches, resource contention may exist for physical resources, such as which commit queue (SQ) to fetch from each command metadata space and how many commands to fetch from each SQ at a time, direct memory access (DMA) engines that can be used to transfer data to host memory or driver memory, access to the driver's dynamic random access memory (DRAM), use of the central processing unit (CPU) within the driver, buffer allocation (for reads, cached reads, prefetched reads, incoming writes, and / or other buffer usage), error correction decoding units, memory dies, memory channels, etc. This resource contention can be caused by operations from programming (e.g., incoming programmed writes, programmed garbage collection, etc.), erase commands, and read commands (e.g., incoming reads, garbage collection reads). In some approaches, the constrained resources can vary with the workload level and drive capability.
[0126] In some of the systems and methods described herein, access arbitration schemes can be implemented for at least a portion of a contested resource. For example, an arbitration scheme granting 30% access to the first tenant and 70% access to the second tenant can be applied to SQ access and / or CPU access. At the SQ level, allocation can be based on the source tenant (e.g., the source tenant's read request, the source tenant's programming request). However, at other resources, the 70 / 30 arbitration split can be applied to different metrics (e.g., die access time). For example, in some cases, programming may be slower than reading. In some examples, SQ entries (SQEs) can appear identical until the SSD brings the SQEs into the drive and resolves them. In some cases, shared arbitration can exist, where tenants blindly acquire 30 or 70. However, there is a possibility that drive SQs can be identified differently, such that an SQ with only read access can be identified differently from an SQ with only write access. If such a separation of SQ types occurs, there can be arbitration for distinctions such as (tenant_1, SQ_rd), (tenant_1, SQ_wr), (tenant_2, SQ_rd), (tenant_2, SQ_wr), etc. Because programming can be slow, when die programming and read access are arbitrated relative to die time (30% read and 70% program), less programming can be completed than expected compared to reads. Additionally or alternatively, because the amount of data being programmed can be more parallel in some memories, the total number of data sectors programmed can exceed the number of individual reads that can be completed in a similar amount of time for such memory.
[0127] In some cases, write amplification or other costs associated with a tenant's access can be attributed to that tenant's resource access allocation. For example, random writes by a tenant can increase that tenant's write amplification. Therefore, other potential activities required for garbage collection to read, write, erase, and perform the requested tenant's activities may stem from that tenant's die allocation time.
[0128] In some implementations, the host can provide physical address access recommendations. Examples of physical address access recommendations include durability groups, NVM sets, reclamation groups, Flexible Data Placement (FDP) reclamation groups, FDP reclamation unit handles, streams, etc. In some examples, at least some reclamation unit handles (RUHs) of the initial isolation type can have a common WAF (or other cost) that determines the frequency of garbage collection. In some examples, one or more (e.g., each) persistent isolation RUHs can have their own individually determined WAF (or other cost). This WAF (or other cost) can be used to apply the consistent arbitration strategy described herein. When calculating the die access time of tenants in initial isolation, it is possible that the WAF of all tenants in initial isolation can be a weighting factor. Thus, the entire group of tenants in initial isolation experiences the same weighting factor in their die access time. Conversely, the driver can track the WAF of persistent isolation RUHs individually. Thus, in the case where a tenant is identified and uses a single persistent isolation RUH, the tenant can be individually weighted by its own WAF. If the WAF of that persistent isolation RUH is relatively high, die access may be blocked in frequency or arbitration may be reduced. Additionally or alternatively, if the WAF of the persistently isolated RUH is relatively low (e.g., very low), die access can be significantly improved for the purpose of programming new data into NAND. The additional reads and writes required to run the garbage collection needed for the initial isolated or persistently isolated RUH's WAF can be allocated within the time of each RUH group provided in each persistently isolated RUH or initial isolated RUH group. RUH is an example, but other identifiers can be considered and / or implemented.
[0129] In some examples, to ensure fairness in access to each tenant's namespace, the system and methods can infer physical information from the RUHs assigned to the namespace. Therefore, arbitration can focus on the RUH or other physical address identifiers previously named in paragraph 102. In some cases, multiple namespaces may exist, and at least some namespaces may be assigned different RUHs and / or provided with the same or similar capacities (e.g., other physical information). Therefore, namespaces can be arbitrated equally or more equally. Similarly, namespaces with RUHs associated with different capacity levels, overprovisioning, RG utilization, and / or different RUH configurations (e.g., shared RUHs) can be treated differently in arbitration.
[0130] The disclosed systems and methods can provide multi-tenant access to resources. In some embodiments, multi-tenancy can refer to more than one entity (e.g., user, application, host, etc.) accessing a shared resource (e.g., storage, software, etc.). In a particular disclosed example, a device (e.g., a solid-state storage device) can be shared by more than one virtual machine (VM) on a host.
[0131] Systems and methods may include a host signaling a device to enable or disable a control mechanism on a specific tenant (e.g., a VM). In some cases, the control mechanism (e.g., a multi-tenancy controller) can adjust (e.g., throttle) the resource consumption of one or more tenants. In some situations, one aspect of this disclosure may involve a device signaling a host to notify that a throttling mechanism has been successfully enabled or disabled. Examples of tenants include VMs, physical functions of storage devices, virtual functions of storage devices, commit queues, namespaces, logical block address (LBA) ranges, etc.
[0132] In some examples, storage devices (e.g., SSDs) may encounter resource contention that may or may not be subject to arbitration (e.g., depending on tenant workload, number of tenants, etc.). One or more hosts (e.g., including hypervisors, VMMs, etc.) may be associated with the storage device. In some cases, each host may include one or more tenants. In one or more examples, the storage device and / or one or more hosts associated with the storage device may include a multi-tenancy controller. In some examples, the control mechanisms of the multi-tenancy controller may be configured to adjust one or more parameters associated with multi-tenancy of a specific storage device. Parameters may include at least one of the following: per-tenant input / output (e.g., maximum) operations per second (IOPS), per-tenant bandwidth (e.g., maximum), reserved (protected service level) IOPS, reserved bandwidth per tenant, per-tenant quality of service (QoS) (e.g., per-tenant performance variation), tenant priority based on the priorities of two or more tenants, arbitration weighting for disputed resources between tenants, scalability to different drives, workload and resource contention variations, write amplification factor (WAF) (e.g., performance adjusted based on WAF), and / or internal device traffic. In some cases, tenants can reside within the device, and arbitration policies can be applied to these tenants. In some cases, control mechanisms can be applied to read traffic, write traffic, garbage collection, NAND flash operations, vendor-specific media operations, deallocation operations, replication operations, and / or other types of operations. In some cases, constrained and / or excessive resources may vary with workload and drive capacity. In some cases, when the multi-tenant controller determines that a resource may be contentious, it can apply an arbitration scheme to that resource. The arbitration scheme implemented by the multi-tenant controller can be configured to conform to one or more standards (e.g., NVMe, PCIe, networking standard principals, ATM, OSI model, etc.).
[0133] Figure 1 An example system 100 according to one or more embodiments as described herein is illustrated. Figure 1The image shows machine 105, which can be referred to as a host, system, or server. Although Figure 1 Machine 105 is described as a tower computer, but embodiments of this disclosure can be extended to machines of any form factor or type. For example, machine 105 may be a rack server, blade server, desktop computer, tower computer, mini-tower computer, desktop server, laptop computer, notebook computer, tablet computer, etc.
[0134] Machine 105 may include processor 110, memory 115, and storage device 120. Processor 110 may be any type of processor. Note that, for ease of illustration, processor 110 and other components discussed below are shown external to the machine: embodiments of this disclosure may include these components internally. Although Figure 1 A single processor 110 is shown, but the machine 105 may include any number of processors, each of which may be a single-core or multi-core processor, each of which may implement a Reduced Instruction Set Computer (RISC) architecture or a Complex Instruction Set Computer (CISC) architecture (and other possibilities), and may be mixed in any desired combination.
[0135] Processor 110 may be coupled to memory 115. Memory 115 may be any type of memory, such as flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), persistent random access memory, ferroelectric random access memory (FRAM), or non-volatile random access memory (NVRAM) (such as magnetoresistive random access memory (MRAM), phase-change memory (PCM), or resistive random access memory (ReRAM)). Memory 115 may include volatile and / or non-volatile memory. Memory 115 may use any desired form factor: for example, single in-line memory module (SIMM), dual in-line memory module (DIMM), non-volatile DIMM (NVDIMM), etc. Memory 115 may be any desired combination of different memory types and may be managed by memory controller 125. Memory 115 may be used to store data that can be referred to as "short-term": that is, data that is not expected to be stored for an extended period of time. Examples of short-term data may include temporary files, data used locally by the application (which may have been copied from other storage locations), etc.
[0136] Processor 110 and memory 115 can support an operating system under which various applications can run. These applications can issue requests (which may be referred to as commands) to read data from memory 115 or storage device 120 or to write data to memory 115 or storage device 120. When storage device 120 is used to support applications that read or write data via a certain file system, device driver 130 can be used to access storage device 120. Although Figure 1 A storage device 120 is shown, but any number (one or more) of storage devices can be present in machine 105. Storage device 120 can support any desired protocol or protocol, including, for example, the Non-Volatile Memory Fast (NVMe) protocol, the Serial Attached Small Computer System Interface (SCSI) (SAS) protocol, or the Serial AT Attach (SATA) protocol. Storage device 120 can include any desired interface, including, for example, a Peripheral Component Interconnect Fast (PCIe) interface or a Compute Fast Link (CXL) interface. Storage device 120 can adopt any desired form factor, including, for example, U.2 form factor, U.3 form factor, M.2 form factor, Enterprise and Data Center Standard Form Factor (EDSFF) (including all its types, such as E1 Short, E1 Long, and E3 types), or Add-in Card (AIC).
[0137] Although Figure 1 The term "storage device" is used, but embodiments of this disclosure may include any storage device format that can benefit from the use of computing storage units, examples of which may include hard disk drives, solid-state drives (SSDs), or persistent memory devices such as PCM, ReRAM, or MRAM. Any references to "storage device" or "SSD" below should be understood to include other embodiments of this disclosure and other types of storage devices (e.g., hard disk drives, etc.). In some cases, the term "storage unit" may encompass both storage device 120 and memory 115.
[0138] Machine 105 may include power supply 135. Power supply 135 provides power to machine 105 and its components. Power supply 135 may have a maximum amount of power that can be used (before exceeding the specifications of power supply 135): this information may be known to machine 105 and may be used, for example, by multi-tenancy controller 140 to determine whether tenants and / or storage devices (e.g., storage device 120) are operating within power constraints. The operating level of power supply 135 may be adjusted (e.g., increasing voltage, decreasing voltage, increasing current, decreasing current, etc.) based on the systems and methods described herein.
[0139] Machine 105 may include a transmitter 145 and a receiver 150. The transmitter 145 or receiver 150 may be used to send or receive data (e.g., between a host and a storage device, between a storage device and one or more tenants, etc.). In some cases, the transmitter 145 and / or receiver 150 may be used to communicate with memory 115 and / or storage device 120. The transmitter 145 may include write circuitry 160, which may be used to write data to registers, memory 115, and / or storage devices. Similarly, the receiver 150 may include read circuitry 165, which may be used to read data from memory 115 and / or storage device 120 from storage such as registers.
[0140] In one or more examples, machine 105 can be implemented using any type of device. Machine 105 can be configured as one or more servers (e.g., a host) such as a computing server, storage server, storage node, network server, supercomputer, data center system, etc., or any combination thereof. Additionally or alternatively, machine 105 can be configured as one or more computers (e.g., a host) such as a workstation, personal computer, tablet computer, smartphone, and / or the like, or any combination thereof. Machine 105 can be implemented using any type of device, which can be configured to include, for example, accelerator devices, storage devices, network devices, memory expansion and / or buffering devices, central processing unit (CPU), graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), etc., or any combination thereof.
[0141] Any communication between devices including machine 105 (e.g., host, compute storage device, and / or any intermediate device) can occur through an interface that can be implemented using any type of wired and / or wireless communication medium, interface, protocol, etc., including PCIe, NVMe, Ethernet, NVMe-oF, Compute Fast Link (CXL) and / or coherent protocols (such as CXL.mem, CXL.cache, CXL.IO, etc.), Gen-Z, Open Coherent Accelerator Processor Interface (OpenCAPI), Cache Coherent Interconnect for Accelerators (CCIX), Advanced Extensible Interface (AXI), etc., or any combination thereof, Transmission Control Protocol / Internet Protocol (TCP / IP), Fibre Channel, InfiniBand, Serial AT Attachment (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, any generation of wireless network (including 2G, 3G, 4G, 5G, etc.), any generation of Wi-Fi, Bluetooth, Near Field Communication (NFC), etc., or any combination thereof. In some embodiments, the communication interface may include a communication structure including one or more links, buses, switches, hubs, nodes, routers, converters, repeaters, and / or the like. In some embodiments, system 100 may include one or more additional devices having one or more additional communication interfaces.
[0142] Any function described herein (including any function of host functions, device functions, multi-lease controller 140 functions, etc.) may be implemented in hardware, software, firmware, or any combination thereof, including, for example, hardware and / or software combinational logic, sequential logic, timers, counters, registers, state machines, volatile memory (such as dynamic random access memory (DRAM) and / or static random access memory (SRAM)), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory, memory with varying volume resistance, phase-change memory (PCM), etc.) and / or any combination thereof, complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), application-specific integrated circuit (ASIC) CPUs (including complex instruction set computer (CISC) processors (such as x86 processors) and / or reduced instruction set computer (RISC) processors (such as RISC-V and / or ARM processors), graphics processing units (GPUs), neural processing units (NPUs), tensor processing units (TPUs), etc.), which run instructions stored in any type of memory. In some embodiments, one or more components of the multi-lease controller 140 may be implemented as a system-on-a-chip (SOC).
[0143] In some examples, the multi-tenant controller 140 may include any one or a combination of logic (e.g., logic circuitry), hardware (e.g., processing units, memory, storage), software, firmware, etc. In some cases, the multi-tenant controller 140 may be combined with processor 110 to perform one or more functions. In some cases, at least a portion of the multi-tenant controller 140 may be implemented in or by processor 110 and / or memory 115. One or more logic circuitry of the multi-tenant controller 140 may include any one or a combination of multiplexers, registers, logic gates, arithmetic logic units (ALUs), caches, computer memory, microprocessors, processing units (CPUs, GPUs, NPUs, and / or TPUs), FPGAs, ASICs, etc., enabling the multi-tenant controller 140 to provide a multi-tenant SSD configuration. In some cases, the multi-tenant controller 140 may include a hypervisor, virtual machine monitor (VMM), or virtualizer configured to perform one or more of the techniques described herein.
[0144] In one or more examples, the multi-tenant controller 140 can identify the identifier of the first tenant of the storage device and assign a first performance level to the first tenant. In some cases, the multi-tenant controller 140 can generate first performance parameters based on the first performance level and send a configuration message to the storage device including the first performance parameters and the identifier of the first tenant. Based on the systems and methods described herein, the multi-tenant controller 140 improves the quality of service for host tenants, thereby providing improved service consistency within defined service variation limits. Based on the systems and methods described herein, the multi-tenant controller 140 reduces SSD latency, thereby improving SSD access among multiple tenants. Therefore, the multi-tenant controller 140 achieves balanced and customized SSD operations, thereby increasing SSD lifetime and reliability. Based on the systems and methods described herein, the multi-tenant controller 140 provides improved application performance and infrastructure efficiency, which improves overall data center system efficiency and utilization. Based on the systems and methods described herein, the multi-tenant controller 140 provides improved control and customization of SSD access, enabling SSDs to adapt to varying tenant loads among multiple tenants. Therefore, the multi-tenancy controller 140 enables the host to specify or request relative settings for various tenants with increased specificity.
[0145] Figure 2 An example based on the description herein is shown. Figure 1Details of machine 105 are provided. In the illustrated example, machine 105 may include one or more processors 110, which may include a memory controller 125 and a clock 205 for coordinating the operation of the machine's components. Processor 110 may be coupled to memory 115, which, as an example, may include random access memory (RAM), read-only memory (ROM), or other state-saving media. Processor 110 may be coupled to storage device 120 and to network connector 210, which may be, for example, an Ethernet connector or a wireless connector. Processor 110 may be connected to bus 215, and user interface 220 and input / output (I / O) interface ports that can be managed using I / O engine 225, as well as other components, may be attached to bus 215. As shown, processor 110 may be coupled to multi-lease controller 230, which may be... Figure 1 Example of a multi-lease controller 140. Alternatively or additionally, processor 110 may be connected to bus 215, and multi-lease controller 230 may be attached to bus 215.
[0146] Figure 3 An example system 300 according to one or more embodiments as described herein is illustrated. In the illustrated example, system 300 depicts aspects of a storage device (e.g., storage device 120, SSD, etc.). As shown, system 300 includes an endurance group 305. In some examples, system 300 indicates components that can be used by a host to assign NAND die time allocations to one or more tenants.
[0147] As shown in the figure, durability group 305 may include one or more recycling unit handles (RUHs) and / or one or more recycling groups (RGs). In the illustrated example, durability group 305 includes RUH 310, RUH 315, RUH 320 (e.g., up to RUH 320), RG 325, RG 330, RG 335, and RG 330. Each RG of durability group 305 may include one or more recycling units (RUs). In the example shown, RG 325 includes RU 345a, RU 345b, 345c, and RU 345d (e.g., up to RU 345d); RG 330 includes RU 350a, RU 350b, 350c, and RU 350d (e.g., up to RU 350d); RG 335 includes RU 355a, RU 355b, RU 355c, and RU 355d (e.g., up to RU 355d); and RG 340 includes RU 360a, RU 360b, RU 360c, and RU 360d (e.g., up to RU 360d). In the example shown, there can be up to M RUHs, where M is a positive integer, and RUH 320 is the Mth RUH. There can be N RGs, where N is a positive integer, and RG 340 is the Nth RG. There can be P RUs in RG 325, where P is a positive integer, and RU 345d is the Pth RU. Similarly, there can be two or more RUs in RG 330, RG335 and / or RG 340.
[0148] In some examples, the systems and methods described herein may include aspects for reducing write amplification (e.g., based on Flexible Data Placement (FDP)). In some examples, persistently isolated RUH tenants may have a write amplification factor (WAF) or other relevant parameters for tracking NAND utilization calculated for each tenant (e.g., over-provisioning (OP), RU fill / sparseness, etc.). A tenant may be defined by an interface identifier and a RUH. However, in some cases, a tenant may use more than one RUH, or several tenants may share a single RUH. In some cases, parameters for tracking NAND utilization may be calculated for each RUH (e.g., based on a tenant-to-RUH relationship when each RUH corresponds to a single tenant). In some cases, a RUH may track NAND utilization and / or parameters used for tracking NAND utilization.
[0149] In some cases, the WAF of a given tenant can affect the level of service provided to that tenant. In some examples, a Persistently Isolated (PI) RUH can correspond to a given tenant. PI_RUH can mean that GC is isolated from and triggered separately from other tenants. In some cases, the controller (e.g., multi-tenant controller 140) can estimate the WAF of the RUH. Assuming random write operations, the number of blocks garbage collected (GC'd) for a tenant can have a valid count of those currently being GC'd. This valid count can be used in a lookup table to correlate the valid count with the WAF and OP. Additionally or alternatively, the controller can have running measurements of incoming writes and GC writes to formulate the WAF estimate. In some cases, a forgetting factor or low-pass filter is used for this estimate. The WAF estimate can be applied as a modifier of the arbitration strength of the tenant and / or RUH. Thus, a tenant with sequential performance of WAF=1 can be unmodified. A tenant with WAF=4 can be modified by one-quarter. The arbitration strength of these two example tenants can be based on the modified (initial arbitration strength value). (WAF modifier) varies. Alternatively, tokens can be assigned to two tenants based on available die time. Tokens for each tenant can be assigned based on the host's assignment value or RUH_ID. Per-tenant-specific GC can attribute all GC reads and GC writes to the tenant and can throttle GC traffic based on the tenant's available token count. The WAF modifier can be based on 1 / WAF, 1 / (alpha) WAF) and / or (alpha) WAF).
[0150] In some examples, a host can leverage the persistent isolation and initial isolation status of RUH tenants to penalize tenants with non-compliant WAFs (e.g., tenants with WAF greater than 1, such as tenants with WAF=4). The host can establish tenant-RUH relationships. As an example, an SSD can have two persistently isolated (PI) and six initially isolated (II) RUHs. In some examples, a host can have ten tenants. The host can determine which tenant(s) are placed on the two PI RUHs and which tenant(s) are placed on each of the six II RUHs. The host can place one or more tenants on each of the RUHs. Based on a given tenant's WAF, systems and methods can include using the WAF to penalize tenants (e.g., based on WAF exceeding a threshold) and / or reward tenants (e.g., based on keeping the WAF below a threshold). Tenants writing to a circular FIFO buffer can cause WAF=1 traffic to be considered for GC, as the FIFO buffer can coincidentally be close to empty on the RU. Therefore, systems and methods can include rewarding tenants based on protecting their data from premature garbage collection (GC). In some examples, tenants can be on the host. For example, a tenant might be performing write operations. The FIFO might fill due to incoming traffic. The SSD can place writes in RUs. When the RUs filled inside the drive reach full capacity, the SSD can select another RU. Then, new RUs continue to be filled. This can be considered the leading edge of the circular FIFO. Meanwhile, the trailing edge of the circular FIFO might become invalid. This is because host tenants might be writing the same LBAs. Therefore, each new LBA entering the end of a RU can invalidate an LBA stored in-place in an older RU. As the older RUs continue to decrease their valid count, they may eventually become invalid due to incoming traffic. However, race conditions can exist. In some cases, an older RU might be about to become completely invalid. But at that same moment, the SSD can trigger GC to free up space. Even if the circular FIFO is WAF=1, the valid count of an older RU might be relatively small and it can be selected for GC. In some cases, a persistently isolated RUH can prevent this race condition (e.g., but not an initially isolated RUH). When selecting an RU for GC, the II RUH can consider all older RUs in the SSD. Therefore, another tenant may be causing GC, but the WAF=1 service is still selected for GC.
[0151] In some examples, an initially isolated RUH tenant group may have a WAF spread across tenants in that group. In some cases, initially isolated (II) RUHs may be associated. SSDs may track II RUHs together. In some examples, RGs and / or durability groups (e.g., protected durability groups) may be associated with one or more tenants. In some cases, an RG may be isolated, but it may be permitted to violate isolation under certain conditions (e.g., the drive experiences error / abnormal conditions, etc.). The NVMe standard defining RGs allows isolation boundaries to be violated by drives under certain conditions. SSDs may still store data, but SSDs may be given the freedom to violate RGs under certain conditions. A tenant routed to several RGs (e.g., to RG 330, RG 335, etc.) may enable that tenant to balance activity across different RGs based on performance characteristics. In some cases, a tenant per RG / RUH (e.g., a tenant isolated to RG 330 and RU 350b) may utilize the RUH count per RG to enable partial isolation, thus providing a measurement of partial isolation. For example, a tenant of RG 325 can be isolated from at least one RUH of RG 325 (e.g., RUH315, RUH 320, etc.). Additionally or alternatively, a tenant of RG 325 can share access to at least one RUH of RG 325 (e.g., isolated on RU 345a, and shared access to RU 345b, etc.). In some cases, a tenant can be routable to a subset of NAND dies. In some examples, performance levels can be provided to a tenant based on performance levels that the tenant has already subscribed to or requested. Any number of performance level tiers can exist (e.g., two or more tiers). In some examples, a first persistently isolated component of system 300 (e.g., persistently isolated RUH, RG, NAND die, etc.) can be assigned to a Tier 1 tenant (e.g., the highest priority tenant); two Tier 2 tenants (e.g., lower service levels) can be assigned to share a second persistently isolated component of system 300; four Tier 3 tenants can be assigned to share a third persistently isolated component of system 300; ten Tier 4 tenants can be assigned to share a fourth persistently isolated component of system 300; and fifteen Tier 5 tenants (e.g., the lowest priority tenants) can be assigned to share a fifth persistently isolated component of system 300. Note that more tenants than RUH can exist. As indicated, RUH provides isolation (e.g., persistent isolation, initial isolation). Therefore, assigning, for example, two tenants to RUH can provide that the two tenants can influence each other, but are not affected by other tenants besides these two. While not completely isolated for a given tenant, it is better isolation than being affected by every single tenant.Although this example involves persistent isolation components, the same arrangement can be applied to components with initial isolation, a configuration of one tenant to a durability group, a configuration of two tenants to a second durability group, one tenant to a die, two tenants to a second die, one tenant to a namespace, etc. In some examples, persistent isolation and initial isolation can be attributes of any combination that can be applied to NAND resource arbitration elements. Examples of NAND resource arbitration elements can include RUH RGs, durability groups, namespaces, etc. In some examples, tenant groups can be associated with one or more RGs (e.g., RG 325 to RG 340). In some cases, high-performance RG tenants get more die access in their respective RGs compared to low-performance RG tenants. In some aspects, the SSD can describe the RG and / or RUH configuration. The host can set tenant relationships (e.g., based on RG and / or RUH configurations). The host can send settings to the SSD based on the arbitration target that the host is requesting from the SSD. For example, the first tenant can be configured to win 70% of the time, while the second tenant can be configured to win 30% of the time. The SSD can be responsible for implementing the arbitration behavior of this request to the best of its capabilities. In some cases, tenants concerned with low latency (e.g., with relatively small I / O) can be assigned dedicated channels (channels to the die of System 300 and / or at least one component) and / or dedicated dies not shared with other tenants. In some aspects, an RG can be a physical group of NAND. The definition of an RG can be defined by the SSD and transferred by the SSD (e.g., to the host, to the customer / tenant). For example, a customer can request that each RG corresponds to a die. A RUH_ID can point to an active filled RU within each RG. In some systems, when the host isolates tenants per RG, the host can fill the RG value on the host before sending a write command to the SSD. The RG value can be a field in the write command. However, based on the systems and methods described herein, the host is able to indicate "VF_1 will go into RG_1". Therefore, the host can define a tenant as a combination of VF, RG, and / or RUH IDs (e.g., three IDs per tenant). In some cases, at least one tenant ID can be used to arbitrate access. In some cases, the host can fill the RUH value within the write command. Therefore, assignment or filling is provided based on the systems and methods described herein. In some examples, cross-channel striped RGs can be assigned to tenants focused on high bandwidth. In other examples, high-performance tenants can be assigned RGs with more chips or fewer other tenants to compete for access to the RG.
[0152] In some examples, tenant groups may be associated with one or more combinations of RGs and RUHs. In some cases, at least one tenant may be assigned exclusive access to RUH 320 (where RUH 320 points to RU 345a, or where new data can be routed for programming and writing). Note that there may be more tenants than RUHs. As indicated, RUHs can provide isolation (e.g., persistent isolation, initial isolation). Thus, assigning, for example, two tenants to RUHs can provide that the two tenants can influence each other, but are not affected by other tenants besides these two. While not completely isolated for a given tenant, it is better isolation than being affected by each tenant individually. In some cases, at least one tenant may be assigned persistently isolated access to RG 330 in combination with RU 350b and RU 350c, etc. In some examples, the RUH type, the number of tenants per RUH, and / or the number of tenants per RG can provide fine-grained control over NAND access granularity via a host (e.g., machine 105). Persistent isolation can be a property of the RUH and can help allow the driver to perform GC (e.g., when GC is used on data in that RUH).
[0153] Figure 4 An example system 400 according to one or more embodiments as described herein is illustrated. In the illustrated example, system 400 includes a host 405, a host 410, and computing resources 415. Computing resources 415 may include at least one of storage devices (e.g., SSDs, compute storage devices), processing units (e.g., CPUs, GPUs, NPUs, ASICs, FPGAs, SoCs, etc.), memory devices (e.g., DRAMs, SRAMs, HBMs, memory controllers, etc.), communication channels, network connections, and / or some other computing resources. In some cases, computing resources 415 are configured by hosts 405 and / or 410 to provide NAND die time allocation to multiple tenants of hosts 405 and / or 410. In some cases, the NAND die may spend a given amount of execution time for each operation. For example, NAND may spend 50µs for reading, 700µs for programming, 3ms for erasing, command overhead, etc. Therefore, tenants can be considered as arbitrators of die time. In some cases, tenants can accumulate tokens to run each operation on the NAND die. With sufficient tokens, tenants are allowed to perform their operations on the NAND die, keeping it busy for a certain amount of time.
[0154] In the example shown, host 410 includes core 445, core 450, virtual machine (VM) 455, container 460, application 465, and / or programming functionality 470. In some cases, host 405 and / or host 410 may be configured as a hypervisor or VMM. As shown, core 445 may include one or more threads (e.g., thread 475). Core 450 may include one or more threads (e.g., thread 480). In some cases, thread 475 and / or thread 480 may be based on file system metadata threads, a cache of host 410, or a database of host 410 (e.g., with hot and cold data). In some aspects, host 410 may include hardware elements that cause at least a portion of host 410 to be configured, managed, docked, and / or viewed (e.g., from the perspective of host 405, host 410, and / or compute resource 415) based on hardware aspects of host 410. Additionally or alternatively, host 410 may include software elements that cause at least a portion of host 410 to be configured, managed, interfacing, and / or viewed (e.g., from the perspective of host 405, host 410, and / or computing resource 415) based on software aspects of host 410. As shown, core 445 and / or core 450 may be configured, managed, interfacing, and / or viewed from a hardware perspective. Additionally or alternatively, VM 455, container 460, application 465, and / or programming functionality 470 may be configured, managed, interfacing, and / or viewed from a software perspective. In some cases, physical functionality 430 may be associated with one or more hardware aspects of host 410. Additionally or alternatively, virtual functionality 435 and / or virtual functionality 440 may be associated with one or more software aspects of host 410.
[0155] In the example shown, compute resource 415 may include one or more ports (e.g., ports 420 and 425). In some cases, ports 420 and / or 425 may include one or more functions (e.g., one or more physical functions and / or one or more virtual functions). As illustrated, port 425 may include physical function 430, virtual function 435, and virtual function 440. In some cases, at least one virtual function of compute resource 415 (e.g., virtual function 435) may be based on an NVMe controller (e.g., a virtual NVMe controller). In some cases, a tenant may be assigned an NVMe controller. The NVMe controller may include a unique identifier for that NVMe controller (e.g., different from the identifiers of other NVMe controllers).
[0156] In one or more examples, computing resource 415 may present or describe its capabilities for supporting multi-tenancy. Computing resource 415 may be configured to support one or more multi-tenancy features. For example, computing resource 415 may be configured to support the features of SR-IOV and FDP in the multi-tenancy feature set. In some cases, host 410 may be configured to read and determine the capabilities of computing resource 415 (e.g., via physical functions, such as Physical Function 0 (PF0)). In some cases, PF0 may be a management and oversight channel for computing resource 415 (e.g., control path access to computing resource 415). In some aspects, host 410 may identify tenants associated with host 410 and establish each tenant to be identified by computing resource 415. In some examples, host 410 may identify core 445, thread 475, core 450, and thread 480 as tenants, and / or any combination thereof as one or more tenants of computing resource 415. Additionally or alternatively, host 410 may identify VM 455, container 460, application 465 and programming function 470 as tenants, and / or any combination thereof as one or more tenants of computing resource 415.
[0157] When a tenant is configured on compute resource 415, host 410 may send one or more messages to compute resource 415. A message may include one or more fields. The fields of a message may include a tenant ID for at least one tenant. For example, a message sent by host 410 may include one or more tenant IDs (e.g., IDs of two or more tenants used to configure each tenant). The tenant ID may be based on one or more multi-tenancy features of compute resource 415 assigned to the tenant. The tenant ID may be associated with one or more identifiers in a drive. For example, tenant VM 455 may be assigned one or more interface identifiers (e.g., virtual function 435) and one or more media access identifiers (e.g., RUH of compute resource 415). Virtual function 435 may include a virtual function ID, and RUH may include a RUH ID. For example, a tenant ID may be associated with RUH 315 of durability group 305, and RUH 315 may include a RUH ID. In some examples, a tenant ID may be associated with RU 350a of RG 330. Therefore, the tenant ID of VM 455 can be based on the virtual function ID and / or RUH ID (e.g., including at least a portion of the virtual function ID and / or RUH ID). Thus, a message from host 410 can instruct tenant VM 455 to link to virtual function 435 and RUH of compute resource 415. In some examples, fields in a message from host 410 may include at least one parameter field indicating parameters associated with a given tenant (e.g., multi-tenancy constraints). Reference Figure 5Other aspects of the messages from host 410 are discussed.
[0158] In some examples, VM 455 may be configured as a tenant by host 410. For example, each VM of host 410 (including VM 455) may be configured with virtualized storage space, virtualized compute resources, and / or virtualized device groups. Virtualized device groups may include virtual functions (VFs) of compute resource 415 (e.g., virtual function 435), where the VF is viewed as a complete compute resource by VM 455. In some cases, host 410 may create namespaces on compute resource 415 and configure RUHs (e.g., default RUH, such as RUH 320) for the namespaces. Therefore, VM 455 may be configured with a virtualized view of compute resource 415. Thus, the virtualized view of compute resource 415 seen by VM 455 may include virtual function 435 and the NVMe controller on virtual function 435. Additionally or alternatively, the virtualized view of compute resource 415 seen by VM 455 may include at least one namespace with at least one associated RUH (such as RUH 310) (e.g., as booted by host 410).
[0159] In some examples, compute resource 415 can receive messages about the NVMe controller on PF0. Compute resource 415 can process messages and implement multi-tenancy settings for the messages. In some cases, the SSD can describe its resources during an initialization phase, and the host can send a requested performance level to the SSD (e.g., the host can indicate that it wants to use the corresponding ID group for a given tenant group, and request the SSD to arbitrate a tenant relative to another tenant based on the corresponding performance level when the resource is contention-ridden).
[0160] In some examples, host 410 may determine whether to change or update the multitenancy settings of at least one tenant. In some cases, host 410 may monitor the services received by each tenant configured by host 410. Host 410 may determine whether the level of services received by a tenant meets expected performance. When host 410 determines that the level of services received by a tenant does not meet expected performance, host 410 may update the multitenancy settings of one or more tenants (e.g., improve the performance of one or more tenants and / or degrade the performance of one or more tenants). In some cases, host 410 may consider WAF when it determines that the level of services received by a tenant does not meet expected performance. In some cases, tenants may structure their operations to be more orderly, which may reduce the tenant's WAF. Additionally or alternatively, tenants may offload space allocation on SSDs to increase OPs and improve WAF. The host's management software may query the SSDs for performance and related logs. Tenants may monitor the number of commands and / or the latency of those commands. For example, sandboxed applications (e.g., eBPF) may implement this monitoring. In some cases, workloads may be completed during drive qualification to check drive compliance with the requested behavior. In some examples, host 410 may determine that a tenant has better performance (e.g., performance is not up to standard and / or the tenant has been upgraded to a higher performance level, etc.).
[0161] In some examples, host 410 can determine the current settings of a tenant (e.g., at least in part based on a previously configured tenant). Host 410 can send updated multi-tenancy settings to compute resource 415 (e.g., according to an updated NVMe standard for that feature). In some cases, the SSD can describe its resources during an initialization phase, and the host can transmit a requested performance level to the SSD (e.g., the host can indicate that it wants to use the appropriate ID group for a given tenant group, and request the SSD to arbitrate a tenant relative to another tenant based on the appropriate performance level when the resource is contentionable). In some cases, host 410 can send one or more messages (e.g., one or more data packets) to compute resource 415. Tenant VM 455 can be assigned a virtual function 435 and a RUH (e.g., RUH 315 of durability group 305) to compute resource 415. Virtual function 435 can include a virtual function ID, and RUH can include a RUH ID. Therefore, the tenant ID of virtual function 435 can be based on the virtual function ID and / or RUH ID (e.g., the RUH ID of RUH 315).
[0162] In some examples, compute resource 415 receives messages from host 410 on an NVMe controller of the control channel (e.g., an NVMe controller on PF0). Upon receiving a message from host 410, compute resource 415 processes the message and implements it accordingly. For example, compute resource 415 may implement parameters relative to each tenant indicated in the message. In some examples, one or more tenants (e.g., VM 455) may have continuous access to compute resource 415. However, to reach compute resource 415, one or more tenants experience virtual functions that are different from or separate from the control functions (e.g., PF0) used by host 410 to deliver messages to compute resource 415. Therefore, one or more tenants are unaware of the control path access to compute resource 415. Consequently, some performance transients may be associated with host 410 sending messages to compute resource 415 and / or compute resource 415 processing messages.
[0163] Then, compute resource 415 attempts to maintain the multi-tenant performance settings as requested. For example, compute resource 415 attempts to satisfy the requested behavior of all tenants based on its various shared resources. Note that some compute resources can receive and understand the content of messages from host 410 (e.g., compute resource 415), but some compute resources may not be able to provide the requested parameters or provide them fully. Therefore, host 410 may adjust the parameters based on the performance level received by each tenant. In the example, the parameters that can be adjusted may include at least one of the following: reservation, limit, arbitration strength (e.g., which may be proportional to the limit value per tenant), and IO consistency value for each read command group, write command group, other command groups, real-time service category, etc.
[0164] In some examples, host 410 can apply a capability library to provide various types of tenant services. In some cases, computing resource 415 can be considered a terminal or endpoint in a given network. Computing resource 415 can be configured to maintain access contracts by implementing access policing (e.g., policing tenant access to computing resource 415). In some cases, computing resource 415 can implement a leaky bucket algorithm and / or other similar tokenization algorithms for implementing access policing. Additionally or alternatively, computing resource 415 can implement a General Cell Rate Algorithm (GCRA) algorithm for implementing access policing.
[0165] When host 410 configures access for a tenant (e.g., VM 455), messages from host 410 can indicate to compute resource 415 the requested service type (e.g., service level, relatively high data rate service, relatively low data rate service, etc.), parameters for each data flow in both directions, and / or requested quality of service (QoS) parameters in each direction. These parameters can form at least a portion of a communication descriptor (e.g., an access descriptor) used for the connection. These service categories can provide a method for associating access characteristics and QoS requirements with compute resource 415. Service categories can be characterized as real-time or non-real-time. Real-time service categories can include at least one of constant bit rate (CBR) and / or real-time variable bit rate (rt-VBR). Non-real-time service categories can include unspecified bit rate (UBR), adaptive bit rate (ABR), and / or non-real-time variable bit rate (nrt-VBR). In some examples, service categories can be transmitted within the context of GCRA parameters. Therefore, when a driver implementation emulates another algorithm of GCRA, the success or failure of that implementation can be compared to a simplified GCRA implementation for compliance. Therefore, access control for computing resource 415 can provide networking standards for transmitting QoS requests from a host (e.g., host 410) to computing resource 415.
[0166] As shown in the figure, compute resource 415 can provide dual-port access via ports 420 and 425. Therefore, compute resource 415 can provide active-active usage for two different hosts without assuming coordination between the hosts. In some examples, two storage heads (e.g., host 405 and host 410) can operate independently in active-active dual-port sharing based on dual-port access to the SSD (e.g., compute resource 415). In some cases, two storage heads can agree to use non-overlapping LBA ranges of the SSD, but each storage head can access the SSD through different ports (e.g., ports 420 and 425 respectively). In some cases, two storage heads can be assigned separate RUHs (e.g., RUH 310) or separate RUH groups (e.g., RUH 310, RUH 315, etc.), where each port provides equal or balanced access to the SSD.
[0167] As shown, port 425 can provide PCIe functionality (e.g., physical function 430, virtual function 435, and / or virtual function 440). The functionality of port 425 can provide tenants with access to the functionality of computing resource 415 (e.g., dedicated access, shared access). In some cases, each VM on host 410 (e.g., VM 455) can be associated with a port and / or functionality of computing resource 415. In some cases, a VM can be assigned its own virtual functionality (e.g., VM 455 is assigned exclusive or shared access to virtual function 435). In some cases, some VMs can share a backend identifier, such as RUH (e.g., RUH 320), while some VMs can have exclusive access to the RUH.
[0168] In some cases, commit queues (SQs) for identified tenants can be created within a specific virtual function (VF) of compute resource 415 (e.g., virtual function 435, virtual function 440, etc.). In some examples, VF1 of compute resource 415 may include an SQ where host 410 is populated with reads (e.g., all reads) sent by tenants of host 410. In some cases, VF1 may be assigned a priority (e.g., highest priority). VF2 may be configured to include an SQ where host 410 is populated with writes (e.g., all writes) sent by tenants, where VF2 may be assigned the second highest priority. VF3 may be configured to include an SQ where host 410 is populated with file system GC reads (e.g., all file system GC reads). VF4 may be configured to include an SQ where host 410 is populated with file system GC writes (e.g., all file system GC writes). VF5 may be configured to include an SQ where host 410 is populated with deallocations (e.g., all deallocations). VF configuration can be based on a persistent key-value store (e.g., RocksDB) and / or a caching engine (e.g., CacheLib, compute / server-layer cache). In some examples, priorities can be set relative to each other's arbitration strength. Arbitration strength can be based on the relationship between Limit_tenan_1 and limit_tenan_2. In some cases, tenants can be reassigned to a new VF and / or a new arbitration strength can be assigned to a tenant without changing the VF relationship. In some aspects, tenant 1's win rate can be based on BW_arb_weight_1 / (BW_arb_weight_2+BW_arb_weight_1), where "BW_arb_weight" refers to the bandwidth arbitration weight. In some cases, tenant 1's win rate can be based on BW_arb_weight_1 / (the sum of all BW_arb_weights of all active tenants). In some cases, determining tenant 1's win rate can be based on active tenants, as doing so removes any tenants with weights provided by the host in the configuration file but who are not currently actively sending commands to the SSD from the computation.
[0169] Figure 5 A message 500 according to one or more embodiments as described herein is illustrated. In some cases, message 500 may be a message or packet for configuring a tenant (e.g., a message sent by host 410 to compute resource 415 to configure a tenant's access to compute resource 415). In some examples, message 500 includes one or more fields. In some cases, the format of message 500 may be based on a non-volatile memory fast packet format (e.g., an NVMe setup feature command).
[0170] In the example shown, the fields of message 500 may include at least one of the following: header 505, tenant identifier 510 (e.g., up to N tenant identifiers), IOPS limit 515, reserved IOPS 520, bandwidth limit 525, reserved bandwidth 530, invariant CRC (ICRC) 560, and / or frame check sequence (FCS) 565. Alternatively, the fields of message 500 may include at least one of the following: arbitration bandwidth strength 535, arbitration IOPS strength 540, variation limit 545, access consistency 550, and / or access latency 555.
[0171] In some examples, the field for IOPS limit 515 may include an indication (e.g., one or more bit values) of the maximum allowed input / output operations per second (IOPS) on the compute resource used to identify the tenant in the field for tenant identifier 510. The field for reserved IOPS 520 may include an indication of the reservation level of IOPS on the compute resource used by the tenant. In some cases, the field for bandwidth limit 525 may include an indication of the maximum available communication bandwidth between the tenant and the compute resource. The field for reserved bandwidth 530 may include an indication of the reservation level of communication bandwidth between the tenant and the compute resource.
[0172] In some cases, the field 535 used for arbitrating bandwidth strength may include an indication of a tenant's proportional bandwidth win rate relative to the bandwidth win rate of at least one other tenant. For example, when a first tenant and a second tenant compete for at least a portion of the bandwidth of the same computing resources (e.g., the same RUH such as RUH 310, the same RG such as RG 325, the same RU such as RU 345b, the same processing resources, the same memory resources, etc.), the computing resources can resolve the conflict by determining whether the first tenant is assigned a higher bandwidth strength than the second tenant. In some instances, an absolute target value may be implemented. The sum of all absolute target values for tenants may be used to ensure that the sum is less than the driver's capacity. The absolute target value may be implemented where a proportional win rate is implemented. In some cases, the driver may limit tenants to not exceeding the target value. In some aspects, the win rate of tenant 1 may be based on BW_arb_weight_1 / (BW_arb_weight_2+BW_arb_weight_1), where "BW_arb_weight_1" may refer to the bandwidth arbitration weight of tenant 1, etc. In some cases, tenant 1's win rate can be based on BW_arb_weight_1 / (the sum of all BW_arb_weights of all active tenants). In some cases, determining tenant 1's win rate can be based on active tenants because doing so removes any tenants with weights provided by the host in the configuration file but who are not currently actively sending commands to the SSD from the calculation.
[0173] When computing resources determine that a first tenant has been assigned a higher bandwidth strength than a second tenant (e.g., assigned a higher bandwidth priority), the computing resources may block or throttle the second tenant and allow the first tenant to use the bandwidth. In some cases, before allowing the second tenant to use the bandwidth for a second set time amount that may be lower than the first set time amount, the computing resources may allow the first tenant to use the bandwidth first based on the first tenant's bandwidth strength and for the first set time amount.
[0174] The field used for arbitrating IOPS intensity 540 may include an indication of the IOPS win rate as a proportion of the tenant's IOPS win rate relative to at least one other tenant. For example, when computing resources determine that a first tenant is assigned a higher IOPS intensity than a second tenant (e.g., assigned a higher IOPS priority), computing resources may block the second tenant and allow the first tenant to use bandwidth. In some cases, computing resources may allow the first tenant to use a first number of IOPS based on the first tenant's IOPS intensity before allowing the second tenant to use a second number of IOPS that may be lower than the first number of IOPS. In some aspects, tenant 1's IOPS win rate may be based on IOPS_arb_weight_1 / (the sum of all IOPS_arb_weights of all active tenants), where "IOPS_arb_weight_1" may refer to tenant 1's IOPS arbitration weight.
[0175] In some cases, the field for Variation Limit 545 may include an indication of the maximum level of variation in access to compute resources over a period of time. For example, Variation Limit 545 may indicate how much each tenant accesses compute resources and / or the extent to which performance received from compute resources can vary over a given time period. The field for Access Consistency 550 may include an indication of the level of access consistency between tenants and compute resources. For example, Access Consistency 550 may indicate the degree of consistency in each tenant's access to compute resources and / or the degree of consistency in performance received from compute resources over a given time period. In some cases, Access Consistency 550 may be based on the completion time of each command. In some cases, command durations may be placed in bins, and histograms and / or probability density functions (PDFs) may be generated based on bin values. In some cases, consistency may be plotted or evaluated as a cumulative distribution function (CDF) and / or transcendental graph.
[0176] The field used for access latency 555 can include an indication of the maximum permissible access latency between a tenant and the compute resource. For example, access consistency 550 could indicate how much latency is permissible for each tenant when accessing and / or receiving performance from the compute resource. The field used for access latency can include an indication of the target maximum permissible latency for a percentage of commands. If a host is sending a target latency of 10ms for 99.99% of commands, the host may be requesting that only one out of 10,000 commands within a measurement period (e.g., 10 minutes) might exceed the 10ms target.
[0177] In some cases, the ICRC field 560 can be configured to include values (e.g., 32-bit values) that override one or more fields of message 500 (e.g., all fields of message 500), which do not change as message 500 travels from the source port to the destination port. The ICRC field 560 can be generated by the link layer of the source port associated with message 500.
[0178] When a host configures a tenant's access to a computing resource, the host can send a message to the computing resource based on message 500. The message from the host can indicate to the computing resource the requested type of service (e.g., service level, relatively high data rate service, relatively low data rate service, etc.), parameters for each data flow in both directions, and / or requested quality of service (QoS) parameters for each direction. Parameters can be transmitted via one or more fields such as IOPS limit 515, reserved IOPS 520, bandwidth limit 525, reserved bandwidth 530, arbitrated bandwidth strength 535, arbitrated IOPS strength 540, variation limit 545, access consistency 550, and / or access latency 555.
[0179] In some examples, the parameters indicated in message 500 may form at least a portion of a tenant's communication descriptor (e.g., an access descriptor). The parameters can provide a method for associating access characteristics and / or QoS requirements with compute resources. In some cases, the value in tenant identifier 510 may include one or more front-end identifiers and / or one or more back-end identifiers to identify the tenant. The interface or front-end identifier may include at least one component of the compute resource, such as a port, physical function, virtual function, NVM controller, SQ, command type, LBA range, stream, RUH (e.g., RUH 310, RUH 315, etc. of durability group 305), zone, sIOV, GPU, portions of GPU cores, FPGA, and / or compute storage threads. The back-end identifier (e.g., a NAND identifier based on SSD compute resources) may include at least one component of the compute resource, such as a namespace, durability group, RG, RUH, zone, stream, command type, and / or command ID. In some cases, any combination of front-end and / or back-end identifiers may be used to identify the tenant. In some examples, more than one identifier may be used. This can be useful when arbitrating DRAM access or computing resource allocation.
[0180] In some examples, storage compute resources (e.g., SSDs) can be configured to interleave NAND management activities (e.g., internal maintenance of compute resources) during steady state using arbitral bandwidth strength 535 and / or arbitral IOPS strength 540. In some cases, the host can identify such internal operations (e.g., via control path PF0), thereby providing the host with improved visibility into the source of performance changes. In some cases, changes in Y performance at time X can be measured. This measurement can be repeated, and changes can occur in the first instance Y1, the second instance Y2, etc. In some cases, peak deviations from the average can be quantified based on performance limits.
[0181] For storage compute resources (e.g., SSDs), writes, reads, and deallocations travel to the compute resource via different paths. In some cases, reads and writes can be arbitrated against NAND access. In other cases, deallocation and writes can be arbitrated against DRAM access (e.g., SSD DRAM). When only two command types attempt to access the resource (e.g., when only reads and deallocations are involved, 20% reads are modified to 40% reads and 30% deallocations are modified to 60%), the requests for X%, Y%, and Z% of these three command types from the host to different tenant partitions can be modified (e.g., 20% for reads, 50% for writes, and 30% for deallocations).
[0182] Figure 6An example system 600 according to one or more embodiments as described herein is illustrated. System 600 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 600 includes tenant 605, tenant 610, tenant 615 (e.g., N tenants, where tenant 615 is the Nth tenant) and SSD 620 (e.g., computing resources such as computing resource 415).
[0183] In the example shown, tenant 605 can be associated with the first tenant load over time with respect to a first performance level and SSD load, tenant 610 can be associated with the second tenant load over time with respect to a second performance level and SSD load, and tenant 615 can be associated with the Nth tenant load over time with respect to the Nth performance level and SSD load. In some cases, tenant load can be a function of the current queue depth and command length of a given tenant at a given time. In some aspects, the WAF of a given tenant can be estimated based on filters, time-weighted averaging, or other techniques. In some aspects, each command can have several LBAs (e.g., the number of logical blocks (NLB)) in a command descriptor (e.g., a Submission Queue Entry (SQE)). The command length of each command can be the NLB for that command. In some cases, each tenant can use one or more Submission Queues (SQs) to submit commands for SSD work. QD can be defined as the number of commands not completed for SSD across all these SQs. For example, the total QD of a tenant can be equal to the sum of (QDs per SQ).
[0184] In the example shown, the first performance level and SSD load can be based on a limit α_1 (e.g., arbitration weights for IOPS limits and / or bandwidth limits) and a reservation β_1 (e.g., reserved IOPS, reserved bandwidth) for tenant 605, the second performance level and SSD load can be based on a limit α_2 and a reservation β_2 for tenant 610, and the nth performance level and SSD load can be based on a limit α_n and a reservation β_n for tenant 615. For example, host 410 can send at least one message (e.g., message 500) to compute resource 415. In some cases, one or more fields of at least one message may include at least one of limit α_1, reservation β_1, limit α_2, reservation β_2, limit α_n, and / or reservation β_n.
[0185] In some cases, the SSD 620 can make arbitration decisions based on appropriate IOPS limits and reserved bandwidth. In some cases, SSD 620 arbitration can be based on: implementing reservations to protect the SSD 620's resource capabilities; the utilization of the SSD 620's resources for each tenant changing over time; removing tenants with in-flight orders (no queue depth) from the arbitration of a given resource (e.g., idle or read-only tenants do not need to arbitrate for write resources); QoS controlled by reservations; and / or permitted (e.g., maximum) resource utilization (e.g., IOPS, bandwidth, etc.) controlled by arbitration weights (e.g., IOPS limits and / or bandwidth limits). Limit values can instruct the SSD that tenants are not allowed to exceed the limit value. IOPS and bandwidth limit values can be assigned independently. Reservation values can instruct the SSD that tenants requesting SSD capabilities below (e.g., IOPS or bandwidth) the reservation value will always be able to access the resource. Arbitration weights can instruct the SSD that when a resource is contested by two tenants (e.g., two or more tenants), the arbitration weights between the two tenants will be compared. Based on each tenant's win ratio, the SSD is requested to attempt the ratio of arbitration weights in an approximate steady-state measurement.
[0186] Based on the multi-tenancy parameter configuration for each tenant (e.g., IOPS and / or bandwidth limits, IOPS and / or bandwidth reservations), tenant 605 may use 10% of the SSD 620's resources for a given time period, tenant 610 may use 55% of the SSD 620's resources for that time period, and tenant 615 may use 35% of the SSD 620's resources for that time period. In some cases, the host (e.g., host 410) may determine whether the resource utilization at a given time period is aligned with the multi-tenancy parameter configuration for each tenant. In some cases, when the host determines that the resource utilization of at least one tenant is inconsistent with the corresponding multi-tenancy parameter configuration, the host may update the multi-tenancy parameter configuration for at least one tenant.
[0187] In some cases, various settings (e.g., limit α, reserve β, etc.) can be relative between tenants. For example, a host can assign various settings to its tenants relative to the performance level assigned to each host. In some cases, a host can monitor the arbitration success rate of different tenants. Note that for tenant 1's IOPS read command group, there can be requested reserve, limit, IO consistency, and / or arbitration strength. For tenant 1's bandwidth read command group, there can be requested reserve, limit, IO consistency, and / or arbitration strength. For tenant 1's IOPS write command group, there can be requested reserve, limit, IO consistency, and / or arbitration strength. For tenant 1's bandwidth write command group, there can be requested reserve, limit, IO consistency, and / or arbitration strength. For tenant 1's other IOPS command groups, there can be requested reserve, limit, IO consistency, and / or arbitration strength. For tenant 1's other bandwidth command groups, there can be requested reserve, limit, IO consistency, and / or arbitration strength.
[0188] In some examples, monitoring can be based on the total IOPS and / or bandwidth for each command type or category, etc. Actions that the SSD 620 can take include temporary adjustments (e.g., increases, decreases) to the limiting or retention behavior of a given tenant, permanent adjustments to the limiting or retention behavior, implementing bounded or unbounded changes affecting another tenant, adjusting the arbitration settings for each tenant, breaking down relatively large commands into smaller commands and / or larger groups, increasing the arbiter cycle, changing the arbiter token quantization, changing the arbiter type, etc. In some examples, a token shortage can trigger one or more actions that the SSD 620 can take. In some cases, the SSD 620 can increase the permissible programming / erasing pauses. In some examples, the SSD 620 can defer internal SSD activities, such as metadata logging, NAND management, etc., based on tenant burst behavior, deferring internal SSD activities until the burst subsides. In some cases, the SSD 620 can implement bypass operations (e.g., based on settings received from the host and / or based on the SSD configuration) to allow hardware blocks to jump to the head-of-the-line-blocking. In some cases, the SSD 620 can implement conventional arbitration modes for arbitrating resources.
[0189] Figure 7An example system 700 according to one or more embodiments as described herein is illustrated. System 700 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 700 includes tenant 705, tenant 710, tenant 715 (e.g., N tenants, where tenant 715 is the Nth tenant), SSD 720 (e.g., computing resources such as computing resource 415), and internal tenant 725. In some cases, SSD 720 may include one or more internal tenants (e.g., internal tenant 725) to monitor and / or manage internal SSD operations (e.g., for internal maintenance of SSD 720).
[0190] As shown in the figure, similar to tenants 705, 710, and 715, internal tenant 725 can be associated with internal tenant load that varies over time with respect to internal performance level and SSD load. As shown in the figure, internal performance level and SSD load can be based on internal tenant 725's limit α_int (e.g., arbitration weights for IOPS limits and / or bandwidth limits) and reservation β_int (e.g., reserved IOPS, reserved bandwidth).
[0191] In some examples, the SSD 720 can be configured with an internal tenant 725 to provide control, logging, and / or system behavior information of the SSD 720 to one or more hosts. The internal tenant 725 can be configured for NAND management activities such as read patrol, write flushing, handling NAND errors, and SSD controller metadata (e.g., logical-to-physical (L2P) logging, firmware updates, etc.). In some cases, the SSD 720 can provide a minimum resource allocation to the internal tenant 725 (e.g., an allocation based on estimated end-of-life or worst-case requirements of the SSD that are not attributable to or associated with host I / O or host tenant SSD activity).
[0192] In some cases, the needs of the internal tenant 725 may be based on NAND management activities such as read inspections, closing erase blocks that have been open for too long and risk data loss, testing the reliability of stored data, latency testing of NAND behavior (e.g., longer / shorter programming time compared to expectations), SSD metadata storage, power failure protection activities, etc. In some cases, the host and / or SSD 720 may allow this activity to occur based on the arbitration win rate determined by the settings of the host and / or SSD 720's arbitration engine. In some cases, commit queue (SQ) arbitration may not be affected when the internal tenant 725 does not insert SQ entries. However, DRAM access may be based on internal needs to implement arbitration strategies to update superblock trace information used for L2P table or GC management of incoming buffers or power failure preparedness.
[0193] In some examples, the host can adjust the quorum strength (e.g., IOPS limits, bandwidth limits) of internal tenant 725 based on one or more tenants and / or SSD 720 settings. A higher quorum strength for internal tenant 725 can imply that the SSD 720 should perform as much background internal work ahead of time. In some cases, a higher quorum strength setting for internal tenant 725 can be used as a mechanism to improve the latency of the SSD 720. In some cases, a lower quorum strength for internal tenant 725 can imply that the SSD 720 should assume that the tenant's workload is bursty (e.g., periods of high and low tenant load). Therefore, the SSD 720 can determine to defer at least some internal activity to improve the QoS of the host tenant. However, deferring too much can risk the SSD 720 reaching a critical threshold where internal activity is mandatory, which may degrade the performance of the host tenant. In some cases, the host and / or SSD 720 can adjust the quorum settings of internal tenant 725 to interleave NAND management activity during steady-state periods.
[0194] In some examples, the internal tenant 725 can be used to set an arbitration rate request within a multi-tenant or QoS environment. In some examples, it can allow the host to configure internal tenant behavior. With the host setting a relatively high level, the host can encourage the drive to lead on internal activities. With the host setting a relatively low level, the host may be requesting the drive to defer internal activities to improve I / O consistency and / or latency. Additionally or alternatively, the internal tenant 725 can be used to manipulate when internal SSD activities occur (defer them, accelerate them, execute them immediately). For example, the internal tenant 725 can execute internal SSD activities, defer them, or accelerate them, thereby improving SSD I / O consistency and latency. Additionally or alternatively, the internal tenant 725 can be used to manipulate the SSD I / O consistency or latency of the SSD 720.
[0195] Based on the multi-tenancy parameter configuration for each tenant (e.g., IOPS and / or bandwidth limits, IOPS and / or bandwidth reservations), tenant 705 may use 27% of the SSD 720's resources for a given period, tenant 710 may use 6% of the SSD 720's resources for that period, tenant 715 may use 18% of the SSD 720's resources for that period, and internal tenant 725 may use 1% of the SSD 720's resources for that period. In some cases, the host (e.g., host 410) may determine whether the resource utilization at any given time is consistent with the multi-tenancy parameter configuration for each tenant. In some cases, when the host determines that the resource utilization of at least one tenant is inconsistent with the corresponding multi-tenancy parameter configuration, the host may update the multi-tenancy parameter configuration for at least one tenant.
[0196] Figure 8 An example system 800 according to one or more embodiments as described herein is illustrated. System 800 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 800 includes tenant 805, tenant 810, internal tenant 815, and SSD 820 (e.g., computing resources such as computing resource 415). In some cases, SSD 820 may include one or more internal tenants (e.g., internal tenant 815) to monitor and / or manage internal SSD operations (e.g., for internal maintenance of SSD 820).
[0197] In some examples, the SSD 820 can be configured with an internal tenant 815 to provide control, logging, and / or system behavior information of the SSD 825 to one or more hosts. The internal tenant 815 can be configured for NAND management activities such as read inspections, write flushes, handling NAND errors, and SSD controller metadata (e.g., logical-to-physical (L2P) logging, firmware updates, etc.). In some cases, the SSD 820 can provide a minimal resource allocation to the internal tenant 815.
[0198] In some examples, tenant 805, tenant 810, and internal tenant 815 can be configured with one or more multi-tenancy configuration parameters. As shown, tenant 805 can be associated with a first limit and a first reservation (e.g., a first performance limit for IOPS and / or bandwidth, a first performance reservation for IOPS and / or bandwidth) relative to its tenant load over time; tenant 810 can be associated with a second limit and a second reservation relative to its tenant load over time; and internal tenant 815 can be associated with an internal reservation (e.g., but not an internal limit) relative to its tenant load over time. In some examples, the host can use commands to set the first limit and / or the first reservation. In some cases, the command can be formatted with fields for each of the values (e.g., first limit, first reservation). In some cases, the command may include fields for reservation, limit, arbitration strength, and / or I / O consistency (e.g., fields for multiple RG / RUH combinations). In some cases, commands can be configured with a first setting for reservation, limitation, I / O variation, and / or arbitration weights to configure the NAND associated with a read command group, a second setting for reservation, limitation, I / O variation, and / or arbitration weights to configure the NAND associated with a write command group, and a third setting for reservation, limitation, I / O variation, and / or arbitration weights to configure the NAND associated with other command groups. As shown in the figure, tenant 805 can receive first-level performance from SSD 820 based on a first limitation and / or a first reservation, tenant 810 can receive second-level performance from SSD 820 based on a second limitation and / or a second reservation, and internal tenant 815 can receive internal-level performance from SSD 820 based on internal reservations.
[0199] In the example shown, the SSD 820 can use the appropriate reservation parameters of tenants to protect the availability of SSD 820 resources for top-level tenants (e.g., for the highest priority tenants). In some examples, the host can set performance levels and send performance level requests to the SSD for SSD maintenance relative to each tenant. The highest priority tenant can have higher limits, higher reservations, stricter I / O change requirements, and / or higher arbitration weights compared to lower priority tenants. In some cases, the SSD 820 can use the appropriate reservation parameters of tenants to generally increase the availability of the SSD 820 (e.g., make the SSD 820 more idle, make the SSD 820 more consistent in performance across tenants). In some aspects, the host can set performance levels and send performance level requests to the SSD for SSD maintenance relative to each tenant. Other tenants can reduce their limit values, reservations, etc., resulting in a relative performance improvement for another tenant. In some cases, I / O change requirement settings can be tightened, and the arbitration weight of tenants can be increased. In some examples, the appropriate parameters can be used based on IOPS and / or bandwidth per VF, RUH, etc. In some cases, the SSD 820 can use the appropriate limit parameters of the tenants to maintain general consistency between tenants, keep performance latency within the expected margin, and / or maintain the relative performance level between tenants.
[0200] In some examples, the SSD 820 can achieve proportional reduction for one or more tenants. For instance, the SSD 820 can reduce the limits and / or retention performance of one or more tenants during relevant periods of high tenant load. As shown in the figure, the performance requirements of tenant 805 and tenant 810 may simultaneously exceed their respective limits (e.g., due to relevant high load). Therefore, the SSD 820 can reduce the limits and / or retention performance of tenant 805 and / or tenant 810 during relevant periods of high tenant load.
[0201] When a tenant's requested performance exceeds its assigned limits (e.g., tenant 805 and / or tenant 810), the SSD 820 can impose a cap on the tenant. As shown in the figure, during periods of high load for tenants 805 and 810, the received performance is capped across both tenants. In some cases, the SSD 820 can allow tenants to exceed their limits, provided that the resources requested by the SSD 820 will not be used by other tenants. For example, when tenant 805's performance request exceeds tenant 805's limit and there is no competition with another tenant (e.g., the performance requests of tenant 810 and internal tenant 815 do not compete with tenant 805's performance request), the SSD 820 can allow tenant 805 to exceed its assigned limits.
[0202] In some examples, the retention level of one or more tenants can be modified by the tenant's host and / or SSD 820. In some cases, the performance retention of one or more tenants can be removed from the SSD 820 (e.g., as a result of the SSD 820's available capacity relative to one or more tenants), thus preventing other tenants from accessing or utilizing the full capacity of the SSD 820. In some cases, the performance retention of one or more tenants can be maintained on the SSD 820 (e.g., not removed). If activity increases from a previously below-reserve value, a given tenant can achieve its retention relatively quickly. In some cases, the performance retention of one or more tenants can be adjusted to modify the change or latency behavior of one or more tenants. For example, the retention of tenant 805 can be adjusted to modify the change or latency behavior of tenant 805 and / or modify the change or latency behavior of tenant 810.
[0203] Figure 9 An example system 900 according to one or more embodiments described herein is illustrated. System 900 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 900 includes tenant 905, tenant 910, internal tenant 915, and SSD 920 (e.g., computing resources such as computing resource 415). In some cases, SSD 920 may include one or more internal tenants (e.g., internal tenant 915) to monitor and / or manage internal SSD operations (e.g., for internal maintenance of SSD 920).
[0204] In some examples, tenants 905, 910, and internal tenant 915 can be configured with one or more multi-tenancy configuration parameters. As shown, tenant 905 can be associated with a first limit and a first reservation relative to its tenant load over time, tenant 910 can be associated with a second limit and a second reservation relative to its tenant load over time, and internal tenant 915 can be associated with an internal reservation (e.g., but not an internal limit) relative to its tenant load over time. As shown, tenant 905 can receive a first level of performance from SSD 920 based on the first limit and / or the first reservation, tenant 910 can receive a second level of performance from SSD 920 based on the second limit and / or the second reservation, and internal tenant 915 can receive internal level performance from SSD 920 based on the internal reservation.
[0205] In the example shown, the reservation setting for tenant 905 can be increased (e.g., via the host with tenant 905 and / or SSD 920). As a result, tenant 910 (e.g., a combination of tenants other than tenant 905 associated with SSD 920) no longer receives the same performance from SSD 920 because the increased reservation setting for tenant 905 protects the performance level from SSD 920, reducing the performance received by tenant 910. The reduced performance received by tenant 910 (e.g., by other tenants) means that tenant 905 receives its requested performance with higher reliability. In some cases, a host can use a tenant's modified reservation setting (e.g., increasing the reservation setting for one or more top-performing high-performance tenants) until variability and consistency goals are achieved for that tenant.
[0206] In some approaches, low-priority and high-priority tenants compete for resources and access to the SSD. Based on the system and method described in this paper, stricter controls are applied to lower-priority tenants while providing more consistent service to high-priority tenants. Furthermore, the system and method keep higher-priority tenants within limits to balance the load sharing of the SSD 920 among tenants.
[0207] Figure 10 An example system 1000 according to one or more embodiments as described herein is illustrated. System 1000 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 1000 includes tenant 1005, tenant 1010, internal tenant 1015, and SSD 1020 (e.g., computing resources such as computing resource 415). In some cases, SSD 1020 may include one or more internal tenants (e.g., internal tenant 1015) to monitor and / or manage internal SSD operations (e.g., for internal maintenance of SSD 1020).
[0208] In some examples, tenants 1005, 1010, and internal tenant 1015 can be configured with one or more multi-tenancy configuration parameters. As shown, tenant 1005 can be associated with a first limit and a first reservation relative to its tenant load over time, tenant 1010 can be associated with a second limit and a second reservation relative to its tenant load over time, and internal tenant 1015 can be associated with an internal reservation (e.g., but not an internal limit) relative to its tenant load over time. As shown, tenant 1005 can receive a first level of performance from SSD 1020 based on the first limit and / or the first reservation, tenant 1010 can receive a second level of performance from SSD 1020 based on the second limit and / or the second reservation, and internal tenant 1015 can receive internal level performance from SSD 1020 based on the internal reservation.
[0209] In the example shown, the reservation setting for internal tenant 1015 can be increased (e.g., via the host of tenant 1005 and / or SSD 1020). As a result, none of the other tenants of SSD 1020 (e.g., tenant 1005, tenant 1010, etc.) are able to obtain the same reservation setting as that increased for internal tenant 1015 (e.g., from a host similar to...). Figure 8 and / or Figure 9 The relatively low internal retention levels shown will have as much performance as before. In some examples, the host can control all tenant settings. In some cases, the host can transfer desired settings to the SSD to maintain internal tenants. In this way, the SSD can provide some internal control over the activities of internal tenants.
[0210] As shown in the figure, the performance received by tenants 1005 and 1010 is reduced due to the increased reservation setting for internal tenant 1015. Even if internal tenant 1015 does not use the full performance reserved for it, the performance of the host tenant is reduced, allowing SSD 1020 to provide up to the increased reservation setting for internal tenant 1015 at any given time. Based on the increased reservation setting for internal tenant 1015, the increased reservation setting for host tenants of SSD 1020 (e.g., tenant 1005, tenant 1010, etc.) mitigates bursts in requested performance. Based on the increased reservation setting for internal tenant 1015, the performance received by the host tenant is smoother.
[0211] In some approaches, low-priority and high-priority tenants compete for resources and access to the SSD. Based on the systems and methods described in this paper, stricter control over changes is imposed on host tenants; however, host tenants may risk not receiving as much average performance as possible because a portion of the performance that the SSD 1020 can provide is reserved for internal tenant 1015, even when internal tenant 1015 does not use all available performance until the increased reservation level.
[0212] Figure 11 An example system 1100 according to one or more embodiments as described herein is illustrated. System 1100 may represent multiple tenants accessing computing resources (e.g., storage computing resources). As shown, system 1100 includes tenant 1105, tenant 1110, internal tenant 1115, and SSD 1120 (e.g., computing resources such as computing resource 415). In some cases, SSD 1120 may include one or more internal tenants (e.g., internal tenant 1115) to monitor and / or manage internal SSD operations (e.g., for internal maintenance of SSD 1120).
[0213] In some examples, tenants 1105, 1110, and internal tenant 1115 can be configured with one or more multi-tenancy configuration parameters. As shown, tenant 1105 can be associated with a first limit and a first reservation relative to its tenant load over time, tenant 1110 can be associated with a second limit and a second reservation relative to its tenant load over time, and internal tenant 1115 can be associated with an internal reservation (e.g., but not an internal limit) relative to its tenant load over time. As shown, tenant 1105 can receive a first level of performance from SSD 1120 based on the first limit and / or the first reservation, tenant 1110 can receive a second level of performance from SSD 1120 based on the second limit and / or the second reservation, and internal tenant 1115 can receive internal level performance from SSD 1120 based on the internal reservation.
[0214] In the example shown, the limit settings for tenant 1105 can be increased (e.g., via the host of tenant 1105 and / or SSD 1120). For example, SSD 1120 can temporarily increase the limit for tenant 1105 based on a temporary request from the host of tenant 1105. In some cases, tenant 1105 may pay an increased level of performance (e.g., a temporary or permanent increase for performance consistency).
[0215] In some cases, the limit settings for tenant 1105 can be increased to improve the consistency of performance received by tenant 1105. For example, increasing the limit settings for tenant 1105 can increase the priority of performance received by tenant 1105 relative to that received by other tenants (e.g., tenant 1110 and / or internal tenant 1115). As shown, increasing the limit settings for tenant 1105 degrades the average variation of other tenants (e.g., tenant 1110 and / or internal tenant 1115). Alternatively or concurrently, the arbitration strength settings for tenant 1105 can be increased to improve performance consistency. Tenants with high arbitration requirements receive more consistent access to the resources of SSD 1120. For example, increasing the arbitration strength for tenant 1105 can increase the priority of performance received by tenant 1105 relative to that received by other tenants (e.g., tenant 1110 and / or internal tenant 1115, etc.). With the increase in arbitration strength, tenant 1105 is configured to more reliably win arbitration (e.g., when competing for the same resources of SSD 1120 with one or more other tenants such as tenant 1110, etc.). As shown, increasing the arbitration victory of tenant 1105 means that the average change for tenants with less arbitration strength (e.g., tenant 1110) is expected to be downgraded.
[0216] In some approaches, low-priority tenants and high-priority tenants compete for resources and access to the SSD. In some examples, the lower-priority limit settings and / or arbitration strength settings can be increased, resulting in stricter control over changes being imposed on lower-priority tenants, where higher-priority tenants may experience performance degradation because lower-priority tenants are allowed to win arbitration at a higher rate.
[0217] Figure 12 A flowchart illustrating an example method 1200 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1200 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1200 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1200 is merely one implementation, and one or more operations of method 1200 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0218] At 1205, method 1200 may include a performance request arriving at a given resource arbitrator at a given time (e.g., such as...). Figure 4 The resource arbitrator of computing resources such as computing resource 415. Performance requests can be based on the performance requested by the tenant. Examples of performance requests may include at least one of IOPS limit 515, reserved IOPS 520, bandwidth limit 525, reserved bandwidth 530, arbitrated bandwidth strength 535, arbitrated IOPS strength 540, variation limit 545, access consistency 550, and / or access latency 555. For example, the resource arbitrator of multi-tenant controller 140 can receive and process performance requests from host tenants.
[0219] At 1210, method 1200 may include determining whether the performance request exceeds the IOPS limit assigned to the host tenant. For example, multitenant controller 140 may determine whether the performance request exceeds the assigned IOPS limit. When it is determined that the performance request exceeds the IOPS limit assigned to the host tenant, method 1200 may proceed to 1215. When it is determined that the performance request does not exceed the IOPS limit assigned to the host tenant, method 1200 may proceed to 1220.
[0220] At 1215, method 1200 may include indicating that a performance request is a non-compliant request. For example, multi-tenancy controller 140 may (e.g., via control path PF0) indicate to a host tenant that a performance request is a non-compliant request. In some cases, multi-tenancy controller 140 may provide a host tenant with a performance level lower than the requested performance.
[0221] At 1220, method 1200 may include determining whether the performance request exceeds the allowed IOPS variance allocated to the host tenant. For example, multitenancy controller 140 may determine whether the requested performance exceeds the allowed IOPS variance, where the allowed IOPS variance indicates the range by which the performance received by the host tenant can vary over time from a baseline level of IOPS performance (e.g., a allowed variation of 25% above or below the baseline level of IOPS performance over time). When it is determined that the requested performance exceeds the allowed IOPS variance allocated to the host tenant, method 1200 may proceed to 1215. When it is determined that the performance request does not exceed the allowed IOPS variance, method 1200 may proceed to 1225. In some cases, the IOPS variance may be requested by the host. The IOPS variance may be related to a real-world metric, such as the number of IOs outside certain variation limits (e.g., 5%) over a period of time (e.g., 1 second). In some cases, the IOPS variance may be an arbitrary ratio. Tenant 1 can have an IO variation of 3 on a scale of 1-10, and tenant 2 can have an IO variation of 7 on the same scale of 1-10. Relatively speaking, 7 can mean that the SSD is trying to maintain a more stringent variation for tenant 2, and therefore, tenant 2 can more reliably win resources due to the higher IOPS variation.
[0222] At 1225, method 1200 may include indicating that the performance request is a compliant request. For example, multitenant controller 140 may indicate to a host tenant that the performance request is a compliant request. In some cases, multitenant controller 140 may provide the host tenant with a performance level consistent with the requested performance.
[0223] At 1230, method 1200 may include determining whether the performance request exceeds the bandwidth limit assigned to the host tenant. For example, multi-tenancy controller 140 may determine whether the performance request exceeds the assigned bandwidth limit. When it is determined that the performance request exceeds the bandwidth limit assigned to the host tenant, method 1200 may proceed to 1235. When it is determined that the performance request does not exceed the bandwidth limit assigned to the host tenant, method 1200 may proceed to 1240.
[0224] At 1235, method 1200 may include indicating that the performance request is a non-compliant request. For example, multi-tenancy controller 140 may (e.g., via control path PF0) indicate to a host tenant that the performance request is a non-compliant request. In some cases, multi-tenancy controller 140 may provide a host tenant with a performance level lower than the requested performance.
[0225] At 1240, method 1200 may include determining whether the performance request exceeds the permissible bandwidth difference assigned to the host tenant. For example, multi-tenancy controller 140 may determine whether the requested performance exceeds the permissible bandwidth difference, wherein the permissible bandwidth difference indicates the range by which the performance received by the host tenant can vary over time from a baseline level of bandwidth (e.g., a permissible difference of 25% above or below the baseline level of bandwidth over time). When it is determined that the requested performance exceeds the permissible bandwidth difference assigned to the host tenant, method 1200 may proceed to 1235. When it is determined that the performance request does not exceed the permissible bandwidth difference, method 1200 may proceed to 1245.
[0226] At 1245, method 1200 may include indicating that the performance request is a compliant request. For example, multitenant controller 140 may indicate to a host tenant that the performance request is a compliant request. In some cases, multitenant controller 140 may provide the host tenant with a performance level consistent with the requested performance.
[0227] Based on the disclosed system and method, a leaky bucket algorithm can be implemented for scheduling tenant access to SSD resources. In some cases, method 1200 can be based on the leaky bucket algorithm. For requests that satisfy both IOPS and bandwidth (e.g., request commands), the request can be allowed to win resources and continue. In some cases, the token count can be modified (e.g., increased, decreased) for the request. In some aspects, when the SSD identifies a tenant at risk of violating its I / O changes, a temporary allocation of additional tokens can be provided to the tenant. In this example, another tenant might lose tokens because that tenant's I / O changes are within permissible limits. For non-compliant requests, non-compliant requests on either branch can be placed in a buffer (e.g., a first-in, first-out buffer). Alternatively or concurrently, the arbitration process can continue until some or all commands are compliant. For example, a command can be blocked, and the tenant can wait for sufficient tokens to accumulate. In some cases, a command can be broken down into smaller parts (e.g., by the SSD, by the host, by the tenant), which can be done in parts. In some cases, a portion of the request can continue (e.g., allowing 25% of the requested performance to continue). In some cases, computing resources (e.g., SSDs) can perform business shaping, access shaping, business policing, and / or access policing. Access policing and / or business policing may include computing resources sending signals to hosts to instruct hosts to modify their behavior and / or the behavior of one or more tenants of the host.
[0228] In some cases, method 1200 can be extended to a variety of performance indicators. For example, method 1200 can be extended to at least one of per-tenant input / output operations per second (IOPS) (e.g., maximum per-tenant IOPS), per-tenant bandwidth (e.g., maximum per-tenant bandwidth), per-tenant reserved IOPS, per-tenant reserved bandwidth, and / or per-tenant quality of service (QoS). In some cases, per-tenant QoS can be based on permissible variations in per-tenant performance (e.g., permissible variations in a tenant's performance relative to one or more other tenants). Additionally or alternatively, method 1200 can be extended to at least one of priority and arbitration weights between tenants on disputed resources, scalability of one or more computing resources (e.g., one or more SSDs), tenant workload, SSD workload, and / or modification (e.g., at least temporary modification) of multi-tenant parameters based on resource contention (e.g., IOPS limits, bandwidth limits, reserved IOPS, reserved bandwidth, etc.). Additionally or alternatively, method 1200 can be extended to at least one of a combination of write amplification factor (WAF) and / or internal device services (e.g., SSD maintenance). In some examples, the arbitration scheme of Method 1200 above can be applied to various command or operation types (e.g., read operations, write operations, data modification operations, copy operations, deallocation operations, garbage collection operations, etc.). In some examples, the tenant's host request arbitration behavior can be implemented based on the General Cell Rate Algorithm (GCRA). For example, Method 1200 can be based on one or more aspects of GCRA.
[0229] Figure 13 A flowchart illustrating an example method 1300 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1300 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1300 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1300 is merely one implementation, and one or more operations of method 1300 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0230] At 1305, method 1300 may include an identifier that identifies a first tenant of the storage device. For example, computing resource 415 may identify (e.g., based on message 500) the identifier of the first tenant of computing resource 415.
[0231] At 1310, method 1300 may include assigning a first performance level to a first tenant. For example, computing resource 415 may assign the first performance level to the first tenant.
[0232] At 1315, method 1300 may include generating a first performance parameter based on a first performance level. For example, host 410 may generate the first performance parameter based on a first performance level (e.g., based on communication from computing resource 415).
[0233] At 1320, method 1300 may include sending a configuration message to the storage device that includes a first performance parameter and an identifier of a first tenant. For example, host 410 may send a configuration message to computing resource 415, wherein the fields of the configuration message include the first performance parameter and the identifier of the first tenant.
[0234] Figure 14 A flowchart illustrating an example method 1400 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1400 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1400 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1400 is merely one implementation, and one or more operations of method 1400 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0235] At 1405, method 1400 may include an identifier that identifies a first tenant of the storage device. For example, computing resource 415 may identify (e.g., based on message 500) the identifier of the first tenant of computing resource 415.
[0236] At 1410, method 1400 may include assigning a first performance level to a first tenant. For example, computing resource 415 may assign the first performance level to the first tenant.
[0237] At 1415, method 1400 may include generating a first performance parameter based on a first performance level. For example, host 410 may generate the first performance parameter based on the first performance level (e.g., based on communication from computing resource 415).
[0238] At 1420, method 1400 may include sending a configuration message to the storage device that includes a first performance parameter and an identifier of a first tenant. For example, host 410 may send a configuration message to computing resource 415, wherein the fields of the configuration message include the first performance parameter and the identifier of the first tenant. In some cases, host 410 may send a sequence of configuration messages for multiple tenants to computing resource 415, wherein each configuration message in the sequence of configuration messages corresponds to one of the multiple tenants. For example, host 410 may send a first configuration message in a sequence of configuration messages, send a second configuration message in a sequence of configuration messages, and so on, wherein the first configuration message corresponds to the first tenant, the second configuration message corresponds to the second tenant, and so on.
[0239] At 1425, method 1400 may include assigning a second performance level to the second tenant based on an identifier that identifies the second tenant. For example, multi-tenancy controller 140 may assign a second performance level to the second tenant based on an identifier that identifies the second tenant.
[0240] Figure 15 A flowchart illustrating an example method 1500 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1300 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1300 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1300 is merely one implementation, and one or more operations of method 1300 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0241] At 1505, method 1500 may include running legacy behavior settings on consecutive SSDs. For example, compute resource 415 may run a first behavior setting.
[0242] At 1510, method 1500 may include the host providing new settings for the arbitration engine specifying IOPS and bandwidth behavior. For example, host 410 may provide a second behavior setting or a behavior setting update. In some cases, the second behavior setting may specify new settings for IOPS and / or bandwidth behavior.
[0243] At 1515, method 1500 may include SSD changing resource arbitration engine settings, limits, etc. For example, computing resource 415 may change resource arbitration engine settings, limits, etc. based on receiving the second-line settings.
[0244] At 1520, method 1500 may include a completion command. For example, based on receiving a second action setting (e.g., a command from host 410), computing resource 415 may complete the implementation of the modifications indicated in the second action setting.
[0245] At 1525, method 1500 may include running new behavior settings on the continuous SSD. For example, compute resource 415 may continue to operate based on the implementation of the modifications indicated in the second behavior settings.
[0246] Figure 16 A flowchart illustrating an example method 1600 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1600 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1600 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1600 is merely one implementation, and one or more operations of method 1600 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0247] At 1605, method 1600 may include running legacy behavior settings on consecutive SSDs. For example, compute resource 415 may run a first behavior setting.
[0248] At 1610, method 1600 may include a host providing new settings for improving the permissible variation / consistency of one or more tenants in a first group and / or reducing the permissible variation / consistency of one or more tenants in a second group. For example, host 410 may provide a second behavioral setting or an updated behavioral setting. In some cases, the second behavioral setting may include settings for consistency. In some examples, the second behavioral setting may include settings for improving the permissible variation and / or consistency of one or more tenants and / or reducing the permissible variation and / or consistency of one or more tenants. In some cases, host 410 may send an aggregated configuration message for multiple tenants to computing resource 415, wherein the configuration message corresponds to settings for multiple tenants. For example, host 410 may send an aggregated configuration message to computing resource 415, wherein the aggregated configuration message indicates the performance level requested by the first tenant and the first tenant (e.g., an updated performance level), indicates the performance level requested by the second tenant and the second tenant (e.g., an updated performance level), and so on.
[0249] At 1615, method 1600 may include SSD changing resource arbitration engine settings, limits, etc. For example, computing resource 415 may change resource arbitration engine settings, limits, etc. based on receiving the second-line settings.
[0250] At 1620, method 1600 may include a completion command. For example, based on receiving a second action setting (e.g., a command from host 410), computing resource 415 may complete the implementation of the modifications indicated in the second action setting.
[0251] At 1625, method 1600 may include running new behavior settings on the contiguous SSD. For example, compute resource 415 may continue operating based on the implementation of the modifications indicated in the second behavior settings.
[0252] Figure 17 A flowchart illustrating an example method 1700 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1700 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1700 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1700 is merely one implementation, and one or more operations of method 1700 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0253] At 1705, method 1700 may include running legacy behavior settings on contiguous SSDs. For example, compute resource 415 may run a first behavior setting.
[0254] At 1710, method 1700 may include new settings provided by the host for change / consistency. For example, host 410 may provide second behavioral settings or updated behavioral settings. In some cases, the second behavioral settings may include settings for consistency and / or change. For example, the second behavioral settings may include settings for indicating the expected consistent performance of one or more tenants. Additionally or alternatively, the second behavioral settings may include settings for indicating the expected performance changes of one or more tenants (e.g., acceptable changes in a tenant's performance relative to one or more other tenants).
[0255] At 1715, method 1700 may include SSD changing resource arbitration engine settings, limits, etc. For example, computing resource 415 may change resource arbitration engine settings, limits, etc. based on receiving the second-line settings.
[0256] At 1720, method 1700 may include a completion command. For example, based on receiving a second action setting (e.g., a command from host 410), computing resource 415 may complete the implementation of the modifications indicated in the second action setting.
[0257] At 1725, method 1700 may include running new behavior settings on the continuous SSD. For example, computing resource 415 may continue to operate based on the implementation of the modifications indicated in the second behavior settings.
[0258] Figure 18 A flowchart illustrating an example method 1800 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1800 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1800 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1800 is merely one implementation, and one or more operations of method 1800 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0259] At 1805, method 1800 may include host 410 identifying one or more tenants and determining the performance requirements of each of the one or more tenants.
[0260] At 1810, method 1800 may include host 410 sending a capability query to computing resource 415. In some cases, host 410 may execute 1805 and 1810 sequentially (e.g., 1805 then 1810, or 1810 then 1805) or simultaneously (e.g., 1805 and 1810 are executed relatively simultaneously).
[0261] At 1815, method 1800 may include computing resource 415 sending a capability response to host 410. In some examples, host 410 may store the capability response. In some cases, the capability response may indicate one or more capabilities of computing resource 415. Examples of one or more capabilities of computing resource 415 may include write speed, read speed, write latency, read latency, random data access latency, cache size, cache bandwidth, cache latency, firmware version, boot time, data bandwidth, power requirements, energy efficiency, number of tenants it can support, identifiers indicating that computing resource 415 supports multi-tenancy, etc.
[0262] At 1820, method 1800 may include host 410 determining a performance level to be assigned to each of one or more tenants. For example, host 410 may determine the ability to use and / or allocate the indicated computing resources 415.
[0263] At 1825, method 1800 may include host 410 sending management commands (e.g., multitenancy setup commands) to computing resource 415. In some cases, host 410 may send management commands to computing resource 415 to establish performance levels for one or more tenants. In some cases, host 410 may send management commands via PF0.
[0264] At 1830, method 1800 may include computing resource 415 parsing management commands. For example, computing resource 415 may receive management commands, identify the fields of the management commands, and store data values in each field.
[0265] At 1835, method 1800 may optionally include computing resource 415 sending a rejection message to host 410. In some examples, host 410 may receive the rejection message from computing resource 415 based on an error or defect detected in a management command by computing resource 415. In some cases, host 410 may send a first management command for a first tenant and a second management command for a second tenant. In some cases, host 410 may send the first and second management commands as separate messages or as an aggregated management command for multiple tenants. In response, host 410 may receive from computing resource 415 an acknowledgment for the first tenant and a rejection message for the second tenant, the acknowledgment indicating that the performance level for the first tenant is accepted and is being implemented, and the rejection message indicating that the request for the second tenant is rejected based on an error detected in the corresponding request. For example, computing resource 415 may not decrypt one of the fields in the second management command. For example, rejection may be based on an error that a field does not have a value, an error that a field has an indeterminate value, an error that a field has an out-of-bounds value (e.g., a value exceeding a value threshold for the field), a transmission error that corrupts one or more fields, etc. In some cases, computing resource 415 may determine that it does not have sufficient granularity and / or sufficient performance capacity to meet the performance level requested by the second tenant (e.g., the requested performance level exceeds the performance availability threshold of computing resource 415).
[0266] At 1840, method 1800 may include computing resource 415 implementing corresponding management commands for each of one or more tenants. In some examples, computing resource 415 may be configured with internal aspects to satisfy multi-tenancy settings requested by host 410 via management commands.
[0267] At 1845, method 1800 may include computing resource 415 sending an acknowledgment to host 410. For example, computing resource 415 may indicate (e.g., via PF0) that the implementation of a management command was successful. If computing resource 415 cannot satisfy or complete the requested performance, and / or if host 410 requests something that violates the capabilities of computing resource 415, computing resource 415 may indicate an error in the acknowledgment instead of indicating successful implementation. In some cases, computing resource 415 may send an acknowledgment that includes at least one acknowledgment and / or at least one rejection. For example, host 410 may receive from computing resource 415 a message that includes an acknowledgment for a first tenant and a rejection message for a second tenant.
[0268] At 1850, method 1800 may include computing resource 415 providing a best-effort performance level to each of one or more tenants.
[0269] Figure 19 A flowchart illustrating an example method 1900 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, method 1900 may be... Figure 1 Multi-rental controller 140 and / or Figure 2 The multi-tenant controller 230 is implemented. In some configurations, method 1900 may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1900 is merely one implementation, and one or more operations of method 1900 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0270] At 1905, method 1900 may include host 410 determining a change in the performance level of one or more tenants. For example, host 410 may determine that the performance level of at least one tenant needs to be changed. In some cases, host 410 may detect changes in operating conditions (e.g., a tenant needs more performance, less performance, etc.), the addition and / or removal of one or more tenants relative to computing resources (e.g., net addition, net reduction of tenants), etc. Therefore, host 410 may determine how to use and / or allocate the capacity of computing resource 415 based on the changed conditions. In some examples, host 410 may store the capacity response of computing resource 415 and determine how to use and / or allocate the capacity of computing resource 415 based on the changed conditions and / or the stored capacity response.
[0271] At 1910, method 1900 may include host 410 determining a performance level to be assigned to each of one or more tenants based on updated conditions. In some examples, host 410 may determine the capability of an indication of how to use and / or allocate computing resources 415 based on updated conditions. For example, host 410 may store capability responses and refer to the stored capability responses to determine the performance level of a tenant (e.g., an updated performance level).
[0272] At 1915, method 1900 may include host 410 sending an update management command (e.g., a multitenancy setting command) to computing resource 415. In some cases, host 410 may send an update management command to computing resource 415 to update the performance levels of one or more tenants. In some cases, host 410 may send the update management command via PF0.
[0273] At 1920, method 1900 may include computing resource 415 resolving update management commands and implementing corresponding update management commands for each of one or more tenants. In some examples, computing resource 415 may be configured with internal aspects to satisfy multi-tenancy settings requested by host 410 via update management commands.
[0274] At 1925, method 1900 may include computing resource 415 sending an acknowledgment to host 410. For example, computing resource 415 may indicate (e.g., via PF0) that the implementation of an update management command was successful. If computing resource 415 cannot meet or complete the requested performance, and / or if host 410 requests something that violates the capabilities of computing resource 415, computing resource 415 may indicate an error in the acknowledgment instead of indicating successful implementation.
[0275] At 1930, method 1900 may include computing resource 415 providing a best-effort performance level to each of one or more tenants. In the examples described herein, the configurations and operations are example configurations and operations, and various additional configurations and operations not explicitly shown may be involved. In some examples, one or more aspects of the configurations and / or operations shown may be omitted. In some embodiments, one or more of the operations may be performed by components other than those shown herein.
[0276] Additionally or alternatively, the order and / or timing of operations may be altered. Some embodiments may be implemented in one or a combination of hardware, firmware, and software. Other embodiments may be implemented as instructions stored on a computer-readable storage device that can be read and executed by at least one processor to perform the operations described herein. A computer-readable storage device may include any non-transitory memory mechanism for storing information in a machine-readable (e.g., computer) form. For example, a computer-readable storage device may include read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, and other storage devices and media.
[0277] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. As used herein, the terms "computing device," "user equipment," "communication station," "station," "handheld device," "mobile device," "wireless device," and "user equipment (UE)" refer to wireless communication devices such as cellular phones, smartphones, tablets, netbooks, wireless terminals, laptops, femtocells, high data rate (HDR) subscriber stations, access points, printers, point-of-sale equipment, access terminals, or other personal communication system (PCS) devices. Such devices may be mobile or stationary.
[0278] As used herein, the term "communication" is intended to include sending, or receiving, or both. This may be particularly useful in the claims when describing an organization of data sent by one device and received by another device, but only the functionality of one of those devices is required to infringe the claims. Similarly, a two-way exchange of data between two devices (where both devices send and receive during the exchange) can be described as "communication" when only the functionality of one of those devices is claimed. The term "transmission," as used herein with respect to wireless communication signals, includes sending and / or receiving wireless communication signals. For example, a wireless communication unit capable of transmitting wireless communication signals may include a wireless transmitter that sends wireless communication signals to at least one other wireless communication unit, and / or a wireless communication receiver that receives wireless communication signals from at least one other wireless communication unit.
[0279] Some embodiments can be used with a variety of devices and systems, such as personal computers (PCs), desktop computers, mobile computers, laptop computers, notebook computers, tablet computers, server computers, handheld computers, handheld devices, personal digital assistant (PDA) devices, handheld PDA devices, external devices, external devices, hybrid devices, in-vehicle devices, non-in-vehicle devices, mobile or portable devices, consumer devices, non-mobile or non-portable devices, wireless communication stations, wireless communication devices, wireless access points (APs), wired or wireless routers, wired or wireless modems, video devices, audio devices, audio-video (A / V) devices, wired or wireless networks, wireless local area networks, wireless video local area networks (WVANs), local area networks (LANs), wireless LANs (WLANs), personal area networks (PANs), wireless PANs (WPANs), etc.
[0280] Some embodiments can be used in conjunction with one-way and / or two-way radio communication systems, cellular wireless telephone communication systems, mobile phones, cell phones, wireless phones, personal communication system (PCS) devices, PDA devices that include wireless communication devices, mobile or portable global positioning system (GPS) devices, devices that include GPS receivers or transceivers or chips, devices that include RFID elements or chips, multiple-input multiple-output (MIMO) transceivers or devices, single-input multiple-output (SIMO) transceivers or devices, multiple-input single-output (MISO) transceivers or devices, devices with one or more internal antennas and / or external antennas, digital video broadcasting (DVB) devices or systems, multi-standard wireless devices or systems, wired or wireless handheld devices (e.g., smartphones), wireless application protocol (WAP) devices, etc.
[0281] Some embodiments can be used in conjunction with one or more types of wireless communication signals and / or systems that conform to one or more wireless communication protocols, such as radio frequency (RF), infrared (IR), frequency division multiplexing (FDM), orthogonal FDM (OFDM), time division multiplexing (TDM), time division multiple access (TDMA), extended TDMA (E-TDMA), General Packet Radio Service (GPRS), extended GPRS, code division multiple access (CDMA), wideband CDMA (WCDMA), CDMA 2000, single-carrier CDMA, multi-carrier CDMA, multi-carrier modulation (MDM), discrete multi-tone (DMT), and Bluetooth. TM Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee TMUltra-wideband (UWB), Global System for Mobile Communications (GSM), 2G, 2.5G, 3G, 3.5G, 4G, fifth-generation (5G) mobile networks, 3GPP, Long Term Evolution (LTE), LTE Advanced, Enhanced Data Rate Evolution of GSM (EDGE), etc. Other embodiments can be used in a variety of other devices, systems, and / or networks.
[0282] Although an example processing system has been described herein, embodiments of the subject matter and functional operation described herein may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed herein and their structural equivalents), or in a combination of one or more of them.
[0283] The embodiments of the subject matter and operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs encoded on a computer storage medium, i.e., one or more components of computer program instructions for execution by or control of an information / data processing device. Alternatively or additionally, program instructions can be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information / data for transmission to a suitable receiver device for execution by the information / data processing device. The computer storage medium can be or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or combinations thereof. Furthermore, while the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included in one or more separate physical components or media.
[0284] The operations described herein can be implemented as operations performed by an information / data processing device on information / data stored on one or more computer-readable storage devices or received from other sources.
[0285] The term "data processing apparatus" includes all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or a combination of the foregoing. The apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0286] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or information / data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more components, subroutines, or code portions). Computer programs can be deployed to execute on one computer or on multiple computers located at a site or distributed across multiple sites and interconnected through a communication network.
[0287] The processes and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input information / data and generating output. Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and information / data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive information / data from or transfer information / data to one or more mass storage devices, or both. However, a computer does not need to have such devices. Suitable devices for storing computer program instructions and information / data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by dedicated logic circuits or incorporated into dedicated logic circuits.
[0288] To provide interaction with the user, embodiments of the subject matter described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information / data to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, voice, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser in response to a request received from a web browser on the user's client device.
[0289] Embodiments of the subject matter described herein can be implemented in a computing system that includes backend components (e.g., as an information / data server), or middleware components (e.g., an application server), or frontend components (e.g., a client computer with a graphical user interface or web browser through which a user can interact with embodiments of the subject matter described herein), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital information / data communication of any form or medium, such as a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (e.g., the Internet) and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).
[0290] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship is established by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends information / data (e.g., HTML pages) to the client device (e.g., for the purpose of displaying information / data to a user interacting with the client device and receiving user input from the user interacting with the client device). Information / data generated at the client device (e.g., the result of user interaction) may be received at the server from the client device.
[0291] While this specification contains numerous details of specific embodiments, these should not be construed as limiting the scope of any embodiment or potentially claimed content, but rather as descriptions of features specific to particular embodiments. Certain features described herein in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described herein as functioning in certain combinations and even initially claimed in this way, one or more features from a claimed combination may be removed from the combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0292] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all of the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described herein should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0293] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
[0294] Benefiting from the teachings presented in the foregoing description and the accompanying drawings, those skilled in the art will conceive of numerous modifications and other examples set forth herein. Therefore, it should be understood that the embodiments are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terminology is used herein, it is used only in a general and descriptive sense and not for limiting purposes.
Claims
1. A method for multiple rentals, the method comprising: Identifier for the first tenant of the storage device; Assign the first performance level to the first tenant; Generate a first performance parameter based on the first performance level; and Send a configuration message to the storage device, including the first performance parameter and the identifier of the first tenant.
2. The method according to claim 1, wherein, The identifier of the first tenant includes the virtual function of the storage device and the recycling unit handle of the storage device.
3. The method according to claim 1, wherein, The identifier of the first tenant includes at least one of the following: the physical function of the storage device, the port of the storage device, the stream of the storage device, the zone of the storage device, the logical block address range of the storage device, the non-volatile memory (NVM) controller of the storage device, the commit queue, or the scalable input / output virtualization.
4. The method according to claim 1, wherein, The identifier of the first tenant includes at least one of the following: compute storage thread, graphics processing unit (GPU), a portion of a GPU, field-programmable gate array, namespace of the storage device, durability group of the storage device, recycling group of the storage device, command type associated with a command generated by the first tenant, or command identifier associated with a command generated by the first tenant.
5. The method according to claim 1, wherein, The configuration message includes one or more tenant identifier fields.
6. The method according to claim 1, wherein, The configuration message includes one or more performance parameter fields based on the first performance level, the one or more performance parameter fields including at least one field for the maximum allowed input / output operations per second (IOPS) for the first tenant on the storage device or for the reservation level of the first tenant's IOPS on the storage device.
7. The method according to claim 1, wherein, The configuration message includes one or more performance parameter fields based on the first performance level, the one or more performance parameter fields including at least one field for the maximum available communication bandwidth between the first tenant and the storage device or for the reservation level of the communication bandwidth between the first tenant and the storage device.
8. The method according to claim 1, wherein: The configuration message includes a performance parameter field for the tenant load of the first tenant, the performance parameter field indicating the performance level of the request based on the first performance level. The tenant load of the first tenant includes at least one of the queue depth (QD) of commands generated by the first tenant or the command length associated with the commands generated by the first tenant.
9. The method according to claim 1, wherein, The configuration message includes one or more performance parameter fields for at least one of the following: the first tenant's proportional bandwidth win rate relative to the bandwidth win rate of at least one other tenant, the first tenant's proportional IOPS win rate relative to the IOPS win rate of the at least one other tenant, the maximum level of change in access to the storage device over a period of time, the access consistency level between the first tenant and the storage device, or the maximum allowed access latency between the first tenant and the storage device.
10. The method according to claim 1, further comprising: Identifier for the second tenant; and Assign a second performance level, which is different from the first performance level, to the second tenant.
11. The method of claim 10, further comprising generating a second performance parameter based on the second performance level, wherein: The second performance parameter is different from the first performance parameter, and The configuration message includes the second performance parameter and the identifier of the second tenant.
12. The method according to claim 1, wherein, The configuration message is formatted based on the fast packet format of non-volatile memory.
13. The method according to claim 1, wherein, The storage device includes a solid-state drive.
14. An apparatus comprising: At least one memory; and At least one processor coupled to the at least one memory is configured to: Identifier for the first tenant of the storage device; Assign the first performance level to the first tenant; Generate a first performance parameter based on the first performance level; and Send a configuration message to the storage device, including the first performance parameter and the identifier of the first tenant.
15. The apparatus of claim 14, further comprising: Identifier for the second tenant; and Assign a second performance level, which is different from the first performance level, to the second tenant.
16. The device of claim 15, wherein the at least one processor is configured to generate a second performance parameter based on the second performance level, wherein: The second performance parameter is different from the first performance parameter, and The configuration message includes the second performance parameter and the identifier of the second tenant.
17. The device according to claim 14, wherein, The identifier of the first tenant includes the virtual function of the storage device and the recycling unit handle of the storage device.
18. A non-transitory computer-readable medium storing code, said code comprising instructions executable by a processor of a device to perform the following operations: Identifier for the first tenant of the storage device; Assign the first performance level to the first tenant; Generate a first performance parameter based on the first performance level; and Send a configuration message to the storage device, including the first performance parameter and the identifier of the first tenant.
19. The non-transitory computer-readable medium according to claim 18, wherein, The code also includes instructions executable by the processor to cause the device to perform the following operations: Identifiers for second tenants; and Assign a second performance level, which is different from the first performance level, to the second tenant.
20. The non-transitory computer-readable medium according to claim 19, wherein, The code also includes instructions executable by the processor to cause the device to generate a second performance parameter based on the second performance level, wherein: The second performance parameter is different from the first performance parameter, and The configuration message includes the second performance parameter and the identifier of the second tenant.