FAIRNESS IN POWER CONSUMPTION IN MULTI-HOST SYSTEMS
By allocating bandwidth in SSD systems based on power consumption rather than data transfer, the method addresses the unfairness in multi-host environments, enhancing system performance and sustainability through a credit-based scheduler.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-03-26
AI Technical Summary
Existing bandwidth allocation methods in multi-host SSD systems fail to ensure fairness and quality of service due to inaccurate assumptions about resource consumption, as they rely solely on the amount of data transferred, neglecting variations in power consumption caused by factors like encryption, decryption, and different workloads.
Implement a fairness control mechanism that allocates bandwidth based on power consumption per command, using a credit-based scheduler to ensure fair distribution of resources across multiple hosts, considering the actual power requirements of each operation.
This approach achieves more accurate and adaptable fairness by ensuring that resources are distributed based on actual power consumption, leading to improved system performance and sustainability, even under varying workloads and data path configurations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF REVELATION Area of Revelation
[0001] Embodiments of the present disclosure generally relate to the improvement of bandwidth allocation in solid-state drives (SSDs). Description of the state of the art
[0002] Quality of Service (QoS) is a parameter in a multi-user system that allows device bandwidth administrators to ensure application throughput, enabling transaction processing within an acceptable timeframe. QoS offers several key benefits, including the ability to grant devices control over bandwidth resources and manage the network using different priorities and bandwidth allocations for each user. Managing bandwidth allocation ensures that time-critical and business-critical applications receive the resources they need, while also allowing other applications to access the available resources. The overall result is an improved user experience and reduced costs through the efficient use of existing resources.
[0003] In a multi-host environment with non-volatile storage (NVM) express (NVMe) and dozens of host applications, a QoS and fairness algorithm is used. This algorithm ensures fairness within the system by guaranteeing that each host application receives a fair share of resources and performance. By implementing an effective QoS and fairness algorithm, the NVMe storage system can optimize resource utilization, prevent conflicts, and deliver balanced and reliable performance across various host applications. The algorithm assumes that transferring the same number of bytes consumes equal resources and performance, which is only true for very specific workloads.
[0004] There is a need in technology for an improvement in bandwidth allocation in SSDs.
[0005] US 2016 / 0 370 841 A1 concerns a storage system and a method for power management using a token bucket that indicates power available for storage operations. US 2023 / 0 393 877 A1 concerns dynamic arbitration mechanisms. US 2020 / 0 209 944 A1 concerns devices and methods for arbitrating operations in a NAND memory system. SUMMARY OF THE REVELATION
[0006] According to the invention, a data storage device is provided having the features of independent claim 1; dependent claims relate to preferred embodiments.
[0007] Instead of allocating bandwidth to different host devices or functions in a multi-device system based on the amount of data transferred, bandwidth is allocated based on the power consumed to execute a command. While two different commands might transfer the same amount of data, they might consume different amounts of power. For example, one command might use encryption, while the other might not. Using encryption requires more power than not using encryption. By considering the power consumption per command, bandwidth allocation can be fair to ensure high Quality of Service (QoS).
[0008] In one embodiment, a data storage device comprises: a storage device; and a controller coupled to the storage device, the controller being configured to: retrieve a first instruction from a first memory location; evaluate the expected power utilization for the first instruction; determine whether the first memory location has sufficient credits to execute the first instruction; and either: execute the first instruction immediately; or hold the first instruction, wait until sufficient credits are allocated to execute the first instruction, and execute the first instruction after the waiting period.
[0009] In another embodiment, a data storage device comprises: a storage device; and a controller coupled to the storage device, wherein the controller includes a fairness control module and wherein the fairness control module is configured to: retrieve one or more instructions from one or more transmission queues (SQs) with an arbiter; manage a power consumption database; determine a power allocation for the one or more instructions based on information from the power consumption database; and determine whether the one or more instructions should be executed or whether the execution of the one or more instructions should be delayed.
[0010] In another embodiment, a data storage device comprises: means for storing data; and a controller coupled to the means for storing data, the controller being configured to: execute a first instruction that transfers a first plurality of bytes to the means for storing data using a first set of parameters and consuming a first quantity of power; execute a second instruction that transfers a second plurality of bytes to the means for storing data using a second set of parameters and consuming a second quantity of power, wherein the first plurality is equal to the second plurality, wherein the first set of parameters differs from the second set of parameters, and wherein the first plurality is smaller than the second plurality;and determine when the first and second instructions should be executed, based on the power consumption credits allocated in relation to the first and second quantities. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To clarify in detail how the aforementioned features of the present disclosure are to be understood, a more detailed description of the disclosure, which has been briefly summarized above, follows, with reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings only illustrate typical embodiments of this disclosure and are therefore not to be regarded as limiting its scope of protection, since the disclosure may also permit other, equally effective embodiments. Fig.Figure 1 is a schematic block diagram illustrating a storage system in which a data storage device can function as a storage device for a host device according to certain embodiments. Fig. Figure 2 is a schematic diagram illustrating a bandwidth limiter according to certain embodiments. Fig. Figure 3 is a schematic diagram illustrating the storage device according to one embodiment, with a focus on the data path components. Fig. Figure 4 is a schematic diagram illustrating the performance fairness control according to one embodiment. Fig. Figure 5 is a schematic representation of a system that includes a fairness control module according to one embodiment. Fig. Figure 6 is a flowchart illustrating a bandwidth allocation procedure according to one embodiment.
[0012] To facilitate understanding, identical reference numerals have been used wherever possible to denote identical elements present in all figures. It is assumed that elements disclosed in one embodiment can also be advantageously used in other embodiments without specific mention. DETAILED DESCRIPTION
[0013] The following refers to embodiments of the disclosure. It is understood, however, that the disclosure is not limited to the specific embodiments described. Instead, any combination of the following features and elements, regardless of whether they relate to different embodiments or not, is intended for the implementation and practical application of the disclosure. Furthermore, although embodiments of the disclosure may offer advantages over other possible solutions and / or over the prior art, the fact that a particular advantage is achieved by a given embodiment does not constitute a limitation of the disclosure.Therefore, the following aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly stated in one or more claims. Likewise, a reference to "the disclosure" is not to be construed as a generalization of any subject matter disclosed herein and is not to be considered part of or a limitation of the appended claims unless expressly stated in one or more claims.
[0014] Instead of allocating bandwidth to different host devices or functions in a multi-device system based on the amount of data transferred, bandwidth is allocated based on the power consumed to execute a command. While two different commands might transfer the same amount of data, they might consume different amounts of power. For example, one command might use encryption, while the other might not. Using encryption requires more power than not using encryption. By considering the power consumption per command, bandwidth allocation can be fair to ensure high Quality of Service (QoS).
[0015] The present disclosure addresses the QoS and fairness challenges in a multi-host environment by introducing a method that offers several advantages over previous approaches. The embodiments generally relate to a system with multiple hosts, but are also applicable to a system comprising a single host with several subclients under the host, such as virtual functions. The motivation is to achieve QoS and fairness among these functions and / or hosts.
[0016] Fig.Figure 1 is a schematic block diagram illustrating a storage system 100 with a data storage device 106, which, according to certain embodiments, can function as a storage device for a host device 104. For example, the host device 104 can utilize non-volatile memory (NVM) 110 enclosed in data storage device 106 for storing and retrieving data. The host device 104 includes dynamic random-access memory (DRAM) 138. In some examples, the storage system 100 can include a plurality of storage devices, such as the data storage device 106, which can function as a storage array. For example, the storage system 100 can include a plurality of data storage devices 106 configured as a redundant array of low-cost / independent disks (RAID) and collectively function as a mass storage device for the host device 104.
[0017] The host device 104 can store data on and / or retrieve data from one or more storage devices, such as the data storage device 106. As described in Fig. As illustrated in Figure 1, the host device 104 can communicate with the data storage device 106 via an interface 114. The host device 104 can encompass a wide range of devices, including computer servers, network attached storage (NAS) units, desktop computers, notebooks (i.e., laptops), tablet computers, set-top boxes, and mobile phone handsets such as so-called "smartphones". t phones”, so-called “Smart Pads”, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices or other devices capable of sending data to or receiving data from a data storage device.
[0018] The host DRAM 138 can optionally include a host memory buffer (HMB) 150. The HMB 150 is a section of the host DRAM 138 that is allocated to the data storage device 106 for the exclusive use of a controller 108 of the data storage device 106. For example, the controller 108 can store mapping data, buffered instructions, logic-physical (L2P) tables, metadata, and the like in the HMB 150. In other words, the HMB 150 can be used by the controller 108 to store data that would normally be stored in volatile memory 112, a buffer 116, internal memory of the controller 108 such as static random-access memory (SRAM), and the like. In examples where the data storage device 106 does not include DRAM (i.e., optional DRAM 118), the controller 108 can use the HMB 150 as the DRAM of the data storage device 106.
[0019] The data storage device 106 includes the controller 108, NVM 110, a power supply 111, volatile memory 112, the interface 114, a write buffer 116, and optional DRAM 118. In some examples, the data storage device 106 may include additional components, which for clarity are shown in Fig.Figure 1 is not shown. For example, the data storage device 106 may include a printed circuit board (PCB) to which components of the data storage device 106 are mechanically attached and which includes electrically conductive traces that electrically connect components of the data storage device 106 or the like. In some examples, the physical dimensions and connection configurations of the data storage device 106 may conform to one or more standard form factors. Some examples of standard form factors include, but are not limited to, 3.5-inch data storage devices (e.g., an HDD or SSD), 2.5-inch data storage devices, 1.8-inch data storage devices, Peripheral Component Interconnect (PCI), PCI-Extended (PCI-X), and PCI Express (PCIe) (e.g., PCIe x1, x4, x8, x16, PCIe Mini Card, MiniPCI, etc.).In some examples, the data storage device 106 can be directly coupled to a mainboard of the host device 104 (e.g., directly soldered or plugged into a connector).
[0020] Interface 114 can include a data bus for data exchange with the host device 104 and / or a control bus for exchanging commands with the host device 104. Interface 114 can operate according to any suitable protocol. For example, interface 114 can operate according to one or more of the following protocols: Advanced Technology Attachment (ATA) (e.g., Serial ATA (SATA) and Parallel ATA (PATA)), Fibre Channel Protocol (FCP), Small Computer System Interface (SCSI), Serially Attached SCSI (SAS), PCI and PCIe, Non-Volatile Memory Express (NVMe), OpenCAPI, GenZ, Cache Coherent Interface Accelerator (CCIX), Open Channel SSD (OCSSD), or the like. Interface 114 (e.g.,The data bus, the control bus, or both are electrically connected to the controller 108, providing an electrical connection between the host device 104 and the controller 108 so that data can be exchanged between them. In some examples, the electrical connection of interface 114 of the data storage device 106 can also allow it to receive power from the host device 104. For example, in... Fig. As illustrated in Figure 1, the power supply 111 can receive power from the host device 104 via interface 114.
[0021] The NVM 110 can include a variety of storage devices or storage units. The NVM 110 can be configured to store and / or retrieve data. For example, a storage unit of the NVM 110 can receive data and a message from Controller 108 instructing the storage unit to store the data. Similarly, the storage unit can receive a message from Controller 108 instructing the storage unit to retrieve data. In some examples, each of the storage units can be referred to as a die. In some examples, the NVM 110 can include a variety of dies (i.e., a variety of storage units). In some examples, each storage unit can be configured to store relatively large amounts of data (e.g., 128 MB, 256 MB, 512 MB, 1 GB, 2 GB, 4 GB, 8 GB, 16 GB, 32 GB, 64 GB, 128 GB, 256 GB, 512 GB, 1 TB, etc.).
[0022] In some examples, each storage unit can include any type of non-volatile storage device, such as flash memory devices, phase-change memory (PCM) devices, resistive random access memory (ReRAM) devices, magnetoresistive random access memory (MRAM) devices, ferroelectric random access memory (FRAM), holographic storage devices, and any other type of non-volatile storage device.
[0023] The NVM 110 can incorporate a variety of flash memory devices or storage units. NVM flash memory devices can include NAND- or NOR-based flash memory devices and can store data based on a charge contained in a floating gate of a transistor for each flash memory cell. In NVM flash memory devices, the flash memory device can be subdivided into a variety of dies, with each die of the variety containing a variety of physical or logical blocks, which can be further subdivided into a variety of pages. Each block of the variety of blocks within a particular memory device can contain a variety of NVM cells. Rows of NVM cells can be electrically connected using a word line to define a page from a variety of pages.Each cell in the multitude of pages can be electrically connected to its respective bit line. Furthermore, NVM flash storage devices can be 2D or 3D and of type: Single Level Cell (SLC), Multi-Level Cell (MLC), Triple Level Cell (TLC), or Quad Level Cell (QLC). The Controller 108 can write and read data to and from NVM flash storage devices at the page level and erase data from NVM flash storage devices at the block level.
[0024] Power supply 111 can provide power to one or more components of the data storage device 106. In standard mode, power supply 111 can provide power to one or more components using power supplied by an external device, such as the host device 104. For example, power supply 111 can provide power to one or more components using power received from the host device 104 via interface 114. In some examples, power supply 111 can include one or more power storage components configured to provide power to the one or more components when they are in shutdown mode, such as when no more power is being received from the external device.In this way, the power supply 111 can function as an integrated backup power source. Some examples of the one or more power storage components include, but are not limited to, capacitors, supercapacitors, batteries, and the like. In some examples, the amount of power that can be stored by the one or more power storage components can be a function of the cost and / or size (e.g., area / volume) of the one or more power storage components. In other words, as the amount of power stored by one or more power storage components increases, so do the cost and / or size of the one or more power storage components.
[0025] The volatile memory 112 can be used by the controller 108 to store information. The volatile memory 112 can include one or more volatile storage devices. In some examples, the controller 108 can use the volatile memory 112 as a cache. For example, the controller 108 can store cached information in the volatile memory 112 until the cached information is written to the NVM 110. As in Fig.As illustrated in Figure 1, volatile memory 112 can consume power received from the power supply 111. Examples of volatile memory 112 include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static RAM (SRAM), and synchronous dynamic RAM (SDRAM (e.g., DDR1, DDR2, DDR3, DDR3L, LPDDR3, DDR4, LPDDR4, and the like)). Likewise, optional DRAM 118 can be used to store mapping data, buffered instructions, logic-physical (L2P) tables, metadata, cached data, and the like. In some examples, the data storage device 106 does not include optional DRAM 118, so the data storage device 106 has no DRAM. In other examples, the data storage device 106 includes the optional DRAM 118.
[0026] The controller 108 can manage one or more operations of the data storage device 106. For example, the controller 108 can manage reading data from and / or writing data to the NVM 110. In some embodiments, when the data storage device 106 receives a write command from the host device 104, the controller 108 can initiate a data storage command to store data in the NVM 110 and monitor the progress of the data storage command. The controller 108 can determine at least one operating property of the storage system 100 and store at least one operating property in the NVM 110. In some embodiments, when the data storage device 106 receives a write command from the host device 104, the controller 108 temporarily stores the data associated with the write command in internal memory or write buffer 116 before sending the data to the NVM 110.The controller 108 can include switching logic or processors configured to run programs to operate the data storage device 106.
[0027] The Controller 108 can include an optional second volatile memory 120. The optional second volatile memory 120 can be similar to the volatile memory 112.
[0028] For example, the optional second volatile memory 120 can be SRAM. The controller 108 can allocate a portion of the optional second volatile memory to the host device 104 as a controller memory buffer (CMB) 122. The CMB 122 can be accessed directly by the host device 104. For example, instead of managing one or more transmission queues in the host device 104, the host device 104 can use the CMB 122 to store the one or more transmission queues that are normally managed in the host device 104. In other words, the host device 104 can generate instructions and store the generated instructions, with or without the associated data, in the CMB 122, with the controller 108 accessing the CMB 122 to retrieve the stored generated instructions and / or the associated data.
[0029] The previous approach was based on performance fairness or the number of bytes transferred within a fixed period of time. Fig. Figure 2 illustrates the core concept of the algorithm. Fig. Figure 2 is a schematic diagram illustrating a bandwidth limiter 200 according to certain embodiments. Upon receiving a command, the command is associated with the corresponding bandwidth limiter vector, and the current bandwidth limiter counter is decremented based on the size of the command. When the counter falls below the lower threshold, the corresponding transmit queue (SQ) identification (ID) is disabled for subsequent retrieval operations until the bandwidth permits retrieval.
[0030] The logic regularly scans the bandwidth limiter groups and allocates bandwidth to each of them. When the upper threshold of the counter is exceeded, all previously deactivated SQ-IDs associated with the vector are reactivated.
[0031] The main limitation of this approach lies in its accuracy. The concept in Fig. The underlying assumption is that transferring X bytes consumes uniform resources and power. However, this assumption proves inaccurate, as it depends on various parameters such as encryption / decryption, CMB, command-specific LDPC decoder operation, and more. The algorithm does not account for the diverse data path configurations, which can vary from command to command. Consequently, this approach only guarantees bandwidth for very specific workloads.
[0032] Fig.Figure 2 illustrates the concept of bandwidth limiting, where all fairness is based on performance. Performance is measured for each individual function, virtual function, or physical function in the system. Each function has a performance requirement, and the controller ensures that the performance for each function does not exceed this requirement. For example, if the target for host number one is one gigabyte per second, the algorithm measures its performance. If the algorithm determines that a command for host number one exceeds the maximum performance of one gigabyte, the controller executes a command throttling mechanism so that no further commands are requested from host one until sufficient credits are available.If there are not enough credits available, the respective host already has maximum performance and must therefore wait to execute the command until more time has passed and host one has sufficient credits, and only then will the command be retrieved for host one.
[0033] The in Fig. The credit system used does not need to be linked to a single host but can be applied to multiple grouped hosts. The controller ensures that the total bandwidth of all group members does not exceed one gigabyte. Bandwidth is allocated to all hosts or functions associated with the respective group based on performance. The controller would measure performance and ensure fairness within the system by allocating or limiting bandwidth accordingly.
[0034] As mentioned above, the problem is that even if the controller ensures performance, allocating up to one gigabyte per second to each host (for example, with four hosts), the apparent fairness is not actually guaranteed. Fairness is lacking because, for instance, the first host might have a different workload than the others. For example, the first host might have a random read workload. A random read workload would consume more resources to complete tasks compared to a sequential read workload. Even if the performance for random and sequential reads is the same compared to the other hosts, there is still no fairness from a performance perspective, despite the identical performance. To achieve the same random read performance as the other hosts' sequential read performance, the first host requires more resources.
[0035] As another example, all hosts might have the same workload, such as sequential reading, but host one has a very high bit error rate (BER). To achieve the same performance for all hosts, host one, due to its high BER, requires more processing power. As yet another example, encryption / decryption might be enabled for one host but not for the others. Enabling encryption / decryption also requires more processing power to achieve the same performance. Therefore, performance is not always a good metric for comparison.
[0036] The disclosure includes a novel approach to addressing the fairness challenge in a multi-host storage device. Instead of relying on a fairness algorithm based on performance and quality of service, the idea is to implement fair power allocation among multiple hosts. The proposal introduces a credit-based power scheduler that ensures fairness based on each host's actual power consumption per operation. This approach is more accurate than previous methods because fairness in performance does not necessarily guarantee fairness in the overall system. Different parameters in the data path can lead to diverse demands on shared resources and power consumption. Because fairness is based on power consumption, the method becomes more accurate and applicable to a wide range of workloads and applications.
[0037] The multi-host environment can take various forms, including: PCIe multiports – with multiple PCIe connections connected to multiple hosts; multiple physical functions sharing the same connection (multi-PF); multiple virtual functions sharing the same link (SR-IOV); multiple NVMe namespaces; multiple delivery queues; and multiple endurance groups. These multi-host options are merely examples, as other possibilities are also considered.
[0038] Power consumption in SSDs per instruction is affected by various factors, including optional pipeline stages such as DRAM, CMB, encryption and decryption operations, and low-density decoder modes with parity checking (LDPC). The following is a breakdown of how each stage can impact power consumption.
[0039] In the DRAM read path, power is consumed when data is retrieved from the NAND flash memory and loaded into the DRAM. The energy required for DRAM read operations contributes to the overall power consumption per read instruction. In the write path, power is consumed when data is temporarily stored in the DRAM by the host before being programmed into the NAND flash memory. DRAM write operations contribute to the power consumption per write instruction. More precisely, there might be a path that uses DRAM as part of the data path and therefore requires more power to achieve the same transfer performance as paths that do not use DRAM.
[0040] When CMB is used in the read path to buffer frequently accessed data, the power required for CMB reads increases the power consumption per read instruction. In the write path, using CMB as a buffer for frequently accessed data results in power consumption during CMB writes, contributing to the overall power consumption per write instruction. More precisely, when CMB is enabled for a particular instruction, its use requires more performance because the data is being copied. Typically, CMB is implemented in DRAM, requiring more overhead and therefore incurring higher power consumption.
[0041] During encryption and decryption operations in the read path, the energy consumed during the decryption process contributes to the power consumption per read command if the data on the NAND flash is encrypted and needs to be decrypted during a read operation. Similarly, if encryption is enabled for data by the host during a write operation in the write path, the power used during the encryption process contributes to the overall power consumption per write command. More precisely, encryption / decryption can consume more power when enabled than when it is disabled.
[0042] In LDPC decoder / encoder modes, LDPC decoding in the read path contributes to the power consumption per read instruction. The power consumed depends on the actual BER (Read-Only Rate). In the write path, LDPC encoding, which is used to add error correction codes during a write operation, increases the power consumption per write instruction. More precisely, if LDPC is enabled due to a high BER, the controller must expend more power and effort to complete a task.
[0043] To protect RAID, additional calculations are required in the read path of RAID configurations, especially those with parity-based RAID levels like RAID 5 or RAID 6, to check and reconstruct data from parity information during read operations. These calculations, which are part of the RAID protection scheme, contribute to increased power consumption per read command. Similarly, parity updates or calculations may be required in the write path during write operations in RAID configurations. More precisely, when RAID is enabled, using RAID requires more power and overhead to complete a task than when RAID is disabled.
[0044] The idea, therefore, is to base the fairness algorithm not on performance, but on output, which leads to better and more accurate results from a fairness perspective.
[0045] In summary, each optional pipeline stage in the SSD data path introduces additional power consumption per command. The specific impact varies depending on the architecture, features, and usage of these optional stages during SSD read and write operations. Relying solely on performance for the fairness algorithm is not a good measure, as the performance of one host might consume X power, while the same performance on another host might consume Y power, where X ≠ Y.
[0046] In other words, the idea is not to compare performance to achieve fairness, but rather to achieve a fair distribution of resources, as such a distribution would be more accurate from a fairness perspective. The process determines how much time is required to complete the task. The aim is to achieve fairness in terms of performance, resource consumption, and resource allocation.
[0047] Fig.Figure 3 illustrates the high-level block diagram of the storage device, focusing on the data path components. Optional components are color-coded. Basic read / write data transfer utilizes the host interface, the NAND interface, and the LDPC. However, several additional operations are part of the data path. These are optional, depend on host requirements, and may occur per command. These operations include DRAM manipulation, security operations, RAID protection, and more. The LDPC decoder supports multiple modes depending on the BER, with each mode consuming a different amount of power. The power consumption of a command depends on the selected operations and is not fixed.
[0048] The System 300 includes the LDPC encoder module, the RAID module in the write path, the encryption module in the write path, the decryption module in the read path, the LDPC decoder module in the read path, and the RAID module in the read path. Depending on the command, any module can be bypassed or disabled, which changes the power consumption.
[0049] Fig. Figure 4 illustrates the concept for achieving performance fairness in the system. System 400 includes an arbiter module, a per-instruction performance allocation module, a performance consumption database, and one or more credit collection modules, where the number of credit collection modules is equal to or greater than the number of SQs (as in Fig. 4 shown), host devices, virtual functions or physical functions.
[0050] Commands are selected from SQs based on the availability of performance credits. The arbiter responsible for determining the next SQ to be processed monitors the available credits per client. The per-command performance allocator then allocates performance to each command, taking into account the specific data path required for that command. Credits are consumed accordingly. The performance consumption database is referenced to determine the number of credits required for each operation.
[0051] For example, if multiple SQs are available, the arbiter decides from which SQ the next instruction will be retrieved. The instruction is retrieved, and then the power assigner determines for each instruction whether it can be completed.
[0052] If the requested command is, for example, a read command, the controller first determines the required power using the power consumption database. Based on the information obtained from the power consumption database, the power allocator can decide for each command whether it can be completed. The information retrieved from the power consumption database can include the command size, whether encryption / decryption is enabled or disabled, and other parameters. The power allocator checks all parameters to determine what is required to execute the command and then calculates the power needed to complete it.
[0053] The system is also credit-based and consumes the credits of the respective customer. In this case, the credits refer to performance, as in performance credits. The required performance corresponds to a number of credits (i.e., performance credits). If the SQ has sufficient credits (i.e., performance credits) allocated, the command can be executed. If there are not enough credits, the command will not be executed until sufficient credits are available, except as specified below.
[0054] For example, if SQ A is linked to Client A and consumes the client's credits, the controller will not retrieve the next command from SQ A if there are insufficient credits. Instead, the controller will wait until SQ A has enough credits and then retrieve the next command from SQ A. In this way, the controller achieves fairness by throttling the retrieval of commands.
[0055] The following table shows the power consumption database. This database contains the power credits that should be consumed for each data path operation, depending on the operating mode. The operating mode includes several parameters, such as clock frequencies and power mode. For each instruction, the required data path is defined, and the required power is calculated accordingly. Table subsystem Operation Mode A Mode B Mode C ..... NAND NAND detection 20 NAND 4KB transfer 5 NAND 16 KB transfer 19 NAND program 30 .... .... DRAM Reading L2P 2 Writing L2P 4 Read 4K 5 Write 4K 5 .... .... .... .... .... LDPC Coding 4 Decoding the ultra-low power mode 4 Decoding low power mode 5 Decoding the full power mode 7 .... .... .... .... .... Security Encryption 3 Decryption 3 .... .... .... .... .... RAID protection RAID 5 encoding 4 RAID 6 encoding 5 .... .... .... .... ....
[0056] It should be noted that the subsystem, operation, and modes in the table are merely examples. Additional subsystems, operations, and modes will be considered.
[0057] Referring to the table, NAND flash memory is considered as an example. The power required for NAND flash memory reading depends on several other parameters, such as the clock frequency, the host, the toggle mode, and so on. Additionally, it depends, for example, on whether DRAM or LDPC is used. The controller adds up all the values for the instruction and checks if enough credits are available to execute it. If not enough credits are available, the instruction is not executed because the controller waits for more credits. Otherwise, the controller uses the credits and executes the instruction. In this way, the mechanism achieves performance-based fairness.
[0058] In other words, the controller calculates the total resources required for a host and executes the command if sufficient credits are available. Otherwise, the controller waits until enough credits are available to execute the command. It's important to note that this calculation is performed on a per-command basis. Upon receiving the command, it is first analyzed, and then the necessary calculations are performed to determine the resources required to execute it.
[0059] In one embodiment, a client can consume more bandwidth than allowed if other clients are inactive. More precisely, the controller can decide to execute a command even if there are insufficient credits due to activity. The controller can execute the command even if there are insufficient credits because another host might be idle and therefore not requiring bandwidth at that time. Therefore, the controller can decide to allow the command to be executed even for a specific group that lacks sufficient credits by permitting the host to borrow credits from an inactive host.
[0060] In other embodiments, predictive logic could be integrated to forecast when a client will be idle or underutilized, allowing the available resources to be allocated to other clients. For example, the logic could predict the workload in a given second and adjust the system accordingly.
[0061] Fig.Figure 5 is a schematic representation of a system 500 which, according to one embodiment, includes a fairness control module. The system includes a plurality of host devices 502A-502N. It is understood that although two host devices are shown, both a single host device and a plurality of host devices can be considered. Similarly, the host devices 502A-502N can be actual, physical host devices, physical functions, virtual functions, or combinations thereof. For example, the host devices 502A-502N could be a single host device with a plurality of virtual functions, a plurality of physical functions, or a mixture of the two. Similarly, the host devices 502A-502N could be a plurality of physical devices.
[0062] The System 500 also includes one or more NVMs 506A-506N and a Controller 504. The Controller 504 is coupled to the Host Devices 502A-502N via a Host Interface Module (HIM) 508, and the NVMs 506 are coupled to the Controller via a Flash Interface Module (FIM) 510. The HIM 508 and the FIM 510 are coupled to a Fairness Control Module 512, such as the one described in Fig. 4 Fairness control modules shown.
[0063] Fig.Figure 6 is a flowchart 600 illustrating a bandwidth allocation procedure according to one embodiment. First, in block 602, an instruction is retrieved from a source. The source could be, for example, a host, a virtual function, or a physical function. Next, in block 604, the power consumption required to execute the instruction is evaluated. This evaluation is performed by querying the power consumption database. After determining the power consumption, the controller, in block 606, determines whether the host, virtual function, or physical function has sufficient credits allocated to execute the instruction. If sufficient credits are available, the instruction is simply executed in block 608, and the process is repeated starting in block 602.
[0064] If insufficient credits are available, block 610 determines whether idle times can be predicted for other hosts, virtual functions, or physical functions to which credits are allocated. If hosts, virtual functions, or physical functions are expected to become inactive soon, the controller, at block 612, prepares to borrow credits from the inactive virtual or physical host function, or it simply allows the host device, virtual function, or physical function to exceed its allocated credits. The host, virtual function, or physical function then borrows or exceeds its allocated credits in block 614 and executes the command in block 608.
[0065] If no prediction is available in block 610, block 616 determines whether any hosts, virtual functions, or physical functions are currently idle. If at least one host, virtual function, or physical function is idle, the process proceeds to block 614. Otherwise, the procedure proceeds to block 618 to wait for sufficient credit allocation and then continues to block 606. It is important to note that blocks 610 and 616 can occur in any order and independently of each other. Furthermore, credits are continuously allocated and managed in the fairness control module.
[0066] This revelation leads to a more efficient, adaptable, and equitable use of energy resources in a multi-host environment, significantly improving the overall performance and sustainability of the system compared to the previous approach, which focused on performance fairness. By ensuring a fair distribution of power across multiple hosts, the algorithm optimizes system performance while addressing issues of unfair resource utilization. The concept's adaptability to varying workloads and data path configurations enhances system flexibility and provides a dynamic approach to power allocation. This adaptability not only promotes balanced resource utilization but also contributes to improved energy efficiency, making the multi-host system more sustainable and cost-effective.
[0067] In one embodiment, a data storage device comprises: a storage device; and a controller coupled to the storage device, the controller being configured to: retrieve a first instruction from a first memory location; evaluate the expected power consumption for the first instruction; determine whether the first memory location has sufficient credits to execute the first instruction; and either: execute the first instruction immediately; or hold the first instruction, wait until sufficient credits are allocated to execute the first instruction, and execute the first instruction after the wait. A second instruction can be retrieved from a second memory location in parallel with the retrieval of the first instruction or after the retrieval of the first instruction.Additionally, the second command can be executed concurrently with the first, after the first has finished executing, or while waiting for sufficient credits to be available to execute the first command—all assuming that enough credits are available to execute the second command. The controller is configured to allocate credits to the first memory location. The expected power usage corresponds to a number of credits. Execution occurs when the number of credits corresponding to the expected power usage is equal to or greater than the number of credits allocated to the first memory location.Execution occurs when a second memory location is idle, the first memory location borrows credits from the second memory location, and the number of credits corresponding to the expected power usage is equal to or less than the combined number of credits allocated to the first memory location and borrowed to the second memory location. The controller is configured to predict that a second memory location will be idle; allocate additional credits to the first memory location; and execute the first instruction. The first memory location and any second memory location each comprise a virtual function, a physical function, or combinations thereof. Evaluation and determination are performed by a fairness control module. The fairness control module includes a per-instruction power allocation module and an arbiter.The controller is configured to manage a power consumption database that includes information on power credit for data path operations depending on the operating mode.
[0068] In another embodiment, a data storage device comprises: a storage device; and a controller coupled to the storage device, the controller comprising a fairness control module and the fairness control module being configured to: retrieve one or more commands from one or more transmission queues (SQs) with an arbiter; manage a power consumption database; determine a power allocation for the one or more commands based on information from the power consumption database; and determine whether the one or more commands should be executed or whether the execution of the one or more commands should be delayed. The execution of the first command of the one or more commands is delayed if the power allocation for the first command of the one or more commands exceeds the credits allocated to a corresponding SQ of the first command.The fairness control module includes an allocation module for each command. The fairness control module is configured to allocate credits to one or more SQs. The one or more SQs comprise a multitude of SQs, with each SQ of the multitude being located in a different physical or virtual function.
[0069] In another embodiment, a data storage device comprises: means for storing data; and a controller coupled to the means for storing data, the controller being configured to: execute a first instruction that transfers a first plurality of bytes to the means for storing data using a first set of parameters and consuming a first quantity of power; execute a second instruction that transfers a second plurality of bytes to the means for storing data using a second set of parameters and consuming a second quantity of power, wherein the first plurality is equal to the second plurality, wherein the first set of parameters differs from the second set of parameters, and wherein the first plurality is smaller than the second plurality;and determine when to execute the first and second instructions based on the power consumption credits allocated relative to the first and second sets. The first instruction is called by a first function, and the second instruction is called by a second function, which is different from the first function. The controller is configured to allow either the first or the second function to exceed its credit allocation if another function is idle. The controller is configured to predict when a function will be idle and adjusts the power allocation based on this prediction. The controller is configured to maintain a power consumption database, where determining involves retrieving the first and second sets from the power consumption database.
[0070] While the foregoing relates to embodiments of the present disclosure, other and further embodiments of the disclosure may be conceived without deviating from its basic scope of protection, and the scope of protection thereof is determined by the following claims.
Claims
[1] Data storage device comprising: a storage device (506A-506N); and a controller (504) coupled to the storage device, wherein the controller is configured to: Retrieve (602) a first command from a first source; Evaluate (604) the expected power usage for the first command, where the expected power usage corresponds to a number of credits; Determine (606) whether the first source has sufficient credits to execute the first command; and either: immediate execution (608) of the first command; or Hold the first command, wait until enough credits are allocated to execute the first command, and execute the first command after the wait, with the execution being (608), when a second source is idle, borrows credits from the second source to the first source (614) and the number of credits corresponding to the expected power usage is equal to or less than the combined number of credits allocated to the first source and borrowed from the second source. [2] Data storage device according to claim 1, wherein the controller (504) is configured to allocate credits to the first source. [3] Data storage device according to claim 1, wherein the execution (608) is carried out when the number of credits corresponding to the expected performance usage is equal to or greater than the number of credits allocated to the first source. [4] Data storage device according to claim 1, wherein the controller (504) is further configured to: Predictions (610) that the second source will remain idle; Assigning additional credits to the first source; and Executing (608) the first instruction. [5] Data storage device according to claim 1, wherein the first source and the second source each comprise a host (502A-502N), a virtual function, a physical function or combinations thereof. [6] Data storage device according to claim 1, wherein the evaluation and determination are carried out by a fairness control module (512). [7] Data storage device according to claim 6, wherein the fairness control module (512) comprises a power allocation module per instruction and an arbiter. [8] Data storage device according to claim 1, wherein the controller (504) is configured to manage a power consumption database that includes power credit information for data path operations depending on the operating mode.
Citation Information
Patent Citations
Memory System and Method for Power Management
US20160370841A1
Arbitration techniques for managed memory
US20200209944A1
Apparatus with dynamic arbitration mechanism and methods for operating the same
US20230393877A1