Adaptive Spare Block Usage in Solid State Drives
By dividing the storage media according to the number of spare blocks of the NAND flash die, the workload allocation strategy is optimized, and the problem of uneven number of good blocks in the SSD is solved, which extends the life of the storage media and improves performance.
Patent Information
- Application Number
- CN202080084020.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-11-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-11-05
AI Technical Summary
In existing solid-state drives (SSDs), the number of good blocks of NAND flash dies is uneven, resulting in frequent use of spare blocks from the flash conversion layer (FTL), which affects the overall performance and life of the storage device.
By partitioning the storage medium according to the number of spare blocks per die, high write intensity workloads are allocated to the set of dies with a higher number of spare blocks, and low write intensity workloads are allocated to the set of dies with a lower number of spare blocks, thereby optimizing wear balance and garbage collection strategies for storage devices.
It effectively extends the life of the storage medium, avoids the storage device becoming read-only due to insufficient spare blocks, and improves the overall performance and reliability of the storage device.
Smart Images

Figure CN114746846B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] A solid state drive (SSD) constructed from NAND flash memory includes a plurality of NAND packages that each contain a plurality of NAND flash memory dies. The NAND flash memory dies contain a plurality of blocks that are used to store data and can be erased, programmed, and read. The NAND flash memory dies are designed to have a fixed number of total blocks. Although the total number of blocks on a die is fixed, the actual number of good blocks available on a die varies from die to die when they are implemented in an SSD. This is due to bad blocks marked during the product development process. This process may leave an uneven number of total good blocks in the dies used in an SSD.
[0002] The flash translation layer (FTL) in an SSD controller is used to manage all of the blocks available to it from the raw NAND flash side. Typically, a NAND flash memory die will have more blocks than are required to meet the SSD capacity requirements. These additional blocks are referred to as spare blocks. Spare blocks are used to replace good blocks within a set of capacity blocks when a good block fails. A block may fail due to a programming, erasing, or read failure. This may cause the FTL to mark these blocks as bad blocks and replace them with blocks available in the spare block list.
[0003] It is in this general technological environment that aspects of the present technology disclosed herein are considered. Further, while a general environment has been discussed, it should be understood that the examples described herein should not be limited to the general environment identified in the background. SUMMARY OF THE INVENTION
[0004] The present Summary is provided to introduce a selection of concepts in a simplified form that will be further described in the Detailed Description section below. The present Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in part in the description that follows and in part will become apparent from the description, or may be learned by practice of the present disclosure.
[0005] Non-limiting examples of the present disclosure describe systems, methods, and devices for partitioning storage dies in a non-volatile storage medium based on spare blocks and workload expectations. The number of dies included in a data storage device can be determined. An additional determination can be made of the number of spare blocks included in each of these dies. The number of spare blocks can be the number of blocks in each die that are above the number required to meet the capacity requirements of the storage device.
[0006] A data storage device can provide storage for multiple workloads executing in a host processing system associated with the data storage device. Each workload can correspond to a virtual machine and / or an application. Workloads can be classified as write-intensive or read-intensive. In some examples, the classification can be associated with each workload for which the data storage device provides storage. For example, a first workload can be expected to require a higher number of writes than a second workload. In another example, a first workload can be expected to require a higher write-to-read ratio than a second workload.
[0007] Die(s) of the data storage device can be partitioned into more than two sets and one or more endurance groups. The sets can be partitioned based on the number of spare blocks each die has. That is, the sets can be partitioned based on each die having a relatively high number of spare blocks and each die having a relatively low number of spare blocks. Workloads determined to have a higher write intensity can be assigned to those sets where each die has a relatively high number of spare blocks, and workloads determined to have a lower write intensity can be assigned to those sets where each die has a relatively low number of spare blocks. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following drawings are described for non-limiting and non-exhaustive examples:
[0009] Figure 1 An exemplary computing environment for partitioning storage die(s) based on spare blocks and workload specifications is illustrated.
[0010] Figure 2 An exemplary computing environment including solid-state storage that has been partitioned into a first set of down-channels and endurance groups for use by a first virtual machine and a second set of down-channels and endurance groups for use by a second virtual machine is illustrated.
[0011] Figure 3 An exemplary computing environment including solid-state storage that has been partitioned into four sets of down-channels and corresponding endurance groups for use by four virtual machines and their corresponding workloads is illustrated.
[0012] Figure 4 An exemplary computing environment including solid-state storage that has been partitioned into two cross-channel sets and corresponding endurance groups for use by two virtual machines and their corresponding workloads is illustrated.
[0013] Figure 5 An exemplary computing environment including solid-state storage that has been partitioned into two channel-agnostic sets and corresponding endurance groups for use by two virtual machines and their corresponding workloads is illustrated.
[0014] Figure 6 is an exemplary method for allocating storage workloads across die sets and endurance groups in solid state storage.
[0015] Figure 7 is an exemplary method for writing data to different die sets in solid state storage based on application workloads. DETAILED DESCRIPTION
[0016] Various embodiments will be described in detail with reference to the accompanying drawings, in which like reference numerals represent like parts and components throughout several views. The reference to various embodiments does not limit the scope of the appended claims. Moreover, any examples set forth in this specification are not intended to be limiting and are merely intended to illustrate some of the many possible embodiments of the appended claims.
[0017] Examples of the present disclosure provide systems, methods, and devices for partitioning storage dies in a non-volatile storage medium based on spare blocks and workload expectations. In an example, the non-volatile storage medium may include an SSD. The SSD may include a NAND storage device. In some examples, the non-volatile storage medium may have a capacity requirement (e.g., 1 / 2TB, 1TB). Each die of the non-volatile storage medium may have a basic number of blocks required to meet the capacity requirement of the non-volatile storage medium, and each block above that basic number may be classified as a spare block.
[0018] According to an example, the dies in a non-volatile storage medium can be divided into more than two sets. As used herein, "set", "set of dies", or "die set" refers to one or more dies in a storage medium. A set of dies can be assigned to a specific workload (e.g., a virtual machine, an application executing on a virtual machine). One or more sets can be associated with a wear leveling policy and / or a garbage collection policy. As used herein, "garbage collection" refers to a background process that allows a storage device to mitigate the performance impact of programming / erase cycles by performing certain tasks in the background. Specifically, garbage collection is a process of freeing "old", partially or fully filled blocks by writing valid data from these blocks to "new" free blocks and erasing the old blocks. Thus, after garbage collection, the stale data from the old blocks is no longer included in the new blocks. Garbage collection can be performed at various thresholds (e.g., when a threshold number of blocks are full or nearly full, when a threshold number of pages of a block are full, when a threshold number or percentage of pages of a threshold number of blocks are full, when a threshold percentage of free spare blocks remain, when a threshold number of free spare blocks remain). Garbage collection can be performed based on the amount of data validity in a given region (e.g., a set, a durability group) that falls under a corresponding management policy. For example, the ratio of the total number of invalid logical block addresses (LBAs) to the total number of available LBAs. Garbage collection policies can vary across different sets and durability groups, as described more fully below. As an example, a write-intensive die set can have a first garbage collection policy applied to it, and a read-intensive die set can have a second, different garbage collection policy applied to it.
[0019] The association of a wear leveling policy and / or a garbage collection policy with one or more sets is referred to herein as a "durability group". Each durability group is a separate storage pool for wear leveling and garbage collection purposes. Separate wear statistics can be reported for each durability group. On a drive with more than one durability group, it is possible for one durability group to become completely exhausted and read-only while the other durability groups remain available.
[0020] The workloads processed by a host associated with a non-volatile storage medium can be individually assigned to sets of dies of the non-volatile storage medium. In an example, a die set with a higher number of spare blocks can be utilized to process a workload with a higher write intensity. Alternatively, a die set with a lower number of spare blocks can be utilized to process a workload with a lower write intensity. That is, a workload that requires a higher number of writes over a period of time can be assigned to a die set where each die has more blocks, and a workload that requires a lower number of writes over a period of time can be assigned to a die set where each die has a lower number of blocks.
[0021] According to an example, a non-volatile storage medium can be connected to a host processing system and communicate via a Non-Volatile Memory Host Controller Interface Specification (NVMHCI) or NVM Express (NVMe). In some examples, the die sets described herein can be partitioned into NVM sets and / or NVM endurance groups. Other specifications and communication interfaces can also be applied to implement the systems, methods, and devices described herein, as described more fully below.
[0022] The systems, methods, and devices described herein provide technical advantages for storing data and extending the life of a storage medium. By allocating high write intensity workloads to die sets of dies having a relatively high number of spare blocks and allocating workloads having a lower write intensity to die sets of dies having a relatively low number of spare blocks, the drive as a whole is less likely to need to become read-only due to a lack of space in the die where a relatively low number of spare blocks are utilized by write-intensive workloads. Additionally, the systems, methods, and devices described herein are more efficient when allocating garbage collection policies based on actual block space and / or page space because garbage collection can be initiated based on specific sets included in corresponding endurance groups rather than applying a single policy to all dies in a storage device.
[0023] Figure 1 An exemplary computing environment 100 for partitioning storage dies based on spare block and workload specifications is illustrated. Computing environment 100 includes a host processing system 102 and a data storage device 104, where data storage device 104 further includes a storage controller 106 operably coupled to a storage medium 108. Host processing system 102 communicates with data storage device 104 via a communication link 103, where communication link 103 can include a Peripheral Component Interconnect Express (PCIe) bus, a Small Computer System Interface (SCSI) bus, a Serial Attached SCSI (SAS) bus, a Serial Advanced Technology Attachment (SATA) bus, a Fibre Channel, or some other similar interface or bus.
[0024] The storage medium 108 represents a non-volatile storage medium, such as a solid-state storage medium and a flash memory medium. The storage medium 108 is used as a form of static memory, and its data is preserved when the computer is turned off or loses its external power supply. The storage medium 108 includes a plurality of dies, planes, blocks, and pages. A die is the smallest unit of the storage medium 108 that can independently execute commands or report status. The NAND flash dies within a NAND flash package are connected to the storage controller 106 through channels. In the illustrated example, there are three separate channels - channel 109, channel 111, and channel 113. It should be understood that the presently disclosed aspects can be applied to various SSD controllers based on the number of back-end channels (e.g., four channels, eight channels, sixteen channels, or any combination thereof). Each of the illustrated channels includes four dies. For illustrative purposes, die 110 from channel 109 is illustrated in more detail as die X 110*, die 112 from channel 111 is illustrated in more detail as die Y 112*, and die 114 from channel 113 is illustrated in more detail as die Z 114*. Each die contains one or more planes. The same concurrent operations can occur on each plane. Each plane contains a number of blocks, which are the smallest units that can be erased. Each block contains a number of pages, which are the smallest units that can be written. Each page contains a number of cells. These cells can be configured as single-level cells (SLCs), multi-level cells (MLCs), triple-level cells (TLCs), quad-level cells (QLCs), and / or penta-level cells (PLCs).
[0025] When maintaining the storage of the storage medium 108, the storage medium 108 is coupled to the storage controller 106. The controller can be implemented in hardware, firmware, software, or a combination thereof, and bridges the memory components of the device to the host. In a hardware implementation, the controller includes control circuitry coupled to the storage medium. The control circuitry includes a receiving circuit that receives write requests from the host. Each write request includes the data that is the subject of the write request, as well as the target address of the data. The target address can be a logical block address (LBA), a physical block address (PBA), or other such identifier that describes where the data is to be stored. The storage controller 106 also includes a controller module 124, which can be implemented in hardware, firmware, software, or a combination thereof. The controller module 124 includes a block analysis engine 126, a set partitioning engine 128, and a workload allocation engine 130.
[0026] The storage controller 106 may include a processing system operatively coupled to the storage system. The processing system may include a microprocessor and other circuitry that retrieves and executes operating software from the storage system. The storage system may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, or other data. The storage system may be implemented as a single storage device, but may also be implemented across multiple storage devices or subsystems. The storage system may include additional elements such as a controller for reading operating software from the storage system. Examples of storage media include random access memory, read only memory, magnetic disks, optical disks, and flash memory, and any combination or variation thereof, or any other type of storage media. It should be understood that storage media are not propagated signals under any circumstances.
[0027] The processing system is typically mounted on a circuit board that may also hold the storage system. The operating software of the storage system includes computer programs, firmware, or some other form of machine-readable program instructions. The operating software of the storage system is capable of providing the erase operations described herein for the storage controller. The operating software on the storage system may also include an operating system, utilities, drivers, network interfaces, applications, or some other type of software or firmware. When read and executed by the processing system, the operating software on the storage system directs the controller to operate as described herein.
[0028] In the current example, each die may have 1200 blocks to meet the capacity requirements of the data storage device 104, and the spare blocks are the blocks beyond those 1200 blocks. In this example, die Z 114* has 1210 blocks, so it has 10 spare blocks. In this example, among all the dies included in the storage medium 108, die Z 114* has the smallest number of blocks. Die X 110* has 1240 blocks, so it has 10 spare blocks and 30 additional spare blocks (i.e., the number of blocks above the number of blocks in the die with the smallest number of spare blocks). Die Y 112* has 1220 blocks, so it has 10 spare blocks and 10 additional spare blocks. Although these blocks are the same, for ease of illustration, they are indicated as capacity blocks, spare blocks, or additional spare blocks. Their names are represented by the symbol illustration 122.
[0029] The storage medium 108 can send data from each die to the storage controller 106. The data can include information identifying each die among the dies included in the storage medium 108, and the number of blocks included in each of those dies. The data can additionally include location data for each die and corresponding block in the die. In some examples, the data can be sent by the storage controller 106 to the host processing system 102 during the SSD initialization process. In an example, using this information, the block analysis engine 126 can determine the number of spare blocks (spare blocks and additional spare blocks) included in each die included in the storage medium 108.
[0030] The host processing system 102 can identify the type of workload associated with the virtual machines and / or applications executed by the host processing system 102. For example, the host processing system 102 can identify whether each virtual machine and / or application executed by the host processing system 102 is read-intensive or write-intensive. The read and / or write intensity can be determined based on the ratio of the virtual machines and / or applications executed by the host processing system 102 relative to each other. In other examples, the read and / or write intensity can be determined based on the absolute value of a score. In some examples, the host processing system 102 can simply receive data specifying the type of workload of the virtual machine and / or application. In other examples, the host processing system 102 can analyze historical and / or real-time data associated with the virtual machine and / or application and determine the type of workload based on that data. The type of workload can be determined based on the number of reads and / or writes performed over a period of time. In other examples, the type of workload can be determined based on the number of writes of a specific amount of data over a period of time. Other variables can be considered when determining the type of workload for a given virtual machine and / or application.
[0031] Using the workload type information of each virtual machine and / or application executing on the host processing system 102, and the spare block information determined by the block analysis engine 126, the set partitioning engine 128 can partition the die of the storage medium 108 into sets and subsets. In an example, the partitioning can be based on software and / or firmware (e.g., not a physical partitioning). In an additional example, the partitioning can be at least partially based on hardware (e.g., by channel, by row). In an example, the die can be partitioned into sets based on the number of blocks they have. For example, a die having a relatively high number of blocks and thus a relatively high number of spare blocks can be partitioned into one or more sets that will be utilized by write-intensive virtual machines and / or applications. Alternatively, a die having a relatively low number of blocks can be partitioned into one or more sets that will be utilized by read-intensive virtual machines and / or applications. In some examples, these sets can be associated with durability-based policies. For example, one set of die can have a first garbage collection policy applied to it (e.g., perform garbage collection when there are X free blocks remaining in the set, perform garbage collection when there are Y invalid pages in the set), and another set of die can have a second garbage collection policy applied to it (e.g., perform garbage collection when there are X + 10 free blocks remaining in the set, perform garbage collection when there are Y + 50 invalid pages in the set). In some examples, the durability-based policies can be implemented in whole or in part by the flash translation layer (FTL). In other examples, the durability-based policies can be implemented in whole or in part by the storage controller 106.
[0032] According to an example, the set partitioning engine 128 can partition the die not only based on the number of blocks but also based on the physical location. In some examples, the partitioning engine 128 can partition the die on the basis of downstream channels. For example, the set partitioning engine 128 can partition a first set of die that includes each die in channel 109, a second set of die that includes each die in channel 111, and a third set of die that includes each die in channel 113. In some examples, a set can include more than one die channel (e.g., a set of die includes channel 109 and channel 111). In an additional example, the partitioning engine 128 can partition the die on the basis of cross-flow channels. For example, the set partitioning engine 128 can partition a first set of die that includes the first die in channel 109, the first die in channel 111, and the first die in channel 113, and the partitioning engine 128 can partition a second set of die that includes the second die in each channel, and so on. In yet some other examples, the set partitioning engine 128 can partition the die regardless of the channel location of the die. That is, the set partitioning engine 128 can partition the die based only on the number of blocks or the number of spare blocks included in each die.
[0033] The workload allocation engine 130 may perform operations associated with the set partitioning engine 128. The workload allocation engine 130 may utilize the workload types of virtual machines and / or applications as indicated by the host processing system 102, and allocate those workloads to the partitioned sets. For example, the workload allocation engine 130 may allocate virtual machines and / or applications with write-intensive workloads to sets with a higher number of blocks per die, and allocate virtual machines and / or applications with read-intensive workloads (or in other words, less intensively written workloads) to sets with a lower number of blocks per die. Thus, in this particular example, die X 110* may be partitioned into a set of dies with a relatively high number of blocks, and that set may be used to process write-intensive workloads; die Y 112* may be partitioned into a set of dies with a relatively medium number of blocks, and that set may be used to process medium workloads (e.g., neither write-intensive nor read-intensive); and die Z 114* may be partitioned into a set of dies with a relatively low number of blocks, and that set may be used to process read-intensive workloads.
[0034] Figure 2 An exemplary computing environment 200 including solid-state storage is illustrated, which has been partitioned into a first set of downstream channels and endurance groups for use by a first virtual machine and a second set of downstream channels and endurance groups for use by a second virtual machine. The computing environment 200 includes virtual machine 1 (VM1 218), virtual machine 2 (VM2 224), host processing system 202, and data storage device 204. The data storage device 204 includes a storage controller 206 and a storage medium 208.
[0035] In this example, the storage medium 208 includes multiple dies connected to the storage controller 206 via three separate channels. The dies have been partitioned to process workloads based on downstream channel partitioning. Specifically, each die in channel 209 has been partitioned into a first set 210 and a first endurance group 212. Additionally, each die in the second channel 211 and the third channel 213 has been partitioned into a second set 214 and a second endurance group 216.
[0036] VM1 218 executes application A 220 with read / write specification A 222. The read / write specification A 222 can include the number of reads and / or writes expected to execute application A over a set duration. In additional examples, the read / write specification A 222 can include a number of reads and / or writes that is more than a threshold amount of data expected to execute application A over a set duration. In an example, the read / write specification A 222 can be provided directly to host processing system 202. In some examples, application A 220 may have previously run on one or more other processing systems and utilized one or more other storage devices, and this historical data from executing application A 220 can be utilized to determine read / write specification A 222. In some examples, one or more machine learning models that have been trained to identify read / write patterns can be applied to the historical data to determine read / write specification A 222.
[0037] Similarly, VM2 224 executes application B 226 with read / write specification B 228. The read / write specification B 228 can include the number of reads and / or writes expected to execute application B over a set duration. In additional examples, the read / write specification B 228 can include a number of reads and / or writes that is more than a threshold amount of data expected to execute application B 226 over a set duration. In some examples, application B 226 may have previously run on one or more other processing systems and utilized one or more other storage devices, and this historical data from executing application B 226 can be utilized to determine read / write specification B 228. In some examples, one or more machine learning models that have been trained to identify read / write patterns can be applied to the historical data to determine read / write specification B 228.
[0038] In this example, when VM1 218 and VM2 224 are assigned to host processing system 202 and data storage device 204, determinations can be made regarding the storage requirements of each respective virtual machine. In this example, it is determined that VM2 224 requires twice the storage bandwidth of VM1 218. Additionally, it is determined that VM2 224 has a relatively less write-intensive workload specification / requirement than VM1 218. As such, the data associated with VM2 224 and application 226 is assigned to second set 214, which includes twice the number of dies as first set 210, but those dies have an average spare block count that is smaller than the average spare block count included in the dies of first set 210. In other examples, each die in first set 210 can have a higher spare block count compared to the die in second set 214 that has the lowest number of blocks.
[0039] Partly because VM1 218 requires half of the storage bandwidth of VM2 224 and partly because VM1 218 has a relatively higher write-intensive workload specification / requirement than VM2 224, the data associated with VM1 218 and application 220 is allocated to the first set 212, which includes half the number of dies of the second set 214, but those dies have a greater average number of spare blocks than the average number of spare blocks included in the dies of the second set 214. Thus, the dies included in set 1 210 can better accommodate VM1 218 because these dies are less likely to experience failures related to the higher write-intensive workload required by VM1 218.
[0040] In this example, each die included in set 1 210 has a first garbage collection strategy applied to them, as indicated by durability group 212. Similarly, each die included in set 2 214 has a second garbage collection strategy applied to them, as indicated by durability group 216.
[0041] Figure 3 An exemplary computing environment 300 including solid state storage is illustrated, which has been partitioned into four down-channel sets and corresponding durability groups for use by four virtual machines and their corresponding workloads. Computing environment 300 includes virtual machine 1 (VM1 312), virtual machine 2 (VM2 314), virtual machine 3 (VM3 316), and virtual machine 4 (VM4 318). Computing environment 300 also includes a host processing system 202 and a data storage device 304, and the host processing system 202 executes each of the virtual machines (VM1-VM4). The data storage device 304 includes a storage controller 306 and a storage medium 308.
[0042] Read / write specifications for each of VM1 312, VM2 314, VM3 316, and VM4 318 may have been determined. Those read / write specifications can include at least an indication of the number of write operations expected over a threshold time period. In some examples, the read / write specifications can include an indication of the amount of data expected to be written per write operation over a threshold time period. In additional examples, the read / write specifications can include the number of read operations expected to be performed over a threshold time period. In still other examples, the read / write specifications can include the number of read operations for a certain amount of data, and those read operations are expected to be performed over a threshold time period.
[0043] In this example, VM1 312, VM2 314, VM3 316, and VM4 318 are assigned to host processing system 302 and data storage device 304. Determinations can be made regarding the storage requirements of each respective virtual machine. In this example, the following determinations are made: Each virtual machine requires approximately the same amount of storage. Additionally, the following determination is made: For storage purposes, VM1 312, VM2 314, VM3 316, and VM4 318 require approximately the same number of dies. The following additional determination is made: VM1 312 has a higher write intensity workload than VM2 314, VM2 314 has a higher write intensity workload than VM3 316, and VM3 316 has a higher write intensity workload than VM4 318. Thus, the data associated with VM1 312 is assigned to first set S1, the data associated with VM2 314 is assigned to second set S2, the data associated with VM3 316 is assigned to third set S3, and the data associated with VM4 318 is assigned to fourth set S4. Each die in set S1 has more than X blocks. Each die in set S2 has more than X minus Y blocks. Each die in set S3 has more than X minus Y* blocks, where Y* is greater than Y. Each die in set S4 has more than X minus Y** blocks, where Y** is greater than Y**. In other examples, the average number of blocks in the dies of set S1 can be higher than the average number of blocks in the dies of set S2, the average number of blocks in the dies of set S2 can be higher than the average number of blocks in the dies of set S3, and the average number of blocks in the dies of set S3 can be higher than the average number of blocks in the dies of set S4. Thus, virtual machines with higher write intensity workloads are assigned to sets of dies with higher numbers of blocks.
[0044] In this example, each set of dies (S1, S2, S3, S4) is also associated with a different garbage collection strategy, as indicated by different durability groups. That is, set S1 has a first garbage collection strategy, as indicated by durability group 1 (E1), set S2 has a second garbage collection strategy, as indicated by durability group 2 (E2), set S3 has a third garbage collection strategy, as indicated by durability group 3 (E3), and set S4 has a fourth garbage collection strategy, as indicated by durability group 4 (E4). The garbage collection strategy for the set with a higher write intensity can provide garbage collection at a different threshold than the set with a lower write intensity.
[0045] As an example, a garbage collection strategy for durability group 1 (E1) can provide for garbage collection to be performed when there are X number of spare blocks remaining in the set, and a garbage collection strategy for durability group 4 (E4) can provide for garbage collection to be performed when there are X+Y or X-Y number of spare blocks remaining in the set. In an additional example, a garbage collection strategy for durability group 1 (E1) can provide for garbage collection to be performed when there are X number of full (or nearly full) blocks in the set, and a garbage collection strategy for durability group 4 (E4) can provide for garbage collection to be performed when there are X-Y or X+Y number of full (or nearly full) blocks in the set. In other examples, a garbage collection strategy for durability group 1 (E1) can provide for garbage collection to be performed when there are X% of spare blocks remaining in the set, and a garbage collection strategy for durability group 4 (E4) can provide for garbage collection to be performed when there are Y% of spare blocks remaining in the set.
[0046] Although in this example, each set of dies is indicated as having its own garbage collection strategy, it should be understood that a single garbage collection strategy can be applied to each set in the set, a first garbage collection strategy can be applied to two sets in the set, and a second garbage collection strategy can be applied to the remaining two sets, or a first garbage collection strategy can be applied to three sets in the set, and a second garbage collection strategy can be applied to the remaining set.
[0047] Figure 4 An exemplary computing environment 400 including solid state storage is illustrated, which has been partitioned into two cross-channel sets and corresponding durability groups for use by two virtual machines and their corresponding workloads. Computing environment 400 includes virtual machines 410, which include virtual machine 1 (VM1 412) and virtual machine 2 (VM2 414). Computing environment 400 also includes a host processing system 402 and a data storage device 404 that execute virtual machines 410. Data storage device 404 includes a storage controller 406 and a storage medium 408.
[0048] Read / write specifications for each of VM1 412 and VM2 414 can have been determined. Those read / write specifications can at least include an indication of the number of write operations expected over a threshold time period. In some examples, the read / write specifications can include an indication of the amount of data expected to be written per write operation over a threshold time period. In additional examples, the read / write specifications can include the number of read operations expected to be performed over a threshold time period. In still some other examples, the read / write specifications can include the number of read operations for a certain amount of data, with those read operations expected to be performed over a threshold time period.
[0049] In this example, VM1 412 and VM2 414 are assigned to host processing system 402 and data storage device 404. Determinations can be made regarding the storage requirements of each respective virtual machine. In this example, the determination is made that VM1 412 requires approximately one-third of the storage capacity of VM2 414. The additional determination can be made that although VM1 412 requires less storage than VM2 414, VM1 412 has a higher write intensity workload than VM2 414. In this example, a first plurality of dies with an average number of spare blocks higher than a second plurality of dies can be selected for VM1 412. The first plurality of dies is assigned to VM1 412 as set S1, and set S1 has fewer dies than set S2 because VM1 412 has lower storage requirements. However, because VM1 412 has a more write-intensive workload than VM2 414, set S1 has a higher average number of spare blocks. As such, VM2 414 is assigned to set S2.
[0050] In this example, each set of dies (S1, S2) is also associated with a different garbage collection strategy, as indicated by different durability groups. That is, set S1 has a first garbage collection strategy, as indicated by durability group 1 (E1), and set S2 has a second garbage collection strategy, as indicated by durability group 2 (E2). Each garbage collection strategy in the garbage collection strategies can be associated with a different threshold for initiating garbage collection for the corresponding set of dies.
[0051] Figure 5 An exemplary computing environment 500 including solid state storage is illustrated, which has been partitioned into two channel-agnostic sets and corresponding durability groups for use by two virtual machines and their corresponding workloads. Computing environment 500 includes virtual machine 1 (VM1 510) and virtual machine 2 (VM2 512). Computing environment 500 also includes host processing system 502 and data storage device 504 that execute each of virtual machines VM1 510 and VM2 512. Data storage device 504 includes storage controller 506 and storage medium 508. Symbolic illustration 514 indicates that first set S1 (dies indicated by diagonal lines) has a higher number of spare blocks than second set S2 (dies indicated as filled).
[0052] Read / write specifications for each of VM1 510 and VM2 512 may already be determined. Those read / write specifications may include at least an indication of the number of write operations expected over a threshold time period. In some examples, the read / write specifications may include an indication of the amount of data expected to be written per write operation over the threshold time period. In additional examples, the read / write specifications may include the number of read operations expected to be performed over the threshold time period. In still other examples, the read / write specifications may include the number of read operations for a certain amount of data, with those read operations expected to be performed over the threshold time period.
[0053] In this example, VM1 510 and VM2 512 are assigned to host processing system 502 and data storage device 504. Determinations may be made regarding the storage requirements of each respective virtual machine. In this example, it is determined that each virtual machine requires approximately the same amount of storage. Additionally, it is determined that VM1 510 and VM2 512 require approximately the same number of dies for storage. An additional determination is made that VM1 510 has a workload with a higher write intensity than VM2 512. As such, the data associated with VM1 510 is assigned to first set S1, and the data associated with VM2 512 is assigned to second set S2. In some examples, each die in set S1 may have a higher number of blocks than each die in set S2. In other examples, the average number of blocks in the dies included in set S1 may have a higher number of blocks than the average number of blocks in the dies included in set S2. In this example, the dies in sets S1 and S2 are distributed among the various channels and rows of storage medium 508. As such, dies may be individually selected based on the number of blocks (or spare blocks) included in each of those dies, rather than assigning sets based on the average number of blocks in a given channel and / or row. Thus, individually selecting dies for assignment to a particular set provides a greater ability to customize wear leveling for individual workloads.
[0054] In this example, sets S1 and S2 are illustrated as being associated with the same garbage collection policy. However, it should be understood that a first garbage collection policy may be applied to set S1, and a different garbage collection policy may be applied to set S2.
[0055] Figure 6 is an exemplary method 600 for allocating storage workloads across die sets and endurance groups in solid state storage. Method 600 begins at the start operation, and the flow moves to operation 602.
[0056] At operation 602, data is received from a non-volatile storage medium. The data may include information identifying each of a plurality of dies included in the non-volatile storage medium and may include the number of blocks included in each of the plurality of dies. In an example, the data may include location information for the plurality of dies.
[0057] The process continues from operation 602 to operation 604, in which a determination is made regarding the number of spare blocks included in each of the plurality of dies. The non-volatile storage medium may have a capacity requirement (e.g., 1 / 2TB, 1TB). Each die may have a basic number of blocks required to meet the capacity requirement of the non-volatile storage medium, and each block above that basic number may be classified as a spare block.
[0058] The process continues from operation 604 to operation 606, in which a first set and a second set of dies are identified. The first set of dies may be identified as having a higher number of spare blocks than the second set of dies. In an example, identifying the first set and the second set may include partitioning the dies based on their spare blocks. For example, a determination may be made regarding a plurality of workloads that will utilize the non-volatile storage medium (e.g., a first workload associated with a first application, a second workload associated with a second application), and the expected write intensity of those respective workloads may be identified. After the expected workloads are identified, the dies may be partitioned into sets based on the number of spare blocks in the dies. Dies having a higher number of spare blocks may be partitioned into the set that will be assigned to the more write-intensive workload, and dies having a lower number of spare blocks may be partitioned into the set that will be assigned to the less write-intensive workload.
[0059] The process continues from operation 606 to operation 608, in which a first workload is assigned to the first set of dies, and the first workload is classified as write-intensive. In some examples, the classification may be related to the read and / or write intensity of one or more additional workloads to be processed by the non-volatile storage medium. In other examples, the classification may be a standalone classification (e.g., independent of other workloads). In some examples, the first workload may be associated with one or more applications that are executed on a host device associated with the non-volatile storage medium. In additional examples, the first workload may be associated with a virtual machine that executes one or more applications that are executed on a host device associated with the non-volatile storage medium.
[0060] The process continues from operation 608 to operation 610, where a second workload is assigned to a second set of dies, and the second workload is classified as read-intensive. In some examples, the classification can be related to the read and / or write intensity of one or more additional workloads to be processed by the non-volatile storage medium. In other examples, the classification can be a standalone classification (e.g., unrelated to other workloads). In some examples, the second workload can be associated with one or more applications that are executed on a host device associated with the non-volatile storage medium. In additional examples, the second workload can be associated with a virtual machine that executes one or more applications that are executed on a host device associated with the non-volatile storage medium.
[0061] From operation 610, method 600 moves to an end operation and method 600 ends.
[0062] Figure 7 Exemplary method 700 is for writing data to different sets of dies in a solid-state storage based on application workloads. Method 700 starts at a start operation, and the process moves to operation 702.
[0063] At operation 702, a command to partition a plurality of dies included in a storage system is received. The command can specify that the dies are to be partitioned into a first set of dies and a second set of dies, where the first set of dies has a higher number of spare blocks than the second set of dies. The number of spare blocks in a die can be determined based on the minimum number of blocks required for each die to meet the capacity requirements of the storage system (e.g., 1 / 2TB, 1TB).
[0064] The process continues from operation 702 to operation 704, where the plurality of dies are partitioned into a first set of dies and a second set of dies. The partitioning can include software partitioning, firmware partitioning, hardware partitioning, or a combination thereof. The partitioning can include assigning each set to a separate workload associated with the storage system. For example, the first workload can include the execution of a first application to be assigned to the first set of dies, and the second workload can include the execution of a second application to be assigned to the second set of dies.
[0065] The process continues from operation 704 to operation 706, where data from a first application executed on a first virtual machine of a host processing system is written to the first set of dies. The first application can be classified as a write-intensive application. In some examples, the write intensity can be relative to a second application executed on a second virtual machine of the host processing system. In other examples, the write intensity can be absolute (e.g., not relative to any other application).
[0066] The process continues from operation 706 to operation 708, in which data from a second application executing on a second virtual machine of the host processing system is written to a second set of dies. The second application can be classified as a read-intensive application. In some examples, the read intensity can be relative to the first application. In other examples, the read intensity can be absolute (e.g., not relative to any other application). In this example, the first application is determined to require a higher number of writes than the second application over a period of time. Thus, the first application is more write-intensive than the second application.
[0067] The process moves from operation 708 to an end operation and method 700 ends.
[0068] For example, aspects of the present disclosure have been described above with reference to block diagrams and / or operational descriptions of methods, systems, and computer program products according to aspects of the present disclosure. The functions / acts labeled in the blocks may not occur in the order shown in any flowchart. For example, depending on the functions / acts involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order.
[0069] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and to enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. The various (both structural and method) features, whether shown and described combinatorially or individually, are intended to be selectively included or omitted to produce embodiments having a particular set of features. Given the description and illustration of this application, those skilled in the art can anticipate variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application and that do not depart from the broad scope of the claimed disclosure.
[0070] The various embodiments described above are provided by way of illustration only and should not be construed as limiting the appended claims. Those skilled in the art will readily recognize that various modifications and changes can be made without following the example embodiments and applications illustrated and described herein and without departing from the true spirit and scope of the appended claims.
Claims
1. A computer-implemented method, comprising: Receiving data from a non-volatile storage medium, the data including: Information identifying each die among a plurality of dies included in the non-volatile storage medium, and The number of blocks included in each die among the plurality of dies; Determining the number of spare blocks included in each die among the plurality of dies; Identifying a first set of the plurality of dies and a second set of the plurality of dies, wherein the first set of dies has a higher number of spare blocks compared to the second set of dies; Allocating a first application or a first virtual machine having a first workload to the first set of dies, the first workload being classified as write-intensive; and Allocating a second application or a second virtual machine having a second workload to the second set of dies, the second workload being classified as read-intensive.
2. The computer-implemented method according to claim 1, wherein determining the number of spare blocks included in each die among the plurality of dies comprises: Determining a minimum number of blocks in each die among the plurality of dies to meet the storage requirements for the non-volatile storage medium; And Identifying the number of blocks above the minimum number of blocks in each die among the plurality of dies as the number of spare blocks included in each die among the plurality of dies.
3. The computer-implemented method according to claim 1, wherein: The first workload is determined based on receiving read-write specifications associated with a first virtual machine to be allocated to the first set of dies; and The second workload is determined based on receiving read-write specifications associated with a second virtual machine to be allocated to the second set of dies.
4. The computer-implemented method according to claim 3, wherein the read-write specifications associated with the first virtual machine indicate that an application associated with the first virtual machine will require the non-volatile storage medium to cache more than a threshold amount of data for more than a threshold duration.
5. The computer-implemented method according to claim 3, wherein the read-write specifications associated with the second virtual machine indicate that an application associated with the first virtual machine will require the non-volatile storage medium to cache more than a threshold amount of data for less than a threshold duration.
6. The computer-implemented method according to claim 3, wherein: The read-write specifications associated with the first virtual machine are determined via the application of a machine learning model that has been trained to identify read-write patterns of data associated with a first application to be executed by the first virtual machine; and The read-write specifications associated with the second virtual machine are determined via applying the machine learning model to data associated with a second application to be executed by the second virtual machine.
7. The computer-implemented method according to claim 1, wherein each die in the first set of dies has a higher number of spare blocks compared to each die in the second set of dies.
8. The computer-implemented method according to claim 7, wherein the dies included in the first set of dies are distributed across multiple channels and multiple rows in the multiple channels.
9. The computer-implemented method according to claim 1, wherein the dies included in the first set of dies are arranged in a first channel, and the dies included in the second set of dies are arranged in a second channel.
10. The computer-implemented method according to claim 1, wherein the dies included in the first set of dies are arranged in a first row of the multiple channels, and the dies included in the second set of dies are arranged in a second row of the multiple channels.
11. The computer-implemented method according to claim 1, further comprising: performing a first garbage collection policy on the first set of dies, wherein the first garbage collection policy specifies that garbage collection for the first set of dies will be initiated when it is determined that there are fewer than a first threshold number of free blocks in the first set of dies; and performing a second garbage collection policy on the second set of dies, wherein the second garbage collection policy specifies that garbage collection for the second set of dies will be initiated when it is determined that there are fewer than a second threshold number of free blocks in the second set of dies.
12. The computer-implemented method according to claim 1, further comprising: determining that there is no additional space in one of the first set or the second set of dies to accept an additional write operation; and converting the set of dies having no additional space to a read-only set.
13. A storage device, comprising: a non-volatile storage medium; a controller coupled to the non-volatile storage medium and configured to: receive data from the non-volatile storage medium, the data including: information identifying each die among a plurality of dies included in the non-volatile storage medium, and the number of blocks included in each die among the plurality of dies; determine the number of spare blocks included in each die among the plurality of dies; identify a first set of the plurality of dies and a second set of the plurality of dies, wherein the first set has a higher number of spare blocks compared to the dies of the second set; allocate a first application or a first virtual machine having a first workload to the first set of dies, the first workload being classified as write-intensive; and allocate a second application or a second virtual machine having a second workload to the second set of dies, the second workload being classified as read-intensive.
14. The storage device according to claim 13, wherein: the first workload is determined based on receiving read-write specifications associated with a first virtual machine to be allocated to the first set of dies; and the second workload is determined based on receiving read-write specifications associated with a second virtual machine to be allocated to the second set of dies.
15. The storage device according to claim 13, wherein each die in the first set of dies has a higher number of spare blocks compared to each die in the second set of dies.
16. The storage device according to claim 15, wherein the dies included in the first set of dies are distributed across multiple channels and multiple rows in the multiple channels.
17. The storage device according to claim 13, wherein the controller is further configured to: perform a first garbage collection policy on the first set of dies; and perform a second garbage collection policy on the second set of dies.
18. The storage device according to claim 17, wherein: the first garbage collection policy provides that garbage collection for the first set of dies is initiated when it is determined that there are fewer than a first threshold number of free blocks in the first set of dies; and the second garbage collection policy provides that garbage collection for the second set of dies is initiated when it is determined that there are fewer than a second threshold number of free blocks in the second set of dies.
19. A computing device, comprising: a host processing system; a storage system coupled to the host processing system and configured to: receive a command to divide a plurality of dies included in the storage system into a first set of dies and a second set of dies, wherein the first set of dies has a higher number of spare blocks compared to the second set of dies; divide the plurality of dies into the first set of dies and the second set of dies; write data from a first application executed on a first virtual machine of the host processing system to the first set of dies, wherein the first application is classified as a write-intensive application; and write data from a second application executed on a second virtual machine of the host processing system to the second set of dies, wherein the second application is classified as a read-intensive application.
20. The computing device according to claim 19, wherein the storage system is further configured to: apply a first garbage collection policy to the first set of dies, the first garbage collection policy providing that garbage collection for the first set of dies is initiated when it is determined that there are fewer than a first threshold number of free blocks in the first set of dies; and apply a second garbage collection policy to the second set of dies, the second garbage collection policy providing that garbage collection for the second set of dies is initiated when it is determined that there are fewer than a second threshold number of free blocks in the second set of dies.
Citation Information
Patent Citations
Control arrangements and methods for accessing block oriented nonvolatile memory
CN103339618A
Memory system and control method
CN108509146A