Dynamic Performance Level Adjustment of Storage Drives
By monitoring and adjusting the aging and wear characteristics of the storage drive, dynamically adjusting its performance rating and reclassifying and reorganizing in the storage environment, the problem of degradation of storage drives after aging and wear is solved, achieving the effect of reduced failure rate and extended life.
Patent Information
- Application Number
- CN202080035143.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-26
- Filing Date
- 2020-06-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-06-11
AI Technical Summary
In the prior art, storage drives fail to provide original performance after aging and wear, resulting in high failure rates and shortened life cycles when used in certain RAID arrays, storage layers or workloads, and lack of an effective dynamic compensation mechanism.
By monitoring the aging and wear characteristics of the storage drive, periodically adjusting its performance ratings, and reclassifying and recombining storage drives in the storage environment based on these characteristics, placing drives of the same performance rating within the same storage group as possible, using the intelligent reconstruction process for data exchange for dynamic compensation.
Reduces the failure rate of storage drives, extends its service life, and optimizes the performance of RAID arrays and storage layers, improving system reliability and efficiency.
Smart Images

Figure CN113811861B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to systems and methods for dynamically compensating for aging and / or wear on storage drives. Background Art
[0002] RAID (i.e., Redundant Array of Independent Disks) is a storage technology that provides increased storage functionality and reliability through redundancy. A RAID is created by combining multiple storage drive components (e.g., disk drives and / or solid state drives) into a logical unit. Data is then distributed across the drives using different techniques, known as "RAID levels". The current standard RAID levels, including levels 1 through 6, are a basic set of RAID configurations that use striping, mirroring, and / or parity to provide data redundancy. Each configuration provides a balance between two key objectives: (1) increased data reliability and (2) increased I / O performance.
[0003] When the storage drives in a RAID are new, the storage drives can have certain performance characteristics or specifications. These characteristics or specifications can be expressed in terms of performance ratings, write classifications per day, storage capacity, amount of over-provisioning, etc. However, as the storage drives age and wear, the storage drives may not be able to provide the same performance characteristics or specifications as when new. This can make the storage drives unsuitable for use in certain RAID arrays, storage tiers, or workloads that may have certain performance requirements. If the wear or lifespan of the storage drives is ignored and the same workload is driven to these storage drives regardless of their lifespan and / or wear, the storage drives may exhibit an unacceptably high failure rate and / or a reduced lifecycle.
[0004] In view of the foregoing, there is a need for systems and methods for dynamically compensating for aging and / or wear on storage drives. Ideally, such systems and methods would periodically reassign storage drives to appropriate RAID arrays, storage tiers, or workloads based on the lifespan and / or wear of the storage drives. Such systems and methods would also ideally reduce the failure rate and increase the lifespan of the storage drives. Summary of the Invention
[0005] The present invention has been developed in response to the current state of the art and, in particular, in response to problems and needs that are not fully addressed by currently available systems and methods in the art. Accordingly, embodiments of the present invention have been developed to dynamically compensate for age and / or wear on storage drives. The features and advantages of the present invention will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the invention as set forth hereinafter.
[0006] Consistent with the foregoing, a method for dynamically changing the performance levels of multiple storage drives is disclosed. In one embodiment, such a method monitors the characteristics (e.g., aging, wear, etc.) of multiple storage drives within a storage environment. Each storage drive has a performance level associated therewith. Based on these characteristics, the method periodically modifies the performance levels of the storage drives. The method then reorganizes the storage drives based on the performance levels of the storage drives within different storage groups (e.g., RAID arrays, storage tiers, workloads, etc.). For example, the method can place storage drives with the same performance level within the same storage group as much as possible.
[0007] Also disclosed and claimed herein are corresponding systems and computer program products. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] To facilitate an understanding of the advantages of the present invention, a more specific description of the present invention briefly described above will be presented by reference to specific embodiments shown in the drawings. It should be understood that these drawings only depict typical embodiments of the present invention and are therefore not considered to limit its scope. The present invention will be described and explained with additional specificity and detail by using the drawings, in which:
[0009] Figure 1 is a high-level block diagram showing an example of a network environment in which the systems and methods according to the present invention can be implemented;
[0010] Figure 2 is a high-level block diagram showing an embodiment of a storage system in which one or more RAID or storage tiers can be implemented;
[0011] Figure 3 is a high-level block diagram showing new and different storage drives and their associated performance levels;
[0012] Figure 4 is a high-level block diagram showing the performance levels of storage drives that degrade as the storage drives age;
[0013] Figure 5 is a high-level block diagram showing the reorganization of storage drives within a RAID based on the performance levels of the storage drives;
[0014] Figure 6 is a high-level block diagram showing the reorganization of storage drives within a RAID based on the daily write classification of the storage drives within the RAID;
[0015] Figure 7 is a high-level block diagram showing the reorganization of storage drives within a RAID based on their logical storage capacity;
[0016] Figure 8is a high - level block diagram showing each sub - module within the drive monitoring module according to the present invention;
[0017] Figure 9 is a high - level block diagram showing each sub - module within the re - classification module according to the present invention;
[0018] Figure 10 is a flowchart showing an embodiment of a method for reorganizing a storage drive based on drive characteristics;
[0019] Figure 11 is a flowchart showing an embodiment of a method for reorganizing a storage drive based on the performance level of the storage drive;
[0020] Figure 12 is a flowchart showing an embodiment of a method for reorganizing a storage drive based on the daily write classification of the storage drive; and
[0021] Figure 13 is a flowchart showing an embodiment of a method for reorganizing a storage drive based on the logical storage capacity of the storage drive. Detailed Description
[0022] It will be readily understood that the components of the present invention, generally described and illustrated in the figures herein, can be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the present invention represented in the figures is not intended to limit the scope of the claimed invention, but merely represents certain examples of the presently contemplated embodiments of the present invention. The presently described embodiments will be best understood by reference to the drawings, wherein like components are designated by like numerals throughout.
[0023] The present invention can be embodied as a system, method, and / or computer program product. The computer program product may include a computer - readable storage medium (or medium) having computer - readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0024] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage system, a magnetic storage system, an optical storage system, an electromagnetic storage system, a semiconductor storage system, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0025] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage system via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0026] The computer-readable program instructions for performing the operations of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages.
[0027] Computer-readable program instructions may be executed entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing aspects of the present invention.
[0028] Aspects of the present invention may be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0029] These computer-readable program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, which, when executed by the processor of the computer or other programmable data processing apparatus, creates a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0030] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing device, or other device that causes a series of operational steps to be performed on the computer, other programmable device, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable device, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0031] See Figure 1, an example of the network environment 100 is shown. The network environment 100 is presented to show an instance of an environment in which the systems and methods according to the present invention can be implemented. The network environment 100 is presented by way of example and not limitation. In fact, in addition to the network environment 100 shown, the systems and methods disclosed herein can be applicable to a wide variety of different network environments.
[0032] As shown, the network environment 100 includes one or more computers 102, 106 interconnected by a network 104. The network 104 can include, for example, a local area network (LAN) 104, a wide area network (WAN) 104, the Internet 104, an intranet 104, etc. In some embodiments, the computers 102, 106 can include both a client computer 102 and a server computer 106 (also referred to herein as "host" 106 or "host system" 106). Generally, the client computer 102 initiates a communication session, while the server computer 106 waits and responds to requests from the client computer 102. In some embodiments, the computer 102 and / or the server 106 can be connected to one or more internal or external directly attached storage systems 112 (e.g., an array of hard storage drives, solid state drives, tape drives, etc.). These computers 102, 106 and the directly attached storage systems 112 can communicate using protocols such as ATA, SATA, SCSI, SAS, Fibre Channel, etc.
[0033] In some embodiments, the network environment 100 can include a storage network 108 behind the server 106, such as a storage area network (SAN) 108 or a LAN 108 (e.g., when using network attached storage). This network 108 can connect the server 106 to one or more storage systems 110, such as an array 110a of hard disk drives or solid state drives, a tape library 110b, individual hard disk drives 110c or solid state drives 110c, tape drives 110d, a CD-ROM library, etc. To access the storage system 110, the host system 106 can communicate through a physical connection from one or more ports on the host 106 to one or more ports on the storage system 110. The connection can be through a switch, fabric, direct connection, etc. In some embodiments, the server 106 and the storage system 110 can communicate using networking standards such as Fibre Channel (FC) or iSCSI.
[0034] See Figure 2, an example of a storage system 110a including an array of hard disk drives 204 and / or solid state drives 204 is shown. The internal components of the storage system 110a are shown because in some embodiments, a RAID array can be implemented, in whole or in part, within such a storage system 110a. As shown, the storage system 110a includes a storage controller 200, one or more switches 202, and one or more storage drives 204, such as hard disk drives 204 and / or solid state drives 204 (e.g., flash-based drives 204). The storage controller 200 can enable one or more host systems 106 (e.g., open systems and / or mainframe servers 106 running operating systems such as z / OS, zVM, etc.) to access data in one or more storage drives 204.
[0035] In selected embodiments, the storage controller 200 includes one or more servers 206a, 206b. The storage controller 200 may also include a host adapter 208 and a device adapter 210 to connect the storage controller 200 to the host system 106 and the storage drives 204, respectively. The multiple servers 206a, 206b can provide redundancy to ensure that data is always available to the connected host systems 106. Thus, when one server 206a fails, another server 206b can pick up the I / O load of the failed server 206a to ensure that I / O can continue between the host system 106 and the storage drives 204. This process can be referred to as "failover".
[0036] In selected embodiments, each server 206 includes one or more processors 212 and memory 214. The memory 214 may include volatile memory (e.g., RAM) as well as non-volatile memory (e.g., ROM, EPROM, EEPROM, hard disk, flash memory, etc.). In some embodiments, the volatile and non-volatile memory may store software modules that run on the processors 212 and are used to access data in the storage drives 204. These software modules can manage all read and write requests to the logical volumes in the storage drives 204.
[0037] Having a similar architecture to Figure 2 an example of the storage system 110a shown is the IBM DS8000 TM Enterprise Storage System. The DS8000 TM is a high-performance, high-capacity storage controller that provides disk and solid state storage designed to support continuous operation. However, the techniques disclosed herein are not limited to the IBM DS8000 TMEnterprise storage system 110a, but can be implemented in any comparable or similar storage system 110, regardless of the manufacturer, product name, or component or component name associated with system 110. Any storage system that can benefit from one or more embodiments of the present invention is considered to fall within the scope of the present invention. Thus, the IBM DS8000 TM is presented only as an example and not as a limitation.
[0038] See Figure 3 , in certain embodiments, the storage drives 204 of storage system 110a can be configured in one or more RAID arrays 304 to provide a desired level of reliability and / or I / O performance. A RAID array 304 is created by combining multiple storage drive components (e.g., disk drives 204 and / or solid state drives 204) into a logical unit. Data is then distributed across the drives using different techniques (referred to as "RAID levels"). The current standard RAID levels, including RAID levels 1 through 6, are a basic set of RAID configurations that use striping, mirroring, and / or parity to provide data redundancy. Each configuration provides a balance between two key objectives: (1) increasing data reliability and (2) increasing I / O performance.
[0039] When the storage drives 204 in the RAID array 304 are new, the storage drives 204 can have certain performance characteristics or specifications. These characteristics or specifications can be expressed in terms of performance grades, write classifications per day, storage capacity, amount of overprovisioning, etc. However, as the storage drives 204 age and wear, the storage drives 204 may not be able to provide the same performance characteristics or specifications that they could when new. This may make the storage drives 204 unsuitable for use in certain RAID arrays 304, storage tiers, or workloads that may have certain performance requirements. If the wear or life of the storage drives 204 is ignored and the same workload is driven to these storage drives 204 regardless of their life and / or wear, the storage drives 204 may exhibit an unacceptably high failure rate and / or a reduced lifecycle.
[0040] Accordingly, there is a need for systems and methods to dynamically compensate for the life and / or wear on the storage drives 204. Ideally, such systems and methods would periodically reassign the storage drives 204 to appropriate RAID arrays 304, storage tiers, or workloads based on the life and / or wear of the storage drives 204. Such systems and methods would also ideally reduce the failure rate and increase the useful life of the storage drives 204.
[0041] As Figure 3As shown, when the storage drive 204 is new, the new storage drive 204 can have a certain performance level associated with it. For example, the storage drive 204 has been assigned a performance level of A, B, or C, where performance level A has better performance (e.g., I / O performance) than performance level B, and performance level B has better performance than performance level C. In the illustrated embodiment, each RAID array 304 starts with storage drives 204 of the same performance level (in this example, performance level A), although this is not required in all embodiments. For example, depending on the performance requirements of the RAID array 304, some RAID arrays 304 can be assigned storage drives 204 with a lower performance level (e.g., performance level B or C).
[0042] Over time, due to the failure and replacement of certain storage drives 204 within the RAID array 304, the RAID array 304 can be composed of storage drives 204 with different aging and / or wear characteristics. However, as the storage drives 204 age and wear, the storage drives 204 may not be able to provide the same performance characteristics or specifications, or may not be able to do so without exhibiting a higher than acceptable failure rate or reduced lifecycle. However, when a storage drive 204 of a certain performance level ages and / or wears, the storage drive 204 can continue to be used in the same manner as the result of its originally assigned performance level.
[0043] In certain embodiments, the systems and methods according to the present invention can monitor the characteristics (e.g., aging and / or wear) of the storage drives 204 in the storage environment and reclassify the storage drives 204 periodically with an appropriate performance level. For example, for a storage drive 204 with an expected lifespan of three years, the storage drive 204 can initially be assigned a performance level of A. After the first year of use, the storage drive 204 can be downgraded to performance level B. After the second year of use, the storage drive 204 can be downgraded to performance level C. Whenever a storage drive 204 is assigned a new performance level, the storage drive 204 can (if not already) be placed in an appropriate storage group (e.g., RAID array 304, a storage tier in a hierarchical storage environment, storage drives 204 with specific workload requirements, etc.). This can be achieved by swapping storage drives 204 of a certain performance level with storage drives 204 of a different performance level in the storage environment so that the storage group (e.g., RAID array 304, storage tier, workload, etc.) contains as many storage drives 204 of the same performance level as possible.
[0044] As Figure 3As shown, in some embodiments, one or more modules 300, 302, 303 can be used to provide different features and functions in accordance with the present invention. For example, the drive monitoring module 300 can be configured to monitor storage drive characteristics such as storage drive age, usage, and / or wear. In contrast, the drive reclassification module 302 can be configured to periodically reclassify the storage drive 204 based on the characteristics of the storage drive 204. For example, the drive reclassification module 302 can lower the performance level of the storage drive 204 when the storage drive 204 ages or deteriorates. Once the storage drive 204 has been reclassified, the drive reorganization module 303 can be configured to reorganize the storage drive 204 in the storage environment based on the classification of the storage drive. For example, in some embodiments, the drive reorganization module 303 can place as many storage drives 204 with the same performance level in the same storage group, such as in the same RAID array 304 or the same storage tier.
[0045] Figure 4 Shown is the Figure 3 RAID array 304 after the storage drive 204 has been reclassified by the drive reclassification module 302. As shown, after a period of time and after the various storage drives 204 within the RAID array 304 have been replaced or swapped with other storage drives 204 of the same or different performance levels, the storage drives 204 in the RAID array 304 can have different life and / or wear characteristics. Based on these aging and / or wear characteristics, the drive reclassification module 302 can modify the performance levels of the storage drives 204 to reflect their aging and / or wear. For example, as Figure 4 shown, the storage environment can include storage drives 204 that are classified as performance levels A, B, or C based on their manufacturer specifications or their aging and / or wear. This creates a scenario where the RAID array 304 contains storage drives 204 of different performance levels, as Figure 4 shown. In some embodiments, the storage drive 204 with the lowest performance level can define the performance of the entire RAID array 304. That is, the RAID array 304 may only be able to perform at the level of its lowest performing storage drive 204 (i.e., the storage drive 204 with the lowest performance level). Thus, in order to maximize the performance of the RAID array 304, the RAID array 304 would ideally contain storage drives 204 with the same performance level.
[0046] See Figure 5, To achieve this, once the drive reclassification module 302 has modified the performance levels of the storage drives 204 to match their age and / or wear, the drive reorganization module 303 can reorganize the storage drives 204 within the storage environment. More specifically, the drive reorganization module 303 can attempt to place storage drives 204 of the same performance level in the same RAID array 304. The higher performance RAID arrays 304 will ideally contain the higher performance level storage drives 204. Similarly, the lower performance RAID arrays 304 can contain the lower performance level storage drives 204.
[0047] To reorganize the storage drives 204, the drive reorganization module 303 can use, for example, a spare storage drive 204 as an intermediate data storage area to exchange the storage drives 204 between the RAID arrays 304 to facilitate the exchange of data. In some embodiments, this is achieved by using an intelligent reconstruction process to copy data from one storage drive 204 to another. The intelligent reconstruction process can reduce the exposure to data loss by maintaining the ability of the storage drive 204 to be used as a spare (even while data is being copied to it). In some embodiments, when copying data from a first storage drive 204 to a second storage drive 204 (e.g., a spare storage drive 204), the intelligent reconstruction process can create a bitmap for the first storage drive 204. Each bit can represent a portion of the storage space on the first storage drive 204 (e.g., a one megabyte region). The intelligent reconstruction process can then begin copying data from the first storage drive 204 to the second storage drive 204. When each section is copied, its associated bit can be recorded in the bitmap.
[0048] If a write to a section of the first storage drive 204 is received while the data copy process is in progress, the intelligent reconstruction process can check the bitmap to determine whether the data in the associated section has already been copied to the second storage drive 204. If not, the intelligent reconstruction process can simply write the data to the corresponding part of the first storage drive 204. Otherwise, after writing the data to the first storage drive 204, the data can also be copied to the second storage drive 204. Once all parts have been copied from the first storage drive 204 to the second storage drive 204, the RAID array 304 can begin using the second storage drive 204 in place of the first storage drive 204. This frees the first storage drive 204 from the RAID array 304.
[0049] Alternatively, the intelligent reconstruction process can utilize watermarks instead of bitmaps to track which data has been copied from the first storage drive 204 to the second storage drive 204. In such an embodiment, segments can be copied from the first storage drive 204 to the second storage drive 204 in a specified order. The watermark can track the extent to which the copy process has advanced through the segments. If a write to a segment of the first storage drive 204 is received during the copy process, the intelligent reconstruction process can check the watermark to determine whether the data in the segment has already been copied to the second storage drive 204. If not, the intelligent reconstruction process can write the data to the first storage drive 204. Otherwise, after writing the data to the first storage drive 204, the intelligent reconstruction process can also copy the data to the second storage drive 204. Once all parts have been copied from the first storage drive 204 to the second storage drive 204, the RAID array 304 can start using the second storage drive 204 in place of the first storage drive 204. This frees the first storage drive 204 from the RAID array 304.
[0050] In other embodiments, the drive reclassification module 302 can change other characteristics of the storage drive 204 within the storage environment. For example, the drive reclassification module 302 can modify the daily write classification, logical storage capacity, and / or the amount of over-provisioning associated with the storage drive 204 based on the aging or wear of the storage drive 204. Then, the drive reorganization module 303 can reorganize the storage drive 204 based on the daily write classification of the storage drive 204 within the RAID array 304 (as Figure 6 shown) or its logical storage capacity (as Figure 7 shown). Thus, the systems and methods according to the present invention can monitor different characteristics of the storage drive 204, reclassify the storage drive 204 based on their characteristics, and reorganize the storage drive 204 after the storage drive 204 has been reclassified. The manner in which this can be achieved will be described in more detail in connection with Figures 8 to 13 more detail.
[0051] Figure 8 is a high-level block diagram showing different sub-modules that can be included within the drive monitoring module 300. The drive monitoring module 300 and associated sub-modules can be implemented in hardware, software, firmware, or a combination thereof. The drive monitoring module 300 and associated sub-modules are presented by way of example and not limitation. In different embodiments, more or fewer sub-modules can be provided. For example, the functionality of some sub-modules can be combined into a single or fewer number of sub-modules, or the functionality of a single sub-module can be distributed across several sub-modules.
[0052] As shown, the drive monitoring module 300 includes one or more of an aging monitoring module 800, a wear monitoring module 802, and an over-provisioning monitoring module 804. The aging monitoring module 800 may be configured to monitor the aging of the storage drive 204 in the storage environment. In some embodiments, this may be accomplished by detecting when the storage drive 204 is newly installed in the storage environment and then tracking the amount of time the storage drive 204 has been in the storage environment from that point forward.
[0053] In contrast, the wear monitoring module 802 may monitor the wear of the storage drive 204 in the storage environment. In some embodiments, wear may be determined based on the use of the storage drive 204, such as the amount of I / O that has been driven to the storage drive 204 during the life of the storage drive 204, the amount of time the storage drive 204 has been active, the storage group (e.g., RAID array 304, storage tier, or workload) to which the storage drive 204 has been associated during its use, and / or the like.
[0054] The over-provisioning monitoring module 804 may be configured to monitor the amount of over-provisioning present within the storage drive 204. Certain storage drives 204 (such as solid-state storage drives 204 (SSDs)) may have a certain percentage of their total storage capacity dedicated to storing data, and the remaining percentage is kept idle in the form of "over-provisioning." This over-provisioning generally improves performance and increases the lifespan of the solid-state storage drive 204. As the solid-state storage drive 204 ages and / or wears, the storage elements within the solid-state storage drive 204 may degrade, which in turn may reduce the amount of over-provisioning within the storage drive 204. This may reduce the performance and / or lifespan of the solid-state storage drive 204.
[0055] Figure 9 is a high-level block diagram showing different sub-modules that may be included within the previously described drive reclassification module 302. As shown, the drive reclassification module 302 may include one or more of a performance level adjustment module 900, a daily write adjustment module 902, and a size / over-provisioning adjustment module 904. The performance level adjustment module 900 may be configured to adjust the performance level of the storage drive 204 based on the characteristics of the storage drive 204 (e.g., aging and / or wear) detected by the drive monitoring module 300. In some embodiments, the performance level adjustment module 900 may adjust the performance level in discrete steps. For example, the performance level adjustment module 900 may reduce the performance level from performance level A to performance level B and from performance level B to performance level C depending on the aging or wear of the storage drive 204.
[0056] The daily write adjustment module 902 can be used to adjust the daily write classification associated with the storage drive 204. Based on the age, wear, and / or amount of the overprovisioning setting associated with the storage drive 204, the daily write adjustment module 902 can reduce the daily write classification associated with the storage drive 204. In some embodiments, this can occur in discrete steps. For example, the daily write classification can drop from a first level (e.g., 200 GB / day) to a second level (150 GB / day), then from the second level to a third level (e.g., 100 GB / day), and so on, in different discrete steps, depending on the characteristics associated with the storage drive 204 (e.g., the amount of overprovisioning, age, etc.).
[0057] The size / overprovisioning adjustment module 904 can be configured to adjust the logical storage capacity and / or the amount of overprovisioning associated with the storage drive 204. As described above, as the storage drive 204 (e.g., the solid-state storage drive 204) ages or is utilized, the sectors or storage elements in the storage drive 204 may degrade. When the amount of overprovisioning within the storage drive 204 decreases to a certain level or threshold, the size / overprovisioning adjustment module 904 can adjust the logical storage capacity and / or the amount of overprovisioning in the storage drive 204. For example, the size / overprovisioning adjustment module 904 can reduce the logical storage capacity of the storage drive 204 in order to increase the amount of overprovisioning. This can improve the performance of the storage drive 204 and / or increase the service life of the storage drive 204. In some embodiments, this can occur in different discrete steps. For example, the size / overprovisioning adjustment module 904 can reduce the logical storage capacity of the storage drive 204 from size A (e.g., 900 GB) to size B (e.g., 800 GB), from size B to size C (e.g., 700 GB), etc., in different discrete steps as the storage drive 204 ages and / or wears. As will be explained in more detail below, in some cases, when the logical storage capacity of the storage drive 204 is reduced, some data may need to be migrated away from the storage drive 204 to facilitate the reduction of the logical storage capacity.
[0058] See Figure 10 , which illustrates a flowchart of an embodiment of a method 1000 for reorganizing a storage drive 204 based on drive characteristics. The method 1000 is intended to broadly cover Figures 11 to 13 the more specific methods shown in
[0059] As shown in the figure, method 1000 initially determines 1002 whether it is time to reclassify and reorganize the storage drive 204 within the storage environment. In some embodiments, method 1000 is intended to be executed periodically, such as weekly, monthly, or every several months. Step 1002 can be configured to determine whether method 1000 should be executed and when method 1000 should be executed.
[0060] If there is a time to reclassify and reorganize the storage drive 204 within the storage environment, then method 1000 determines (1004) the drive characteristics of the storage drive 204 in the storage environment, such as aging, wear, amount of overprovisioning, etc. In some embodiments, once certain characteristics are observed, method 1000 actually modifies the drive characteristics. For example, in the case where the amount of overprovisioning in the storage drive 204 drops below a specified level, method 1000 can reduce the logical storage capacity, thereby increasing the amount of overprovisioning of the storage drive 204.
[0061] Method 1000 then reclassifies 1008 the storage drive 204 within the storage environment based on the determined characteristics. For example, if the storage drive 204 has reached a certain age, then method 1000 can reclassify (1008) the storage drive 204 from performance level A to performance level B, or reclassify (1008) the storage drive 204 from performance level B to performance level C. In another example, if the storage drive 204 has reached a specific lifespan or amount of wear, then method 1000 can reclassify (1008) the storage drive 204 from a first daily write classification to a second daily write classification. In yet another example, if the storage drive 204 has reached a specific lifespan or amount of wear, then method 1000 can reclassify (1008) the storage drive 204 from having a first logical storage capacity to having a second logical storage capacity.
[0062] Method 1000 then determines (1010) the requirements for certain storage groups that include the storage drive 204 (e.g., RAID array 304, storage tier, storage drives 204 that support certain workloads, etc.). For example, method 1000 may determine the performance requirements for the RAID array 304 within the storage environment. Based on the requirements of the storage group and the classification of the storage drive 204, method 1000 reorganizes (1012) the storage drives 204 within the storage group. For example, method 1000 may attempt to reorganize (1012) the storage drives 204 in the RAID array 304 such that the higher performance RAID array 304 or storage tier contains storage drives 204 with a higher performance level (e.g., performance level A), and the lower performance RAID array 304 or storage tier contains storage drives 204 with a lower performance level. In some embodiments, this may be achieved by using an intelligent reconstruction process that exchanges data between the storage drives 204 to swap the storage drives 204 in the RAID array 304 or storage tier.
[0063] See Figure 11 , an embodiment of method 1100 for reorganizing the storage drive 204 based on the performance level of the storage drive 204 is shown. If it is time to reclassify and reorganize the storage drives 204 within the storage environment at step 1102, method 1100 determines (1104) the aging of the storage drives 204 in the storage environment. Then method 1100 reduces (1106) the performance level of the storage drives 204 whose age has reached a specified threshold. For example, once a storage drive 204 with a three-year projected lifespan has reached one year old, it can be reduced (1106) from performance level A to performance level B. When the same storage drive 204 has reached two years old, it can be reduced (1106) from performance level B to performance level C.
[0064] Method 1100 then determines (1108) the requirements for different storage groups (e.g., RAID array 304, storage tier, storage drives 204 that support certain workloads, etc.) in the storage environment. For example, method 1100 may determine that a first RAID array 304 or storage tier in the storage environment requires higher performance, and thus higher performance storage drives 204, and a second RAID array 304 or storage tier in the storage environment can utilize lower performance storage drives 204. Based on the requirements of the storage group and the characteristics of the storage drive 204, method 1100 reorganizes (1110) the storage drives 204 within the storage group. For example, method 1100 may place as many storage drives 204 with the same performance level in the same storage group as possible. In some embodiments, this may be achieved by using an intelligent reconstruction process to swap the storage drives 204 in the RAID array 304 or storage tier.
[0065] In some embodiments, reorganization step 1110 works as follows. Assume that storage drives 204 are reorganized according to their performance levels, storage drives 204 are classified into performance levels A, B, or C, and the storage group is RAID array 304: Reorganization step 1110 can first generate a "count" of all storage drives 204 in the storage environment of performance level A. Then reorganization step 1110 can find the RAID array 304 with the most storage drives 204 of performance level A in the storage environment. Reorganization step 1110 reduces the "count" by the number of storage drives 204 of performance level A in RAID array 304. Then, reorganization step 1110 uses an intelligent reconstruction process to swap storage drives 204 of performance level A from other RAID arrays 304 into RAID array 304 until RAID array 304 contains all storage drives 204 of performance level A. Reorganization step 1110 reduces the "count" by the number of storage drives 204 that are swapped. If the "count" is zero, the reorganization of storage drives 204 of performance level A stops. Otherwise, the reorganization step 1110 is repeated for the next highest number of RAID arrays 304 with storage drives 204 of performance level A or until the "count" becomes zero. This reorganization step 1110 is repeated for storage drives 204 of performance level B and performance level C. Reorganization step 1110 places as many storage drives 204 of the same performance level in the same RAID array 304 as possible.
[0066] See Figure 12 FIG. shows an embodiment of method 1200 for reorganizing storage drives based on the daily write classification of the storage drives. If, at step 1202, it is time to reclassify and reorganize the storage drives 204 within the storage environment, method 1200 determines (1204) the age and / or amount of over-provisioning of the storage drives 204 in the storage environment. Then, method 1200 reduces (1206) the daily write classification of the storage drives 204 whose age or amount of over-provisioning has reached a specified threshold. For example, once a storage drive 204 has reached a first age (e.g., one year) or a specified amount of over-provisioning (e.g., less than ten percent over-provisioning), the storage drive 204 can be reduced (1206) from a level 1 daily write classification to a level 2 daily write classification. Similarly, once a storage drive 204 has reached a second age (e.g., two years) or again reaches the specified amount of over-provisioning, the storage drive 204 can be reduced (1206) from a level 2 writes / day classification to a level 3 writes / day classification.
[0067] Then, method 1200 determines (1208) the requirements of different storage groups (e.g., RAID array 304, storage tier, storage drive 204 that supports certain workloads, etc.) in the storage environment. For example, method 1200 may determine that the first RAID array 304 or storage tier in the storage environment requires higher performance, and thus storage drives 204 with a higher daily write classification and the second RAID array 304 or storage tier in the storage environment can utilize storage drives 204 with a lower daily write classification. Based on the requirements of the storage group and the daily write classification of storage drives 204, method 1200 reorganizes (1210) the storage drives 204 within the storage group. For example, method 1200 may place storage drives 204 with the same daily write classification in the same storage group as much as possible. In some embodiments, this can be achieved by using an intelligent reconstruction process to swap storage drives 204 in the RAID array 304 or storage tier. In some embodiments, the reorganization step 1210 may work in substantially the same manner as the reorganization step 1110 described in conjunction with Figure 11 works in substantially the same manner as the reorganization step 1110 described in conjunction with
[0068] See Figure 13 , which shows an embodiment of method 1300 for reorganizing storage drives based on the logical storage capacity of the storage drives. If it is time to reclassify and reorganize the storage drives 204 within the storage environment at step 1302, method 1300 determines (1304) the age of the storage drives 204 in the storage environment. Method 1300 then determines (1306) an appropriate reduction in the logical storage capacity and an increase in the amount of overprovisioning of the storage drives 204 based on the age of the storage drives 204. If the storage drive 204 contains an amount of data that exceeds the amount of data that can be accommodated in the newly determined logical storage capacity, method 1300 migrates (1308) the data from the storage drive 204 to another location. The storage drive 204 can then be reconfigured to dedicate the released logical storage capacity to overprovisioning. Method 1300 then reduces (1310) the logical storage capacity according to the amount determined in step 1306 and increases (1310) the amount of overprovisioning in the storage drives 204.
[0069] Then, method 1300 reorganizes (1310) the storage drives 204 within a storage group (e.g., RAID array 304, storage tier, etc.). For example, method 1300 may place storage drives 204 with the same logical storage capacity in the same storage group as much as possible. In some embodiments, this can be achieved by using an intelligent reconstruction process to swap storage drives 204 in the storage group. In some embodiments, the reorganization step 1310 may be combined with Figure 11 and 12The described recombination steps 1110, 1210 work in substantially the same manner.
[0070] The flowcharts and / or block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-usable media according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, depending on the functionality involved, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for dynamically changing the performance levels of multiple storage drives, the method comprising: Monitoring the characteristics of multiple storage drives within a storage environment, each storage drive having a performance level associated therewith; Periodically performing the following steps: Identifying a selected storage drive among the multiple storage drives, the amount of overprovisioning of which has fallen below a specific level; Changing the amount of overprovisioning in the selected storage drive to modify the logical storage capacity of the selected storage drive; Modifying the performance level of the storage drive based on the change in the amount of overprovisioning; And Reorganizing the storage drives based on the performance levels of the storage drives within different storage groups.
2. The method according to claim 1, wherein the storage group is a RAID array.
3. The method according to claim 1, wherein the storage group is a storage tier in a hierarchical storage environment.
4. The method according to claim 1, wherein The storage group is a group of storage drives configured to support a certain I / O workload.
5. The method according to claim 1, wherein, Monitoring the characteristics includes monitoring the age of the storage drive.
6. The method according to claim 1, wherein Monitoring the characteristics includes monitoring the wear of the storage drive.
7. The method according to claim 1, wherein Reorganizing the storage drives includes placing as many storage drives with the same performance level as possible in the same storage group.
8. A computer program product for dynamically changing the performance levels of multiple storage drives, the computer program product comprising a computer-readable medium having computer-usable program code embodied therein, the computer-usable program code being configured to perform the following when executed by at least one processor: Monitoring the characteristics of multiple storage drives within a storage environment, each storage drive having a performance level associated therewith; Periodically performing the following steps: Identifying a selected storage drive among the multiple storage drives, the amount of overprovisioning of which has fallen below a specific level; Changing the amount of overprovisioning in the selected storage drive to modify the logical storage capacity of the selected storage drive; Modifying the performance level of the storage drive based on the change in the amount of overprovisioning; And Reorganizing the storage drives based on the performance levels of the storage drives within different storage groups.
9. The computer program product according to claim 8, wherein, The storage group is a RAID array.
10. The computer program product according to claim 8, wherein, The storage group is a storage tier in a hierarchical storage environment.
11. The computer program product according to claim 8, wherein, The storage group is a group of storage drives configured to support a specific I / O workload.
12. The computer program product according to claim 8, wherein, Monitoring the characteristics includes monitoring the age of the storage drive.
13. The computer program product according to claim 8, wherein, Monitoring the characteristics includes monitoring the wear of the storage drive.
14. The computer program product according to claim 8, wherein, Reorganizing the storage drives includes placing as many storage drives with the same performance level as possible in the same storage group.
15. A system for dynamically changing the performance levels of multiple storage drives, the system comprising: At least one processor; At least one memory device coupled to the at least one processor and storing instructions for execution on the at least one processor, the instructions causing the at least one processor to: Monitor the characteristics of multiple storage drives within a storage environment, each storage drive having a performance level associated therewith; Periodically perform the following steps: Identify a selected storage drive among the plurality of storage drives, the amount of overprovisioning of which has fallen below a specific level; Change the amount of overprovisioning in the selected storage drive to modify the logical storage capacity of the selected storage drive; Modify the performance level of the storage drive based on the change in the amount of overprovisioning; And Reorganize the storage drives based on the performance levels of the storage drives within different storage groups.
16. The system according to claim 15, wherein, The storage group is a RAID array.
17. The system according to claim 15, wherein, The storage group is a storage tier in a hierarchical storage environment.
18. The system according to claim 15, wherein, The storage group is a group of storage drives configured to support a specific I / O workload.
19. The system according to claim 15, wherein, Monitoring the characteristic includes monitoring the age of the storage drive.
20. The system according to claim 15, wherein, Monitoring the characteristic includes monitoring the wear of the storage drive.
Citation Information
Patent Citations
System and method for extending service life of nonvolatile memory, and computer program product
CN107436847A