Performing wear leveling between storage systems of a storage cluster

By migrating storage objects between storage systems in storage clusters, the problem of loss imbalance between storage systems is solved, and the efficiency and performance of storage devices are improved.

CN114860150BActive Publication Date: 2025-07-25DELL PROD LP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110158120.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-04
Publication Date
2025-07-25
Estimated Expiration
2041-02-04

AI Technical Summary

Technical Problem

The prior art cannot effectively solve the problem of loss imbalance between storage systems in storage clusters, resulting in inefficiency of storage devices and performance impacts.

Method used

By moving storage objects between storage systems of the storage cluster, using the loss balance service module to collect usage information, identify loss level imbalance, and migrate storage objects between storage systems to achieve loss balance.

Benefits of technology

It improves the overall efficiency of storage devices in the storage cluster, reduces the performance impact of loss imbalance, and optimizes the utilization of storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860150B_ABST
    Figure CN114860150B_ABST
Patent Text Reader

Abstract

An apparatus includes at least one processing device, and the at least one processing device includes a processor coupled to a memory. The at least one processing device is configured to obtain usage information of each of two or more storage systems of a storage cluster, and determine a wear level of each of the storage systems of the storage cluster at least partially based on the obtained usage information. The at least one processing device is further configured to identify a wear level imbalance of the storage cluster at least partially based on the determined wear level of each of the storage systems of the storage cluster. The at least one processing device is further configured to move storage objects between the storage systems of the storage cluster in response to the identified wear level imbalance of the storage cluster being greater than an imbalance threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This field generally relates to information processing, and more particularly to storage devices in an information processing system. Background Art

[0002] Storage arrays and other types of storage systems are typically shared by multiple host devices over a network. Applications running on the host devices each include one or more processes that perform the application functionality. Such processes issue input-output (IO) operation requests to be delivered to the storage system. The storage controller of the storage system services such IO operation requests. In some information processing systems, multiple storage systems may be used to form a storage cluster. Summary of the Invention

[0003] Exemplary embodiments of the present disclosure provide techniques for performing wear leveling between storage systems of a storage cluster.

[0004] In one embodiment, a device includes at least one processing device that includes a processor coupled to a memory. The at least one processing device is configured to perform the following steps: obtain usage information for each of two or more storage systems of a storage cluster; determine a wear level for each of the two or more storage systems of the storage cluster at least in part based on the obtained usage information; and identify a wear level imbalance in the storage cluster at least in part based on the determined wear levels for each of the two or more storage systems of the storage cluster. The at least one processing device is further configured to, in response to the identified wear level imbalance in the storage cluster being greater than an imbalance threshold, perform the following steps: move one or more storage objects between the two or more storage systems of the storage cluster.

[0005] These and other exemplary embodiments include, but are not limited to, methods, devices, networks, systems, and processor-readable storage media. Brief Description of the Drawings

[0006] Figure 1 is a block diagram of an information processing system for performing wear leveling between storage systems of a storage cluster in an exemplary embodiment.

[0007] Figure 2 is a flowchart of an exemplary process for performing wear leveling between storage systems of a storage cluster in an exemplary embodiment.

[0008] Figure 3 illustrates a storage cluster implementing storage cluster-wide wear leveling in an exemplary embodiment.

[0009] Figure 4Illustrates a process flow for rebalancing wear levels between storage systems of a storage cluster in an illustrative embodiment.

[0010] Figure 5 Illustrates a table of storage cluster array status information and wear levels in an illustrative embodiment.

[0011] Figure 6 Illustrates the wear levels of a storage array of a storage cluster prior to rebalancing in an illustrative embodiment.

[0012] Figure 7 Illustrates a table of storage objects stored on a storage array of a storage cluster in an illustrative embodiment.

[0013] Figure 8 Illustrates a table of iterations of rebalancing in a storage cluster and storage cluster array status information prior to rebalancing in an illustrative embodiment.

[0014] Figure 9 Illustrates the wear levels of a storage array of a storage cluster after rebalancing in an illustrative embodiment.

[0015] Figure 10 and Figure 11 Illustrates an example of a processing platform that can be utilized to implement at least a portion of an information processing system in an illustrative embodiment. DETAILED DESCRIPTION

[0016] Exemplary information processing systems and associated computers, servers, storage devices, and other processing devices will be referred to herein to describe illustrative embodiments. However, it should be understood that the embodiments are not limited to use with the specific illustrative system and device configurations shown. Thus, the term "information processing system" as used herein is intended to be broadly interpreted to encompass, for example, processing systems including cloud computing and storage systems and other types of processing systems including various combinations of physical and virtual processing resources. Thus, an information processing system can include, for example, at least one data center or other type of cloud-based system that includes one or more cloud-hosted tenants accessing cloud resources.

[0017] Figure 1FIG. 0 illustrates an information processing system 100 that is configured to provide functionality for storage cluster-wide loss balancing among storage systems for a storage cluster according to an illustrative embodiment. The information processing system 100 includes one or more host devices 102-1, 102-2,......, 102-N (collectively referred to as host devices 102) that communicate with one or more storage arrays 106-1, 106-2,......, 106-M (collectively referred to as storage arrays 106) via a network 104. The network 104 may include a storage area network (SAN).

[0018] As Figure 1 shown, the storage array 106-1 includes a plurality of storage devices 108, each of which stores data utilized by one or more applications running on the host device 102. The storage devices 108 are illustratively arranged in one or more storage pools. The storage array 106-1 also includes one or more storage controllers 110 that facilitate IO processing to the storage devices 108. The storage array 106-1 and its associated storage devices 108 are examples of what is more generally referred to herein as a "storage system". Such a storage system in this embodiment is shared by the host device 102 and is thus also referred to herein as a "shared storage system". In an embodiment where there is only a single host device 102, the host device 102 may be configured to use the storage system exclusively.

[0019] The host device 102 illustratively includes a corresponding computer, server, or other type of processing device capable of communicating with the storage array 106 via the network 104. For example, at least a subset of the host devices 102 may be implemented as corresponding virtual machines of a computing service platform or other type of processing platform. The host devices 102 in such an arrangement illustratively provide computing services, such as implementing one or more applications on behalf of each of one or more users associated with the corresponding host device among the host devices 102.

[0020] The term "user" herein is intended to be broadly interpreted to cover numerous arrangements of people, hardware, software, or firmware entities, and combinations of such entities.

[0021] Although computing and / or storage services may be provided to users according to a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, and / or a function as a service (FaaS) model, it should be understood that numerous other cloud infrastructure arrangements may be used. Moreover, in the case of stand-alone computing and storage systems implemented within a given enterprise, the illustrative embodiments may be implemented outside the context of a cloud infrastructure.

[0022] The storage devices 108 of the storage array 106-1 can implement logical units (LUNs) configured to store objects of a user associated with the host device 102. These objects can include files, blocks, or other types of objects. The host device 102 interacts with the storage array 106-1 using read and write commands and other types of commands transmitted over the network 104. In some embodiments, although such commands more specifically include Small Computer System Interface (SCSI) commands, in other embodiments, other types of commands can also be used. As used widely herein, the term given IO operation illustratively includes one or more such commands. References to terms such as "input-output" and "IO" herein should be understood to refer to input and / or output. Thus, an IO operation involves at least one of input and output.

[0023] Moreover, the term "storage device" as used herein is intended to be broadly interpreted to cover, for example, logical storage devices such as LUNs or other logical storage volumes. A logical storage device can be defined in the storage array 106-1 to include different portions of one or more physical storage devices. Thus, the storage device 108 can be considered to include the corresponding LUN or other logical storage volume.

[0024] In Figure 1 the information processing system 100, multiple arrays in the storage array 106 are assumed to be part of a storage cluster, and the host device 102 is assumed to submit IO operations to be processed by the storage cluster. At least one of the storage controllers in the storage array 106 (e.g., the storage controller 110 of the storage array 106-1) is assumed to implement storage cluster management functionality for the storage cluster. Alternatively, this storage cluster management functionality can be implemented external to the storage array 106 of the storage cluster (e.g., on a dedicated server, on the host device 102, etc.). The information processing system 100 also includes a storage cluster wear leveling service 112 configured to provide functionality for implementing wear leveling between the storage arrays 106 of the storage cluster. The storage cluster wear leveling service 112 includes a usage information collection module 114, a cluster-wide wear state determination module 116, and a storage object migration module 118.

[0025] The usage information collection module 114 is configured to obtain usage information for each of the storage arrays 106 that are part of a storage cluster. The cluster-wide loss status determination module 116 is configured to determine a loss level for each of the storage arrays 106 of the storage cluster at least in part based on the obtained usage information, and to identify a loss level imbalance of the storage cluster at least in part based on the determined loss levels of each of the storage arrays 106 of the storage cluster. The storage object migration module 118 is configured to move one or more storage objects between the storage arrays 106 of the storage cluster in response to the identified loss level imbalance of the storage cluster being greater than an imbalance threshold.

[0026] At least part of the functionality of the usage information collection module 114, the cluster-wide loss status determination module 116, and the storage object migration module 118 may be implemented at least in part in the form of software stored in a memory and executed by a processor.

[0027] Although shown in the embodiments of Figure 1 as being external to the host device 102 and the storage arrays 106, it should be understood that in other embodiments, the storage cluster loss balancing service 112 may be implemented at least in part inside one or more of the host devices 102 and / or one or more of the storage arrays 106 (e.g., on the storage controller 110 of the storage array 106-1).

[0028] Figure 1 The host device 102, the storage arrays 106, and the storage cluster loss balancing service 112 in the embodiments of

[0029] are assumed to be implemented using at least one processing platform, where each processing platform includes one or more processing devices, and each processing device has a processor coupled to a memory. Such processing devices may illustratively include a particular arrangement of computing, storage, and network resources. For example, in some embodiments, in an arrangement where Docker containers or other types of LXCs are configured to run on a VM, the processing device is implemented at least in part using virtual resources such as virtual machines (VMs) or Linux containers (LXCs) or a combination of both.

[0030] Network 104 can be implemented using a variety of different types of networks to interconnect storage system components. For example, network 104 can include a SAN, which is part of a global computer network such as the Internet, but other types of networks can also be part of the SAN, including wide area networks (WANs), local area networks (LANs), satellite networks, telephone or cable networks, cellular networks, wireless networks (such as WiFi or WiMAX networks), or various portions or combinations of these networks and other types of networks. Thus, in some embodiments, network 104 includes a combination of multiple different types of networks, each network including processing devices configured to communicate using Internet Protocol (IP) or other relevant communication protocols.

[0031] As a more specific example, some embodiments can utilize one or more high-speed local area networks, where the associated processing devices communicate with each other using high-speed peripheral component interconnect (PCIe) cards of these devices and networking protocols such as InfiniBand, Gigabit Ethernet, or Fibre Channel. As will be appreciated by those skilled in the art, numerous alternative networking arrangements are possible in a given embodiment.

[0032] Although in some embodiments, certain commands used by host device 102 to communicate with storage array 106 illustratively include SCSI commands, other types of commands and command formats can also be used in other embodiments. For example, some embodiments can utilize command features and functionality associated with Non-Volatile Memory Express (NVMe) as described in the May 2017 revision 1.3 of the NVMe specification, which is incorporated herein by reference. Other storage protocols of this type that can be utilized in the illustrative embodiments disclosed herein include: Fabric-based NVMe, also known as NVMeoF; and Transmission Control Protocol (TCP)-based NVMe, also known as NVMe / TCP.

[0033] The storage array 106-1 in this embodiment is assumed to include persistent memory implemented using flash memory or other types of non-volatile memory of the storage array 106-1. More specific examples include NAND-based flash memory or other types of non-volatile memory, such as resistive RAM, phase change memory, spin transfer torque magnetoresistive RAM (STT-MRAM), and Intel Optane based on 3D XPoint TM memory TMDevice. Although the persistent memory is also assumed to be separate from the storage devices 108 of the storage array 106-1, in other embodiments, the persistent memory may be implemented as one or more designated portions of one or more of the storage devices 108. For example, in some embodiments, the storage device 108 may include a flash-based storage device, such as in embodiments involving all-flash storage arrays, or may be implemented in whole or in part using other types of non-volatile memory.

[0034] As described above, the communication between the host device 102 and the storage array 106 may utilize a PCIe connection or other types of connections implemented over one or more networks. For example, illustrative embodiments may use interfaces such as Internet Small Computer System Interface (iSCSI), Serial Attached SCSI (SAS), and Serial ATA (SATA). In other embodiments, numerous other interfaces and associated communication protocols may be used.

[0035] In some embodiments, the storage array 106 may be implemented as part of a cloud-based system.

[0036] The storage devices 108 of the storage array 106-1 may be implemented using solid-state drives (SSDs). Such SSDs are implemented using non-volatile memory (NVM) devices such as flash memory. Other types of NVM devices that may be used to implement at least a portion of the storage device 108 include non-volatile random access memory (NVRAM), phase change RAM (PC-RAM), and magnetic RAM (MRAM). These and various combinations of different types of NVM devices or other storage devices may also be used. For example, hard disk drives (HDDs) may be used in combination with or instead of SSDs or other types of NVM devices. Thus, numerous other types of electronic or magnetic media may be used when implementing at least one subset of the storage device 108.

[0037] The storage array 106 may alternatively or additionally be configured to implement multiple different storage tiers of a multi-tier storage system. For example, a given multi-tier storage system may include a fast or performance tier implemented using flash storage devices or other types of SSDs and a capacity tier implemented using HDDs, where one or more such tiers may be server-based. It will be apparent to those skilled in the art that a wide variety of other types of storage devices and multi-tier storage systems may be used in other embodiments. The particular storage devices used in a given storage tier may vary according to the specific requirements of a given embodiment, and multiple different storage device types may be used within a single storage tier. As indicated previously, the term "storage device" as used herein is intended to be interpreted broadly and thus may encompass, for example, SSDs, HDDs, flash drives, hybrid drives, or other types of storage products and devices or portions thereof, and illustratively includes logical storage devices such as LUNs.

[0038] As another example, the storage array 106 may be used to implement one or more storage nodes in a clustered storage system that includes multiple storage nodes interconnected by one or more networks.

[0039] Thus, it should be apparent that the term "storage array" as used herein is intended to be interpreted broadly and may encompass multiple different instances of commercially available storage arrays.

[0040] Other types of storage products that may be used to implement a given storage system in an illustrative embodiment include software-defined storage, cloud storage, object-based storage, and scale-out storage. In an illustrative embodiment, combinations of multiple of these and other storage types may also be used to implement a given storage system.

[0041] In some embodiments, the storage system includes a first storage array and a second storage array arranged in an active-active configuration. For example, such an arrangement may be used to ensure that data stored in one of the storage arrays is replicated to the other storage array using a synchronous replication process. Such data replication across multiple storage arrays may be used to facilitate fault recovery in system 100. Thus, one of the storage arrays may act as a production storage array relative to the other storage array that acts as a backup or recovery storage array.

[0042] However, it should be understood that the embodiments disclosed herein are not limited to an active-active configuration or any other particular storage system arrangement. Thus, the illustrative embodiments herein may be configured using a variety of other arrangements, including for example, active-passive arrangements, active-active Asymmetric Logical Unit Access (ALUA) arrangements, and other types of ALUA arrangements.

[0043] These and other storage systems may be part of what is more generally referred to herein as a processing platform, which includes one or more processing devices, each of which includes a processor coupled to a memory. A given such processing device may correspond to one or more virtual machines or other types of virtualization infrastructure, such as Docker containers or other types of LXC. As indicated above, communication between such elements of system 100 may occur over one or more networks.

[0044] As used herein, the term "processing platform" is intended to be broadly interpreted to cover (e.g., but not limited to) a collection of multiple processing devices configured to communicate over one or more networks and one or more associated storage systems. For example, a distributed implementation of host device 102 is possible, where some of the host devices in host device 102 reside in a first data center at a first geographic location, while other host devices in host device 102 reside in one or more other data centers at one or more other geographic locations that may be remote from the first geographic location. Storage array 106 and storage cluster loss balancing service 112 may be implemented at least in part in the first geographic location, the second geographic location, and one or more other geographic locations. Thus, in some implementations of system 100, different host devices in host device 102, storage array 106, and storage cluster loss balancing service 112 may reside in different data centers.

[0045] Numerous other distributed implementations of host device 102, storage array 106, and storage cluster loss balancing service 112 are possible. Thus, host device 102, storage array 106, and storage cluster loss balancing service 112 may also be implemented in a distributed manner across multiple data centers.

[0046] The following will be combined with Figure 10 and Figure 11 to more fully describe additional examples of the processing platform utilized to implement portions of system 100 in the illustrative embodiments.

[0047] It should be understood that Figure 1The specific set of elements shown for performing wear leveling between storage systems of a storage cluster is presented only by way of illustrative example, and in other embodiments, additional or alternative elements may be used. Thus, another embodiment may include additional or alternative systems, devices, and other network entities, as well as different arrangements of modules and other components.

[0048] It should be understood that these and other features of the illustrative embodiments are presented only by way of example and should not be construed in any way as being restrictive.

[0049] Reference will now be made Figure 2 to the flowchart of to describe in more detail an exemplary process for performing wear leveling between storage systems of a storage cluster. It should be understood that this specific process is only an example, and in other embodiments, additional or alternative processes for performing wear leveling between storage systems of a storage cluster may be used.

[0050] In this embodiment, the process includes steps 200 through 206. It is assumed that these steps are performed by the storage cluster wear leveling service 112 using the usage information collection module 114, the cluster-wide wear state determination module 116, and the storage object migration module 118. The process begins at step 200: obtaining usage information for each of two or more storage systems of a storage cluster. The obtained usage information may include: capacity usage information for each of two or more storage systems of the storage cluster; I / O temperature information characterizing the number of input-output requests within a specified threshold of the current time for each of two or more storage systems of the storage cluster; and cumulative write request count information for each of two or more storage systems of the storage cluster.

[0051] In step 202, a wear level for each of two or more storage systems of the storage cluster is determined at least in part based on the obtained usage information. This may include, for a given storage system among the two or more storage systems, calculating a weighted sum of the capacity usage information of the given storage system, the input-output temperature information of the given storage system, and the cumulative write request count information of the given storage system. A first weight assigned to the capacity usage information of the given storage system may be lower than a second weight assigned to the input-output temperature information and a third weight assigned to the cumulative write request count information of the given storage system.

[0052] Figure 2The process continues in step 204: identifying an imbalance in the wear level of the storage cluster based at least in part on the determined wear levels of each of two or more storage systems of the storage cluster. Step 204 may include: determining an average of the wear levels of two or more storage systems of the storage cluster; determining a standard deviation of the wear levels of two or more storage systems of the storage cluster; and determining the wear level imbalance of the storage cluster as a ratio of the standard deviation and the average of the wear levels of two or more storage systems of the storage cluster.

[0053] In step 206, in response to the identified wear level imbalance of the storage cluster being greater than an imbalance threshold, moving one or more storage objects between two or more storage systems of the storage cluster. The first of the two or more storage systems may be part of a first distributed file system, and the second of the two or more storage systems may be part of a second distributed file system different from the first distributed file system. In some embodiments, the first of the two or more storage systems utilizes block-based storage, and the second of the two or more storage systems utilizes file-based storage, and the first storage system and the second storage system independently provide block and file storage services to each other.

[0054] Step 206 may include: selecting the first of the two or more storage systems of the storage cluster as a source storage system, and selecting the second of the two or more storage systems of the storage cluster as a destination storage system; and selecting a given storage object stored on the first storage system that will be moved to the second storage system, wherein the first storage system has a higher determined wear level than the second storage system. Selecting the given storage object may include: determining, for each of at least one subset of the storage objects stored on the first storage system, heat information characterizing the number of write requests per unit capacity; and selecting the given storage object from the subset of the storage objects stored on the first storage system based at least in part on the determined heat information. Moving one or more storage objects between two or more storage systems of the storage cluster may further include: assuming that the given storage object is moved from the first storage system to the second storage system, re-determining the wear levels of the first storage system and the second storage system; and moving the given storage object from the first storage system to the second storage system in response to the re-determined wear level of the first storage system being less than or equal to the re-determined wear level of the second storage system.

[0055] In some embodiments, step 206 includes performing two or more iterations of the following: selecting a first and a second of two or more storage systems of a storage cluster as respective source and destination storage systems to move a given storage object of one or more storage objects; determining whether a wear level of the first storage system will be less than or equal to a wear level of the second storage system if the given storage object is moved from the first storage system to the second storage system; and moving the given storage object from the first storage system to the second storage system in response to determining that the wear level of the first storage system will be less than or equal to the wear level of the second storage system. The given storage object may be selected based at least in part on: (i) the amount of input-output requests for storage objects stored on the first storage system; and (ii) the size of storage objects stored on the first storage system. The two or more iterations may continue until a given iteration in which it is determined that if the given storage object is moved from the first storage system to the second storage system, the wear level of the first storage system will be greater than the second storage system.

[0056] A storage array may implement a wear leveling mechanism to balance wear among the storage devices of the storage array (e.g., balance SSD wear among the disks in a storage array). However, from the perspective of a storage array cluster, there may be many different storage arrays. When the data stored on the storage array cluster changes, even if each of the storage arrays individually performs wear leveling for the storage devices in those storage arrays, this may result in uneven wear leveling among the different arrays in the storage array. In other words, there may be uneven wear leveling among the different arrays in the storage arrays of the cluster even if wear leveling is performed locally at each storage array in the cluster. For example, this may cause the storage devices of some of the storage arrays in the cluster to be depleted while the storage devices of other arrays in the storage arrays of the cluster are not fully utilized. As a result, the overall storage device (e.g., SSD drive) efficiency in the cluster is poor. Current wear leveling mechanisms cannot handle these scenarios and may thus have a negative impact on performance. Manual balancing across storage arrays in the cluster is also problematic because it is difficult to determine wear balance across multiple storage arrays (especially as the number of storage arrays in the cluster increases) and there is a lack of expertise.

[0057] Exemplary embodiments provide techniques for implementing inter - storage - array wear leveling in a storage array cluster. In some embodiments, the wear - leveling state of storage devices in all storage arrays in the cluster and the I / O metrics (e.g., I / O temperature) of all storage arrays in the cluster are monitored. The I / O metrics or I / O temperature can be measured or determined at various granularity levels (such as LUNs, file systems, virtual machine file systems (VMFS), etc.). Based on this data, wear leveling between the storage arrays of the cluster is implemented (e.g., by moving LUNs, file systems, VMFS, etc. between the storage arrays of the cluster) to achieve a uniformly distributed wear leveling among all storage arrays. The data migration can be performed automatically or a data migration can be recommended to a storage device administrator or other authorized users of the storage cluster in order to achieve cluster - wide wear leveling.

[0058] In some embodiments, the storage cluster - wide wear leveling (also known as inter - storage - array wear leveling) is implemented as a low - level service of a storage cluster management component that obtains all wear - leveling and I / O information of the storage arrays of the storage cluster. In this way, the storage cluster management component can use the inter - storage - array wear - leveling functionality to improve the overall efficiency of the storage devices used across different storage arrays of the storage cluster as a whole. The inter - storage - array wear - leveling functionality can be used for various different types of storage clusters, including those utilizing different types of storage arrays (e.g., different storage device products from one or more storage device vendors). Additionally, the inter - storage - array wear - leveling functionality does not require that all storage nodes or storage arrays in the storage cluster belong to the same distributed file system (e.g., Coda, Lustre, Hadoop Distributed File System (HDFS), Ceph). The inter - storage - array wear - leveling functionality supports block and file storage arrays for which there is no mapping table, where such storage arrays independently provide block and file storage services.

[0059] Figure 3 An example configuration is shown where the storage cluster wear - leveling service 112 is implemented internally within one of the storage arrays 106 - 1 that acts as a storage cluster controller for a storage cluster 300 including storage arrays 106 - 1 through 106 - M. However, it should be understood that in some embodiments, the storage cluster wear - leveling service 112 can be implemented externally to the storage arrays 106 - 1 through 106 - M (e.g., as shown in Figure 1 ), or can be at least partially implemented internally within one or more of the host devices 102, or it can itself be implemented or distributed across multiple of the host devices 102 and / or storage arrays 106. In Figure 3In an example, the storage array 106-1 implements the storage cluster wear leveling service 112 as a low-level service of the cluster management component of the storage cluster 300. In this example, the storage array 106-1 is selected from the cluster 300 to act as the cluster controller and thus runs the storage cluster wear leveling service 112.

[0060] The storage cluster wear leveling service 112 implements a series of steps (e.g., using the usage information collection module 114, the cluster-wide wear leveling determination module 116, and the storage object migration module 118) for implementing storage cluster wear leveling. In step 301, the storage cluster wear leveling service 112 uses the usage information collection module 114 to collect various usage information data from each of the storage arrays 106 that are part of the storage cluster. The collected usage information data may include, but is not limited to: the wear level status of the storage devices in each of the storage arrays 106; the capacity usage of the storage devices in each of the storage arrays 106; the IO temperature of the storage devices in each of the storage arrays 106; the sizes of various storage objects (e.g., LUNs, file systems, VMFS, etc.) stored on the storage devices in each of the storage arrays; the IO temperature of the storage objects, etc.

[0061] In step 303, the storage cluster wear leveling service 112 uses the cluster-wide wear leveling determination module 116 to use the data collected in step 301 to calculate and evaluate whether the wear level in the storage cluster 300 is an undesirable condition (e.g., a certain specified threshold uneven wear level). If the result is yes, step 303 also includes determining how to achieve cluster-wide wear leveling across the storage arrays 106 of the storage cluster 300 using storage objects or other data migration solutions. Rebalancing algorithms described in further detail elsewhere in this document may be used to determine the storage objects or other data migration solutions.

[0062] In step 305, the storage cluster wear leveling service 112 uses the storage object migration module 118 to migrate storage objects between the storage arrays 106 of the storage cluster 300 to achieve the desired cluster-wide wear leveling. Step 305 may include: automatically moving storage objects between the storage arrays 106; providing guidance to the storage device administrator of the storage cluster 300 on how to perform manual data migration to achieve the desired cluster-wide wear leveling; a combination thereof, etc. Storage object migration may be handled via federated live migration (FLM), replication, in-band migration tool (IMT), or other data migration tools.

[0063] An algorithm for evaluating and balancing wear leveling of a storage cluster such as storage cluster 300 will now be described. Assume that the storage cluster includes M storage arrays, and each storage array has N storage devices (also referred to as disks). However, it should be understood that the value of "N" may be different for each of the storage arrays in the storage cluster. In other words, it is not required that different storage arrays have an equal number of disks. Capacity usage is represented by C, IO temperature is represented by T, and wear level (e.g., write request count of an SSD disk) is represented by W. These three criteria C, T, and W can be used to measure the wear level of a storage array. The array wear level state in the storage cluster can be calculated by combining C, T, and W. Then, the standard deviation of the wear levels of the storage arrays in the storage cluster is calculated and used to measure the imbalance rate of the storage cluster. Once the storage cluster-wide wear level reaches a specified imbalance rate threshold, candidate source and destination storage arrays for data or storage object migration and data migration targets are determined.

[0064] C 磁盘 Is used to represent the capacity usage of a disk, and C 阵列 Represents the sum of the disk capacity usages of the disks in a storage array, which is calculated as The higher the capacity usage of a storage array, the more wear the storage array is considered to have, and thus fewer storage objects should be moved to the storage array. As described above, N represents the number of disks on a storage array. T 磁盘 Represents the IO temperature of a disk, which can be calculated by the current MCR component, and T 阵列 Represents the sum of the IO temperatures of the disks in a storage array, and the sum of the IO temperatures is calculated according to W 磁盘 Represents the write request count of a disk, which reflects the wear level of the disk, and W 阵列 Represents the sum of the write request counts of the disks in a storage array. W 磁盘 The larger W is, the closer the disk is to being exhausted. W 磁盘 Can be calculated by the current MCR component. W 阵列 Can be calculated according to To calculate.

[0065] The wear level of a storage array can be measured according to the following:

[0066]

[0067] The value R of storage array i i The larger R is, the more aged storage array i is. R i Is determined as a combination of the three criteria of capacity usage, IO temperature, and degree of wear level. ω c ω T And ω Wrespectively represent the capacity usage, IO temperature, and wear level criteria, and ω c + ω T + ω W = 1. By tuning the weights ω c , ω T and ω W values, a better overall balance result can be achieved. R 平均 represents the average wear level of the storage arrays in the storage cluster and can be calculated according to .

[0068] The standard deviation of the wear level of the storage arrays in the storage cluster is denoted as σ and can be calculated according to . The standard deviation σ can be regarded as or used as a measure of the imbalance rate of the storage cluster. A low standard deviation indicates that the wear levels of the arrays tend to be close to the average of the set (also known as the expected value), while a high standard deviation indicates wear imbalance in the storage cluster and thus the need for rebalancing. λ represents the imbalance rate of the storage cluster and is calculated according to . θ represents the acceptable threshold of the imbalance rate of the storage cluster. If in the storage cluster, the value of λ is equal to or greater than θ, rebalancing is required. S 对象 represents the size of the storage object, and T 对象 represents the temperature of the storage object measured by write requests. H 对象k represents the "hotness" of the storage object measured by write requests per unit capacity and is calculated according to .

[0069] In some embodiments, the storage cluster wear leveling service 112 is configured to periodically evaluate the current cluster-wide wear level by calculating the imbalance rate λ of the storage cluster. After the imbalance rate λ exceeds the specified threshold θ (e.g., λ ≥ θ), the storage cluster wear leveling service 112 will select a source storage array and a destination storage array and move optimized storage objects (e.g., LUNs, file systems, VMFS, etc.) with high IO temperatures from the source storage array to the destination storage array to balance the wear level of the storage cluster. The value of θ can be tuned as needed (such as based on the expected wear level of the storage cluster). If it is expected or desired that the storage cluster has a low imbalance wear level, θ can be set to a small value (e.g., about 10% to 20%) to trigger the rebalancing algorithm earlier. If it is not desired to trigger the rebalancing algorithm earlier, θ can be set to a larger value (e.g., about 20% to 40%).

[0070] The storage device wear level (e.g., SSD or flash drive wear level) is the result of long-term I / O loads and represents the wear state of the storage device rather than a transient load state. When the storage cluster reaches the imbalance rate threshold, the rebalancing algorithm is triggered. After the rebalancing algorithm is completed (e.g., after storage objects are moved between storage arrays of the storage cluster), the imbalance rate of the storage cluster will drop to a small value (e.g., typically less than 5%). Due to the efficiency of the rebalancing algorithm and the fact that the rebalancing algorithm is triggered based on cumulative I / O loads rather than transient I / O loads, the rebalancing algorithm may be triggered relatively infrequently (e.g., the time taken from a small imbalance rate of 5% to an imbalance threshold of 20%). As described above, the imbalance threshold can be tuned as needed to avoid frequent triggering of rebalancing.

[0071] It should be understood that the storage cluster wear leveling service 112 can be deployed on the storage cluster when the storage cluster is first created, or can be deployed at some point after it is expected that rebalancing the wear level may be beneficial (e.g., deployed to a mid-life storage cluster). However, the rebalancing algorithm is only triggered when the imbalance rate threshold is reached. To avoid the rebalancing algorithm affecting the user I / O of the storage cluster, the storage object movement task can be set to a lower priority and / or run as a background process on the cluster management node to reduce or eliminate such effects. In this way, the storage cluster can be kept in a useful state while the wear level rebalancing is in progress. Additionally, as detailed above, the combination of the efficiency of rebalancing and appropriately setting the imbalance rate threshold results in infrequent triggering of the rebalancing algorithm (e.g., after rebalancing, it takes time for the wear leveling to reach the imbalance rate threshold).

[0072] The steps for evaluating and balancing the storage cluster-wide wear level will now be described in detail. The first step includes calculating the wear level R of the arrays in the storage cluster and obtaining the source storage array and the destination storage array for storage object migration. For each storage array in the storage cluster i Calculate R i . Select the array with the maximum wear level R 最大 as the source storage array, which is the most severely worn storage array. If there are multiple storage arrays with the same maximum value, one of such storage arrays can be randomly selected as the source storage array. Select the array with the minimum wear level R 最小 as the destination storage array. If there are multiple storage arrays with the same minimum value, one of such storage arrays can be randomly selected as the destination storage array. Then, as described above, calculate R using the normalization criterion according to the following equation i :

[0073]

[0074] The next step is to calculate the heat of the source storage objects and obtain the target storage objects that will be moved from the source storage array to the destination storage array. Sort the storage objects residing in the source storage array and then select the object with the maximum heat H 最大 as the target storage object to be moved from the source storage array to the destination storage array. According to Calculate the heat criterion H by dividing the IO temperature (e.g., write request) of the storage object by the size of the storage object. The value of H 对象k is the write request per unit capacity. The larger the value of H 对象k , the hotter the storage object k, such that moving this storage object yields better results at lower cost.

[0075] Assume that a storage object with H 最大 is moved from the source storage array to the destination storage array, then recalculate the loss level R 源 of the source storage array and the loss level R 目的地 of the target storage array. If R 源 ≤R 目的地 , then there is no need to migrate the storage object, and the rebalancing algorithm can be ended to avoid over-adjustment. If R 源 >R 目的地 , then move the object with H 最大 from the source storage array to the destination storage array. Then repeat this process as needed to recalculate the new source storage array and destination storage array, and then recalculate the heat of the storage objects on the new source storage array to potentially migrate the hottest storage object from the new source storage array to the new destination storage array. In this algorithm, the loss levels of the storage arrays in the storage cluster are evaluated, and hot objects are moved from the storage array with a higher loss level to the storage array with a lower loss level. The result of this algorithm is to balance the cluster-wide loss level to ensure the durability of the storage devices across the storage arrays of the storage cluster.

[0076] Figure 4 The process flow of the above algorithm is shown. This process flow starts in step 401 or is initiated in response to a certain condition (such as a specified threshold time period since the last rebalancing of the storage cluster), in response to an explicit user request to rebalance the storage cluster, according to a schedule, etc. (e.g., initiated by the storage cluster loss balancing service 112). In step 403, calculate the loss level R of the storage arrays in the storage cluster and the imbalance rate λ of the storage cluster. In step 405, determine whether λ≥θ. If the result determined in step 405 is "yes", the process continues to step 407. If the result determined in the step is "no", the process ends in step 417.

[0077] In step 407, select the storage array with R 最大 in the storage cluster as the source storage array, and select the storage array with R 最小 in the storage cluster as the destination storage array. On the source storage array, in step 409, calculate the heat of each storage object, and select the hottest storage object with H 最大 as the target storage object. In step 411, assuming that the target storage object with H 最大 is moved to the destination storage array, recalculate R 源 and R 目的地 . In step 413, determine whether R 源 ≤R 目的地 in the recalculation in step 411. If the result determined in step 413 is "yes", the process ends in step 417. If the result determined in step 413 is "no", the process proceeds to step 415, where the target storage object is moved or migrated from the source storage array to the destination storage array. Then, in step 417, the process ends. It should be understood that the rebalancing algorithm may include iterating through steps 401 to 417 multiple times to move or migrate multiple storage objects between different groups of source storage arrays and destination storage arrays.

[0078] An example of storage cluster-wide loss balancing using the Figure 4 process will now be described. Assume there is a storage cluster with three storage arrays (e.g., three all-flash storage arrays) that are interconnected with each other using a SAN. Also assume that after a period of IO activity, the Figure 4 algorithm is applied, where the current capacity usage, IO temperature, and loss status are shown in Figure 5 Table 500. The first storage array (Array 1) acts as the cluster controller (running the storage cluster loss balancing service 112). Also assume that the capacity of each of the storage arrays is 80 terabytes (TB), and the imbalance rate threshold is set to Θ = 20%. Considering that the temperature of the storage array will affect future losses, and the loss of the storage array reflects the current loss status, ω T and ω W are set to have relatively high weights, and ω C is set to have a relatively small weight (e.g., ω C = 20%, ω T = 40%, and ω W = 40%). According to the above equation, calculate R i to obtain as shown in Figure 5The normalized storage array loss levels shown in Table 505. Based on these results, it is determined that the third storage array (Array 3) is the array with the greatest loss, while the first storage array (Array 1) is the array with the least loss. After calculating σ = 15.27%, the imbalance rate of the storage cluster is λ = 41.33%. Since this value is greater than the imbalance rate threshold θ = 20%, rebalancing is triggered for the storage arrays of the storage cluster.

[0079] Figure 6 Shows the storage array loss status of the storage cluster before rebalancing the storage cluster range loss level. Figure 6 Shows three storage arrays 601-1, 601-2, and 601-3, and the associated storage array loss statuses 610-1, 610-2, and 610-3. Figure 7 Shows Tables 700, 705, and 710, illustrating the distribution of storage objects across three storage arrays, including the sizes, IO temperatures, and hotness of different storage objects stored on each of the storage arrays. As Figure 6 and Figure 7 indicated, the third storage array (Array 3, 601-3) is the array with the greatest loss. There are nine storage objects on the third storage array, and the hotness H of each of the storage objects is calculated 对象 , and then the storage objects are sorted from largest to smallest by the value of H 对象 .

[0080] Then, storage cluster range loss equalization begins rebalancing. The source storage array, destination storage array, and target moving object for five iterations of rebalancing will now be described.

[0081] In the first iteration, the source storage array is the third storage array (Array 3, 601-3), the destination storage array is the first storage array (Array 1, 601-1), and the target object is Object 6, which is proposed to be moved from the third storage array to the first storage array, resulting in the recalculated values R 阵列1 = 0.24139, R 阵列2 = 0.34344, R 阵列3 = 0.41516.

[0082] In the second iteration, the source storage array is the third storage array (Array 3, 601-3), the destination storage array is the first storage array (Array 1, 601-1), and the target object is Object 8, which is proposed to be moved from the third storage array to the first storage array, resulting in the recalculated values R 陈列1 = 0.26411, R 陈列2 = 0.34344, R 陈列3 = 0.39244.

[0083] In the third iteration, the source storage array is the third storage array (Array 3, 601-3), the destination storage array is the first storage array (Array 1, 601-1), and the target object is Object 2. It is proposed to move this object from the third storage array to the first storage array, resulting in a recalculated value of R 阵列1 = 0.29242, R 阵列2 = 0.34344, R 阵列3 = 0.36413.

[0084] In the fourth iteration, the source storage array is the third storage array (Array 3, 601-3), the destination storage array is the first storage array (Array 1, 601-1), and the target object is Object 4. It is proposed to move this object from the third storage array to the first storage array, resulting in a recalculated value of R 阵列1 = 0.31944, R 阵列2 = 0.343442, R 阵列3 = 0.337118.

[0085] In the fifth iteration, the source storage array is the second storage array (Array 2, 601-2), the destination storage array is the first storage array (Array 1, 601-1), and the target object is Object 4. It is proposed to move this object from the second storage array to the first storage array, resulting in a recalculated value of R 阵列1 = 0.36356, R 阵列2 = 0.30596, R 阵列3 = 0.33048.

[0086] After the fifth iteration, R 阵列2 < R 阵列1 And thus the algorithm reaches the end condition and there is no need to perform the storage object movement. Figure 8 Table 800 showing the guidance for rebalancing the storage cluster is given, presenting the iteration number, source storage array, destination storage array, and storage object to be moved. Figure 8 Table 805 is also shown, illustrating the loss level status of the storage arrays in the storage cluster after rebalancing. The imbalance rate after rebalancing is recalculated as λ = 0.03239565 = ~3.2%. Figure 9 The storage array loss status of the storage cluster after rebalancing the storage cluster range loss level is shown, where the three storage arrays 601-1, 601-3, and 601-3 have updated storage array loss statuses 910-1, 910-2, and 910-3 respectively.

[0087] Some storage clusters may mirror data across storage arrays or storage nodes. If moving a storage object will affect such data mirroring, then the contribution to wear is counted, but when selecting a LUN, file system, or other storage object for cluster-wide rebalancing, such mirrored LUNs, file systems, or other storage objects will not be selected as targets. If moving a mirrored LUN, file system, or other storage object will not affect the mirroring, then the algorithm can proceed as normal.

[0088] The techniques described herein also support scenarios for moving storage arrays between storage clusters. If a new storage cluster is created using several different storage arrays that previously resided in other storage clusters, then the storage arrays of the new storage cluster will accurately share information about their respective wear levels. If a storage array is being moved from a first storage cluster to a second storage cluster, then the second storage cluster will obtain information about the storage array (e.g., LUN, file system, and other storage object IO temperature and storage device wear status). Then, the algorithm can be run to perform storage cluster-wide wear leveling on the second storage cluster using the information obtained from the storage array newly introduced to the second storage cluster. For the scenario of creating a new storage cluster from several storage arrays that previously resided in other storage clusters, the algorithm will be used to obtain information (e.g., LUN, file system, or other storage object IO temperature and storage device wear status) from each of the storage arrays in the newly created storage cluster. Then, the algorithm can be run to perform storage cluster-wide wear leveling on the newly created storage cluster using the information obtained from each of the storage arrays.

[0089] Advantageously, the illustrative embodiments utilize a statistics-based method for optimizing wear at the storage cluster level, which significantly improves the overall cluster-wide storage device (e.g., SSD) efficiency. The techniques described are not limited to storage clusters where all storage arrays or nodes belong to the same distributed file system. Rather, the techniques described support block and file storage arrays that have no requirement for a mapping table and that independently provide block and file storage services. Performing wear leveling on these and other types of storage clusters is difficult.

[0090] It should be understood that the specific advantages described above and elsewhere in this document are associated with specific illustrative embodiments and may not necessarily exist in other embodiments. Moreover, the specific types of information processing system features and functionality shown in the figures and described above are merely exemplary, and numerous other arrangements may be used in other embodiments.

[0091] Reference will now be made to Figure 10 and Figure 11Descriptive embodiments of a processing platform are described in more detail, which is utilized to implement functionality for performing wear leveling between storage systems of a storage cluster. Although described in the context of system 100, in other embodiments, these platforms can also be used to implement at least several parts of other information processing systems.

[0092] Figure 10 An example processing platform including cloud infrastructure 1000 is shown. Cloud infrastructure 1000 includes a combination of physical processing resources and virtual processing resources, which can be utilized to implement Figure 1 at least a part of the information processing system 100 in. Cloud infrastructure 1000 includes a plurality of virtual machines (VMs) and / or container sets 1002-1, 1002-2,......, 1002-L implemented using virtualization infrastructure 1004. Virtualization infrastructure 1004 runs on physical infrastructure 1005 and illustratively includes one or more hypervisors and / or operating system-level virtualization infrastructure. Operating system-level virtualization infrastructure illustratively includes kernel control groups of a Linux operating system or other types of operating systems.

[0093] Cloud infrastructure 1000 also includes application sets 1010-1, 1010-2,......, 1010-L that run on corresponding VM / container sets in VM / container sets 1002-1, 1002-2,......, 1002-L under the control of virtualization infrastructure 1004. VM / container sets 1002 can include corresponding VMs, corresponding one or more container sets, or corresponding one or more container sets running in VMs.

[0094] In Figure 10 some implementations of the embodiments of, VM / container sets 1002 include corresponding VMs implemented using virtualization infrastructure 1004 including at least one hypervisor. A hypervisor platform can be used to implement a hypervisor within virtualization infrastructure 1004, where the hypervisor platform has an associated virtual infrastructure management system. The underlying physical machine can include one or more distributed processing platforms, which include one or more storage systems.

[0095] In Figure 10 other implementations of the embodiments of, VM / container sets 1002 include corresponding containers implemented using virtualization infrastructure 1004 that provides operating system-level virtualization functionality (such as support for Docker containers running on bare metal hosts or Docker containers running on VMs). Containers are illustratively implemented using corresponding kernel control groups of the operating system.

[0096] As is apparent from above, one or more of the processing modules or other components of system 100 may each operate on a computer, server, storage device, or other processing platform element. A given such element may be considered an example of what is more generally referred to herein as a "processing device." Figure 10 The cloud infrastructure 1000 shown in Figure 11 may represent at least a portion of a processing platform. Another example of such a processing platform is

[0097] the processing platform 1100 shown in. In this embodiment, the processing platform 1100 includes a portion of system 100 and includes a plurality of processing devices represented as 1102-1, 1102-2, 1102-3,......, 1102-K, which communicate with each other via a network 1104.

[0098] The network 1104 may include any type of network, including, for example, a global computer network (such as the Internet), WAN, LAN, satellite network, telephone or wired network, cellular network, wireless network (such as a WiFi or WiMAX network), or various portions or combinations of these and other types of networks.

[0099] The processing device 1102-1 in the processing platform 1100 includes a processor 1110 coupled to a memory 1112.

[0100] The processor 1110 may include a microprocessor, microcontroller, application specific integrated circuit (ASIC), field programmable gate array (FPGA), central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), video processing unit (VPU), or other type of processing circuit, as well as portions or combinations of such circuit elements.

[0101] The memory 1112 may include random access memory (RAM), read only memory (ROM), flash memory, or other types of memory in any combination. The memory 1112 and other memories disclosed herein should be considered illustrative examples of what is more generally referred to as a "processor-readable storage medium" that stores executable program code for one or more software programs.

[0102] An article of manufacture including such a processor-readable storage medium is considered an illustrative embodiment. A given such article of manufacture may include, for example, a storage array, a storage disk, or an integrated circuit containing RAM, ROM, flash memory, or other electronic memory, or any of a variety of other types of computer program products. As used herein, the term "article of manufacture" should be understood to exclude transient propagated signals. Numerous other types of computer program products including a processor-readable storage medium may be used.

[0103] The processing device 1102-1 further includes a network interface circuit 1114, which is used to interface the processing device with the network 1104 and other system components, and may include a conventional transceiver.

[0104] It is assumed that the other processing devices 1102 of the processing platform 1100 are configured in a manner similar to that shown for the processing device 1102-1 in the figure.

[0105] Moreover, the specific processing platform 1100 shown in the figure is presented only by way of example, and the system 100 may include additional or alternative processing platforms, as well as numerous different processing platforms in any combination, where each such platform includes one or more computers, servers, storage devices, or other processing devices.

[0106] For example, other processing platforms for implementing the illustrative embodiments may include converged infrastructure.

[0107] Therefore, it should be understood that in other embodiments, different arrangements of additional or alternative elements may be used. At least a subset of these elements may be collectively implemented on a common processing platform, or each such element may be implemented on a separate processing platform.

[0108] As indicated above, the components of the information processing system as disclosed herein may be at least partially implemented in the form of one or more software programs stored in a memory and executed by a processor of a processing device. For example, at least a part of the functionality for performing wear leveling between storage systems of a storage cluster as disclosed herein is illustratively implemented in the form of software running on one or more processing devices.

[0109] It should be emphasized again that the above embodiments are presented for illustrative purposes only. Many variations and other alternative embodiments may be used. For example, the disclosed technology may be applicable to a wide variety of other types of information processing systems, storage systems, storage clusters, etc. Moreover, the specific configurations of the system and device elements illustratively shown in the drawings and the associated processing operations may vary in other embodiments. In addition, the various assumptions made above in the process of describing the illustrative embodiments should also be regarded as exemplary, rather than requirements or limitations of the present disclosure. Numerous other alternative embodiments within the scope of the appended claims will be apparent to those skilled in the art.

Claims

1. A device, the device comprising: at least one processing device, the at least one processing device including a processor coupled to a memory; the at least one processing device being configured to perform the following steps: obtain usage information of each of two or more storage systems of a storage cluster; determine a wear level of each of the two or more storage systems of the storage cluster at least in part based on the obtained usage information, wherein the wear level of a given storage system among the two or more storage systems is determined at least in part based on a combination of the following: (i) capacity usage information of the given storage system; (ii) input-output temperature information characterizing the number of input-output requests of the given storage system within a specified time period; and (iii) cumulative write request count information of the given storage system; identify a wear level imbalance of the storage cluster at least in part based on the determined wear levels of each of the two or more storage systems of the storage cluster; and move one or more storage objects between the two or more storage systems of the storage cluster at least in part based on the identified wear level imbalance of the storage cluster being greater than an imbalance threshold; wherein moving the one or more storage objects between the two or more storage systems of the storage cluster is also at least in part based on a determination of the expected changes in the following: a first wear level of a first storage system among the two or more storage systems of the storage cluster due to moving the one or more storage objects; and a second wear level of a second storage system among the two or more storage systems of the storage cluster due to moving the one or more storage objects; and wherein identifying the wear level imbalance of the storage cluster includes: determining an imbalance rate of the storage cluster at least in part according to a function of a first statistic of the wear levels of the two or more storage systems of the storage cluster and a second statistic of the wear levels of the two or more storage systems of the storage cluster, the second statistic being different from the first statistic.

2. The apparatus according to claim 1, wherein determining the wear level of the given storage system comprises: Calculate a weighted sum of the capacity usage information of the given storage system, the input-output temperature information of the given storage system, and the cumulative write request count information of the given storage system.

3. The device according to claim 2, wherein a first weight assigned to the capacity usage information of the given storage system is lower than a second weight assigned to the input-output temperature information and a third weight assigned to the cumulative write request count information of the given storage system.

4. The device according to claim 1, wherein moving the one or more storage objects between the two or more storage systems of the storage cluster comprises: Select the first storage system among the two or more storage systems of the storage cluster as a source storage system, and select the second storage system among the two or more storage systems of the storage cluster as a destination storage system; and select a given storage object stored on the first storage system to be moved to the second storage system.

5. The device according to claim 4, wherein the first storage system has a higher determined loss level than the second storage system.

6. The device according to claim 4, wherein selecting the given storage object comprises: determining, for each of at least one subset of storage objects stored on the first storage system, heat information representative of the number of write requests per unit of capacity; and selecting the given storage object from the subset of storage objects stored on the first storage system at least in part based on the determined heat information.

7. The device according to claim 4, wherein moving the one or more storage objects between the two or more storage systems of the storage cluster further comprises: determining a first loss level of the first storage system and a second loss level of the second storage system assuming that the given storage object is moved from the first storage system to the second storage system; and moving the given storage object from the first storage system to the second storage system at least in part based on the determined first loss level of the first storage system being greater than the determined second loss level of the second storage system.

8. The device according to claim 1, wherein moving the one or more storage objects between the two or more storage systems comprises performing two or more iterations of: selecting a first storage system and a second storage system from the two or more storage systems of the storage cluster as a respective source storage system and destination storage system to move at least one of the one or more storage objects; determining whether a first loss level of the first storage system will be less than or equal to a second loss level of the second storage system if the at least one storage object is moved from the first storage system to the second storage system; and moving the at least one storage object from the first storage system to the second storage system at least in part based on determining that the first loss level of the first storage system will be greater than the second loss level of the second storage system.

9. The device according to claim 8, wherein the at least one storage object is selected at least in part based on: (i) the amount of input-output requests for the storage objects stored on the first storage system; and (ii) the size of the storage objects stored on the first storage system.

10. The device according to claim 8, wherein the two or more iterations are continued until a given iteration in which it is determined that if the at least one storage object is moved from the first storage system to the second storage system, the first loss level of the first storage system will be greater than the second loss level of the second storage system.

11. The apparatus according to claim 1, wherein a first storage system among the two or more storage systems is part of a first distributed file system, and a second storage system among the two or more storage systems is part of a second distributed file system different from the first distributed file system.

12. The apparatus according to claim 1, wherein a first storage system among the two or more storage systems utilizes block-based storage, and a second storage system among the two or more storage systems utilizes file-based storage, and wherein the first storage system and the second storage system independently provide block and file storage services to each other.

13. An apparatus, the apparatus comprising: at least one processing device, the at least one processing device including a processor coupled to a memory; the at least one processing device is configured to perform the following steps: obtain usage information of each of two or more storage systems of a storage cluster; determine a wear level of each of the two or more storage systems of the storage cluster at least in part based on the obtained usage information; identify a wear level imbalance of the storage cluster at least in part based on the determined wear levels of each of the two or more storage systems of the storage cluster; and move one or more storage objects between the two or more storage systems of the storage cluster at least in part based on the identified wear level imbalance of the storage cluster being greater than an imbalance threshold; wherein moving the one or more storage objects between the two or more storage systems of the storage cluster is also at least in part based on a determination of expected changes in: a first wear level of a first storage system among the two or more storage systems of the storage cluster due to moving the one or more storage objects; and a second wear level of a second storage system among the two or more storage systems of the storage cluster due to moving the one or more storage objects; wherein identifying the wear level imbalance of the storage cluster includes: determining an imbalance rate of the storage cluster at least in part according to a function of a first statistic of the wear levels of the two or more storage systems of the storage cluster and a second statistic of the wear levels of the two or more storage systems of the storage cluster, the second statistic being different from the first statistic; wherein the first statistic includes: an average value of the wear levels of the two or more storage systems of the storage cluster; wherein the second statistic includes: a standard deviation of the wear levels of the two or more storage systems of the storage cluster; and wherein the function includes: a ratio of the standard deviation and the average value of the wear levels of the two or more storage systems of the storage cluster.

14. A computer program product, the computer program product comprising a non-transitory processor-readable storage medium storing program code of one or more software programs, wherein the program code, when executed by at least one processing device, causes the at least one processing device to perform the following steps: Obtain usage information of each of two or more storage systems of a storage cluster; Determine a wear level of each of the two or more storage systems of the storage cluster at least in part based on the obtained usage information, wherein the wear level of a given storage system among the two or more storage systems is determined at least in part based on a combination of the following: (i) capacity usage information of the given storage system; (ii) input-output temperature information characterizing the number of input-output requests of the given storage system within a specified time period; And (iii) cumulative write request count information of the given storage system; Identify a wear level imbalance of the storage cluster at least in part based on the determined wear levels of each of the two or more storage systems of the storage cluster; And Move one or more storage objects between the two or more storage systems of the storage cluster at least in part based on the identified wear level imbalance of the storage cluster being greater than an imbalance threshold; Wherein moving the one or more storage objects between the two or more storage systems of the storage cluster is also at least in part based on a determination of the following predicted changes: a first wear level of a first storage system among the two or more storage systems of the storage cluster resulting from moving the one or more storage objects; and a second wear level of a second storage system among the two or more storage systems of the storage cluster resulting from moving the one or more storage objects; and Wherein identifying the wear level imbalance of the storage cluster includes: determining an imbalance rate of the storage cluster at least in part according to a function of a first statistic of the wear levels of the two or more storage systems of the storage cluster and a second statistic of the wear levels of the two or more storage systems of the storage cluster, the second statistic being different from the first statistic.

15. The computer program product according to claim 14, wherein moving the one or more storage objects between the two or more storage systems includes performing two or more iterations of the following: Select a first storage system and a second storage system among the two or more storage systems of the storage cluster as a corresponding source storage system and destination storage system to move at least one storage object among the one or more storage objects; Determine whether a first wear level of the first storage system will be less than or equal to a second wear level of the second storage system if the at least one storage object is moved from the first storage system to the second storage system; And Moving the at least one storage object from the first storage system to the second storage system is at least partially based on determining that the first wear level of the first storage system will be greater than the second wear level of the second storage system.

16. The computer program product according to claim 15, wherein the at least one storage object is selected at least partially based on: (i) the amount of input-output requests for the storage objects stored on the first storage system; and (ii) the size of the storage objects stored on the first storage system.

17. A method, the method comprising: Obtaining usage information for each of two or more storage systems of a storage cluster; Determining a wear level for each of the two or more storage systems of the storage cluster at least partially based on the obtained usage information, wherein the wear level for a given storage system of the two or more storage systems is determined at least partially based on a combination of: (i) capacity usage information for the given storage system; (ii) input-output temperature information characterizing the number of input-output requests for the given storage system over a specified period of time; and (iii) cumulative write request count information for the given storage system; Identifying a wear level imbalance for the storage cluster at least partially based on the determined wear levels for each of the two or more storage systems of the storage cluster; and Moving one or more storage objects between the two or more storage systems of the storage cluster at least partially based on the identified wear level imbalance for the storage cluster being greater than an imbalance threshold; wherein moving the one or more storage objects between the two or more storage systems of the storage cluster is also at least partially based on determining a projected change in: a first wear level for a first storage system of the two or more storage systems of the storage cluster resulting from moving the one or more storage objects; and a second wear level for a second storage system of the two or more storage systems of the storage cluster resulting from moving the one or more storage objects; and wherein identifying the wear level imbalance for the storage cluster comprises: determining an imbalance rate for the storage cluster at least partially based on a function of a first statistic of the wear levels of the two or more storage systems of the storage cluster and a second statistic of the wear levels of the two or more storage systems of the storage cluster, the second statistic being different from the first statistic; wherein the method is performed by at least one processing device, the at least one processing device comprising a processor coupled to a memory.

18. The method according to claim 17, wherein moving the one or more storage objects between the two or more storage systems comprises performing two or more iterations of the following: Select the first storage system and the second storage system among the two or more storage systems of the storage cluster as the corresponding source storage system and destination storage system respectively to move at least one of the one or more storage objects; Determine whether the first loss level of the first storage system will be less than or equal to the second loss level of the second storage system if the at least one storage object is moved from the first storage system to the second storage system; And Move the at least one storage object from the first storage system to the second storage system at least partially based on determining that the first loss level of the first storage system is greater than the second loss level of the second storage system.

19. The method according to claim 18, wherein the at least one storage object is selected at least partially based on: (i) the amount of input-output requests for the storage objects stored on the first storage system; and (ii) the size of the storage objects stored on the first storage system.

Citation Information

Patent Citations

  • Storage device

    US20170003889A1