Management of storage system with wear balancing

By assessing and managing wear conditions across nodes in hyperconverged storage systems, the method balances wear patterns, extends device lifespan, and reduces maintenance costs.

WO2025091359A1PCT designated stage expired Publication Date: 2025-05-08LENOVO ENTERPRISE SOLUTIONS (SINGAPORE) PTE LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/129218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing hyperconverged storage systems (HCSS) do not effectively manage device wear across nodes in the cluster, leading to uneven wear patterns and unnecessary maintenance and costs due to premature drive failures.

Method used

A method that assesses the wear condition of monitored storage devices, generates a target level of accessing activities based on the wear condition, and controls data movement within the compute cluster to balance wear patterns across nodes.

Benefits of technology

This approach extends the lifespan of storage devices by evenly distributing wear patterns, reducing maintenance costs, and preventing premature failures within the hyperconverged storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023129218_08052025_PF_FP_ABST
    Figure CN2023129218_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A method includes assessing wear condition of a monitored storage device (10) within a compute cluster (100) (102), generating a target level of accessing activities for accessing at least the monitored storage device (10) according to at least the wear condition of the monitored storage device (10) (104), and controlling data movement within the compute cluster (100) according to the target level of accessing activities (106).
Need to check novelty before this filing date? Find Prior Art

Description

MANAGEMENT OF STORAGE SYSTEM WITH WEAR BALANCINGTECHNICAL FIELD

[0001] The present disclosure generally relates to the field of compute cluster technologies and, more particularly, to the management of computer storage devices.BACKGROUND

[0002] Computer storage devices such as solid-state and spinning disk drives all suffer from wear-out effects. To optimize and balance storage workloads, drive controllers and Redundant Array of Individual Devices / Drives (RAID) controllers are used to implement control algorithms to move data periodically to ensure even wear and therefore longest possible drive lifespan. In hyperconverged environments, software moves data between drives across a set of servers for locality. This can result in “hot spots” where data movements occur frequently within the hyperconverged cluster, resulting in locally increased wear-out in comparison to peer nodes.

[0003] In the existing practice, the management of hyperconverged storage systems (HCSS) often considers storage workloads or capacity for workload provisioning but does not factor device wear across nodes in the cluster of storage devices directly. However, device wear often is related to the process of writing and reading, particularly of writing activities, rather than the storage workload itself. Techniques for balancing workload placement in the cluster act as a proxy for wear placement but can result in imbalance when workloads within the cluster have varying disk input-output (IO) profiles. Since hyperconverged storage software provides for data redundancy within the cluster, drives that wear out before the cluster is retired are simply replaced. This incurs unnecessary maintenance and material cost.

[0004] SUMMARY OF THE DISCLOSURE

[0005] One aspect of the present disclosure provides a method including assessing wear condition of a monitored storage device within a compute cluster, generating a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device, and controlling data movement within the compute cluster according to the target level of accessing activities.

[0006] Another aspect of the present disclosure provides a computer storage system including a plurality of storage devices and a storage controller coupled with the plurality of storage devices and configured to assess wear condition of a monitored storage device of the plurality of storage devices, generate a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device, and control data movement of the computer storage system according to the target level of assessing activities.

[0007] Yet another aspect of the present disclosure provides a non-transitory computer readable storage medium having program instructions stored therewith, the program instructions being executable by a processor to cause the processor to perform a method including assessing wear condition of a monitored storage device within a compute cluster, generating a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device, and controlling data movement within the compute cluster according to the target level of accessing activities.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The following drawings are provided for illustrative purposes according to various disclosed embodiments and are not intended to limit the scope of the present disclosure.

[0009] FIG. 1A is a schematic diagram of a hyperconverged storage system according to an embodiment of the present disclosure.

[0010] FIG. 1B is a flowchart showing a method of storage management according to an embodiment of the present disclosure.

[0011] FIG. 2 is a flowchart showing a method of storage management by identifying storage exceeding an I / O activity threshold according to an embodiment of the present disclosure.

[0012] FIG. 3 is a flowchart showing a method of storage management by generating a profile of I / O activities across a plurality of storage devices in a plurality of nodes according to an embodiment of the present disclosure.

[0013] FIG. 4 is a flowchart showing a method of storage management by identifying a pattern of wear conditions among storage devices according to an example embodiment of the present disclosure.

[0014] FIG. 5 is a flowchart showing a method of storage management by reducing storage I / O activities in storage devices identified as having excessive wear conditions according to an example embodiment of the present disclosure.

[0015] FIG. 6 is a flowchart showing generating types of threshold used by the method of storage management according to an embodiment of the present disclosure.

[0016] FIG. 7 is a block diagram of an example apparatus of the compute cluster with the storage management according to an example embodiment of the present disclosure.

[0017] FIG. 8 is a block diagram of a computer program product according to an example embodiment of the present disclosure.DETAILED DESCRIPTION

[0018] The computer storage device referred to in the present disclosure can be any type of storage device suitable for storing data, such as a random access memory, a non-volatile read-only memory, a hard drive disk, a solid-state disk, a compact disk, a floppy disk, or a magnet tape. In this disclosure, “computer storage device, ” “storage device, ” and “storage” are all interchangeably used.

[0019] FIG. 1A shows a hyper-converged storage system (HCSS) 1, also referred to as a “storage pool, ” consistent with embodiments of the present disclosure. As shown in FIG. 1A, the HCSS 1 includes a plurality of storage nodes 100 and a storage workload server or controller 50, and at least one router 52 connecting the plurality of nodes 100 with the storage workload controller 50. In some embodiments, as shown in FIG. 1A, the HCSS 1 further includes file servers located on a compute cluster 200 (such as the Internet) , and the at least one router 52 is further connected to the compute cluster 200.

[0020] Each of the plurality of storage nodes 100 may be a storage system in a larger cluster of shared resources and, as shown in FIG. 1A, includes one or more storage devices 10, such as one or more hard drives and / or one or more solid state drives. Each storage node 100 further includes a local area network (LAN) connector 30, and a node input / output (I / O) log 20,  also referred to as a write / read log, for recording I / O activities of the one or more storage devices 10 of node 100. In this disclosure, “storage node, ” “node of storage, ” and “node” are used interchangeably.

[0021] The operation of the storage workload controller 50 may be enabled by a computer / personal electronic device 250. The computer / personal electronic device 250 can be connected to the storage workload controller 50 directly, or can be connected to the storage workload controller 50 via a network, e.g. the compute cluster 200 (such as the Internet) , as shown in FIG. 1A. In some embodiments, the storage workload controller 50 may include the computer or personal electronic device 250, i.e., the computer or personal electronic device 250 may function as the storage workload controller 50 or be part of the storage workload controller 50. In some embodiments, the storage workload controller 50 may be part of the computer or personal electronic device 250.

[0022] As shown in FIG. 1A, each storage device 10 further includes a drive controller 40. In some embodiments of the present disclosure, the drive controllers 40 are configured to communicate with the storage workload controller 50 to, e.g., execute the movement of data according to the storage management method of the present disclosure. The drive controllers 40 can be any storage unit drive controllers. For example, a drive controller 40 can be a RAID controller used for a data storage virtualization technology that combines multiple physical disk drive components into one or more logical units, where RAID stands for “Redundant Array of Independent Disks. ”

[0023] The drive controllers 40 can be configured to function differently for different types of storage devices. For example, a storage device 10 can be, e.g., a hard drive disk, and the  drive controller thereof can be configured to function as self-monitoring, analysis, and reporting technology (S.M.A.R.T.) to monitor and detect signs of wear condition and communicate the same with the storage workload controller 50. As another example, a storage device 10 can be, e.g., a solid-state disk, such as an NVMe SSD, and the drive controller thereof can be configured to monitor terabytes written (TBW) as a metric and communicate same with the storage workload controller 50.

[0024] According to the present disclosure, the storage workload controller 50 is configured to provision storage workload according to workload capacities of each storage device 10. Further, the storage workload controller 50 includes an I / O monitor 60 monitoring the I / O activities of each storage device 10 of each storage node 100. In some embodiments, as shown in FIG. 1A, the I / O monitor 60 is coupled to the node I / O logs 20 through the at least one router 52, and is configured to monitor I / O activities of the nodes 100 through the corresponding node I / O logs 20. In some embodiments, as shown in FIG. 1A, one or more of the storage devices 10 each include a drive-wear controller 62 coupled with the corresponding node I / O log 20 via the I / O of the corresponding storage device 10.

[0025] Each storage node 100 can be a storage facility located in a geographic location, such as New York City, Tokyo, or Paris, and the storage facility can be managed by a cloud computing company. The storage devices 10 can be the same type of storage devices rented by the same Cloud customer for some specified types of files, or at least some of the storage devices 10 can be different types of storage devices belonging to different customers having different types of files. For example, the management of a storage device 10 of a node 100 can be “file based, ” “customer based, ” and / or “device based. ”

[0026] Each node I / O log 20 may be part of a corresponding node 100 and is also part of the overall storage pool, i.e., the HCSS 1. Each node I / O log 20 also functions to communicate with the storage controller 50 that is built into the corresponding node 100.

[0027] Those skilled in the art would appreciate that the structure of the HCSS 1 shown in FIG. 1A may be presented or designed in many ways. For example, the HCSS 1 can include a plurality of routers 52 and / or a plurality of storage workload controllers 50. As another example, the connection between the storage workload controller 50 and each node I / O log 20 may include another layer in addition to the router 52. Variations of such are all within the scope of the present disclosure.

[0028] According to the present disclosure, the storage I / O activities can be managed or leveled based on the assessment of wear condition of each of the storage devices according to the storage management of the whole HCSS 1.

[0029] The storage workload controller 50 may be configured to assess the wear condition of each storage device, generate target accessing activities for accessing the storage devices, including, e.g., target I / O activities to the storage devices, according to the wear condition, and control data movement among the storage devices according to target accessing activities. The target accessing activities can include, e.g., threshold or profile (pattern) of the relevant I / O activities.

[0030] In some embodiments, the storage workload controller 50 can be placed on a non-local drive in order to balance wear patterns, i.e., distribution or profile of wears among the storage devices 10. If a critical threshold is reached, the storage management functions consistent with the  disclosure perform a wholesale swap of data between a node with drives in an “over-wear” state and a node with drives in an “under-wear” state.

[0031] In some embodiments, the wear condition of a storage device 10 may be measured by a frequency of input and output activities (the I / O activities) of the storage device 10. The wear condition of a node 100 of storage devices may be measured according to frequencies of input and output activities of the storage devices 10 of the node 100, such as an average frequency of such input and output activities over the storage devices 10 of the node 100.

[0032] In some embodiments, the wear condition of a storage device 10 may be measured by a frequency of input activities of the storage device 10. The wear condition of a node 100 of storage devices 10 may be measured according to frequencies of input activities of the storage devices 10 of the node 100, such as an average frequency of such input activities over the storage devices 10 of the node 100.

[0033] In some embodiments, the wear condition of a storage device 10 may be measured by a pattern of frequencies of input and output activities of the storage device 10 over a period of time. The wear condition of a node 100 of storage devices 10 may be measured by a pattern of frequencies of input and output activities of the storage devices 10 of the node 100.

[0034] In some embodiments, the wear condition of a storage device 10 may be measured by a pattern of frequencies of input activities of the storage device 10 over a period of time. The wear condition of a node 100 of storage devices 10 may be measured by a pattern of frequencies of input activities of the storage devices 10 of the node 100.

[0035] In some embodiments, the wear condition of a storage device 10 may include wear condition associated with at least one physical wear parameter of the storage device 10. The  wear condition of a node 100 of storage devices 10 may include wear condition associated with at least one physical wear parameter of the storage devices 10 of the node 100 collectively, such as an average value of the at least one physical wear parameter over the storage devices 10 of the node 100.

[0036] In some embodiments, the wear condition of a storage device 10 may include a pattern of the wear condition associated with the at least one physical wear parameter of the storage device 10 over a period of time. The wear condition of a node 100 of storage devices 10 may include a pattern of the wear condition associated with the at least one physical wear parameter of the storage devices 10 of the node 100 collectively, such as a pattern of an average value of the at least one physical wear parameter over the storage devices 10 of the node 100.

[0037] Example methods used by the storage workload controller 50 to manage the HCSS 1 are described in detail below in connection with FIGs. 1B, 2-6 and Table 1.

[0038] FIG. 1B is a flowchart showing a method of storage management according to an embodiment of the present disclosure.

[0039] As shown in FIG. 1B, at 102, the method of storage management includes assessing wear condition of a monitored storage device, such as storage device 10 within a compute cluster 100 in FIG. 1A. At 104, the method includes generating a target level of accessing activities, such as I / O activities, for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device. At 106, the method of storage management includes controlling data movement within the compute cluster according to the target level of accessing activities and the assessed wear condition of the monitored storage device 10. Details of the method of storage management are described in the description as follows.

[0040] FIG. 2 is a flowchart showing a method for storage management by identifying storage devices exceeding an I / O activity threshold according to an example embodiment of the present disclosure. The example method shown in Fig. 2 can be referred to as an “I / O threshold method, ” and can be implemented, e.g., by the HCSS 1 of FIG. 1A, such as by the storage workload controller 50 of the HCSS 1.

[0041] As shown in FIG. 2, at 62, a wear-control task is defined. Defining the wear-control task may include specifying what storage devices are monitored and controlled. The monitored and controlled storage devices can include, for example, all the storage devices 10 of one or more of the nodes 100, etc. In some embodiments, the monitored and controlled storage devices can include some of the storage devices 10 of one or more of the nodes 100. A storage device being monitored is also referred to as a “monitored storage device, ” and a node including one or more monitored storage device is also referred to as a “monitored node. ”

[0042] The task definition also may include whether the task is an I / O monitoring task or a data collection task. The data collection task refers to a task related to collecting data associated with one or more storage devices 10. The I / O monitoring task can include monitoring collected data related to I / O activities reported by, e.g., each node I / O log 20. The task definition can also include a time range of the I / O monitoring or data collection, and a predetermined threshold pertaining to I / O activities (I / O activity threshold) or data collection (data collection threshold) . For example, the time range of an I / O monitoring can vary between a few minutes to hours depending on how the task is defined.

[0043] The task definition may also include whether a task is an intermittent or always-on task, for example, whether an I / O activity task is an intermittent I / O activity monitoring task or  an always-on I / O activity monitoring task, and / or whether a data collection task is an intermittent data collection task or an always-on data collection task.

[0044] The intermittent I / O monitoring task could be performed at certain specified timings or periodically and can be performed automatically or manually. For example, the intermittent I / O monitoring task can be performed a few times a day, e.g., the I / O monitoring is turned on for 30 minutes and turned off for 90 minutes for every 120 minutes of time span. The always-on I / O monitoring task can be an I / O monitoring task that remains active, i.e., remains turned on, once the task is started, e.g., by a user, until the task is turned off, e.g., by a user. Intermittent and always-on data collection tasks can be defined in a similar manner.

[0045] At 64, it is determined whether the storage management is based on the on-going I / O activities or the historical I / O activities. If the storage management is based on the statistics of the on-going I / O activities, the process proceeds to 66, at which data for the on-going I / O activities is collected from all the storage devices of all the monitored nodes according to the wear-control task defined at 62. Monitoring the statistics of the on-going I / O activities is an example of monitoring concurrent data of a present operation. If the storage management is based on the statistics of the historical I / O activities, the process proceeds to 67, at which data for the historical I / O activities is collected from all the storage devices of all the monitored nodes according to the wear-control task defined at 62. Data, whether for the on-going I / O activities or for the historical I / O activities, collected from the storage devices of the monitored nodes is also referred to as “collected data. ”

[0046] At 68, the collected data is analyzed to identify one or more over-worn storage devices. An over-worn storage device refers to a storage device having I / O activities that have passed the threshold discussed above.

[0047] At 69, storage workload placement pattern is modified or adjusted to reduce I / O activities of the identified one or more over-worn storage devices to a reduced level. The reduced level can be, for example, 1 / 3 of the I / O activity threshold. The storage workload placement pattern characterizes how storage workloads placement are distributed among various storage devices, such as among the monitored storage devices or among the storage devices of the monitored nodes.

[0048] FIG. 3 is a flowchart showing a method for storage management by generating a profile of I / O activities across a plurality of storage devices in one or more nodes according to an example embodiment of the present disclosure. The example method shown in FIG. 3 can be referred to as an “I / O profile method, ” and can be implemented, e.g., by the HCSS 1 of FIG. 1A, such as by the storage workload controller 50 of the HCSS 1.

[0049] As shown in FIG. 3, at 72, a wear-control task is defined. The wear-control task definition used in the method shown in FIG. 3 can be similar to that used in the method shown in FIG. 2, and detailed description thereof is omitted. In some embodiments, as compared to the wear-control task definition used in the method shown in FIG. 2, the wear-control task definition used in the method shown in FIG. 3 can omit certain information not needed, such as the I / O activity threshold.

[0050] At 74, it is determined if the storage management is based on the on-going I / O activities or the historical I / O activities. If the storage management is based on statistics of the on- going I / O activities, the process of the storage management method proceeds to 76, at which data for the on-going I / O activities is collected from all the storage devices of all the monitored nodes according to the wear-control task defined at 72. If the storage management is based on the statistics of the historical activities, the process of the storage management method proceeds to 77, at which data for the historical I / O activities is collected from all the storage devices of all the monitored nodes according to the wear-control task defined at 72.

[0051] At 78, the data collected at 76 or 77 is analyzed to generate an I / O activity profile among the monitored storage devices of the monitored nodes. The I / O activity profile characterizes how I / O activities are distributed across the monitored storage devices of the monitored nodes, e.g., what the I / O activities are at various ones of the monitored storage devices. The I / O activity at a monitored storage device can be expressed, e.g., as the percentage difference of the number of I / O events at the monitored storage device above or below the I / O activity threshold. For example, if the number of I / O events at a monitored storage device is 10% higher than the I / O activity threshold, the I / O activity at the monitored storage device can be expressed as +10%. As another example, if the number of I / O events at a monitored storage device is 10% lower than the I / O activity threshold, the I / O activity at the monitored storage device can be expressed as -10%.

[0052] At 79, a storage workload placement pattern is modified or adjusted by inverting the I / O activity profile to balance the I / O activities of the monitored storage devices. Inverting the I / O activity profile can refer to “inverting” the percentage differences with respect to the I / O activity threshold that represent the I / O activities at the monitored storage devices. For example, assume that before the modification or adjustment, the I / O activity at a monitored storage device is  +10% (i.e., 10% higher than the I / O activity threshold) , then the I / O activity at the monitored storage device can be modified / adjusted to -10% (i.e., 10% lower than the I / O activity threshold) .

[0053] FIG. 4 is a flowchart showing a method for storage management by identifying a pattern of wear conditions among storage devices according to an example embodiment of the present disclosure. The example method shown in Fig. 4 can be referred to as a “wear condition profile method, ” and can be implemented, e.g., by the HCSS 1 of FIG. 1A, such as by the storage workload controller 50 of the HCSS 1.

[0054] As shown in FIG. 4, at 82, a wear-control task is defined. The definition of the wear-control task used in the method shown in FIG. 4 can be similar to those used in the methods shown in FIGs. 2 and 3, except that in the method shown in FIG. 4, wear conditions, rather than I / O activities, of the storage devices are monitored. For example, the task definition used in the method shown in FIG. 4 may include whether the task is an intermittent or always-on wear condition monitoring task, and may also include the time range of wear condition monitoring, and a predetermined threshold pertaining to wear condition (wear condition threshold) . In some embodiments, information not needed, such as the wear condition threshold, can be omitted. For other details of the wear-control task definition, reference can be made to the descriptions above related to processes 62 and 72.

[0055] At 84, it is determined if the storage management is based on the statistics of the on-going storage wear monitoring or the historical storage wear monitoring. If the storage management method is based on the statistics of the on-going storage wear monitoring, the process of the storage management method proceeds to 86, at which data for the on-going storage wear monitoring is collected from all the storage devices of all the monitored nodes according to the  task defined at 82. If the storage management is based on the statistics of the historical storage wear monitoring, the process of the storage management method proceeds to 87, at which data for the historical storage wear monitoring is collected from all the storage devices of all the monitored nodes according to the task defined at 82.

[0056] At 88, data collected at 86 or 87 is analyzed to generate storage wear profile on the monitored storage devices of the monitored nodes. The storage wear profile includes the wear conditions of the monitored storage devices and characterizes how badly various ones of the monitored storage devices are worn out. The storage wear profile characterizes how wear conditions are distributed across the monitored storage devices of the monitored nodes, e.g., what the wear conditions are at various ones of the monitored storage devices. In some embodiments, the wear condition of a storage device may include wear condition associated with the value of at least one physical wear parameter of the storage devices 10 of the node 100 collectively, such as an average value of the at least one physical wear parameter over the storage devices 10 of the node 100. For example, the physical wear parameter may be the “G” value to indicate the "growth" list of bad sectors.

[0057] The wear condition at a monitored storage device can be expressed, e.g., as the percentage difference of wear condition at the monitored storage device above or below the wear condition threshold. For example, if the value of the wear parameter at a monitored storage device is 10% higher than the wear parameter threshold, the value of the wear parameter at the monitored storage device can be expressed as +10%. As another example, if the number of I / O events at a monitored storage device is 10% lower than the I / O activity threshold, the I / O activity at the  monitored storage device can be expressed as -10%. The pattern of the wear parameter forms the wear profile in at 88.

[0058] At 89, the storage controller is modified / adjusted to invert the storage wear profile to reduce I / O activities at certain monitored storage devices. The storage controller can be, e.g., the controller 50 shown in FIG. 1A. In some embodiments, the storage controller can be modified by, e.g., adjusting a work mode of the storage controller or a program (software or a firmware) installed at the storage controller.

[0059] In some embodiments, the storage management method shown in FIG. 4 further includes checking if the storage wear profile generated at process 88 indicates the wear condition of any storage device reaches a hazardous level (80) . If so (80: Yes) , an alert is sent out to remove the storage device (s) having a wear condition reaching the hazardous level (81) . If the wear conditions of all the monitored storage devices are below the hazardous level (80: No) , the process proceeds to 89 to modify / adjust the storage controller as described above.

[0060] FIG. 5 is a flowchart showing a process of a storage management method by reducing storage I / O activities in storage devices identified as having excessive wear conditions according to an example embodiment of the present disclosure. The example method shown in FIG. 5 can be referred to as a “wear condition threshold method, ” and can be implemented, e.g., by the HCSS 1 of FIG. 1A, such as by the storage workload controller 50 of the HCSS 1.

[0061] As shown in FIG. 5, at 92, a wear-control task is defined. The wear-control task definition used in the method shown in FIG. 5 can be similar to that used in the method shown in FIG. 4, and detailed description thereof is omitted. For example, the task definition may include the wear condition threshold.

[0062] At 94, it is determined if the storage management is based on the statistics of the on-going storage wear monitoring, or the historical storage wear monitoring. If the storage management method is based on the statistics of the on-going storage wear monitoring, the process of the storage management method goes to 96, at which data for the on-going storage wear conditions is collected from all the storage devices of all the monitored nodes according to the task defined at 92. If the storage management is based on the statistics of the historical storage wear monitoring, the process of the storage management method proceeds to 97, at which data for the historical storage wear monitoring is collected from all the storage devices of all the monitored nodes according to the task defined at 92.

[0063] At 98, data collected at 96 or 97 is analyzed to identify the storage device (s) with wear condition exceeding the wear condition threshold defined at 92.

[0064] At 99, the storage controller is modified / adjusted to reduce I / O activities for the storage device (s) having wear condition exceeding the threshold wear condition.

[0065] In some embodiments, the storage management method shown in FIG. 5 further includes checking if the wear condition of any storage device as identified at process 98 reaches a hazardous level (90) . If so (90: Yes) , an alert is sent out to remove the storage device (s) having a wear condition reaching the hazardous level (91) . If the wear conditions of all the monitored storage devices are below the hazardous level (90: No) , the process proceeds to 99 to modify / adjust the storage controller as described above.

[0066] FIG. 6 is a flowchart showing a method of generating types of threshold used by the storage management method according to an embodiment of the present disclosure.

[0067] At 1010, the types of storage devices and their IDs and / or locations in nodes of a compute cluster are retrieved. As described above, a storage device in the HCSS 1 may be a random access memory, a non-volatile read-only memory, a hard drive disk, a solid-state disk, a compact disk, a floppy disk, or a magnet tap, etc. Different types of storage devices may have different thresholds associated with the wear condition (s) or I / O activities.

[0068] At 1020, whether thresholds are related to I / O activities or wear condition (s) is determined. If the thresholds are related to the I / O activities (1020: I / O activities) , the process proceeds to 1030, while if the thresholds are related to the wear conditions (1020: Wear conditions) , the process proceeds to 1035.

[0069] At 1030, the thresholds for I / O activities are determined for various types of storage devices.

[0070] At 1035, the thresholds for device wear conditions are determined for various types of storage devices.

[0071] At 1040, the thresholds determined at 1030 and 1035 are provided to the storage controller 50.

[0072] Writing and reading activities both have impact on the wear-out of storage devices. However, depending on the type of a storage device, writing or making input into the storage device often is the most predominant factor affecting the rate of wear-out for the storage device. Therefore, in some embodiments, instead of monitoring and adjusting the I / O activities (which include both writing and reading activities) , monitoring and adjusting writing activities can be used to level wear-out effect of a hyperconverged storage system. The methods of using writing activities for leveling wear-out effect are similar to the methods of using I / O activities for leveling  wear-out effect described above, except that the writing activities are tracked instead of I / O activities, and thus detailed description thereof is omitted.

[0073] Table 1 shows an example where both I / O and writing activities are tracked, while the profile of writing activities is used for adjusting and controlling writing in the HCSS 1 to level the wear-out effect. The first column in Table 1 lists storage devices being monitored by their drive IDs. The I / O activities relative to a predetermined threshold and “writing” activities relative to another predetermined threshold in the last 1000 hours, for example, are analyzed and listed. The data movement control in the HCSS 1 is implemented by the storage workload controller 50 by making a simple inversion of the profile of “writing activities, ” as shown in the last column of Table 1. In Table 1, the activities are expressed as percentage differences relative to the threshold (I / O activity threshold for I / O activities or writing activity threshold for writing activities) .

[0074] Table 1. Example of Leveling Storage Device Wear-out, Writing Profile Method

[0075] FIG. 7 is a block diagram of an example system 700 for assessing wear condition and controlling the usage placement of storage devices in the HCSS 1 according to an example embodiment of the present disclosure. The system 700 can implement any of the storage workload controller 50, the I / O monitor 60, the drive-wear controller 62, and the node I / O log 20 in FIG. 1A, or any combination thereof. The system 700 includes a storage 710 storing a program logic 715, a processor 720 for executing a process 725, and a communications I / O interface 730, connected via a bus 735. The exemplary system 700 is described only for illustrative purpose and should not be construed as a limitation to the embodiments or scope of the present disclosure. In some cases, some devices may be added to or removed from the system 700 based on specific situations. The process 725 may be a process for implementing a method consistent with the disclosure, such as any of the exemplary methods described above in connection with FIGs. 2-6. The processor 720 can be configured to execute the program logic 715 to implement the process 725.

[0076] Processing may be implemented in hardware, software, or a combination of the two. Processing may be implemented in computer programs executed on programmable computers / machines that each includes a processor, a storage medium or other article of manufacture that is readable by the processor (including volatile and non-volatile storage and / or storage elements) , at least one input device, and one or more output devices. Program code may be  applied to data entered using an input device to perform processing and to generate output information.

[0077] FIG. 8 is a block diagram of a computer program product 800 including program logic, upon being executed, implements the storage management method according to an example embodiment of the present disclosure, such as one of the above-described exemplary methods.

[0078] As shown in FIG. 8, the computer program product 800 includes the program logic 715 of FIG. 7, stored on a computer-readable medium 860 in computer-executable code configured for assessing wear condition and controlling the usage placement of storage devices according to an example embodiment of the present disclosure. The logic for carrying out the method may be embodied as part of the aforementioned system, which is useful for carrying out a method described with reference to embodiments shown. In one embodiment, the program logic 715 may be loaded into a storage device and executed by a processor.

[0079] Although the foregoing disclosure has been described in some detail for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the disclosure. The disclosure encompasses numerous alternatives, modifications, and equivalents. Specific details are set forth in the above description in order to provide a thorough understanding of the disclosure. These details are provided as example and the disclosure may be practiced without some or all of these specific details. For the purpose of clarity, technical material that is known in the technical fields related to the disclosure has not been described in detail so that the disclosure is not unnecessarily obscured. Accordingly, the above implementations are to be considered as illustrative and not restrictive, and the disclosure is not to  be limited to the details given herein but may be modified within the scope and equivalents of the disclosure.

[0080] Various exemplary embodiments of the present disclosure have been described with reference to the accompanying drawings. It may be appreciated that these example embodiments are provided only for enabling those skilled in the art to better understand and implement the present disclosure and not intended to limit the scope of the present disclosure in any manner. It should be noted that these drawings and description are only presented as exemplary embodiments and, based on this description, alternative embodiments may be conceived that may have a structure and method disclosed as herein, and such alternative embodiments may be used without departing from the principle of the disclosure.

[0081] It may be noted that the flowcharts and block diagrams in the figures may illustrate the apparatus, method, as well as architecture, functions, and operations executable by a computer program product according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, a program segment, or a part of code, which may contain one or more executable instructions for performing specified logic functions. It should be further noted that, in some alternative implementations, functions indicated in blocks may occur in an order differing from the order as illustrated in the figures. For example, two blocks shown consecutively may be performed in parallel substantially or in an inverse order sometimes, which depends on the functions involved. It should be further noted that each block and a combination of blocks in the block diagrams or flowcharts may be implemented by a dedicated, hardware-based system for performing specified functions or operations or by a combination of dedicated hardware and computer instructions.

[0082] The terms “comprise (s) , ” “include (s) , ” their derivatives, and like expressions used herein should be understood to be open (i.e., “comprising / including, but not limited to” ) . The term “based on” means “at least in part based on, ” the term “one embodiment” means “at least one embodiment, ” and the term “another embodiment” indicates “at least one further embodiment. ” Relevant definitions of other terms have been provided.

[0083] The foregoing description of the embodiments of the disclosure have been presented for the purpose of illustration. It is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the above disclosure.

[0084] Some portions of this description describe the embodiments of the disclosure in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Furthermore, it has also proven convenient at times, to refer to these arrangements of operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combinations thereof.

[0085] Any of the steps, operations, or processes described herein may be performed or implemented with one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented with a computer program product comprising a computer-readable medium containing computer program code, which can be  executed by a computer processor for performing any or all of the steps, operations, or processes described.

[0086] Embodiments of the disclosure may also relate to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, and / or it may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory, tangible computer readable medium, or any type of media suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing systems referred to in the specification may include a single processor or may be architectures employing multiple processor designs for increased computing capability.

[0087] Embodiments of the disclosure may also relate to a product that is produced by a computing process described herein. Such a product may comprise information resulting from a computing process, where the information is stored on a non-transitory, tangible computer readable medium and may include any embodiment of a computer program product or other data combination described herein.

[0088] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the disclosure be limited not by this detailed description. Accordingly, the disclosure of the embodiments is intended to be illustrative, but not limiting, of the scope of the disclosure.

Claims

1.A method, comprising:assessing wear condition of a monitored storage device within a compute cluster;generating a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device; andcontrolling workload placement within the compute cluster according to the target level of accessing activities.2.The method of claim 1, wherein assessing the wear condition comprises assessing a frequency of input and output of the monitored storage device.3.The method of claim 1, wherein assessing the wear condition comprises assessing a frequency of input of the monitored storage device.4.The method of claim 1, wherein assessing the wear condition comprises assessing a pattern of frequencies of input and output of the monitored storage device.5.The method of claim 1, wherein assessing the wear condition comprises assessing a pattern of frequencies of input of the monitored storage device.6.The method of claim 1, wherein assessing the wear condition comprises assessing wear condition associated with at least one physical wear parameter of the monitored storage device.7.The method of claim 1, wherein assessing the wear condition comprises assessing a pattern of wear condition associated with at least one physical wear parameter of the monitored storage device.8.The method of claim 1, wherein assessing the wear condition comprises assessing historical input / output activities related to the monitored storage device.9.The method of claim 1, wherein assessing the wear condition comprises monitoring concurrent data of a present operation of the monitored storage device.10.The method of claim 1, wherein the monitored storage device is one of a plurality of monitored storage device in the compute cluster, and generating the target level of accessing activities includes defining an inversion of a pattern of the wear conditions of the plurality of monitored storage devices.11.The method of claim 1, wherein generating the target level of accessing activities includes defining a predetermined threshold of accessing activities.12.The method of claim 1, wherein the workload placement including data movement is controlled further according to a workload placement pattern according to the target level of accessing activities of the compute cluster.13.The method of claim 1, wherein the compute cluster includes a hyperconverged storage system.14.The method of claim 1, wherein the monitored storage device includes at least one of a random access memory, a non-volatile read-only memory, a hard drive disk, a solid-state disk, a compact disk, a floppy disk, or a magnet tap.15.The method of claim 1, wherein the plurality of storage devices are grouped into storage nodes.16.The method of claim 15, wherein the controlling workload placement including a swap of respective data and workload between two nodes of the storage nodes according to the wear condition of the monitored storage device.17.A computer storage system comprising:a plurality of storage devices; anda storage controller coupled with the plurality of storage devices and configured to:assess wear condition of a monitored storage device of the plurality of storage devices;generate a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device; andcontrol data movement of the computer storage system according to the target level of assessing activities.18.The computer storage system of claim 17, wherein the storage controller assesses the wear condition by assessing a frequency of input into the monitored storage device.19.The computer storage system of claim 17, wherein the storage controller assesses the wear condition by assessing wear condition associated with at least one physical wear parameters of the monitored storage device.20.A non-transitory computer readable storage medium having program instructions stored therewith, the program instructions being executable by a processor to cause the processor to perform a method comprising:assessing wear condition of a monitored storage device within a compute cluster;generating a target level of accessing activities for accessing at least the monitored storage device according to at least the wear condition of the monitored storage device; andcontrolling data movement within the compute cluster according to the target level of accessing activities.

Citation Information

Patent Citations

  • Dynamic load balancing of hardware threads in clustered processor cores using shared hardware resources, and related circuits, methods, and computer-readable media

    CN106462394A

  • Workload management with data access awareness in a computing cluster

    CN112005219A

  • Performing wear leveling between storage systems of storage cluster

    CN114860150A

  • Data layout selection between storage devices associated with nodes of distributed file system cluster

    CN116414796A

  • Enhanced application performance optimized using storage system

    CN116529695A