Storage management method, device, and computer program product
By adjusting the recovery rate and reserved time difference, the problem of the long recovery process of Uber storage unit is solved, and recovery is completed within a limited time, improving the reliability and performance of the storage system.
Patent Information
- Application Number
- CN202110088242.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-22
AI Technical Summary
When the prior art restores Uber storage units, the recovery process is too long, making it difficult to ensure the reliability of the storage system, and the recovery process occupies system resources and affects performance.
By determining the first recovery rate and the second recovery rate, combined with the reserved time difference, the recovery rate is adjusted to ensure that all disk sets are restored within a limited time, reducing the impact on system performance.
Recovery of all disk sets within a limited time, reducing the impact on system performance, and improving the reliability and recovery efficiency of the storage system.
Smart Images

Figure CN114816221B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of data storage, and more particularly, to storage management methods, devices, and computer program products. Background Art
[0002] A disk array, such as a redundant array of independent disks (RAID), is a disk group formed by combining multiple independent disks in a certain manner. In the view of a user, a redundant array of independent disks is similar to a single disk, but it can provide higher storage capacity than a single hard disk and can also provide data backup. When data in a disk area is damaged, the damaged data can be restored using the data backup, thereby protecting the security of user data.
[0003] An Uber storage unit has a structure and function similar to that of RAID and can be considered a lightweight RAID. When one or more disks in an Uber go offline due to reasons such as poor contact or failure, the Uber needs to be recovered (or referred to as "rebuild"). If the recovery process is too long, it is difficult to ensure the reliability of the storage system. In addition, since the recovery process needs to occupy system resources, it is desirable that the impact of the recovery process on system performance be as small as possible. Summary of the Invention
[0004] Embodiments of the present disclosure provide a storage management method, device, and computer program product.
[0005] According to a first aspect of the present disclosure, there is provided a storage management method. The method may include determining a first recovery rate for recovering at least a part of a plurality of disk sets based at least on an upper limit duration for recovering a predetermined number of disk sets in the plurality of disk sets. The method may further include determining the number of disk sets in the plurality of disk sets that are not recovered based on the first recovery rate. In addition, the method may further include performing data recovery on the unrecovered disk sets in the plurality of disk sets based on a predetermined second recovery rate according to the determined number being less than or equal to a predetermined number, the second recovery rate being lower than the first recovery rate and associated with the upper limit duration.
[0006] According to a second aspect of the present disclosure, an electronic device is provided. The electronic device includes: a processor; and a memory storing computer program instructions, and the processor runs the computer program instructions in the memory to control the electronic device to perform operations, and the operations include: determining a first recovery rate based at least on an upper limit duration for recovering a predetermined number of disk sets in a plurality of disk sets, for recovering at least a part of the plurality of disk sets; determining the number of disk sets in the plurality of disk sets that are not recovered based on the first recovery rate; and based on determining that the number is less than or equal to a predetermined number, performing data recovery on the unrecovered disk sets in the plurality of disk sets based on a predetermined second recovery rate, where the second recovery rate is lower than the first recovery rate and is associated with the upper limit duration.
[0007] According to a third aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-volatile computer-readable medium and includes machine-executable instructions, and the machine-executable instructions, when executed, cause the machine to perform the steps of the method in the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. Among them, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0009] Figure 1 A schematic diagram showing a plurality of disk sets to be recovered according to an embodiment of the present disclosure;
[0010] Figure 2 A schematic diagram showing computing resources according to an embodiment of the present disclosure;
[0011] Figure 3 A schematic diagram showing a storage management process according to an embodiment of the present disclosure;
[0012] Figure 4 A schematic diagram showing a process of determining a first recovery rate according to an embodiment of the present disclosure;
[0013] Figure 5 A schematic diagram showing a process of determining the number of recovery jobs according to an embodiment of the present disclosure;
[0014] Figure 6 A schematic block diagram of an exemplary device suitable for implementing embodiments of the present disclosure.
[0015] In each of the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] The principles of the present disclosure will be described below with reference to several exemplary embodiments shown in the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the description of these embodiments is only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way.
[0019] A RAID-based disk set is a storage disk group formed by combining multiple independent disks in different ways. When multiple disks are stored in association as a disk set, if some disks are unavailable, the disk set needs to be recovered. The recovery duration is usually an important indicator of system reliability. For example, in currently widely used storage systems, the recovery duration is usually limited within a certain range, such as within 4.4 hours. Given the acceptable recovery duration of a known storage system, the recovery rate can be determined based on this recovery duration and the disk capacity of the disk set to be recovered, so as to ensure that all disk sets to be recovered can be completed within the recovery duration.
[0020] To adjust the recovery rate, the storage system can adjust the number of recovery jobs for parallel processing of multiple disk sets. For example, when the actual progress of the recovery lags behind the predetermined progress, multiple recovery jobs can be initiated in parallel on one or more computing nodes, and each recovery job is used to handle the recovery task of a disk set. Therefore, in order to meet the requirements of the recovery duration, a considerable amount of system computing resources can be occupied to initiate enough recovery jobs in parallel. It should be understood that usually, the system computing resources will not be overly occupied to increase the recovery rate, because this will affect the normal use of the system and thus degrade the user experience. Therefore, the adjustment of the recovery rate for the disk set should not only meet the requirements of the limited recovery duration but also minimize the impact of the recovery operation on the system performance.
[0021] However, there are some risks in the above recovery method. For example, in some cases, when calculating the recovery rate in the above manner, the maximum number of concurrent recovery jobs selected can just complete the recovery operation of the disk set within the limited recovery duration. However, when the number of the last remaining disk sets is less than the maximum number of concurrent jobs, even if the maximum number of concurrent recovery jobs is selected, the recovery rate cannot reach the expected rate. For example, assume that the maximum number of concurrent recovery jobs is 8. When the number of the last remaining disk sets to be recovered is 4, even if 8 recovery jobs are selected to recover these 4 disk sets in parallel, the recovery rate cannot reach the rate of 8 recovery jobs recovering 8 disk sets in parallel. Therefore, the above recovery method may exceed the limited recovery duration, thus affecting the reliability of the system.
[0022] To solve the above problems, the present disclosure proposes a new storage management scheme. A ceiling duration is subtracted from the total recovery duration, and the recovery rate is determined based on this time difference. The ceiling duration is sufficient to enable the last remaining several disk sets to complete the recovery operation at the slowest or slower recovery rate without exceeding the limited recovery duration. To better understand the process of storage management according to the embodiments of the present disclosure, the following will first refer to Figure 1 Describe the disk sets to be recovered.
[0023] Figure 1 FIG. shows a schematic diagram of a plurality of disk sets 100 to be recovered according to an embodiment of the present disclosure. In Figure 1 , a plurality of disk sets 101-1, 101-2, 101-3…101-N, 102-1, 102-2, 102-3…102-N, 103-1, 103-2, 103-3…103-N…M-1, M-2, M-3…M-N (hereinafter collectively referred to as “disk sets 100”) are all disk sets to be recovered, where N is an integer greater than 1 and less than N, and M is an integer greater than 1 and less than N. It should be understood that Figure 1 each disk set in contains a plurality of storage disks. Since the disk sets 100 are all disk sets to be recovered, there are a small number of offline storage disks in each disk set.
[0024] It should be understood that the above storage disks may include various types of devices with storage functions, including but not limited to, hard disk drives (HDDs), solid state drives (SSDs), removable disks, compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, Blu-ray discs, serial attached small computer system interface (SCSI) storage disks (SAS), serial advanced technology attachment (SATA) storage disks, any other magnetic storage devices and any other optical storage devices, or any combination thereof.
[0025] It should also be understood that in order to avoid showing the idea of the present disclosure in a more complex way,Figure 1 The specific structure within each disk set is not shown. In fact, Figure 1 each disk set in [reference] can have its own different type to provide different levels of data redundancy and recovery capabilities. The types of RAID include RAID 2, RAID 3, RAID 4, RAID 5, RAID 6, RAID 7, RAID 10, etc. In addition, disk sets of the same RAID type can also have different widths. For example, RAID6 can include RAID6 of 8 + 2 and RAID6 of 16 + 2. The above examples are only for illustrating the present disclosure and are not intended to limit the present disclosure.
[0026] In addition, as Figure 1 shown, the disk set 100 may further include a remaining disk set 120. The remaining disk set 120 is usually generated due to multiple recovery jobs recovering the disk set 100 in parallel. The remaining disk set 120 and its impact on the recovery duration will be described in detail below in conjunction with Figure 2 [reference].
[0027] Figure 2 FIG. [reference] shows a schematic diagram of a computing resource 200 according to an embodiment of the present disclosure. It should be understood that when there is sufficient time, each disk set 100 can be recovered one by one in a single recovery job, and in this case, no remaining disk set 120 will be generated. In addition, when the number of disk sets 100 is large and the disk capacity to be recovered is large, multiple recovery jobs can be initiated in parallel to recover the disk sets 100 in parallel. In Figure 2 [reference], it is assumed that the maximum concurrent number of recovery jobs is 8, Figure 2 FIG. [reference] shows an example of initiating 4 recovery jobs (i.e., recovery job group 210) in parallel to recover the disk sets 100. At the same time, the remaining computing resources (i.e., potential recovery job group 220) can be used for normal system or host IO. As Figure 2 shown, 4 disk sets can be recovered in parallel in 4 recovery jobs in the recovery job group 210. If the number of disk sets to be recovered in the disk set 100 cannot be evenly divided by 4, then a remaining disk set 120 will inevitably be generated. Obviously, at this time, the number of the remaining disk sets 120 is less than 4 and greater than 0. Since the number of the remaining disk sets 120 is less than the number of parallel recovery jobs, the recovery rate of the remaining disk sets 120 will be less than the expected rate, resulting in the recovery duration not meeting the expected requirements.
[0028] For the above problems, the system of the present disclosure can ensure that the recovery duration meets the expected requirements by executing a process for storage management as shown in Figure 3 [reference]. The flowchart of the process for storage management will be described in detail below in conjunction with Figure 3 [reference].
[0029] Figure 3FIG. 300 is a schematic diagram of a storage management process according to an embodiment of the present disclosure. In some embodiments, process 300 may be implemented in the Figure 6 device shown. For ease of understanding, the specific data mentioned in the following description are exemplary and are not used to limit the protection scope of the present disclosure.
[0030] At 301, a first recovery rate may be determined based at least on an upper limit duration for recovering a predetermined number of disk sets (e.g., Figure 1 disk set 100 in FIG. 100) in a plurality of disk sets, for recovering at least a portion of the plurality of disk sets. As an example, the predetermined number may be the concurrent number of the largest recovery jobs. To further minimize the impact on system performance, the predetermined number may also be slightly smaller than the concurrent number of the largest recovery jobs, e.g., the concurrent number of the largest recovery jobs minus one.
[0031] It should be understood that the reason for determining the first recovery rate based on the upper limit duration of the recovery of a predetermined number of disk sets is to reserve time for the recovery operation of the remaining disk set 120 at a lower recovery rate. Specifically, the upper limit duration may be subtracted from the total threshold duration, and thus the first recovery rate may be determined based on the updated time. Figure 4 FIG. 400 is a schematic diagram of a process for determining a first recovery rate according to an embodiment of the present disclosure.
[0032] At 401, a total threshold duration for recovering the plurality of disk sets may be determined in advance. For currently widely used storage systems, the threshold duration may be set to 4.4 hours. It should be understood that this duration is only exemplary and may be changed according to different criteria, requirements or instructions. In addition, an upper limit duration for recovering a predetermined number of disk sets in the plurality of disk sets is also determined in advance. As described above, the predetermined number may be associated with the concurrent number of the largest recovery jobs. In addition, the disk capacity of these disk sets to be recovered also needs to be determined in advance.
[0033] At 403, the difference between the total threshold duration and the upper limit duration may be determined. That is, the upper limit duration may be subtracted from the total threshold duration, so as to reserve time for the recovery operation of the remaining disk set 120 that is finally recovered. Even if the recovery operation of the remaining disk set 120 can only be performed at the lowest recovery rate that the system can provide, the actual duration consumed for recovering disk set 100 will not be longer than the total threshold duration determined in advance.
[0034] At 405, the first recovery rate can be further determined by calculating the ratio of the disk capacity to the above time difference. For example, a first threshold rate can be determined based on this ratio first. Then, at 407, the first recovery rate can be determined based on this first threshold rate. To ensure that the actual consumption duration of the recovery disk set 100 is not longer than the predetermined total threshold duration, the first recovery rate needs to be set to be greater than or equal to this first threshold rate. Preferably, the first recovery rate can be directly set to be equal to this first threshold rate, so as to occupy as few system resources as possible and thus reduce the impact on system performance.
[0035] Back to Figure 3 , after determining the first recovery rate, at 303, the number of disk sets in the above multiple disk sets that have not been recovered based on this first recovery rate can be determined. Further, at 305, the determined number of unrecovered disk sets can be compared with the above predetermined number. If the number of unrecovered disk sets is greater than or equal to this predetermined number, data recovery can be performed on the unrecovered disk sets in the multiple disk sets based on a predetermined second recovery rate. It should be understood that this second recovery rate is lower than this first recovery rate, and this second recovery rate is associated with the above upper limit duration. As an example, this second recovery rate can be the lowest recovery rate that the system can provide.
[0036] In some embodiments, in order to perform data recovery on the unrecovered disk sets in the multiple disk sets based on a predetermined second recovery rate, data recovery can be performed on the unrecovered disk sets in a single recovery job. Since the above upper limit duration is reserved in advance, the recovery operation on the remaining disk sets can still be completed within the predetermined total threshold duration with a single recovery job.
[0037] In some embodiments, the number of recovery jobs for parallel processing of the above multiple disk sets can be further determined at least based on the determined first recovery rate. That is to say, in order to make the recovery process reach the first recovery rate, the number of recovery jobs for parallel processing of the above multiple disk sets can be adjusted. Figure 5 A schematic diagram of a process 500 for determining the number of recovery jobs according to an embodiment of the present disclosure is shown.
[0038] At 501, a section of the recovery process can be executed with a first number of parallel jobs in a relatively short time period, so as to detect the disk capacity processed by this section of the recovery process, and thus determine a third recovery rate based on this disk capacity and this time period. As an example, the recovery process can be executed with 2 parallel jobs for 30 seconds to determine the third recovery rate. It should be understood that both this time period and the first number of parallel jobs can be set arbitrarily as needed, and only their linear relationship is used to determine the third recovery rate.
[0039] At 503, determine the occupancy of the processing nodes for performing the recovery job based at least on the third recovery rate. As an example, the occupancy of the processing nodes for performing the recovery job can be determined based on the third recovery rate, the type and width of the set of disks being recovered, and the first number of parallel jobs.
[0040] At 505, determine the number of recovery jobs based at least on the determined resource occupancy and the first recovery rate. As an example, the number of recovery jobs can be determined based on the first recovery rate, the type and width of the set of disks being recovered, and the determined resource occupancy.
[0041] As an example, the above process can be performed by looking up a table. An exemplary lookup table is shown in Table 1 below:
[0042] Table 1
[0043]
[0044] For example, the recovery process can be performed with 2 parallel jobs for a period of time to determine the third recovery rate. Assume that the third recovery rate is determined to be 850 (MB / s) and the set of disks being recovered is 8 + 2 RAID6. Then, the position with a parallel job number of 2 and a recovery rate close to 850 can be found in Table 1, and the occupancy of the processing nodes for performing the recovery job is determined to be 25% occupancy. After that, assume that the first recovery rate is determined to be 1050 (MB / s), the set of disks being recovered is 8 + 2 RAID6, and the resource occupancy has been determined to be 25%. Then, by looking up Table 1, the number of recovery jobs can be determined to be 4.
[0045] Table 2
[0046]
[0047]
[0048] Again, for example, the recovery process can be performed with 2 parallel jobs for a period of time to determine the third recovery rate. Assume that the third recovery rate is determined to be 550 (MB / s) and the set of disks being recovered is 16 + 2 RAID6. Then, the position with a parallel job number of 2 and a recovery rate close to 550 can be found in Table 1, and the occupancy of the processing nodes for performing the recovery job is determined to be 25% occupancy. After that, assume that the first recovery rate is determined to be 1250 (MB / s), the set of disks being recovered is 16 + 2 RAID6, and the resource occupancy has been determined to be 25%. Then, by looking up Table 2, the number of recovery jobs can be determined to be 8.
[0049] In this way, the first recovery rate for recovering most of the disk sets can be quickly determined. In addition, it is also possible to be compatible with the recovery of disk sets with different RAID types and different RAID widths. It should be understood that due to changes in factors such as the occupancy of processing nodes, the above process needs to be periodically executed to adjust the first recovery rate in a timely manner.
[0050] In addition, alternatively or additionally, a machine learning model can also be trained based on multiple calibrated training data sets to determine the first recovery rate for recovering most of the disk sets. For example, the machine learning model is trained through multiple sets of calibrated data including the first recovery rate, disk set type, resource occupancy, number of recovery jobs, etc., so that the number of concurrent jobs can be determined more precisely to adjust the first recovery rate in a timely manner.
[0051] In addition, in some embodiments, the disk sets to be recovered can also be sorted, and the disk sets with a larger number of unavailable storage disks and a larger RAID width are preferentially recovered, so that the disk sets with more unrecoverable risks can be preferentially recovered within the specified total threshold duration.
[0052] Through the above embodiments, by reserving time for the recovery operation with a lower recovery rate of the remaining disk sets, it can be ensured that all the disk sets to be recovered are completed within the recovery duration. In addition, due to the reasonable utilization of computing resources, the impact on system performance can be minimized. In addition, in the present disclosure, by creating and maintaining several lookup tables, the RAID types and RAID widths of the recovered disk sets can be different from each other, thus improving the compatibility of the recovery operation.
[0053] Figure 6 FIG. shows a schematic block diagram of an exemplary device 600 suitable for implementing the embodiments of the present disclosure. As shown in the figure, the device 600 includes a central processing unit (CPU) 601, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 602 or the computer program instructions loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0054] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0055] Each of the processes and treatments described above, such as methods 300, 400, and / or 500, may be executed by processing unit 601. For example, in some embodiments, methods 300, 400, and / or 500 may be implemented as computer software programs tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by CPU 601, one or more actions of methods 300, 400, and / or 500 described above may be performed.
[0056] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0057] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example (but not limited to), an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions, and any suitable combination of the foregoing. A computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0058] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0059] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0060] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0061] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0062] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0063] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0064] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art in the field to understand the embodiments disclosed herein.
Claims
1. A storage management method, comprising: Determining a first recovery rate based at least on an upper limit duration for recovering a predetermined number of disk sets in a plurality of disk sets, for recovering at least a part of the plurality of disk sets; Determining the number of disk sets in the plurality of disk sets that are not recovered based on the first recovery rate; And Based on determining that the number is less than or equal to the predetermined number, performing data recovery on the unrecovered disk sets in the plurality of disk sets based on a predetermined second recovery rate, the second recovery rate being lower than the first recovery rate and associated with the upper limit duration.
2. The method according to claim 1, wherein determining the first recovery rate includes: Determining a total threshold duration for recovering the plurality of disk sets, the upper limit duration, and the disk capacities of the plurality of disk sets; Determining the difference between the total threshold duration and the upper limit duration; Determining a first threshold rate by calculating a ratio of the disk capacity to the difference; And Determining the first recovery rate based on the first threshold rate, the first recovery rate being greater than or equal to the first threshold rate.
3. The method according to claim 1, further comprising: Determining, based at least on the first recovery rate, the number of recovery jobs for parallel processing of the recovery of the plurality of disk sets.
4. The method according to claim 3, wherein determining the number of recovery jobs based at least on the first recovery rate includes: Determining the resource occupancy of a processing node for executing the recovery jobs; And Determining the number of recovery jobs based at least on the resource occupancy and the first recovery rate.
5. The method according to claim 4, wherein determining the resource occupancy includes: Determining a third recovery rate based on the disk capacities of the disk sets parallel processed by a first number of jobs and the parallel processing time; And Determining the occupancy of the processing node for executing the recovery jobs based at least on the third recovery rate.
6. The method according to claim 1, wherein performing the data recovery includes: Performing data recovery on the unrecovered disk sets in a single recovery job.
7. An electronic device, comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing machine-executable instructions that, when executed by the at least one processing unit, cause the device to perform actions, the actions including: Determining a first recovery rate based at least on an upper limit duration for recovering a predetermined number of disk sets in a plurality of disk sets, for recovering at least a part of the plurality of disk sets; Determining the number of disk sets in the plurality of disk sets that are not recovered based on the first recovery rate; and Based on determining that the number is less than or equal to the predetermined number, performing data recovery on the unrecovered disk sets in the plurality of disk sets based on a predetermined second recovery rate, the second recovery rate being lower than the first recovery rate and associated with the upper limit duration.
8. The device according to claim 7, wherein determining the first recovery rate includes: Determine a total threshold duration for restoring the plurality of disk sets, the upper limit duration, and the disk capacity of the plurality of disk sets; Determine the difference between the total threshold duration and the upper limit duration; Determine a first threshold rate by calculating a ratio of the disk capacity to the difference; and Determine the first recovery rate based on the first threshold rate, the first recovery rate being greater than or equal to the first threshold rate.
9. The apparatus according to claim 7, wherein the action further comprises: Determine the number of recovery jobs for parallel processing of the plurality of disk sets at least based on the first recovery rate.
10. The apparatus according to claim 9, wherein determining the number of recovery jobs at least based on the first recovery rate comprises: Determine the resource occupancy of the processing nodes for executing the recovery jobs; and Determine the number of recovery jobs at least based on the resource occupancy and the first recovery rate.
11. The apparatus according to claim 10, wherein determining the resource occupancy comprises: Determine a third recovery rate based on the disk capacity of the disk sets processed in parallel by a first number of jobs and the parallel processing time; and Determine the occupancy of the processing nodes for executing the recovery jobs at least based on the third recovery rate.
12. The apparatus according to claim 7, wherein performing the data recovery comprises: Performing data recovery on the un-recovered disk sets in a single recovery job.
13. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Storage controller having dynamic voltage for regulating super capacitance
CN101203825A
System and method for improved rebuild in RAID
CN103246480A