Redundant array of independent disks (RAID) reconstruction method and electronic equipment

By monitoring physical disk and service I/O performance metrics, and dynamically adjusting resource allocation during the RAID reconstruction process, the problem of resource contention between reconstruction I/O and service I/O was solved, reconstruction efficiency was improved, service performance loss was reduced, and service quality was guaranteed.

CN120892264AActive Publication Date: 2025-11-04LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511423226.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-04
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing RAID reconstruction technology lacks a fine-grained dynamic control mechanism when reconstructing I/O and competing for business I/O resources, which affects business continuity in response time-sensitive scenarios such as financial databases.

Method used

By monitoring physical disk performance metrics, obtaining business I/O performance metrics, selecting the target reconstruction mode, and generating resource allocation decisions based on real-time collected reconstruction performance metrics, the system dynamically adjusts the input/output bandwidth and processor usage during the reconstruction process, thereby achieving intelligent adjustment of RAID reconstruction.

Benefits of technology

It improves RAID reconstruction efficiency, reduces service performance loss, enables resource allocation selection under different service conditions, and balances reconstruction efficiency and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892264A_ABST
    Figure CN120892264A_ABST
Patent Text Reader

Abstract

The invention discloses a redundant array of independent disks (RAID) reconstruction method and electronic equipment, and relates to the technical field of storage. According to the scheme, multiple reconfiguration modes are preset and used for specifying different input and output bandwidth occupancy amounts and processor occupancy amounts in the reconfiguration process, multiple resource occupancy selections are provided for RAID reconfiguration, and the method is suitable for different service conditions; according to the method, the target reconfiguration mode is selected from the reconfiguration modes based on the current service input / output performance index when the fact that the faulty physical disk exists in the RAID is confirmed, the influence of RAID reconfiguration on the current service is considered, the optimal reconfiguration mode is selected according to the actual service requirement, the reconfiguration efficiency is improved, and the service performance loss is reduced; and according to the target reconstruction mode and the reconstruction performance index acquired in real time, a reconstruction resource allocation decision is generated, and reconstruction is executed according to the decision, so that RAID reconstruction efficiency and business service quality are effectively considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of storage, and in particular to a redundant array of independent disks reconstruction method and electronic equipment. BACKGROUND

[0002] Redundant array of independent disks (RAID) reconstruction is a process of recovering lost data to a hot spare disk or a new disk using redundant information after a member disk fails, and its speed and efficiency directly affect the data availability window period, which is a key indicator of measuring the reliability of a storage system.

[0003] However, existing RAID reconstruction techniques mainly focus on accelerating the reconstruction process, but ignore the resource competition problem between reconstruction input / output (I / O) and business I / O, and lack a fine-grained dynamic control mechanism to guarantee business service quality during reconstruction. In particular, in financial databases and other scenarios sensitive to response time, the reconstruction peak period may cause a significant increase in the response time of critical transactions, seriously affecting business continuity. There is currently no intelligent strategy to adjust the reconstruction speed based on business priority and system load in real time to effectively balance the relationship between reconstruction efficiency and business performance.

[0004] In view of the above, how to solve the problems of the current RAID reconstruction technology, such as resource competition between reconstruction I / O and business I / O, lack of business service quality guarantee and intelligent adjustment strategy, is a problem that needs to be solved by technical personnel in this field. SUMMARY

[0005] The present application provides a redundant array of independent disks reconstruction method and electronic equipment to at least solve the problems of the current RAID reconstruction technology, such as resource competition between reconstruction I / O and business I / O, lack of business service quality guarantee and intelligent adjustment strategy.

[0006] The present application provides a redundant array of independent disks reconstruction method, comprising: monitoring the performance indicators of each physical disk and determining whether there is a faulty physical disk in each physical disk according to the performance indicators; if so, obtaining business input / output performance indicators; wherein the business input / output performance indicators at least include the total throughput, average delay and quantile delay of business input / output; The target reconstruction mode is selected in each reconstruction mode according to the size relationship between the total throughput, average delay and quantile delay of the service input / output and the corresponding threshold value; the resource occupation of the service is in a negative correlation relationship with the resource occupation in the target reconstruction mode; the reconstruction mode is used to determine the input / output bandwidth occupation and processor occupation of the reconstruction process; the reconstruction mode includes a first reconstruction mode, a second reconstruction mode and a third reconstruction mode; the maximum input / output bandwidth of the reconstruction in the first reconstruction mode is less than the maximum input / output bandwidth of the reconstruction in the third reconstruction mode, and the maximum input / output bandwidth of the reconstruction in the third reconstruction mode is less than the maximum input / output bandwidth of the reconstruction in the second reconstruction mode; the processor priority weight in the first reconstruction mode is less than the processor priority weight in the third reconstruction mode, and the processor priority weight in the third reconstruction mode is less than the processor priority weight in the second reconstruction mode; According to the target reconstruction mode and the real-time collected reconstruction performance index, a reconstruction resource allocation decision is generated. Based on the reconstruction resource allocation decision, the reconstruction of each stripe in the faulty physical disk is performed.

[0007] The application further provides an electronic device, comprising a memory for storing a computer program and a processor for executing the computer program to realize the steps of any one of the RAID reconstruction methods.

[0008] The application has the advantages that multiple reconstruction modes are set in advance, different input / output bandwidth occupation and processor occupation in the reconstruction process are limited respectively, multiple resource occupation options are provided for the RAID reconstruction, and the application is suitable for different service conditions; when it is confirmed that there is a faulty physical disk in the RAID according to the performance index of the physical disk, the target reconstruction mode is selected in each reconstruction mode based on the current service input / output performance index, the influence of the RAID reconstruction on the current service is considered, the best reconstruction mode is selected according to the actual service demand, the reconstruction efficiency is improved, and the service performance loss is reduced; the reconstruction resource allocation decision is generated according to the target reconstruction mode and the real-time collected reconstruction performance index, and the reconstruction is performed according to the decision, the system resources allocated to the reconstruction task are dynamically and accurately fine-tuned, and the RAID reconstruction efficiency and the service quality are effectively balanced.

[0009] In addition, the application further provides an electronic device, and the effects are the same as above. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A flow chart of a method for reconstructing a redundant array of independent disks according to an embodiment of the present application is shown in FIG. 1. Figure 2 A schematic diagram of a device for reconstructing a redundant array of independent disks according to an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, any other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0013] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0014] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0015] Currently, the RAID reconstruction technology mainly focuses on accelerating the reconstruction process, but ignores the resource competition between reconstruction I / O and business I / O, and lacks a fine dynamic control mechanism to guarantee the quality of service during reconstruction. In particular, in the financial database and other scenarios sensitive to response time, the reconstruction peak period may cause a significant increase in the response time of critical transactions, seriously affecting business continuity. There is currently no intelligent strategy to adjust the reconstruction speed in real time based on business priority and system load to effectively balance the relationship between reconstruction efficiency and business performance. Therefore, in order to solve the above problems, the present application provides a method for reconstructing a redundant array of independent disks.

[0016] Figure 1 A flow chart of a method for reconstructing a redundant array of independent disks according to an embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the method comprises the following steps: S10: Monitor the performance indicators of each physical disk, and determine whether there is a faulty physical disk in each physical disk according to the performance indicators; if yes, go to step S11; if no, end.

[0017] Specifically, the storage system is mainly composed of a plurality of physical disks and a RAID controller. During the system operation, the performance indicators of each physical disk are monitored in real time, including but not limited to the I / O utilization rate, read-write bandwidth, average response time and queue depth of the physical disk, etc.

[0018] In order to determine whether to trigger the RAID reconstruction, in this embodiment, according to the performance indicators of the physical disks, it is judged whether there is a faulty physical disk in each physical disk. It can be understood that in this embodiment, the reconstruction is mainly performed on the stripes in the faulty physical disk, and the offline of the faulty physical disk is the core condition for the reconstruction trigger. Therefore, in order to determine whether the RAID reconstruction needs to be performed, it is necessary to first judge whether there is a faulty physical disk. Specifically, the physical disk with abnormal performance indicators in each physical disk can be determined as a faulty physical disk through the self-monitoring analysis and reporting technology (Self-Monitoring, Analysis, and Reporting Technology, SMART). In this embodiment, the specific content of the abnormal performance indicators is not limited, for example, when the average response time of a physical disk is continuously timed out, or the disk I / O utilization rate is abnormally zeroed or stuck and the read-write bandwidth is simultaneously reduced to 0 MB / s, or the disk queue depth is abnormally stacked and not released, the physical disk can be considered as a faulty physical disk. In addition, the offline physical disk in each physical disk can also be directly or through the I / O timeout detection technology determined as a faulty physical disk.

[0019] S11: Obtain service input / output performance indicators.

[0020] The service input / output performance indicators at least include total throughput, average delay and quantile delay of the service input / output.

[0021] It should be noted that the quantile delay includes 95 quantile delay and 99 quantile delay, which are two key indicators for measuring system performance, respectively representing that 95% and 99% of all requests are completed within the time. In short, they represent the response speed of the system in most cases, helping to understand the performance of the system under high load or extreme conditions. In this embodiment, the specific selected quantile delay is not limited.

[0022] S12: According to the size relationship between the total throughput, average delay and quantile delay of the service input / output and the corresponding threshold value, respectively, select the target reconstruction mode in each reconstruction mode.

[0023] When it is confirmed that there is a faulty physical disk, the RAID reconstruction process is started. Specifically, according to the size relationship between the total throughput, average delay and quantile delay of the service input / output and the corresponding threshold value, respectively, the target reconstruction mode is selected in each reconstruction mode.

[0024] It should be noted that the reconstruction mode is used to determine the I / O bandwidth occupation and the CPU occupation of the reconstruction process, that is, the reconstruction mode indicates the I / O bandwidth occupation and the CPU resource occupation allowed to be used when the reconstruction is performed using the mode, and the resource occupation of the service is negatively correlated with the resource occupation under the target reconstruction mode. It should be noted that the I / O bandwidth occupation and the processor occupation corresponding to each reconstruction mode are different, that is, the amount of reconstruction resources allowed to be used under each reconstruction mode is different. In this embodiment, the reconstruction mode includes a first reconstruction mode, a second reconstruction mode, and a third reconstruction mode; the maximum input / output bandwidth of the reconstruction under the first reconstruction mode is less than the maximum input / output bandwidth of the reconstruction under the third reconstruction mode, and the maximum input / output bandwidth of the reconstruction under the third reconstruction mode is less than the maximum input / output bandwidth of the reconstruction under the second reconstruction mode; the processor priority weight under the first reconstruction mode is less than the processor priority weight under the third reconstruction mode, and the processor priority weight under the third reconstruction mode is less than the processor priority weight under the second reconstruction mode. The specific setting mode of the three reconstruction modes in this embodiment is not limited.

[0025] In order to take into account the normal service I / O and realize intelligent adjustment of the service resource occupation and the RAID reconstruction resource occupation, in this embodiment, the target reconstruction mode is selected from the pre-configured reconstruction modes according to the collected service I / O performance indicators, so as to reduce the influence of the RAID reconstruction process on the service. The specific selection mode of the target reconstruction mode in this embodiment is not limited, and is determined according to the specific implementation condition.

[0026] S13: generating a reconstruction resource allocation decision according to the target reconstruction mode and the real-time collected reconstruction performance indicators.

[0027] S14: performing reconstruction on each stripe in the faulty physical disk based on the reconstruction resource allocation decision.

[0028] Subsequently, to adapt to real-time resource changes during the reconstruction process, while improving reconstruction efficiency and further reducing the impact of reconstruction on services, this embodiment also needs to generate a reconstruction resource allocation decision based on the target reconstruction mode and real-time collected reconstruction performance indicators. This reconstruction resource allocation decision essentially determines the I / O bandwidth and CPU resource usage for reconstruction. It's important to note that this embodiment does not limit the specific content of the reconstruction performance indicators; for example, they may include reconstruction progress, throughput, and resource consumption (CPU resource consumption, memory resource consumption, and I / O bandwidth consumption). The generation process of the reconstruction resource allocation decision is also not limited and depends on the specific implementation. Finally, based on the reconstruction resource allocation decision, the reconstruction of each stripe in the faulty physical disk is performed, achieving dynamic and precise fine-tuning of the system resources allocated to the reconstruction task, effectively balancing RAID reconstruction efficiency and service quality. This embodiment does not limit the specific process of reconstructing each stripe in the faulty physical disk and depends on the specific implementation.

[0029] In this embodiment, multiple reconstruction modes are pre-set, each limiting different input / output bandwidth and processor usage during the reconstruction process. This provides multiple resource usage options for RAID reconstruction, suitable for different business scenarios. When a faulty physical disk is confirmed in the RAID based on the physical disk's performance indicators, a target reconstruction mode is selected from among the various reconstruction modes based on the current business input / output performance indicators. This considers the impact of RAID reconstruction on the current business, enabling the selection of the optimal reconstruction mode based on actual business needs, improving reconstruction efficiency and reducing business performance loss. Reconstruction resource allocation decisions are generated based on the target reconstruction mode and real-time collected reconstruction performance indicators, and reconstruction is executed accordingly. This achieves dynamic and precise fine-tuning of system resources allocated to the reconstruction task, effectively balancing RAID reconstruction efficiency and business service quality.

[0030] To better manage the stripes in a RAID array and adapt to RAID refactoring mechanisms, based on the above embodiments, some embodiments further include the following before monitoring the performance metrics of each physical disk: S101: Divide the logical volume space into multiple consecutive stripes based on data access popularity information.

[0031] S102: Collect the real-time disk utilization rate of each physical disk.

[0032] S103: Determine the disk weight of each physical disk based on the real-time utilization of each disk and the standard scalable hash-controlled replication weight.

[0033] S104: Store the data blocks and parity blocks in each stripe on different physical disks according to the weight of each disk.

[0034] S105: Obtain the position information of each stripe.

[0035] The position information at least includes a stripe identifier, a physical disk identifier of a physical disk where the stripe is located, a logical block range, and a state tag. The logical block range is a starting logical block address and a size of the stripe on the corresponding physical disk.

[0036] S106: Generate a dynamic stripe mapping table according to the position information of each stripe.

[0037] Specifically, in the embodiment, the logical volume space is divided into multiple continuous stripes according to the data access heat information. The sizes of the stripes range from 16 KB to 1 MB, and the specific size is dynamically adjusted based on the data access heat. For example, a hot area uses a smaller stripe (such as 16 KB) to improve random access efficiency, and a cold data area uses a larger stripe (such as 1 MB) to reduce metadata overhead.

[0038] Further, for each stripe, the data blocks (Data Blocks) and the parity block (Parity Block) contained therein are stored on different physical disks according to the RAID level (such as RAID5). It should be noted that, in order to avoid placing the associated data (data blocks and parity blocks) of the same stripe on the disks that may be overloaded at the same time, the improved controlled replication under scalable hashing (CRUSH) algorithm is used in the embodiment to realize the storage of stripe data.

[0039] Specifically, the real-time utilization rates of the physical disks are collected, and the disk weights of the physical disks are determined according to the real-time utilization rates of the disks and the standard CRUSH weight, as follows: Disk weight = standard CRUSH weight × (1 - real-time utilization rate of disk / 100); As can be seen from the above formula, the higher the real-time utilization rate of the disk, the lower the corresponding disk weight, thereby avoiding data allocation, that is, data is preferentially stored in the physical disk with a low real-time utilization rate. Finally, the data blocks and the parity blocks in each stripe are stored in different physical disks according to the disk weights. It should be noted that, in the data allocation process, the data blocks + parity blocks of the same stripe must satisfy: physical dispersion (RAID standard), and the real-time utilization rates of all disks are less than the disk utilization rate threshold. In this way, the traditional fixed-size stripe is subdivided into smaller and dynamically adjustable stripes, each of which independently contains its data blocks and redundancy information, and the improved CRUSH algorithm is used to store them on multiple physical disks, which greatly reduces the basic granularity of reconstruction and increases the parallel potential.

[0040] Meanwhile, in order to determine the position information and the usage state of each stripe in the RAID, after the data blocks and the check blocks in each stripe are respectively stored in different physical disks according to the weight of each disk, the position information of each stripe needs to be obtained. It should be noted that the position information at least includes a stripe identifier (ID), a physical disk ID of a physical disk where the stripe is located, a logical block address (LBA) range, and a state flag. Among them, the logical block address range is the starting LBA and size of the stripe on the corresponding physical disk, and the state flag represents whether the stripe is in a normal state or a frozen state. It can be understood that after a fault is triggered, all the stripes of the fault disk are marked as frozen and are prohibited from being written. Finally, a dynamic stripe mapping table is generated according to the position information of each stripe, which is stored in the memory of the RAID controller and is persisted to the metadata area. In this way, the dynamic stripe mapping table is generated based on the position information of each stripe, which can determine the position information and the usage state of each stripe in the RAID, and is beneficial to the unified management of each stripe.

[0041] On the basis of the above-mentioned embodiments, in some embodiments, the configuration process of the reconstruction mode includes: S111: setting the reconstruction maximum input / output bandwidth in the first reconstruction mode to be not greater than a first percentage of the total bandwidth, and the processor priority weight to be not greater than a second percentage of the total processor occupancy.

[0042] S112: setting the reconstruction maximum input / output bandwidth in the second reconstruction mode to be not less than a third percentage of the total bandwidth, and the processor priority weight to be not less than a fourth percentage of the total processor occupancy.

[0043] S113: setting the reconstruction maximum input / output bandwidth in the third reconstruction mode to be within a first range, and the processor priority weight to be within a second range.

[0044] Among them, the first percentage is less than the third percentage, and the second percentage is less than the fourth percentage; the first range is determined based on the first percentage of the total bandwidth and the third percentage of the total bandwidth, and the second range is determined based on the second percentage of the total processor occupancy and the fourth percentage of the total processor occupancy.

[0045] The core of the reconstruction mode is the corresponding rule of the mode target and the resource allocation threshold, which ensures that the resource allocation does not deviate from the core demand of the mode. In the embodiment, three reconstruction modes are set, and the fixed upper and lower limits of the reconstruction maximum I / O bandwidth and the CPU priority weight are set for each mode to avoid extreme resource allocation. It can be understood that the reconstruction maximum I / O bandwidth is the maximum I / O bandwidth allowed to be used in the reconstruction process. The CPU priority weight is a proportional parameter to measure the "priority" of the reconstruction task in the CPU resource competition, which is essentially the "proportion of CPU time slice that the reconstruction task can occupy". Its value range is 0%-100%, and the sum of the business task CPU weight is 100%; at the same time, the lower the CPU priority weight, the lower the CPU scheduling priority, that is, when the CPU resource is tight, the operating system will preferentially allocate time slices to business tasks, and the reconstruction task only occupies resources when the business CPU is idle.

[0046] Specifically, in the embodiment, the reconstruction maximum I / O bandwidth in the first reconstruction mode is set to be not greater than the first percentage of the total bandwidth, and the processor priority weight is set to be not greater than the second percentage of the total processor occupancy. The first reconstruction mode can be considered as a business high-performance mode, in which the business CPU and bandwidth are preferentially guaranteed. The reconstruction maximum I / O bandwidth in the second reconstruction mode is set to be not less than the third percentage of the total bandwidth, and the processor priority weight is set to be not less than the fourth percentage of the total processor occupancy. The second reconstruction mode can be considered as a fast reconstruction mode, in which the reconstruction resource is preferentially met. The reconstruction maximum I / O bandwidth in the third reconstruction mode is set to be within the first range, and the processor priority weight is set to be within the second range. The third reconstruction mode can be considered as a balanced mode, which takes into account the business resource occupancy and the reconstruction resource occupancy.

[0047] It should be noted that the sizes of the first percentage, the second percentage, the third percentage and the fourth percentage are not limited in the embodiment, but it is necessary to ensure that the first percentage is less than the third percentage, and the second percentage is less than the fourth percentage. For example, the first percentage is 30%, the second percentage is 20%, the third percentage is 60%, and the fourth percentage is 50%, then: the reconstruction maximum I / O bandwidth in the first reconstruction mode is set to be not greater than 30% of the total bandwidth, and the processor priority weight is set to be not greater than 20% of the total processor occupancy, the reconstruction maximum I / O bandwidth in the second reconstruction mode is set to be not less than 60% of the total bandwidth, and the processor priority weight is set to be not less than 50% of the total processor occupancy.

[0048] Meanwhile, the first range is determined based on a first percentage of the total bandwidth and a third percentage of the total bandwidth, and the second range is determined based on a second percentage of the total processor occupancy and a fourth percentage of the total processor occupancy. That is, the reconstruction maximum I / O bandwidth and the CPU priority weight in the third reconstruction mode can be selected within the corresponding range. For example, the reconstruction maximum I / O bandwidth can be adjusted to 20% to 50%, and the CPU priority weight can be adjusted to 20% to 40%.

[0049] In the embodiment, by setting three different reconstruction modes, different resource occupancy options are provided for the reconstruction process, which can be applied to different service conditions.

[0050] On the basis of the above embodiment, in some embodiments, according to the size relationship between the total throughput, the average delay and the quantile delay of the service input / output and the corresponding threshold, the target reconstruction mode is selected in each reconstruction mode, including: S121: When the total throughput of the service input / output is not less than the first daily throughput threshold, or the quantile delay is greater than the quantile delay threshold, the first reconstruction mode is selected as the target reconstruction mode.

[0051] S122: When the total throughput of the service input / output is not greater than the second daily throughput threshold, and the average delay is less than the first average delay threshold, the second reconstruction mode is selected as the target reconstruction mode.

[0052] S123: When the average delay is less than the second average delay threshold, and the disk input / output utilization of each physical disk is less than the disk input / output utilization threshold, the third reconstruction mode is selected as the target reconstruction mode.

[0053] Wherein, the first daily throughput threshold is greater than the second daily throughput threshold; the first average delay threshold is less than the second average delay threshold.

[0054] In order to accurately select the target reconstruction mode, the core demand of the current service needs to be determined first, which is preset by the user or the system, and then it is judged by the real-time performance index collected which mode can be implemented under the demand, forming a closed loop of demand-parameter verification-mode matching.

[0055] Therefore, when the total throughput of the business I / O is not less than the first daily throughput threshold, or the quantile delay is greater than the quantile delay threshold (the 95th quantile delay is greater than the 95th quantile delay threshold, or the 99th quantile delay is greater than the 99th quantile delay threshold), the first reconstruction mode is selected as the target reconstruction mode, that is, the business high-performance mode is selected. The applicable scenario of this mode is that the business is extremely sensitive to delay / throughput, the core requirement is that the business cannot be blocked during reconstruction, and the extension of reconstruction time is acceptable. In this embodiment, the sizes of the first daily throughput threshold, the 95th quantile delay threshold, and the 99th quantile delay threshold are not limited, for example, the first daily throughput threshold is 80% of the daily peak, and the 95th quantile delay threshold and the 99th quantile delay threshold are 2.8 ms.

[0056] When the total throughput of the business I / O is not greater than the second daily throughput threshold, and the average delay is less than the first average delay threshold, the second reconstruction mode is selected as the target reconstruction mode, that is, the fast reconstruction mode is selected. The applicable scenario of this mode is that the business is in the "low period", or the data security requirement is extremely high, such as medical data and core backup; the core requirement is to get out of the degraded state as soon as possible to avoid data loss caused by double disk failure, and a slight decrease in business performance is acceptable. In this embodiment, the sizes of the second daily throughput threshold and the first average delay threshold are not limited, as long as the first daily throughput threshold is greater than the second daily throughput threshold, for example, the second daily throughput threshold is 30% of the daily peak, and the first average delay threshold is 1.2 ms.

[0057] When the average delay is less than the second average delay threshold, and the disk I / O utilization rate of each physical disk is less than the disk I / O utilization rate threshold, the third reconstruction mode is selected as the target reconstruction mode, that is, the balanced mode is selected. The applicable scenario of this mode is that the business requirement has no extreme tendency, that is, neither absolute non-jamming nor fastest reconstruction is required, and the core target is that the reconstruction does not significantly affect the business. In this embodiment, the second average delay threshold and the disk I / O utilization rate threshold are not limited, as long as the first average delay threshold is less than the second average delay threshold, for example, the second average delay threshold is 3 ms, and the disk I / O utilization rate threshold is 85%.

[0058] In this embodiment, the target reconstruction mode is selected according to the actual business requirement, so that the business performance loss is within a controllable range.

[0059] On the basis of the above-mentioned embodiments, in some embodiments, after the target reconstruction mode is selected, the method further includes: S131: updating the business input / output performance index according to a preset period.

[0060] S132: selecting a new target reconstruction mode in the reconstruction mode according to the updated business input / output performance index.

[0061] It should be noted that the target reconstruction mode selected in the above embodiments is not fixed, but can be dynamically switched in real time according to the business I / O performance indicators, to ensure that it is always adapted to the current state. Specifically, the business input / output performance indicators are updated according to a preset period, including but not limited to the total throughput, average delay, and quantile delay (95th quantile delay or 99th quantile delay) of the business I / O. Then, according to the updated business input / output performance indicators, a new target reconstruction mode is selected in the reconstruction mode.

[0062] For example, in the third reconstruction mode, if the business suddenly enters a peak period, the business delay increases from 1.8 ms to 2.9 ms, close to the threshold, then switch to the first reconstruction mode, reduce the reconstruction bandwidth to 30%, and ensure that the business does not freeze. In the second reconstruction mode, if the business suddenly has a temporary high load, the throughput increases from 100 MB / s to 500 MB / s, then temporarily switch to the third reconstruction mode to avoid a loss of more than 10% in business performance.

[0063] In this embodiment, a new target reconstruction mode is selected in combination with the real-time changing business I / O performance indicators, to ensure the adaptability of reconstruction to the business.

[0064] On the basis of the above embodiments, in some embodiments, after selecting the target reconstruction mode, the following steps are further included: S133: When the third reconstruction mode is selected as the target reconstruction mode, it is determined whether the average delay is less than the second average delay threshold and the disk input / output utilization of each physical disk is less than the disk input / output utilization threshold after a preset time; if yes, the reconstruction maximum input / output bandwidth is increased; if no, the reconstruction maximum input / output bandwidth is reduced.

[0065] In a specific implementation, when the third reconstruction mode (i.e., the balanced mode) is selected as the target reconstruction mode, since the reconstruction maximum I / O bandwidth in the third reconstruction mode is in the first range and the CPU priority weight is in the second range, to determine the specific reconstruction maximum I / O bandwidth and CPU priority weight, specific resource adjustment can also be made based on the mode.

[0066] Specifically, it is determined whether the average delay is less than the second average delay threshold and the disk I / O utilization of each physical disk is less than the disk I / O utilization threshold after a preset time. If yes, the reconstruction maximum I / O bandwidth is gradually increased, for example, by 10%. If no, the reconstruction maximum I / O bandwidth is immediately reduced, for example, by 20%. In this way, the exclusive adjustment strategy added for the third reconstruction mode is the core of the "dynamic balance" of the third reconstruction mode, and the scene-based adaptation of reconstruction acceleration and business guarantee is realized through real-time business performance indicator feedback.

[0067] On the basis of the above embodiments, in some embodiments, according to the target reconstruction mode and the real-time collected reconstruction performance index, a reconstruction resource allocation decision is generated, including: S141: Obtain the initial reconstruction maximum input / output bandwidth and the initial processor priority weight specified by the target reconstruction mode.

[0068] S142: Determine the fine-tuning amount of the reconstruction maximum input / output bandwidth according to the real-time collected reconstruction performance index.

[0069] S143: Calculate the target reconstruction maximum input / output bandwidth according to the initial reconstruction maximum input / output bandwidth and the fine-tuning amount.

[0070] S144: Calculate the target processor priority weight according to the target reconstruction maximum input / output bandwidth and the initial processor priority weight.

[0071] In order to generate the reconstruction resource allocation decision, in the embodiment, the initial reconstruction maximum I / O bandwidth and the initial CPU priority weight specified by the target reconstruction mode are specifically obtained. Subsequently, the adjustment rule is triggered, the fine-tuning amount of the reconstruction maximum I / O bandwidth is determined according to the real-time collected reconstruction performance index, and the target reconstruction maximum I / O bandwidth is calculated according to the initial reconstruction maximum I / O bandwidth and the fine-tuning amount.

[0072] For example, when the initial reconstruction maximum I / O bandwidth is 30% of the total bandwidth, the real-time detection shows that the service average delay is less than the average delay threshold value, and the current disk utilization rate is less than the disk utilization rate threshold value, the upgrade bandwidth rule is met, the fine-tuning amount is determined to be increased by 10%, and the final target reconstruction maximum I / O bandwidth is 40% of the total bandwidth. The real-time detection shows that the service average delay is greater than the average delay threshold value, and the single disk utilization rate 88% is greater than the disk utilization rate threshold value, the downgrade bandwidth rule is met, and the final target reconstruction maximum I / O bandwidth quota is 10% of the total bandwidth by fine-tuning with a step of 20%.

[0073] Finally, the target CPU priority weight is calculated according to the target reconstruction maximum I / O bandwidth and the initial CPU priority weight, and the formula is specifically: target CPU priority weight=(target reconstruction maximum I / O bandwidth / total bandwidth) x CPU weight coefficient of the current reconstruction mode. In the embodiment, the size of the CPU weight coefficient under different reconstruction modes is not limited, for example, the CPU weight coefficient of the first reconstruction mode is 0.8, the CPU weight coefficient of the second reconstruction mode is 1.5, and the CPU weight coefficient of the third reconstruction mode is 1.2.

[0074] In the embodiment, the initial reconstruction maximum input / output bandwidth and initial processor priority weight specified for the target reconstruction mode are fine-tuned to generate the final target reconstruction maximum input / output bandwidth and target processor priority weight for reconstruction, so as to ensure that the resource allocation meets the mode target and is adapted to the real-time system state.

[0075] On the basis of the above embodiment, in some embodiments, the reconstruction of each stripe in the failed physical disk is performed based on the reconstruction resource allocation decision, including: S151: Determine the position information of each stripe in the failed physical disk according to the dynamic stripe mapping table, and generate a corresponding to-be-reconstructed stripe list.

[0076] S152: Determine the source physical disk associated with the failed physical disk based on the to-be-reconstructed stripe list.

[0077] S153: Determine the target disk in each physical disk based on each performance index.

[0078] S154: Divide the to-be-reconstructed stripe list into a plurality of task groups according to the load information of the source physical disk and the target disk; wherein, each task group contains a preset number of to-be-reconstructed stripes.

[0079] S155: Decompose each task group into a plurality of subtasks, and put each subtask into a shared task queue.

[0080] S156: Parallelly acquire and execute each subtask in the shared task queue through a plurality of processing units, calculate the corresponding reconstruction data block through the Redundant Array of Independent Disks algorithm, and store it to the buffer.

[0081] S157: When the data in the buffer reaches the corresponding storage threshold, sequentially write each reconstruction data block in the buffer to the target disk.

[0082] In order to realize the RAID reconstruction, in the embodiment, first, the position information of each stripe in the failed physical disk is determined according to the dynamic stripe mapping table, and a corresponding to-be-reconstructed stripe list is generated. It can be understood that the to-be-reconstructed stripe list contains the position information of the to-be-reconstructed stripe, including stripe, belonging to the failed disk ID, LBA range and stripe size.

[0083] Further, based on the list of to-be-reconstructed stripes, a source physical disk associated with the failed physical disk is determined. It should be noted that the source physical disk is an associated surviving disk of the to-be-reconstructed stripe in the RAID. In order to determine the source physical disk, in some embodiments, the non-failed physical disks corresponding to the failed physical disk are queried based on the dynamic stripe mapping table. It should be noted that the data blocks in the non-failed physical disks and the data blocks in the failed physical disk are associated data; the abnormal physical disks in each non-failed physical disk, such as the disks that have been offline or failed, are removed to determine the source physical disk. For example, the stripe M001 is originally stored in the failed physical disk Disk3, the associated data blocks are in Disk0, and the check blocks are in Disk4, so Disk0 and Disk4 are the source physical disks.

[0084] Subsequently, based on the performance indicators, a target disk is determined in each physical disk. It should be noted that the target disk is a physical disk used to store the data of the to-be-reconstructed stripe after reconstruction. In order to determine the target disk, in some embodiments, the non-failed physical disks other than the failed physical disk in each physical disk are specifically determined; the disk I / O utilization rate and the remaining storage space of each non-failed physical disk are determined, and the non-failed physical disk corresponding to the disk I / O utilization rate less than the disk I / O utilization rate threshold, the remaining storage space greater than the remaining storage space threshold, and not in the same load domain as the source physical disk is determined as the target disk.

[0085] Further, according to the load information of the source physical disk and the target disk, the list of to-be-reconstructed stripes is divided into a plurality of task groups (Task Group). It should be noted that a task group contains a predetermined number of to-be-reconstructed stripes, for example, 32. For each task group, it is further decomposed into a plurality of sub-tasks (sub-task), and each sub-task is placed in a shared task queue. It can be understood that each sub-task contains the position information of the corresponding to-be-reconstructed stripe, the information of the source physical disk, and the information of the target disk.

[0086] Finally, each sub-task in the shared task queue is acquired and executed in parallel by a plurality of processing units, the corresponding reconstruction data block is calculated by the RAID algorithm, and stored to a buffer. When the data in the buffer reaches the corresponding storage threshold, each reconstruction data block in the buffer is sequentially written to the target disk, thereby completing the RAID reconstruction process.

[0087] It should be noted that the specific process of calculating the corresponding reconstruction data block by the RAID algorithm in this embodiment is not limited, for example, when using RAID5 level, the XOR algorithm is used to calculate the reconstruction data block; when using RAID6 level, the Reed-Solomon algorithm is used to calculate the reconstruction data block, which is determined according to the specific implementation.

[0088] Further, when a sub-task is successfully executed, the system automatically triggers a series of key operations to ensure data consistency and optimize reconstruction efficiency. First, the system enforces data integrity verification by calculating the CRC value of the target disk data and comparing it with the source data and the expected value of the RAID algorithm, ensuring that the reconstructed data is consistent with the original data. This real-time verification at the sub-task level avoids the accumulation of errors that may occur with the traditional RAID full post-verification. After the verification passes, the system updates the dynamic stripe mapping table, updating the original location information of the stripes on the failed physical disk to the new physical address of the target disk, and marking the micro-stripe state as "available", thereby shortening the system degradation time. At the same time, the successful feedback updates the reconstruction progress and dynamically adjusts the I / O bandwidth quota based on the remaining reconstruction amount, optimizing the disk selection for subsequent sub-tasks. Finally, the system releases resources such as pre-read cache and write-merge buffer to maintain pipeline operation.

[0089] When a sub-task fails, the system triggers an automatic retry mechanism, reinserting the sub-task into the shared task queue at the tail and possibly adjusting the target disk to reduce invalid I / O. If the number of retries exceeds a threshold, the system triggers an alarm and freezes the relevant stripes. Subsequently, based on the continuous failure feedback, the load is re-evaluated, the load weight of high-load disks is reduced, and resource quotas are dynamically adjusted to prioritize business I / O. If the failure rate within the same task group exceeds a threshold, the task group is re-divided, and the source / target disk combination is adjusted to reduce the impact of a single failure. In terms of fault isolation and logging, the system freezes the logical addresses of the associated stripes and records failure logs for subsequent analysis.

[0090] In summary, whether the sub-task is successful or fails, the above-mentioned closed-loop optimization process is triggered, including load model updating, shared task queue scheduling optimization, and business performance guarantee, to achieve an execution-verification-optimization closed loop, ensuring data consistency while dynamically optimizing reconstruction efficiency and business performance. It has higher intelligence and flexibility.

[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0092] Figure 2 A schematic diagram of an independent disk redundant array reconstruction device provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the device includes: Figure 2 A judgment module 10 is used to monitor the performance indicators of each physical disk and determine whether there is a failed physical disk in each physical disk according to the performance indicators; if so, a retrieval module 11 is triggered; ​The acquisition module 11 is configured to acquire service input / output performance indicators, wherein the service input / output performance indicators at least include total throughput, average delay and quantile delay of service input / output; The selection module 12 is configured to select a target reconstruction mode in each reconstruction mode according to the size relationship between the total throughput, average delay and quantile delay of service input / output and corresponding threshold values, respectively, wherein the resource occupancy of the service is in a negative correlation with the resource occupancy in the target reconstruction mode; the reconstruction mode is used to determine input / output bandwidth occupancy and processor occupancy in a reconstruction process; the reconstruction mode includes a first reconstruction mode, a second reconstruction mode and a third reconstruction mode; the reconstruction maximum input / output bandwidth in the first reconstruction mode is smaller than the reconstruction maximum input / output bandwidth in the third reconstruction mode, and the reconstruction maximum input / output bandwidth in the third reconstruction mode is smaller than the reconstruction maximum input / output bandwidth in the second reconstruction mode; the processor priority weight in the first reconstruction mode is smaller than the processor priority weight in the third reconstruction mode, and the processor priority weight in the third reconstruction mode is smaller than the processor priority weight in the second reconstruction mode; The generation module 13 is configured to generate a reconstruction resource allocation decision according to the target reconstruction mode and the real-time collected reconstruction performance indicators; The execution module 14 is configured to execute reconstruction of each stripe in the faulty physical disk based on the reconstruction resource allocation decision.

[0093] In some embodiments, further comprising: The division module is configured to divide the logical volume space into a plurality of continuous stripes according to data access heat information; The disk real-time utilization rate acquisition module is configured to acquire disk real-time utilization rates of each physical disk; The disk weight determination module is configured to determine disk weights of each physical disk according to the disk real-time utilization rates and standard extensible hash controlled replication weights; The storage module is configured to store data blocks and check blocks in each stripe in different physical disks according to the disk weights; The position information acquisition module is configured to acquire position information of each stripe, wherein the position information at least includes a stripe identifier, a physical disk identifier of a physical disk where the stripe is located, a logical block range and a state marker; the logical block range is a starting logical block address and size of the stripe on the corresponding physical disk; The mapping table generation module is configured to generate a dynamic stripe mapping table according to the position information of each stripe.

[0094] In some embodiments, the configuration process of the reconstruction mode comprises: setting the reconstruction maximum input / output bandwidth in the first reconstruction mode to be not greater than a first percentage of the total bandwidth, and the processor priority weight to be not greater than a second percentage of the total processor occupancy; setting the reconstruction maximum input / output bandwidth in the second reconstruction mode to be not less than a third percentage of the total bandwidth, and the processor priority weight to be not less than a fourth percentage of the total processor occupancy; setting the reconstruction maximum input / output bandwidth in the third reconstruction mode to be within a first range, and the processor priority weight to be within a second range; wherein the first percentage is less than the third percentage, and the second percentage is less than the fourth percentage; the first range is determined based on the first percentage of the total bandwidth and the third percentage of the total bandwidth, and the second range is determined based on the second percentage of the total processor occupancy and the fourth percentage of the total processor occupancy.

[0095] In some embodiments, the selection module 12 comprises: a first selection submodule configured to select the first reconstruction mode as the target reconstruction mode when the total throughput of the service input / output is not less than a first daily throughput threshold, or the quantile delay is greater than a quantile delay threshold; a second selection submodule configured to select the second reconstruction mode as the target reconstruction mode when the total throughput of the service input / output is not greater than a second daily throughput threshold, and the average delay is less than a first average delay threshold; a third selection submodule configured to select the third reconstruction mode as the target reconstruction mode when the average delay is less than a second average delay threshold, and the disk input / output utilization of each physical disk is less than a disk input / output utilization threshold; wherein the first daily throughput threshold is greater than the second daily throughput threshold, and the first average delay threshold is less than the second average delay threshold.

[0096] In some embodiments, the system further comprises: an updating submodule configured to update the service input / output performance indicators according to a preset period; a re-selection submodule configured to re-select a new target reconstruction mode for the reconstruction mode according to the updated service input / output performance indicators.

[0097] In some embodiments, the system further comprises: a fine-tuning submodule configured to, when the third reconstruction mode is selected as the target reconstruction mode, determine whether the average delay is less than the second average delay threshold, and the disk input / output utilization of each physical disk is less than the disk input / output utilization threshold after a preset time; if yes, increase the reconstruction maximum input / output bandwidth; and if no, decrease the reconstruction maximum input / output bandwidth.

[0098] In some embodiments, the generation module 13 comprises: The acquisition sub-module is configured to acquire an initial reconstruction maximum input / output bandwidth and an initial processor priority weight specified by a target reconstruction mode. The fine adjustment amount determination sub-module is configured to determine a fine adjustment amount of the reconstruction maximum input / output bandwidth according to a real-time collected reconstruction performance index. The target reconstruction maximum input / output bandwidth calculation sub-module is configured to calculate the target reconstruction maximum input / output bandwidth according to the initial reconstruction maximum input / output bandwidth and the fine adjustment amount. The target processor priority weight calculation sub-module is configured to calculate the target processor priority weight according to the target reconstruction maximum input / output bandwidth and the initial processor priority weight.

[0099] In some embodiments, the execution module 14 includes: The to-be-reconstructed stripe list generation module is configured to determine position information of each stripe in the failed physical disk according to the dynamic stripe mapping table, and generate a corresponding to-be-reconstructed stripe list. The source physical disk determination module is configured to determine, based on the to-be-reconstructed stripe list, a source physical disk that has a data association with the failed physical disk. The target disk determination module is configured to determine, based on each performance index, a target disk in each physical disk. The task group division module is configured to divide the to-be-reconstructed stripe list into a plurality of task groups according to load information of the source physical disk and the target disk; wherein, each task group contains a preset number of to-be-reconstructed stripes. The subtask division module is configured to divide each task group into a plurality of subtasks, and put each subtask into a shared task queue. The reconstruction data block calculation module is configured to acquire and execute each subtask in the shared task queue in parallel through a plurality of processing units, calculate a corresponding reconstruction data block through a redundant array of independent disks algorithm, and store the reconstruction data block to a buffer. The writing module is configured to sequentially write each reconstruction data block in the buffer to the target disk when data in the buffer reaches a corresponding storage threshold.

[0100] In some embodiments, the source physical disk determination module includes: The query module is configured to query, based on the dynamic stripe mapping table, a non-failed physical disk corresponding to the failed physical disk; wherein, a data block in the non-failed physical disk and a data block in the failed physical disk are associated data. The processing module is configured to remove an abnormal physical disk in each non-failed physical disk to determine the source physical disk.

[0101] The target disk determination module includes: The third determination sub-module is configured to determine a non-failed physical disk in each physical disk except the failed physical disk. a fourth determining sub-module, configured to determine the disk input / output utilization and the remaining storage space of each non-faulty physical disk; a fifth determining sub-module, configured to determine a non-faulty physical disk as a target disk if the corresponding disk input / output utilization is less than the disk input / output utilization threshold, the remaining storage space is greater than the remaining storage space threshold, and the non-faulty physical disk is not in the same load domain as the source physical disk.

[0102] The description of the features in the embodiments of the Redundant Array of Independent Disks reconstruction device can refer to the related description of the embodiments of the Redundant Array of Independent Disks reconstruction method, which will not be repeated here.

[0103] The embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-mentioned embodiments of the Redundant Array of Independent Disks reconstruction method.

[0104] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned embodiments of the Redundant Array of Independent Disks reconstruction method when running.

[0105] In an exemplary embodiment, the above-mentioned computer readable storage medium can include but is not limited to: a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0106] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the Redundant Array of Independent Disks reconstruction method.

[0107] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the Redundant Array of Independent Disks reconstruction method.

[0108] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of the claimed invention and are not intended to limit the scope of the claimed invention to these embodiments. Therefore, embodiments described herein are not meant to be limiting, but merely representative. Further, the routines executed to implement the embodiments of the invention, individually or collectively, need not be limited to any specific combination of hardware and software. Various embodiments can also be implemented using more conventional components, such as microprocessors, digital signal processors, microcomputers or microcontrollers, or using one or more analog or digital hardware circuits, or using a combination of hardware, software, and / or firmware components.

[0109] The independent disk redundant array reconstruction method and electronic device provided by the present application are described in detail above. The principles and implementation manners of the present application are described by applying specific examples in the present article. The above description of the embodiments is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A method for reconstructing an independent disk redundant array, characterized in that, include: Monitor the performance metrics of each physical disk, and determine whether there are any faulty physical disks among them based on the performance metrics. If so, then obtain the service input / output performance metrics; wherein, the service input / output performance metrics include at least the total throughput, average latency, and quantile latency of the service input / output; Based on the relationship between the total throughput of the service input / output, the average latency, and the quantile latency and their corresponding thresholds, a target reconstruction mode is selected from each reconstruction mode. The resource consumption of the service is negatively correlated with the resource consumption under the target reconstruction mode. The reconstruction mode is used to determine the input / output bandwidth consumption and processor consumption during the reconstruction process. The reconstruction modes include a first reconstruction mode, a second reconstruction mode, and a third reconstruction mode. The maximum reconstruction input / output bandwidth under the first reconstruction mode is less than that under the third reconstruction mode, and the maximum reconstruction input / output bandwidth under the third reconstruction mode is less than that under the second reconstruction mode. The processor priority weight under the first reconstruction mode is less than that under the third reconstruction mode, and the processor priority weight under the third reconstruction mode is less than that under the second reconstruction mode. Based on the target reconstruction mode and the real-time collected reconstruction performance indicators, a reconstruction resource allocation decision is generated; Based on the reconstructed resource allocation decision, the reconstructing of each stripe in the faulty physical disk is performed.

2. The independent disk redundancy array reconfiguration method according to claim 1, characterized in that, Before monitoring the performance metrics of each physical disk, the following is also included: The logical volume space is divided into multiple consecutive stripes based on data access frequency information; Collect the real-time disk utilization rate of each physical disk; The disk weight of each physical disk is determined based on the real-time utilization of each disk and the standard scalable hash-controlled replication weight. According to the disk weights, the data blocks and check blocks in each stripe are stored on different physical disks respectively; Obtain the location information of each stripe; wherein, the location information includes at least a stripe identifier, a physical disk identifier of the physical disk where the stripe is located, a logical block range, and a status flag; the logical block range is the starting logical block address and size of the stripe on the corresponding physical disk; A dynamic strip mapping table is generated based on the position information of each strip.

3. The independent disk redundancy array reconfiguration method according to claim 1, characterized in that, The configuration process for the refactoring mode includes: In the first reconstruction mode, the maximum input / output bandwidth of reconstruction is set to be no greater than the first percentage of the total bandwidth, and the processor priority weight is set to be no greater than the second percentage of the total processor utilization. In the second refactoring mode, the maximum input / output bandwidth of the refactoring is set to be no less than the third percent of the total bandwidth, and the processor priority weight is set to be no less than the fourth percent of the total processor utilization. In the third refactoring mode, the maximum input / output bandwidth for refactoring is set to the first range, and the processor priority weight is set to the second range. The first percentage is less than the third percentage, and the second percentage is less than the fourth percentage; the first range is determined based on the first percentage and the third percentage of the total bandwidth, and the second range is determined based on the second percentage and the fourth percentage of the total processor utilization.

4. The independent disk redundancy array reconfiguration method according to claim 3, characterized in that, Based on the relationships between the total throughput of the service input / output, the average latency, and the quantile latency and their corresponding thresholds, a target reconstruction mode is selected from each reconstruction mode, including: When the total throughput of the business input / output is not less than the first daily throughput threshold, or the quantile delay is greater than the quantile delay threshold, the first reconstruction mode is selected as the target reconstruction mode. When the total throughput of the business input / output is not greater than the second daily throughput threshold and the average latency is less than the first average latency threshold, the second reconstruction mode is selected as the target reconstruction mode. When the average latency is less than the second average latency threshold and the disk input / output utilization of each physical disk is less than the disk input / output utilization threshold, the third reconstruction mode is selected as the target reconstruction mode. Among them, the first daily throughput threshold is greater than the second daily throughput threshold; the first average latency threshold is less than the second average latency threshold.

5. The independent disk redundancy array reconfiguration method according to claim 4, characterized in that, After selecting the target reconstruction mode, the following is also included: The service input / output performance metrics are updated according to a preset period. Based on the updated business input / output performance metrics, a new target reconstruction mode is selected from each of the reconstruction modes.

6. The independent disk redundancy array reconfiguration method according to claim 4, characterized in that, After selecting the target reconstruction mode, the following is also included: When the third reconstruction mode is selected as the target reconstruction mode, it is determined whether the average latency is less than the second average latency threshold after a preset time, and whether the disk input / output utilization of each physical disk is less than the disk input / output utilization threshold. If so, then increase the maximum input / output bandwidth of the reconstruction; If not, then reduce the maximum input / output bandwidth of the reconstruction.

7. The independent disk redundancy array reconfiguration method according to claim 3, characterized in that, Based on the target reconstruction mode and the real-time collected reconstruction performance indicators, a reconstruction resource allocation decision is generated, including: Obtain the initial maximum input / output bandwidth and the initial processor priority weight specified by the target reconstruction mode; The fine-tuning amount of the maximum input / output bandwidth of the reconstruction is determined based on the reconstruction performance indicators collected in real time; Calculate the target reconstructed maximum input / output bandwidth based on the initial reconstructed maximum input / output bandwidth and the fine-tuning amount; The target processor priority weight is calculated based on the target reconstructed maximum input / output bandwidth and the initial processor priority weight.

8. The independent disk redundancy array reconstruction method according to claim 2, characterized in that, Based on the reconstructed resource allocation decision, the reconstructing of each stripe in the faulty physical disk is performed, including: The location information of each stripe in the faulty physical disk is determined according to the dynamic stripe mapping table, and a corresponding list of stripes to be reconstructed is generated. Based on the list of stripes to be reconstructed, the source physical disks that have data associations with the faulty physical disks are identified; Based on the aforementioned performance metrics, the target disk is determined from among the aforementioned physical disks; Based on the load information of the source physical disk and the target disk, the list of stripes to be reconstructed is divided into multiple task groups; wherein, each task group contains a preset number of stripes to be reconstructed. Each of the task groups is decomposed into multiple subtasks, and each subtask is placed into a shared task queue; The subtasks in the shared task queue are acquired and executed in parallel by multiple processing units, and the corresponding reconstructed data blocks are calculated by the independent disk redundancy array algorithm and stored in the buffer. When the data in the buffer reaches the corresponding storage threshold, the reconstructed data blocks in the buffer are sequentially written to the target disk.

9. The independent disk redundancy array reconfiguration method according to claim 8, characterized in that, Based on the list of stripes to be reconstructed, source physical disks that have data associations with the faulty physical disk are identified, including: Based on the dynamic stripe mapping table, query the non-faulty physical disk corresponding to the faulty physical disk; wherein, the data blocks in the non-faulty physical disk and the data blocks in the faulty physical disk are related data; Remove the abnormal physical disks from each of the non-faulty physical disks to determine the source physical disk; Correspondingly, based on the aforementioned performance metrics, the target disk is determined among the aforementioned physical disks, including: Identify the non-faulty physical disks among the physical disks, excluding the faulty physical disk; Determine the disk I / O utilization and remaining storage space of each of the non-faulty physical disks; The non-faulty physical disk that corresponds to a disk input / output utilization rate less than the disk input / output utilization threshold, a remaining storage space greater than the remaining storage space threshold, and is not in the same load domain as the source physical disk is identified as the target disk.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the independent disk redundant array reconfiguration method as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Disk reconfiguration method and disk reconfiguration device

    CN103049400A

  • Controlling data storage in an array of storage devices

    CN104123100A

  • Disk reconfiguration method based on RAID and related apparatus

    CN104536698A

  • Data reconstruction method, device, equipment, medium and product

    CN119883713A

  • Fault reconstruction method and device for mechanical hard disk, equipment and storage medium

    CN120233949A