A multi-level data recovery system and method based on redundant storage

Through a multi-level storage system and dynamic recovery strategy, the congestion problems caused by slow core data recovery and node failures during data recovery are solved, achieving high efficiency and reliability of data recovery.

CN120578539BActive Publication Date: 2025-10-10TOPDISK ENTERPRISE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511081305.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-10
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing technologies are unable to divert data according to its importance, resulting in slow recovery of core data. In addition, when a node fails during the data recovery process, it is easy to cause congestion in the entire process and a low recovery speed.

Method used

By detecting data usage parameters and node CPU occupancy, a multi-level storage system is adopted, including hot data, warm data, and cold data storage units. In combination with erasure codes and blockchain technology, the data storage location and recovery mode are dynamically adjusted, the node failure rate is predicted, and the data recovery strategy is optimized.

Benefits of technology

It achieves rapid recovery of core data and efficient processing of non-critical data, reduces delays and congestion in the data recovery process, and improves data recovery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578539B_ABST
    Figure CN120578539B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field, in particular to a multi-level data recovery system and method based on redundant storage, the system comprises: a detection module used for detecting the use parameters of data and the CPU occupancy rate of each node; a storage module; an analysis module used for judging the storage position of data according to the data heat value, in response to data recovery, the analysis module predicts the failure rate of each node according to the historical failure information, and adjusts the data recovery proportion of each node according to the failure rate of each node and the CPU occupancy rate. By detecting the use parameters of data and the CPU occupancy rate of each node, the analysis module can accurately judge the data heat and match the corresponding redundant storage and recovery mode, reduce the storage and recovery cost of non-critical data, and thus increase the data recovery efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a multi-level data recovery system and method based on redundant storage. BACKGROUND

[0002] Due to the high-speed data processing ability and large-capacity storage ability of the computer, the computer is widely used in various fields of production and life. Once the server is down, many user terminals will not be able to communicate normally, and once the information processing system of the bank card organization is down, a large number of cardholders and merchants will not be able to carry out bank card business.

[0003] The application with the application publication number CN1801107A discloses a data recovery method, which comprises: dividing the file to be backed up into a plurality of data blocks, and saving the data blocks of the file on at least one backup computer at the same time; when the data of the data blocks of the file changes, the main computer sends the data blocks with changed data to the backup computer for backup; determining the damaged data blocks, and obtaining the backup data blocks from the backup computer for data recovery. It can be seen that the data recovery method cannot shunt the data according to the importance in the process of recovering the data, cannot guarantee the rapid recovery of the core data, and if a node fails and the amount of data to be recovered by the node is large during the data recovery process, the whole data recovery process will be congested, thereby causing the low speed of data recovery. SUMMARY

[0004] The purpose of the present application is to provide a multi-level data recovery system and method based on redundant storage, so as to solve the problem that the prior art cannot shunt the data according to the importance, cannot guarantee the rapid recovery of the core data, and if a node fails and the amount of data to be recovered by the node is large during the data recovery process, the whole data recovery process will be congested, thereby causing the low speed of data recovery.

[0005] The present application provides a multi-level data recovery system based on redundant storage, which comprises:

[0006] The detection module is used to detect the use parameters of the data, the CPU occupancy rate and the failure rate of each node in the multi-level data recovery, and the use parameters comprise: access frequency, recent access time and read-write ratio of the data;

[0007] The storage module is connected with the detection module, and comprises: a hot data storage unit, a warm data storage unit, a cold data storage unit and a running storage unit, wherein: the running storage unit is used to store the historical failure information of each node;

[0008] An analysis module is connected to the detection module and the storage module respectively, and is used to calculate the data heat value according to the usage parameters and determine the storage location of the data. In response to data recovery, the analysis module predicts the current failure rate of each node according to the historical fault information, and adjusts the data recovery ratio of each node according to the current failure rate of each node and the CPU occupancy rate.

[0009] As an optimal technical solution for a multi-level data recovery system based on redundant storage, the analysis module determines the storage location of the data based on the data heat value and calculates the data heat value based on usage parameters, wherein: the data heat value is positively correlated with the access frequency, the most recent access time and the read-write ratio, and the write operation weight coefficient is higher than the read operation weight coefficient.

[0010] As a preferred technical solution for a multi-level data recovery system based on redundant storage, the analysis module determines the storage location of data based on the data heat value, including:

[0011] In response to the data heat value being greater than a first threshold, determining that the data storage location is the hot data storage unit;

[0012] In response to the data heat value being less than or equal to the first threshold and greater than or equal to a second threshold, determining that the storage location of the data is the warm data storage unit;

[0013] In response to the data heat value being less than the second threshold, it is determined that the storage location of the data is a cold data storage unit.

[0014] As a preferred technical solution for a multi-level data recovery system based on redundant storage, the analysis module predicts the current failure rate of each node based on the historical fault information, and uses the current moment as the starting point of the window to divide into multiple short-term analysis windows and long-term analysis windows. Each analysis window has a different weight, and the weight value decays over time. The number of failures of the node in each time window is counted. For a single node, the product of the number of failures of the single node in each time window and the weight value corresponding to the window is accumulated to obtain the comprehensive failure amount of the single node. The ratio of the comprehensive failure amount of the single node to the comprehensive failure amount of all nodes is recorded as the failure rate.

[0015] The fault information includes: fault occurrence time, duration and fault type.

[0016] As an optimal technical solution for a multi-level data recovery system based on redundant storage, the analysis module adjusts the data recovery ratio of each node according to the current failure rate of each node and the CPU occupancy rate, and records the product of the current failure rate of each node and the CPU occupancy rate as the stability coefficient. The data recovery ratio of each node is inversely proportional to the ratio of the stability coefficient.

[0017] As a preferred technical solution for a multi-level data recovery system based on redundant storage, the erasure code of the warm data storage unit adopts a local repair code, including:

[0018] Global checksum block: used for overall data recovery;

[0019] Local check block: used to reconstruct data through local nodes;

[0020] During the process of using the local repair code for the erasure code of the warm data storage unit, the CPU occupancy rate of each node in the storage cluster is collected in real time, and the number of subtasks to be rebuilt of each node is corrected according to the CPU occupancy rate;

[0021] In response to the CPU usage of a node being higher than a load threshold, the number of subtasks to be rebuilt allocated to the node is reduced, and the reduction ratio is positively correlated with the extent to which the CPU usage exceeds the threshold.

[0022] As an optimal technical solution for a multi-level data recovery system based on redundant storage, the storage module also includes a blockchain storage unit for storing the hash values ​​corresponding to the hot data storage unit, the warm data storage unit, and the cold data storage unit. Each time data is written, the data hash value is recorded to the blockchain through a distributed node consensus mechanism. After data recovery, the consistency of the hash values ​​is verified. In response to inconsistent hash values, the data is determined to be incomplete. If the verification fails, multi-level rollback recovery is triggered.

[0023] As an optimal technical solution for a multi-level data recovery system based on redundant storage, the detection module periodically performs consistency checks on the three copies of the hot data storage unit. In response to the inconsistency between the hash value of any copy and the other two, one of the remaining two valid copies is selected and a new copy is generated on other high-speed storage media.

[0024] The present invention also provides a multi-level data recovery method based on redundant storage, comprising:

[0025] Obtain the usage parameters of the data and calculate the data heat value based on the usage parameters;

[0026] Periodically determining the redundant storage location of the data based on the data heat value, and storing the data in the corresponding location after the determination is completed;

[0027] In response to a storage node failure requiring data recovery, after selecting a corresponding recovery mode based on the level of the data to be recovered, the failure rate of each node is predicted based on the historical failure information, and the data recovery ratio of each node is adjusted based on the current failure rate and CPU usage of each node;

[0028] The data is restored according to the data recovery ratio.

[0029] As an optimal technical solution for the multi-level data recovery method based on redundant storage, the method selects the corresponding recovery mode according to the level at which the data needs to be recovered, including: using a hot data storage unit for real-time recovery, using erasure codes to perform parallel reconstruction of data, and selecting corresponding parts from the storage medium to restore data.

[0030] Compared with the existing technology, the beneficial effect of the present invention is that by detecting the data usage parameters (access frequency, last access time, read-write ratio) and node CPU occupancy, the analysis module can accurately judge the data popularity and match the corresponding redundant storage and recovery mode, thereby ensuring the rapid recovery of core data while reducing the storage and recovery costs of non-critical data, avoiding the need to restore all data before using it, thereby ensuring that core data can be quickly recovered and non-critical data can be efficiently processed, thereby increasing data recovery efficiency.

[0031] Furthermore, the present invention divides the time windows into multiple short-term / long-term time windows and introduces a time decay weight mechanism, so that recent fault data has a greater impact on the failure rate calculation, effectively capturing the real-time changes in node status, thereby making more accurate predictions on the occurrence of faults, reducing the data retransmission rate when a node fails, and further increasing data recovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a structural block diagram of a multi-level data recovery system based on redundant storage according to an embodiment of the present invention;

[0033] Figure 2 This is a flowchart of the steps of a multi-level data recovery method based on redundant storage according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0035] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0036] See also Figure 1 As shown in FIG, it is a structural block diagram of a multi-level data recovery system based on redundant storage according to an embodiment of the present invention, including:

[0037] The detection module is used to detect data usage parameters and the CPU usage of each node in multi-level data recovery. The usage parameters include: access frequency, last access time, and data read / write ratio;

[0038] The storage module includes: a hot data storage unit, a warm data storage unit, a cold data storage unit, and an operational storage unit. The operational storage unit is used to store historical fault information of each node. The hot data storage unit uses three copies of redundant storage in a high-speed storage medium. The warm data storage unit uses erasure coding for redundant storage in a distributed storage cluster. The cold data layer uses a single copy of compressed storage in the storage medium.

[0039] The analysis module is connected to the detection module and the storage module respectively, and is used to calculate the data heat value according to the usage parameters and determine the storage location of the data. In response to data recovery, the analysis module predicts the current failure rate of each node based on historical fault information, and adjusts the data recovery ratio of each node according to the current failure rate of each node and the CPU occupancy rate; the analysis module is also used to select the corresponding data recovery mode according to the location of the data to be recovered in the storage module, including: using a hot data storage unit for real-time recovery, using erasure codes for parallel reconstruction of data, and selecting corresponding parts from the storage medium to recover data.

[0040] Furthermore, the detection module collects data access frequency (for example, counting the number of data accesses within the past 24 hours), the time of last access (calculating the difference in hours between the current time and the last access time), and the data read / write ratio in real time through system logs or monitoring tools. For multi-tiered data recovery, CPU usage for each node can be collected every five minutes using monitoring commands provided by the operating system.

[0041] Specifically, the data heat value is positively correlated with the access frequency, the most recent access time, and the read-write ratio, where the write operation has a higher weight than the read operation.

[0042] Furthermore, the data heat value can be obtained by establishing a model or calculating a mathematical formula. An embodiment of the present invention provides a method for calculating the data heat value, including:

[0043] Get the access frequency per unit time (times / 24 hours);

[0044] When calculating the data heat value, the weight of the read operation is set to 1, the weight of the write operation is set to 2, and the read-write ratio is set. ;

[0045] The weight of the most recent access time is calculated as follows: if the most recent access time is within 1 hour, the weight is 1; if it is between 1 and 6 hours, the weight is 0.8; if it is between 6 and 12 hours, the weight is 0.6; if it is between 12 and 24 hours, the weight is 0.4; if it is more than 24 hours, the weight is 0.2;

[0046] The calculation formula for setting the data heat value is: Heat value = access frequency × 0.3 + recent access time weight × 0.4 + read-write ratio × 0.3. Substitute the access frequency, recent access time weight, and read-write ratio into the above formula to obtain the data heat value.

[0047] Specifically, on the one hand, the present invention sets the weight of write operations to be higher than that of read operations, so that the data heat value is more in line with actual business needs. On the other hand, by clarifying the weighted calculation logic of read and write operations and the units and calculation steps of each parameter in the formula, technical personnel in this field can clearly realize the quantitative evaluation of data heat value, ensure the accuracy of hot data identification, and more accurately identify hot data that really needs to be updated frequently, so as to ensure that hot data is stored in high-speed storage media first. When recovery is needed, since it is stored in high-speed media and is accurately identified as hot data, recovery delays caused by incorrect classification are avoided, thereby further increasing data recovery efficiency.

[0048] Furthermore, the storage location of the data is determined based on the data heat value, including:

[0049] In response to the data heat value being greater than a first threshold, determining the data storage location as a hot data storage unit;

[0050] In response to the data heat value being less than or equal to a first threshold and greater than or equal to a second threshold, determining that the storage location of the data is a warm data storage unit;

[0051] In response to the data heat value being less than a second threshold, it is determined that the storage location of the data is a cold data storage unit.

[0052] Specifically, the minimum value of the heat value of data recovered in real time in the past 30 days is selected as the first threshold, which is usually in the range of 80-100. Preferably, the value of the first threshold is 90; to ensure that the temperature layer data is "medium frequency access, non-core but requires rapid reconstruction", the value of the second threshold is 30% to 50% of the first threshold. Preferably, 40% of the first threshold is selected, that is, the value is 36.

[0053] Specifically, the present invention establishes a three-tier storage system: hot, warm, and cold, by setting first and second thresholds. Data can be dynamically relocated based on real-time data popularity. When data popularity changes, it can be promptly migrated to the appropriate storage unit, ensuring the most appropriate recovery mode is used during recovery. This ensures that hot data maintains high-speed recovery, warm data can take advantage of parallel reconstruction, and cold data can be restored on demand, avoiding slow recovery times caused by inappropriate storage locations and further increasing data recovery efficiency.

[0054] Specifically, the failure rate of each node is predicted based on historical failure information. Multiple short-term and long-term analysis windows are divided. Each analysis window has a different weight, and the weight value decays over time. The number of node failures in each time window is counted. For each node, the product of the number of failures of the node in each time window and the weight value corresponding to the window is accumulated to obtain the comprehensive failure rate of the node. The ratio of the comprehensive failure rate of the node to the comprehensive failure rate of all nodes is recorded as the failure rate.

[0055] Fault information includes: fault occurrence time, duration, and fault type (hardware / software / network).

[0056] In implementation, multiple overlapping time analysis windows are set up. The short-term time window is used to capture recent faults. The duration of the short-term time window is selected from 1 to 12 hours. Preferably, the duration of the short-term time window is 1 hour, 6 hours, and 12 hours; the long-term time window is used to analyze the overall fault trend. The duration of the long-term time window must be greater than 24 hours to better analyze the fault trend. Preferably, the duration of the long-term time window is 24 hours, 48 ​​hours, and 240 hours.

[0057] An embodiment of the present invention provides a method for calculating a failure rate, comprising:

[0058] Set the analysis window weight: ,in: is the attenuation coefficient, γ∈[0.1,0.3]. Preferably, in this embodiment, the value is 0.1. t is the duration of the time window, in hours. That is, if there is a short-term time window, the time window duration is 1 hour, then the weight coefficient w k =e -0.1×1 ≈0.905.

[0059] For node i, traverse all analysis windows and count the number of failures N of node i in each analysis window k. i,k .

[0060] Calculate the weighted number of failures in window k: N i,k ×w k .

[0061] The weighted failure counts of all windows are accumulated to obtain the comprehensive failure count of node i.

[0062] Repeat the above steps to calculate the comprehensive fault quantity corresponding to each node.

[0063] The failure rate of node i = the comprehensive failure volume of node i / the sum of the comprehensive failure volumes corresponding to each node.

[0064] Furthermore, by dividing the time windows into multiple short-term / long-term time windows and introducing a time decay weight mechanism, the recent fault data has a greater impact on the failure rate calculation, effectively capturing the real-time changes in node status. Compared with the static historical average method, the fault prediction accuracy is significantly improved, thereby making more accurate predictions on the occurrence of faults and reducing the data retransmission rate when a node fails, thereby further increasing data recovery efficiency.

[0065] Furthermore, the data recovery ratio for each node is adjusted based on the current failure rate and CPU utilization of each node. The product of the current failure rate and CPU utilization is calculated and recorded as the stability coefficient. The data recovery ratio of each node is inversely proportional to the ratio of the stability coefficient. For example, if node A has a failure rate of 0.1 and a CPU utilization of 80%, and node B has a failure rate of 0.05 and a CPU utilization of 40%, the stability coefficient ratio is 4:1, and the data recovery ratio of nodes A and B is 1:4.

[0066] Furthermore, the data recovery ratio is the ratio of the amount of data recovered by the node to the total amount of data that needs to be recovered during data recovery.

[0067] Specifically, the data that needs to be recovered is allocated by calculating the stability coefficient and determining the data ratio based on the stability coefficient. During the allocation process, a small number of data recovery tasks are allocated to nodes with high CPU occupancy and high failure rate, thereby reducing the data retransmission rate when a node fails. For nodes with low failure rate and low CPU occupancy, more data recovery tasks are allocated, thereby further increasing the data recovery speed.

[0068] Specifically, the erasure code of the warm data storage unit adopts a local repair code, including:

[0069] Global checksum block: used for overall data recovery;

[0070] Local check block: used to rebuild data through local nodes to reduce network bandwidth usage during recovery.

[0071] Among them: the process of using local repair code as erasure code to recover data belongs to the existing technology and will not be described in detail here.

[0072] Furthermore, warm data storage utilizes a local repair code, consisting of global and local parity blocks. When some nodes fail, data can be reconstructed from a small number of nearby nodes using the local parity blocks, reducing network bandwidth usage and data transmission during recovery. In a distributed cluster environment, this reduces network congestion, speeds up data reconstruction, and further improves data recovery efficiency.

[0073] Specifically, when the erasure code of the warm data storage unit adopts the local repair code, the CPU occupancy rate of each node in the storage cluster is collected in real time, and the number of subtasks to be rebuilt of each node is corrected according to the CPU occupancy rate.

[0074] In response to the CPU usage of a node being higher than a load threshold, the number of subtasks to be rebuilt allocated to the node is reduced, and the reduction ratio is positively correlated with the extent to which the CPU usage exceeds the threshold.

[0075] Furthermore, the load threshold is selected based on the performance and stability of the node, which can not only ensure that the node has sufficient computing power to process tasks, but also avoid system failures caused by excessive load. Preferably, the preset load threshold is 70%. The detection module collects the CPU occupancy of each node every 10 seconds and obtains it through the monitoring interface of the operating system in units of "%". An embodiment of the present invention provides a method for reducing the number of subtasks to be rebuilt, setting the reduction in the number of currently allocated subtasks = 2 (CPU occupancy - load threshold) × the number of subtasks to be rebuilt.

[0076] Furthermore, the storage module also includes a blockchain storage unit for storing hash values ​​corresponding to the hot data storage unit, the warm data storage unit, and the cold data storage unit. Each time data is written, the data hash value is recorded to the blockchain through a distributed node consensus mechanism. After the data is restored, the consistency of the hash value is verified. In response to inconsistent hash values, the data is determined to be incomplete. If the verification fails, a multi-level rollback recovery is triggered; in response to consistent hash values, the data is determined to be complete.

[0077] Specifically, by storing hash values ​​of data at each layer on the blockchain and performing hash value verification after data recovery, data integrity can be quickly determined. If data incompleteness is found, a multi-level rollback recovery can be triggered immediately, retrieving data from a more reliable storage layer for recovery. This prevents incorrect data from being written to the system, reducing repeated recovery processes caused by data errors, increasing data recovery speed, and ensuring the accuracy and reliability of recovered data.

[0078] Specifically, the detection module periodically performs consistency checks on the three copies of the hot data storage unit. In response to any copy hash value being inconsistent with the other two, one of the remaining two valid copies is selected and a new copy is generated on other high-speed storage media.

[0079] Specifically, regular consistency checks are performed on the three replicas of hot data to promptly detect replica corruption caused by storage media failures. Once a corrupted replica is discovered, a new replica is immediately generated from a valid replica, ensuring that hot data always maintains three copies of redundancy. During data recovery, since the hot data replica is always healthy, rapid recovery can be performed directly from the reliable replica, further improving data recovery efficiency.

[0080] See also Figure 2 As shown, it is a multi-level data recovery method based on redundant storage according to an embodiment of the present invention, including:

[0081] Step S1, obtaining usage parameters of data and calculating data heat value according to the usage parameters;

[0082] Step S2: Periodically determine the redundant storage location of the data based on the data heat value, and store the data in the corresponding location after the determination is completed;

[0083] Step S3: In response to a storage node failure requiring data recovery, after selecting a corresponding recovery mode based on the level of the data to be recovered, the failure rate of each node is predicted based on historical failure information, and the data recovery ratio of each node is adjusted based on the current failure rate of each node and the CPU usage;

[0084] Step S4: Recover the data according to the data recovery ratio.

[0085] Furthermore, in step S3, a corresponding recovery mode is selected according to the level of the data to be recovered, including: using a hot data storage unit for real-time recovery, using erasure codes to parallel rebuild the data, and selecting a corresponding portion of the data from the storage medium for recovery. Using erasure codes to parallel rebuild the data includes:

[0086] Split the data block to be recovered into multiple subtasks and number them (for example, each subtask processes 1MB of data);

[0087] The sub-tasks are distributed to different computing nodes or servers for simultaneous execution by a task scheduler, and the finite field mathematical operations required for erasure code are accelerated using the SIMD instruction set of the CPU or GPU;

[0088] The data at the end of the reconstruction is sorted by number to complete the reconstruction of the data.

[0089] Specifically, the selection of the computing nodes or servers is determined based on the physical distance between nodes and the real-time network delay (measured by Ping value), and the node closest in topology distance is preferentially selected to transmit the data shards; when there is no available shard in the local node, the node is automatically switched to a suboptimal node (such as other machine rooms in the same city).

[0090] Specifically, when using erasure code to perform parallel reconstruction of data, the data block to be recovered is split into multiple sub-tasks, which are distributed to different nodes for simultaneous execution by a task scheduler, and the finite field mathematical operations are accelerated using the SIMD instruction set of the CPU or GPU. This parallel processing and accelerated operation greatly improves the speed of data reconstruction, thereby further increasing the data recovery efficiency.

[0091] Obviously, the above embodiments of the present application are only examples for clarity, and are not intended to limit the embodiments of the present application. Based on the above description, those skilled in the art can make other different forms of changes or variations. Here, it is not necessary and impossible to exhaust all the embodiments. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the claims of the present application.

Claims

1. A multi-level data recovery system based on redundant storage, characterized in that: include: A detection module is used to detect data usage parameters, the CPU occupancy rate of each node in multi-level data recovery, and whether each node has failed. The usage parameters include: data access frequency, recent access time, and data read / write ratio; a storage module connected to the detection module, comprising: a hot data storage unit, a warm data storage unit, a cold data storage unit, and an operation storage unit, wherein: the operation storage unit is used to store historical fault information of each of the nodes; an analysis module, connected to the detection module and the storage module, respectively, for calculating a data heat value based on the usage parameters and determining a storage location of the data. In response to data recovery, the analysis module predicts a current failure rate of each node based on the historical failure information and adjusts a data recovery ratio of each node based on the current failure rate and CPU occupancy of each node; The analysis module predicts the current failure rate of each node based on the historical fault information, and uses the current time as the starting point of the window to divide into multiple short-term analysis windows and long-term analysis windows. Each analysis window has a different weight, and the weight value decays over time. The number of failures of the node in each time window is counted. For a single node, the product of the number of failures of the single node in each time window and the weight value corresponding to the window is accumulated to obtain the comprehensive failure amount of the single node. The ratio of the comprehensive failure amount of the single node to the sum of the comprehensive failure amounts of all nodes is recorded as the failure rate; The fault information includes: fault occurrence time, duration and fault type; The erasure code of the warm data storage unit adopts a local repair code, including: Global checksum block: used for overall data recovery; Local check block: used to reconstruct data through local nodes; During the process of using the local repair code for the erasure code of the warm data storage unit, the CPU occupancy rate of each node in the storage cluster is collected in real time, and the number of subtasks to be rebuilt of each node is corrected according to the CPU occupancy rate; In response to the CPU usage of a node being higher than a load threshold, the number of subtasks to be rebuilt allocated to the node is reduced, and the reduction ratio is positively correlated with the extent to which the CPU usage exceeds the threshold.

2. The multi-level data recovery system based on redundant storage according to claim 1, characterized in that: The analysis module determines the storage location of the data based on the usage parameters and calculates the data heat value based on the usage parameters, wherein: the data heat value is positively correlated with the access frequency, the most recent access time and the read-write ratio, and the write operation weight coefficient is higher than the read operation weight coefficient.

3. The multi-level data recovery system based on redundant storage according to claim 2, characterized in that: The analysis module determines the storage location of the data according to the data heat value, including: In response to the data heat value being greater than a first threshold, determining that the data storage location is the hot data storage unit; In response to the data heat value being less than or equal to the first threshold and greater than or equal to a second threshold, determining that the storage location of the data is the warm data storage unit; In response to the data heat value being less than the second threshold, it is determined that the storage location of the data is a cold data storage unit.

4. The multi-level data recovery system based on redundant storage according to claim 1, characterized in that: The analysis module adjusts the data recovery ratio of each node according to the current failure rate and CPU occupancy of each node, and records the product of the current failure rate of each node and the CPU occupancy as the stability coefficient. The data recovery ratio of each node is inversely proportional to the stability coefficient.

5. The multi-level data recovery system based on redundant storage according to claim 1, characterized in that: The storage module also includes a blockchain storage unit for storing hash values ​​corresponding to the hot data storage unit, the warm data storage unit, and the cold data storage unit. Each time data is written, the data hash value is recorded in the blockchain through a distributed node consensus mechanism. After data recovery, the consistency of the hash values ​​is verified. In response to inconsistent hash values, the data is determined to be incomplete. If the verification fails, a multi-level rollback recovery is triggered.

6. The multi-level data recovery system based on redundant storage according to claim 1, characterized in that: The detection module periodically performs consistency checks on the three copies of the hot data storage unit, and in response to a hash value of any copy being inconsistent with the other two, selects one of the remaining two valid copies and generates a new copy on other high-speed storage media.

7. A multi-level data recovery method based on redundant storage, applied to the multi-level data recovery system based on redundant storage according to any one of claims 1 to 6, characterized in that: include: Obtain the usage parameters of the data and calculate the data heat value based on the usage parameters; Periodically determining the redundant storage location of the data based on the data heat value, and storing the data in the corresponding location after the determination is completed; In response to a storage node failure requiring data recovery, after selecting a corresponding recovery mode based on the level of the data to be recovered, the failure rate of each node is predicted based on the historical failure information, and the data recovery ratio of each node is adjusted based on the current failure rate and CPU usage of each node; The data is restored according to the data recovery ratio.

8. The multi-level data recovery method based on redundant storage according to claim 7, characterized in that: The method of selecting a corresponding recovery mode according to the level at which the data needs to be restored includes: using a hot data storage unit for real-time recovery, using erasure codes to parallel rebuild data, and selecting a corresponding portion from a storage medium to restore data.

Citation Information

Patent Citations

  • Data recovery method

    CN1801107A

  • Heterogeneous cloud storage cluster fault automatic repair method, system, medium and terminal

    CN113535474A

  • Data recovery method, system and equipment based on distributed storage and medium

    CN115934420A