Data reconstruction method and device, electronic equipment and storage medium

By monitoring and dynamically calculating disk redundancy, adjusting the erasure code type, and adopting an incremental recovery mechanism, the problems of resource consumption and network bandwidth usage during data reconstruction are solved, achieving a balance between efficient data recovery and low cost.

CN120687035AActive Publication Date: 2025-09-23JINAN INSPUR DATA TECH CO LTD

Patent Information

Application Number
CN202510803415.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-23
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The existing technology cannot dynamically balance data recovery performance, resource consumption and network bandwidth resource usage during data reconstruction, especially in large file scenarios, which affects the efficiency and resource consumption of real-time data transmission tasks.

Method used

By monitoring the status data of the disk to be reconstructed, dynamically calculating its redundancy, adjusting the erasure code type according to different redundancy levels, and using the incremental recovery mechanism to transmit the minimum amount of data, an efficient balance is achieved in the data recovery process.

Benefits of technology

In the event of a disk failure or expansion, redundancy and erasure code types are dynamically adjusted to reduce network bandwidth usage, improve data recovery efficiency and speed, maintain a balance between high performance and low cost, and ensure front-end business stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687035A_ABST
    Figure CN120687035A_ABST
Patent Text Reader

Abstract

The invention discloses a data reconstruction method and device, electronic equipment and a storage medium, and relates to the technical field of data storage and recovery, and the data reconstruction method comprises the steps that under the condition that a to-be-reconstructed disk exists in a distributed storage system, the redundancy of the to-be-reconstructed disk is determined based on state data of the to-be-reconstructed disk; determining a target redundancy level corresponding to the redundancy of the to-be-reconstructed disk from a plurality of preset redundancy levels; different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rates and storage resource occupation proportions; and switching the erasure code type of the to-be-reconstructed disk to a target erasure code type corresponding to the target redundancy level, and according to the target erasure code type, transmitting the fault data block in the to-be-reconstructed disk to a specified disk for data reconstruction. The problem that data recovery performance, resource consumption and network bandwidth resource occupation cannot be dynamically balanced in the data reconstruction process in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data storage and recovery technology, and in particular to a data reconstruction method, device, electronic device and storage medium. Background Art

[0002] A distributed storage system provides distributed storage services to external cloud platforms. This distributed storage system can offer other distributed storage services, such as object storage, block storage, and file system storage. The most fundamental element of a distributed storage system is the disk, where all data is stored. When a disk fails or storage capacity expansion is required, the data on the failed disk or the disk to be expanded must be reconstructed. This involves restoring the data on other functioning hard drives.

[0003] Related data reconstruction techniques involve copying all the data from the failed disk or the disk to be expanded to other functioning disks. However, this method consumes a lot of network bandwidth, especially in large file scenarios (such as satellite remote sensing), which can interfere with real-time data downloads, resulting in low recovery efficiency and high resource consumption. Summary of the Invention

[0004] The present application provides a data reconstruction method, device, electronic device and storage medium to at least solve the problem in the related art that data recovery performance, resource consumption and network bandwidth resource occupancy cannot be dynamically balanced during data reconstruction.

[0005] The present application provides a data reconstruction method, comprising: in a case where a disk to be reconstructed exists in a distributed storage system, monitoring status data of the disk to be reconstructed, and determining the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed; determining a target redundancy level corresponding to the redundancy of the disk to be reconstructed from a plurality of preset redundancy levels; different redundancy levels among the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio; switching the erasure code type of the disk to be reconstructed to a target erasure code type corresponding to the target redundancy level, transferring faulty data blocks in the disk to be reconstructed to a designated disk according to the target erasure code type, and performing data reconstruction on the designated disk; the designated disk is a normal disk in the distributed storage system.

[0006] The present application also provides a data reconstruction device, including: a redundancy calculation module, which is used to monitor the status data of the disk to be reconstructed in a distributed storage system, and determine the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed; a redundancy level determination module, which is used to determine a target redundancy level corresponding to the redundancy of the disk to be reconstructed from a plurality of preset redundancy levels; different redundancy levels in the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio; a reconstruction module, which is used to switch the erasure code type of the disk to be reconstructed to a target erasure code type corresponding to the target redundancy level, transfer the faulty data blocks in the disk to be reconstructed to a designated disk according to the target erasure code type, and perform data reconstruction on the designated disk; the designated disk is a normal disk in the distributed storage system.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data reconstruction methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data reconstruction methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned data reconstruction methods when executed by a processor.

[0010] Through the present application, multiple redundancy levels are pre-set, different redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types. The erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio, and based on the real-time status data of the disk to be reconstructed, the redundancy of the disk to be reconstructed is dynamically calculated. According to the calculated redundancy, the target redundancy level corresponding to the redundancy of the disk to be reconstructed is dynamically determined from the preset multiple redundancy levels. Among them, compared with the static redundancy strategy, the dynamic redundancy calculation can not only ensure the rapid recovery of data in the event of a failure, but also avoid storage waste caused by excessive redundancy, effectively balancing data reliability and resource efficiency, and the dynamic grading mechanism can be based on the real-time status data of the disk to be reconstructed. , automatically adjusts the data redundancy level and erasure code type, and can flexibly match the optimal recovery strategy within different redundancy ranges; unlike full copy, the incremental recovery mechanism of the erasure code is adopted in the embodiment of the present application. When the disk to be reconstructed is reconstructed, it does not simply copy the data in full. Instead, based on the redundant information in the erasure code, only the minimum amount of data required to recover the faulty data block is calculated and transmitted, reducing the amount of transmitted data during the data recovery process, thereby reducing the consumption of storage and transmission resources, improving the efficiency and speed of data recovery, ensuring the optimal balance between high performance and low cost during the data reconstruction process of the disk to be reconstructed, maintaining the stability of the front-end business, and solving the problem in the related technology that data recovery performance, resource consumption and network bandwidth resource occupancy cannot be dynamically balanced during the data reconstruction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 An application scenario diagram of a data reconstruction method provided in an embodiment of the present application.

[0013] Figure 2 A flowchart of a data reconstruction method provided in an embodiment of the present application.

[0014] Figure 3 A flowchart of another data reconstruction method provided in an embodiment of the present application.

[0015] Figure 4 A structural diagram of another data reconstruction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0018] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] According to one aspect of the embodiment of the present application, a data reconstruction method is provided. Optionally, in this embodiment, the data reconstruction method can be applied to, but is not limited to, Figure 1 In the distributed storage system shown, multiple disks can exchange data. The multiple disks include a disk to be reconstructed 102 and multiple normal disks 104. Disk to be reconstructed 102 refers to a disk among the multiple disks that requires data reconstruction, such as a failed hard drive or a disk to be expanded. Normal disk 104 refers to a disk among the multiple disks that is storing data normally.

[0020] The above disks include but are not limited to hard disk drives (HDD), solid state drives (SSD), hybrid hard drives (SSHD), etc.

[0021] The data reconstruction method of the embodiment of the present application can be executed by a distributed storage system. Figure 2 is a flow chart of an optional data reconstruction method according to an embodiment of the present application, such as Figure 2 As shown, the process of the method may include the following steps:

[0022] Step S202 : When there is a disk to be reconstructed in the distributed storage system, monitor the status data of the disk to be reconstructed, and determine the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed.

[0023] The data reconstruction method of the present application can be applied to the field of data storage and recovery technology, and is specifically applied to scenarios such as satellite remote sensing data management, large-scale distributed storage clusters, high-performance computing environments, Internet of Things (IoT) data storage, video monitoring and storage. For example, when a disk in a satellite storage system fails, the data reconstruction method of the embodiment of the present application can quickly and efficiently restore satellite remote sensing data, reduce the risk of data loss, and reduce the bandwidth occupancy of the satellite-to-ground link. It is particularly suitable for scenarios where satellite remote sensing platforms need to process a large number of remote sensing images in real time or near real time. For example, in a large distributed storage system composed of thousands or even tens of thousands of disks, when some disks in the large distributed storage system fail or are expanded, the data reconstruction method of the embodiment of the present application can dynamically adjust redundancy and resource allocation according to the real-time system status and data characteristics to balance recovery speed and resource efficiency. It is suitable for large-scale data storage environments such as data centers and cloud computing platforms.

[0024] Distributed storage systems use algorithms to organize these disks into storage pools, providing secure and reliable storage services. However, a key characteristic of distributed storage is that large-scale clusters can consist of large-capacity storage pools comprised of tens of thousands of disks, potentially holding petabytes or even exabytes of data. When a distributed storage system encounters a failed disk or a disk requiring capacity expansion, data reconstruction is necessary to redistribute the data from the failed disk or disk to a new disk to ensure data integrity and consistency. Data reconstruction is a form of data recovery, and recovery speeds vary depending on the data type. For example, in satellite remote sensing scenarios, satellite remote sensing data is typically large in volume, high in resolution, and requires high real-time performance. It may also be distributed across different orbits and sensors. If a storage device fails, recovering this data requires an efficient method that takes into account the data's importance and timeliness. However, related data reconstruction techniques rely on copying all the data from the failed disk or disk requiring capacity expansion to other functioning disks. This method consumes significant network bandwidth, especially in large file scenarios. For example, in satellite remote sensing scenarios, satellite-to-ground link bandwidth utilization can exceed 80%, disrupting real-time data downlinks and resulting in low recovery efficiency and high resource consumption.

[0025] Therefore, in order to solve this problem, an embodiment of the present application determines the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed. The redundancy changes dynamically with the status data of the disk to be reconstructed, and automatically adjusts the erasure code parameters according to the dynamic redundancy to achieve a dynamic balance between data recovery performance and network bandwidth resource occupancy.

[0026] The disk to be reconstructed refers to a disk in the distributed storage system that has failed or is about to undergo data migration for capacity expansion. The status data of the disk to be reconstructed refers to a set of information used to describe the health status of the disk to be reconstructed and the current operating status of the distributed storage system. It is used to calculate the redundancy of the disk and determine the subsequent reconstruction strategy. For example, in the scenario of satellite remote sensing data management, the status data of the disk to be reconstructed can be obtained from the storage manager, fault log, and satellite-to-ground link status monitor. The status data of the disk to be reconstructed includes but is not limited to the remaining storage capacity of the node (C free ), historical node failure rate (p fail ), satellite-to-ground link bandwidth constraint (B limit ) and other parameters.

[0027] The redundancy of the disk to be reconstructed refers to the ratio of the amount of additional data stored in the distributed storage system to the amount of original data in order to ensure the integrity and reliability of the data. It is used to guide the selection of the erasure code type to achieve optimal data recovery performance and resource efficiency. In an embodiment of the present application, the redundancy of the disk to be reconstructed is calculated based on real-time status data, that is, the redundancy of the disk to be reconstructed changes dynamically based on real-time status data. For example, the health index of the disk to be reconstructed can be predicted based on the status data of the disk to be reconstructed, and the target health level of the health index can be determined from multiple health levels based on the health index. Different health levels correspond to different redundancies, and the redundancy corresponding to the target health level is determined as the redundancy of the disk to be reconstructed. For example, the status data of the disk to be reconstructed is input into a pre-trained prediction model to obtain the redundancy of the disk to be reconstructed, wherein the prediction model is used to determine the redundancy corresponding to the disk based on the input status data of the disk.

[0028] In step S204, a target redundancy level corresponding to the redundancy of the disk to be reconstructed is determined from a plurality of preset redundancy levels; different redundancy levels in the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio.

[0029] Multiple redundancy levels are used to measure data redundancy levels in distributed storage systems, corresponding to different data recovery strategies. Each level is associated with a specific redundancy interval and erasure code type. Because the redundancy of the disk to be reconstructed changes dynamically, the redundancy level of the disk to be reconstructed also changes dynamically. The target redundancy level is a level selected from multiple preset redundancy levels based on the current state of the distributed storage system and the redundancy calculation results of the disk to be reconstructed.

[0030] Erasure coding is a data redundancy technology that protects stored data from corruption or loss by generating parity blocks. Unlike full replication, this embodiment uses an incremental recovery mechanism with erasure coding. When reconstructing data on the disk to be reconstructed, instead of simply replicating the entire data, only the minimum amount of data required to recover the failed data block is calculated and transmitted based on the redundant information in the erasure code. This reduces the amount of data transmitted during the data recovery process, thereby reducing network bandwidth usage. Different erasure code types (such as RS(6,2), RS(8,3), and RS(10,4)) represent different encoding methods and correspond to different redundant data ratios. The redundant data ratio represents the ratio of the number of parity blocks to the total number of data blocks after encoding. The numbers in different erasure code types represent the total number of data blocks and the number of parity blocks after encoding. Different encoding methods result in different data recovery speeds and storage resource usage ratios. For example, the "6" in RS(6,2) represents a total of six data blocks after encoding, while the "2" indicates that two of them are redundant parity blocks. This means that the distributed storage system generates two additional check blocks based on the original data, bringing the total number of data blocks to six. Under the RS(6,2) encoding method, the original data is divided into four data blocks, and two redundant check blocks are generated. This means that the redundant data ratio is 33% (the two check blocks account for one-third of the total number of six blocks). Therefore, during data recovery, only a small number of data blocks are required to quickly reconstruct the complete data. For example, the division of multiple redundancy levels and their corresponding explanations are shown in Table 1 below:

[0031] Table 1

[0032]

[0033] The redundancy levels in Table 1 are categorized into high, medium, and low levels, where γ represents redundancy. High redundancy levels (e.g., RS(6,2), with a redundancy range of ≥0.5) indicate a high level of redundant data and are suitable for scenarios with a surge in storage node failure rates or unstable links. Medium redundancy levels (e.g., RS(8,3), with a redundancy range of 0.3–0.5) balance data integrity and resource efficiency and are suitable for scenarios with normal storage pressures and ample bandwidth. Low redundancy levels (e.g., RS(10,4), with a redundancy range of <0.3) are suitable for scenarios with sufficient storage redundancy and limited link bandwidth. This redundancy grading strategy ensures that the distributed storage system automatically selects the most appropriate erasure code type under different conditions, achieving an optimal balance between data recovery speed and storage resource utilization.

[0034] In order to solve the problem in related technologies that data recovery performance, resource consumption and network bandwidth resource usage cannot be dynamically balanced during data reconstruction, in this implementation, the higher the redundancy level, the greater the required bandwidth, and the erasure code type will have a higher ratio of check blocks to data blocks (i.e., the redundant data ratio).

[0035] In step S206, the erasure code type of the disk to be reconstructed is switched to the target erasure code type corresponding to the target redundancy level. According to the target erasure code type, the faulty data blocks in the disk to be reconstructed are transferred to the designated disk, and data reconstruction is performed on the designated disk; the designated disk is a normal disk in the distributed storage system.

[0036] The target erasure code type refers to the erasure code type corresponding to the target redundancy level of the disk to be reconstructed.

[0037] The designated disk specifically refers to a normal disk that is selected to receive and store faulty data blocks from the disk to be reconstructed during the data reconstruction process. The number of the designated disk can be one or more, depending on actual needs. The designated disk usually has sufficient storage space and is in good health to carry out data reconstruction tasks and ensure smooth data recovery. In the embodiment of the present application, the designated disk is screened out from the distributed storage system through a certain screening mechanism. For example, by comprehensively evaluating the storage capacity, read and write speed, network bandwidth, and current task load of the disk, the disk with the most abundant resources and the lowest load is dynamically selected as the designated disk for the data reconstruction task.

[0038] Optionally, failed data blocks on the disk to be reconstructed are marked and prepared for migration. Using a network transmission mechanism, these failed data blocks are transferred from the disk to a designated disk selected through a load balancing algorithm, ensuring that the designated disk has sufficient storage space and network bandwidth to receive and process the data. Data reconstruction is initiated on the designated disk, using the switched target erasure code type to decode and reconstruct the data blocks and restore data integrity.

[0039] Through this embodiment, multiple redundancy levels are pre-set, different redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types. The erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio, and based on the real-time status data of the disk to be reconstructed, the redundancy of the disk to be reconstructed is dynamically calculated. According to the calculated redundancy, the target redundancy level corresponding to the redundancy of the disk to be reconstructed is dynamically determined from the preset multiple redundancy levels. Compared with the static redundancy strategy, the dynamic redundancy calculation can not only ensure the rapid recovery of data in the event of a failure, but also avoid storage waste caused by excessive redundancy, effectively balancing data reliability and resource efficiency, and the dynamic grading mechanism can be based on the real-time status data of the disk to be reconstructed. According to the data, the data redundancy level and the erasure code type are automatically adjusted, and the optimal recovery strategy can be flexibly matched within different redundancy ranges; unlike full copy, the incremental recovery mechanism of the erasure code is adopted in the embodiment of the present application. When the disk to be reconstructed is reconstructed, it does not simply copy the data in full. Instead, based on the redundant information in the erasure code, only the minimum amount of data required to recover the faulty data block is calculated and transmitted, which reduces the amount of data transmitted during the data recovery process, thereby reducing the consumption of storage and transmission resources, improving the efficiency and speed of data recovery, ensuring the optimal balance between high performance and low cost during the data reconstruction process of the disk to be reconstructed, maintaining the stability of the front-end business, and solving the problem in the related technology that data recovery performance, resource consumption and network bandwidth resource occupancy cannot be dynamically balanced during the data reconstruction process.

[0040] In an exemplary embodiment, the status data of the disk to be reconstructed is used to indicate the following information of the disk to be reconstructed: remaining storage capacity, and historical failure rate;

[0041] Among them, the remaining storage capacity (C free ) refers to the amount of storage space currently unused by data on the disk to be reconstructed, typically measured in terabytes or larger. In satellite remote sensing scenarios, the onboard storage management module can calculate the remaining storage capacity of the disk to be reconstructed in real time.

[0042] Historical failure rate (p fail ) represents the frequency of failures of the disk to be reconstructed in the past period of time (e.g., the past 24 hours), including but not limited to device downtime, disk failure, network interruption, etc. In an embodiment, the historical failure rate is calculated by counting the failure events in a specific time window. The historical failure rate can be calculated using the following formula (1):

[0043]

[0044] In some embodiments, based on the status data of the disk to be reconstructed, the redundancy of the disk to be reconstructed is determined, including: determining the proportion of used storage capacity of the disk to be reconstructed according to the total storage capacity and the remaining storage capacity of the disk to be reconstructed; according to a first weight coefficient of the proportion of used storage capacity and a second weight coefficient of the historical failure rate, weighted summing the proportion of used storage capacity and the historical failure rate to obtain the redundancy of the disk to be reconstructed; the first weight coefficient and the second weight coefficient are dynamically adjusted according to the recovery efficiency index of the disk to be reconstructed.

[0045] The used storage capacity percentage of the disk to be reconstructed refers to the ratio of the amount of data currently stored on the disk to its total storage capacity. This percentage can be calculated by monitoring the total storage capacity and remaining storage capacity of the disk in real time.

[0046] The first weight coefficient is used to quantify the importance of the proportion of used storage capacity in calculating redundancy. The second weight coefficient is used to quantify the importance of the historical failure rate in calculating redundancy. The first and second weight coefficients are dynamically adjusted based on the recovery efficiency index of the disk to be reconstructed. The recovery efficiency index refers to multiple parameters that comprehensively consider the quality, speed, and resource consumption of data recovery under the current state of the disk to be reconstructed and the actions performed. For example, the recovery efficiency index can be shown in Table 2 below:

[0047] Table 2

[0048] Indicator Type Key parameters Acquisition frequency Storage pressure <![CDATA[C free / C total ]]> 1 minute Link Status Bandwidth utilization, packet loss rate 5 minutes Recovery performance Data reconstruction speed and success rate real time

[0049] This embodiment monitors the recovery efficiency index and dynamically adjusts the first weight coefficient and the second weight coefficient to optimize the redundancy calculation and implement a more efficient and intelligent data recovery strategy. The redundancy of the disk to be reconstructed can be expressed using the following formula (2):

[0050]

[0051] Among them, α represents the first weight coefficient, β represents the second weight coefficient, and p fail represents the historical failure rate, C free Indicates the remaining storage capacity of the disk to be reconstructed, C total Indicates the total storage capacity of the disk to be reconstructed; Indicates the remaining storage ratio, reflecting storage pressure.

[0052] Through this embodiment, the proportion of used storage capacity and the historical failure rate of the disk to be reconstructed are included as core parameters in the redundancy calculation, and the storage pressure and failure risk are comprehensively evaluated by weighted summation of the proportion of used storage capacity and the historical failure rate. The redundancy can be intelligently adjusted according to the actual status of different disks, thereby determining the accurate redundancy, which not only ensures the security of data in high failure risk conditions, but also avoids unnecessary resource consumption in low storage pressure environments.

[0053] In one exemplary embodiment, the first and second weight coefficients are dynamically adjusted using reinforcement learning technology. The data reconstruction method further includes: monitoring the recovery efficiency index of the disk to be reconstructed within a specified time period to obtain an index parameter corresponding to the specified time period; inputting the index parameter into a pre-trained reinforcement learning model to obtain the first and second weight coefficients; and the reinforcement learning model is used to adjust the first and second weight coefficients based on the input index parameter.

[0054] The indicator parameter of the specified time period specifically refers to the specific value of the recovery efficiency indicator collected by the distributed storage system within the current time window.

[0055] In this embodiment, the first and second weight coefficients are dynamically optimized based on the recovery efficiency index through a reinforcement learning model. The reinforcement learning model takes the recovery efficiency index parameters as input and, after learning and iteration, outputs the optimized first and second weight coefficients. These are used to adjust the proportion of used storage capacity and the impact of historical failure rates in subsequent redundancy calculations. This enables the distributed storage system to adaptively optimize resource allocation strategies, improving data recovery efficiency while reducing resource waste.

[0056] In this embodiment, the first and second weight coefficients are not fixed but are dynamically optimized using a reinforcement learning algorithm based on the recovery efficiency index of the disk to be reconstructed. This means that the distributed storage system automatically adjusts the first and second weight coefficients based on the recovery efficiency index to find the redundancy that best suits the current environment, achieving efficient resource allocation and intelligent optimization of data recovery strategies.

[0057] In an exemplary embodiment, a method for training a reinforcement learning model includes the following steps:

[0058] Perform multiple rounds of training operations until the reinforcement learning model meets the training end conditions, and obtain a trained reinforcement learning model: wherein, in the process of performing the current round of training operations, select the current state from a pre-constructed state space, and select the current action from a pre-constructed action space, wherein the state space includes multiple states, and the state parameter of one of the multiple states is a sample indicator parameter; the action space includes multiple actions; the action parameter of one of the multiple actions includes an action instruction for adjusting the redundancy level; the current action is executed through the current reinforcement learning model, and the immediate reward value corresponding to the current action is determined according to the reward function; the action value between the current state and the current action is determined according to the immediate reward value corresponding to the current action and the value function, and the current reinforcement learning model is adjusted according to the action value between the current state and the current action; the action value includes a first expected long-term benefit and a second expected long-term benefit; the value function represents the mapping relationship between the action value and the first weight coefficient, the second weight coefficient, the first expected long-term benefit and the second expected long-term benefit.

[0059] Among them, the state space (s) contains a series of possible state sets of the distributed storage system, and the state parameter of each state is a sample indicator parameter. For example, the state parameter of each state includes historical redundancy, node failure rate, and network bandwidth fluctuation coefficient.

[0060] The action space (a) defines the set of actions the reinforcement learning model can take, such as adjusting the redundancy level to high, medium, or low. These actions determine the resource allocation and data recovery strategy of the distributed storage system. The reinforcement learning module explores different actions to find the optimal resource allocation solution for the current state.

[0061] The immediate reward value is the feedback value calculated based on the reward function after the current action is executed. It reflects the direct impact of the action on system performance (such as data recovery success rate, bandwidth utilization, task time consumption, etc.) in the current state. In this embodiment, the reward function can be calculated according to the following formula (3):

[0062]

[0063] Where ω1, ω2, and ω3 are weight coefficients used to balance reliability, efficiency, and cost; the recovery success rate is the proportion of failed data blocks that are successfully recovered; the bandwidth usage is the actual bandwidth occupied by redundant data transmission; and the task duration is the time required to rebuild the data (in seconds).

[0064] Action value is the expected long-term benefit calculated using the immediate reward value and the value function, including a first expected long-term benefit and a second expected long-term benefit. The first expected long-term benefit specifically refers to the expected level of data integrity, consistency, and high availability that a distributed storage system can maintain over a long period of time. It measures the long-term ability of the data recovery strategy to effectively recover data, ensure business continuity, and ensure user data security in the face of various failures and anomalies. The second expected long-term benefit refers to the long-term performance of the distributed storage system's resource utilization efficiency and data processing speed. It reflects the ability of the distributed storage system to complete data recovery with minimal resource consumption, avoid excessive bandwidth usage and waste of computing resources, and ensure the speed and quality of data recovery while ensuring data security. The value function is used to characterize the mathematical relationship between the action value and the first weight coefficient, the second weight coefficient, the first expected long-term benefit, and the second expected long-term benefit, guiding the reinforcement learning model to learn and select behaviors that maximize long-term benefits. In this embodiment, the action value can be the sum of the first weighted result and the second weighted result, where the first weighted result refers to the weighted result of the first weight coefficient and the first expected long-term benefit, and the second weighted result refers to the weighted result of the second weight coefficient and the second expected long-term benefit. Thus, the value function can be expressed as follows:

[0065] Q(s,a;α,β)=α·Q reliability (s,a)+β·Q efficiency (s,a) (4)

[0066] Among them, Q(s,a;α,β) represents the action value; Q reliability (s,a) represents the first expected long-term benefit of the current state-current action pair; Q efficiency (s, a) represents the second expected long-term benefit of the current state-current action pair. In this embodiment, the initial value of α is 0.6 and the initial value of β is 0.4. α and β are dynamically adjusted through training of the reinforcement learning model.

[0067] Through this embodiment, the first expected long-term benefit and the second expected long-term benefit are respectively associated with the reliability and resource efficiency of data recovery, and the first weight coefficient and the second weight coefficient are embedded in the value function. By adjusting the values ​​of the first weight coefficient and the second weight coefficient, the value function can better seek a dynamic balance between reliability (first expected long-term benefit) and efficiency (second expected long-term benefit).

[0068] In an exemplary embodiment, according to the target erasure code type, the faulty data blocks in the disk to be reconstructed are transferred to the designated disk, including: dividing the faulty data blocks in the disk to be reconstructed into multiple sub-fault blocks according to the target erasure code type; constructing multiple subtasks; the multiple subtasks correspond one-to-one to the multiple sub-fault blocks; the subtasks in the multiple subtasks are used to transfer the corresponding sub-fault blocks to the designated disk for data reconstruction; according to the redundancy of the disk to be reconstructed, adjusting the computing resources of the multiple subtasks, and using the adjusted multiple subtasks to transfer the multiple sub-fault blocks to the designated disk for data reconstruction.

[0069] A faulty data block refers to a data unit stored on a faulty disk that is inaccessible or incomplete due to a disk failure. A faulty subblock is a smaller data unit further divided into by the target erasure code type. By dividing a faulty data block into multiple faulty subblocks, data recovery can be performed simultaneously on multiple designated disks, thereby shortening the overall reconstruction time. Multiple faulty subblocks include multiple data blocks and multiple parity blocks, where the ratio of multiple parity blocks to multiple data blocks is determined by the redundant data ratio of the target erasure code type. For example, if the target erasure code type is RS(8,3), the faulty data block is divided into 8 faulty subblocks, each of which includes 5 data blocks and 3 parity blocks.

[0070] During the data recovery process, a subtask is a reconstruction task assigned to each faulty subblock. Each subtask corresponds to the recovery process for a faulty subblock, including data transmission and reconstruction calculations. By executing multiple subtasks in parallel, the overall efficiency of data recovery can be significantly improved.

[0071] Computing resources refer to the hardware resources, such as CPU and memory, required to execute a subtask. Based on the redundancy level of the disk to be reconstructed, the system dynamically adjusts the computing resources allocated to the subtask to balance the recovery speed of redundant data with the bandwidth utilization of the satellite-to-ground link, ensuring optimal resource utilization.

[0072] Optionally, Figure 3 A flowchart of another data reconstruction method provided in an embodiment of the present application is shown in FIG. Figure 3As shown, the distributed storage system includes a status monitoring module, a dynamic decision module, a resource scheduling module and a feedback optimization module, wherein the status monitoring module is used to monitor the status data of the disk to be reconstructed. The feedback optimization module is used to monitor the recovery efficiency index during the data reconstruction process, and use the reinforcement learning model to perform reinforcement learning on the indicator parameters of the monitored recovery efficiency index to obtain a first weight coefficient and a second weight coefficient, and feed the first weight coefficient and the second weight coefficient back to the dynamic decision module. The dynamic decision module calculates the redundancy of the disk to be reconstructed based on the first weight coefficient, the second weight coefficient and the status data of the disk to be reconstructed, determines the target redundancy level corresponding to the redundancy of the disk to be reconstructed, and switches the erasure code type of the disk to be reconstructed to the target erasure code type corresponding to the target redundancy level. The resource scheduling module divides the faulty data blocks in the disk to be reconstructed into multiple sub-fault blocks according to the target erasure code type; based on the calculated redundancy of the disk to be reconstructed, the erasure code reconstruction task is divided into multiple sub-tasks, with multiple sub-tasks corresponding to multiple sub-fault blocks one-to-one. The data recovery rate and the ratio of network bandwidth occupied resources are adjusted, and the multiple sub-fault blocks are transferred to the designated disk for data reconstruction using the adjusted sub-tasks, avoiding resource overload through dynamic threads.

[0073] This embodiment divides the faulty data block on the disk to be reconstructed into multiple sub-faulty blocks based on the target erasure code type, and constructs an independent subtask for each sub-faulty block. This strategy not only facilitates parallel processing and improves data recovery speed, but also flexibly allocates resources based on the specific needs of the data block, avoiding resource waste. Dynamically adjusting the computing resources required to execute subtasks significantly saves network bandwidth and balances data recovery performance with resource consumption.

[0074] In some embodiments, computing resources of multiple subtasks are adjusted according to the redundancy of the disk to be reconstructed, including: collecting status data of each designated disk in real time, and calculating the disk redundancy of each designated disk based on the status data of each designated disk; determining the risk level of each designated disk based on the disk redundancy of each designated disk; different risk levels correspond to different resource allocation strategies and priorities; determining the priority and resource allocation strategy corresponding to each subtask based on the risk level of the designated disk to which each subtask belongs.

[0075] Optionally, a status monitoring module is started to collect status data of each designated disk in real time, including remaining storage capacity, historical failure rate, network bandwidth status, etc. The status data of each designated disk is analyzed in real time, and the disk redundancy of each designated disk at the current moment is obtained according to a preset redundancy calculation formula (such as the above formula (2)). By comparing historical data, disks with redundancy lower than a first preset threshold (which can be an average redundancy or set according to actual needs) are identified from multiple designated disks and marked as high-risk disks. Disks with redundancy greater than the first preset threshold and less than the second preset threshold (which can be set according to actual needs) are identified from multiple designated disks and marked as medium-risk disks. Disks with redundancy greater than the third preset threshold are identified from multiple designated disks and marked as low-risk disks (which can be set according to actual needs). For subtasks on high-risk disks, the highest priority and the most abundant computing resources are allocated; tasks on medium-risk disks are ranked second; tasks on low-risk disks have the lowest priority and are only allocated basic computing resources. When reconstructing data on the disk to be reconstructed, each subtask is automatically assigned a priority based on the redundancy of its associated disk. Tasks with higher priorities are prioritized for computing resources. The resource scheduling module dynamically adjusts the computing resource allocation for each subtask based on the current disk redundancy and task priority of each designated disk. High-priority tasks are prioritized for resources on high-performance computing nodes to accelerate data recovery, while low-priority tasks are assigned to nodes with average performance to avoid wasted resources.

[0076] Through this embodiment, the risk level of each specified disk is determined based on its disk redundancy, and specific resource allocation strategies and task priorities are formulated according to different risk levels. Subtasks corresponding to high-risk disks obtain the highest priority and the most abundant computing resources, while subtasks corresponding to low-risk disks have lower priority and more conservative resource allocation. This ensures that resources are preferentially allocated to the disks that need them most, avoids waste of resources in high-risk situations, and improves the efficiency and success rate of data recovery. Secondly, by dynamically adjusting the priority of each subtask, it is possible to intelligently respond to real-time changes in disk status, thereby optimizing resource usage while ensuring data integrity, reducing excessive occupation of satellite-to-ground link bandwidth, and ensuring business stability and continuity.

[0077] In an exemplary example, the above-mentioned data reconstruction method also includes: reserving specified storage capacity in a specified disk; the specified storage capacity refers to a storage capacity of a preset proportion in the remaining storage capacity of the specified disk; the specified storage capacity is used to cache redundant data generated by the specified disk during the data reconstruction process.

[0078] The designated storage capacity is the portion of the remaining storage space on a specified disk that is reserved for storing redundant data. The designated storage capacity is determined based on a preset percentage (for example, 10% of the total remaining storage capacity). This ensures that even if new failures or anomalies occur during the data reconstruction process, there is sufficient space to store additional redundant data, thereby improving data recovery integrity and system robustness.

[0079] Redundant data specifically refers to one or more data copies or coded information that are intentionally generated and saved during the data reconstruction process to cope with possible secondary failures of a specified disk or potential risks in the future.

[0080] Through this embodiment, designated storage capacity is reserved as redundant data cache. Even if a new disk failure occurs during or immediately after data reconstruction, the pre-stored redundant data can be used for rapid recovery, avoiding secondary data loss, thereby enhancing the overall stability and fault tolerance of the distributed storage system.

[0081] In an exemplary embodiment, before data reconstruction is performed on a disk to be reconstructed, the erasure code type of the disk to be reconstructed is used to indicate that a first redundancy level is assigned to data in a first data tier and a second redundancy level is assigned to data in a second data tier; the first data tier is used to store data with an access frequency greater than a first frequency value; the second data tier is used to store data with an access frequency less than a second frequency value; and the first redundancy level is greater than the second redundancy level.

[0082] Before the erasure coding type of the disk to be reconstructed is switched to the target erasure coding type corresponding to the target redundancy level, the erasure coding type of the disk to be reconstructed is to divide the disk data into a first data tier and a second data tier based on data access frequency. Different tiers have different redundancy levels, i.e., different tiers have different data recovery rates and storage resource utilization ratios. The first data tier is a storage tier designed specifically for storing hot data with an access frequency exceeding a first frequency threshold. Hot data generally refers to data that is frequently accessed or updated. A higher redundancy level is used for hot data to ensure rapid recovery and high availability in the event of storage anomalies. The second data tier is a storage tier for storing cold data with an access frequency below the second frequency threshold. Cold data is less frequently accessed than hot data. Using a lower redundancy level can reduce storage space usage and resource consumption while ensuring data security. In this embodiment, the first redundancy level of the first data tier is greater than the second redundancy level of the second data tier. This means that hot data has higher redundancy and is more resilient to storage failures, while cold data optimizes storage space usage while meeting basic security requirements.

[0083] When the disk to be reconstructed is undergoing data reconstruction, the erasure code type of the disk to be reconstructed needs to be switched to the target erasure code type corresponding to the target redundancy level.

[0084] In some embodiments, not only the disk to be reconstructed can adopt the erasure code type of this embodiment, but any disk in the distributed storage system can also adopt the solution of setting different erasure code types according to access frequency in this embodiment.

[0085] Through this embodiment, before data reconstruction, each disk in the distributed storage system divides the disk data into a first data layer and a second data layer according to different data access frequencies, and the first redundancy level of the first data layer is greater than the second redundancy level of the second data layer. This ensures that high-frequency access data (first data layer) is quickly recovered while reducing the occupation of network bandwidth. It can be quickly located and restored during data reconstruction without having to transmit the entire data, significantly improving the speed of data recovery while reducing the consumption of bandwidth resources. For the second data layer with a lower access frequency, a lower redundancy level is adopted, which not only saves storage space but also avoids unnecessary computing and bandwidth overhead, thereby realizing refined resource management.

[0086] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0087] The embodiment of the present application also provides a data reconstruction device, such as Figure 4 As shown, including:

[0088] The redundancy calculation module 402 is configured to monitor the status data of the disk to be reconstructed when there is a disk to be reconstructed in the distributed storage system, and determine the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed;

[0089] Redundancy level determination module 404 is configured to determine a target redundancy level corresponding to the redundancy of the disk to be reconstructed from a plurality of preset redundancy levels. Different redundancy levels in the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types. The erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource usage ratio.

[0090] The reconstruction module 406 is used to switch the erasure code type of the disk to be reconstructed to the target erasure code type corresponding to the target redundancy level, transfer the faulty data blocks in the disk to be reconstructed to the designated disk according to the target erasure code type, and reconstruct the data on the designated disk; the designated disk is a normal disk in the distributed storage system.

[0091] In an exemplary embodiment, the status data of the disk to be reconstructed is used to indicate the following information of the disk to be reconstructed: the remaining storage capacity, and the historical failure rate; the redundancy calculation module 402 is also used to determine the proportion of used storage capacity of the disk to be reconstructed based on the total storage capacity and the remaining storage capacity of the disk to be reconstructed; according to the first weight coefficient of the proportion of used storage capacity and the second weight coefficient of the historical failure rate, the proportion of used storage capacity and the historical failure rate are weighted and summed to obtain the redundancy of the disk to be reconstructed; the first weight coefficient and the second weight coefficient are dynamically adjusted according to the recovery efficiency index of the disk to be reconstructed.

[0092] In an exemplary embodiment, the redundancy calculation module 402 is also used to monitor the recovery efficiency index of the disk to be reconstructed within a specified time period to obtain index parameters corresponding to the specified time period; the index parameters are input into a pre-trained reinforcement learning model to obtain a first weight coefficient and a second weight coefficient; the reinforcement learning model is used to adjust the first weight coefficient and the second weight coefficient according to the input index parameters.

[0093] In an exemplary embodiment, the redundancy calculation module 402 is also used to perform multiple rounds of training operations until the reinforcement learning model meets the training end conditions to obtain a trained reinforcement learning model; wherein, in the process of performing the current round of training operations, the current state is selected from a pre-constructed state space, and the current action is selected from a pre-constructed action space, wherein the state space includes multiple states, and the state parameter of one of the multiple states is a sample indicator parameter; the action space includes multiple actions; the action parameter of one of the multiple actions includes an action instruction for adjusting the redundancy level; the current action is executed through the current reinforcement learning model, and the immediate reward value corresponding to the current action is determined according to the reward function; the action value between the current state and the current action is determined according to the immediate reward value corresponding to the current action and the value function, and the current reinforcement learning model is adjusted according to the action value between the current state and the current action; the action value includes a first expected long-term benefit and a second expected long-term benefit; the value function represents the mapping relationship between the action value and the first weight coefficient, the second weight coefficient, the first expected long-term benefit and the second expected long-term benefit.

[0094] In an exemplary embodiment, the reconstruction module 406 is also used to divide the faulty data blocks in the disk to be reconstructed into multiple sub-fault blocks according to the target erasure code type; construct multiple subtasks; the multiple subtasks correspond one-to-one to the multiple sub-fault blocks; the subtasks in the multiple subtasks are used to transfer the corresponding sub-fault blocks to the designated disk for data reconstruction; according to the redundancy of the disk to be reconstructed, the computing resources of the multiple subtasks are adjusted, and the adjusted multiple subtasks are used to transfer the multiple sub-fault blocks to the designated disk for data reconstruction.

[0095] In an exemplary embodiment, the reconstruction module 406 is also used to reserve a specified storage capacity in a specified disk; the specified storage capacity refers to a storage capacity that is a preset percentage of the remaining storage capacity of the specified disk; the specified storage capacity is used to cache redundant data generated by the specified disk during the data reconstruction process.

[0096] In an exemplary embodiment, the erasure code type of the disk to be reconstructed is used to indicate that a first redundancy level is assigned to data in a first data tier and a second redundancy level is assigned to data in a second data tier; the first data tier is used to store data with an access frequency greater than a first frequency value; the second data tier is used to store data with an access frequency less than a second frequency value; and the first redundancy level is greater than the second redundancy level.

[0097] For the description of the features in the embodiment corresponding to the data reconstruction device, reference can be made to the relevant description of the embodiment corresponding to the data reconstruction method, which will not be repeated here.

[0098] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above data reconstruction method embodiments.

[0099] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data reconstruction method embodiments when run.

[0100] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0101] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above data reconstruction method embodiments are implemented.

[0102] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data reconstruction method embodiments are implemented.

[0103] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0104] The above is a detailed introduction to a data reconstruction method, device, electronic device, and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data reconstruction method, characterized in that: include: In a case where a disk to be reconstructed exists in the distributed storage system, monitoring status data of the disk to be reconstructed, and determining the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed; Determining a target redundancy level corresponding to the redundancy of the disk to be reconstructed from a plurality of preset redundancy levels; different redundancy levels in the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio; The erasure code type of the disk to be reconstructed is switched to the target erasure code type corresponding to the target redundancy level. According to the target erasure code type, the faulty data blocks in the disk to be reconstructed are transferred to a designated disk, and data reconstruction is performed on the designated disk; the designated disk is a normal disk in the distributed storage system.

2. The method according to claim 1, characterized in that The status data of the disk to be reconstructed is used to indicate the following information of the disk to be reconstructed: remaining storage capacity, and historical failure rate; The determining the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed includes: Determining a proportion of used storage capacity of the disk to be reconstructed according to the total storage capacity of the disk to be reconstructed and the remaining storage capacity; According to a first weight coefficient of the proportion of used storage capacity and a second weight coefficient of the historical failure rate, a weighted sum is taken of the proportion of used storage capacity and the historical failure rate to obtain the redundancy of the disk to be reconstructed; the first weight coefficient and the second weight coefficient are dynamically adjusted according to the recovery efficiency index of the disk to be reconstructed.

3. The method according to claim 2, characterized in that The method further comprises: Monitoring the recovery efficiency index of the disk to be reconstructed within a specified time period to obtain an index parameter corresponding to the specified time period; The indicator parameters are input into a pre-trained reinforcement learning model to obtain the first weight coefficient and the second weight coefficient; the reinforcement learning model is used to adjust the first weight coefficient and the second weight coefficient according to the input indicator parameters.

4. The method according to claim 3, characterized in that The method further comprises: Perform multiple rounds of training operations until the reinforcement learning model meets the training end condition, thereby obtaining the trained reinforcement learning model; In the process of performing the current round of training operation, a current state is selected from a pre-constructed state space, and a current action is selected from a pre-constructed action space, wherein the state space includes a plurality of states, a state parameter of one of the plurality of states is a sample indicator parameter; the action space includes a plurality of actions; and an action parameter of one of the plurality of actions includes an action instruction for adjusting a redundancy level; Executing the current action through the current reinforcement learning model and determining an immediate reward value corresponding to the current action according to a reward function; According to the immediate reward value and value function corresponding to the current action, the action value between the current state and the current action is determined, and according to the action value between the current state and the current action, the current reinforcement learning model is adjusted; the action value includes a first expected long-term benefit and a second expected long-term benefit; the value function represents the mapping relationship between the action value and the first weight coefficient, the second weight coefficient, the first expected long-term benefit and the second expected long-term benefit.

5. The method according to claim 1, wherein The step of transferring the failed data block in the disk to be reconstructed to a designated disk according to the target erasure code type includes: Dividing the faulty data block in the disk to be reconstructed into a plurality of sub-faulty blocks according to the target erasure code type; Constructing a plurality of subtasks; wherein the plurality of subtasks correspond one-to-one to the plurality of sub-fault blocks; a subtask in the plurality of subtasks is used to transfer the corresponding sub-fault block to the designated disk for data reconstruction; The computing resources of the plurality of subtasks are adjusted according to the redundancy of the disk to be reconstructed, and the plurality of sub-failure blocks are transferred to the designated disk for data reconstruction using the adjusted plurality of subtasks.

6. The method according to claim 1, characterized in that The method further comprises: A specified storage capacity is reserved in the specified disk; the specified storage capacity refers to a storage capacity of a preset proportion in the remaining storage capacity of the specified disk; the specified storage capacity is used to cache redundant data generated by the specified disk during the data reconstruction process.

7. The method according to any one of claims 1 to 6, characterized in that The erasure code type of the disk to be reconstructed is used to indicate that a first redundancy level is assigned to data in a first data layer, and a second redundancy level is assigned to data in a second data layer; the first data layer is used to store data with an access frequency greater than a first frequency value; the second data layer is used to store data with an access frequency less than a second frequency value; and the first redundancy level is greater than the second redundancy level.

8. A data reconstruction device, characterized in that: include: a redundancy calculation module configured to monitor the status data of a disk to be reconstructed when there is one in the distributed storage system, and determine the redundancy of the disk to be reconstructed based on the status data of the disk to be reconstructed; a redundancy level determination module, configured to determine a target redundancy level corresponding to the redundancy of the disk to be reconstructed from a plurality of preset redundancy levels; different redundancy levels in the plurality of redundancy levels correspond to different redundancy intervals, and different redundancy levels correspond to different erasure code types; the erasure code types corresponding to different redundancy levels represent different combinations of data recovery rate and storage resource occupancy ratio; A reconstruction module is used to switch the erasure code type of the disk to be reconstructed to the target erasure code type corresponding to the target redundancy level, transfer the faulty data blocks in the disk to be reconstructed to a designated disk according to the target erasure code type, and reconstruct the data on the designated disk; the designated disk is a normal disk in the distributed storage system.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data reconstruction method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data reconstruction method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data reconstruction method and device, storage medium and program product

    CN119336536A

  • Medical image management method and system based on photo-electromagnetic hybrid hierarchical storage

    CN120066420A

  • Method and device for reconstructing stored data based on erasure coding, and storage node

    WO2018001110A1

  • Data management method, apparatus and system for distributed storage system, and electronic device

    WO2025113088A1

Cited By

  • Data reconstruction method and electronic equipment

    CN120929298A