Data repair method and apparatus

By building and updating the full-rank check matrix of distributed storage node clusters, the problems of low data repair efficiency and high bandwidth consumption in the existing technology are solved, and efficient data repair and storage optimization are achieved.

WO2025138971A1PCT designated stage expired Publication Date: 2025-07-03HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/115449
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-08-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing data repair scheme based on erasure codes has problems such as low storage efficiency, complex repair process and large bandwidth consumption in distributed storage systems, making it difficult to effectively reduce the amount of redundant data and repair bandwidth consumption.

Method used

Build the initial check matrix and update it according to the matrix full rank strategy to generate a full rank check matrix. Through the check matrix elements, determine the target node in the distributed storage node cluster and calculate the repair data of the failed nodes, reducing the node repair bandwidth and hard disk consumption.

Benefits of technology

It realizes data repair that meets the MDS nature in a distributed storage node cluster, improves repair capabilities, reduces node repair bandwidth and hard disk consumption, and optimizes system storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115449_03072025_PF_FP_ABST
    Figure CN2024115449_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a data repair method and apparatus. The data repair method comprises: for a distributed storage node cluster, constructing an initial check matrix for data repair; updating a target block matrix in the initial check matrix according to a matrix full-rank strategy, so as to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix meets a full-rank condition of the matrix full-rank strategy; when the distributed storage node cluster has a faulty node, selecting from the check matrix a matrix element associated with the faulty node; determining a target node from the distributed storage node cluster on the basis of the matrix element, and acquiring node storage data corresponding to the target node; and calculating repair data of the faulty node on the basis of the matrix element and the node storage data, and uploading the repair data to the faulty node. A repair function for faulty nodes is implemented while reducing node repair bandwidth and hard disk consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Data repair method and device

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 28, 2023, with application number 202311846063.4 and application name “Data Repair Method and Device,” the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The present disclosure relates to the field of data processing technology, and in particular to a data repair method and device. Background Art

[0003] With the development of computer technology, the amount of network information data has become increasingly large, leading to a surge in demand for massive data storage, making distributed data storage systems the mainstream storage technology. Distributed data storage systems are prone to data loss due to storage node failure. Therefore, redundancy technologies such as replicas or erasure codes can be used to ensure that the entire system's services are not interrupted. Compared to simply copying data, erasure codes have relatively high space utilization. Current solutions for repairing data based on erasure codes result in low system storage efficiency and require a large amount of disk I / O during the repair process, resulting in high data repair bandwidth and a relatively complex process. Therefore, there is an urgent need for a data repair method that can reduce the amount of redundant data, improve system storage efficiency, and reduce bandwidth consumption during data repair.

[0004] Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a data repair method. One or more embodiments of the present disclosure also relate to a data repair apparatus, a data repair system, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a data repair method, comprising:

[0007] Construct an initial check matrix for data repair for a distributed storage node cluster;

[0008] Updating the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy;

[0009] In the case where there is a faulty node in the distributed storage node cluster, selecting a matrix element associated with the faulty node in the check matrix;

[0010] Determine a target node in the distributed storage node cluster based on the matrix elements, and obtain node storage data corresponding to the target node;

[0011] The repair data of the faulty node is calculated according to the matrix elements and the node storage data, and the repair data is uploaded to the faulty node.

[0012] According to a second aspect of an embodiment of the present disclosure, there is provided a data repair device, comprising:

[0013] A construction module is configured to construct an initial check matrix for data repair for a distributed storage node cluster;

[0014] An update module is configured to update a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies a full rank condition of the matrix full rank strategy;

[0015] A selection module is configured to select, in the case where there is a faulty node in the distributed storage node cluster, a matrix element associated with the faulty node in the check matrix;

[0016] an acquisition module configured to determine a target node in the distributed storage node cluster based on the matrix elements, and acquire node storage data corresponding to the target node;

[0017] A calculation module is configured to calculate repair data of the faulty node according to the matrix elements and the node storage data, and upload the repair data to the faulty node.

[0018] According to a third aspect of an embodiment of the present disclosure, there is provided a data repair system, the system comprising a server and a distributed storage node cluster;

[0019] The server constructs an initial check matrix for data repair for the distributed storage node cluster; updates the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; in the case where there is a faulty node in the distributed storage node cluster, selects a matrix element associated with the faulty node in the check matrix; determines a target node in the distributed storage node cluster based on the matrix element, and sends a data acquisition request to the target node;

[0020] The target node, in response to the data acquisition request, sends the node storage data to the server;

[0021] The server calculates repair data of the faulty node according to the matrix elements and the node storage data, and uploads the repair data to the faulty node.

[0022] According to a fourth aspect of an embodiment of the present disclosure, there is provided a computing device, including:

[0023] memory and processor;

[0024] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data repair method are implemented.

[0025] According to a fifth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data repair method are implemented.

[0026] According to a sixth aspect of an embodiment of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data repair method.

[0027] The present disclosure provides a data repair method, comprising constructing an initial check matrix for data repair for a distributed storage node cluster; updating a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; in the case where there is a faulty node in the distributed storage node cluster, selecting a matrix element associated with the faulty node in the check matrix; determining a target node in the distributed storage node cluster based on the matrix element, and obtaining node storage data corresponding to the target node; calculating repair data of the faulty node based on the matrix element and the node storage data, and uploading the repair data to the faulty node.

[0028] An embodiment of the present disclosure implements the construction of an initial check matrix for data repair for a distributed storage node cluster, and updates the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, thereby realizing the construction of a repair matrix for a distributed storage node cluster with any number of nodes, and the constructed check matrix satisfies the MDS (maximum distance separable) property, thereby improving the data repair capability. In the case of a faulty node in the distributed storage node cluster, the repair data of the faulty node can be calculated using the matrix elements in the check matrix and the node storage data of the target node, thereby realizing the repair function of the faulty node while reducing the node repair bandwidth and hard disk consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG1 is a schematic diagram of a scenario of a data repair method provided by an embodiment of the present disclosure;

[0030] FIG2 is a flow chart of a data repair method provided by an embodiment of the present disclosure;

[0031] FIG3 is a schematic diagram of an initial check matrix in a data repair method provided by an embodiment of the present disclosure;

[0032] FIG4 is a flowchart of a data repair method according to an embodiment of the present disclosure;

[0033] FIG5 is a schematic diagram of a check matrix in a data repair method provided by an embodiment of the present disclosure;

[0034] FIG6 is a schematic structural diagram of a data repair device provided by an embodiment of the present disclosure;

[0035] FIG7 is a schematic structural diagram of a data repair system provided by an embodiment of the present disclosure;

[0036] FIG8 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.

[0038] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0039] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0040] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0041] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0042] Distributed storage node cluster: A distributed storage system is a system cluster that stores large amounts of data in a dispersed manner across multiple storage nodes.

[0043] Erasure coding: Erasure coding is a fault-tolerant coding technology that divides stored data into shards and uses a certain checksum calculation method to generate k+r copies of data from k copies of the original data. The system can then restore the original data using any k of the k+r copies. This allows the system to recover the original data even if some data is lost.

[0044] Maximum distance separable code: A maximum distance separable code (MDS code) satisfies the MDS property, exhibits strong fault tolerance, and achieves optimal storage efficiency. This is known as a k+r code, which can tolerate the loss of any k nodes and still recover data.

[0045] With the development of social networks, mobile payments, and online audio and video platforms, people are increasingly active online. These activities generate vast amounts of data, which is growing exponentially. To address the storage and analysis of this massive data, a large number of storage nodes are typically connected via a network to form a large-scale distributed storage system. Node failure is a common occurrence in large-scale distributed storage systems. Distributed storage systems typically use replication and erasure coding techniques to introduce redundant data, ensuring that data can be maintained even in the event of node failure. RS (Reed-Solomon) codes are a type of maximum distance separable code (MDS code) widely used in distributed storage systems due to their excellent storage efficiency. When a node fails, the system needs to promptly reconstruct the data stored on the failed node and store it on a new node. This process is called node repair. To achieve node repair, the new node must download data from other nodes (also called helper nodes) to solve the required data. The total amount of data transmitted from all helper nodes to the new node is called the repair bandwidth. The repair bandwidth of traditional RS codes is k times the amount of data stored on a single node. Therefore, locally repairable codes (LRC) and minimum storage regenerating codes (MSR) are often used to reduce the repair bandwidth. LRC lacks the MDS property, resulting in suboptimal system storage efficiency. MSR codes require exponential subpacketization levels, which not only complicates the management of the system's original data but also introduces a large amount of non-continuous data reads and writes during the repair process on some nodes, consuming significant disk I / O. This causes the actual performance of MSR codes to fall far short of theoretical expectations.

[0046] Based on this, the present disclosure provides a data repair method. The present disclosure also relates to a data repair device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0047] Refer to Figure 1, which is a scenario diagram of a data repair method provided according to an embodiment of the present disclosure, wherein a distributed storage node cluster includes multiple storage nodes, and each storage node stores its corresponding storage data. An initial check matrix for data repair is constructed for the distributed storage node cluster, and the target block matrix in the initial check matrix is ​​updated according to the matrix full rank strategy to obtain the check matrix H. The matrix full rank strategy is used to make any target block matrix in the check matrix full rank, ensuring that the check matrix conforms to the MDS property. In the event of a faulty node in the distributed storage node cluster, the matrix elements associated with the faulty node are selected from the check matrix, and based on the node storage data of the remaining normally working target nodes in the distributed storage node cluster, data repair is performed on the faulty node, the repair data corresponding to the faulty node is calculated, and the repair data is uploaded to the faulty node, thereby reconstructing the lost data corresponding to the faulty node.

[0048] Referring to FIG. 2 , FIG. 2 shows a flow chart of a data repair method provided according to an embodiment of the present disclosure, which specifically includes the following steps.

[0049] Step 202: Construct an initial check matrix for data repair for the distributed storage node cluster.

[0050] A distributed storage node cluster can be understood as a node cluster that stores data on multiple storage nodes. To prevent data loss due to storage node failure, which could render the distributed storage node cluster's corresponding services inoperable, a check matrix needs to be constructed for the distributed storage node cluster. The initial check matrix can be understood as the check matrix used to repair data on the distributed storage nodes. However, some elements in the initial check matrix need to be updated. After the update, the check matrix is ​​obtained, and the check matrix can be used to repair data on faulty distributed storage nodes.

[0051] In practical applications, in order to construct a parity check matrix for data repair in a distributed storage node cluster, an initial parity check matrix must be constructed first. The initial parity check matrix must be determined based on the number of distributed storage nodes.

[0052] In a specific embodiment of the present disclosure, an initial check matrix is ​​constructed for a distributed storage node cluster, and the initial check matrix needs to be further updated according to relevant information of the distributed storage node cluster to obtain a check matrix for data repair corresponding to the distributed storage node cluster.

[0053] Furthermore, in order to correctly construct the initial check matrix corresponding to the distributed storage node cluster, it is necessary to obtain the node quantity and coding information of the distributed storage node cluster. Specifically, the initial check matrix for data repair is constructed for the distributed storage node cluster, including: determining the node quantity information and coding information corresponding to the distributed storage node cluster; encoding the storage nodes in the distributed storage node cluster based on the coding information, and constructing the initial check matrix for data repair according to the node quantity information and the coding results.

[0054] Among them, the node quantity information can be understood as the number of nodes corresponding to the storage nodes in the distributed storage node cluster, and the coding information can be understood as the information encoded for the distributed storage nodes when the matrix is ​​constructed. The coding information includes the number of subpackets, the number of information blocks, the number of check blocks, etc.

[0055] In practical applications, the matrix parameters (n, k = nr, l) are constructed, where n is the number of nodes in the distributed storage node cluster, k represents the number of data blocks, r represents the number of check blocks, and l represents the number of equal divisions of each data block when encoding. For example, for a 1MB data block, l = 4, it is divided into 4 256KB data blocks during encoding, which corresponds to 4 columns of matrix elements in the matrix. In one embodiment of the present disclosure, the size of the constructed initial check matrix is ​​rl×nl, which can also be regarded as a block matrix of size r×n, with each block of size l×l, as shown in Figure 3. Figure 3 is a schematic diagram of the initial check matrix in a data repair method provided in one embodiment of the present disclosure, wherein the first l rows contain n l×l unit block matrices, the lth row is row 0, the rlth row is row rl-1, and the same applies to the columns. λ0,λ1……λ n-1 are n different non-zero elements in the finite field F.

[0056] In a specific embodiment of the present disclosure, the node number information corresponding to the distributed storage node cluster is determined to be n=8, the coding information is determined to be k=5, l=2, and the initial node check matrix corresponding to the distributed storage node is constructed based on the above information. Subsequently, it is necessary to update the elements of the initial node check matrix to obtain the check matrix corresponding to the distributed storage node.

[0057] Based on this, an initial check matrix of corresponding size is constructed by using the node quantity information and coding information corresponding to the distributed storage node cluster, which facilitates the subsequent construction of a check matrix for data repair.

[0058] Furthermore, in the process of constructing the initial check matrix, it is also necessary to encode the data blocks corresponding to each node. Specifically, the storage nodes in the distributed storage node cluster are encoded based on the encoding information, including: determining the data blocks corresponding to the storage nodes in the distributed storage node cluster; encoding the data blocks based on the encoding information to obtain the encoding results.

[0059] The data blocks corresponding to storage nodes can be understood as data blocks associated with each storage node in the distributed storage node cluster. These data blocks contain mapping data for the associated node's storage data. Each storage node stores its own data blocks, and some of the data blocks in the corresponding distributed storage nodes serve as check blocks. For example, if there are eight storage nodes in a distributed storage node cluster, five of them store data blocks and three store check blocks.

[0060] In practical applications, encoding the storage nodes in the distributed storage node cluster can be understood as dividing the data block D and the check block P into l equal parts, denoted as D ij or P ij , represents the jth equal part of the i-th data block or the check block, so that the number of columns in the initial check matrix can be determined.

[0061] In a specific embodiment of the present disclosure, data blocks corresponding to storage nodes in a distributed storage node cluster are encoded and divided into two equal parts, and an initial check matrix is ​​constructed based on the encoding results.

[0062] Furthermore, in order to correctly construct the initial check matrix, after determining the coding result, an initial check matrix for data repair can be constructed based on the node quantity information and the coding result, including: determining matrix parameters based on the node quantity information; and constructing an initial check matrix for data repair based on the matrix parameters and the coding result.

[0063] The matrix parameters are the parameters of the initial check matrix, including the number of rows and columns, and the locations of the nonzero elements. The initial check matrix can be constructed based on the matrix parameters and the encoding result.

[0064] In a specific embodiment of the present disclosure, matrix parameters are determined based on the node number information. The matrix parameters determine the number of rows, columns, and positions of non-zero elements of the initial check matrix. The initial check matrix can be constructed based on the matrix parameters and the encoding results. The matrix size of the initial check matrix is ​​6×16.

[0065] Step 204: Update the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy.

[0066] The matrix full rank strategy can be understood as a strategy for updating the initial check matrix. The matrix full rank strategy is used to update the initial check matrix into a check matrix so that any block matrix in the updated check matrix satisfies the full rank condition. The target block matrix can be understood as any r×r block matrix in the initial check matrix.

[0067] In practical applications, updating the target block matrix in the initial parity check matrix according to the matrix full rank strategy can be understood as adding non-zero elements to the elements at specified positions in the initial parity check matrix, so that any target block matrix in the updated parity check matrix satisfies the full rank condition. In specific implementations, the elements in the initial parity check matrix can also be adjusted with matrix full rank as the goal, as long as any target block matrix in the updated parity check matrix satisfies the full rank condition.

[0068] In a specific embodiment of the present disclosure, the target block matrix in the initial check matrix is ​​updated according to the matrix full rank strategy so that the target block matrix satisfies the full rank condition, thereby obtaining the check matrix.

[0069] Based on this, the initial check matrix is ​​updated through the matrix full rank strategy, so that the updated check matrix satisfies the MDS property, and the k+r ratio can be achieved. If no more than r stored data are lost, the data can be recovered.

[0070] Furthermore, updating the target block matrix is ​​actually updating the elements in the target block matrix, so it is necessary to determine the elements that need to be updated in the target block matrix. Specifically, the target block matrix in the initial check matrix is ​​updated according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, including: determining the target block matrix in the initial check matrix according to the matrix full rank strategy, and selecting the initial element in the target block matrix; replacing the initial element with the target element, and determining the check matrix corresponding to the distributed storage node cluster according to the replacement result.

[0071] Among them, the initial element can be understood as the element that needs to be updated in the target block matrix. After the initial element is replaced by the target element, it can be considered that the target block matrix has been updated. After all target block matrices are updated, the check matrix corresponding to the distributed storage node cluster can be determined.

[0072] In practical applications, when updating the target block matrix in the initial check matrix, in order to improve the matrix update speed, the initial elements can be selected without following the format of the target block matrix. Instead, the initial elements that need to be replaced can be determined from the entire initial check matrix. After replacing the initial elements with the target elements, it can be ensured that any target block matrix in the initial check matrix meets the full rank condition, thereby obtaining the check matrix corresponding to the distributed storage node cluster.

[0073] In a specific embodiment of the present disclosure, the target block matrix in the initial check matrix is ​​determined according to the matrix full rank strategy, the initial element a is selected in the target block matrix, and the initial element is replaced with the target element b. After all the initial elements are replaced, the check matrix corresponding to the distributed storage node is obtained according to the replacement result.

[0074] Based on this, by updating the initial check matrix according to the matrix full rank strategy, a check matrix that meets the MDS property can be obtained, thereby improving the fault tolerance of subsequent data repair.

[0075] Furthermore, in order to improve the matrix update speed, the element position of the initial element can be selected in the matrix first. Specifically, the initial element is selected in the target block matrix, including: determining the position information of the element to be updated in the target block matrix according to the matrix full rank strategy; and determining the initial element in the target block matrix based on the position information of the element to be updated.

[0076] The position information of the element to be updated can be understood as the position information of the element to be updated, and the position information of the element to be updated is the relative position information of the element in the matrix. The corresponding initial element can be determined from the target block matrix based on the position information of the element to be updated.

[0077] In specific implementation, the position information of the element to be updated can be determined in the following way: first, n nodes can be evenly divided into g groups, and the number of nodes in each group is in The initial check matrix is ​​divided into groups of l×(l-1) rows from l+1 to r×l rows in order. Each group is responsible for l groups, with a total of ls nodes. Then, for each of the t groups, add non-zero elements ψ i First, determine that the i-th group is currently being processed, 0≤i≤t-1; there are l×(l-1) rows in the i-th group, with l rows as a set, a total of l-1 sets, and the q-th set is currently being processed, 0≤q≤l-2; there are l rows in the q-th set, and the p-th row is currently being processed, 0≤p≤l-1; on the p-th row, non-zero elements need to be added to s columns, and the j-th column is currently being processed, 0≤j≤s-1. Then, according to formula (1), the non-zero element ψ can be determinedi That is, the location of the initial element, formula (1) is as follows:

[0078] Among them, row is the row information of the initial element in the matrix, and col is the column information of the initial element in the matrix.

[0079] Based on this, by for each initial element ψ i , select appropriate non-zero values ​​on the finite field F so that any r×r row block matrix in the updated matrix is ​​full rank, then the check matrix construction is completed, and the check matrix also satisfies the MDS property.

[0080] After obtaining the check matrix, the corresponding check block can be calculated according to the data block and the check matrix. The specific calculation process is: k data blocks and r data blocks are divided into l equal parts, recorded as D ij or P ij . Generate vector v by combining data block and check block, v={D 0,0 …D 0,l-1 ,…D k-1,0 …D k-1,l-1 ,P 0,0 …P 0,l-1 ,…P r-1,0 …P r-1,l-1}, vector v and check matrix Satisfaction relationship Where D is known, P can be solved to obtain the check block and the encoding is completed.

[0081] Step 206: When there is a faulty node in the distributed storage node cluster, select a matrix element associated with the faulty node in the check matrix.

[0082] Among them, when there is a faulty node in the distributed storage node cluster, it can be understood that the storage data corresponding to the distributed storage node cluster has been lost. At this time, the matrix elements associated with the faulty node can be selected from the check matrix, that is, the check coefficient elements corresponding to the faulty node can be obtained from the check matrix.

[0083] In practical applications, the matrix elements associated with the faulty nodes are the check coefficient elements corresponding to the faulty nodes in the check matrix. Since the matrix size of the check matrix is ​​constructed according to the relevant parameter information of the storage nodes in the distributed storage node cluster, the matrix elements in the check matrix have a corresponding relationship with the storage nodes. When it is determined that a storage node has a fault, the matrix element corresponding to the storage node, i.e., the check coefficient element, can be selected from the check matrix.

[0084] Furthermore, in order to correctly repair the data, it is necessary to select the correct matrix elements from the check matrix. Specifically, the matrix elements associated with the faulty node in the check matrix are selected, including: determining the element association information of the faulty node in the correction matrix; and selecting the matrix elements associated with the faulty node in the check matrix according to the element association information.

[0085] The element association information may be understood as the position information of the fault element corresponding to the fault node in the matrix. According to the element association information, the matrix element associated with the fault node may be selected from the check matrix.

[0086] In practical applications, each storage node has a corresponding matrix element in the check matrix. The matrix element corresponding to each storage node can be determined based on the element association information of each storage node.

[0087] In a specific embodiment of the present disclosure, matrix elements associated with the faulty node are selected from a check matrix according to element association information of the faulty node, and data repair can be performed subsequently based on the matrix elements.

[0088] Based on this, through the correspondence between matrix elements in the check matrix and data blocks, and the mapping relationship between data blocks and stored data, data repair can be performed through the check matrix and stored data.

[0089] Step 208: Determine a target node in the distributed storage node cluster based on the matrix elements, and obtain node storage data corresponding to the target node.

[0090] Among them, the target node can be understood as a node that is still working normally in the distributed storage node cluster. According to the matrix elements corresponding to the faulty node, the target node that provides storage data during subsequent data repair can be determined. Therefore, after determining the target node, it is necessary to obtain the node storage data corresponding to the target node.

[0091] In practical applications, the storage data of the faulty node can be treated as an unknown number to be solved. The solution equation is constructed using the matrix elements corresponding to the faulty node and the known storage data and matrix elements of the target node to calculate the repair data corresponding to the faulty node.

[0092] In a specific embodiment of the present disclosure, a target node that is still working normally is selected in a distributed storage node cluster based on the matrix elements of the faulty node to provide known storage data. Subsequently, the corresponding lost data can be calculated based on the known storage data and the matrix elements in the check matrix.

[0093] Furthermore, different numbers of target nodes can be selected according to the number of different faulty nodes, thereby reducing the repair bandwidth. Specifically, the target nodes are determined in the distributed storage node cluster based on the matrix elements, including: determining a node selection strategy based on the quantity information corresponding to the faulty nodes; and determining the target nodes in the distributed storage node cluster according to the node selection strategy and the matrix elements.

[0094] The node selection strategy can be understood as a strategy for selecting target nodes from a distributed storage node cluster. When there are multiple failed nodes, multiple different target nodes can be selected from the distributed storage node cluster. For example, when r out of n nodes fail, k target nodes need to be selected from n nodes, where k = nr. That is, when there are multiple failed nodes, all remaining normally functioning nodes need to be used as target nodes. In this case, the node selection strategy is to use all remaining normally functioning nodes as target nodes. When the failed node is a single node, some target nodes can be selected from the distributed storage node cluster to reduce the repair bandwidth.

[0095] In actual applications, when repairing data on a single-fault node, data is actually provided to each target node. However, some target nodes do not read the data completely, but only read part of the data, thereby reducing the data repair bandwidth.

[0096] In a specific embodiment of the present disclosure, a node selection strategy is determined based on the number information corresponding to the faulty nodes. If the node selection strategy is a full selection strategy, a target node is selected from the distributed storage node cluster according to the full selection strategy and matrix elements.

[0097] Step 210: Calculate the repair data of the faulty node according to the matrix elements and the node storage data, and upload the repair data to the faulty node.

[0098] Repair data can be understood as the lost data obtained after performing data repair calculations on the failed node. Calculating the repair data for the failed node based on the matrix elements and the node's stored data can actually be understood as reconstructing the lost data from the failed node based on the existing stored data and the check matrix.

[0099] In specific implementation, data repair includes two cases: one is repairing multiple faulty nodes, and the other is repairing a single faulty node. In the case of repairing multiple faulty nodes, if there are n nodes and r nodes are faulty, then k readable nodes are determined and a decoding vector u is constructed, such as u = {D 0,0 …D 0,l-1 ,…X k-1,0 …X k-1,l-1 ,P 0,0 …P 0,l-1,…Y r-1,0 …Y r-1,l-1}, set the position data corresponding to the fault node to unknown numbers x and y, because the encoding process ensures Right now If there are unknown data corresponding to r faulty nodes in vector u, the repair data of r faulty nodes can be calculated. In the case of repairing a single faulty node, due to the encoding process, n nodes are evenly divided into g groups, g groups are divided into t groups, and each group has l groups; suppose the faulty node X is in the jth group, belongs to the i=j / lth group, and the p=j%lth group; 0≤j≤g-1,0≤j≤t-1,0≤p≤l-1. Then, according to the following method, the matrix Pick l rows, p rows, and the following l-1 rows, row =l+i×l×(l-1)+c×l+p, c=0,1…l-2. Using the above method, we can construct l systems of equations. For the lost node X, the coefficient matrix in the equation system is reversible, as are the other nodes in the same group. Furthermore, nodes in other groups only need to read one copy of the data (each node reads l equal copies), thereby reducing the data repair bandwidth for a single faulty node.

[0100] In a specific embodiment of the present disclosure, 8 nodes are divided into g = 4 groups, each group has s = 2 nodes, the faulty node is the 0th node, and the 0th and 2nd rows of the check matrix are taken to form a linear equation system, which is shown in formula (2):

[0101] Since the data D stored in node D0 in the above linear equations 0,0 and D 0,1 The corresponding coefficient matrix is full rank, so when downloading data D 1,0 、D 1,1 、D 2,0 、D 3,0 、D 4,0 、P 0,0 、P 1,0 、P 2,0 After that, we can solve for D 0,0 and D 0,1 , the amount of data that the repair node 1 needs to download is 8, which is 80% of the traditional RS code repair bandwidth.

[0102] Furthermore, in order to correctly calculate the repair data, it is necessary to construct a corresponding repair equation, and calculate the repair data of the faulty node based on the matrix elements and the node storage data, including: selecting the matrix repair elements corresponding to the node storage data in the check matrix; constructing the repair equation corresponding to the faulty node based on the matrix elements and the matrix repair elements; and calculating the repair data of the faulty node based on the node storage data and the repair equation.

[0103] Among them, the matrix repair element is the matrix element corresponding to the target node. After determining the corresponding matrix element, a repair equation can be constructed based on the matrix element of the fault node, the matrix repair element of the target node, and the node storage data, and the corresponding repair data can be calculated according to the repair equation.

[0104] In practical applications, after calculating the lost data corresponding to the failed node, repair data can be uploaded to the failed node if the failed node is able to operate normally, so that the failed node can be put back into use. If the failed node is unavailable, or if it is necessary to ensure uninterrupted normal service of the distributed node cluster, a backup node can be used to replace the failed node. Specifically, the method further includes: determining a backup node corresponding to the failed node and uploading the repair data to the backup node.

[0105] Among them, the backup node can be understood as a node used to replace the faulty node. In actual circumstances, when the faulty node cannot be repaired and put into use temporarily, the backup node can be used to replace the faulty node. At this time, the repair data corresponding to the faulty node can be uploaded to the backup node so that the backup node can replace the faulty node to avoid the problem of the corresponding service of the cluster being stopped.

[0106] In one embodiment of the present disclosure, a backup node A1 corresponding to the faulty node A is determined, and repair data is uploaded to the backup node A1, so that the backup node A1 works in place of the faulty node.

[0107] The present disclosure provides a data repair method, comprising constructing an initial check matrix for data repair for a distributed storage node cluster; updating a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; in the case where there is a faulty node in the distributed storage node cluster, selecting a matrix element associated with the faulty node in the check matrix; determining a target node in the distributed storage node cluster based on the matrix element, and obtaining node storage data corresponding to the target node; calculating repair data of the faulty node based on the matrix element and the node storage data, and uploading the repair data to the faulty node.

[0108] The data repair method provided by the present disclosure can achieve the goal of constructing an initial check matrix for data repair for a distributed storage node cluster, and updating the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, thereby achieving the goal of constructing a repair matrix for a distributed storage node cluster with any number of nodes, and the constructed check matrix satisfies the MDS (maximum distance separable) property, thereby improving the data repair capability. In the case of a faulty node in a distributed storage node cluster, the repair data of the faulty node can be calculated using the matrix elements in the check matrix and the node storage data of the target node, thereby achieving the repair function of the faulty node while reducing the node repair bandwidth and hard disk consumption.

[0109] The following further illustrates the data repair method provided by the present disclosure using the application of the data repair method in node repair as an example, in conjunction with FIG4 . FIG4 shows a process flow chart of a data repair method provided by an embodiment of the present disclosure, which specifically includes the following steps.

[0110] Step 402: Determine the node quantity information and encoding information corresponding to the distributed storage node cluster.

[0111] In an implementable manner, it is determined that the node number information n=8, the coding information k=5, and l=2.

[0112] Step 404: Encode the storage nodes in the distributed storage node cluster based on the encoding information, and construct an initial check matrix for data repair according to the node quantity information and the encoding result.

[0113] In one achievable method, a data block corresponding to a storage node in a distributed storage node cluster is determined, and the data block is encoded based on the encoding information to obtain an encoding result. Matrix parameters are determined based on the node number information, and an initial check matrix for data repair is constructed based on the matrix parameters and the encoding result.

[0114] Step 406: Determine a target block matrix in the initial check matrix according to the matrix full rank strategy, and select an initial element in the target block matrix.

[0115] In one feasible manner, position information of elements to be updated where non-zero elements need to be added is determined from the initial check matrix according to the matrix full rank strategy, and the initial element is determined according to the position information of the elements to be updated.

[0116] Step 408: Replace the initial element with the target element, and determine the check matrix corresponding to the distributed storage node cluster according to the replacement result.

[0117] In one feasible manner, the initial element is replaced with the target element, and the check matrix corresponding to the distributed storage node cluster is determined based on the replacement result. The check matrix is ​​shown in FIG5 , which is a schematic diagram of the check matrix in a data repair method provided by the present disclosure.

[0118] Step 410: Determine element association information of the faulty node in the calibration matrix, and select a matrix element associated with the faulty node in the calibration matrix according to the element association information.

[0119] In one possible implementation, the faulty node is node 0, and the matrix element associated with the faulty node is determined based on the element association information of the faulty node in the correction matrix.

[0120] Step 412: Select a matrix repair element corresponding to the node storage data in the check matrix, and construct a repair equation corresponding to the faulty node based on the matrix element and the matrix repair element.

[0121] In one implementable manner, a matrix repair element corresponding to the node storage data of the target node is selected, and a corresponding repair equation is constructed according to the matrix element, the matrix repair element, and the node storage element.

[0122] Step 414: Calculate the repair data of the faulty node based on the node storage data and the repair equation, and upload the repair data to the faulty node.

[0123] In one achievable manner, repair data of the faulty node 0 is calculated based on the node storage data and the repair equation, and the repair data is uploaded to the faulty node 0.

[0124] The present disclosure provides a data repair method that constructs an initial check matrix for data repair for a distributed storage node cluster, and updates the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, thereby enabling the construction of a repair matrix for a distributed storage node cluster with any number of nodes, and the constructed check matrix satisfies the MDS (maximum distance separable) property, thereby improving data repair capabilities. In the event that there is a faulty node in the distributed storage node cluster, the repair data of the faulty node can be calculated using the matrix elements in the check matrix and the node storage data of the target node, thereby achieving the repair function of the faulty node while reducing the node repair bandwidth and hard disk consumption.

[0125] Corresponding to the above method embodiment, the present disclosure also provides a data repair device embodiment. FIG6 shows a schematic diagram of the structure of a data repair device provided by an embodiment of the present disclosure. As shown in FIG6, the device includes:

[0126] A construction module 602 is configured to construct an initial check matrix for data repair for a distributed storage node cluster;

[0127] An updating module 604 is configured to update a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies a full rank condition of the matrix full rank strategy;

[0128] A selection module 606 is configured to select, in the case where there is a faulty node in the distributed storage node cluster, a matrix element associated with the faulty node in the check matrix;

[0129] An acquisition module 608 is configured to determine a target node in the distributed storage node cluster based on the matrix elements, and acquire node storage data corresponding to the target node;

[0130] The calculation module 610 is configured to calculate the repair data of the faulty node according to the matrix elements and the node storage data, and upload the repair data to the faulty node.

[0131] Optionally, the construction module 602 is further configured to determine the node quantity information and coding information corresponding to the distributed storage node cluster; encode the storage nodes in the distributed storage node cluster based on the coding information, and construct an initial check matrix for data repair based on the node quantity information and the coding results.

[0132] Optionally, the construction module 602 is further configured to determine a data block corresponding to a storage node in the distributed storage node cluster; and encode the data block based on the encoding information to obtain an encoding result.

[0133] Optionally, the construction module 602 is further configured to determine matrix parameters according to the node quantity information; and construct an initial check matrix for data repair based on the matrix parameters and the encoding result.

[0134] Optionally, the update module 604 is further configured to determine the target block matrix in the initial check matrix according to the matrix full rank strategy, and select the initial element in the target block matrix; replace the initial element with the target element, and determine the check matrix corresponding to the distributed storage node cluster based on the replacement result.

[0135] Optionally, the updating module 604 is further configured to determine the position information of the element to be updated in the target block matrix according to the matrix full rank strategy; and determine the initial element in the target block matrix based on the position information of the element to be updated.

[0136] Optionally, the selection module 606 is further configured to determine element association information of the faulty node in the correction matrix; and select a matrix element associated with the faulty node in the check matrix according to the element association information.

[0137] Optionally, the acquisition module 608 is further configured to determine a node selection strategy according to the quantity information corresponding to the faulty nodes; and determine a target node in the distributed storage node cluster according to the node selection strategy and the matrix elements.

[0138] Optionally, the calculation module 610 is further configured to select a matrix repair element corresponding to the node storage data in the check matrix; construct a repair equation corresponding to the faulty node based on the matrix element and the matrix repair element; and calculate the repair data of the faulty node based on the node storage data and the repair equation.

[0139] Optionally, the device further includes a backup module configured to determine a backup node corresponding to the faulty node and upload the repair data to the backup node.

[0140] The present disclosure provides a data repair device, including: a construction module, configured to construct an initial check matrix for data repair for a distributed storage node cluster; an update module, configured to update a target block matrix in the initial check matrix according to a matrix full rank strategy, and obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; a selection module, configured to, when a faulty node exists in the distributed storage node cluster, select a matrix element associated with the faulty node in the check matrix; an acquisition module, configured to determine a target node in the distributed storage node cluster based on the matrix element, and obtain node storage data corresponding to the target node; a calculation module, configured to calculate repair data of the faulty node based on the matrix element and the node storage data, and upload the repair data to the faulty node. This paper constructs an initial check matrix for data repair in a distributed storage node cluster. The target block matrix in the initial check matrix is ​​updated according to the matrix full-rank strategy to obtain the check matrix corresponding to the distributed storage node cluster. This allows the construction of a repair matrix for distributed storage node clusters with any number of nodes. The constructed check matrix satisfies the MDS (maximum distance separable) property, improving data repair capabilities. In the event of a faulty node in a distributed storage node cluster, the repair data for the faulty node can be calculated using the matrix elements in the check matrix and the node storage data of the target node. This allows the repair function of the faulty node to be achieved while reducing node repair bandwidth and hard disk consumption.

[0141] The above is a schematic diagram of a data repair device according to this embodiment. It should be noted that the technical solution of the data repair device and the technical solution of the above-mentioned data repair method are based on the same concept. For details not described in detail in the technical solution of the data repair device, please refer to the description of the technical solution of the above-mentioned data repair method.

[0142] Corresponding to the above method embodiments, the present disclosure also provides a data repair system embodiment, and Figure 7 shows a schematic diagram of the structure of a data repair system provided by an embodiment of the present disclosure. As shown in Figure 7, the system includes a server 702 and a distributed storage node cluster 704.

[0143] The server 702 constructs an initial check matrix for data repair for the distributed storage node cluster; updates the target block matrix in the initial check matrix according to the matrix full rank strategy to obtain the check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; in the case where there is a faulty node in the distributed storage node cluster, selects a matrix element associated with the faulty node in the check matrix; determines a target node in the distributed storage node cluster based on the matrix element, and sends a data acquisition request to the target node;

[0144] The target node, in response to the data acquisition request, sends the node storage data to the server;

[0145] The server 704 calculates repair data of the faulty node according to the matrix elements and the node storage data, and uploads the repair data to the faulty node.

[0146] The above is a schematic diagram of a data repair system according to this embodiment. It should be noted that the technical solution of the data repair system and the technical solution of the above-mentioned data repair method are based on the same concept. For details not described in detail in the technical solution of the data repair system, please refer to the description of the technical solution of the above-mentioned data repair method.

[0147] Figure 8 shows a block diagram of a computing device 800 according to one embodiment of the present disclosure. Components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.

[0148] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0149] In one embodiment of the present disclosure, the aforementioned components of the computing device 800 and other components not shown in FIG8 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG8 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0150] Computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 800 may also be a mobile or stationary server.

[0151] The processor 820 is configured to execute the following computer executable instructions, which implement the steps of the above-mentioned data repair method when executed by the processor.

[0152] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above-mentioned data repair method are of the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned data repair method.

[0153] An embodiment of the present disclosure further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned data repair method.

[0154] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above-mentioned data repair method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-mentioned data repair method.

[0155] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data repair method.

[0156] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above-mentioned data repair method are based on the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the above-mentioned data repair method.

[0157] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0158] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0159] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.

[0160] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0161] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A data repair method, comprising: Constructing an initial check matrix for data repair for a distributed storage node cluster; Updating a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full rank condition of the matrix full rank strategy; When there are faulty nodes in the distributed storage node cluster, selecting matrix elements associated with the faulty nodes in the check matrix; Determining target nodes in the distributed storage node cluster based on the matrix elements and obtaining the node storage data corresponding to the target nodes; Calculating repair data for the faulty nodes according to the matrix elements and the node storage data, and uploading the repair data to the faulty nodes.

2. The method according to claim 1, wherein constructing an initial check matrix for data repair for a distributed storage node cluster comprises: Determining the number of nodes information and coding information corresponding to the distributed storage node cluster; Encoding the storage nodes in the distributed storage node cluster based on the coding information, and constructing an initial check matrix for data repair according to the number of nodes information and the coding result.

3. The method according to claim 2, wherein encoding the storage nodes in the distributed storage node cluster based on the coding information comprises: Determining the data blocks corresponding to the storage nodes in the distributed storage node cluster; Encoding the data blocks based on the coding information to obtain a coding result.

4. The method according to claim 2, wherein constructing an initial check matrix for data repair according to the number of nodes information and the coding result comprises: Determining matrix parameters according to the number of nodes information; Constructing an initial check matrix for data repair based on the matrix parameters and the coding result.

5. The method according to any one of claims 1-4, wherein updating a target block matrix in the initial check matrix according to a matrix full rank strategy to obtain a check matrix corresponding to the distributed storage node cluster comprises: Determining a target block matrix in the initial check matrix according to the matrix full rank strategy, and selecting initial elements in the target block matrix; Replacing the initial elements with target elements, and determining a check matrix corresponding to the distributed storage node cluster according to the replacement result.

6. The method according to claim 5, wherein selecting initial elements in the target block matrix comprises: Determining the position information of elements to be updated in the target block matrix according to the matrix full rank strategy; Determining initial elements in the target block matrix based on the position information of the elements to be updated.

7. The method according to any one of claims 1-6, wherein selecting matrix elements associated with the faulty nodes in the check matrix comprises: Determining the element association information of the faulty nodes in the correction matrix; Selecting matrix elements associated with the faulty nodes in the check matrix according to the element association information.

8. The method according to any one of claims 1-7, wherein determining target nodes in the distributed storage node cluster based on the matrix elements comprises: Determine a node selection strategy according to the quantity information corresponding to the faulty node; Determine a target node in the distributed storage node cluster according to the node selection strategy and the matrix element.

9. The method according to any one of claims 1-8, wherein calculating repair data of the faulty node according to the matrix element and the data stored in the node comprises: Select a matrix repair element corresponding to the data stored in the node in the check matrix; Construct a repair equation corresponding to the faulty node according to the matrix element and the matrix repair element; Calculate the repair data of the faulty node according to the data stored in the node and the repair equation.

10. The method according to any one of claims 1-9, wherein the method further comprises: Determine a standby node corresponding to the faulty node, and upload the repair data to the standby node.

11. A data repair device, comprising: A construction module configured to construct an initial check matrix for data repair for a distributed storage node cluster; An update module configured to update a target block matrix in the initial check matrix according to a matrix full-rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full-rank condition of the matrix full-rank strategy; A selection module configured to select a matrix element associated with the faulty node in the check matrix when there is a faulty node in the distributed storage node cluster; An acquisition module configured to determine a target node in the distributed storage node cluster based on the matrix element, and acquire the data stored in the node corresponding to the target node; A calculation module configured to calculate the repair data of the faulty node according to the matrix element and the data stored in the node, and upload the repair data to the faulty node.

12. A data repair system, the system comprising a server and a distributed storage node cluster; The server constructs an initial check matrix for data repair for the distributed storage node cluster; Update a target block matrix in the initial check matrix according to a matrix full-rank strategy to obtain a check matrix corresponding to the distributed storage node cluster, wherein the target block matrix satisfies the full-rank condition of the matrix full-rank strategy; when there is a faulty node in the distributed storage node cluster, select a matrix element associated with the faulty node in the check matrix; determine a target node in the distributed storage node cluster based on the matrix element, and send a data acquisition request to the target node; The target node, in response to the data acquisition request, sends the data stored in the node to the server; The server calculates the repair data of the faulty node according to the matrix element and the data stored in the node, and uploads the repair data to the faulty node.

13. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

15. A computer program that, when executed on a computer, causes the computer to execute the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Minimum storage regeneration code encoding method and system for improving data restoration performance

    CN110750382A

  • Method for constructing and repairing maximum distance separable code and related device

    CN115858230A

  • Non-binary LDPC code construction method and device, equipment and storage medium

    CN115865102A

  • Code generating, coding and decoding method and device

    CN117271199A

  • Efficient repair of erasure coded data based on coefficient matrix decomposition

    US20180121286A1