Data reconstruction method, medium, device, program product and distributed storage system

Through asynchronous transmission and partial reconstruction, the problem of low data reconstruction efficiency in distributed storage systems is solved, and more efficient data block transmission and data recovery are achieved.

CN120704588APending Publication Date: 2025-09-26HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410349751.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the data reconstruction efficiency of distributed storage systems is low, mainly because the throughput is limited by the synchronous transmission of data blocks by each data node. In particular, when the devices are heterogeneous or there are device failures, the cluster bandwidth cannot be effectively utilized.

Method used

Data blocks are transmitted asynchronously. Each data node sends the next data block immediately after the previous data block is sent, and partially reconstructs the data based on the currently obtained data block at the receiving end until all data blocks are obtained and the complete data is reconstructed.

Benefits of technology

The data block transmission efficiency and data reconstruction efficiency are improved, especially when the device performance is inconsistent or fails, the cluster bandwidth can still be effectively utilized, and the stability and throughput of data recovery are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704588A_ABST
    Figure CN120704588A_ABST
Patent Text Reader

Abstract

A data reconstruction method, a medium, a device, a program product and a distributed storage system are applied to the distributed storage system comprising a plurality of data nodes, each data node stores a data block used for reconstructing original data, and each piece of original data is obtained by reconstructing a plurality of data blocks respectively stored on the plurality of data nodes. The method comprises the following steps: acquiring a plurality of data blocks asynchronously sent by a plurality of data nodes at present; the data node responds to the completion of sending of the previous data block on the node and sends the current data block on the node; performing partial data reconstruction based on each data block and the intermediate reconstruction result of the original data reconstructed by the data block to obtain a current reconstruction result of the original data reconstructed by the data block; and for any part of original data, if the number of the obtained data blocks used for reconstructing the original data is smaller than the minimum number of the data blocks required for reconstructing the original data, returning to the step of obtaining the plurality of data blocks sent by the plurality of data nodes respectively at present.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technology, and in particular to a data reconstruction method, medium, device, program product, and distributed storage system. Background Art

[0002] In a distributed storage system, to improve data fault tolerance, multiple data blocks used to reconstruct the original data can be stored on multiple data nodes in the distributed storage system. When the original data is lost or damaged, data blocks can be read from multiple data nodes and reconstructed based on the read data blocks. In related art, each data node synchronously sends the data blocks required to reconstruct the same piece of original data. Only after all data blocks required to restore the same piece of original data have been transmitted will each data node continue to send the data blocks required to reconstruct the next piece of original data. This approach results in low data throughput, resulting in low data reconstruction efficiency. Summary of the Invention

[0003] In a first aspect, an embodiment of the present disclosure provides a data reconstruction method, which is applied to a distributed storage system including multiple data nodes, each data node stores a data block for reconstructing original data, and each copy of original data is reconstructed from multiple data blocks respectively stored on multiple data nodes; the method includes: obtaining multiple data blocks currently sent asynchronously by the multiple data nodes; performing data reconstruction based on each data block and the intermediate reconstruction result of the original data reconstructed by the data block to obtain the current reconstruction result of the original data reconstructed by the data block; the intermediate reconstruction result of the original data is obtained by partially reconstructing the data blocks historically sent by at least one data node for reconstructing the original data; for any copy of original data, if the number of data blocks obtained for reconstructing the original data is less than the minimum number of data blocks required to reconstruct the original data, returning to the step of obtaining the multiple data blocks currently sent by the multiple data nodes respectively.

[0004] In a second aspect, an embodiment of the present disclosure provides a data reconstruction method, which is applied to a distributed storage system including multiple data nodes, each of which stores data blocks for reconstructing original data, and each copy of original data is reconstructed from multiple data blocks stored on multiple data nodes respectively; the method includes: if a data block for reconstructing the original data is stored in this node, sending the data block to the reconstruction node; wherein this node and other data nodes in the distributed storage system asynchronously send data blocks for reconstructing the same copy of original data, so that each time the reconstruction node receives a data block, it partially reconstructs the original data based on the currently received data block until the original data is reconstructed; and if this node is used to receive the reconstructed original data, obtaining the reconstructed original data returned by the reconstruction node.

[0005] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the present disclosure.

[0006] In a fourth aspect, an embodiment of the present disclosure provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any embodiment of the present disclosure when executing the program.

[0007] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in any embodiment of the present disclosure when executed by a processor.

[0008] In a sixth aspect, an embodiment of the present disclosure provides a distributed storage system, comprising a metadata node and multiple data nodes, wherein the multiple data nodes store data blocks for reconstructing original data, and each copy of original data is reconstructed from multiple data blocks respectively stored on multiple data nodes; wherein the metadata node is used to distribute data reconstruction tasks to reconstruction nodes among the multiple data nodes; and the data node is used to execute the method described in the first aspect in response to the data reconstruction task when the current node is a reconstruction node, and execute the method described in the second aspect when the current node is a data node other than the reconstruction node.

[0009] In the disclosed embodiment, each data node uses an asynchronous method to send the data blocks used to reconstruct the same original data on the node. As long as the previous data block on the node is sent, the current data block can be sent immediately, without having to wait for all the data blocks required to reconstruct the same original data to be sent before continuing to send data blocks. This improves the transmission efficiency of the data blocks and thus improves the efficiency of data reconstruction. Since the data blocks that are sent first cannot reconstruct the complete original data, at the receiving end of the data blocks, each time a data block is obtained, the corresponding original data can be partially reconstructed based on the currently obtained data block to obtain the current reconstruction result corresponding to the original data. After obtaining all the data blocks required to reconstruct the original data, the complete original data can be reconstructed.

[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings herein are incorporated into the specification and constitute a part of the present disclosure. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0012] Figure 1 Schematic diagram of a distributed storage system according to an embodiment of the present disclosure.

[0013] Figure 2 It is a schematic diagram of the data reconstruction process of an embodiment of the present disclosure.

[0014] Figure 3 It is a schematic diagram of the synchronous transmission mode of an embodiment of the present disclosure.

[0015] Figure 4 4 is a flow chart of a data reconstruction method according to an embodiment of the present disclosure.

[0016] Figure 5 Schematic diagram of the asynchronous transmission mode of the embodiment of the present disclosure.

[0017] Figure 6 It is a schematic diagram of the reconstruction process and memory allocation of an embodiment of the present disclosure.

[0018] Figure 7 Schematic diagram of the difference between data block transmission progress of data nodes according to an embodiment of the present disclosure.

[0019] Figure 8 is a flowchart of a data reconstruction method according to another embodiment of the present disclosure.

[0020] Figure 9 is a schematic diagram of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0022] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the" and "the" used in this disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein means any combination of at least two of any one or more of a plurality of.

[0023] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0024] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure are further described in detail below with reference to the accompanying drawings.

[0025] A distributed storage system is a system that stores data on multiple independent data nodes to achieve distributed storage and management of data. The distributed storage system needs to ensure that the state presented to the outside is consistent and that state rollback does not occur. Figure 1 As shown, the distributed storage system 10 includes a metadata node (MetaNode) 102 and multiple data nodes (ChunkServers) 104. Metanode 102 is a centralized metadata storage node, typically used to store data status information, storage location information, data length information, etc. Data node 104 is a node in the distributed storage system 10 used to store data, and is generally responsible for operations such as writing, storing, reading, and deleting data copies.

[0026] In the distributed storage system 10, to improve data fault tolerance, multiple data blocks used to reconstruct original data can be stored on multiple data nodes 104 of the distributed storage system 10. To improve the reliability of data in the distributed storage system 10, when some data is lost or damaged, the distributed storage system 10 can obtain data blocks from multiple data nodes 104 to reconstruct the lost or damaged data, and write the reconstructed data to other data nodes 104, thereby ensuring data security.

[0027] One data reconstruction method is a data reconstruction method based on erasure codes. In the data reconstruction method based on erasure codes, a section of data is divided into multiple data blocks of equal length, and the multiple data blocks include original data blocks and at least one check block. The original data block includes the original data content that needs to be stored or transmitted, and can be obtained by dividing the original data into blocks according to a certain division method. The check block is redundant check information calculated on the data block, which is used to detect and recover errors or losses in the data block. The check block does not contain the original user data, but is redundant information calculated based on the content of the data block. When any one or more data blocks are lost, the lost data blocks can be reconstructed and recovered based on other data blocks. Different data blocks can be stored on different data nodes 104. As Figure 2As shown, the example of data nodes 104 being equal to 4 is used for illustration. For ease of distinction, the different data nodes 104 are respectively denoted as data node CS1, data node CS2, data node CS3, and data node CS4. Data node CS1 is used to store data block A1, data node CS2 is used to store data block A2, data node CS3 is used to store data block A3, and data node CS4 is used to store data block A4. Data blocks A1, A2, and A3 can be original data blocks, and data block A4 can be a check block. If some of the data blocks A1, A2, and A3 are lost or damaged, the data blocks can be reconstructed using other data blocks. For example, if data block A1 is damaged, data block A1 can be reconstructed using data blocks A2, A3, and A4. Specifically, data nodes CS2, CS3, and CS4 can respectively send data block A2, data block A3, and data block A4 to reconstruction node Y. Reconstruction node Y can reconstruct data based on the received data block A2, data block A3, and data block A4 to obtain reconstructed data block A1, and send the reconstructed data block A1 to data node CS1 for storage. In some embodiments, when the number of original data blocks is k and the number of check blocks is m, the total number of data blocks and check blocks required to reconstruct any original data block is k. Assuming k is 8 and m is 3, the total number of data blocks and check blocks required to reconstruct any original data block is 8. For example, data reconstruction can be performed using 7 original data blocks and 1 check block, or using 6 original data blocks and 2 check blocks.

[0028] Related technologies generally read data blocks on multiple data nodes synchronously to perform data reconstruction. After all data blocks used to reconstruct the same original data are read, the data blocks used to reconstruct the next original data will be read. This method is called synchronous mode. In synchronous mode, the throughput is limited by the data node with the slowest transmission speed among multiple data nodes. When there are situations such as device heterogeneity, partial device failure, high pressure on the device front end, and a large amount of data to be copied from the device, the transmission speed of some data nodes in the distributed storage system remains at a low level for a long time, and the cluster bandwidth of the distributed storage system cannot be effectively utilized. There are disadvantages such as low data reconstruction throughput and poor throughput stability.

[0029] like Figure 3As shown, it is assumed that the number of data nodes 104 is 4, and each data node 104 is respectively denoted as CS1, CS2, CS3 and CS4. Each rounded rectangular box represents a data block required for reconstructing a copy of the original data. The original data reconstructed by the data block is shown as the mark in the solid rounded rectangular box. For example, the rounded rectangular box marked with "original data 1" represents the data block used to reconstruct original data 1, the rounded rectangular box marked with "original data 2" represents the data block used to reconstruct original data 2, and so on. The dotted rounded rectangular box indicates that the corresponding data node 104 is currently in an idle waiting state. The height of the rounded rectangular box represents the time required to transmit the data block. t1 to t8 represent different moments. For the sake of convenience, it is assumed that the time interval between adjacent moments is 1s, that is, t1 represents the 1st second, t2 represents the 2nd second, and so on. As can be seen, when reconstructing original data 1, data nodes CS1 and CS3 complete the transmission of the data blocks required for reconstructing original data 1 on their own nodes within the first second, while data nodes CS2 and CS4 complete the transmission of the data blocks required for reconstructing original data 1 on their own nodes within the second second. Therefore, after data nodes CS1 and CS3 complete the transmission of the data blocks required for reconstructing original data 1 on their own nodes, they need to wait for 1 second before continuing to transmit the data blocks required for reconstructing original data 2 on their own nodes. When reconstructing original data 2, data node CS1 requires 3 seconds to complete the transmission of the data blocks required for reconstructing original data 2 on their own nodes, while data nodes CS2, CS3, and CS4 only require 1 second to complete the transmission of the data blocks required for reconstructing original data 2 on their own nodes. Therefore, after data nodes CS2, CS3, and CS4 complete the transmission of the data blocks required for reconstructing original data 2 on their own nodes, they need to wait for 2 seconds before continuing to transmit the data blocks required for reconstructing original data 3 on their own nodes. The transmission method for other data blocks is similar to the above process and will not be repeated here. It can be seen that since each data node 104 is in an idle waiting state for a long time, only the data blocks required for reconstructing three copies of the original data can be transmitted in the first 8 seconds, and the cluster throughput is low.

[0030] Based on this, the embodiment of the present disclosure provides a data reconstruction method, which is applied to a distributed storage system 10 including multiple data nodes 104. Each data node 104 stores a data block for reconstructing original data. Each piece of original data is reconstructed from multiple data blocks stored on multiple data nodes 104. Figure 4 , the method comprising:

[0031] Step S12: obtaining multiple data blocks currently asynchronously sent by multiple data nodes 104;

[0032] Step S14: Reconstructing data based on each data block and the intermediate reconstruction results of the original data reconstructed by the data block to obtain a current reconstruction result of the original data reconstructed by the data block; the intermediate reconstruction result of the original data is obtained by partially reconstructing the data blocks previously sent by at least one data node 104 for reconstructing the original data;

[0033] Step S16: For any piece of original data, if the number of data blocks obtained for reconstructing the original data is less than the minimum number of data blocks required to reconstruct the original data, return to step S12.

[0034] In the disclosed embodiment, each data node 104 transmits data blocks on the node used to reconstruct the same original data asynchronously. Each data node 104 transmits the current data block on the node in response to the completion of the previous data block transmission on the node. There is no need to wait for all data blocks required to reconstruct the same original data to be transmitted before continuing to transmit data blocks. This improves the transmission efficiency of data blocks and, in turn, the efficiency of data reconstruction. At the receiving end of the data blocks, each time a data block is obtained, the corresponding original data is partially reconstructed based on the currently obtained data block to obtain the current reconstruction result corresponding to the original data. After all the data blocks required to reconstruct the original data are obtained, the complete original data can be reconstructed.

[0035] The specific implementation of the embodiment of the present disclosure is described below with examples.

[0036] The method of the disclosed embodiment can be executed by a reconstruction node. The reconstruction node can be selected from multiple data nodes 104. For example, the source data node (the data node where the data block for reconstructing the original data is located) can be determined as the reconstruction node, or the destination data node (the data node for storing the reconstructed original data) can be determined as the reconstruction node. Any other data node can also be determined as the reconstruction node.

[0037] Each piece of original data can be reconstructed based on multiple data blocks, which may include both original data blocks and check blocks. Each data node 104 can store one or more data blocks. During data reconstruction, each data node 104 can send at least some of its data blocks to a reconstruction node for data reconstruction. For ease of description, the following example uses the storage of one data block on each data node 104 as an example.

[0038] In step S12, when data is missing, the metadata node 102 may send a data reconstruction task to the reconstruction node. The reconstruction node may receive the data reconstruction task from the metadata node 102. The data reconstruction task may include information about the original data to be reconstructed and information about each data block used to reconstruct the original data. The original data information includes the storage location and length of the original data, and the data block information includes the storage location and length of the data block. The storage location may include the identification information of the data node 104 storing the original data or data block, as well as the starting storage location of the original data or data block on the corresponding data node 104. For each piece of original data requiring data reconstruction, the data reconstruction task may be used to determine the multiple target data blocks required to reconstruct the original data and determine the candidate data nodes where these multiple target data blocks are located. The reconstruction node may send a data request to each target data node among the candidate data nodes, so that the target data node returns the data block in response to the data request. The reconstruction node may then reconstruct the original data based on the data blocks returned by each target data node and return the reconstructed original data to the data node 104 storing the original data. The data node 104 used to store the reconstructed original data may be the same as or different from the original data node where the original data is located.

[0039] Any data node 104 can send the current data block on the node in response to the completion of sending the previous data block on the node. Figure 5 As shown, taking data node CS1 as an example, in response to the completion of sending the data block for reconstructing original data 1 on this node, data node CS1 can immediately send the data block for reconstructing original data 2 on this node, without having to wait for the completion of sending the data blocks for reconstructing original data 1 on data nodes CS2, CS3, and CS4 before sending the data block for reconstructing original data 2 on this node. Similarly, in response to the completion of sending the data block for reconstructing original data 2 on this node, data node CS1 can also immediately send the data block for reconstructing original data 3 on this node, without having to wait for the completion of sending the data blocks for reconstructing original data 2 on data nodes CS2, CS3, and CS4 before sending the data block for reconstructing original data 3 on this node. The process of data nodes CS2, CS3, and CS4 sending data blocks is similar and will not be repeated here.

[0040] It can be seen that in Figure 5 In the embodiment of the asynchronous mode shown, each data node can immediately transmit the next data block when the transmission of a data block on the node is completed. The time required for each data node CS1, CS2, CS3 and CS4 to complete the transmission of all the data blocks used to reconstruct the original data 1, original data 2 and original data 3 is only 5 seconds, which is much faster than the original data 3. Figure 3The synchronization mode shown requires 8 seconds to transmit all the data blocks used to reconstruct the original data 1, original data 2, and original data 3. Figure 5 The asynchronous mode shown can effectively improve the transmission efficiency of data blocks. In situations such as foreground request fluctuations and HDD random reads, background data transmission is prone to random read latency fluctuations. When random read latency fluctuations occur, the latency for the same data node 104 to read different data blocks is often inconsistent. The latency for different data nodes 104 to read data blocks used to reconstruct the same original data is also often inconsistent. Using the asynchronous mode of the disclosed embodiment can effectively improve the transmission efficiency of data blocks.

[0041] In some embodiments, a cache queue can be set for each data node 104. Multiple data nodes 104 correspond to multiple cache queues one by one. The cache queue corresponding to each data node 104 is used to cache data requests received by the data node 104. The data requests received by the data node 104 are used to request data blocks stored by the data node for reconstructing the original data. Figure 5 Taking the illustrated system as an example, a cache queue corresponding to data node CS1, denoted as cache queue Q1, can be set up to store data requests for data blocks on data node CS1. A cache queue corresponding to data node CS2, denoted as cache queue Q2, can also be set up to store data requests for data blocks on data node CS2. A cache queue corresponding to data node CS3, denoted as cache queue Q3, can also be set up to store data requests for data blocks on data node CS3. A cache queue corresponding to data node CS4, denoted as cache queue Q4, can also be set up to store data requests for data blocks on data node CS4. The reconstruction node can send data requests to multiple data nodes 104. After receiving a data request, each data node 104 can add the data request to its corresponding cache queue. Data node 104 can sequentially read each data request cached in its corresponding cache queue. After each data request is read, it transmits the data block stored on its node to the reconstruction node in response to the data request. After the data block is successfully transmitted, the corresponding data request is deleted from the cache queue. The aforementioned cache queues can be asynchronous queues.

[0042] In some embodiments, each piece of original data is a data block in the data to be reconstructed (referred to as a data block to be reconstructed). The data block to be reconstructed may be a data block of a preset size (e.g., 1 MB). Each data block stored on any data node 104 is sent sequentially by the corresponding data node 104 according to the position of the original data corresponding to each data block in the data to be reconstructed. For example, assuming that the size of the data to be reconstructed is 10 MB and includes 10 1 MB data blocks to be reconstructed, each data node 104 first sends a data block for reconstructing the first 1 MB data block to be reconstructed in the data to be reconstructed, then sends a data block for reconstructing the second 1 MB data block to be reconstructed in the data to be reconstructed, and so on.

[0043] In step S14, since the data blocks used to reconstruct the same original data are not sent synchronously, it is possible that the reconstruction node only obtains part of the data blocks used to reconstruct the original data, resulting in an inability to fully reconstruct the original data. Based on this, it is possible to first perform partial data reconstruction (referred to as partial reconstruction) on the original data based on the obtained data blocks, and store the intermediate reconstruction results obtained by the partial reconstruction. When other data blocks used to reconstruct the original data are subsequently obtained, the original data is further reconstructed until the complete original data is reconstructed. Partial data reconstruction refers to reconstructing a portion of the original data.

[0044] Specifically, when one or more data blocks for reconstructing a certain piece of original data are obtained for the first time, a memory space can be allocated for the original data in the memory to store the intermediate reconstruction result corresponding to the original data. Data reconstruction can be performed based on the data blocks obtained for the first time to reconstruct the original data to obtain the current reconstruction result corresponding to the original data, and the current reconstruction result corresponding to the original data is stored in the memory space corresponding to the original data as the intermediate reconstruction result corresponding to the metadata data. When the data blocks for reconstructing the original data are obtained again, the intermediate reconstruction result corresponding to the original data can be obtained from the memory, and data reconstruction can be performed again based on the data blocks obtained again and the intermediate reconstruction result corresponding to the original data to obtain a new intermediate reconstruction result, and the intermediate reconstruction result in the memory space corresponding to the original data is updated to the new intermediate reconstruction result. And so on, until the complete original data is reconstructed.

[0045] like Figure 6As shown in the figure, when reconstructing original data 1, assuming that data nodes CS3 and CS4 have already received the data blocks used to reconstruct original data 1, while data nodes CS1 and CS2 have not, a memory space can be allocated for original data 1, assuming that this is a memory space between 0MB and 1MB in memory, denoted as memory space 1. Based on the data blocks used to reconstruct original data 1 on data nodes CS3 and CS4, data reconstruction can be performed on original data 1, obtaining the current reconstruction result corresponding to original data 1. This current reconstruction result is then written into memory space 1 as the intermediate reconstruction result corresponding to original data 1. Since data nodes CS3 and CS4 have already sent the data blocks used to reconstruct original data 1, data nodes CS3 and CS4 can continue to send data blocks used to reconstruct original data 2. After obtaining the data blocks used to reconstruct original data 1 on data nodes CS1 and CS2, data reconstruction can be further performed based on the intermediate results in memory space 1 and the data blocks sent by data nodes CS1 and CS2 to obtain the complete original data 1.

[0046] Similarly, at some point later, assuming that data nodes CS1, CS3, and CS4 have already received the data blocks used to reconstruct original data 2, while data node CS2 has not, a memory space can be allocated for original data 2. Assume this is 1MB to 2MB of memory space, denoted as memory space 2. Based on the data blocks used to reconstruct original data 2 on data nodes CS1, CS3, and CS4, data reconstruction of original data 2 can be performed to obtain the current reconstruction result corresponding to original data 2. This current reconstruction result of original data 2 is then written into memory space 2 as the intermediate reconstruction result corresponding to original data 2. Since data nodes CS1, CS3, and CS4 have already sent the data blocks used to reconstruct original data 2, data nodes CS1, CS3, and CS4 can continue to send data blocks used to reconstruct original data 3. After receiving the data blocks used to reconstruct original data 2 on data node CS2, data reconstruction can be further performed based on the intermediate results in memory space 2 and the data blocks sent by data node CS2 to obtain the complete original data 2.

[0047] Similarly, at some point later, assuming that the data blocks used to reconstruct original data 3 have been obtained on data node CS1, but the data blocks used to reconstruct original data 3 have not been obtained on data nodes CS2, CS3, and CS4, a memory space can be allocated in the memory for original data 3. Assume that this is a 2MB to 3MB memory space in the memory, which is recorded as memory space 3. Then, based on the data blocks used to reconstruct original data 3 on data node CS1, data reconstruction of original data 3 can be performed to obtain the current reconstruction result corresponding to original data 3. The current reconstruction result corresponding to original data 3 is then written into memory space 3 as the intermediate reconstruction result corresponding to original data 3.

[0048] In step S16, assuming that the minimum number of data blocks required to reconstruct the original data is k, if the number of data blocks obtained for reconstructing the original data is less than k, it means that the original data cannot be completely reconstructed at present. Therefore, you can return to step S12 to continue obtaining data blocks for reconstructing the original data.

[0049] If the number of data blocks obtained for reconstructing the original data reaches k, it means that the original data can be completely reconstructed. The current reconstruction result of the original data can be determined as the reconstructed original data, and the memory allocated for the original data can be released, thereby reducing memory usage.

[0050] Continue to see Figure 6 Taking original data 1 as an example, if the minimum number k of data blocks required to reconstruct original data 1 is 4, then the data blocks required to reconstruct original data 1 on data nodes CS1, CS2, CS3, and CS4 must be obtained to fully reconstruct original data 1. Assuming that at a certain point in time, the data blocks required to reconstruct original data 1 on data nodes CS1, CS2, and CS3 have already been obtained, the process returns to step S12 and continues to obtain the data blocks required to reconstruct original data 1 on data node CS4. If all the data blocks required to reconstruct original data 1 on data nodes CS1, CS2, CS3, and CS4 have been obtained, then original data 1 can be fully reconstructed and the memory allocated for original data 1 can be released.

[0051] In some embodiments, if the transmission performance of a data node 104 is continuously poor, while the transmission performance of another data node 104 is continuously good, more and more original data may be partially reconstructed. Figure 7Taking two data nodes CS1 and CS2 as an example, data node CS1's transmission performance is consistently good, while data node CS2's transmission performance is consistently poor. At time t3, data node CS2 has only completed transmitting the data block used to reconstruct original data 1, while data node CS1 has already completed transmitting the data block used to reconstruct original data 3. Data node CS2's transmission progress is two data blocks behind that of data node CS1. At time t6, data node CS2 has only completed transmitting the data block used to reconstruct original data 2, while data node CS1 has already completed transmitting the data block used to reconstruct original data 6. Data node CS2's transmission progress is four data blocks behind that of data node CS1. Similarly, at time t8, data node CS2's transmission progress is five data blocks behind that of data node CS1. If data node 1's transmission is not restricted, the difference between data node CS1's and data node CS2's transmission progress will continue to grow. This results in the need to continuously reserve new memory to store the intermediate reconstruction results obtained from the data blocks that data node CS1 has already transmitted, resulting in excessive memory overhead.

[0052] To solve the above problem, the maximum number difference of data blocks sent by multiple data nodes 104 can be obtained. If the maximum number difference is greater than a first preset threshold, stop requesting data blocks from the first data node that has sent the most data blocks among the multiple data nodes 104, so that the first data node stops sending data blocks on the node. Figure 7 Taking data nodes CS1 and CS2 in the example, if data node CS1 has sent the most data blocks among multiple data nodes 104, and data node CS2 has sent the fewest data blocks among multiple data nodes 104, then the number p1 of data blocks sent by data node CS1 and the number p2 of data blocks sent by data node CS2 can be obtained. The maximum difference in number can be recorded as p1-p2. If the value of p1-p2 is greater than a first preset threshold, it indicates that the progress of data block transmission between data nodes CS1 and CS2 differs significantly. To reduce memory usage, data requests can be stopped from data node CS1, that is, data block requests can be stopped from data node CS1, thereby causing data node CS1 to stop sending data blocks. In this way, data node CS2 can gradually catch up with the data block transmission progress of data node CS1, thereby gradually freeing up the memory occupied by the partially reconstructed original data.

[0053] Furthermore, after ceasing to request data blocks from the first data node that has sent the most data blocks among the multiple data nodes 104, if the maximum number difference drops from the first preset threshold to below the second preset threshold, data block requests can be continued from the first data node, so that the first data node can continue to send data blocks on this node, and the second preset threshold is less than the first preset threshold. If the maximum number difference drops from the first preset threshold to below the second preset threshold, it indicates that data node CS2 has gradually caught up with the data block transmission progress of data node CS1 to a certain extent, the memory occupied by the partially reconstructed portion of the original data has been released, and the memory space usage has returned to a tolerable range. Therefore, data requests can be continued to be sent to data node CS1, that is, data block requests can be continued from data node CS1, so that data node CS1 can continue to send data blocks.

[0054] In the above embodiment, the first preset threshold and the second preset threshold may both be fixed values ​​set by the user. Alternatively, the first preset threshold and the second preset threshold may also be set based on the total memory of the data node 104, for example, they may be positively correlated with the total memory of the data node 104.

[0055] See also Figure 8 The present disclosure also provides a data reconstruction method, which is applied to a distributed storage system 10 including multiple data nodes 104. Each of the multiple data nodes 104 stores data blocks for reconstructing original data. Each piece of original data is reconstructed from multiple data blocks stored on the multiple data nodes. The method includes:

[0056] Step S22: If the node stores a data block for reconstructing the original data, the node sends the data block to the reconstruction node; wherein the node and other data nodes in the distributed storage system asynchronously send data blocks for reconstructing the same original data, so that each time the reconstruction node receives a data block, it partially reconstructs the original data based on the currently received data block until the original data is reconstructed; and

[0057] Step S24: If the node is used to receive the reconstructed original data, obtain the reconstructed original data returned by the reconstruction node.

[0058] The method of the embodiment of the present disclosure can be executed by any data node 104 in the distributed storage system 10. In this embodiment, each data node 104 can asynchronously send the data block on the node, and as long as the previous data block on the node is sent, the current data block on the node can be sent. Assume that the i-th data block on the data node 104 is used to reconstruct the i-th original data, and the i+1-th data block is used to reconstruct the i+1-th original data. When any data node 104 sends the i+1-th data block, as long as the i-th data block on the node is sent, the i+1-th data block on the node can be sent, without waiting for other data nodes 104 to send the data blocks used to reconstruct the i-th original data, thereby improving the sending efficiency of the data blocks. After obtaining any data block used to reconstruct the original data, the reconstruction node can perform data reconstruction based on the currently obtained data block and the intermediate reconstruction result of the original data reconstructed by the data block, and obtain the current reconstruction result of the original data reconstructed by the data block. The intermediate reconstruction result of the original data is obtained by performing partial data reconstruction on the data blocks historically received by the reconstruction node (ie, historically sent by each data node 104 ) for reconstructing the original data.

[0059] In some embodiments, the multiple data nodes correspond one-to-one to multiple asynchronous queues, and the asynchronous queue corresponding to each data node is used to cache the data requests received by the data node, and the data requests received by the data node are used to request the data blocks stored on the data node; the method also includes: obtaining the data request sent by the reconstruction node; caching the data request to the asynchronous queue corresponding to this node; sending the data block to the reconstruction node includes: weighting the current data request in the asynchronous queue, and sending the data block requested by the current data request to the reconstruction node.

[0060] In some embodiments, each piece of original data is a data block in the data to be reconstructed, and sending the data block to the reconstruction node includes: sending the data blocks stored on the current node to the reconstruction node in sequence according to the position of the original data corresponding to the data block stored on the current node in the data to be reconstructed.

[0061] The specific details of the embodiments of the present disclosure are detailed in the aforementioned embodiments of the method executed by the reconstruction node, which will not be repeated here.

[0062] An embodiment of the present disclosure further provides a computer device, which includes at least a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any of the aforementioned embodiments when executing the program.

[0063] Figure 91 shows a more specific hardware structure diagram of a computing device provided by an embodiment of the present disclosure. The device may include: a processor 20, a memory 22, an input / output interface 24, a communication interface 26, and a bus 28. The processor 20, the memory 22, the input / output interface 24, and the communication interface 26 are connected to each other within the device via the bus 28.

[0064] The processor 20 can be implemented using a general-purpose central processing unit, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure. The processor 20 may also include a graphics card, such as an Nvidia Titan X graphics card or an 1080Ti graphics card.

[0065] The memory 22 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 22 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented through software or firmware, the relevant program codes are stored in the memory 22 and are called and executed by the processor 20.

[0066] The input / output interface 24 is used to connect input / output modules to enable information input and output. The input / output modules can be configured as components within the device (not shown) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0067] The communication interface 26 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).

[0068] The bus 28 comprises a pathway for transmitting information between the various components of the device, such as the processor 20 , the memory 22 , the input / output interface 24 , and the communication interface 26 .

[0069] It should be noted that although the above device only shows the processor 20, memory 22, input / output interface 24, communication interface 26, and bus 28, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0070] The present disclosure also provides a distributed storage system 10. Figure 1 The distributed storage system 10 includes a metadata node 102 and multiple data nodes 104, and the multiple data nodes 104 store data blocks for reconstructing original data, and each piece of original data is reconstructed from multiple data blocks stored respectively on the multiple data nodes 104; wherein the metadata node 102 is used to distribute data reconstruction tasks to reconstruction nodes among the multiple data nodes 104, and the data node 104 is used to, when the node is a reconstruction node, respond to the data reconstruction task and execute the method executed by the reconstruction node in any embodiment of the present disclosure, as well as the method executed by other data nodes when the node is a data node other than the reconstruction node.

[0071] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in any embodiment of the present disclosure when executed by a processor.

[0072] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any of the aforementioned embodiments when the program is executed by a processor.

[0073] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0074] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments of the present disclosure.

[0075] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer device or entity, or by a product having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0076] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and when implementing the embodiment of the present disclosure, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0077] The above is only a specific implementation of the embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present disclosure.

Claims

1. A data reconstruction method, applied to a distributed storage system comprising multiple data nodes, each data node storing a data block for reconstructing original data, wherein each piece of original data is reconstructed from multiple data blocks stored on the multiple data nodes; the method comprising: Acquire multiple data blocks currently asynchronously sent by the multiple data nodes; Performing data reconstruction based on each data block and an intermediate reconstruction result of the original data reconstructed by the data block to obtain a current reconstruction result of the original data reconstructed by the data block; the intermediate reconstruction result of the original data is obtained by partially reconstructing data blocks historically sent by at least one data node for reconstructing the original data; For any piece of original data, if the number of data blocks obtained for reconstructing the original data is less than the minimum number of data blocks required to reconstruct the original data, return to the step of obtaining multiple data blocks currently sent by the multiple data nodes respectively.

2. The method according to claim 1, further comprising: For any piece of original data, when a data block for reconstructing the original data is first obtained, memory is allocated for the original data; The original data is reconstructed based on the data block obtained for the first time to obtain a current reconstruction result of the original data. The current reconstruction result of the original data is used as an intermediate reconstruction result of the original data and stored in the memory allocated for the original data.

3. The method according to claim 2, further comprising: If the number of data blocks acquired for reconstructing the original data reaches the minimum number of data blocks required for reconstructing the original data, the current reconstruction result of the original data is determined as the reconstructed original data, and the memory allocated for the original data is released.

4. The method according to claim 1, further comprising: Obtaining a maximum number difference of data blocks sent by the plurality of data nodes; If the maximum number difference is greater than a first preset threshold, stop requesting data blocks from a first data node that has sent the most data blocks among the multiple data nodes.

5. The method according to claim 4, further comprising, after stopping requesting data blocks from a first data node that has sent the most data blocks among the plurality of data nodes: If the maximum number difference drops from the first preset threshold to below a second preset threshold, continue to request data blocks from the first data node, and the second preset threshold is less than the first preset threshold.

6. A data reconstruction method, applied to a distributed storage system comprising a plurality of data nodes, each of which stores data blocks for reconstructing original data, wherein each piece of original data is reconstructed from a plurality of data blocks stored on the plurality of data nodes; the method comprising: If the local node stores a data block for reconstructing the original data, the node sends the data block to the reconstruction node; wherein the local node and other data nodes in the distributed storage system asynchronously send data blocks for reconstructing the same original data, so that each time the reconstruction node receives a data block, it partially reconstructs the original data based on the currently received data block until the original data is reconstructed; and If this node is used to receive the reconstructed original data, obtain the reconstructed original data returned by the reconstruction node.

7. The method according to claim 6, wherein the plurality of data nodes correspond one-to-one to the plurality of asynchronous queues, the asynchronous queue corresponding to each data node is used to cache data requests received by the data node, and the data requests received by the data node are used to request data blocks stored on the data node; The method further comprises: Obtaining a data request sent by the reconstruction node; Cache the data request to the asynchronous queue corresponding to this node; The sending the data block to the reconstruction node includes: The current data requests in the asynchronous queue are weighted, and the data blocks requested by the current data requests are sent to the reconstruction node.

8. The method according to claim 6, wherein each piece of original data is a data block in the data to be reconstructed, and sending the data block to the reconstruction node comprises: The data blocks stored on the node are sent to the reconstruction node in sequence according to the positions of the original data corresponding to the data blocks stored on the node in the data to be reconstructed.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 9 when executing the program.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. A distributed storage system comprising a metadata node and a plurality of data nodes, wherein the plurality of data nodes store data blocks for reconstructing original data, and each piece of original data is reconstructed from a plurality of data blocks stored on the plurality of data nodes; wherein: The metadata node is used to distribute data reconstruction tasks to reconstruction nodes among the multiple data nodes; The data node is used to execute the method described in any one of claims 1 to 5 in response to the data reconstruction task when the current node is a reconstruction node, and to execute the method described in any one of claims 6 to 8 when the current node is a data node other than the reconstruction node.

Citation Information

Patent Citations

  • Method for improving erasure code based storage cluster recovery performance

    CN103209210A

  • Distributed storage CEPH based erasure correction code overwriting method

    CN105930103A

  • Storage system, computer-readable recording medium, and system control method

    CN110383251A

  • Data processing method and device of distributed assembly line and storage medium

    CN114428786A

  • Remote mirroring for data storage systems using cloud backup

    US10896200B1