Data synchronization method, medium, computer device and program product
By obtaining and determining the length information of the target erasure code block through the client, the interaction between the metadata node and the data node is reduced, which solves the problem of high communication overhead for metadata node synchronization and improves data synchronization efficiency.
Patent Information
- Application Number
- CN202410551480.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-11-07
AI Technical Summary
In a distributed storage system, metadata nodes need to access multiple data nodes to synchronize the original data length information, resulting in significant communication overhead.
The client obtains the length information of multiple erasure coding blocks and determines the target erasure coding block. It only sends this length information to the metadata node for synchronization, reducing the direct interaction between the metadata node and the data node.
This reduces the communication overhead of metadata nodes and improves the efficiency and performance of data synchronization.
Smart Images

Figure CN120909493A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of distributed storage, and particularly relates to a data synchronization method, a medium, a computer device and a program product. BACKGROUND
[0002] When a client writes data to a distributed storage system, one way is to directly write the erasure code corresponding to the original data to the data node of the distributed storage system, and synchronize the length information of the original data written to the distributed storage system to the metadata node of the distributed storage system. In order to improve the write performance, not every time the erasure code is written, the length information of the original data is synchronized in real time, but when a certain condition is met, the synchronization is performed. In the related technology, when the length information of the original data is synchronized, the metadata node needs to access multiple data nodes, causing the communication overhead of the metadata node to be relatively large. SUMMARY
[0003] Embodiments of the present disclosure provide a data synchronization method, a medium, a computer device and a program product, which are used to solve the problem of large communication overhead of the metadata node when synchronizing the length information of the original data.
[0004] In a first aspect, embodiments of the present disclosure provide a data synchronization method applied to a metadata node of a distributed storage system, multiple data nodes of the distributed storage system are respectively used to store multiple erasure code blocks, any one erasure code block includes at least one erasure code corresponding to at least one piece of original data, respectively, the erasure codes corresponding to the same original data in the multiple erasure code blocks are used to reconstruct the data of the original data, and the metadata node is used to store the length information of the original data written to the distributed storage system. The method comprises: obtaining at least one piece of length information sent by a client connected to the distributed storage system, the at least one piece of length information is used to indicate the length of at least one piece of original data corresponding to at least one erasure code in a target erasure code block, the target erasure code block is the kth longest erasure code block in the multiple erasure code blocks, k is greater than or equal to the minimum number of erasure codes required for reconstructing the original data; determining target length information from the at least one piece of length information, and synchronizing the length information of the original data written to the distributed storage system based on the target length information.
[0005] In a second aspect, the embodiments of the present disclosure provide a data synchronization method, applied to a client connected to a distributed storage system, the distributed storage system comprising a metadata node and a plurality of data nodes, the plurality of data nodes respectively used for storing a plurality of erasure code blocks, any one of the erasure code blocks comprising at least one erasure code corresponding to at least one piece of original data respectively, erasure codes corresponding to a same piece of original data in the plurality of erasure code blocks being used for data reconstruction of the original data, and the metadata node being used for storing length information of the original data written into the distributed storage system; the method comprising: determining a target erasure code block of a kth length in the plurality of erasure code blocks, k being greater than or equal to a minimum number of erasure codes required for reconstructing the original data; and sending at least one piece of length information to the metadata node, so that the metadata node determines target length information from the at least one piece of length information, and synchronizes lengths of the original data written into the distributed storage system based on the target length information; the at least one piece of length information being used for representing lengths of at least one piece of original data corresponding to at least one erasure code in the target erasure code block respectively.
[0006] In a third aspect, the embodiments of the present disclosure provide a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method of any of the embodiments of the present disclosure.
[0007] In a fourth aspect, the embodiments of the present disclosure provide a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implementing the method of any of the embodiments of the present disclosure when executing the program.
[0008] In a fifth aspect, the embodiments of the present disclosure provide a computer program product, comprising a computer program, which, when executed by a processor, implements the method of any of the embodiments of the present disclosure.
[0009] In a sixth aspect, the embodiments of the present disclosure provide a distributed storage system, comprising: a plurality of data nodes respectively used for storing a plurality of erasure code blocks, any one of the erasure code blocks comprising at least one erasure code corresponding to at least one piece of original data respectively, erasure codes corresponding to a same piece of original data in the plurality of erasure code blocks being used for data reconstruction of the original data; and a metadata node used for executing the method of the first aspect; wherein the distributed storage system is connected to a client, and the client is used for executing the method of the second aspect.
[0010] In the embodiment of the present disclosure, the client determines the k-th longest erasure code block (i.e., a target erasure code block) from the plurality of erasure code blocks stored by the plurality of data nodes respectively, and obtains at least one piece of length information, each piece of length information indicating the length of the original data corresponding to one erasure code in the target erasure code block. The client sends the obtained at least one piece of length information to the metadata node, so that the metadata node selects the target length information, and synchronizes the length information of the original data written into the distributed storage system based on the target length information. In the above scheme, the interaction process with each data node is implemented by the client, and the metadata node only needs to determine the target length information from the at least one piece of length information sent by the client to realize synchronization, thereby reducing the communication overhead of the metadata node in the synchronization process.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are incorporated into the specification and constitute a part of the present disclosure, illustrate embodiments consistent with the present disclosure, and together with the specification serve to explain the technical solutions of the present disclosure.
[0013] Figure 1 is a schematic diagram of a distributed storage system according to an embodiment of the present disclosure.
[0014] Figure 2 is a schematic diagram of a process of writing erasure codes by a client into a distributed storage system according to an embodiment of the present disclosure.
[0015] Figure 3 is a schematic diagram of the progress of writing erasure codes into each data node according to an embodiment of the present disclosure.
[0016] Figure 4 is a flowchart of a data synchronization method according to an embodiment of the present disclosure.
[0017] Figure 5 is a schematic diagram of the physical length and the logical length of original data according to an embodiment of the present disclosure.
[0018] Figure 6 is a schematic diagram of splitting original data according to an embodiment of the present disclosure.
[0019] Figure 7 is a schematic diagram of an interaction process between a client, a metadata node and a data node according to an embodiment of the present disclosure.
[0020] Figure 8 is a timing diagram of a data synchronization method according to an embodiment of the present disclosure.
[0021] Figure 9 is a schematic diagram of a data synchronization method according to another embodiment of the present disclosure.
[0022] Figure 10 is a schematic diagram of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The description of the exemplary embodiments is intended to apply to various alternative embodiments as well. It is to be understood that the description of any exemplary embodiment is intended only to be used to interpret and understand the scope of the disclosure. Therefore, the exemplary embodiments are not intended to be exhaustive or to be limited to the precise form disclosed. In some instances, detailed descriptions of well-known methods, ingredients, and equipment are omitted so as not to obscure the disclosure in detail.
[0024] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the description of the present disclosure and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It also will be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. In addition, the term "at least one of' as used herein means any one of or any combination of any two or more of the listed items.
[0025] It should be understood that although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a particular order or hierarchy. These terms are used only to distinguish one from another. For example, a first information can be termed a second information, and, similarly, a second information can be termed a first information, without departing from the scope of the present disclosure. As used herein, the term "if' can be interpreted to mean "when" or "upon" or "in response to determining" taking into account the context in which the term is used.
[0026] In order to enable persons skilled in the art to better understand the technical solutions in the embodiments of the present disclosure, and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be described in further detail below with reference to the drawings.
[0027] The distributed storage system refers to a system that stores data on multiple independent data nodes to realize distributed storage and management of data. The distributed storage system needs to ensure that the state presented to the outside is consistent and does not occur state rollback. For example, the distributed storage system needs to ensure that the data read by the client is consistent and does not occur state rollback. Figure 1As shown, the distributed storage system 10 includes a MetaNode 102 and a plurality of ChunkServers 104. The MetaNode 102 is a centralized metadata storage node, which is usually used to store state information of data, storage location information, length information of data, etc. The ChunkServer 104 is a node in the distributed storage system 10 for storing data, which is usually responsible for operations such as writing, storing, reading, deleting, etc. of data. The distributed storage system 10 can interface with a Client 20, which can write data to each ChunkServer 104 in the distributed storage system 10 to which the Client 20 interfaces, and synchronize length information of the data written to the distributed storage system 10 to the MetaNode 102. In some embodiments, the Client 20 writes data to each ChunkServer 104 in an appending manner, i.e., the Client 20 appends new data to the data already in the ChunkServer 104, instead of modifying the data already in the ChunkServer 104. This manner can effectively reduce data write conflicts, improve write performance, and ensure data order.
[0028] To improve the reliability of data in the distributed storage system 10, the distributed storage system 10 can employ an erasure code technique. On this basis, there are two ways for the Client 20 to write data to the ChunkServer 104, one way is to directly write raw data to the ChunkServer 104, and then the ChunkServer 104 converts the written data copy into an erasure code in the background. The other way is to convert the raw data into an erasure code corresponding to the raw data by the Client, and then write the erasure code to the ChunkServer 104. Each erasure code corresponding to the same raw data can be used to reconstruct the raw data. The latter way is called a direct write erasure code way, which can effectively reduce write amplification compared to the former way.
[0029] In the direct write erasure code way, the Client 20 usually writes the erasure code corresponding to each raw data to the ChunkServer 104 in the order of the raw data to be written. For example, as shown in FIG. 1, the Client 20 writes the erasure code corresponding to the first raw data to the ChunkServer 104a, and then writes the erasure code corresponding to the second raw data to the ChunkServer 104b, and so on. Figure 2As shown, assuming that original data 1, original data 2 and original data 3 are three pieces of data to be written in data in turn, the client 20 first generates the corresponding erasure code for the first piece of data to be written (i.e. original data 1), assuming that the number of erasure codes corresponding to each piece of original data is 3, then the three erasure codes corresponding to original data 1 can be denoted as erasure code 11, erasure code 12 and erasure code 13 respectively, and the client 20 can send the erasure code 11, erasure code 12 and erasure code 13 to the data node 104 for storage. Assuming that the multiple erasure codes corresponding to the same piece of original data are stored in different data nodes 104, in order to facilitate distinction, the data nodes 104 for storing the three erasure codes can be denoted as CS1, CS2 and CS3 respectively. That is, the erasure code 11 is stored in the data node CS1, the erasure code 12 is stored in the data node CS2, and the erasure code 13 is stored in the data node CS3. After generating the erasure code 11, erasure code 12 and erasure code 13, the client 20 can continue to generate the erasure codes corresponding to original data 2, i.e. erasure code 21, erasure code 22 and erasure code 23, and send the erasure code 21, erasure code 22 and erasure code 23 to the data nodes CS1, CS2 and CS3 respectively. After generating the erasure code 21, erasure code 22 and erasure code 23, the client 20 can continue to generate the erasure codes corresponding to original data 3, i.e. erasure code 31, erasure code 32 and erasure code 33, and send the erasure code 31, erasure code 32 and erasure code 33 to the data nodes CS1, CS2 and CS3 respectively. And so on. Among them, the client 20 sends the erasure codes corresponding to the next piece of original data after sending the erasure codes corresponding to a piece of original data, that is, the sending process of the erasure codes corresponding to different pieces of original data is serial. But the process of sending the erasure codes corresponding to a piece of original data and the process of generating the erasure codes corresponding to the next piece of original data can be performed asynchronously through different threads on the client 20.
[0030] Since the write bandwidths of the various data nodes 104 are different, the progress of writing erasure codes to different data nodes 104 is also different. For example, in the example shown in FIG. 4, the write bandwidth of the data node CS1 is greater than that of the data node CS2, and the write bandwidth of the data node CS2 is greater than that of the data node CS3. Therefore, the progress of writing erasure codes to the data node CS1 is faster than that of the data node CS2, and the progress of writing erasure codes to the data node CS2 is faster than that of the data node CS3. Figure 3In the illustrated embodiment, the client 20 has successfully written the erasure code 11, the erasure code 21 and the erasure code 31 to the data node CS1, but has only successfully written the erasure code 12 and the erasure code 22 to the data node CS2, and has only successfully written the erasure code 13 to the data node CS3. Since the erasure codes can be used to reconstruct data, for any original data, if the number of erasure codes successfully written to the data nodes 104 among the erasure codes corresponding to the original data reaches the minimum number of erasure codes required for reconstructing the original data, it is considered that the original data has been successfully written to the distributed storage system 10, and the erasure codes that have not been successfully written can be reconstructed from the erasure codes that have been successfully written. For example, assuming that the number of erasure codes required for reconstructing an original data is greater than or equal to 2, when at least two erasure codes corresponding to a piece of original data are successfully written to the data nodes 104, it is considered that the original data has been successfully written to the distributed storage system 10. In the illustrated embodiment, since the three erasure codes corresponding to the original data 1 have all been successfully written to the data nodes 104, and the two erasure codes corresponding to the original data 2 have all been successfully written to the data nodes 104, it is considered that the original data 1 and the original data 2 have been successfully written to the distributed storage system 10. However, the erasure code corresponding to the original data 3 has only been successfully written to one data node 104, and thus it is considered that the original data 3 has not been successfully written to the distributed storage system 10. Figure 3
[0031] In addition to writing the erasure codes corresponding to the original data to the data nodes 104 of the distributed storage system 10, the client 20 can also synchronize the length information of the original data written to the distributed storage system 10 to the metadata node 102 of the distributed storage system 10. There are two ways to synchronize the length information of the data to the metadata node 102. One way is to synchronize the length information of each piece of original data to the metadata node 102 every time the erasure code corresponding to the original data is written. The other way is to synchronize the length information of each piece of original data that has been written but not yet synchronized to the metadata node 102 together only when a certain condition is met. The above-mentioned condition can be that the length of the original data that has been written but not yet synchronized reaches a preset length, a certain data node 104 fails to write, the client 20 currently writing data to the distributed storage system 10 changes, etc. Compared to the former way, the latter way can effectively reduce the write amplification.
[0032] In the example of triggering data synchronization by a specific condition, when the specific condition is not met, there can be some length information of data in the raw data written into the distributed storage system 10 that has not been synchronized to the metadata node 102. However, when the client 20 performing data writing (referred to as a writer) changes, the new writer needs to start writing from the offset after the previous writer writes successfully, and therefore, needs to obtain the length information of the raw data that has been written into the distributed storage system 10. When the client 20 performing data reading (referred to as a reader) reads data, it is necessary to ensure that the length of the data read by the reader cannot be rolled back, and at this time, the reader also needs to obtain the length information of the raw data that has been written into the distributed storage system 10.
[0033] In the related art, when the client 20 has not synchronized the length information of the raw data to the metadata node 102, the metadata node 102 needs to query the length information of the erasure code written on each data node 104 respectively, and then determine the raw data that has been successfully written into the distributed storage system 10 according to the length information of the erasure code queried, so as to determine the length information of the raw data. In the above process, the metadata node 102 needs to access multiple data nodes 104, and the communication overhead of the metadata node 102 is relatively large.
[0034] Based on this, the embodiment of the present disclosure provides a data synchronization method applied to a metadata node 102 of a distributed storage system 10, multiple data nodes 104 of the distributed storage system 10 are respectively used to store multiple erasure code blocks, any one of the erasure code blocks includes at least one erasure code corresponding to at least one raw data respectively, the erasure codes corresponding to the same raw data in the multiple erasure code blocks are used to reconstruct the raw data, and the metadata node 102 is used to store the length information of the raw data written into the distributed storage system 10. Referring to Figure 4 , the method comprises:
[0035] Step S12: Obtain at least one length information sent by a client 20 connected to the distributed storage system, the at least one length information is used to indicate the length of at least one raw data corresponding to at least one erasure code in a target erasure code block, the target erasure code block is the kth longest erasure code block in the multiple erasure code blocks, and k is greater than or equal to the minimum number of erasure codes required for reconstructing the raw data;
[0036] Step S14: Determine target length information from the at least one length information, and synchronize the length information of the raw data written into the distributed storage system 10 based on the target length information.
[0037] The specific scheme of the embodiment of the present disclosure is illustrated below.
[0038] In step S12, the data block composed of the erasure codes written by each data node 104 can be referred to as an erasure code block. One or more erasure codes are included in each erasure code block, and each erasure code in any erasure code block can be written to the data node 104 of the distributed storage system 10 by the client 20 interfacing with the distributed storage system 10. The length information of the erasure code block is the total length of each erasure code included in the erasure code block, and the length information can be the physical length information of the erasure code block, indicating the actual storage space occupied by the erasure code block. For example, in the example shown in FIG. 1, the data node CS1 has written three erasure codes (including erasure code 11, erasure code 21, and erasure code 31), and thus the erasure code block on the data node CS1 includes the erasure code 11, the erasure code 21, and the erasure code 31, and the length of the erasure code block on the data node CS1 is the total length of the erasure code 11, the erasure code 21, and the erasure code 31. Similarly, the erasure code block on the data node CS2 includes two erasure codes, the erasure code 12 and the erasure code 22, and the length of the erasure code block on the data node CS2 is the total length of the erasure code 12 and the erasure code 22. The erasure code block on the data node CS3 includes one erasure code, the erasure code 13, and the length of the erasure code block on the data node CS3 is the length of the erasure code 13. Figure 3
[0039] In some embodiments, the length information of the corresponding erasure code block can be requested by the client 20 from each data node 104 respectively. Each data node 104 can return the length information of the erasure code block on the node to the client 20. Continuing to refer to FIG. 1, the client 20 can request the length information of the erasure code block from the data nodes CS1, CS2, and CS3 respectively, and the data node CS1 can return the length information of the erasure code block including the erasure code 11, the erasure code 21, and the erasure code 31 to the client 20 in response to the request of the client 20. Similarly, the data node CS2 can return the length information of the erasure code block including the erasure code 12 and the erasure code 22 to the client 20 in response to the request of the client 20. The data node CS3 can return the length information of the erasure code block including the erasure code 13 to the client 20 in response to the request of the client 20. Since the operation of requesting the length information of the erasure code block from each data node 104 is performed by the client 20, the number of communications between the metadata node 102 and each data node 104 is reduced, thereby reducing the communication overhead of the metadata node 102. Figure 3
[0040] In some embodiments, the client 20 can only obtain the length information of part of the erasure code blocks due to communication failure, data loss or damage, etc. For ease of description, the erasure code block to which the length information that the client 20 fails to obtain belongs is referred to as a first erasure code block, and the data node 104 used to reconstruct the erasure code in the first erasure code block is referred to as a first data node. At this time, the length information of the reconstructed erasure code block corresponding to the first erasure code block can be requested from the plurality of data nodes 104, and the reconstructed erasure code block can be obtained by respectively reconstructing each erasure code included in the first erasure code block. Assuming that the first erasure code block includes erasure codes 12 and 22 as shown in FIG. 12, the erasure codes 12 and 22 can be respectively reconstructed, and the length information of the reconstructed erasure code block corresponding to the first erasure code block can be obtained according to the reconstructed erasure codes 12 and 22. The client 20 can determine the length information of the reconstructed erasure code block as the length information of the first erasure code block. Figure 3
[0041] The above process of data reconstruction can be performed on the client 20 or on any one of the data nodes 104 (referred to as a reconstruction node) of the distributed storage system 10. In the example of data reconstruction by the client 20, each first data node can return an erasure code to the client 20, and the client can reconstruct the reconstructed erasure code block corresponding to the first erasure code block according to the erasure codes returned by each first data node, and then obtain the length information of the reconstructed erasure code block. In the example of data reconstruction by the reconstruction node, each first data node can return an erasure code to the reconstruction node, and the reconstruction node can reconstruct the reconstructed erasure code block corresponding to the first erasure code block according to the erasure codes returned by each first data node, and then return the length information of the reconstructed erasure code block to the client 20.
[0042] If any erasure code included in the first erasure code block fails to be reconstructed, the client 20 can re-request the length information of the plurality of erasure code blocks from the plurality of data nodes 104. Among them, the failure of any erasure code included in the first erasure code block to be reconstructed can be that the number of other erasure codes for reconstructing the erasure code that have been written in the distributed storage system 10 is less than k. For example, in the embodiment shown in FIG. 10, assuming that the first erasure code block is the erasure code block on the data node CS1, the erasure code block includes the erasure code 31, and each data node 104 has failed to successfully write other erasure codes for reconstructing the erasure code 31, so that the erasure code 31 cannot be reconstructed at the current time, and thus the first erasure code block on the data node CS1 cannot be reconstructed. In this case, the length information of the plurality of erasure code blocks can be re-requested from the plurality of data nodes 104 through multiple attempts. Figure 3
[0043] After obtaining the length information of the erasure coding blocks on each data node 104, the client 20 can determine the k-th longest erasure coding block from the erasure coding blocks on each data node 104 based on the length information of the erasure coding blocks on each data node 104. For example, in Figure 3 In the example shown, assuming k equals 2, the k-th longest erasure code block is the erasure code block including erasure code 12 and erasure code 22. For another example, assuming the number of erasure code blocks is 11 and k equals 8, the k-th longest erasure code block is the 8th longest erasure code block among the 11 blocks, in descending order of length. If the lengths of the 11 erasure code blocks are 1M, 2M, 3M, ..., and so on, then the k-th longest erasure code block is a block with a length of 4M. The client 20 can sort the length information of the erasure code blocks returned by each data node 104 and determine the k-th longest erasure code block based on the sorting result.
[0044] Where k is greater than or equal to the minimum number of erasure codes required to reconstruct the original data. For example, assuming the number of erasure codes corresponding to each piece of original data is 11, and the minimum number of erasure codes required to reconstruct the original data is 8, then the value of k can be 8, 9, 10, or 11. The writing process of the original data is limited by the data nodes 104, which have relatively small write bandwidth. The ranking of the lengths of the erasure code blocks on each data node 104 reflects the write bandwidth of each data node 104. If the length of the erasure code block on a data node 104 ranks higher, it indicates that the writing speed of the erasure code on that data node 104 is faster, and thus the write bandwidth of that data node 104 is larger. Conversely, if the length of the erasure code block on a data node 104 ranks lower, it indicates that the writing speed of the erasure code on that data node 104 is slower, and thus the write bandwidth of that data node 104 is smaller. For a given piece of original data, if the erasure code corresponding to that original data has been successfully written to data node 104 with a smaller write bandwidth, it can be inferred that the erasure code corresponding to that original data has also been successfully written to data node 104 with a larger write bandwidth. Therefore, it can be assumed that erasure code blocks ranked from 1 to k0 (where k0 is the minimum number of erasure codes required for data reconstruction) can reconstruct the erasure codes that have not yet been successfully written from erasure code blocks ranked from k0+1 to n (where n is the total number of erasure codes corresponding to a piece of original data). Therefore, the target erasure code block of length k (k≥k0) can be determined.
[0045] For example, in Figure 3In the example shown, the write bandwidths of the data nodes CS1, CS2 and CS3 decrease in turn. For the original data 1, a certain erasure code (here, erasure code 13) corresponding to the original data 1 has been successfully written to the data node CS3, and it can be inferred that other erasure codes (here, erasure codes 11 and 12) corresponding to the original data 1 have also been successfully written to the data nodes CS1 and CS2 respectively, which have greater write bandwidths. Similarly, the erasure code 22 corresponding to the original data 2 has been successfully written to the data node CS2, and it can be inferred that the other erasure code (here, erasure code 21) corresponding to the original data 2 has also been successfully written to the data node CS1, which has greater write bandwidth. Assuming k = 2, the top 2 longest erasure code blocks are the erasure code block on the data node CS1 and the erasure code block on the data node CS2, and the erasure codes corresponding to the original data 2 in these two erasure code blocks can reconstruct the erasure code corresponding to the original data 2 on the data node CS3, which has not yet been transmitted. Therefore, the second longest target erasure code block can be determined.
[0046] After determining the target erasure code block, the client 20 can send at least one piece of length information to the metadata node 102. The at least one piece of length information is used to indicate the lengths of at least one piece of original data corresponding to at least one erasure code in the target erasure code block. Then Figure 5 In the example, in the case where the target erasure code block includes the erasure code 21 and the erasure code 22, the at least one piece of length information includes the length information of the original data (i.e., the original data 1) corresponding to the erasure code 21 (referred to as length information 1) and the length information of the original data (i.e., the original data 2) corresponding to the erasure code 22 (referred to as length information 2). The length information 1 is used to indicate the length of the original data 1, and the length information 2 is used to indicate the length of the original data 2.
[0047] In some embodiments, an erasure code can include a data block and a check block. The data block can be obtained by splitting the original data, for example, the original data can be split into a plurality of data blocks of equal length. The check block can be obtained by performing erasure code encoding on each data block, and the specific erasure code encoding algorithm is not limited by the present disclosure. Since the erasure code requires data alignment to generate the check block, the original data is usually padded data. The data before padding is the data actually written by the user (referred to as user data), and after obtaining each piece of user data, the client 20 will first pad the obtained user data to a preset length using padding data, and maintain the mapping relationship between the physical length information and the logical length information of the original data. The physical length information is used to indicate the total length of the user data and the padding data in the original data; and the logical length information is used to indicate the length of the user data in the original data.
[0048] As Figure 6As shown, the client 20 can fill the user data 1 to obtain the original data 1, which includes the user data 1, the filling data, and the mapping relationship (denoted as mapping relationship 1) between the physical length information and the logical length information corresponding to the original data 1. The user data 2 and the user data 3 are processed to obtain the original data 2 and the original data 3, respectively, in a manner similar to that of filling the user data 1 to obtain the original data 1, which will not be described herein again. The client 20 can generate the erasure code based on the original data obtained after filling, and write the generated erasure code to the data node 104. Meanwhile, the client 20 can also synchronize the length information (including the physical length information and the logical length information) of the original data after filling to the metadata node 102.
[0049] Since the data blocks in the erasure code are obtained by splitting the original data, and the original data includes the length information, a certain data block (referred to as a target data block) obtained by splitting the original data includes the length information. The client 20 can obtain the target data block including the length information from the data node 104, parse the length information (including the above-mentioned physical length information and logical length information) from the target data block, and send the parsed length information to the metadata node 102. As shown, Figure 3 As shown, one piece of original data is split into data block 1, data block 2, data block 3, and data block 4, wherein the data block 1, the data block 2, and the data block 3 are data blocks corresponding to the original user data, and the data block 4 is a data block including filling data and length information. Therefore, the client 20 can obtain the data block 4 from the data node 104 storing the data block 4, and parse the length information from the data block 4.
[0050] If the length information of the original data returned by the data node 104 storing the target data block cannot be obtained, it may be caused by damage or loss of the target data block, and the like. Therefore, the client 20 can obtain the length information of the original data determined from the reconstructed data block corresponding to the target data block, which can be reconstructed based on other data blocks in the plurality of data blocks except the target data block. For example, in the example shown, Figure 3 As shown in the example, the target data block corresponding to the original data 1 is the erasure code 13, which is written to the data node CS3. Therefore, the client 20 can obtain the erasure code 13 from the data node CS3, and obtain the length information of the original data 1 from the erasure code 13. If the length information of the original data 1 cannot be obtained, the erasure code 13 can be reconstructed using the erasure code 11 and the erasure code 12 to obtain a reconstructed data block, and the length information of the original data 1 can be obtained from the reconstructed data block.
[0051] In some embodiments, the raw data written into the distributed storage system 10 includes unsynchronized data and synchronized data. Accordingly, the metadata node 102 can maintain length information of the unsynchronized data and length information of the synchronized data at the same time. In a case where the client 20 needs to obtain the latest length information of the raw data and does not need synchronization, the length information of the unsynchronized data can be synchronized. In a case where the client 20 needs to synchronize the latest length information of the raw data, the length information of the synchronized data can be synchronized.
[0052] The distributed storage system 10 can include one or more clients 20, and at the same time, only one client 20 can be allowed to write data into each data node 104, and one or more clients 20 can be allowed to read data from each data node 104. This data access mode is called "one-write multi-read", which can effectively ensure data consistency and avoid write conflicts, while allowing multiple readers to read data concurrently and improving the read performance of the system. The metadata node 102 can assign a writer lock to the client 20, thereby assigning data writing permission to the client 20. The client 20 assigned with the writer lock can perform data writing, and the client 20 not assigned with the writer lock can only perform data reading. Among them, the case of needing to obtain the latest length information of the raw data and not needing synchronization can be the case where the client 20 reads the erasure code from multiple data nodes 104 in the process of data writing by other clients 20. The case of needing to synchronize the latest length information of the raw data can be the case where the client 20 writes the erasure code to multiple data nodes 104 by replacing other clients 20 as a new writer.
[0053] In step S14, the metadata node 102 can determine the target length information from at least one piece of length information sent by the client 20. The at least one piece of length information sent by the client 20 includes length information of at least one piece of raw data corresponding to each of at least one target erasure code block, see Figure 3, assuming that the target erasure code block is the erasure code block on the data node CS2, the length information sent by the client 20 includes the length information of the original data (i.e., original data 1) corresponding to the erasure code 12, and the length information of the original data (i.e., original data 2) corresponding to the erasure code 22. The original data 1 and the original data 2 can be considered as data that have been successfully written into the distributed storage system 10. The metadata node 102 can select target length information from the length information of the original data 1 and the length information of the original data 2 according to a fault tolerance configuration. The fault tolerance configuration is used to indicate the number of erasure codes that can be lost or damaged for any original data under the condition of ensuring the integrity of the original data. For example, assuming that the number of erasure codes corresponding to one piece of original data is 11, and the minimum number of erasure codes required for data reconstruction is 8, the number indicated by the fault tolerance configuration can be any integer between 0 and 3. When the number indicated by the fault tolerance configuration is 0, it means that any one erasure code in the original data is not allowed to be lost or damaged, so that the length information of the original data can be synchronized only after all the erasure codes corresponding to the original data are successfully written into the data nodes 104. When the number indicated by the fault tolerance configuration is 1, it means that one erasure code in the original data is allowed to be lost or damaged, so that the length information of the original data can be synchronized as long as 10 erasure codes corresponding to the original data are successfully written into the data nodes 104, and the remaining one erasure code can be obtained through data reconstruction.
[0054] In the synchronization of the length information of the unsynchronized data, the metadata node 102 can update the length information of the most recently acquired unsynchronized data based on the target length information. The length information of the unsynchronized data can represent the total length of the unsynchronized data in the original data, and the length information of the most recently acquired unsynchronized data can be directly replaced by the currently acquired target length information, so as to update the length information of the most recently acquired unsynchronized data. The metadata node 102 can also return the updated length information of the unsynchronized data to the client 20 in response to a read request of the original data written into the distributed storage system 10 by the client 20. In this way, the original data written into the distributed storage system 10 can be ensured to be real-time visible when the client 20 reads the data. The client 20 can be a client that has not acquired the data write permission to write data into the distributed storage system 10. Further, if the currently acquired target length information is greater than or equal to the length information of the most recently acquired unsynchronized data by the metadata node 102, the metadata node 102 can return the currently acquired target length information as the updated length information of the unsynchronized data to the client 20. If the currently acquired target length information is less than the length information of the most recently acquired unsynchronized data by the metadata node 102, the metadata node 102 can return the length information of the most recently acquired unsynchronized data to the client 20. In this way, the length information acquired by the client 20 can be prevented from being rolled back. If the length represented by the updated length information of the unsynchronized data reaches a preset length, the preset length can also be added to the length information of the most recently acquired synchronized data (i.e., the length information of the most recently acquired synchronized data and the updated length information of the unsynchronized data are summed), so as to update the length information of the most recently acquired synchronized data and clear the length information of the unsynchronized data.
[0055] In synchronizing the length information of the synchronized data, the metadata node 102 can update the length information of the synchronized data obtained last time based on the target length information. The length information of the synchronized data can represent the total length of the synchronized data in the original data, and the metadata node 102 can sum the target length information and the length information of the synchronized data obtained last time to obtain the updated length information of the synchronized data. The metadata node 102 can also return the updated length information of the synchronized data to the client 20 in response to a write request of the client 20 for writing the erasure code to the plurality of data nodes. The write request can be sent by the changed client 20 after the client 20 for writing the original data to the distributed storage system 20 is changed. That is, the metadata node 102 can return the updated length information of the synchronized data to the changed client 20 in response to a write request of the changed client 20 for writing the erasure code to the plurality of data nodes after the client 20 for writing the original data to the distributed storage system 20 is changed. In this way, it can be ensured that the changed client 20 can continue to write data after the original data written by the client 20 before the change.
[0056] In some embodiments, after determining the target erasure code block of the kth length, the second erasure code block can be obtained by data reconstruction so that the length of the second erasure code block is equal to the length of the target erasure code block. The second erasure code block is an erasure code block on the plurality of data nodes 104 whose length is less than the length of the target erasure code block. Specifically, if the erasure code corresponding to a piece of original data is included in the target erasure code block but not included in the second erasure code block, the erasure code corresponding to the original data in the second erasure code block can be reconstructed based on the erasure code corresponding to the original data on each data node 104. Figure 7, still assuming k = 2, the target erasure code block is the erasure code block on the data node CS2, and the second erasure code block is the erasure code block on the data node CS3. Since the erasure code corresponding to the original data 2 has been written into the erasure code block in the data node CS2, but has not been written into the erasure code block in the data node CS3, the erasure code corresponding to the original data 2 written in each data node 104 (including the data nodes CS1 and CS2) other than the data node CS3 can be obtained, and the erasure code corresponding to the original data 2 obtained from the data nodes CS1 and CS2 is used to reconstruct the erasure code corresponding to the original data 2 in the erasure code block on the data node CS3. In the distributed storage system 10, when there is a data node 104 with a small write bandwidth, if the length information of the original data is synchronized only after each erasure code corresponding to the original data is successfully written into the data node 104, the data synchronization process will be limited by the data node 104 with the smallest write bandwidth, thereby reducing the efficiency. The embodiment of the present disclosure combines data synchronization with data reconstruction, provides a certain data redundancy through data reconstruction, thereby improving the fault tolerance of the data. As long as the number of successfully written erasure codes in the erasure code corresponding to one piece of original data is greater than or equal to the minimum number of erasure codes required for data reconstruction, the length information of the original data can be synchronized, and the erasure code that has not been successfully written can be recovered through data reconstruction. In this way, the synchronization process does not need to be limited by the data node 104 with the smallest write bandwidth, and the integrity of the original data is ensured.
[0057] If the second erasure code block reconstruction fails, the step of determining the target erasure code block with the kth length in the plurality of erasure code blocks is returned. In the embodiment, the second erasure code block reconstruction failure can be that the number of erasure codes used to reconstruct any one or more erasure codes in the second erasure code block is less than the minimum number of erasure codes required for data reconstruction. That is, the current number of erasure codes is not sufficient to reconstruct the second erasure code block. This can be caused by data loss or damage, or data node 104 failure, etc. In this case, the target erasure code block can be re-determined, thereby reducing the risk of data loss.
[0058] Figure 8A schematic diagram showing the interaction process between the client 20, the metadata node 102 and the data node 104 is shown. The metadata node 102 divides the original data into unsynchronized data and synchronized data, the data block corresponding to the synchronized data cannot be written again, and when writing is needed, the metadata node 102 allocates a new group of data nodes 104 for writing. The data being written by the new data node 104 is all unsynchronized data. The metadata node 102 can simultaneously maintain the length information of the synchronized data (i.e., synchronized length information) and the length information of the unsynchronized data (i.e., unsynchronized length information), and the above-mentioned synchronized length information and unsynchronized length information both include physical length information and logical length information.
[0059] In step S20, the client 20 can write the erasure code corresponding to the original data to each data node 104, and other clients can read the written erasure code from each data node 104. Taking two clients 20 as an example below, for the convenience of distinction, the two clients 20 are respectively denoted as client 1 and client 2, wherein the client 2 has data writing authority, and the erasure code corresponding to the original data can be written to each data node 104 by the client 2. The client 1 only has data reading authority.
[0060] In step S22, when the client 1 needs to obtain the unsynchronized length information of the original data and does not need to synchronize, it queries the length information of the erasure code block on each data node 104, sorts the length information from large to small, takes the kth largest length information, and further queries the length information of the original data (i.e., at least one length information in the foregoing embodiment, including physical length information and logical length information) to the corresponding data node 104. The length information can be recorded in a mapping table, and the mapping table can record the mapping relationship between the physical length information of the original data and the logical length information of the original data, and the mapping table can be included in the original data. When the client 1 writes the erasure code to the data node 104, it can split the original data to obtain data blocks, and the mapping table can be recorded in one of the data blocks and sent to the data node 104. When the client 20 needs to obtain the synchronized length information, it also queries the length information of the erasure code block on the data node 104, and then queries the length information of the original data (i.e., at least one length information in the foregoing embodiment). The subsequent processing mode is similar and will not be described here.
[0061] In step S24, the client 1 can also obtain the length information of the unsynchronized data obtained last time and the length information of the synchronized data obtained last time from the metadata node 102.
[0062] In step S26, the client 1 uploads the queried physical length information and logical length information to the metadata node 102, so that the metadata node 102 decides the target length information (the target length information can also include the physical length information and the logical length information).
[0063] In step S28, after the metadata node 102 decides the target length information, the metadata node 102 adds the target length information to the length information of the latest acquired unsynchronized data and the length information of the latest acquired synchronized data, so as to maintain the latest length information of the unsynchronized data and the latest length information of the synchronized data respectively. The metadata node 102 can also send the latest length information of the unsynchronized data and the latest length information of the synchronized data to the client 1, so that the client 1 determines the total length of the raw data currently being written.
[0064] Figure 8 A timing diagram of a data synchronization process is shown. The flow of acquiring the length information of the latest data of the unsynchronized data and the length information of the synchronized data is as follows:
[0065] Step S30: The client 20 requests the synchronized length information from the metadata node 102.
[0066] Step S32: The metadata node 102 returns the synchronized length information to the client 20. If there is unsynchronized data on a data node 104, the data node 104 is a writing data node, and the metadata node 102 can also return the positions of the writing data nodes at the current time to the client 20.
[0067] Step S34: The client 20 requests the length information of the erasure code block currently being written in the data node 104 from the plurality of data nodes 104.
[0068] Steps S36-1 to S36-3: Each data node 104 (assuming three data nodes here) returns the length information of the erasure code block on the node to the client 20 in turn.
[0069] Step S38: The client 20 obtains the k-th largest length information from the length information.
[0070] Step S40: The client 20 queries the length information of the raw data from the specific data node 104 by using the obtained k-th largest length information of the target erasure code block, including the length information of the raw data corresponding to each erasure code in the target erasure code block, which is recorded in the mapping table.
[0071] Step S42: The data node 104 returns the mapping table to the client 20.
[0072] Step S44: The client 20 reports the acquired length information as candidate length information to the metadata node 102.
[0073] Step S46: The metadata node 102 decides target length information according to the fault-tolerant configuration. The metadata node 102 can update the length information of the unsynchronized data according to the target length information. If the updated length information of the unsynchronized data is greater than the length information of the unsynchronized data before the update, the metadata node 102 notifies the client 20 of the updated length information of the unsynchronized data, otherwise, the metadata node 102 notifies the client 20 of the length information of the unsynchronized data before the update. The metadata node 102 can also update the length information of the synchronized data according to the target length information. If the updated length information of the synchronized data is greater than the length information of the synchronized data before the update, the metadata node 102 notifies the client 20 of the updated length information of the synchronized data, otherwise, the metadata node 102 notifies the client 20 of the length information of the synchronized data before the update.
[0074] Step S48: After the client 20 acquires the length information returned by the metadata node 102, the client 20 determines whether the length information returned by the metadata node 102 is greater than the length information maintained locally. If yes, the client 20 accepts the length information returned by the metadata node 102, otherwise, the client 20 requests retry or failure. If the metadata node 102 returns the length information of the unsynchronized data, the client 20 can add the length information of the synchronized data acquired in step S30 to the length information of the unsynchronized data returned by the metadata node 102 to obtain the total length of the raw data being written into the distributed storage system 10. If the metadata node 102 returns the length information of the synchronized data, the client 20 can directly determine the length information returned by the metadata node 102 as the total length of the raw data being written into the distributed storage system 10.
[0075] In the entire processing flow, there are a large number of interactions between the client 20 and the metadata node 102 and the data node 104. When some components of the distributed storage system 10 fail or are temporarily lost, fault-tolerant processing is required. Corresponding to the fault-tolerant configuration, the metadata node 102 can decide the target length information according to the fault-tolerant configuration. Figure 9Each step, in steps S34 to S36-1, S36-2 and S36-3, in the event of a failure to access the data node 104, as long as the length information returned by the k data nodes 104 can be guaranteed, the subsequent process can continue. For step S40, the mapping table is generally stored in one of the data nodes 104, if this data node 104 cannot be accessed or the length information stored is not enough, the mapping table stored in the data node 104 is reconstructed through other data nodes 104. For steps S30, S32, S44, etc., if the execution fails, it can be retried several times, and if it still fails, a prompt information for prompting the processing failure can be returned. Between steps S34 and S40, different data nodes 104 may be disconnected, which can cause the erasure code block with a length smaller than the target erasure code block to be unable to be obtained through data reconstruction, at this time, the client 20 can execute the above process from the beginning to ensure the integrity of the data.
[0076] The embodiments of the present disclosure multiplex the logic of the client 20 to obtain the length information from each data node 104, and leave the important length decision logic in the metadata node 102, on the one hand, avoiding the metadata node 102 to query all the write data nodes 104, reducing the communication burden of the metadata node 102, and improving the read-write performance of the distributed storage system 10; on the other hand, keeping the decision logic in the metadata node 102, reducing the situation that the length information of the data is determined incorrectly or the state is rolled back due to the inconsistent behaviors of different clients 20.
[0077] Referring to Figure 10 The embodiments of the present disclosure also provide a data synchronization method, applied to a client 20 connected to a distributed storage system 10, the distributed storage system 10 including a metadata node 102 and a plurality of data nodes 104, the plurality of data nodes 104 respectively used for storing a plurality of erasure code blocks, any one of the erasure code blocks including at least one erasure code respectively corresponding to at least one piece of original data, the erasure codes corresponding to the same original data in the plurality of erasure code blocks used for data reconstruction of the original data, the metadata node 102 used for storing length information of the original data written into the distributed storage system 10; the method includes:
[0078] Step S52: determining a target erasure code block with the kth length in the plurality of erasure code blocks, k being greater than or equal to the minimum number of erasure codes required for reconstructing the original data;
[0079] Step S54: sending at least one piece of length information to the metadata node 102, so that the metadata node 102 determines target length information from the at least one piece of length information, and synchronizes the length of the original data written into the distributed storage system 10 based on the target length information; the at least one piece of length information is used to indicate the length of at least one piece of original data corresponding to at least one erasure code in the target erasure code block.
[0080] In some embodiments, the determining the target erasure code block of the kth length in the plurality of erasure code blocks comprises: requesting the length information of the plurality of erasure code blocks from the plurality of data nodes 104 respectively; sorting the length information of the plurality of erasure code blocks to obtain a sorting result; and determining the target erasure code block of the kth length in the plurality of erasure code blocks according to the sorting result.
[0081] In some embodiments, the plurality of erasure code blocks comprises a first erasure code block, and the length information of the first erasure code block fails to be obtained from the plurality of data nodes 104, and the method further comprises: requesting the length information of a reconstructed erasure code block corresponding to the first erasure code block from the plurality of data nodes 104, wherein the reconstructed erasure code block is obtained by reconstructing each erasure code included in the first erasure code block.
[0082] In some embodiments, the method further comprises: if any one of the erasure codes included in the first erasure code fails to be reconstructed, returning to the step of requesting the length information of the plurality of erasure code blocks from the plurality of data nodes 104 respectively.
[0083] In some embodiments, the original data written into the distributed storage system 10 comprises unsynchronized data, and the target length information is used by the metadata node 102 to determine the total length information of the unsynchronized data; and the method further comprises: obtaining the total length information of the unsynchronized data returned by the metadata node 102, wherein the total length information of the unsynchronized data is returned in response to a read request of the client 20 to the distributed storage system 10.
[0084] In some embodiments, the method further comprises: if the length indicated by the total length information of the unsynchronized data returned by the metadata node 102 is greater than the length indicated by the total length information of the unsynchronized data currently stored on the client 20, replacing the total length information of the unsynchronized data currently stored with the total length information of the unsynchronized data returned by the metadata node 102.
[0085] In some embodiments, the original data written into the distributed storage system 10 comprises synchronized data, and the target length information is used by the metadata node 102 to determine the total length information of the synchronized data; and the method further comprises: obtaining the total length information of the synchronized data returned by the metadata node 102, wherein the total length information of the synchronized data is returned in response to a write request of the client 20 to the plurality of data nodes 104 to write erasure codes.
[0086] In some embodiments, for any one piece of original data, the erasure code corresponding to the original data comprises a plurality of data blocks obtained by splitting the original data, and the length information of the original data is recorded in a target data block in the plurality of data blocks; and the method further comprises: requesting the length information of the original data from a data node 104 storing the target data block.
[0087] In some embodiments, the method further comprises: if the length information of the original data returned by the data node 104 storing the target data block is not acquired, acquiring the length information of the original data determined from the reconstructed data block corresponding to the target data block, the reconstructed data block being reconstructed based on the data blocks other than the target data block.
[0088] In some embodiments, the original data comprises user data and padding data, and the length information of the original data comprises physical length information and logical length information, the physical length information being used to indicate the total length of the user data and the padding data, and the logical length information being used to indicate the length of the user data.
[0089] In some embodiments, the at least one erasure code included in the second erasure code block is reconstructed by data, and the second erasure code block is an erasure code block on the data node 104 with a length smaller than the target erasure code block.
[0090] In some embodiments, the method further comprises: if the reconstruction of the second erasure code block fails, returning to the step of determining the target erasure code block with the kth length in the plurality of erasure code blocks.
[0091] The embodiments of the method executed by the client 20 are specific details, which can refer to the embodiments of the method executed by the metadata node 102, which are not described herein.
[0092] The embodiments of the present disclosure also provide a computer device, which at least includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of any of the preceding embodiments when executing the program.
[0093] A more specific hardware structure of a computer device provided by the embodiments of the present disclosure is shown, which can include a processor 30, a memory 32, an input / output interface 34, a communication interface 36, and a bus 38. The processor 30, the memory 32, the input / output interface 34, and the communication interface 36 are connected to each other through the bus 38 for communication within the device.
[0094] The processor 30 can be implemented in the form of a general-purpose central processing unit, a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present disclosure. The processor 30 can also include a graphics card, which can be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0095] The memory 32 can be implemented in the form of a read only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 32 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented by software or firmware, the related program codes are stored in the memory 32 and are invoked and executed by the processor 30.
[0096] The input / output interface 34 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0097] The communication interface 36 is configured to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as a USB, a network cable, etc.) or a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).
[0098] The bus 38 includes a path for transmitting information between various components (such as the processor 30, the memory 32, the input / output interface 34, and the communication interface 36) of the device.
[0099] It should be noted that although the above device only shows the processor 30, the memory 32, the input / output interface 34, the communication interface 36, and the bus 38, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include the components necessary for implementing the solutions of the embodiments of the present disclosure, and does not have to include all the components shown in the figure.
[0100] The embodiments of the present disclosure provide a computer program product, including a computer program, which is executed by a processor to implement the method described in any of the embodiments of the present disclosure.
[0101] The embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method described in any of the preceding embodiments.
[0102] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0103] The embodiments of the present disclosure also provide a distributed storage system 10, which comprises:
[0104] a metadata node 102, configured to perform the method performed by the metadata node 102 in any of the preceding embodiments; and
[0105] a plurality of data nodes 104, respectively configured to store a plurality of erasure code blocks, any one of the erasure code blocks comprising at least one erasure code corresponding to at least one piece of original data respectively, and the erasure codes corresponding to the same piece of original data in the plurality of erasure code blocks being used to reconstruct data of the piece of original data.
[0106] The distributed storage system 10 interfaces with a client 20, which is configured to perform the method performed by the client 20 in any of the preceding embodiments.
[0107] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus necessary universal hardware platforms. Based on such understanding, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the various embodiments or some parts of the embodiments of the present disclosure.
[0108] The systems, apparatuses, modules or units in the above embodiments can be implemented by computer devices or entities, or by products with certain functions. A typical implementation device is a computer, which can be specifically a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0109] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the device embodiments are described more simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the part of the method embodiments. The device embodiments described above are merely illustrative, and the modules described as separate components can or can not be physically separated. In the implementation of the embodiments of the present disclosure, the functions of each module can be implemented in one or more software and / or hardware. Part or all of the modules can be selected to achieve the purpose of the embodiments of the present disclosure according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0110] The above is only a specific implementation of the embodiments of the present disclosure. It should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the embodiments of the present disclosure, and these improvements and refinements should also be considered within the protection scope of the embodiments of the present disclosure.
Claims
1. A data synchronization method applied to a metadata node of a distributed storage system, wherein a plurality of data nodes of the distributed storage system are respectively configured to store a plurality of erasure code blocks, any one of the erasure code blocks comprises at least one erasure code corresponding to at least one piece of original data respectively, erasure codes corresponding to a same piece of original data in the plurality of erasure code blocks are configured to reconstruct the original data, and the metadata node is configured to store length information of the original data written into the distributed storage system; the method comprises: obtaining at least one piece of length information sent by a client connected to the distributed storage system, wherein the at least one piece of length information is configured to represent lengths of at least one piece of original data corresponding to at least one erasure code in a target erasure code block, the target erasure code block is a kth longest erasure code block in the plurality of erasure code blocks, and k is greater than or equal to a minimum number of erasure codes required for reconstructing the original data; determining target length information from the at least one piece of length information, and synchronizing the length information of the original data written into the distributed storage system based on the target length information. 2.The method of claim 1, wherein the original data written into the distributed storage system comprises unsynchronized data; and the synchronizing the length information of the original data written into the distributed storage system based on the target length information comprises: updating length information of the unsynchronized data obtained last time based on the target length information. The method further comprises: in response to a read request of the original data written into the distributed storage system from the client, returning the updated length information of the unsynchronized data to the client. 3.The method of claim 2, further comprising: if a length represented by the updated length information of the unsynchronized data is greater than or equal to a length represented by the length information of the unsynchronized data obtained last time, returning the updated length information of the unsynchronized data to the client. 4.The method of claim 3, further comprising: if the length represented by the updated length information of the unsynchronized data is less than the length represented by the length information of the unsynchronized data obtained last time, returning the length represented by the length information of the unsynchronized data obtained last time to the client. 5.The method of claim 1, wherein the original data written into the distributed storage system comprises synchronized data; and the synchronizing the length information of the original data written into the distributed storage system based on the target length information comprises: updating length information of the synchronized data obtained last time based on the target length information. The method further comprises: in response to a write request of the erasure code to the plurality of data nodes from the client, returning the updated length information of the synchronized data to the client.
6. The method of claim 1, wherein the original data comprises user data and padding data, and the length information of the original data comprises physical length information and logical length information, the physical length information being used to indicate a total length of the user data and the padding data, and the logical length information being used to indicate a length of the user data. The at least one length information sent by the client comprises at least one physical length information and logical length information corresponding to the at least one physical length information respectively; and the target length information is determined from the at least one length information, and the length information of the original data written into the distributed storage system is synchronized based on the target length information, comprising: target physical length information is determined from the at least one physical length information, and target logical length information corresponding to the target physical length information is determined, and the physical length information and the logical length information of the original data written into the distributed storage system are synchronized based on the target physical length information and the target logical length information respectively.
7. A data synchronization method applied to a client interfacing with a distributed storage system, the distributed storage system comprising a metadata node and a plurality of data nodes, the plurality of data nodes respectively configured to store a plurality of erasure code blocks, any one of the erasure code blocks comprising at least one erasure code corresponding to at least one piece of original data respectively, erasure codes corresponding to a same piece of original data in the plurality of erasure code blocks being configured to reconstruct the original data, and the metadata node being configured to store length information of the original data written into the distributed storage system. The method comprises: determining a target erasure code block of the kth length in the plurality of erasure code blocks, k being greater than or equal to a minimum number of erasure codes required for reconstructing the original data; determining at least one length information based on the target erasure code block; sending the at least one length information to the metadata node, so that the metadata node determines target length information from the at least one length information, and synchronizes the length of the original data written into the distributed storage system based on the target length information; the at least one length information being used to indicate the length of at least one original data corresponding to at least one erasure code in the target erasure code block respectively.
8. The method of claim 7, wherein the determining the target erasure code block of the kth length in the plurality of erasure code blocks comprises: requesting length information of the plurality of erasure code blocks from the plurality of data nodes respectively; sorting the length information of the plurality of erasure code blocks to obtain a sorting result; determining the target erasure code block of the kth length in the plurality of erasure code blocks according to the sorting result.
9. The method of claim 8, wherein the plurality of erasure code blocks comprise a first erasure code block, and the length information of the first erasure code block fails to be obtained from the plurality of data nodes, and the method further comprises: requesting length information of a reconstructed erasure code block corresponding to the first erasure code block from the plurality of data nodes, the reconstructed erasure code block being obtained by reconstructing each erasure code included in the first erasure code block respectively.
10. The method of claim 9, further comprising: if any one erasure code included in the first erasure code block fails to be reconstructed, returning to the step of requesting the length information of the plurality of erasure code blocks from the plurality of data nodes respectively.
11. The method of claim 7, wherein the original data written into the distributed storage system comprises unsynchronized data, and the target length information is used by the metadata node to determine total length information of the unsynchronized data; and the method further comprises: obtaining the total length information of the unsynchronized data returned by the metadata node, wherein the total length information of the unsynchronized data is returned in response to a read request of the client for writing the distributed storage system; if the length indicated by the total length information of the unsynchronized data returned by the metadata node is greater than the length indicated by the total length information of the unsynchronized data currently stored on the client, replacing the total length information of the unsynchronized data currently stored with the total length information of the unsynchronized data returned by the metadata node.
12. The method of claim 7, wherein the original data written into the distributed storage system comprises synchronized data, and the target length information is used by the metadata node to determine the total length information of the synchronized data; the method further comprises: obtaining the total length information of the synchronized data returned by the metadata node, wherein the total length information of the synchronized data is returned in response to a write request of the client for writing erasure codes to the plurality of data nodes.
13. The method of claim 7, wherein for any one piece of original data, the erasure codes corresponding to the original data comprise a plurality of data blocks obtained by splitting the original data, and the length information of the original data is recorded in a target data block in the plurality of data blocks; the method further comprises: requesting the length information of the original data from a data node storing the target data block.
14. The method of claim 13, the method further comprises: if the length information of the original data returned by the data node storing the target data block is not obtained, obtaining the length information of the original data determined from a reconstructed data block corresponding to the target data block, the reconstructed data block being reconstructed based on other data blocks in the plurality of data blocks except the target data block.
15. The method of claim 7, wherein the original data comprises user data and padding data, and the length information of the original data comprises physical length information and logical length information, the physical length information being used to indicate the total length of the user data and the padding data, and the logical length information being used to indicate the length of the user data; the at least one piece of length information comprises at least one piece of physical length information and logical length information corresponding to the at least one piece of physical length information; and the sending of the at least one piece of length information to the metadata node, so that the metadata node determines target length information from the at least one piece of length information and synchronizes the length of the original data written into the distributed storage system based on the target length information, comprises: sending the at least one piece of physical length information and the at least one piece of logical length information to the metadata node, so that the metadata node determines target physical length information from the at least one piece of physical length information, and determines logical length information corresponding to the target physical length information as target logical length information, and synchronizes physical length information and logical length information of original data written into the distributed storage system based on the target physical length information and the target logical length information respectively.
16. The method of claim 7, wherein the at least one erasure code included in the second erasure code block is obtained through data reconstruction, and the second erasure code block is an erasure code block having a length smaller than the target erasure code block among the plurality of data nodes.
17. The method of claim 16, further comprising: if the second erasure code block fails to be reconstructed, returning to the step of determining the target erasure code block having the kth length among the plurality of erasure code blocks.
18. A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 17.
19. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method of any one of claims 1 to 17 when executing the program.
20. A computer program product comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 17.
21. A distributed storage system, comprising: a plurality of data nodes respectively configured to store a plurality of erasure code blocks, wherein any one erasure code block includes at least one erasure code corresponding to at least one piece of original data respectively, and erasure codes corresponding to the same original data among the plurality of erasure code blocks are used for data reconstruction of the original data; and a metadata node configured to implement the method of any one of claims 1 to 6; and wherein the distributed storage system interfaces with a client, and the client is configured to implement the method of any one of claims 7 to 17.
22. The distributed storage system of claim 21, wherein the client is configured to: request length information of the plurality of erasure code blocks from the plurality of data nodes respectively; sort the length information of the plurality of erasure code blocks to obtain a sorting result; and determine a target erasure code block having the kth length among the plurality of erasure code blocks according to the sorting result.
23. The distributed storage system of claim 22, wherein the plurality of erasure code blocks include a second erasure code block, and length information of the second erasure code block fails to be obtained from the plurality of data nodes, and the client is further configured to: request length information of a reconstructed erasure code block corresponding to the second erasure code block from the plurality of data nodes, wherein the reconstructed erasure code block is obtained through data reconstruction of each erasure code included in the second erasure code block respectively.