An erasure code storage method, system and electronic device

By introducing virtual data blocks and dynamic update mechanisms in distributed storage systems, the problem of difficult reduction in the storage cost of erasure codes is solved, and the multiplexing of verification blocks and saving of storage costs is achieved.

CN114003174BActive Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111283184.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-06-24
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

As the amount of data increases, the storage cost of erasure codes in distributed storage systems is difficult to effectively reduce, especially when cluster nodes are expanded.

Method used

By introducing virtual data blocks into the storage node, using the dynamic update mechanism of virtual data blocks and verification blocks, the replacement of erasure data blocks and the multiplexing of verification blocks are realized, and the storage cost is reduced.

Benefits of technology

The multiplexing of the verification block is realized, the storage cost is reduced, and the erasure code storage method is dynamically adjusted to adapt to the expansion of cluster nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114003174B_ABST
    Figure CN114003174B_ABST
Patent Text Reader

Abstract

The present application discloses an erasure code storage method, system and electronic device. Based on the first original data to be stored, a first number of first erasure data blocks and a second number of first parity blocks are determined. The third number of data blocks are respectively stored in different storage nodes. The third number of data blocks includes: the first number of first erasure data blocks, the second number of first parity blocks and at least one virtual data block. The second original data to be stored is obtained. At least one second erasure data block determined based on the second original data is stored in the storage node where the at least one virtual data block is stored. The at least one virtual data block is deleted. Based on the first number of first erasure data blocks, the second number of first parity blocks and at least one second erasure data block, a second number of second parity blocks are determined. The second number of first parity blocks are updated to the second number of second parity blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data storage, and particularly to an erasure code storage method, system, and electronic device. Background Art

[0002] Erasure codes are widely used in distributed storage systems. The data to be stored is divided into different data blocks, and the parity blocks are stored in different nodes simultaneously, so that when a data block fails, the data in the failed data block can be determined jointly based on the parity blocks and the non-failed data blocks.

[0003] However, with the increase in data volume, the storage demand for data is also increasing. When the cluster nodes are expanded, in order to reduce the storage cost of erasure codes, it is necessary to adjust the storage method of erasure codes. Summary of the Invention

[0004] In view of this, this application provides an erasure code storage method, system, and electronic device, and the specific solutions are as follows:

[0005] An erasure code storage method includes:

[0006] Determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store data blocks of a third number in different storage nodes respectively. The data blocks of the third data include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block;

[0007] Obtain the second original data to be stored;

[0008] Store at least one second erasure data block determined based on the second original data in the storage node where the at least one virtual data block is stored, and delete the at least one virtual data block;

[0009] Determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0010] Further, the storing the data blocks of the third number in different storage nodes respectively includes:

[0011] Store the first number of first erasure data blocks and the second number of first parity blocks in different cluster nodes respectively, and store the at least one virtual data block in different virtual nodes respectively.

[0012] Storing the at least one second erasure data block into the storage nodes of the at least one virtual data block storage includes:

[0013] When obtaining the second original data, determining that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, and determining the newly added expansion nodes;

[0014] Storing at least one second erasure data block determined based on the second original data into the expansion nodes, and deleting the virtual nodes matching the number of the expansion nodes.

[0015] Further, storing the data blocks of the third quantity into different storage nodes respectively includes:

[0016] Storing the first quantity of first erasure data blocks and the second quantity of first parity blocks into different first cluster nodes respectively, and storing the at least one virtual data block into at least one second cluster node respectively;

[0017] Wherein, the second cluster node is a physical storage node or a virtual node.

[0018] Further, it further includes:

[0019] Obtaining the third original data to be stored;

[0020] If it is determined that the number of current cluster nodes is less than the number of cluster nodes when obtaining the second original data, storing the third erasure data block and the third parity block determined based on the third original data into different cluster nodes respectively, and storing at least one virtual data block related to the third original data into different virtual nodes respectively.

[0021] Further, it further includes:

[0022] Receiving a deletion instruction from a first cluster node;

[0023] Determining at least one first data block stored in the first cluster node to be deleted based on the deletion instruction, and storing the at least one first data block into other cluster nodes in the cluster except the first cluster node respectively;

[0024] Updating the parity blocks stored in other cluster nodes in the cluster except the first cluster node.

[0025] Further, storing the data blocks of the third quantity into different storage nodes respectively includes:

[0026] Storing the data blocks of the third quantity into the first stripes of each storage node;

[0027] Respectively store data blocks of a third quantity determined based on fourth original data to be stored to a second stripe of each storage node, where the data blocks of the third quantity determined based on the fourth original data to be stored include: a first quantity of first erasure data blocks related to the fourth original data, a second quantity of first parity blocks related to the fourth original data, and at least one virtual data block.

[0028] Further, storing the at least one second erasure data block to the storage node storing the at least one virtual data block includes:

[0029] Sequentially store the at least one second erasure data block to positions corresponding to different stripes in the storage node for storing the virtual data block.

[0030] Further, sequentially storing the at least one second erasure data block to positions corresponding to different stripes in the storage node for storing the virtual data block includes:

[0031] If the quantity of the second erasure data blocks is greater than the quantity of stripes storing data blocks, store a quantity of second erasure data blocks equal to the quantity of stripes storing data blocks to positions of different stripes in the storage node for storing the virtual data block;

[0032] Store the second erasure data blocks not stored to the storage node for storing the virtual data block sequentially to a third stripe in each storage node where no data block is stored;

[0033] Generate a parity block based on the data blocks stored to the third stripe, and store it to the third stripe.

[0034] An electronic device includes:

[0035] A processor, configured to determine a first quantity of first erasure data blocks and a second quantity of first parity blocks based on first original data to be stored, store data blocks of a third quantity to different storage nodes respectively, where the data blocks of the third quantity include: the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one virtual data block; obtain second original data to be stored, store at least one second erasure data block determined based on the second original data to the storage node storing the at least one virtual data block, and delete the at least one virtual data block; determine a second quantity of second parity blocks based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and the at least one second erasure data block, and update the second quantity of first parity blocks to the second quantity of second parity blocks;

[0036] A memory, configured to store a program for the processor to execute the above processing procedure.

[0037] An erasure code storage system, comprising:

[0038] A first storage unit, configured to determine a first number of first erasure data blocks and a second number of first parity blocks based on first original data to be stored, and store data blocks of a third number to different storage nodes respectively, where the data blocks of the third data include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block;

[0039] An obtaining unit, configured to obtain second original data to be stored;

[0040] A second storage unit, configured to store at least one second erasure data block determined based on the second original data to the storage node where the at least one virtual data block is stored, and delete the at least one virtual data block;

[0041] A determining unit, configured to determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0042] As can be seen from the above technical solution, in the erasure code storage method, system and electronic device disclosed in this application, a first number of first erasure data blocks and a second number of first parity blocks are determined based on first original data to be stored, and data blocks of a third number are stored to different storage nodes respectively, where the data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block, second original data to be stored is obtained, at least one second erasure data block determined based on the second original data is stored to the storage node where the at least one virtual data block is stored, the at least one virtual data block is deleted, a second number of second parity blocks are determined based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and the second number of first parity blocks are updated to the second number of second parity blocks. In this solution, relevant data blocks of a certain data are stored through virtual data blocks, parity blocks and erasure data blocks, and after there is new data to be stored, the second erasure data blocks corresponding to the new data to be stored are stored to the storage nodes of the virtual data blocks, so as to realize replacing the virtual data blocks with the second erasure data blocks, only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data blocks again, realizing the reuse of the parity blocks and saving the storage cost. Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a flowchart of an erasure code storage method disclosed in an embodiment of the present application;

[0045] Figure 2 It is a schematic diagram of storing data blocks in a distributed storage system in the prior art;

[0046] Figure 3 It is a schematic diagram of storing data blocks in a distributed storage system disclosed in an embodiment of the present application;

[0047] Figure 4 It is a schematic diagram of storing data blocks in a distributed storage system disclosed in an embodiment of the present application;

[0048] Figure 5 It is a flowchart of an erasure code storage method disclosed in an embodiment of the present application;

[0049] Figure 6 It is a schematic diagram of storing data blocks in a distributed storage system disclosed in an embodiment of the present application;

[0050] Figure 7 It is a flowchart of an erasure code storage method disclosed in an embodiment of the present application;

[0051] Figure 8 It is a flowchart of an erasure code storage method disclosed in an embodiment of the present application;

[0052] Figure 9 It is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application;

[0053] Figure 10 It is a schematic diagram of the structure of an erasure code storage system disclosed in an embodiment of the present application. Detailed implementation manners

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0055] The present application discloses an erasure code storage method, and its flowchart is asFigure 1 As shown, it includes:

[0056] Step S11: Determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store data blocks of a third number in different storage nodes respectively. The data blocks of the third data include: a first number of first erasure data blocks, a second number of first parity blocks, and at least one virtual data block;

[0057] Step S12: Obtain the second original data to be stored;

[0058] Step S13: Store at least one second erasure data block determined based on the second original data in the storage nodes where at least one virtual data block is stored, and delete at least one virtual data block;

[0059] Step S14: Determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0060] As the amount of data in the distributed system increases, the storage requirements for data are also getting larger. Currently, static erasure codes are usually used for storage. However, when the cluster expands, the storage cost of data will increase.

[0061] Among them, static erasure code storage is to keep the erasure code configuration unchanged when the cluster expands. When there is data to be stored, directly store the erasure data blocks corresponding to the data and the generated parity blocks.

[0062] For example: The erasure data blocks of the first stored data include D1 and D2, the erasure data blocks of the second stored data include D3 and D4, the parity blocks corresponding to the erasure data blocks of the first stored data are P1 and P2, and the parity blocks corresponding to the erasure data blocks of the second stored data are P3 and P4. When storing the relevant data blocks of the first stored data and the relevant data blocks of the second stored data, as Figure 2 shown, store the first erasure data block D1 of the first stored data in node 1, store the second erasure data block D2 of the first stored data in node 2, store the first parity block P1 of the first stored data in node 3, and store the second parity block P2 of the first stored data in node 4; store the first erasure data block D3 of the second stored data in node 1, store the second erasure data block D4 of the second stored data in node 2, store the first parity block P3 of the second stored data in node 3, and store the second parity block P4 of the second stored data in node 4.

[0063] That is, when storing the relevant data blocks of the first stored data and the relevant data blocks of the second stored data, the corresponding data blocks are sequentially stored in the corresponding nodes, which makes when taking Figure 2 as an example for storage, the storage cost is: (k + m) / k = (2 + 2) / 2 = 2, where k is the number of erasure code data blocks, m is the number of parity blocks, and k + m is the number of storage nodes, that is, the storage cost is 2 times.

[0064] In this solution, when storing data, a separate virtual data block is generated for each data, which is used as one of the erasure data blocks, and the erasure data block is also stored on a storage node. So that when there is new data to be stored, the erasure data block of the new data can replace the virtual data block and be stored on this storage node, and the parity block is updated so that the parity block is obtained based on the previously stored erasure data blocks and the newly stored erasure data blocks together, thereby realizing the reuse of the storage nodes of the parity blocks.

[0065] Specifically, when storing the first original data, first generate the first erasure data block of the first original data and the first parity block. There are a total of the first number of the first erasure data blocks and a total of the second number of the first parity blocks. When one or several data blocks in the first number of the first erasure data blocks are abnormal and the complete first original data cannot be obtained based on all the first erasure data blocks, the first original data can be determined jointly by using the first parity block and the first erasure data blocks that are not abnormal, where the number of abnormal data blocks is less than or equal to the second number.

[0066] When storing, the first number of the first erasure data blocks and the second number of the first parity blocks can be directly stored in different storage nodes. For example, if the first number is 3 and the second number is 2, then store the first first erasure data block in the first node, the second first erasure data block in the second node, the third first erasure data block in the third node, the first first parity block in the fourth node, and the second first parity block in the fifth node.

[0067] In addition, when storing the first original data to be stored, a virtual data block can be added and stored on an independent node. By using the virtual data block to occupy the position of the storage node, when there is a new data block to be stored, the virtual data block can be directly replaced, and the parity block is updated, thereby saving the storage node for storing the parity block of the new data block.

[0068] Among them, the virtual data block can be 0. When determining the first parity block based on the virtual data block and the first erasure data block, the virtual data block will not affect the data of the finally generated first parity block. That is, the first parity block determined based on the virtual data block with a value of 0 and the first erasure data block is the same as the first parity block determined based on the first erasure data block.

[0069] When there is a second original data to be stored, determine the second erasure data block of the second original data, store at least one second erasure data block to the storage nodes where at least one virtual data block is stored, and at the same time, delete at least one virtual data block.

[0070] Specifically, determine the number of virtual data blocks. If the number of second erasure data blocks generated by the second original data is not greater than the number of virtual data blocks, directly store all the second erasure data blocks to the storage nodes of different virtual data blocks respectively; if the number of second erasure data blocks generated by the second original data is greater than the number of virtual data blocks, select the data blocks with the same number as the number of virtual data blocks from the second erasure data blocks, and store the selected data blocks to the storage nodes of different virtual data blocks respectively.

[0071] Among them, the storage node that can store the second erasure data block must be an actual storage node, which can be the storage node after expanding the node.

[0072] For example: when storing the first original data, the first original data includes the first erasure data blocks D1 and D2, and the first parity blocks P1 and P2. When storing, store D1 in node 1, store D2 in node 2, store P1 in node 3, store P2 in node 4. In addition, a virtual data block can be set and virtually stored in node 5. Among them, the virtual data block is a data block that actually does not exist, so its storage is also an actual non-existent storage. Only for the convenience of explanation and illustration, the virtual data block and node 5 are added here. As Figure 3 shown, the virtual data block is represented by a square box in the figure.

[0073] Before storing the second original data, multiple first original data can be stored in the above storage nodes. Taking storing 2 original data as an example for illustration, as Figure 3As shown, it includes related data blocks D1, D2, P1, and P2 of the first first original data, and also includes the second first original data, that is, the fourth original data. Among them, the second first original data includes: the first erasure data blocks D3 and D4, the first parity blocks P3 and P4. Store D1 and D3 in node 1, store D2 and D4 in node 2, store P1 and P3 in node 3, and store P2 and P4 in node 4. Among them, each related data block of the first original data also includes at least one virtual data block, and the virtual data block is virtually stored in node 5;

[0074] When the second original data is to be stored, first determine the second erasure data blocks of the second original data, that is, determine D5 and D6. When storing D5 and D6, since among the multiple nodes storing multiple first original data, there is at least a storage node for storing 2 virtual data blocks, that is, node 5. As long as node 5 is an actual node rather than a virtual node, the second erasure data blocks can be preferentially stored in node 5. As Figure 4 shown, store the second erasure data block D5 in the first stripe in node 5, that is, at the position where the virtual data block of the first first original data is stored, and store the second erasure data block D6 in the second stripe in node 5, that is, at the position where the virtual data block of the second first original data is stored. This eliminates the need to sequentially store the related data blocks of the second original data at the relevant positions of nodes 1 - 4.

[0075] In addition, since the second erasure data block D5 is stored at the storage position of the virtual data block of the first first original data, the second erasure data block D5 needs to share the same two parity blocks with the first erasure data blocks D1 and D2. Then, these two parity blocks need to be updated, that is, generate new first parity blocks P1' and new first parity blocks P2' based on the first erasure data block D1, the first erasure data block D2, and the second erasure data block D5, so that when one or two data blocks in the first stripe are abnormal, based on the non - abnormal data blocks and parity blocks, the normal data of the abnormal data blocks can be determined, ensuring data security;

[0076] At the same time, since the second erasure data block D6 is stored at the storage position of the virtual data block of the second first original data, the second erasure data block D6 needs to share the same two parity blocks with the first erasure data blocks D3 and D4. Then, these two parity blocks need to be updated, that is, generate new first parity blocks P3' and new first parity blocks P4' based on the first erasure data block D3, the first erasure data block D4, and the second erasure data block D6, so that when one or two data blocks in the second stripe are abnormal, based on the non - abnormal data blocks and parity blocks, the normal data of the abnormal data blocks can be determined, ensuring data security.

[0077] Based on this, when storing the first original data, its storage cost is: (k + m) / k = (2 + 2) / 2 = 2. Then, after storing the second original data through this solution, its storage cost is: (k + m) / k = (3 + 2) / 3 = 1.67. Its storage cost is reduced, achieving dynamic erasure code storage, that is, the parity blocks can change based on the change in the amount of stored data.

[0078] In addition, when storing the second original data, the first parity block is updated to obtain the second parity block, which can be: as shown in the above example, the second quantity of second parity blocks is jointly determined based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one second erasure data block. It can also be: since the second quantity of first parity blocks is obtained based on the first quantity of first erasure data blocks, when determining the second parity block, it can be determined only based on the second quantity of first parity blocks and at least one second erasure data block, without the need to re-obtain the first quantity of first erasure data blocks.

[0079] The erasure code storage method disclosed in this embodiment determines the first quantity of first erasure data blocks and the second quantity of first parity blocks based on the first original data to be stored, and stores the third quantity of data blocks to different storage nodes respectively. The third quantity of data blocks includes: the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one virtual data block. The second original data to be stored is obtained, and at least one second erasure data block determined based on the second original data is stored to the storage node where at least one virtual data block is stored. At least one virtual data block is deleted, and the second quantity of second parity blocks is determined based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one second erasure data block. The second quantity of first parity blocks is updated to the second quantity of second parity blocks. In this solution, the relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks. After there is new data to be stored, the second erasure data block corresponding to the new data to be stored is stored to the storage node of the virtual data block, so as to realize replacing the virtual data block with the second erasure data block. Only the parity block needs to be updated, and there is no need to set a new parity block for the second erasure data block again, achieving the reuse of the parity block and saving storage costs.

[0080] This embodiment discloses an erasure code storage method, and its flowchart is as Figure 5 shown, including:

[0081] Step S51, determine the first quantity of first erasure data blocks and the second quantity of first parity blocks based on the first original data to be stored;

[0082] Step S52: Store the first number of first erasure data blocks and the second number of first parity blocks in different cluster nodes respectively, and store at least one virtual data block in different virtual nodes;

[0083] Step S53: Obtain the second original data to be stored, determine that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, and determine the newly added expansion nodes;

[0084] Step S54: Store at least one second erasure data block determined based on the second original data in the expansion nodes, and delete the virtual nodes matching the number of expansion nodes;

[0085] Step S55: Determine the second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks and at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0086] Store the third number of data blocks, where the third number of data blocks at least includes: the first number of first erasure data blocks, the second number of first parity blocks and at least one virtual data block.

[0087] When storing the data blocks of the third data, it is usually to store the first number of first erasure data blocks and the second number of first parity blocks in different first cluster nodes respectively, where the first cluster nodes are physical storage nodes, that is, the first cluster nodes are actual existing cluster nodes; store at least one virtual data block in at least one second cluster node respectively, where the second cluster nodes can be physical storage nodes or virtual nodes.

[0088] Specifically, when the second cluster node is a physical storage node, actually when storing the first original data, there are already expansion nodes in the distributed storage system for storing data, but the first erasure data blocks and first parity blocks generated based on the first original data are not stored in the expansion nodes, but in the original nodes. For example: there are originally 4 storage nodes, and then 1 expansion node is added to the distributed storage system. When storing the first original data, the first erasure data blocks and first parity blocks of the first original data are stored in the original 4 storage nodes respectively, and a virtual data block is set and stored in this expansion node. Then, when there is second original data to be stored, directly store the second erasure data blocks of the second original data in the expansion node to replace the virtual data block;

[0089] In addition, if the second cluster node is a virtual node, then the second cluster node does not actually exist. At least when storing the first original data, the second cluster node is still a non-existent node. Storing the actual erasure data blocks and parity blocks in the physical nodes and storing the virtual data blocks in the virtual nodes will not have any impact on the stored erasure data blocks and parity blocks.

[0090] When there is second original data to be stored, first determine whether there are new expansion nodes in the current distributed storage system, that is, whether the difference between the number of current storage nodes in the system and the number of storage nodes when storing the first original data is greater than 0. If it is determined that the difference between the number of current storage nodes in the system and the number of storage nodes in the system when storing the first original data is equal to 0, it indicates that no new expansion nodes have been added to the system after storing the first original data. At this time, the relevant data blocks of the second original data to be stored can only be stored in the original storage nodes; if it is determined that the difference between the number of current storage nodes in the system and the number of storage nodes in the system when storing the first original data is greater than 0, it indicates that there are new expansion nodes in the system after storing the first original data. At this time, the relevant data blocks of the second original data can be stored in the expansion nodes.

[0091] Specifically, if it is determined that there are new expansion nodes when storing the second original data, then when storing the first original data, it is actually: storing the first erasure data block and the first parity block of the first original data in their actual existing cluster nodes respectively, and setting at least one virtual data block, which is respectively stored virtually in at least one virtual node; when there are new expansion nodes in the distributed storage system, replace the virtual nodes with the expansion nodes. When there is second original data to be stored, store the second erasure data block of the second original data in the expansion nodes, thereby realizing the replacement of the virtual nodes and virtual data blocks, and sharing the same parity block for the data blocks stored different times through different nodes in the same stripe, avoiding the problem of occupying storage space and increasing storage costs caused by setting parity blocks separately for each stored data.

[0092] In addition, when the number of newly added expansion nodes is relatively large in a certain time, and the number of expansion nodes is greater than or equal to the number of second erasure data blocks of the second original data to be stored, then when storing the second erasure data blocks, preferentially store the second erasure data blocks in the first stripe. When the first stripe is full, if there is other newly added original data, it will only then occupy the storage positions in the second stripe of each storage node.

[0093] For example: Figure 6Taking the example shown below, before there are newly added extended nodes, the distributed storage system only includes Node 1, Node 2, Node 3, and Node 4, and the first stripe and the second stripe of the above 4 nodes are both full; when there are 2 newly added extended nodes and there is second original data to be stored, the second erasure data block D5 of the second original data is preferentially stored at the position of the first stripe of extended node 1, and the second erasure data block D6 of the second original data is stored at the position of the first stripe of extended node 2, while the positions of the second stripes of extended node 1 and extended node 2 are still virtual data blocks, and the parity blocks P1 and P2 are updated to P1' and P2' based on D1, D2, D5, and D6.

[0094] Preferentially occupy the storage positions in the first stripe, which enables when updating the parity blocks, if the newly added stored data blocks are few, only the parity blocks in the first stripe need to be updated, without having to update the parity blocks in multiple stripes, reducing the data processing volume of the system.

[0095] The erasure code storage method disclosed in this embodiment determines a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, stores the third number of data blocks in different storage nodes respectively, and the third number of data blocks includes: a first number of first erasure data blocks, a second number of first parity blocks, and at least one virtual data block, obtains the second original data to be stored, stores at least one second erasure data block determined based on the second original data in the storage node where at least one virtual data block is stored, deletes at least one virtual data block, determines a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and at least one second erasure data block, and updates the second number of first parity blocks to the second number of second parity blocks. In this solution, the relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks, and after there is new data to be stored, the second erasure data block corresponding to the new data to be stored is stored in the storage node of the virtual data block, so as to realize replacing the virtual data block with the second erasure data block, only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data block again, realizing the reuse of parity blocks and saving storage costs.

[0096] This embodiment discloses an erasure code storage method, and its flowchart is as Figure 7 shown, including:

[0097] Step S71, determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store the third number of data blocks in different storage nodes respectively, and the data blocks of the third data include: a first number of first erasure data blocks, a second number of first parity blocks, and at least one virtual data block;

[0098] Step S72: Obtain the second original data to be stored;

[0099] Step S73: Store at least one second erasure data block determined based on the second original data into the storage nodes of at least one virtual data block storage, and delete at least one virtual data block;

[0100] Step S74: Determine the second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks;

[0101] Step S75: Obtain the third original data to be stored;

[0102] Step S76: If it is determined that the number of current cluster nodes is less than the number of cluster nodes when obtaining the second original data, store the third erasure data block and the third parity block determined based on the third original data into different cluster nodes respectively, and store at least one virtual data block related to the third original data into different virtual nodes respectively.

[0103] When obtaining the third original data to be stored, it is first necessary to determine whether the number of current storage nodes in the distributed storage system when obtaining the third original data is less than the number of cluster nodes when obtaining the second original data, or less than the number of cluster nodes when obtaining the first original data.

[0104] If it is determined that when obtaining the third original data to be stored, the number of storage nodes in the system is the same as the number of cluster nodes when obtaining the second original data, then store the relevant data blocks of the third original data in the same way as the relevant data blocks of the second original data;

[0105] If it is determined that when obtaining the third original data to be stored, the number of storage nodes in the system is greater than the number of cluster nodes when obtaining the second original data, then the relevant data blocks of the third original data can also be stored in the same way as the relevant data blocks of the second original data;

[0106] If it is determined that when obtaining the third original data to be stored, the number of storage nodes in the system is less than the number of cluster nodes when obtaining the second original data, then the relevant data blocks of the third original data can be stored into the existing cluster nodes respectively, and virtual data blocks and virtual nodes are set, and the virtual data blocks are stored into the virtual nodes, so that when there are newly added expansion nodes, the virtual nodes can be replaced by the newly added expansion nodes, and when there are other data blocks to be stored continuously, the other data blocks are preferentially stored into the newly added expansion nodes.

[0107] When obtaining the third original data to be stored, if the number of storage nodes in the system is greater than the number of cluster nodes when storing the first original data, store the relevant data blocks of the third original data according to the storage method of the relevant data blocks of the second original data; if it is equal to the number of cluster nodes when storing the first original data, store the relevant data blocks of the third original data according to the storage method of the relevant data blocks of the first original data; if it is less than the number of cluster nodes when storing the first original data, store the relevant data blocks of the third original data into the existing cluster nodes respectively.

[0108] Furthermore, it also includes:

[0109] Receive the deletion instruction of the first cluster node; determine no less than one first data block stored in the first cluster node to be deleted based on the deletion instruction, and store no less than one first data block into other cluster nodes in the cluster except the first cluster node; at the same time, update the parity blocks stored in other cluster nodes in the cluster except the first cluster node.

[0110] Specifically, when receiving the deletion instruction, it indicates that there is a cluster node to be deleted in the distributed storage system, and it is necessary to determine whether there are data blocks stored in the cluster node to be deleted. If there are no data blocks stored in the cluster node to be deleted, the cluster node can be directly deleted;

[0111] If there are data blocks stored in the cluster node to be deleted, it is necessary to first transfer the data blocks it stores to the cluster nodes that do not need to be deleted, that is, transfer the data blocks stored in the cluster node to be deleted to other cluster nodes in the distributed storage system except the cluster node to be deleted. After the transfer, since the data structure stored in each storage node in the distributed storage system changes, it is necessary to update the parity blocks stored in each storage node. For example: one storage location is reduced in the first stripe of each storage node in the distributed storage system, that is, one data block is reduced, then update the parity block based on the existing erasure-coded data blocks in the first stripe; if one data block is added in the third stripe of the distributed storage system, it is necessary to update the parity block based on the existing erasure-coded data blocks in the third stripe, so as to ensure that the corresponding original data can be obtained based on at least part of the data in each stripe.

[0112] The erasure code storage method disclosed in this embodiment determines a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and stores data blocks of a third number in different storage nodes respectively. The data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block. Obtain the second original data to be stored, store at least one second erasure data block determined based on the second original data in the storage node where the at least one virtual data block is stored, delete the at least one virtual data block, determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks. In this solution, the relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks. After there is new data to be stored, the second erasure data blocks corresponding to the new data to be stored are stored in the storage nodes of the virtual data blocks, so as to replace the virtual data blocks with the second erasure data blocks. Only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data blocks again, realizing the reuse of parity blocks and saving storage costs.

[0113] This embodiment discloses an erasure code storage method, and its flowchart is as Figure 8 shown, including:

[0114] Step S81: Determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store data blocks of a third number in different storage nodes respectively. The data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block;

[0115] Step S82: Obtain the second original data to be stored;

[0116] Step S83: Store at least one second erasure data block determined based on the second original data in the positions corresponding to different stripes in the storage node for storing virtual data blocks in sequence;

[0117] Step S84: Determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0118] When storing at least one second erasure data block of the second original data, it is sequentially stored at positions corresponding to different stripes in the storage nodes for storing virtual data blocks. That is, if there is 1 storage node for storing virtual data blocks, and there are 2 positions corresponding to stripes in this node for storing virtual data blocks, then 2 second erasure data blocks are determined from all the second erasure data blocks, and the determined 2 second erasure data blocks are respectively used to replace the virtual data blocks and stored at the 2 positions corresponding to the 2 stripes in this node;

[0119] Or, if there are 2 storage nodes for storing virtual data blocks, and there are 2 positions corresponding to stripes in each of the 2 nodes for storing virtual data blocks, and if the number of second erasure data blocks is at least 4, that is, the number of second erasure data blocks is greater than the number of stripes storing data blocks, then first 4 are selected from all the second erasure data blocks to replace the virtual data blocks and are respectively stored at the 4 positions corresponding to the 2 stripes in the 2 nodes;

[0120] If the number of second erasure data blocks is 2 or 3, then first 2 or 3 are selected from all the second erasure data blocks. 2 of the second erasure data blocks are stored at the 2 positions corresponding to the 2 stripes in the first node, and then the remaining third second erasure data block is stored at the 1 position corresponding to 1 of the stripes in the second node, that is, the first node is preferentially stored, and after the first node is stored, the second node is stored.

[0121] Or, if the number of second erasure data blocks is greater than the number of stripes storing data blocks, the same number of second erasure data blocks as the number of stripes storing data blocks are stored at positions of different stripes in the storage nodes for storing virtual data blocks, and the second erasure data blocks not stored in the storage nodes for storing virtual data blocks are sequentially stored in the third stripes in each storage node where no data blocks are stored, and parity blocks are generated based on the data blocks stored in the third stripes and stored in the third stripes.

[0122] That is, if there is 1 expansion node, when the number of second erasure data blocks is greater than the number of stripes storing data blocks, the same number of second erasure data blocks as the number of stripes are stored at positions in different stripes in this expansion node, and the remaining second erasure data blocks are stored at positions in the storage node corresponding to the new stripes.

[0123] For example: Figure 3Taking it as an example, the strip where D1, D2, P1, and P2 are located is the first strip, the strip where D3, D4, P3, and P4 are located is the second strip, node 5 is an extended node, and two squares in node 5 store virtual data blocks. When there is new raw data to be stored, a second erasure data block is generated. If the number of second erasure data blocks is 2, the 2 second erasure data blocks are respectively stored in the two squares in node 5; if the number of second erasure data blocks is greater than 2, 2 of the second erasure data blocks are respectively stored in the two squares in node 5, and the remaining second erasure data blocks that are not stored in the above two strips are stored in the third condition, that is, in other strips that have not stored data except the above first strip and second strip.

[0124] In addition, if there are multiple extended nodes, it is necessary to determine the number of positions where all strips intersect with all extended nodes, that is, the number of virtual data blocks existing in the strips storing data. When the number of second erasure data blocks is not greater than the number of virtual data blocks, all the second erasure data blocks are stored in the storage positions of the virtual data blocks; when the number of second erasure data blocks is greater than the number of virtual data blocks, the same number of second erasure data blocks as the number of virtual data blocks are stored in the positions in the extended nodes in different strips, and the remaining second erasure data blocks are stored in the positions in the storage nodes corresponding to the new strips.

[0125] Among them, the position in the storage node corresponding to the new strip, and the new strip is the strip that has not stored data before storing the second erasure data.

[0126] If the number of positions in the new strip where data blocks can be stored is greater than the number of remaining second erasure data blocks, after all the remaining second erasure data blocks are stored in the new strip, the remaining positions in the strip are used to store virtual data blocks, and the parity blocks in the strip are determined based on the virtual data blocks and the second erasure data blocks stored in the strip.

[0127] The erasure code storage method disclosed in this embodiment determines a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and stores data blocks of a third number in different storage nodes respectively. The data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block. The second original data to be stored is obtained, and at least one second erasure data block determined based on the second original data is stored in the storage node where at least one virtual data block is stored, and at least one virtual data block is deleted. A second number of second parity blocks are determined based on the first number of first erasure data blocks, the second number of first parity blocks, and at least one second erasure data block, and the second number of first parity blocks are updated to the second number of second parity blocks. In this solution, the relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks. After there is new data to be stored, the second erasure data blocks corresponding to the new data to be stored are stored in the storage nodes of the virtual data blocks, so as to replace the virtual data blocks with the second erasure data blocks. Only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data blocks again, realizing the reuse of parity blocks and saving storage costs.

[0128] This embodiment discloses an electronic device, and its structural schematic diagram is as Figure 9 shown, including:

[0129] a processor 91 and a memory 92.

[0130] Among them, the processor 91 is used to determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store data blocks of a third number in different storage nodes respectively. The data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block; obtain the second original data to be stored, store at least one second erasure data block determined based on the second original data in the storage node where at least one virtual data block is stored, and delete at least one virtual data block; determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks;

[0131] The memory 92 is used to store the program for the processor to execute the above processing process.

[0132] Further, the processor stores data blocks of the third number in different storage nodes respectively, including:

[0133] The processor stores the first quantity of first erasure data blocks and the second quantity of first parity blocks in different cluster nodes respectively, and stores at least one virtual data block in different virtual nodes.

[0134] Then the processor stores at least one second erasure data block in the storage nodes where at least one virtual data block is stored, including:

[0135] When obtaining the second original data, the processor determines that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, and determines the newly added expansion nodes; stores at least one second erasure data block determined based on the second original data in the expansion nodes, and deletes the virtual nodes matching the number of expansion nodes.

[0136] Further, the processor stores the third quantity of data blocks in different storage nodes respectively, including:

[0137] The processor stores the first quantity of first erasure data blocks and the second quantity of first parity blocks in different first cluster nodes respectively, and stores at least one virtual data block in at least one second cluster node; wherein, the second cluster node is a physical storage node or a virtual node.

[0138] Further, the processor is also used for:

[0139] Obtain the third original data to be stored; if it is determined that the number of current cluster nodes is less than the number of cluster nodes when obtaining the second original data, store the third erasure data blocks and the third parity blocks determined based on the third original data in different cluster nodes respectively, and store at least one virtual data block related to the third original data in different virtual nodes.

[0140] Further, the processor is also used for:

[0141] Receive the deletion instruction of the first cluster node; determine no less than one first data block stored in the first cluster node to be deleted based on the deletion instruction, store no less than one first data block in other cluster nodes in the cluster except the first cluster node respectively; update the parity blocks stored in other cluster nodes in the cluster except the first cluster node.

[0142] Further, the processor stores the third quantity of data blocks in different storage nodes respectively, including:

[0143] The processor stores the third quantity of data blocks into the first stripe of each storage node respectively; and stores the third quantity of data blocks determined based on the fourth original data to be stored into the second stripe of each storage node, wherein the third quantity of data blocks determined based on the fourth original data to be stored includes: the first quantity of first erasure data blocks related to the fourth original data, the second quantity of first parity blocks related to the fourth original data, and at least one virtual data block.

[0144] Further, the processor stores at least one second erasure data block into the storage node where at least one virtual data block is stored, including:

[0145] The processor sequentially stores at least one second erasure data block at positions corresponding to different stripes in the storage node for storing virtual data blocks.

[0146] Further, the processor sequentially stores at least one second erasure data block at positions corresponding to different stripes in the storage node for storing virtual data blocks, including:

[0147] If the quantity of the second erasure data blocks is greater than the quantity of the stripes storing data blocks, the processor stores the same quantity of second erasure data blocks as the quantity of the stripes storing data blocks at positions corresponding to different stripes in the storage node for storing virtual data blocks; stores the second erasure data blocks not stored into the storage node for storing virtual data blocks sequentially into the third stripe of each storage node where no data block is stored; generates parity blocks based on the data blocks stored into the third stripe, and stores them into the third stripe.

[0148] The electronic device disclosed in this embodiment is implemented based on the erasure code storage method disclosed in the above embodiment, which will not be elaborated here.

[0149] The electronic device disclosed in this embodiment determines a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, stores data blocks of a third number in different storage nodes respectively, and the data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block, obtains the second original data to be stored, stores at least one second erasure data block determined based on the second original data in the storage node where the at least one virtual data block is stored, deletes the at least one virtual data block, determines a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and updates the second number of first parity blocks to the second number of second parity blocks. In this solution, the relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks, and after there is new data to be stored, the second erasure data blocks corresponding to the new data to be stored are stored in the storage nodes of the virtual data blocks, so as to replace the virtual data blocks with the second erasure data blocks. Only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data blocks again, realizing the reuse of parity blocks and saving storage costs.

[0150] This embodiment discloses an erasure code storage system, and its structural schematic diagram is as Figure 10 shown, including:

[0151] A first storage unit 101, an obtaining unit 102, a second storage unit 103, and a determining unit 104.

[0152] Among them, the first storage unit 101 is used to determine a first number of first erasure data blocks and a second number of first parity blocks based on the first original data to be stored, and store data blocks of a third number in different storage nodes respectively. The data blocks of the third number include: the first number of first erasure data blocks, the second number of first parity blocks, and at least one virtual data block;

[0153] The obtaining unit 102 is used to obtain the second original data to be stored;

[0154] The second storage unit 103 is used to store at least one second erasure data block determined based on the second original data in the storage node where the at least one virtual data block is stored, and delete the at least one virtual data block;

[0155] The determining unit 104 is used to determine a second number of second parity blocks based on the first number of first erasure data blocks, the second number of first parity blocks, and the at least one second erasure data block, and update the second number of first parity blocks to the second number of second parity blocks.

[0156] The erasure code storage system disclosed in this embodiment is implemented based on the erasure code storage method disclosed in the above embodiment, and will not be elaborated here.

[0157] The erasure code storage system disclosed in this embodiment determines a first quantity of first erasure data blocks and a second quantity of first parity blocks based on the first original data to be stored, and stores data blocks in a third quantity in different storage nodes respectively. The data blocks in the third quantity include: the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one virtual data block. The second original data to be stored is obtained, at least one second erasure data block determined based on the second original data is stored in the storage node where the at least one virtual data block is stored, the at least one virtual data block is deleted, a second quantity of second parity blocks is determined based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one second erasure data block, and the second quantity of first parity blocks is updated to the second quantity of second parity blocks. In this solution, relevant data blocks of a certain data are stored through virtual data blocks, parity blocks, and erasure data blocks. After there is new data to be stored, the second erasure data block corresponding to the new data to be stored is stored in the storage node of the virtual data block, so as to replace the virtual data block with the second erasure data block. Only the parity blocks need to be updated, and there is no need to set parity blocks for the second erasure data blocks again, realizing the reuse of parity blocks and saving storage costs.

[0158] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0159] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0160] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0161] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An erasure code storage method, comprising: Determining a first quantity of first erasure data blocks and a second quantity of first parity blocks based on first original data to be stored, and storing data blocks of a third quantity in different storage nodes respectively, where the data blocks of the third quantity include: the first quantity of first erasure data blocks, the second quantity of first parity blocks, and at least one virtual data block; wherein, the storage nodes storing the virtual data blocks include extended nodes; Obtaining second original data to be stored, determining that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, and determining newly added extended nodes; Storing at least one second erasure data block determined based on the second original data in the extended nodes, and deleting virtual nodes matching the number of the extended nodes; Determining a second quantity of second parity blocks based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and the at least one second erasure data block, and updating the second quantity of first parity blocks to the second quantity of second parity blocks.

2. The method according to claim 1, wherein, The storing the data blocks of the third quantity in different storage nodes respectively includes: Storing the first quantity of first erasure data blocks and the second quantity of first parity blocks in different cluster nodes respectively, and storing the at least one virtual data block in different virtual nodes respectively.

3. The method according to claim 1, wherein, The storing the data blocks of the third quantity in different storage nodes respectively includes: Storing the first quantity of first erasure data blocks and the second quantity of first parity blocks in different first cluster nodes respectively, and storing the at least one virtual data block in at least one second cluster node respectively; wherein, the second cluster node is a physical storage node or a virtual node.

4. The method according to claim 1, wherein, It further includes: Obtaining third original data to be stored; If it is determined that the number of current cluster nodes is less than the number of cluster nodes when obtaining the second original data, storing third erasure data blocks and third parity blocks determined based on the third original data in different cluster nodes respectively, and storing at least one virtual data block related to the third original data in different virtual nodes respectively.

5. The method according to claim 1, wherein It further includes: Receiving a deletion instruction from a first cluster node; Determining no less than one first data block stored in the first cluster node to be deleted based on the deletion instruction, and storing the no less than one first data block in other cluster nodes in the cluster except the first cluster node respectively; Updating the parity blocks stored in other cluster nodes in the cluster except the first cluster node.

6. The method according to claim 1, wherein, The storing the data blocks of the third quantity in different storage nodes respectively includes: Storing the data blocks of the third quantity in the first stripe of each storage node; Storing data blocks of the third quantity determined based on fourth original data to be stored in the second stripe of each storage node respectively, where the data blocks of the third quantity determined based on the fourth original data to be stored include: the first quantity of first erasure data blocks related to the fourth original data, the second quantity of first parity blocks related to the fourth original data, and at least one virtual data block.

7. The method according to claim 6, further comprising: Store the at least one second erasure-coded data block sequentially at positions corresponding to different stripes in the storage nodes for storing the virtual data block.

8. The method according to claim 7, wherein The storing the at least one second erasure-coded data block sequentially at positions corresponding to different stripes in the storage nodes for storing the virtual data block includes: If the number of the second erasure-coded data blocks is greater than the number of stripes storing data blocks, store the same number of second erasure-coded data blocks as the number of stripes storing data blocks at positions of different stripes in the storage nodes for storing the virtual data block; Sequentially store the second erasure-coded data blocks not stored in the storage nodes for storing the virtual data block into the third stripes in each storage node where no data block is stored; Generate a parity block based on the data blocks stored in the third stripe and store it in the third stripe.

9. An electronic device, comprising: A processor, configured to determine a first number of first erasure-coded data blocks and a second number of first parity blocks based on first original data to be stored, store a third number of data blocks in different storage nodes respectively, where the third number of data blocks includes: the first number of first erasure-coded data blocks, the second number of first parity blocks and at least one virtual data block; wherein, the storage nodes storing the virtual data block include expansion nodes; obtain second original data to be stored, determine that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, determine newly added expansion nodes, store at least one second erasure-coded data block determined based on the second original data in the expansion nodes, and delete virtual nodes matching the number of the expansion nodes; determine a second number of second parity blocks based on the first number of first erasure-coded data blocks, the second number of first parity blocks and the at least one second erasure-coded data block, and update the second number of first parity blocks to the second number of second parity blocks; A memory, configured to store a program for the processor to execute the above processing procedure.

10. An erasure code storage system, comprising: A first storage unit, configured to determine a first number of first erasure-coded data blocks and a second number of first parity blocks based on first original data to be stored, store a third number of data blocks in different storage nodes respectively, where the third number of data blocks includes: the first number of first erasure-coded data blocks, the second number of first parity blocks and at least one virtual data block; wherein, the storage nodes storing the virtual data block include expansion nodes; An obtaining unit, configured to obtain second original data to be stored, determine that the number of current cluster nodes is greater than the number of cluster nodes when obtaining the first original data, and determine newly added expansion nodes; A second storage unit, configured to store at least one second erasure-coded data block determined based on the second original data in the expansion nodes, and delete virtual nodes matching the number of the expansion nodes; A determination unit is configured to determine a second quantity of second parity blocks based on the first quantity of first erasure data blocks, the second quantity of first parity blocks, and the at least one second erasure data block, and update the second quantity of first parity blocks to the second quantity of second parity blocks.

Citation Information

Patent Citations

  • Data storage method, device and system and storage medium

    CN112019788A