Erasure code data storage method and device and cloned volume data operation method and device

By writing data existence flags in the erasure code redundant group and storing non-full striped data in the form of a copy, the problems of low efficiency of conditional read and write query and waste of storage space in the prior art are solved, and efficient data storage and query are realized.

CN120216253APending Publication Date: 2025-06-27CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256405.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing erasure code redundancy strategy has insufficient conditions in terms of read and write query efficiency, and it is easy to lead to serious waste of storage capacity and space.

Method used

By writing data existence flags to each shard in the erasure code redundant group and storing them in the form of a copy when writing original data with non-full stripes, avoiding zero-compensation operations, thereby improving the query efficiency of conditional read and write.

Benefits of technology

While avoiding wasting storage capacity and space, it improves the efficiency of conditional read, write and query, and improves the performance and reliability of the data storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216253A_ABST
    Figure CN120216253A_ABST
Patent Text Reader

Abstract

The invention provides an erasure code data storage method and a clone volume data operation method and device.The method comprises the steps that when original data of a full stripe is written into a redundancy group, each original data block is written into a corresponding data fragment in the redundancy group, and each check data block is written into a corresponding check fragment in the redundancy group; when original data of a non-full stripe is written into the redundancy group, each original data block is written into a corresponding data fragment in the redundancy group, and all the original data blocks are written into each verification fragment in the redundancy group; for any original data written into the redundancy group, a corresponding data existence mark is written into each fragment in the redundancy group, and the data existence mark corresponds to a key of the original data. According to the method and the device, the condition read-write query efficiency can be improved on the premise of avoiding serious storage capacity space waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed storage technology, and in particular, to an erasure code data storage method, a clone volume data operation method, and an apparatus therefor. Background Art

[0002] With the increasing penetration of informatization, the amount of data generated by the whole society every day shows an explosive growth. Therefore, people's demands for the reliability and availability of data storage have become increasingly urgent. Replication and erasure code are two common redundancy strategies in the field of distributed storage. Among them, erasure code is used in business scenarios with low performance requirements and high storage capacity utilization requirements, such as video, archiving, and other storage scenarios.

[0003] In some application scenarios, before reading and writing data, it is necessary to first query whether the data corresponding to the key exists to ensure the correct reading and writing of the data. Such reading and writing operations are called conditional reading and writing. For the replication redundancy strategy, it is only necessary to query any one shard in the redundancy group, ensuring the query efficiency of conditional reading and writing. For the erasure code redundancy strategy, existing solutions either cause serious waste of storage capacity space or reduce the query efficiency of conditional reading and writing. Summary of the Invention

[0004] This application provides an erasure code data storage method, a clone volume data operation method, and an apparatus therefor, which can improve the query efficiency of conditional reading and writing on the premise of avoiding serious waste of storage capacity space.

[0005] In a first aspect, an embodiment of this application provides an erasure code data storage method, and the erasure code data storage method includes:

[0006] When writing full-strip original data into a redundancy group, write each original data block into a corresponding data shard in the redundancy group, and write each parity data block into a corresponding parity shard in the redundancy group;

[0007] When writing non-full-strip original data into a redundancy group, write each original data block into a corresponding data shard in the redundancy group, and write all original data blocks into each parity shard in the redundancy group;

[0008] For any original data written into the redundancy group, write a corresponding data existence flag into each shard in the redundancy group, where the data existence flag corresponds to the key of the original data.

[0009] Further, in an embodiment, the step of writing a corresponding data existence flag into each shard in the redundancy group for any original data written into the redundancy group includes:

[0010] When writing any original data to a redundant group, query whether there is a corresponding data existence flag on any target shard in the redundant group, where the target shards include check shards and data shards corresponding to the original data blocks;

[0011] If it is queried that there is no corresponding data existence flag, write the corresponding data existence flag to each shard in the redundant group.

[0012] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group for any original data written to the redundant group includes:

[0013] When writing any original data to a redundant group, query whether there is a corresponding data existence flag on any shard in the redundant group;

[0014] If it is queried that there is no corresponding data existence flag, write the corresponding data existence flag to each shard in the redundant group.

[0015] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group for any original data written to the redundant group includes:

[0016] When writing any original data to a redundant group, write the corresponding data existence flag to each shard in the redundant group.

[0017] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group includes:

[0018] When writing an original data block or a check data block to each target shard in the redundant group, write the corresponding data existence flag to the corresponding target shard, where the target shards include check shards and data shards corresponding to the original data blocks;

[0019] When the number of target shards is less than the number of shards in the redundant group, specify at most the number of target shards of shards from the first shards to be written as the forwarding targets for this round, specify a target shard as the forwarding source for this round for each forwarding target for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the first shards to be written include other shards in the redundant group except the target shards;

[0020] When twice the number of forwarding sources in the previous round is less than the number of redundant component shards, select at most twice the number of forwarding sources in the previous round from the second shards to be written as the forwarding targets in this round, and assign an upper round forwarding source or an upper round forwarding target as the forwarding source in this round for each forwarding target in this round, so that each forwarding source in this round forwards the corresponding data existence flag to a corresponding forwarding target in this round for the forwarding target in this round to write the corresponding data existence flag, where the second shards to be written include other shards in the redundant group except the upper round forwarding sources and upper round forwarding targets.

[0021] In a second aspect, an embodiment of the present application further provides an erasure code data storage device, where the erasure code data storage device includes:

[0022] A first data writing module, configured to write each original data block into a corresponding data shard in the redundant group and write each check data block into a corresponding check shard in the redundant group when writing full-strip original data into the redundant group;

[0023] A second data writing module, configured to write each original data block into a corresponding data shard in the redundant group and write all original data blocks into each check shard in the redundant group when writing non-full-strip original data into the redundant group;

[0024] A flag writing module, configured to write a corresponding data existence flag into each shard in the redundant group for any original data written into the redundant group, where the data existence flag corresponds to the key of the original data.

[0025] Further, in one embodiment, the flag writing module is configured to:

[0026] When writing any original data into the redundant group, query whether there is a corresponding data existence flag on any target shard in the redundant group, where the target shard includes the check shard and the data shard corresponding to the original data block;

[0027] If it is queried that there is no corresponding data existence flag, write the corresponding data existence flag into each shard in the redundant group.

[0028] Further, in one embodiment, the flag writing module is configured to:

[0029] When writing an original data block or a check data block into each target shard in the redundant group, write the corresponding data existence flag into the corresponding target shard;

[0030] When the number of target shards is less than the number of redundant component shards, at most the number of target shards is specified from the first shards to be written as the forwarding targets for this round. A target shard is specified for each forwarding target for this round as the forwarding source for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the first shards to be written include the other shards in the redundant group except the target shards;

[0031] When twice the number of forwarding sources in the previous round is less than the number of shards in the redundant group, at most twice the number of shards of the number of forwarding sources in the previous round is selected from the second shards to be written as the forwarding targets for this round. An upper-round forwarding source or an upper-round forwarding target is specified for each forwarding target for this round as the forwarding source for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the second shards to be written include the other shards in the redundant group except the upper-round forwarding sources and the upper-round forwarding targets.

[0032] In a third aspect, an embodiment of the present application further provides a method for operating clone volume data. The original volume and the clone volume store data using the above erasure code data storage method. The method for operating clone volume data includes:

[0033] When a read / write instruction for the clone volume is received during the copy operation from the original volume to the clone volume, it is queried whether there is a corresponding data existence flag on any shard in the redundant group corresponding to the clone volume;

[0034] If there is no corresponding data existence flag on the queried shard in the clone volume, it is queried whether there is a corresponding data existence flag on any shard in the redundant group corresponding to the original volume;

[0035] If there is a corresponding data existence flag on the queried shard in the clone volume, or there is no corresponding data existence flag on the queried shards in both the clone volume and the original volume, subsequent operations are performed according to the read / write instruction;

[0036] If there is no corresponding data existence flag on the queried shard in the clone volume and there is a corresponding data existence flag on the queried shard in the original volume, the corresponding original data is copied from the original volume to the clone volume, and subsequent operations are performed according to the read / write instruction.

[0037] In a fourth aspect, an embodiment of the present application further provides a device for operating clone volume data. The original volume and the clone volume store data using the above erasure code data storage method. The device for operating clone volume data includes:

[0038] The first query module is used to query whether there is a corresponding data existence flag on any one of the shards in the redundant group corresponding to the clone volume when a read / write instruction for the clone volume is received during the copy operation from the original volume to the clone volume;

[0039] The second query module is used to query whether there is a corresponding data existence flag on any one of the shards in the redundant group corresponding to the original volume if there is no corresponding data existence flag on the shard queried in the clone volume;

[0040] The first execution module is used to perform subsequent operations according to the read / write instruction if there is a corresponding data existence flag on the shard queried in the clone volume, or if there is no corresponding data existence flag on the shards queried in both the clone volume and the original volume;

[0041] The second execution module is used to copy the corresponding original data from the original volume to the clone volume and perform subsequent operations according to the read / write instruction if there is no corresponding data existence flag on the shard queried in the clone volume and there is a corresponding data existence flag on the shard queried in the original volume.

[0042] In this application, when writing non-full-strip original data into the redundant group, each original data block is written into a corresponding data shard in the redundant group, and all original data blocks are written into each parity shard in the redundant group. The conventional zero-padding operation is not adopted, avoiding serious waste of storage capacity space. For any original data written into the redundant group, a data existence flag corresponding to the key of the original data is written into each shard in the redundant group. During conditional read / write, it is only necessary to query whether there is a corresponding data existence flag on any one of the shards in the redundant group to determine whether the original data corresponding to the key exists in the redundant group. Through this application, the conditional read / write query efficiency can be improved on the premise of avoiding serious waste of storage capacity space. Description of the Drawings

[0043] Figure 1 It is a schematic flowchart of an erasure code data storage method in an embodiment of this application;

[0044] Figure 2 It is a schematic flowchart of writing full-strip original data into the redundant group in an embodiment of this application;

[0045] Figure 3 It is a schematic flowchart of writing non-full-strip original data into the redundant group in an embodiment of this application;

[0046] Figure 4 It is a schematic flowchart of data existence flag forwarding in an embodiment of this application;

[0047] Figure 5 It is a schematic functional module diagram of an erasure code data storage device in an embodiment of this application;

[0048] Figure 6 This is a schematic flowchart of the method for operating clone volume data in an embodiment of the present application.

[0049] Figure 7 This is a schematic diagram of the functional modules of the device for operating clone volume data in an embodiment of the present application. Detailed implementation manners

[0050] To enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0051] First, some technical terms in the present application are explained to facilitate the understanding of the present application by those skilled in the art.

[0052] Distributed storage system: A data storage system that disperses data storage on multiple independent devices. Traditional network storage systems use a centralized storage server to store all data. The storage server becomes the bottleneck of system performance and the focus of reliability and security, and cannot meet the needs of large-scale storage applications. The distributed network storage system adopts an extensible system structure, uses multiple storage servers to share the storage load, and uses a location server to locate storage information. It not only improves the reliability, availability, and access efficiency of the system, but also is easy to expand.

[0053] Erasure code data storage: A data protection method that divides the original data into k original data blocks according to a specified size, then encodes them into p parity data blocks according to the redundancy level, and finally stores the original data blocks and parity data blocks in different locations to achieve the purpose of data protection.

[0054] Replica data storage: A data protection method that takes the original data as a whole and stores n identical copies of the data at n different locations according to the redundancy level to achieve the purpose of data protection.

[0055] Clone volume: In the storage field, a clone volume is an integrity copy of the data of the original volume at a certain point in time. The clone volume allows users to start a new branch for reading and writing operations on the copied data without affecting the data of the original volume. This feature makes clone volumes widely used in data backup, recovery, and testing.

[0056] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0057] In a first aspect, an embodiment of this application provides an erasure code data storage method.

[0058] Figure 1 The flowchart of the erasure code data storage method in an embodiment of this application is shown.

[0059] Referring to Figure 1 , in one embodiment, the erasure code data storage method includes the following steps:

[0060] S11. When writing full-strip original data to a redundancy group, write each original data block to a corresponding data shard in the redundancy group, and write each parity data block to a corresponding parity shard in the redundancy group.

[0061] Specifically, after the foreground write IO arrives, the control flow splits the original data into strips according to information such as the IO offset, data length, chunksize, and redundancy level. Among them, the strip size can be determined according to the redundancy level and chunksize.

[0062] Figure 2 The flowchart of writing full-strip original data to a redundancy group in an embodiment of this application is shown.

[0063] Referring to Figure 2 , the redundancy level of the erasure code is 4+2, the chunksize is equal to 4K. Correspondingly, the redundancy group includes 4 data shards and 2 parity shards, and the strip size is 16K.

[0064] When writing full-strip original data to a redundancy group, this embodiment is the same as the conventional solution. The original data will be split into 4 original data blocks A, B, C, D with a size of 4K. At the same time, the original data is encoded according to the erasure code algorithm to calculate 2 parity data blocks P, Q with a size of 4K. Write the original data blocks A, B, C, D to the data shards 0, 1, 2, 3 one by one, and write the parity data blocks P, Q to the parity shards P, Q one by one.

[0065] S12. When writing non-full-strip original data to a redundancy group, write each original data block to a corresponding data shard in the redundancy group, and write all the original data blocks to each parity shard in the redundancy group.

[0066] Figure 3 The flowchart of writing non-full-strip original data to a redundancy group in an embodiment of this application is shown.

[0067] Referring to Figure 3, the redundancy level of the erasure code and the chunksize are the same as Figure 2 .

[0068] When writing non-full-strip original data to a redundancy group, this embodiment is inconsistent with the conventional scheme. The traditional scheme will pad the original data with zeros to make it fill the strip, and then adopt the Figure 2 shown full-strip writing process. In this embodiment, the original data will be sliced into at least one original data block according to the actual size. Figure 3 In the example, the 10K-sized original data is sliced into 2 4K-sized original data blocks A, B, and 1 2K-sized original data block C. Without calculating the parity data block, the original data blocks A, B, and C are written to data shards 0, 1, and 2 one by one. Taking the original data blocks A, B, and C as a whole, they are written to parity shards P and Q in the form of replicas. In this way, any 4 shards in the redundancy group can be read to obtain the original data, meeting the redundancy level requirement of 4+2. At the same time, since there is no need to pad with zeros and calculate parity shards, the writing performance is improved.

[0069] S13. For any original data written to the redundancy group, write the corresponding data existence flag to each shard in the redundancy group, where the data existence flag corresponds to the key of the original data.

[0070] In the prior art, corresponding metadata is also written when writing original data, and the metadata of the original data includes the key of the original data. For the full-strip original data written to the redundancy group, the metadata of the original data is stored on each shard of the redundancy group. For the non-full-strip original data written to the redundancy group, if the conventional zero-padding operation is adopted, the metadata of the original data is stored on each shard of the redundancy group. If the Figure 3 shown replica operation is adopted, the metadata of the original data is stored only on the parity shards of the redundancy group and the data shards corresponding to the original data blocks.

[0071] In the conventional zero-padding scheme, there is no requirement for the query object during conditional reading and writing. The key of the original data can be queried on any shard in the redundancy group. Although the query efficiency of conditional reading and writing is high, it causes serious waste of storage capacity space.

[0072] In the replica scheme, there are certain requirements for the query object during conditional reading and writing. It is necessary to first query on the parity shards in the redundancy group. If all parity shards are unavailable, since it is not clear which data shard or shards the original data block is written to, only when the key of the original data is queried on a certain data shard or the key of the original data is not queried on all data shards, can a conclusion be drawn. Although it avoids serious waste of storage capacity space, it reduces the query efficiency of conditional reading and writing.

[0073] In this embodiment, referring toFigure 2 and Figure 3 Regardless of whether the original data fills the stripe, the data existence flag corresponding to the key of the original data is written to each shard of the redundant group. In this way, there is no requirement for the query object during conditional read and write. It is possible to query any shard in the redundant group to check whether there is a corresponding data existence flag. If a corresponding data existence flag is found, it is determined that the original data corresponding to the key exists in the redundant group. If no corresponding data existence flag is found, it is determined that the original data corresponding to the key does not exist in the redundant group. If the originally scheduled shard for querying is unavailable, it can be replaced with other available shards.

[0074] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group for any original data written to the redundant group includes:

[0075] When writing any original data to the redundant group, query whether there is a corresponding data existence flag on any shard in the redundant group;

[0076] If it is found that there is no corresponding data existence flag, write the corresponding data existence flag to each shard in the redundant group.

[0077] In this embodiment, considering that the data existence flag writing logic has additional overhead, the data existence flag is written only when the original data corresponding to each key is written for the first time. If it is not the first write, for example, new data of the same key overwrites old data, since the corresponding data existence flags on all shards in the redundant group already exist and do not need to be changed, there is no need to perform additional repeated writing of the data existence flag.

[0078] Specifically, if it is found that there is a corresponding data existence flag on any shard in the redundant group, it means that the original data corresponding to the current key has been written to the redundant group, and the corresponding data existence flags exist on all shards in the redundant group, and no data existence flag is written. If it is found that there is no corresponding data existence flag on any shard in the redundant group, it means that the original data corresponding to the current key is written to the redundant group for the first time, and the corresponding data existence flags do not exist on all shards in the redundant group, and it is necessary to write the corresponding data existence flag to each shard in the redundant group.

[0079] Through this embodiment, unnecessary system overhead caused by repeated writing of data existence flags can be avoided.

[0080] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group for any original data written to the redundant group includes:

[0081] When writing any original data to a redundant group, query whether there is a corresponding data existence flag on any target shard in the redundant group, where the target shards include the parity shards and the data shards corresponding to the original data blocks;

[0082] If it is queried that there is no corresponding data existence flag, write the corresponding data existence flag to each shard in the redundant group.

[0083] The difference between this embodiment and the previous embodiment is that the query object is limited within the range of the parity shards and the data shards corresponding to the original data blocks, that is, the shards that need to write data blocks. Writing data blocks to these shards will necessarily involve data interaction, which is convenient for query operations. However, query operations on other data blocks require additional establishment of data interaction relationships, increasing unnecessary system overhead. Therefore, through this embodiment, unnecessary system overhead can be further avoided.

[0084] Specifically Figure 3 In the embodiment shown, the target shards include parity shards P, Q and data shards 0, 1, 2.

[0085] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group for any original data written to the redundant group includes:

[0086] When writing any original data to the redundant group, write the corresponding data existence flag to each shard in the redundant group.

[0087] In this embodiment, the data existence flag is written once every time the original data is written. The operation logic is simple, but the repeated writing of the data existence flag will cause unnecessary system overhead.

[0088] Further, in one embodiment, the step of writing the corresponding data existence flag to each shard in the redundant group includes:

[0089] When writing an original data block or a parity data block to each target shard in the redundant group, write the corresponding data existence flag to the corresponding target shard, where the target shards include the parity shards and the data shards corresponding to the original data blocks;

[0090] When the number of target shards is less than the number of shards in the redundant group, specify at most the number of target shards of shards from the first shards to be written as the forwarding targets for this round, and specify a target shard as the forwarding source for this round for each forwarding target for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the first shards to be written include the other shards in the redundant group except the target shards;

[0091] When twice the number of forwarding sources in the previous round is less than the number of redundant component slices, select at most twice the number of forwarding sources in the previous round of slices from the second slices to be written as the forwarding targets in this round. Designate an upper-round forwarding source or an upper-round forwarding target as the forwarding source for each forwarding target in this round, so that each forwarding source in this round forwards the corresponding data existence flag to a corresponding forwarding target in this round for the forwarding target in this round to write the corresponding data existence flag, where the second slices to be written include other slices in the redundant group except the upper-round forwarding sources and upper-round forwarding targets.

[0092] In this embodiment, for the target slices that are to write data blocks themselves, the data existence flag is written relying on the writing process of the data block, similar to the writing of metadata in the prior art. For the writing of non-full stripes, a diffusion-based forwarding strategy is adopted to write the data existence flag to other slices. The forwarding sources in the first round can be selected from the target slices, and the forwarding sources in each subsequent round can be selected from the forwarding sources and forwarding targets in the previous round. In this way, the forwarding efficiency will increase exponentially, there is no single-point performance bottleneck, and the advantage is more obvious in redundant groups with more data slices.

[0093] Figure 4 The flowchart of the forwarding of the data existence flag in an embodiment of the present application is shown.

[0094] Refer to Figure 4 , the redundancy level of the erasure code is 8 + 2, the chunk size is equal to 4K, correspondingly, the redundant group includes 8 data slices and 2 parity slices, and the stripe size is 32K.

[0095] Figure 4 The 8K-sized original data in is cut into two 4K-sized original data blocks A and B. Without calculating the parity data blocks, the original data blocks A and B are written into data slices 0 and 1 one by one. Taking the original data blocks A and B as a whole, they are written into parity slices P and Q in the form of replicas. Relying on the writing process of the data block, the corresponding data existence flags are written into data slices 0 and 1 and parity slices P and Q. The number of target slices is 4, the number of redundant component slices is 10, and the ceiling of log2(10 / 4) is 2. It can be seen that at least two rounds of forwarding are required.

[0096] Exemplarily, data slices 2, 3, 6, and 7 are used as the forwarding targets in the first round, and data slices 0, 1, and parity slices P and Q are respectively designated as the corresponding forwarding sources, so that the data existence flag is written into data slices 2, 3, 6, and 7. Data slices 4 and 5 are used as the forwarding targets in the second round, and data slices 0 and 1 are respectively designated as the corresponding forwarding sources, so that the data existence flag is written into data slices 4 and 5.

[0097] In a second aspect, an erasure code data storage device is further provided in an embodiment of the present application.

[0098] Figure 5 Shows a schematic diagram of the functional modules of an erasure code data storage device in an embodiment of the present application.

[0099] Referring to Figure 5 , in one embodiment, the erasure code data storage device includes:

[0100] The first data writing module 11 is configured to, when writing full-strip original data into a redundant group, write each original data block into a corresponding data shard in the redundant group, and write each parity data block into a corresponding parity shard in the redundant group;

[0101] The second data writing module 12 is configured to, when writing non-full-strip original data into a redundant group, write each original data block into a corresponding data shard in the redundant group, and write all the original data blocks into each parity shard in the redundant group;

[0102] The flag writing module 13 is configured to, for any original data written into the redundant group, write a corresponding data presence flag into each shard in the redundant group, where the data presence flag corresponds to the key of the original data.

[0103] Further, in one embodiment, the flag writing module 13 is configured to:

[0104] When writing any original data into the redundant group, query whether there is a corresponding data presence flag on any target shard in the redundant group, where the target shard includes the parity shard and the data shard corresponding to the original data block;

[0105] If it is queried that there is no corresponding data presence flag, write the corresponding data presence flag into each shard in the redundant group.

[0106] Further, in one embodiment, the flag writing module 13 is configured to:

[0107] When writing any original data into the redundant group, query whether there is a corresponding data presence flag on any shard in the redundant group;

[0108] If it is queried that there is no corresponding data presence flag, write the corresponding data presence flag into each shard in the redundant group.

[0109] Further, in one embodiment, the flag writing module 13 is configured to:

[0110] When writing any original data into the redundant group, write the corresponding data presence flag into each shard in the redundant group.

[0111] Further, in one embodiment, the flag writing module 13 is configured to:

[0112] When writing the original data block or the parity data block to each target shard in the redundancy group, write the corresponding data existence flag to the corresponding target shard, where the target shards include the parity shards and the data shards corresponding to the original data blocks;

[0113] When the number of target shards is less than the number of shards in the redundancy group, specify at most the number of target shards of shards from the first shards to be written as the forwarding targets for this round, and specify a target shard as the forwarding source for this round for each forwarding target for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the first shards to be written include the other shards in the redundancy group except the target shards;

[0114] When twice the number of forwarding sources in the previous round is less than the number of shards in the redundancy group, select at most twice the number of shards of the number of forwarding sources in the previous round from the second shards to be written as the forwarding targets for this round, and specify a forwarding source in the previous round or a forwarding target in the previous round as the forwarding source for this round for each forwarding target for this round, so that each forwarding source for this round forwards the corresponding data existence flag to the corresponding forwarding target for this round for the forwarding target for this round to write the corresponding data existence flag, where the second shards to be written include the other shards in the redundancy group except the forwarding sources in the previous round and the forwarding targets in the previous round.

[0115] Wherein, the function implementation of each module in the above erasure code data storage device corresponds to each step in the above erasure code data storage method embodiment, and its function and implementation process will not be elaborated here one by one.

[0116] In a third aspect, the embodiments of the present application further provide a method for operating clone volume data, and the original volume and the clone volume store data using the above erasure code data storage method.

[0117] Figure 6 The flowchart of the method for operating clone volume data in an embodiment of the present application is shown.

[0118] Refer to Figure 6 , in an embodiment, the method for operating clone volume data includes the following steps:

[0119] S21. When a read / write instruction for the clone volume is received during the copy operation from the original volume to the clone volume, query whether there is a corresponding data existence flag on any shard in the redundancy group corresponding to the clone volume.

[0120] Specifically, the original volume and the clone volume are concepts at the user level, and the redundancy group is a concept at the storage system level. The original volume and the clone volume will correspond to multiple redundancy groups, and for the read / write instructions of the user, one of the redundancy groups will be located.

[0121] Allowing users to read and write to the clone volume before the copy operation from the original volume to the clone volume is completed can reduce the waiting time perceived by users and improve the user experience. Correspondingly, conditional read and write needs to be adopted to ensure the correctness of read and write operations. First, it is necessary to query whether there is the original data (referred to as the target data for short) corresponding to the key involved in the user's read and write instruction in the clone volume. Since the clone volume uses the above-mentioned erasure code data storage method for data storage, it is only necessary to query whether there is a corresponding data existence flag on any one shard in the redundant group corresponding to the clone volume to determine whether there is target data in the clone volume, and the query efficiency is high.

[0122] S22. If there is no corresponding data existence flag on the shard queried in the clone volume, then query whether there is a corresponding data existence flag on any one shard in the redundant group corresponding to the original volume.

[0123] Specifically, the non-existence of target data in the clone volume is divided into two cases. The first case is that the target data does not exist in the original volume either, and the second case is that the target data exists in the original volume but has not been copied to the clone volume yet. Therefore, it is necessary to further query whether there is target data in the original volume. Since the original volume uses the above-mentioned erasure code data storage method for data storage, it is only necessary to query whether there is a corresponding data existence flag on any one shard in the redundant group corresponding to the original volume to determine whether there is target data in the original volume, and the query efficiency is high.

[0124] S23. If there is a corresponding data existence flag on the shard queried in the clone volume, or there is no corresponding data existence flag on the shards queried in both the clone volume and the original volume, then perform subsequent operations according to the read and write instructions.

[0125] Specifically, when there is target data in the clone volume, for a read instruction, return the target data in the clone volume to the user. For a write instruction, based on the target data in the clone volume, write the original data newly input by the user into the clone volume, and the new data will overwrite the corresponding old data. When there is no target data in both the clone volume and the original volume, for a read instruction, return a prompt indicating that the target data does not exist to the user. For a write instruction, write the original data newly input by the user into the clone volume.

[0126] S24. If there is no corresponding data existence flag on the shard queried in the clone volume and there is a corresponding data existence flag on the shard queried in the original volume, then copy the corresponding original data from the original volume to the clone volume and perform subsequent operations according to the read and write instructions.

[0127] Specifically, when there is no target data in the clone volume and there is target data in the original volume, first copy the target data from the original volume to the clone volume. Then, for a read instruction, return the target data in the clone volume to the user. For a write instruction, based on the target data in the clone volume, write the original data newly input by the user into the clone volume, and the new data will overwrite the corresponding old data.

[0128] Through this embodiment, during the copy operation from the original volume to the cloned volume, the correctness and execution efficiency of read and write operations can be ensured.

[0129] It should be noted that this embodiment describes an application scenario that requires conditional read and write, and it does not mean that the above erasure code data storage method can only be applied in this scenario.

[0130] Figure 7 The schematic diagram of the functional modules of the cloned volume data operation device in an embodiment of the present application is shown.

[0131] Referring to Figure 7 , in an embodiment, the cloned volume data operation device includes:

[0132] A first query module 21, configured to query whether there is a corresponding data existence flag on any one of the shards in the redundant group corresponding to the cloned volume when a read / write instruction for the cloned volume is received during the copy operation from the original volume to the cloned volume;

[0133] A second query module 22, configured to query whether there is a corresponding data existence flag on any one of the shards in the redundant group corresponding to the original volume if there is no corresponding data existence flag on the shard queried in the cloned volume;

[0134] A first execution module 23, configured to perform subsequent operations according to the read / write instruction if there is a corresponding data existence flag on the shard queried in the cloned volume, or if there is no corresponding data existence flag on the shards queried in both the cloned volume and the original volume;

[0135] A second execution module 24, configured to copy the corresponding original data from the original volume to the cloned volume and perform subsequent operations according to the read / write instruction if there is no corresponding data existence flag on the shard queried in the cloned volume and there is a corresponding data existence flag on the shard queried in the original volume.

[0136] Among them, the function implementation of each module in the above cloned volume data operation device corresponds to each step in the above embodiment of the cloned volume data operation method, and its function and implementation process will not be elaborated here one by one.

[0137] It should be noted that the serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0138] In the description of the specification, claims and the above-mentioned drawings of this application, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. The descriptions such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are different types.

[0139] In the description of the embodiments of this application, "exemplary", "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary", "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary", "for example" or "for instance" is intended to present related concepts in a specific manner.

[0140] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0141] In some processes described in the embodiments of this application, there are a plurality of operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to enable a terminal device to execute the methods described in the various embodiments of this application.

[0143] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.

Claims

1. A method for storing erasure coded data, characterized in that: The erasure code data storage method comprises: When writing full stripe original data to a redundant group, each original data block is written to a corresponding data shard in the redundant group, and each check data block is written to a corresponding check shard in the redundant group; When writing the original data of a non-full stripe to a redundant group, each original data block is written to a corresponding data shard in the redundant group, and all original data blocks are written to each check shard in the redundant group; For any original data written into the redundancy group, a corresponding data existence flag is written into each shard in the redundancy group, wherein the data existence flag corresponds to the key of the original data.

2. The erasure code data storage method according to claim 1, characterized in that: The step of writing a corresponding data existence mark to each slice in the redundant group for any original data written into the redundant group comprises: When writing any original data to the redundancy group, check whether there is a corresponding data existence mark on any target shard in the redundancy group, where the target shard includes the check shard and the data shard corresponding to the original data block; If no corresponding data existence mark is found in the query, the corresponding data existence mark is written into each shard in the redundant group.

3. The erasure code data storage method according to claim 1, characterized in that: The step of writing a corresponding data existence mark to each slice in the redundant group for any original data written into the redundant group comprises: When writing any original data to the redundancy group, check whether there is a corresponding data existence flag on any shard in the redundancy group; If no corresponding data existence mark is found in the query, the corresponding data existence mark is written into each shard in the redundant group.

4. The erasure code data storage method according to claim 1, characterized in that: The step of writing a corresponding data existence mark to each slice in the redundant group for any original data written into the redundant group comprises: When any original data is written to the redundancy group, the corresponding data existence flag is written to each shard in the redundancy group.

5. The erasure coded data storage method according to any one of claims 2 to 4, characterized in that: The step of writing the corresponding data existence flag into each slice in the redundancy group comprises: When writing the original data block or the check data block to each target shard in the redundancy group, the corresponding data existence flag is written to the corresponding target shard, wherein the target shard includes the check shard and the data shard corresponding to the original data block; When the number of target shards is less than the number of shards in the redundant group, specify shards of at most the number of target shards from the first shards to be written as forwarding targets for this round, specify a target shard for each forwarding target for this round as a forwarding source for this round, and make each forwarding source for this round forward a corresponding data existence mark to a corresponding forwarding target for this round, so that the forwarding target for this round writes the corresponding data existence mark, wherein the first shards to be written include other shards in the redundant group except the target shard; When twice the number of forwarding sources in the previous round is less than the number of shards in the redundant group, select shards at most twice the number of forwarding sources in the previous round from the second shards to be written as forwarding targets for this round, and designate a previous round forwarding source or a previous round forwarding target as the forwarding source for this round for each forwarding target in this round, so that each forwarding source in this round forwards a corresponding data existence flag to a corresponding forwarding target in this round, so that the forwarding target in this round can write the corresponding data existence flag, wherein the second shards to be written include other shards in the redundant group except the previous round forwarding sources and the previous round forwarding targets.

6. An erasure coded data storage device, characterized in that: The erasure code data storage device comprises: A first data writing module is used to write each original data block into a corresponding data slice in the redundant group and write each check data block into a corresponding check slice in the redundant group when writing the original data of the full stripe to the redundant group; A second data writing module is used to write each original data block into a corresponding data slice in the redundant group when writing the original data of a non-full stripe to the redundant group, and write all the original data blocks into each check slice in the redundant group; The flag writing module is used to write a corresponding data existence flag to each fragment in the redundant group for any original data written into the redundant group, wherein the data existence flag corresponds to the key of the original data.

7. The erasure coded data storage device according to claim 6, characterized in that: The flag writing module is used to: When writing any original data to the redundancy group, check whether there is a corresponding data existence mark on any target shard in the redundancy group, where the target shard includes the check shard and the data shard corresponding to the original data block; If no corresponding data existence mark is found in the query, the corresponding data existence mark is written into each shard in the redundant group.

8. The erasure coded data storage device according to claim 7, characterized in that: The flag writing module is used to: When writing the original data block or the check data block to each target shard in the redundancy group, the corresponding data existence flag is written to the corresponding target shard; When the number of target shards is less than the number of shards in the redundant group, specify shards of at most the number of target shards from the first shards to be written as forwarding targets for this round, specify a target shard for each forwarding target for this round as a forwarding source for this round, and make each forwarding source for this round forward a corresponding data existence mark to a corresponding forwarding target for this round, so that the forwarding target for this round writes the corresponding data existence mark, wherein the first shards to be written include other shards in the redundant group except the target shard; When twice the number of forwarding sources in the previous round is less than the number of shards in the redundant group, select shards at most twice the number of forwarding sources in the previous round from the second shards to be written as forwarding targets for this round, and designate a previous round forwarding source or a previous round forwarding target as the forwarding source for this round for each forwarding target in this round, so that each forwarding source in this round forwards a corresponding data existence flag to a corresponding forwarding target in this round, so that the forwarding target in this round can write the corresponding data existence flag, wherein the second shards to be written include other shards in the redundant group except the previous round forwarding sources and the previous round forwarding targets.

9. A method for operating clone volume data, characterized in that: The original volume and the clone volume use the erasure code data storage method according to any one of claims 1 to 5 for data storage, and the clone volume data operation method includes: When receiving a read / write instruction from a user for a clone volume during the copy operation from the original volume to the clone volume, query whether there is a corresponding data existence mark on any slice in the redundancy group corresponding to the clone volume; If there is no corresponding data existence mark on the queried slice in the clone volume, check whether there is a corresponding data existence mark on any slice in the redundancy group corresponding to the original volume; If the corresponding data exists mark on the queried slice in the clone volume, or if the corresponding data exists mark on both the queried slice in the clone volume and the original volume, the subsequent operation is performed according to the read and write instructions; If there is no corresponding data existence mark on the queried slice in the clone volume, and there is a corresponding data existence mark on the queried slice in the original volume, the corresponding original data is copied from the original volume to the clone volume, and subsequent operations are performed according to the read and write instructions.

10. A clone volume data operation device, characterized in that: The original volume and the clone volume use the erasure code data storage method according to any one of claims 1 to 5 for data storage, and the clone volume data operation device includes: A first query module is used to query whether there is a corresponding data existence mark on any slice in the redundancy group corresponding to the clone volume when receiving a user's read and write instruction for the clone volume during the copy operation from the original volume to the clone volume; The second query module is used to query whether there is a corresponding data existence mark on any slice in the redundancy group corresponding to the original volume if there is no corresponding data existence mark on the queried slice in the clone volume; A first execution module, configured to execute subsequent operations according to the read / write instruction if there is a corresponding data existence mark on the queried slice in the clone volume, or if there is no corresponding data existence mark on the queried slice in both the clone volume and the original volume; The second execution module is used to copy the corresponding original data from the original volume to the clone volume if there is no corresponding data existence mark on the queried slice in the clone volume, and perform subsequent operations according to the read and write instructions.