Distributed storage method, node, system, computer program product and storage medium

By piggybacking the rewriting of fragments of failing nodes in a distributed storage system, the performance degradation caused by full copying is solved, thus improving data reliability while ensuring read and write performance.

CN120928997APending Publication Date: 2025-11-11ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410570280.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

When a storage node failure is detected, the existing distributed storage system uses full copy erasure coding fragmentation, which leads to a decrease in read and write performance. Furthermore, false alarms may cause invalid copies, affecting system performance and efficiency.

Method used

When a missing erasure coding shard group is detected, a shard rewrite task is created to rewrite the shards on the faulty node without performing a full copy. This increases the write throughput slightly by utilizing the existing shard rewrite task, ensuring data reliability and performance.

Benefits of technology

It effectively reduces read and write traffic caused by failing nodes, ensures the read and write performance of the distributed storage system, saves data copy traffic after node failure, and improves data reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928997A_ABST
    Figure CN120928997A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a distributed storage method, a node, a system, a computer program product and a storage medium. In the embodiment of the invention, under the condition that fragment missing occurs in the erasure code fragment group corresponding to the target data, the fragment rewriting task can be created for the target data, however, when the fragment rewriting task is created, the missing fragment does not only serve as a rewriting object any more, and the rewriting efficiency is improved. However, fragments, located on the temporary bad nodes in the distributed storage system, in the erasure code fragment group are also used as rewriting objects. Therefore, after the fragment rewriting task is executed, not only can the missing fragments be rewritten, but also the fragments located on the damaged nodes can be rewritten incidentally. Based on the incidentally rewriting mechanism, the data reliability in the distributed storage system can be improved under the condition that the read-write performance of the distributed storage system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular to a distributed storage method, node, system, computer program product, and storage medium. Background Technology

[0002] In distributed storage systems, erasure coding (EC) mechanisms can be used for data protection. According to the erasure coding mechanism, a segment of data can be divided into k data blocks, and m parity blocks are generated based on these k data blocks to produce k+m erasure-coded fragments corresponding to that data segment. In a distributed storage system, these k+m erasure-coded fragments can be distributed and stored across different storage nodes. Based on this, the data segment can be recovered using any k fragments from the k+m erasure-coded fragments. Therefore, in the erasure coding mechanism, the erasure-coded fragments constructed for the data are a crucial foundation for achieving data protection.

[0003] Therefore, currently, if it is predicted that a storage node in a distributed storage system may be about to fail, the erasure coding fragments stored in this failing node are fully copied to other storage nodes to ensure the reliability of these fragments. However, this full copy method results in a large amount of read and write traffic, and there is a possibility that the failing node may be falsely reported. This is equivalent to performing a large amount of useless copying work. Therefore, although the original intention of this full copy method is to ensure the data reliability of the distributed storage system, it has a negative impact on the read and write performance of the distributed storage system. Summary of the Invention

[0004] This application provides a distributed storage method, node, system, computer program product, and storage medium to improve data reliability in a distributed storage system while ensuring read and write performance.

[0005] This application provides a distributed storage method applicable to metadata nodes in a distributed storage system. The method includes:

[0006] If a missing fragment is detected in the erasure coding fragment group corresponding to the target data, it is determined whether there is a first-type fragment located on a faulty node in the distributed storage system in the erasure coding fragment group.

[0007] If they exist, the first type of fragment and the missing second type of fragment in the erasure coding fragment group are identified as the target fragments to be rewritten;

[0008] A shard rewrite task is created for the target data, the shard rewrite task being used to instruct the target shard to be rewritten for the target data in the distributed storage system.

[0009] This application also provides a distributed storage method applicable to a target node in a distributed storage system, wherein the target node is a non-critical node, and the method includes:

[0010] In response to receiving a fragment rewrite task for target data, the target fragment to be rewritten corresponding to the target data is determined, wherein the target fragment includes a first type of fragment and a second type of fragment, the first type of fragment is the fragment located on the dying node in the erasure coding fragment group corresponding to the target data, and the second type of fragment is the missing fragment in the erasure coding fragment group;

[0011] Read k available fragments from the erasure coding fragment group from the distributed storage system;

[0012] Based on the k available shards, the shard rewrite task is executed to rewrite the target shard for the target data in the distributed storage system.

[0013] This application embodiment also provides a distributed storage system, including a metadata node and multiple storage nodes. The metadata storage node is used to execute the aforementioned distributed storage method applicable to the metadata node to create a sharding rewrite task. The sharding rewrite task is used to trigger a target node in the distributed storage system to execute the aforementioned distributed storage method applicable to the target node, wherein the metadata node or any storage node serves as the target node.

[0014] This application embodiment also provides a metadata node, including a memory, a processor, and a communication component;

[0015] The memory is used to store one or more computer instructions;

[0016] The processor is coupled to the memory and the communication component to execute one or more computer instructions for performing the aforementioned distributed storage method applicable to metadata nodes.

[0017] This application embodiment also provides a node, including a memory, a processor, and a communication component;

[0018] The memory is used to store one or more computer instructions;

[0019] The processor is coupled to the memory and the communication component to execute one or more computer instructions for performing the aforementioned distributed storage method applicable to the target node.

[0020] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by one or more processors, causes the one or more processors to perform the aforementioned distributed storage method.

[0021] This application also provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to execute the aforementioned distributed storage method.

[0022] In this embodiment of the application, it is proposed that when a fragment is missing within the erasure coding shard group corresponding to the target data, a shard rewrite task can be created for the target data. However, when creating the shard rewrite task, instead of only considering the missing fragment as the rewrite object, it is proposed to also consider the fragments located on the faulty nodes in the distributed storage system within the erasure coding shard group as the rewrite object. This ensures that after such a shard rewrite task is executed, not only can the missing fragment be rewritten, but also the fragments located on the faulty nodes can be rewritten incidentally. Based on this piggybacking rewrite mechanism, only a very small amount of write traffic is added to the already required fragment rewrite task. Thus, even if a failing node is falsely reported, the read / write traffic caused by the failing node is effectively reduced because a full copy is no longer performed, ensuring the read / write performance of the distributed storage system. Furthermore, if a failing node is not falsely reported, the mechanism allows for the pre-writing of some data on the failing node, effectively saving the read / write traffic required for data copying after a true failure. This not only ensures data reliability in the distributed storage system but also guarantees its read / write performance. Therefore, in this embodiment, data reliability in the distributed storage system can be improved while ensuring its read / write performance. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0024] Figure 1 A schematic diagram of the structure of a distributed storage system provided in an exemplary embodiment of this application;

[0025] Figure 2 A schematic diagram of an erasure code fragment group corresponding to target data provided in an exemplary embodiment of this application;

[0026] Figure 3 A schematic diagram illustrating the execution logic of a fragmented rewriting task, provided as an embodiment of this application;

[0027] Figure 4A schematic diagram illustrating another fragmented rewrite task execution logic provided in an embodiment of this application;

[0028] Figure 5 A flowchart illustrating a distributed storage method provided as another exemplary embodiment of this application;

[0029] Figure 6 A flowchart illustrating another distributed storage method provided as another exemplary embodiment of this application;

[0030] Figure 7 This is a schematic diagram of the structure of a node in a distributed storage system, which is another exemplary embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] Before providing a detailed description of the technical solutions provided in the various embodiments of this application, the following explanations are given for several technical concepts involved in this application.

[0033] A distributed storage system is a storage system that distributes data across multiple storage nodes. A distributed storage system typically includes a metadata node and multiple storage nodes.

[0034] Metadata nodes are nodes in a distributed storage system used to store metadata information and can perform management and control functions. Metadata information may include, but is not limited to, data storage status, data storage location, and various attribute information used to describe the data.

[0035] A storage node is a node in a distributed storage system used to store data and can perform tasks such as data reading, writing, and deletion.

[0036] Erasure coding (EC) is a data protection mechanism. It divides data into blocks and creates redundant check blocks, generating a set of erasure-coded fragments, often called an erasure-coded fragment group (EC group). These fragments are then stored on different storage nodes. This way, even if some erasure-coded fragments are missing, the original data can be recovered by combining the remaining fragments.

[0037] As described in the background section, erasure coding shards, constructed for data, are a crucial foundation for data protection in distributed storage systems. Reliable erasure coding shards better support the recovery of original data, thereby ensuring data reliability in the distributed storage system. Therefore, currently, when a failing node appears in a distributed storage system, as described in the background section, the erasure coding shards stored on the failing node are fully copied to other storage nodes to ensure the security of these shards. While the initial intention of this full copy method is to ensure data reliability in the distributed storage system, it can negatively impact the read and write performance of the distributed storage system. This is because a full copy consumes a significant amount of read and write bandwidth in the distributed storage system, placing considerable read and write pressure on it.

[0038] Furthermore, the inventors discovered during their research that potentially failing nodes in a distributed storage system can be falsely identified. This is because potentially failing nodes are typically determined through fault prediction, and therefore, the fault confidence level corresponding to a potentially failing node is not necessarily 100%. During fault prediction, a storage node may be identified as a potentially failing node when its fault confidence level reaches a certain threshold (e.g., 50%). Clearly, the lower the fault confidence level, the lower the probability that a potentially failing node will actually fail. This is why such storage nodes are described as potentially failing nodes; in other words, a potentially failing node refers to a storage node that has not yet failed but is predicted to be likely to fail soon.

[0039] Thus, according to the full copy method mentioned in the background technology, if a faulty node is falsely reported, the full copy of the faulty node becomes a useless task. Therefore, although the original intention of this full copy method is to ensure the data reliability of the distributed storage system, it will have a negative impact on the read and write performance of the distributed storage system.

[0040] Therefore, this embodiment proposes a distributed storage method to improve data reliability in a distributed storage system while ensuring read and write performance.

[0041] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0042] Figure 1 This is a schematic diagram of the structure of a distributed storage system provided for an exemplary embodiment of this application. For example... Figure 1 As shown, the system includes a metadata node and multiple storage nodes. The metadata node communicates with each of the multiple storage nodes, and the multiple storage nodes also communicate with each other.

[0043] This embodiment proposes adding the ability to detect nodes on the verge of failure to the metadata node. This allows the metadata node to promptly identify which storage nodes in the distributed storage system have been predicted to be on the verge of failure.

[0044] This embodiment can employ various implementation methods to support the metadata node's ability to detect nodes on the verge of failure. In one optional implementation: the metadata node can receive fault prediction information provided by a fault prediction tool; based on the fault prediction information, it marks nodes on the verge of failure in the distributed storage system. Here, the fault prediction tool refers to a tool component with device fault prediction capabilities, which can be implemented as software, hardware, or a combination of software and hardware. In this optional implementation, the metadata node can establish a communication connection with the fault prediction tool, which can perform fault prediction on storage nodes in the distributed storage system. The fault prediction information provided by the fault prediction tool may include, but is not limited to, information such as the identifier of nodes on the verge of failure. Thus, the metadata node can parse the fault prediction information to detect nodes on the verge of failure in the distributed storage system and mark them.

[0045] It should be understood that, in addition to the optional implementation methods described above, other implementation methods can also be used in this embodiment to support the ability of the metadata node to detect nodes on the verge of failure. For example, a fault prediction module can be added to the metadata node, and the fault prediction module can perform fault prediction on the storage nodes in the distributed storage system. In this way, the metadata node itself can have fault prediction capabilities, thereby being able to detect nodes on the verge of failure in the distributed storage system in a timely manner and mark them.

[0046] Furthermore, this embodiment does not limit the fault prediction technology used in the fault prediction process; it can be any existing or future available fault prediction technology. For example, fault prediction can be based on the Regularized Greedy Forest (RGF) algorithm combined with a transfer learning algorithm. No further examples are provided here.

[0047] Building upon this, this embodiment no longer performs a full copy of the failing node, but instead proposes a piggyback rewrite mechanism for failing nodes. In this piggyback rewrite mechanism, the fragment rewrite task that is already required in the distributed storage system can be used to piggyback the rewrite of the fragments stored by the failing node.

[0048] During their research, the inventors discovered that under the erasure coding mechanism, not only can the original data be recovered based on other erasure coding fragments when some erasure coding fragments are missing, as mentioned earlier, but it also supports reconstructing the missing fragments based on these other erasure coding fragments and writing the reconstructed fragments into the distributed storage system to ensure the completeness of the erasure coding fragments, thereby better supporting the recovery of the original data. In this embodiment, the process of reconstructing the missing erasure coding fragments and rewriting them into the distributed storage system is described as fragment rewriting.

[0049] During their research, the inventors also discovered that in a distributed storage system, it is possible to monitor whether a fragment is missing within the erasure coding fragment group corresponding to the target data. If so, a fragment rewriting task can be created for the target data. In this embodiment, this triggering method for fragment rewriting tasks is used, but the creation process of fragment rewriting tasks has been modified.

[0050] It should be understood that the target data in this embodiment can be any data stored in a distributed storage system using an erasure coding mechanism. The metadata information recorded in the metadata node for this type of data may include, but is not limited to, the identifiers of the corresponding erasure coding fragments, missing markers used to indicate that erasure coding fragments are missing, and the storage nodes where each erasure coding fragment is located, etc., without further examples here.

[0051] Based on this, this embodiment proposes that the metadata node can also determine whether there is a first-type fragment located on a failing node within the erasure coding fragment group corresponding to the target data after detecting fragment loss of the target data. As mentioned earlier, the metadata node records metadata information for the target data and has already marked the failing nodes in the distributed storage system. Based on this, in this embodiment, the metadata node can search for the storage nodes where each fragment in its erasure coding fragment group is located based on the metadata information corresponding to the target data. If a failing node exists among the searched storage nodes, it is determined that there is a first-type fragment located on a failing node within the erasure coding fragment group. Therefore, in this embodiment, the metadata node can easily determine whether there is a first-type fragment located on a failing node within the erasure coding fragment group corresponding to the target data. In this embodiment, a group of erasure coding fragments generated for a piece of data in the erasure coding mechanism is described as an erasure coding fragment group (EC group). It is understood that the erasure coding fragment group corresponding to the target data includes all the additional erasure coding fragments generated for the target data, including the aforementioned data blocks and check blocks. Thus, the first type of fragment in this embodiment may be a data block or a verification block.

[0052] Figure 2 This is a schematic diagram of an erasure coding fragment group corresponding to target data, provided as an exemplary embodiment of this application. (Reference) Figure 2 The erasure coding fragment group corresponding to the target data contains 6 erasure coding fragments, which are distributed across different storage nodes. Erasure coding fragment 4 is missing, and erasure coding fragment 5 is stored on node 5, which has been marked as a potentially faulty node. Figure 2 Erasure coding fragment 4 in this embodiment is the second type of fragment. Figure 2 The erasure coding fragment 5 in this embodiment is the first type of fragment.

[0053] This embodiment proposes that if the metadata node detects a first-type fragment within the erasure coding fragment group corresponding to the target data, then in addition to treating the missing second-type fragment within that erasure coding fragment group as the rewriting target, the detected first-type fragment will also be treated as the rewriting target. That is, in this case, the metadata node can determine both the first-type and second-type fragments within the erasure coding fragment group as target fragments to be rewritten. Figure 2 The erasure coding fragment shown in the example, Figure 2 Erasure coding fragment 4 and erasure coding fragment 5 in this embodiment are both used as rewriting objects (i.e., target fragments).

[0054] Based on this, the metadata node can create a fragment rewrite task for the target data and carry information such as the identifier of the determined target fragment in the fragment rewrite task. In this way, the fragment rewrite task can instruct the rewrite of these target fragments for the target data in the distributed storage system. As a result, the first type of fragments corresponding to the target data located on the faulty node will be piggybacked and rewritten during the execution of the fragment rewrite task.

[0055] It is understood that this embodiment will no longer perform a full copy of the failing node, nor will it actively initiate a data copy task for the failing node. Instead, it will utilize the shard rewrite task that is already required to be executed in the distributed storage system to incidentally rewrite the relevant shards stored on the failing node. That is, some shards stored on the failing node will be rewritten in the shard rewrite task that is already required to be executed in the distributed storage system.

[0056] Furthermore, it is worth emphasizing that the rewriting mentioned in this embodiment is essentially a write operation. Distributed storage systems already have mature mechanisms to ensure the reasonable scheduling of write operations, guaranteeing that data is written to reliable locations. That is, in this embodiment, through the fragment rewriting task, it can be ensured that the target fragment can be written to a reliable storage node in the distributed storage system. Typically, full and / or failing storage nodes are unreliable storage nodes; therefore, the target fragment will not be rewritten to these unreliable storage nodes. In an optional implementation: the metadata node can allocate the destination node for the target fragment during rewriting from the remaining storage nodes after excluding failing nodes in the distributed storage system. Different target fragments can be allocated to different destination nodes; and the identifier of the destination node allocated to the target fragment can be carried in the fragment rewriting task to indicate that the target fragment should be rewritten to its allocated destination node. Of course, this is only optional, and this embodiment is not limited to this.

[0057] In this embodiment, after the metadata node creates a fragmented rewrite task for the target data, the execution subject of the fragmented rewrite task is not limited.

[0058] refer to Figure 1 Optionally, in this embodiment, a storage node in the distributed storage system can be used as the execution entity for the fragment rewrite task. In this embodiment, the metadata node can select a target node from the remaining storage nodes after excluding failing nodes in the distributed storage system; and send the fragment rewrite task created for the target data to the target node to trigger the target node to execute the fragment rewrite task. In a preferred embodiment: the metadata node can use any node from the destination nodes allocated for the target fragment as the target node; or, it can select any available fragment used in the rewriting process of the target fragment from the aforementioned remaining storage nodes as the target node. In this preferred embodiment, the storage node containing any erasure code fragment that needs to be read / written in the fragment rewrite task is selected as the target node. Since the erasure code fragment itself is located on the target node, the read / write operation of the erasure code fragment on the target node can be an intra-node operation, without causing inter-node traffic. That is, the process corresponding to the erasure code fragment can be eliminated during the execution of the fragment rewrite task, which can effectively save the traffic caused by the fragment rewrite task.

[0059] Of course, in this embodiment, the metadata node can also serve as the execution entity for the sharding rewrite task; that is, the metadata node can serve as the target node. A dedicated node can also be added to the distributed storage system to serve as the execution entity for the sharding rewrite task; this dedicated node can serve as the target node. Further examples of execution entities are not provided here.

[0060] Based on this, the target node can, in response to receiving a shard rewrite task for the target data, determine the target shard to be rewritten corresponding to the target data. As mentioned earlier, in this embodiment, the target shard may include not only the missing second-type shard corresponding to the target data, but also the first-type shard corresponding to the target data located on the failing node.

[0061] refer to Figure 1 In this embodiment, the target node can read k available fragments from the erasure coding fragment group in the distributed storage system; based on the k available fragments, it can execute a fragment rewrite task to rewrite the target fragment for the target data in the distributed storage system.

[0062] It is understandable that the target node's sharding rewrite task generally involves two stages: a read stage and a write stage.

[0063] During the read phase, the target node needs to read k available fragments corresponding to the target data. Here, k is the number of erasure coding fragments required for missing fragment reconstruction in the erasure coding mechanism. That is, based on the k erasure coding fragments corresponding to the target data, any other erasure coding fragment corresponding to the target data can be reconstructed. This embodiment does not limit the implementation logic of the reconstruction. One exemplary implementation logic could be: recovering the target data based on the read k erasure coding fragments, regenerating erasure coding fragments for the target data, and the generated erasure coding fragments containing the target fragments. Further examples of the reconstruction implementation logic are not provided here.

[0064] In this embodiment, available fragments can be understood as fragments that are not missing. During the read operation, the target node can send read commands to the storage nodes where the required k available fragments are located, in order to read the k available fragments.

[0065] Thus, in this embodiment, during the read phase, the fragmented rewrite task created for the target data will result in k read traffic.

[0066] In this embodiment, based on the k available fragments read, target fragments that need to be rewritten can be generated for the target data. As mentioned above, the target fragments may include first-type fragments and second-type fragments.

[0067] Based on this, during the write phase, the target node can rewrite the target fragment for the target data in the distributed storage system. The target node can send a write command to the destination node to which the target fragment is allocated, so as to write the target fragment to the allocated destination node.

[0068] Thus, in this embodiment, during the write phase, the fragment rewrite task created for the target data will also cause write traffic. Among them, the write traffic caused by rewriting the second type of fragment is the write traffic that the fragment rewrite task originally needed to occupy, while only the write traffic caused by rewriting the first type of fragment is the additional write traffic caused in this embodiment.

[0069] It is worth emphasizing that, compared to traditional fragmented rewrite tasks, although the fragmented rewrite task created in this embodiment causes additional write traffic for the first type of fragments, the inventors found during the research process that the number of first type fragments involved in a single fragmented rewrite task is usually very small. Therefore, the additional write traffic for the first type of fragments caused in this embodiment is also very small, and this type of write traffic accounts for a very low proportion of the read and write traffic caused by the fragmented rewrite task.

[0070] For example, if the erasure coding shard group of the target data contains k+m erasure coding shards, including one missing shard and one shard located on a faulty node, then the shard rewrite task will result in k read traffic, one write traffic for the missing shard, and one write traffic for the shard located on a faulty node. The k read traffic and the one write traffic corresponding to the missing shard are the traffic that would occur in a traditional shard rewrite task. Therefore, the shard rewrite task created in this embodiment only results in an additional one write traffic.

[0071] Therefore, the piggyback rewrite mechanism proposed in this embodiment will not negatively impact the read and write performance of the distributed storage system.

[0072] From a data reliability perspective, if the critical node of the first type of shard detected for the target data is a false alarm, then the data on that critical node will not affect the data reliability of the distributed storage system even if it is not rewritten. In this embodiment, a small number of shards on that critical node were rewritten, which also does not affect the data reliability of the distributed storage system. Moreover, as mentioned above, the additional traffic caused by the critical node in this embodiment is very small. Therefore, in this case, it will not have a negative impact on the read and write performance of the distributed storage system.

[0073] If the critical node containing the first type of shard detected by the target data is not falsely reported, since this embodiment has already used a shard rewrite task to piggyback some shards on the critical node, after the critical node actually fails, only the shards on the critical node that have not yet been rewritten need to be copied. This effectively reduces the read / write traffic caused by copying after the critical node actually fails. It can be seen that in this case, the piggyback rewrite stage does not negatively impact the read / write performance of the distributed storage system, nor does the copying stage after the critical node actually fails negatively impact the read / write performance of the distributed storage system. Moreover, the shards in the critical node can be rewritten through the aforementioned piggyback rewrite stage and the copying stage after the failure, thus ensuring the data reliability of the distributed storage system.

[0074] In summary, this embodiment proposes that when a fragment is detected to be missing within the erasure coding shard group corresponding to the target data, a shard rewrite task can be created for the target data. However, when creating the shard rewrite task, instead of only considering the missing fragment as the rewrite target, it proposes to also consider the fragments located on the failing nodes in the distributed storage system within the erasure coding shard group as the rewrite targets. This ensures that after such a shard rewrite task is executed, not only can the missing fragment be rewritten, but the fragments located on the failing nodes can also be rewritten incidentally. Based on this piggybacking rewrite mechanism, only a very small amount of write traffic is added to the already required fragment rewrite task. Thus, even if a failing node is falsely reported, the read / write traffic caused by the failing node is effectively reduced because a full copy is no longer performed, ensuring the read / write performance of the distributed storage system. Furthermore, if a failing node is not falsely reported, the mechanism allows for the pre-writing of some data on the failing node, effectively saving the read / write traffic required for data copying after a true failure. This not only ensures data reliability in the distributed storage system but also guarantees its read / write performance. Therefore, in this embodiment, data reliability in the distributed storage system can be improved while ensuring its read / write performance.

[0075] In the above or below embodiments, the metadata node can specify the k available shards to be used for the shard rewrite task.

[0076] As mentioned earlier, in this embodiment, available fragments can be understood as fragments that are not missing. Therefore, the first type of fragments in this embodiment also belong to available fragments. In this regard, this embodiment proposes that when selecting k available fragments, the first type of fragments should be avoided as much as possible. However, if necessary, the k available fragments may also include the first type of fragments.

[0077] Based on this, this embodiment proposes an exemplary scheme for selecting k available fragments: if the number of remaining fragments after removing the first and second types of fragments within the erasure coding fragment group is greater than or equal to k, then k available fragments can be randomly selected from the remaining fragments. That is, if there are sufficient remaining fragments after removing the first and second types of fragments, the first type of fragments will not be selected. This avoids affecting the read phase of the fragment rewriting task due to problems such as slow transmission rates, partial data loss, or inability to copy data from failing nodes.

[0078] In this exemplary scheme: if the number of remaining fragments is less than k, then the required number of fragments can be selected from the first type of fragments corresponding to the target data to make up to k available fragments. That is, the total number of remaining fragments and the selected first type of fragments is k.

[0079] After selecting k available shards, the metadata node can carry the identifiers of the k available shards in the shard rewrite task. In this way, the shard rewrite task can instruct the generation of target shards based on the k available shards for the target data.

[0080] It is worth noting that, besides the metadata node specifying the k available shards to be used for the sharding rewrite task, the target node can also autonomously select the k available shards. In this case, the metadata node can provide sufficient information to the target node to support its autonomous selection. This embodiment does not limit the entity performing the operation of selecting k available shards.

[0081] Thus, in this embodiment, the k available fragments selected for the fragment rewrite task may fall into two categories: one is that the k available fragments do not contain any first-type fragments, and the other is that the k available fragments contain one or more first-type fragments.

[0082] In response to the two possible scenarios involving k available shards, this embodiment proposes that the execution logic implemented by the target node can differ when performing a shard rewrite task based on k available shards. The following provides illustrative examples of the execution logic implemented by the target node for each of these two scenarios.

[0083] Figure 3 This is a schematic diagram illustrating the execution logic of a fragmented rewriting task, provided as an embodiment of this application. (Reference) Figure 3 In one scenario, if none of the first-type shards are included in the k available shards, the target node can reconstruct the first-type and second-type shards for the target data based on the k available shards; and can write the reconstructed first-type and second-type shards to their respective assigned destination nodes to complete the shard rewriting task.

[0084] refer to Figure 3 Four available fragments were selected for the fragment rewriting task corresponding to the target data: fragment 1, fragment 2, fragment 3, and fragment 6. None of these four available fragments are of the first type in this embodiment. (See reference...) Figure 3 The target node can reconstruct shard 4 (second type of shard) and shard 5 (first type of shard) based on the four available shards, and can write the reconstructed shard 4 and shard 5 into their respective assigned target nodes.

[0085] It can be seen that in this case, the target node will reconstruct the first type of shards and the second type of shards during the shard rewriting task.

[0086] Figure 4 This is a schematic diagram illustrating another fragmented rewrite task execution logic provided in an embodiment of this application. (See reference...) Figure 4In another case, if the k available shards include p first-class shards, the target node can, after reading the p first-class shards, write the p first-class shards to its assigned destination node, where p is an integer greater than 0; based on the k available shards, reconstruct the second-class shards and the remaining first-class shards for the target data; and write the reconstructed second-class shards and the remaining first-class shards to their respective assigned destination nodes.

[0087] In this scenario, for the first type of shards included in the k available shards, the target node essentially uses a copying method to rewrite these first-type shards. If the target data also stores the remaining first-type shards not included in the k available shards, then these first-type shards can be rewritten using the method of reconstruction followed by writing, as in the previous scenario. For the second type of shards, the rewriting method remains consistent with the previous scenario (i.e., reconstruction followed by writing).

[0088] refer to Figure 4 This also selects four available shards for the sharded rewrite task corresponding to the target data: shard 1, shard 2, shard 3, and shard 5. Shard 5 is the first type of shard. (See reference) Figure 4 The target node can reconstruct shard 4 (second type of shard) based on the four available shards, without having to reconstruct shard 5 (first type of shard). The reconstructed shard 4 and the read shard 5 can be written to their respective assigned target nodes.

[0089] It should be understood that the execution logic in the two cases described above is merely exemplary, and this embodiment is not limited thereto. For example, rewriting can also be achieved by reconstructing all target shards before writing. No further examples will be provided here.

[0090] In summary, in this embodiment, k available fragments can be selected for the fragment rewriting task corresponding to the target data, and the selection of first-type fragments can be avoided as much as possible to prevent affecting the read phase of the fragment rewriting task. Even if one or more first-type fragments are selected from the k available fragments, it will not affect the rewriting of the target fragment in this embodiment. Moreover, if first-type fragments are selected, the selected first-type fragments can be rewritten by copying during the fragment rewriting task execution, which can save some of the computational cost of fragment reconstruction, thereby offsetting the negative impact of the failing node on transmission rate and other aspects.

[0091] Figure 5 This is a flowchart illustrating a distributed storage method as another exemplary embodiment of this application, which can be executed by a metadata node in a distributed storage system. (See reference...) Figure 5 The method may include:

[0092] Step 500: If a fragment is missing in the erasure coding fragment group corresponding to the target data, determine whether there is a first type fragment located on a faulty node in the distributed storage system in the erasure coding fragment group.

[0093] Step 501: If it exists, then the first type of fragment and the missing second type of fragment in the erasure coding fragment group are identified as the target fragments to be rewritten;

[0094] Step 502: Create a shard rewrite task for the target data. The shard rewrite task is used to instruct the target shard to be rewritten for the target data in the distributed storage system.

[0095] In an optional embodiment, the step of determining whether there is a first-type shard located on a critical node in the distributed storage system within the erasure coding shard group may include:

[0096] Based on the metadata information corresponding to the target data, locate the storage node where each fragment in the erasure coding fragment group is located;

[0097] If a faulty node is found among the storage nodes, then it is determined that there is a first-type fragment located on a faulty node in the erasure coding fragment group.

[0098] In an optional embodiment, the method may further include:

[0099] Select the target node from the remaining storage nodes after excluding the faulty nodes in the distributed storage system;

[0100] The shard rewrite task is sent to the target node to trigger the target node to execute the shard rewrite task.

[0101] In an optional embodiment, the step of selecting a target node from the remaining storage nodes after excluding failing nodes in the distributed storage system may include:

[0102] From the remaining storage nodes, allocate the destination node for the rewrite of the target fragment;

[0103] Take any of the assigned destination nodes as the target node; or...

[0104] From the remaining storage nodes, select the storage node containing any available fragment used in the process of rewriting the target fragment, and use it as the target node.

[0105] In an optional embodiment, the method may further include:

[0106] Receive fault prediction information provided by the fault prediction tool;

[0107] Based on the fault prediction information, the nodes in the distributed storage system that are about to fail are marked.

[0108] In an optional embodiment, the method may further include:

[0109] If the number of remaining fragments after removing the first type of fragments and the second type of fragments in the erasure coding fragment group is greater than or equal to k, then k usable fragments are randomly selected from the remaining fragments.

[0110] The identifiers of the k available shards are carried in the shard rewrite task, which is used to instruct the generation of the target shard for the target data based on the k available shards.

[0111] In an optional embodiment, the method may further include:

[0112] If the number of remaining fragments is less than k, select the required number of fragments from the first type of fragments to make up to k available fragments.

[0113] In summary, the distributed storage method proposed in this embodiment can create a shard rewrite task for the target data when a shard is detected to be missing within the erasure coding shard group corresponding to the target data. However, when creating the shard rewrite task, instead of only considering the missing shards as the rewrite targets, it proposes to also consider the shards located on the faulty nodes in the distributed storage system within the erasure coding shard group as the rewrite targets. This ensures that after such a shard rewrite task is executed, not only can the missing shards be rewritten, but the shards located on the faulty nodes can also be rewritten incidentally. Based on this piggybacking rewrite mechanism, only a very small amount of write traffic is added to the already required fragment rewrite task. Thus, even if a failing node is falsely reported, the read / write traffic caused by the failing node is effectively reduced because a full copy is no longer performed, ensuring the read / write performance of the distributed storage system. Furthermore, if a failing node is not falsely reported, the mechanism allows for the pre-writing of some data on the failing node, effectively saving the read / write traffic required for data copying after a true failure. This not only ensures data reliability in the distributed storage system but also guarantees its read / write performance. Therefore, in this embodiment, data reliability in the distributed storage system can be improved while ensuring its read / write performance.

[0114] It is worth noting that the technical details of the above embodiments of the distributed storage method can be found in the description of the metadata node in the foregoing system embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0115] Figure 6This is a flowchart illustrating another distributed storage method provided as an exemplary embodiment of the present application. This method can be executed by a target node in a distributed storage system. The distributed storage system may include a metadata node and multiple storage nodes, wherein the metadata node or any storage node can serve as the target node in this embodiment. (Reference) Figure 6 The method may include:

[0116] Step 600: In response to receiving a fragment rewrite task for target data, determine the target fragment to be rewritten corresponding to the target data, wherein the target fragment includes a first type of fragment and a second type of fragment, the first type of fragment is the fragment located on the dying node in the erasure coding fragment group corresponding to the target data, and the second type of fragment is the missing fragment in the erasure coding fragment group;

[0117] Step 601: Read k available fragments from the erasure coding fragment group from the distributed storage system;

[0118] Step 602: Based on the k available shards, execute the shard rewrite task to rewrite the target shard for the target data in the distributed storage system.

[0119] In an alternative embodiment, step 602 may include:

[0120] If none of the first type of shards are included in the k available shards, then based on the k available shards, the first type of shards and the second type of shards are reconstructed for the target data;

[0121] The reconstructed first-type and second-type shards are written to their respective assigned destination nodes to complete the shard rewriting task;

[0122] The destination node is a non-damaging storage node in the distributed storage system.

[0123] In an optional embodiment, the method may further include:

[0124] If the k available shards include p first-type shards, then after reading the p first-type shards, the p first-type shards are written to their assigned destination nodes, where p is an integer greater than 0;

[0125] Based on the k available shards, reconstruct the second type of shards and the remaining first type of shards for the target data;

[0126] The reconstructed second-type fragments and the remaining first-type fragments are written to their respective assigned destination nodes.

[0127] In summary, the distributed storage method proposed in this embodiment no longer only considers missing fragments as rewrite targets in the received fragment rewrite task. Instead, it proposes to also consider fragments located on failing nodes in the distributed storage system within the erasure coding fragment group as rewrite targets. This allows the execution of such fragment rewrite tasks to not only rewrite missing fragments but also to piggyback on the rewrite of fragments located on failing nodes. Based on this piggyback rewrite mechanism, only a very small amount of write traffic is added to the fragment rewrite task that would otherwise need to be executed. Thus, even if a failing node is falsely reported, since a full copy is no longer performed, the read and write traffic caused by the failing node can be effectively reduced, ensuring the read and write performance of the distributed storage system. If the failing node is not falsely reported, since the mechanism allows for the pre-writing of some data on the failing node, the read and write traffic required for data copying after a true failure of the failing node can be effectively saved. This not only ensures the data reliability of the distributed storage system but also guarantees its read and write performance. Therefore, in this embodiment, the data reliability in the distributed storage system can be improved while ensuring the read and write performance of the distributed storage system.

[0128] It is worth noting that the technical details of the above embodiments of the distributed storage method can be found in the description of the target node in the foregoing system embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0129] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 501, 502, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different fragmentation types, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0130] Figure 7 This is a schematic diagram of the node structure in a distributed storage system, provided as another exemplary embodiment of this application. For example... Figure 7 As shown, a node may include: a memory 70, a processor 71, and a communication component 72. The processor 71 may be coupled to the memory 70 and the communication component 72.

[0131] In one case, Figure 7The node shown can serve as a metadata node in a distributed storage system. In this case, processor 71 can execute a computer program in memory 70 for:

[0132] If a missing fragment is detected in the erasure coding fragment group corresponding to the target data, it is determined whether there is a first-type fragment located on a faulty node in the distributed storage system in the erasure coding fragment group.

[0133] If they exist, the first type of fragment and the missing second type of fragment in the erasure coding fragment group are identified as the target fragments to be rewritten;

[0134] A shard rewrite task is created for the target data, the shard rewrite task being used to instruct the target shard to be rewritten for the target data in the distributed storage system.

[0135] In an optional embodiment, when determining whether a first-type fragment located on a failing node in the erasure coding fragment group exists, the processor 71 may specifically be used to:

[0136] Based on the metadata information corresponding to the target data, locate the storage node where each fragment in the erasure coding fragment group is located;

[0137] If a faulty node is found among the storage nodes, then it is determined that there is a first-type fragment located on a faulty node in the erasure coding fragment group.

[0138] In an alternative embodiment, the processor 71 may also be used for:

[0139] Select the target node from the remaining storage nodes after excluding the faulty nodes in the distributed storage system;

[0140] The shard rewrite task is sent to the target node to trigger the target node to execute the shard rewrite task.

[0141] In an optional embodiment, when process 71 selects a target node from the remaining storage nodes after excluding failing nodes in the distributed storage system, it may specifically be used to:

[0142] From the remaining storage nodes, allocate the destination node for the rewrite of the target fragment;

[0143] Use any of the assigned destination nodes as the target node; or...

[0144] From the remaining storage nodes, select the storage node containing any available fragment used in the process of rewriting the target fragment, and use it as the target node.

[0145] In an alternative embodiment, the processor 71 may also be used for:

[0146] Receive fault prediction information provided by the fault prediction tool;

[0147] Based on the fault prediction information, the nodes in the distributed storage system that are about to fail are marked.

[0148] In an alternative embodiment, the processor 71 may also be used for:

[0149] If the number of remaining fragments after removing the first type of fragments and the second type of fragments in the erasure coding fragment group is greater than or equal to k, then k usable fragments are randomly selected from the remaining fragments.

[0150] The identifiers of the k available shards are carried in the shard rewrite task so that the shard rewrite task is used to instruct the generation of the target shard for the target data based on the k available shards.

[0151] In an alternative embodiment, the processor 71 may also be used for:

[0152] If the number of remaining fragments is less than k, select the required number of fragments from the first type of fragments to make up to k available fragments.

[0153] It is worth noting that the technical details of the above-mentioned embodiments of the metadata node can be referred to the relevant descriptions of the metadata node in the foregoing system embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0154] In another case, Figure 7 The node shown can serve as a target node in a distributed storage system. The target node can be a metadata node or any storage node in the distributed storage system. In this case, the processor 71 can execute a computer program in the memory 70 for:

[0155] In response to receiving a fragment rewrite task for target data, the target fragment to be rewritten corresponding to the target data is determined, wherein the target fragment includes a first type of fragment and a second type of fragment, the first type of fragment is the fragment located on the dying node in the erasure coding fragment group corresponding to the target data, and the second type of fragment is the missing fragment in the erasure coding fragment group;

[0156] Read k available fragments from the erasure coding fragment group from the distributed storage system;

[0157] Based on the k available shards, the shard rewrite task is executed to rewrite the target shard for the target data in the distributed storage system.

[0158] In an optional embodiment, when processor 71 performs the shard rewrite task based on the k available shards, it may specifically be used to:

[0159] If none of the first type of shards are included in the k available shards, then based on the k available shards, the first type of shards and the second type of shards are reconstructed for the target data;

[0160] The reconstructed first-type and second-type shards are written to their respective assigned destination nodes to complete the shard rewriting task;

[0161] The destination node is a non-damaging storage node in the distributed storage system.

[0162] In an alternative embodiment, the processor 71 may also be used for:

[0163] If the k available shards include p first-type shards, then after reading the p first-type shards, the p first-type shards are written to their assigned destination nodes, where p is an integer greater than 0;

[0164] Based on the k available shards, reconstruct the second type of shards and the remaining first type of shards for the target data;

[0165] The reconstructed second-type fragments and the remaining first-type fragments are written to their respective assigned destination nodes.

[0166] It is worth noting that the technical details of the above-mentioned embodiments of the target node can be referred to the relevant descriptions of the target node in the foregoing system embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0167] Furthermore, such as Figure 7 As shown, this node also includes other components such as power supply component 73. Figure 7 The diagram only shows a portion of the components and does not imply that the node includes only these components. Figure 7 The components shown.

[0168] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps in the above method embodiments.

[0169] Accordingly, this application also provides a computer program product, which, when executed, can implement the steps in the above method embodiments.

[0170] The above Figure 7The memory in a computer is used to store computer programs and can be configured to store various other data to support operation on a computing platform. Examples of this data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disks, or optical disks.

[0171] The above Figure 7 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0172] The above Figure 7 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.

[0173] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0175] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0176] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0177] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0179] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A distributed storage method, characterized in that, The method, applicable to metadata nodes in a distributed storage system, includes: If a missing fragment is detected in the erasure coding fragment group corresponding to the target data, it is determined whether there is a first-type fragment located on a faulty node in the distributed storage system in the erasure coding fragment group. If they exist, the first type of fragment and the missing second type of fragment in the erasure coding fragment group are identified as the target fragments to be rewritten; A shard rewrite task is created for the target data, the shard rewrite task being used to instruct the target shard to be rewritten for the target data in the distributed storage system.

2. The method according to claim 1, characterized in that, Determining whether a first-type shard exists in the erasure coding shard group on a failing node in the distributed storage system includes: Based on the metadata information corresponding to the target data, locate the storage node where each fragment in the erasure coding fragment group is located; If a faulty node is found among the storage nodes, then it is determined that there is a first-type fragment located on a faulty node in the erasure coding fragment group.

3. The method according to claim 1, characterized in that, Also includes: Select the target node from the remaining storage nodes after excluding the faulty nodes in the distributed storage system; The shard rewrite task is sent to the target node to trigger the target node to execute the shard rewrite task.

4. The method according to claim 3, characterized in that, From the remaining storage nodes in the distributed storage system after excluding the faulty nodes, the target node is selected, including: From the remaining storage nodes, allocate the destination node for the rewrite of the target fragment; Take any of the assigned destination nodes as the target node; or... From the remaining storage nodes, select the storage node containing any available fragment used in the process of rewriting the target fragment, and use it as the target node.

5. The method according to claim 1, characterized in that, Also includes: Receive fault prediction information provided by the fault prediction tool; Based on the fault prediction information, the nodes in the distributed storage system that are about to fail are marked.

6. The method according to claim 1, characterized in that, Also includes: If the number of remaining fragments after removing the first type of fragments and the second type of fragments in the erasure coding fragment group is greater than or equal to k, then k usable fragments are randomly selected from the remaining fragments. The identifiers of the k available shards are carried in the shard rewrite task, which is used to instruct the generation of the target shard for the target data based on the k available shards.

7. The method according to claim 6, characterized in that, Also includes: If the number of remaining fragments is less than k, select the required number of fragments from the first type of fragments to make up to k available fragments.

8. A distributed storage method, characterized in that, The method, applicable to target nodes in a distributed storage system where the target node is not in a state of imminent failure, includes: In response to receiving a fragment rewrite task for target data, the target fragment to be rewritten corresponding to the target data is determined, wherein the target fragment includes a first type of fragment and a second type of fragment, the first type of fragment is the fragment located on the dying node in the erasure coding fragment group corresponding to the target data, and the second type of fragment is the missing fragment in the erasure coding fragment group; Read k available fragments from the erasure coding fragment group from the distributed storage system; Based on the k available shards, the shard rewrite task is executed to rewrite the target shard for the target data in the distributed storage system.

9. The method according to claim 8, characterized in that, Based on the k available shards, the shard rewrite task is executed, including: If none of the first type of shards are included in the k available shards, then based on the k available shards, the first type of shards and the second type of shards are reconstructed for the target data; The reconstructed first-type and second-type shards are written to their respective assigned destination nodes to complete the shard rewriting task; The destination node is a non-damaging storage node in the distributed storage system.

10. The method according to claim 9, characterized in that, Also includes: If the k available shards include p first-type shards, then after reading the p first-type shards, the p first-type shards are written to their assigned destination nodes, where p is an integer greater than 0; Based on the k available shards, reconstruct the second type of shards and the remaining first type of shards for the target data; The reconstructed second-type fragments and the remaining first-type fragments are written to their respective assigned destination nodes.

11. A distributed storage system, characterized in that, The system includes a metadata node and multiple storage nodes. The metadata storage node is used to execute the method described in any one of claims 1-7 to create a sharded rewrite task. The sharded rewrite task is used to trigger a target node in the distributed storage system to execute the method described in any one of claims 8-10, wherein the metadata node or any storage node serves as the target node, and the target node is a non-critical node.

12. A metadata node, characterized in that, Includes memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute one or more computer instructions to perform the distributed storage method according to any one of claims 1-7.

13. A node, characterized in that, Includes memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute one or more computer instructions to perform the distributed storage method according to any one of claims 8-10.

14. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by one or more processors, the one or more processors perform the distributed storage method according to any one of claims 1-10.

15. A computer program product, characterized in that, Includes a computer program that, when executed by one or more processors, causes the one or more processors to perform the distributed storage method according to any one of claims 1-10.