Hierarchical erasure correction encoding method, apparatus and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI BILIBILI TECH CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-07
AI Technical Summary
[0010]利用本公开提供的一个或多个实施例,通过分层纠删编码与分级修复机制,实现局部故障零跨AZ流量修复,有效降低修复时延与带宽开销,同时保障AZ级故障的高效恢复,兼顾高可靠性与动态适配性。
Smart Images

Figure CN122533702A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of storage technology, specifically to a hierarchical erasure coding method and apparatus, electronic equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of cloud computing infrastructure, modern cloud storage systems are evolving from single data centers to multi-Availability Zone (Multi-AZ) architectures to provide greater disaster recovery capabilities and service availability. In a multi-AZ architecture, different AZs are highly isolated in terms of physical facilities, power supply systems, and network connections, forming independent fault domains.
[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0004] This disclosure provides a layered erasure coding method and apparatus, electronic device, computer-readable storage medium, and computer program product.
[0005] According to one aspect of this disclosure, a layered erasure coding method is provided for a multi-availability zone storage system. The method includes: dividing the original data to be stored into multiple original data blocks, wherein the original data blocks are allocated to multiple physically isolated availability zones, and each availability zone stores at least one original data block; performing first erasure coding on the original data blocks in each availability zone to generate a first check block; performing second erasure coding on the original data blocks in all availability zones to generate a second check block; detecting access anomalies of original data blocks, storage nodes, or availability zones in the multi-availability zone storage system to determine the fault impact range and fault level corresponding to the access anomalies; performing first decoding repair using the first check block for first-level faults; and performing second decoding repair using the second check block for scenarios where the first decoding repair fails, or for second-level faults with an impact range greater than that of the first-level fault.
[0006] According to one aspect of this disclosure, a hierarchical erasure coding apparatus is provided for a multi-availability zone storage system. The apparatus includes: a data processing module for dividing raw data to be stored into multiple raw data blocks, wherein the raw data blocks are allocated to multiple physically isolated availability zones, and each availability zone stores at least one raw data block; a first coding module deployed within each availability zone for performing first erasure coding on the raw data blocks within that availability zone to generate a first check block; a second coding module for performing second erasure coding on the raw data blocks of all availability zones to generate a second check block; a fault management module for detecting access anomalies in raw data blocks, storage nodes, or availability zones in the multi-availability zone storage system, determining the fault impact range and fault level corresponding to the access anomaly, and issuing a repair instruction; and a decoding and repair module for performing first decoding and repair using the first check block for first-level faults according to the repair instruction, and performing second decoding and repair using the second check block for scenarios where the first decoding and repair fails, or for second-level faults with an impact range greater than that of the first-level fault.
[0007] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program that, when executed by the at least one processor, implements the method described above.
[0008] According to another aspect of this disclosure, a computer-readable storage medium storing a computer program is also provided, wherein the computer program implements the above-described method when executed by a processor.
[0009] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program implements the above-described method when executed by a processor.
[0010] By utilizing one or more embodiments provided in this disclosure, local fault repair with zero cross-AZ traffic can be achieved through layered erasure coding and hierarchical repair mechanisms, effectively reducing repair latency and bandwidth overhead, while ensuring efficient recovery of AZ-level faults, taking into account both high reliability and dynamic adaptability.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 This is a schematic diagram of the architecture of a layered erasure coding system 100 for multi-availability zone cloud storage according to some embodiments of this disclosure.
[0014] Figure 2 This is a flowchart illustrating the layered erasure coding method 200 in some embodiments of this disclosure.
[0015] Figure 3 This is a schematic diagram of the data writing and hierarchical encoding generation process 300 of the hierarchical erasure coding method for multi-availability zone cloud storage in some embodiments of this disclosure.
[0016] Figure 4 This is a schematic diagram of a hierarchical fault repair process 400 for a hierarchical erasure coding method for multi-availability zone cloud storage in some embodiments of this disclosure.
[0017] Figure 5 This is a schematic diagram of a layered erasure coding device 500 in some embodiments of this disclosure.
[0018] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0019] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0020] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0021] The terminology used in the description of the various examples in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0022] It should be noted that, in any part of this disclosure involving the collection, storage, use, transmission, and processing of data, each stage strictly adheres to the laws, regulations, industry standards, and regulatory requirements of the data source, usage location, and relevant countries and regions to ensure the legality and compliance of data activities. In the collection stage, the purpose, method, and scope of collection are clearly communicated to the data subject in a prominent manner. Collection is conducted only after obtaining the data subject's legal authorization, ensuring that the collection process follows the "minimum necessary" principle and does not exceed the scope of data collection. In the storage stage, storage periods are limited, and data is promptly deleted or anonymized / encrypted after the storage purpose is achieved. In the usage stage, a strict data security protection mechanism is implemented, using field-level desensitization technology and processing the original data according to preset desensitization rules. For different types of data, multiple desensitization strategies, such as data generalization, data anonymization, and data encryption, are employed to effectively mitigate the risk of sensitive information leakage and ensure that all data used is securely processed and desensitized, comprehensively protecting the rights and interests of data subjects and data security. In the transmission and processing stages, the confidentiality and security of data are ensured during transmission and processing.
[0023] As mentioned earlier, with the rapid development of cloud computing infrastructure, distributed cloud storage systems have gradually evolved from single data center architectures to multi-availability zone (Multi-AZ) architectures. Multiple availability zones are isolated from each other at the power supply, network, and physical infrastructure levels, forming independent fault domains, which can significantly improve the disaster recovery capabilities and service availability of cloud storage services. Under a multi-availability zone architecture, system failures exhibit significant hierarchical characteristics: localized failures of single nodes or single data blocks occur frequently but have limited impact, while extreme failures involving the entire availability zone occur with a low probability but are extremely destructive. Simultaneously, there is a significant difference between network bandwidth within an availability zone and network bandwidth across availability zones, resulting in higher latency and operational costs for cross-availability zone data transmission.
[0024] Current erasure coding redundancy technologies for distributed storage systems suffer from numerous technical shortcomings in multi-availability zone scenarios. While traditional RS codes offer superior storage efficiency, single-block repair requires reading large amounts of data across availability zones, resulting in significant bandwidth waste and increased repair latency. Local repair codes, while reducing local repair overhead, employ a fixed code rate design, making it difficult to adapt to the dynamically changing network conditions, fault modes, and business requirements of cloud storage systems. Hierarchical regeneration codes optimize repair bandwidth, but their high computational complexity due to high-dimensional finite-field matrix operations can easily become a system performance bottleneck. Traditional fountain codes, while possessing no code rate adaptive feature, fail to consider the topological heterogeneity of multiple availability zones in their degree distribution design, making it impossible to distinguish the cost differences between intra-zone and cross-zone traffic. Furthermore, existing solutions generally suffer from the coupling of local and global repair strategies, making independent optimization for different fault scenarios difficult and failing to achieve an ideal balance between low-overhead local repair and global extreme fault disaster recovery.
[0025] To overcome the aforementioned deficiencies of existing technologies, this disclosure provides a hierarchical erasure coding method and corresponding apparatus for multi-availability zone storage systems. By explicitly dividing the coding structure into mutually independent local coding layers and global coding layers, local repair and global recovery are decoupled. At the same time, optimized degree distributions are designed for single-availability zone and cross-availability zone scenarios, and a hierarchical decoding and repair mechanism with local priority and global fallback is adopted. This effectively reduces the overhead of local fault repair while ensuring efficient recovery of extreme faults at the availability zone level. Furthermore, dynamic adaptive capability is achieved based on the zero-rate characteristic, balancing high data reliability, low repair overhead, and low computational complexity.
[0026] The present invention will now be described in further detail with reference to the accompanying drawings.
[0027] Figure 1 This is a schematic diagram of the architecture of a layered erasure coding system 100 for multi-availability zone cloud storage according to some embodiments of this disclosure. As an example and not a limitation, system 100 implements the layered coding architecture design of this disclosure, and is divided into three functional layers, from top to bottom: a client proxy layer 110, a metadata server 120, and a multi-availability zone storage layer 130. In some embodiments, the client proxy layer 110 is the execution unit for data preprocessing and global coding, and its functions include three parts: data segmentation, global coding, and data distribution. In some embodiments, the client proxy layer 110 receives the raw data 111 to be stored, first performs a data segmentation operation, and divides the input raw data object into k equal-length system data blocks 112 of a fixed size, with the data block numbers d1, d2, ..., d... kEach block is typically 64MB-256MB in size. It is understood that k here is a positive integer and can be flexibly configured according to the total size of the original data and the preset fixed block size. As an example rather than a limitation, the client proxy layer 110 integrates a global encoding engine 113, which integrates the hierarchical encoding engine of this disclosure. According to the system configuration parameters (k data blocks, p availability zones), it performs global encoding block generation based on AZ-Raptor link degree distribution (AZ-RLD) 115 to generate cross-zone check blocks 114, which are the global encoding blocks defined in this disclosure. In some implementations, the global encoding blocks are generated as follows: (1) from the global degree distribution Medium sampling degree d; (2) From the complete set of original data blocks A set S is formed by uniformly and randomly selecting d data blocks. j (3) Generate global coding blocks through linear combinations over finite fields. , The linear combination ⊕ can be an XOR operation or a more general linear operation in the implementation. Unlike local coding, global coding blocks span data dependencies across multiple availability zones, providing redundancy across fault domains.
[0028] In some implementations, the degree distribution design of the global layer needs to satisfy four characteristics: 1. High average degree, facilitating rapid connection between different availability zones; 2. Low degree initiation, with a small number of low-degree symbols used to initiate Belief Propagation (BP) decoding; 3. Heavy-tailed distribution, with high-order symbols to quickly mix variables across availability zones; 4. Controllable sparsity, controlling complexity and avoiding an average degree reaching the k level. Based on the above objectives, this disclosure designs AZ-RLD 115, defining the global degree distribution as follows:
[0029] Where c1, c2, and c3 are constants used to normalize the probabilities. The range of values for can be 2≤ ≤p, which can be flexibly adjusted according to the scale of multiple availability zones to adapt to data center scenarios of different sizes. It can be understood that the AZ-RLD degree distribution combines the low-degree portion and power-law tail of the Robust Soliton distribution. Through the heavy-tailed distribution characteristic, each global encoded symbol can simultaneously mix variables from multiple availability zones, increasing the information density of the encoded symbol. During availability zone-level fault recovery, only a smaller number of global encoded symbols are needed to complete decoding, thus significantly reducing the amount of cross-availability zone transmission. After the global encoded block is generated, the client proxy layer 110 distributes the segmented original data blocks and the generated cross-zone check blocks to the corresponding availability zones of the multi-availability zone storage layer 130 according to a preset placement strategy, completing the cross-fault domain distribution of data.
[0030] In some implementations, the metadata server 120 communicates with the client agent layer 110 and the multi-availability zone storage layer 130 via a bidirectional control path 121 to maintain the system's global state information. Its functions include namespace management, data placement mapping table maintenance, fault detection, and repair scheduling. Specifically, namespace management maintains the mapping relationship between files and data blocks; the data placement mapping table records the storage location of each data block and encoded block, including availability zone number, node ID, etc.; fault detection monitors the health status of each availability zone and storage node through a heartbeat mechanism; and repair scheduling issues repair instructions to storage nodes based on the fault level (local / global) when a fault is detected. It is understood that the control process of the metadata server 120 can be implemented through control signaling, such as the Monitors Health State & Issues Repair Instructions shown in the figure, which are based on local or global fault judgment. In this example, the control signaling does not affect the actual data flow, i.e., it does not impact the performance of the data transmission path, as shown in the figure. In some implementations, the system of this disclosure consists of p independent availability zones, denoted as p. Each availability zone is isolated from the others at the power supply, network, and physical infrastructure levels, serving as an independent primary fault domain, corresponding to the physical deployment architecture of the multi-availability zone storage layer 130. The original data object is divided into k equal-sized data blocks, denoted as... At the stripe level, data is evenly divided across availability zones, with each availability zone initially storing... Let there be data blocks, and let the set of data blocks in the i-th availability zone be denoted as . As an example and not a limitation, the system disclosed herein considers three types of typical failures with hierarchical structures. The metadata server 120 can accurately detect and classify the three types of failures, as follows: (1) Block-level failures: A single data block or coded block is inaccessible, caused by bad disk blocks, temporary I / O errors, or short-term node anomalies. This type of failure is frequent, accounting for 60-70% of online failures, but its impact is limited; (2) Node-level failures: A single storage node fails, which is equivalent to the simultaneous loss of a limited number of data blocks on that node. This type of failure can be abstracted as the loss of multiple blocks at the coding level, but it is still limited to the same availability zone, accounting for 20-30% of online failures; (3) Availability zone-level failures: The entire availability zone is inaccessible, caused by power outages, network isolation, natural disasters, etc. At this time, all data blocks in the availability zone are unavailable, which is equivalent to the simultaneous loss of k l This type of failure, involving a single data block and its local coded blocks, has a low probability of occurring, accounting for less than 5% of online failures, but it is considered a high-impact event.
[0031] In some implementations, the multi-availability zone storage layer 130 consists of p physically isolated availability zones, including a first availability zone AZ1 131-1, a second availability zone AZ2 131-2, and so on up to the p-th availability zone AZp 131-p. Each availability zone is equipped with several storage nodes, including storage node 132-1 in AZ1, storage node 132-2 in AZ2, and so on up to storage node 132-p in AZp. The functions of the storage nodes include: data storage, local encoding engine, decoding engine, and data repair. Specifically, the storage nodes provide basic data block read / write functions, storing the original data blocks within the corresponding availability zones, including data block 133-1 in AZ1, data block 133-2 in AZ2, and so on up to data block 133-p in AZp. Each storage node in each availability zone has a built-in local encoding engine, including local encoding engine 134-1 in AZ1, local encoding engine 134-2 in AZ2, and so on up to local encoding engine 134-p in AZp. The local encoding engine incorporates a local encoding generator based on AZ-Optimized Bi-Modal Distribution (AZ-OBMD), independently generating local check blocks within the availability zone, including local check block 135-1 in AZ1, local check block 135-2 in AZ2, and so on up to local check block 135-p in AZp. Simultaneously, the storage nodes have a built-in decoding engine, supporting local BP decoding and global BP decoding, and can respond to repair commands from the metadata server 120 to perform data recovery tasks. In some implementations, local encoded blocks are generated using a fountain code method, and for each local encoded block... Its generation process is as follows: (1) From degree distribution (2) Randomly sample a degree d from set D; i A set S is formed by uniformly and randomly selecting d data blocks. j (i) (3) Generate coding blocks through linear combinations over finite fields.
[0032] Here, ⊕ represents a linear combination over a finite field, which in the implementation can be an XOR operation or a more general linear operation. It is understood that a local coded block depends only on data blocks within the same availability zone; therefore, its recovery process can be completed entirely within the availability zone, without requiring cross-availability zone communication, achieving local repair with zero cross-zone traffic.
[0033] In some implementations, the optimization objective of the degree distribution of traditional fountain codes (such as the Robust Soliton distribution) is to minimize the total number of received symbols, assuming that the acquisition cost of all symbols is the same. However, in multi-availability zone scenarios, the acquisition cost of symbols within an availability zone differs greatly from that of symbols across availability zones. Therefore, the local degree distribution optimization objective of this disclosure is to maximize the probability of successful decoding using only symbols within the availability zone. Based on this, AZ-OBMD 136-1, AZ-OBMD 136-2, and up to AZ-OBMD136-p are proposed. This distribution is a bimodal local degree distribution, defined as follows:
[0034] Wherein: the first two terms α and β explicitly control ripple generation, the latter term is the truncated Soliton tail, Z is the normalization constant, and D is the preset maximum degree value. It can be understood that the AZ-OBMD degree distribution adopts a bimodal structure, which includes three parts: (1) degree 1 nodes of proportion α, to ensure that BP decoding can be started; (2) degree 2 nodes of proportion β, to ensure that the ripple continues without collapsing; (3) the truncated Soliton tail, to control the average degree and coding complexity. It can be understood that the parameter selection of this distribution is based on mathematical derivation, which satisfies the local stability decoding condition. By controlling the proportion of low-degree nodes, it ensures that the ripple can evolve stably during small-scale local decoding, and will not be exhausted or stagnant too early, thereby achieving local independent decoding with a high probability. It is easy to understand that the local coding generation process of each availability zone is independent of each other and does not interfere with each other. For example, it can be executed in parallel without waiting for the coding operation of other availability zones to complete, thus completely decoupling from the global coding process. Figure 1The architecture shown achieves functional decoupling between local and global encoding through the distributed deployment of a layered encoding engine. Furthermore, based on a targeted, optimized dual-degree distribution design, it balances low-overhead, rapid repair of local faults with efficient disaster recovery from extreme faults across availability zones. Leveraging the rate-independent nature of fountain codes, the system can generate encoded symbols on demand based on real-time network conditions and fault scenarios, eliminating the need for a pre-determined fixed code rate, thus providing adaptability to dynamic cloud environments. While ensuring high data reliability, this effectively reduces the system's repair bandwidth overhead and computational complexity, adapting to the heterogeneous network characteristics and operational requirements of multi-availability zone cloud storage systems.
[0035] Figure 2 This is a flowchart illustrating a layered erasure coding method 200 in some embodiments of this disclosure. In some embodiments, the layered erasure coding method 200 can be used in a multi-availability zone storage system (i.e., a storage system with multiple availability zones). As shown in the figure, method 200 may include: step 210, dividing the original data to be stored into multiple original data blocks, wherein the original data blocks are allocated to multiple physically isolated availability zones, and each availability zone stores at least one original data block; step 220, performing first erasure coding on the original data blocks in each availability zone to generate a first check block; step 230, performing second erasure coding on the original data blocks in all availability zones to generate a second check block; step 240, detecting access anomalies of original data blocks, storage nodes, or availability zones in the multi-availability zone storage system to determine the fault impact range and fault level corresponding to the access anomaly; step 250, performing first decoding repair using the first check block for a first-level fault; and step 260, performing second decoding repair using the second check block for scenarios where the first decoding repair fails, or for a second-level fault with an impact range greater than that of the first-level fault.
[0036] Optionally, the execution flow of this method 200 may begin with obtaining the raw data to be stored, so as to provide a basic data source for subsequent hierarchical redundancy coding and cross-availability zone storage.
[0037] In some embodiments, in step 210, data partitioning and cross-availability zone allocation can be performed. Specifically, the raw data to be stored can be partitioned into multiple (e.g., equal-length) raw data blocks. In the example, the partitioned raw data blocks can be allocated to multiple physically isolated availability zones, where each availability zone can store at least one raw data block. It is understood that the availability zones here are physically isolated independent fault domains in the cloud storage system, independent of each other at the power supply, network, and infrastructure levels. By allocating raw data blocks across availability zones, basic fault isolation capabilities can be achieved at the data distribution level. As an example and not a limitation, the raw data blocks can be evenly distributed to each availability zone according to striped storage logic, or non-uniformly distributed according to the storage capacity and real-time load status of each availability zone. When an availability zone experiences a temporary access failure, the raw data blocks can also be written to other normally accessible availability zones. This disclosure does not limit this.
[0038] In some embodiments, after the availability zone allocation and write operation of the original data blocks is completed in step 220, a first erasure coding for a single availability zone can be performed. Specifically, a first erasure coding can be performed on the original data blocks in each availability zone to generate a first check block. It is understood that the generated first check block can rely solely on the original data blocks within its own availability zone without needing to obtain any data across availability zones, and the generated first check block can be stored within its own availability zone, thereby providing local redundancy support for subsequent fault repair within the single availability zone. It is easy to grasp that the first erasure coding processes in each availability zone are independent and do not interfere with each other; they can be executed in parallel and asynchronously without waiting for the coding operations of other availability zones to complete, and will not block the main write process of the original data, effectively improving the response efficiency of data writing.
[0039] In some embodiments, in step 230, after all original data blocks have completed the availability zone write operation, a second erasure coding operation can be performed across all availability zones. Specifically, a second erasure coding operation can be performed on the original data blocks of all availability zones to generate a second check block. As an example and not a limitation, the generated second check block relies on original data blocks from at least two availability zones to achieve cross-fault domain redundancy. The generated second check block is stored in each availability zone according to a preset placement rule, providing fallback redundancy support for cross-availability zone data recovery in extreme availability zone failure scenarios.
[0040] In some embodiments, in step 240, fault detection and level determination can be continuously performed throughout the entire lifecycle of data storage. Specifically, access anomalies of raw data blocks, storage nodes, or availability zones in a multi-availability zone storage system can be detected (e.g., in real time), and when an access anomaly is detected, the fault impact range and fault level corresponding to the access anomaly are determined. It is understood that the fault level can be divided into first-level faults and second-level faults, where a first-level fault is a fault whose impact range is limited to within a single availability zone, including block-level faults such as the inaccessibility of a single data block and node-level faults such as the failure of a single storage node, which are frequent small-scale faults; a second-level fault is a fault whose impact range covers all storage nodes of at least one availability zone, i.e., an extreme fault at the availability zone level.
[0041] In some embodiments, after determining the fault level in steps 250 and 260, a graded decoding and repair step 260 can be executed to match the optimal repair path for different fault levels. Specifically, for a first-level fault, a first decoding and repair can be performed using a first verification block; for scenarios where the first decoding and repair fails, or for a second-level fault (e.g., whose impact is greater than that of the first-level fault), a second decoding and repair can be performed using a second verification block. It is understood that the first decoding and repair only uses the original data blocks accessible within the availability zone to which the fault belongs and the first verification block, and the repair process is completed entirely within a closed loop within a single availability zone without the need to transmit any data across availability zones; the second decoding and repair uses the original data blocks and the second verification block that are normally accessible within the availability zone to achieve extreme fault data recovery across availability zones, following the repair principle of prioritizing local faults while providing global fallback.
[0042] Understandably, this method 200, through the deep integration of a layered coding architecture and a hierarchical repair mechanism, achieves complete decoupling between low-overhead, rapid repair of local faults and efficient disaster recovery for extreme faults across availability zones. This effectively reduces bandwidth overhead and latency for cross-availability zone repair in high-frequency local fault scenarios, while providing reliable fallback recovery capabilities for extreme faults at the availability zone level. Based on the layered coding design, the encoding and decoding process can be implemented using simple XOR linear combination operations, resulting in low computational complexity. Furthermore, redundancy can be flexibly adjusted according to the real-time operating status of the system, providing dynamic cloud environment adaptability. This significantly reduces system operation and maintenance overhead and operating costs while ensuring high data reliability.
[0043] Simulation and synthesis results show that the unit redundancy recovery capability index using method 200 reaches 0.93, which is higher than LRC (0.74), AZ-Code (0.81), and ACH-Code (0.86). This demonstrates that this disclosure provides stronger fault tolerance under the same redundancy overhead. Here, unit redundancy recovery capability refers to the maximum proportion of data loss that the system can successfully recover under a given storage redundancy overhead. In the example, its calculation formula is: Unit redundancy recovery capability = Maximum recoverable lost blocks / Total number of coding redundancy blocks, where the numerator is the maximum number of lost blocks that can be tolerated when the decoding success rate is ≥99%, and the denominator is the total number of coding redundancy blocks additionally stored by the system.
[0044] Figure 3 This is a schematic diagram of the data writing and hierarchical encoding generation process 300 of the hierarchical erasure coding method for multi-availability zone cloud storage according to some embodiments of this disclosure. In some embodiments, the entire process follows the principle of local isolation and global correlation, covering five steps: data segmentation, data allocation and writing, local encoding generation, global encoding generation, and metadata persistence. Corresponding to the design of the hierarchical encoding architecture of this disclosure, it achieves mutual independence and functional decoupling between the local encoding layer and the global encoding layer. In some embodiments, the encoding writing process begins with the original data 310 to be stored. First, the data segmentation step 320 is executed, dividing the input original data object into k equal-length system data blocks. A unique identifier is assigned to each data block, and the data block numbers are d1, d2, ..., d... k As an example and not a limitation, each block size can typically be 64MB-256MB. It is understood that k here is a positive integer and can be flexibly configured based on the total size of the original data and the preset fixed block size. This step provides the data foundation for subsequent hierarchical coding and cross-availability zone distribution. As an example and not a limitation, the availability zones involved in this disclosure are multiple physically isolated data centers or server rooms in a cloud storage system, independent of each other in terms of power supply, network, and infrastructure, serving as independent fault domains. Therefore, cross-availability zone distribution of data is the foundation for achieving high disaster recovery capabilities.
[0045] In some implementations, after the original data blocks are segmented, data allocation and writing step 330 is performed. Following striped storage logic, k data blocks are uniformly mapped to p physically isolated availability zones, where p is a positive integer greater than or equal to 2. The i-th availability zone receives the corresponding data block. Each availability zone initially stores k l = k / pEach raw data block is written to the target storage node of the corresponding availability zone after availability zone allocation is completed. This includes writing to the first availability zone AZ-1 340-1, writing to the second availability zone AZ-2 340-2, and so on, until writing to the p-th availability zone AZ-1. p 340-p, and the storage location information of each data block is simultaneously recorded by the metadata server. It's easy to understand that the uniform allocation here is the general implementation method in most scenarios. Additionally or alternatively, non-uniform allocation can also be performed based on the storage capacity of the availability zone, real-time load status, and historical availability metrics. For example, when an availability zone experiences access anomalies, the original data block can be written to another normally accessible availability zone; this disclosure does not limit this. In the example, after the original data block is written, the storage node of each availability zone returns a write success response to the client, providing a timing trigger for subsequent tiered encoding execution.
[0046] In some implementations, after the original data block in each availability zone completes the write operation and the storage node returns a write success response, a local encoding generation step is asynchronously executed within each availability zone. Specifically, this includes executing local encoding generation 350-1 in AZ-1, local encoding generation 350-2 in AZ-2, and so on, up to executing local encoding generation 350-p in AZ-p. It can be understood that this step is triggered asynchronously or synchronously by the storage node within the corresponding availability zone, performing local encoding (i.e., the first erasure coding) on the original data block within the availability zone, encoding the original data block within the availability zone based on AZ-Optimized Bimodal Attitude Distribution (AZ-OBMD), and generating n. l The system generates a local coding block (i.e., the first check block) and stores it within the local availability zone, synchronously updating the metadata mapping table. As used in this paper, when referring to the AZ-OBMD degree distribution used for local coding, it can refer to a bimodal locality distribution optimized for single availability zone scenarios. This distribution explicitly controls the proportion of low-degree nodes, ensuring stable evolution of ripples during decoding and adapting to the rapid decoding and repair needs within a single availability zone. When generating local coding blocks, XOR operations or Galois field operations are used to perform linear combinations, eliminating the need for complex high-dimensional finite field matrix operations and achieving extremely low computational complexity. It is easy to understand that the local coding generation processes in each availability zone are independent and do not interfere with each other, allowing for parallel execution without waiting for the coding operations of other availability zones to complete. Furthermore, each local coding block only depends on the original data block within its own availability zone, without needing to obtain data across availability zones. The coding process is entirely completed within the availability zone, providing a foundation for zero-cross-availability zone traffic repair of subsequent local faults.
[0047] In some implementations, after the client receives a successful write response for all original data blocks, the client agent or dedicated computing node executes global encoding generation step 360. This involves performing global encoding (i.e., second erasure encoding) on all original data blocks in all availability zones, generating a global encoded block (i.e., a second check block) based on the AZ-Raptor link degree distribution (AZ-RLD). Specifically, for each global encoded block to be generated, a degree d is sampled from the AZ-RLD degree distribution. Then, d data blocks are randomly selected from all k original data blocks. These d data blocks are read and XORed to generate the global encoded block, ultimately generating n... g A global encoded block is generated. After the global encoded block is generated, it is distributed and stored in each availability zone 370 in a round-robin manner, and the metadata mapping table is updated synchronously. As an example and not a limitation, when selecting the original data blocks, the original data blocks corresponding to each global encoded block come from at least two different availability zones to achieve redundancy across fault domains. Additionally or alternatively, the global encoded block can be preferentially stored in availability zones that do not contain the original data blocks it depends on to further improve the disaster recovery capability across availability zone faults. As used in this document, when referring to the AZ-RLD degree distribution adopted by the global encoding, it can refer to the degree distribution with heavy-tail characteristics designed for cross-availability zone scenarios. This distribution achieves rapid mixing of original data blocks across availability zones through high-order encoding symbols, which can minimize the amount of cross-zone data transmission in availability zone-level fault scenarios and provide support for efficient data recovery in extreme fault scenarios. Of course, the global encoded blocks in this disclosure can also adopt a flexible placement strategy, which can be placed in p-1 availability zones or evenly distributed across all p availability zones, without strictly ensuring that the number of storage blocks in each availability zone is completely consistent. When the total number of data blocks k is not divisible by the number of available zones p, a non-uniform allocation rule can be adopted.
[0048] In some implementations, after both local and global encoding are completed and the storage information of all encoded blocks has been reported, a metadata persistence step can be optionally performed. Specifically, the metadata server can persistently store the location information of all data blocks and encoded blocks, while also recording the original data block information that each encoded block depends on—that is, the degree information corresponding to each encoded block—for subsequent decoding and repair operations. It can be understood that the mapping table maintained by the metadata server not only contains the physical storage location and dependencies between data blocks and encoded blocks, but also maintains the mapping relationship between original data and original data blocks, providing a metadata foundation for subsequent fault detection and repair scheduling. It should be noted that in this solution, the encoding block dependencies are recorded by default using a complete dependency list. The metadata server records a complete list of dependent data blocks for each local and global encoded block, specifically including the encoding block ID, type, AZ number, degree value, and a list of dependent data block IDs. The dependencies are stored in structured metadata form, and the metadata overhead is negligible compared to the data itself. This is easy to understand. Figure 3 The encoding and writing process 300 shown follows a layered architecture of prioritizing local encoding and providing global encoding as a fallback, achieving decoupling between rapid repair within a single availability zone and disaster recovery from extreme faults across availability zones. In this process 300, local encoding is executed asynchronously, which does not block the main data writing process, improving data writing efficiency and system response performance. Furthermore, based on the characteristics of rateless encoding, the system can adjust the number of encoded blocks generated as needed according to real-time network conditions, fault modes, and business requirements, without needing to predetermine a fixed bitrate, thus possessing adaptive capabilities to cope with dynamic cloud environments.
[0049] Figure 4This is a schematic diagram of a hierarchical fault repair process 400 for a hierarchical erasure coding method for multi-availability zone cloud storage, as described in some embodiments of this disclosure. In some embodiments, this process 400 implements a cost-oriented hierarchical decoder, executing a local priority recovery strategy, corresponding to the local coding layer and global coding layer design of the hierarchical coding architecture, matching the best repair path for different levels of faults. In some embodiments, the hierarchical fault repair process 400 begins with a fault detection step, i.e., a fault 401 is detected. The faults here may include, but are not limited to, three typical types of hierarchical faults: block-level faults, node-level faults, and availability zone (AZ)-level faults. Block-level faults are faults where a single data block or coding block is inaccessible; node-level faults are faults where a limited number of data blocks are lost simultaneously due to the failure of a single storage node; and AZ-level faults are faults where the entire availability zone is inaccessible. After a fault is detected, a fault type determination 402 is performed to determine the scope and level of the fault. As an example, not a limitation, fault types can be categorized into two types: node-level faults and AZ-level faults. Node-level faults are those whose impact is limited to a single availability zone, including block-level faults where a single raw data block becomes inaccessible, and faults where a single storage node fails and all inaccessible raw data blocks are located in the same availability zone. These faults occur frequently, accounting for 60-70% of online faults. AZ-level faults, on the other hand, affect all storage nodes in at least one availability zone. These faults occur less frequently but are still considered high-impact events. In some implementations, when the fault type determination 402 results in a node-level fault, local repair 403 is initiated, which corresponds to the first decoding repair of the local encoding layer. It is understood that local repair follows a closed-loop execution rule within a single availability zone, using only available data within the faulty availability zone to complete the repair, without needing to transmit data across availability zones. Therefore, it can effectively reduce repair bandwidth overhead and repair latency, adapting to cloud storage environments where network bandwidth within an availability zone is significantly higher than cross-availability zone network bandwidth. In the example, the trigger condition for the partial repair path is: the monitoring system detects a block-level or node-level failure, and the rest of the availability zone where the damaged block is located is still accessible. After initiating a partial repair 403, the fault detection and location steps are executed first. The metadata server detects that a data block or node is inaccessible through a heartbeat mechanism, determining the availability zone number i where the fault occurs and the set of damaged data blocks F. D iThe system determines that the number of damaged blocks does not exceed the number of local coding blocks, thus meeting the condition for attempting local repair. Subsequently, it collects available original data blocks and local verification blocks within the damaged availability zone (404), which involves reading all surviving original data blocks and all local coding blocks (i.e., the first verification blocks) within the availability zone. Using original data blocks as variable nodes and local coding blocks as verification nodes, a local decoding bipartite graph is constructed to provide the data and structural foundation for decoding repair. After completing the collection of available data and the construction of the decoding bipartite graph, iterative erasure decoding based on the AZ-OBMD degree distribution (405) is performed, which is the local BP decoding operation defined in this disclosure. As used herein, when referring to the AZ-OBMD degree distribution, it can refer to the AZ-optimized bimodal attitude distribution. This degree distribution adopts a bimodal structure and can be an optimized degree distribution adapted to local decoding scenarios within a single availability zone, thereby ensuring the stable evolution of ripples during decoding and effectively improving the success rate and efficiency of local decoding. In the example, the execution process of iterative erasure decoding is as follows: First, all known surviving original data blocks are marked as "decoded". Then, the following steps are repeated: a. Find a check node with a degree of 1, i.e., a check node that is only connected to a node with an unknown variable; b. Recover the unknown original data block through an XOR operation; c. Mark the newly recovered original data block as "decoded"; d. Update the degree of all check nodes associated with the newly recovered original data block, subtracting the contribution of the decoded variable; e. Repeat the above steps until all damaged original data blocks are recovered, or the lack of check nodes with a degree of 1 causes decoding to stall. After completing the iterative erasure decoding operation, a decoding success judgment 406 is executed, which is a branch node of process 400. Easy to understand, the decoding success judgment 406 corresponds to the decoding double termination conditions defined in this disclosure: when the decoding result is successful, that is, when all blocks in the damaged data block set F are recovered, the repair data write-back step is executed, the recovered data blocks are written to a new storage location, the metadata mapping table is updated, and the affected local coding blocks can be regenerated to restore redundancy, and then the repair completion 411 is directly entered; when the decoding result is unsuccessful, that is, when the decoding stalls due to the absence of a check node with a degree of 1, the global repair 407 is executed, triggering the fallback global repair operation.
[0050] In some implementations, when the fault type determination 402 results in an AZ-level fault, global repair 408 can be directly initiated, which corresponds to the second decoding repair of the global coding layer. Additionally or alternatively, a global repair operation is also triggered when a local repair operation fails to decode. It is understood that global repair is a cross-availability zone fault recovery operation, used to address extreme fault scenarios where repair cannot be completed within a single availability zone, providing a safety net for data reliability. The triggering conditions for the global repair path include two categories: first, local decoding failure, i.e., the number of damaged blocks exceeds the local redundancy capacity; second, an AZ-level fault occurs, making the entire availability zone inaccessible. After initiating global repair (408), the fault assessment step is first executed to determine the set of inaccessible availability zones, calculate the total number of lost original data blocks, and determine whether the recovery conditions are met. Then, the process of collecting available original data blocks and global verification blocks within surviving availability zones (409) is performed. This involves reading all normally accessible original data blocks and all normally accessible global encoded blocks (i.e., second verification blocks) within all surviving availability zones and constructing a global decoding bipartite graph. The variable nodes are all original data blocks, surviving original data blocks are marked as known states, and lost original data blocks are marked as unknown states. The verification nodes are all accessible global encoded blocks, providing the data and structural foundation for global decoding. After triggering the global repair operation, iterative erasure decoding (410) based on the AZ-RLD degree distribution is executed, which is the global BP decoding operation defined in this disclosure. As used herein, when referring to the AZ-RLD degree distribution, it can refer to the AZ-Raptor link degree distribution. This degree distribution has heavy-tailed characteristics and can achieve rapid mixing of original data blocks across availability zones through a preset proportion of high-order encoded symbols, adapting to global decoding scenarios across availability zones and accelerating the decoding convergence process. In the example, the execution process of global iterative erasure decoding is consistent with the local decoding logic, but it covers all original data blocks and global encoded blocks: First, all original data blocks in the surviving availability zones are marked as "decoded". Then, the steps of searching for check nodes with a degree of 1, recovering unknown data blocks, updating the decoding status, and updating the degree of check nodes are executed iteratively until all lost original data blocks are recovered, or the lack of check nodes with a degree of 1 causes decoding to stall. After completing global decoding, data reconstruction and placement steps are performed. The recovered original data blocks are written to the surviving availability zones according to the load balancing strategy, the affected local encoded blocks and global encoded blocks are regenerated, the metadata mapping table is updated, and the system is restored to a redundant state. Additionally or alternatively, if an availability zone is inaccessible for a long period of time, a data migration operation can be triggered to redistribute the original data corresponding to that availability zone to other availability zones and adjust the redundancy parameters to adapt to the new availability zone configuration.
[0051] After completing iterative erasure decoding based on AZ-RLD degree distribution (410) and data reconstruction and metadata updates, all lost original data blocks were recovered, system redundancy was restored to normal levels, and process 400 entered repair completion (411), achieving full-scenario fault data recovery. This is understandable. Figure 4 The illustrated tiered fault repair process 400 follows the layered design philosophy of prioritizing local faults while providing global fallback, achieving complete decoupling between rapid repair within a single availability zone and disaster recovery from extreme cross-availability zone faults, deeply aligning with the design of the layered coding architecture. In simple terms, process 400 optimizes repair costs through pre-classification of fault levels and accurate matching of repair paths: for block-level and node-level local faults accounting for more than 60%, data recovery is achieved through low-overhead local repair, realizing data repair with zero cross-availability zone traffic, effectively reducing repair latency and bandwidth costs; high-overhead global repair is only triggered when local repair fails or AZ-level extreme faults occur, effectively reducing the system's average repair bandwidth overhead and repair latency while ensuring high data reliability. Alternatively or additionally, both local and global decoding in this process 400 can be implemented using the iterative erasure decoding defined in this disclosure, without being limited to a specific BP decoding algorithm, and have strong engineering adaptability and scalability; at the same time, based on the characteristics of rateless encoding, the decoding strategy can be dynamically adjusted according to real-time network conditions and fault modes, and thus has the ability to adapt to dynamic cloud environments.
[0052] Embodiments of other aspects of this disclosure are described below.
[0053] In some embodiments, during the entire process of data writing and encoding generation, a mapping table of storage information and dependencies of the original data blocks can be obtained (e.g., generated and / or recorded). After the first and second check blocks are generated, the mapping table is updated (e.g., synchronously) based on the storage information and dependencies of the first and second check blocks. When the storage location or dependencies of the original data blocks or check blocks (e.g., including the first and / or second check blocks) change, the mapping table is updated in real time. Through the maintenance of the mapping table throughout its entire lifecycle, the location and dependencies of all data blocks and check blocks can be accurately tracked, providing complete and accurate metadata support for subsequent decoding and repair.
[0054] In some embodiments, when performing first erasure coding on raw data blocks within each availability zone, the storage node within that availability zone performs the first erasure coding after the raw data block is allocated to the corresponding availability zone and the write operation is completed. When performing second erasure coding on raw data blocks across all availability zones, the client agent or dedicated compute node performs the second erasure coding after all raw data blocks have been allocated to availability zones and the write operation is completed. By separating the execution subject and timing of hierarchical coding, distributed parallel execution of local coding is achieved, while ensuring global data consistency of global coding, thus improving the efficiency and reliability of coding execution.
[0055] In some embodiments, when the raw data to be stored is divided into multiple raw data blocks, the raw data is first divided into k equal-length raw data blocks according to a preset fixed block size, and each raw data block is assigned a unique identifier; then, according to striped storage logic, the k raw data blocks are mapped to p physically isolated availability zones, where p ≥ 2. Each availability zone stores at least one raw data block, and the number of raw data blocks stored in each availability zone can be the same or different. Furthermore, when an access anomaly occurs in an availability zone, the raw data block is written to another availability zone that can be accessed normally. Through data partitioning and flexible striped allocation strategies, the standardization of data distribution across availability zones is ensured, and the real-time status of availability zones can be adapted, improving the flexibility and disaster recovery capability of data writing.
[0056] In some embodiments, when the raw data to be stored is divided into multiple raw data blocks, the availability zone number, storage node identifier, physical storage location, and unique identifier corresponding to each raw data block can also be recorded in the mapping table. By accurately recording the full-dimensional information of the raw data blocks through the mapping table, rapid location and traceability of the data blocks can be achieved, providing accurate metadata for fault detection and repair.
[0057] In some embodiments, when performing first erasure coding on raw data blocks within each availability zone to generate a first check block, a preset first degree distribution can be configured for each availability zone, the first degree distribution adopting a bimodal structure. For each first check block to be generated, a coding degree value is sampled from the first degree distribution, raw data blocks matching the number of coding degree values are selected from the availability zone, and a linear combination operation is performed on the selected raw data blocks to generate the first check block. In the example, the generated first check block can be stored in a storage node within the availability zone. By adapting the bimodal structure degree distribution design to a single availability zone scenario, the stable evolution of ripples during decoding can be guaranteed, improving the success rate and decoding efficiency of decoding within a single availability zone.
[0058] In some embodiments, when performing first erasure coding on the original data blocks in each availability zone to generate a first check block, the dependencies and storage locations of the original data blocks corresponding to the first check block can also be recorded in a mapping table. By recording the dependencies and storage locations of the first check block in real time, the corresponding check block can be quickly retrieved when a local failure occurs, shortening the response latency of local repair.
[0059] In some embodiments, when selecting a raw data block matching the number of coding degree values from the availability zone, a raw data block matching the number of coding degree values is selected from all raw data blocks in the availability zone; simultaneously, the selection result of the raw data block corresponding to each first check block is recorded as a dependency basis for the decoding process. By recording the selection result of the raw data block of the first check block, accurate matching of dependencies during the decoding process can be ensured, decoding errors can be avoided, and the accuracy of local decoding can be improved.
[0060] In some embodiments, when performing second erasure coding on the original data blocks of all availability zones to generate a second check block, a preset second degree distribution can be configured first. For each second check block to be generated, a coding degree value is sampled from the second degree distribution. Original data blocks matching the number of coding degree values are selected from the original data blocks of all availability zones. A linear combination operation is performed on the selected original data blocks to generate the second check block. In the example, the generated second check blocks are stored in the storage nodes of each availability zone according to preset rules. By adapting the second degree distribution design to cross-availability zone scenarios, efficient mixing of cross-availability zone data can be achieved, improving the information density of global coding and reducing the bandwidth overhead of cross-availability zone repair.
[0061] In some embodiments, when performing second erasure coding on the original data blocks of all availability zones to generate a second check block, the dependencies and storage locations of the original data blocks corresponding to the second check block can also be recorded in the mapping table. By recording the dependencies and storage locations of the second check block, the corresponding global check block can be quickly retrieved when global repair is triggered, ensuring the efficient execution of the global repair process.
[0062] In some embodiments, the second degree distribution has a heavy-tailed characteristic. When performing second erasure coding on the original data blocks of all availability zones to generate the second check block, a preset proportion of higher-order coding symbols is also set. The coding degree values corresponding to the higher-order coding symbols cover the original data blocks of at least two availability zones. In the example, when selecting the original data blocks, the original data blocks corresponding to each second check block come from at least two different availability zones. Alternatively or additionally, the generated second check blocks can be stored in availability zones that do not contain their dependent original data blocks and distributed to each availability zone in a round-robin manner. By using the degree distribution with heavy-tailed characteristics and the check block placement strategy across availability zones, the cross-fault domain disaster recovery capability of global coding can be further improved, while minimizing the amount of cross-availability zone data transmission during global repair.
[0063] In some embodiments, when selecting a raw data block matching the number of coding degree values from the raw data blocks of all available zones, a raw data block matching the number of coding degree values is selected from all raw data blocks of all available zones. Simultaneously, the selection result of the raw data block corresponding to each second check block and its associated available zone information are recorded as a dependency basis for the decoding process. By recording the selection result of the raw data block for the second check block and the available zone information, accurate matching of dependencies during the global decoding process can be ensured, improving the accuracy and convergence speed of global decoding.
[0064] In some embodiments, when performing linear combination operations on selected original data blocks, XOR operations or Galois field operations are used to perform linear combination; when generating the first check block, generating the second check block, performing the first decoding repair, and performing the second decoding repair, XOR operations or Galois field operations are all used to perform linear combination. It can be understood that the encoding and decoding implementation based on pure XOR operations can effectively reduce the computational complexity of the encoding and decoding process, improve the system's throughput performance, and is easy to implement with hardware acceleration.
[0065] In some embodiments, when performing the first decoding repair using the first verification block for a first-level fault, the availability zone to which the access anomaly belongs, and the set of inaccessible original data blocks within that availability zone, can be located first. After determining that the number of inaccessible original data blocks does not exceed the number of first verification blocks within the availability zone, the first decoding repair is initiated. All accessible original data blocks and all first verification blocks within the availability zone are read, and a first decoding bipartite graph is constructed using original data blocks as variable nodes and first verification blocks as verification nodes. All accessible original data blocks are marked as decoded, and iterative erasure decoding operations are performed until all inaccessible original data blocks are recovered or decoding stalls. When decoding stalls, the first decoding repair is deemed to have failed. In the example, at least one of the following operations can be performed to write the recovered original data blocks to the target storage node of the availability zone: updating the storage location of the affected first verification blocks; regenerating the affected first verification blocks. Through the local decoding repair process, closed-loop repair of faults within a single availability zone can be achieved without transmitting data across availability zones, effectively reducing the bandwidth overhead and repair latency of local fault repair.
[0066] In some embodiments, when performing the iterative erasure decoding operation corresponding to the first decoding repair, the following steps are repeated until all inaccessible original data blocks are marked as decoded, or the decoding stalls due to the absence of a check node with a degree of 1: A check node with a degree of 1 is searched in the first decoding bipartite graph, where each check node connects only to one inaccessible original data block; when a check node with a degree of 1 is found, the inaccessible original data block corresponding to that check node is recovered through a linear combination operation, the newly recovered original data block is marked as decoded, and the degree of all check nodes associated with the newly recovered original data block is updated; when no check node with a degree of 1 is found, the pruning or backtracking operation corresponding to the iterative erasure decoding is performed to continue decoding. The iterative decoding process based on check nodes with a degree of 1 ensures the stable progress of the local decoding process, while pruning or backtracking operations further improve the success rate of local decoding.
[0067] In some embodiments, for scenarios where the first decoding repair fails, or for second-level faults with a wider impact than the first level of fault, when performing the second decoding repair using the second verification block, the impact range of the access anomaly can be assessed first to determine the set of availability zones that cannot be accessed normally, and all lost original data blocks corresponding to that set; read all original data blocks and all accessible second verification blocks in all accessible availability zones, using all original data blocks as variable nodes and accessible second verification blocks as verification nodes, to construct a second decoding bipartite graph; mark all original data blocks in all accessible availability zones as decoded, and perform iterative erasure decoding operations until all lost original data blocks are recovered or decoding stops. In the example, the recovered original data blocks can be written to the storage nodes of accessible availability zones to regenerate the affected first and second verification blocks. Through the global decoding repair process, efficient recovery from extreme failures at the availability zone level can be achieved, providing a safety net for data reliability, and the system redundancy can be quickly rebuilt after recovery.
[0068] In some embodiments, when performing the iterative erasure decoding operation corresponding to the second decoding repair, the following steps are repeated until all lost original data blocks are marked as decoded, or the decoding stalls due to the absence of a check node with a degree of 1: A check node with a degree of 1 is searched in the second decoding bipartite graph, where each check node with a degree of 1 connects to only one lost original data block; when a check node with a degree of 1 is found, the lost original data block corresponding to that check node is recovered through a linear combination operation, the newly recovered original data block is marked as decoded, and the degree of all check nodes associated with the newly recovered original data block is updated; when no check node with a degree of 1 is found, the pruning or backtracking operation corresponding to the iterative erasure decoding is performed to continue decoding. By adapting the iterative decoding process to the global decoding scenario, rapid convergence of the global decoding process can be guaranteed, while the success rate of global decoding under extreme failure scenarios can be improved through pruning or backtracking operations.
[0069] In some embodiments, operational status data of the multi-availability zone storage system is also collected, including network bandwidth data, access anomaly frequency data, and node load data. Based on the operational status data, the number of generated first and second check blocks is adjusted. Based on the adjusted number of check blocks, new check blocks are generated. Dynamic adjustment of redundancy based on the system's real-time operational status allows the system to adapt to dynamically changing cloud environments, achieving a dynamic balance between storage efficiency and data reliability.
[0070] In some embodiments, the first level of failure includes block-level failure or node-level failure. A block-level failure is a failure where a single raw data block is inaccessible, and a node-level failure is a failure where a single storage node fails and all inaccessible raw data blocks are located in the same availability zone. The second level of failure is a failure where all storage nodes in at least one availability zone are inaccessible. By finely classifying the failure levels, optimal repair paths can be matched to different failures, thereby optimizing the repair cost.
[0071] In some embodiments, when mapping k original data blocks to p physically isolated availability zones according to striped storage logic, firstly, stripe units corresponding to the original data are created. Each stripe unit contains k original data blocks and corresponding parity blocks. A unique stripe identifier is assigned to each stripe unit, and the association between the stripe identifier and all original data blocks and parity blocks within the stripe unit is recorded. When an original data block within a stripe unit is modified, the version information of the corresponding stripe unit is updated, and the associated parity blocks within the stripe unit are also updated. Through stripe unit management, unified management of data and parity blocks can be achieved. At the same time, version control ensures the consistency of encoding redundancy after data modification, avoiding decoding errors caused by data inconsistency.
[0072] In some embodiments, the execution timing of the first erasure coding and the second erasure coding follows these rules: the first erasure coding is triggered after the original data block is written to the corresponding availability zone and a write success response is returned; the second erasure coding is triggered after all original data blocks have been allocated to availability zones and written; and the original data writing process is marked as complete after both the first and second erasure coding are completed. Through the timing control of layered coding, both rapid response to original data writing and orderly generation of layered redundancy are ensured, preventing the coding process from blocking the main write process and improving the system's write performance.
[0073] In some embodiments, the first level of fault is a fault whose impact is limited to a single availability zone, and the second level of fault is a fault whose impact covers all storage nodes in at least two availability zones; the first decoding repair uses only the original data blocks and the first check block that are accessible within the availability zone to which the fault belongs, and the second decoding repair uses the original data blocks and the second check block that are normally accessible within the availability zone.
[0074] To verify the effectiveness of the layered AZ-Fountain coding proposed in this disclosure in a multi-availability zone distributed storage environment, a complete distributed storage prototype system was implemented based on the Golang language, and a corresponding physical test cluster was built. This cluster deploys 24 physical storage nodes, logically divided into 4 logical availability zones, each containing 6 nodes. Each physical node mounts 36 enterprise-grade HDDs, totaling 864 physical disks, which can realistically simulate the concurrent I / O contention scenarios of a large-scale storage pool. The system follows a strict disk-level isolation strategy when placing data; all shards within the same stripe are forcibly distributed discretely across different physical disks, ensuring that data loss within a stripe is independent in the event of a single disk failure.
[0075] This test performed engineering parameter optimization for both the local coding layer and the global coding layer. For the local layer, the AZ-OBMD degree distribution was optimized with a compromise of α=0.15, β=0.20, D=12, and an average degree of 4.2, balancing decoding stability and decoding success rate in multi-block loss scenarios. The repair success rate was further optimized through precoding. The global layer AZ-RLD degree distribution was optimized with parameters, and the final average degree was 4.7, which is suitable for the core design requirements of cross-availability zone coding and ensures decoding convergence efficiency in AZ-level failure scenarios.
[0076] This test fairly compared the proposed solution with mainstream industry coding schemes such as RS-Code, LRC, AZ-Code, and ACH-Code under an equivalent uniform storage redundancy rate of 1.45x. All compared schemes were configured with equivalent redundancy conditions to ensure the comparability and rigor of the test results. The test evaluated the solution's three core capabilities: local fault repair overhead, local decoding success rate, and AZ-level disaster recovery efficiency. The test results verified the technical advantages of the proposed solution in multi-availability zone storage scenarios.
[0077] Figure 5 This is a schematic diagram of a layered erasure coding device 500 in some embodiments of this disclosure.
[0078] In some embodiments, the hierarchical erasure coding apparatus 500 can be used in a multi-availability zone storage system, the apparatus comprising: The data processing module 510 is used to divide the raw data to be stored into multiple raw data blocks, wherein the raw data blocks will be allocated to multiple physically isolated availability zones, and each availability zone stores at least one raw data block. The first encoding module 520 is deployed within each availability zone and is used to perform first erasure encoding on the original data blocks within that availability zone to generate a first verification block. The second encoding module 530 is used to perform second erasure encoding on the original data blocks of all availability zones to generate a second check block; The fault management module 540 is used to detect access anomalies in raw data blocks, storage nodes or availability zones in a multi-availability zone storage system, determine the fault impact range and fault level corresponding to the access anomaly, and issue repair instructions. The decoding and repair module 550 is used to perform a first decoding and repair for a first-level fault using the first verification block according to the repair instruction; and to perform a second decoding and repair for a scenario where the first decoding and repair fails, or for a second-level fault whose impact is greater than that of the first-level fault, using the second verification block.
[0079] In some embodiments, the hierarchical erasure coding apparatus 500 further includes a metadata management module, the metadata management module being used for: Obtain the mapping table of storage information and dependencies of the original data blocks; After the first and second verification blocks are generated, the mapping table is updated based on the storage information and dependencies of the first and second verification blocks. When the storage location or dependency of the original data block, the first check block, or the second check block changes, the mapping table is updated; Maintain the mapping relationship between the original data and the original data blocks; Provides metadata query and update services for other modules.
[0080] In some embodiments, the data processing module 510 is further configured to: The raw data is divided into k equal-length raw data blocks according to a preset fixed block size, and each raw data block is assigned a unique identifier. According to the striped storage logic, k original data blocks are mapped to p physically isolated availability zones, where p ≥ 2. Each availability zone stores at least one original data block. The number of original data blocks stored in each availability zone may be the same or different. When an access anomaly occurs in an availability zone, the original data block is written to another availability zone that can be accessed normally.
[0081] In some embodiments, the first encoding module 520 is further configured to: After the original data block is allocated to the corresponding availability zone and the write operation is completed, the first erasure coding is performed; Each availability zone is configured with a preset first degree distribution, which adopts a bimodal structure. For each first check block to be generated, the coding degree value is sampled from the first degree distribution, and the number of original data blocks matching the number of coding degree values is selected from the available area. A linear combination operation is performed on the selected original data blocks to generate the first check block. The first verification block generated is stored in the storage node within the availability zone.
[0082] In some embodiments, the second encoding module 530 is further configured to: After all raw data blocks have been allocated to the availability zone and written, the second erasure coding is performed. Configure a second-degree distribution with heavy-tailed characteristics; For each second check block to be generated, the coding degree value is sampled from the second degree distribution, and the original data blocks that match the number of coding degree values are selected from the original data blocks of all available areas. A linear combination operation is performed on the selected original data blocks to generate the second check block. The generated second verification block is stored in an availability zone that does not contain the original data block it depends on, and is distributed to each availability zone in a round-robin manner.
[0083] In some embodiments, the first encoding module 520 and the second encoding module 530 perform a linear combination using an XOR operation or a Galois field operation.
[0084] In some embodiments, when the decoding and repair module 550 performs the first decoding and repair, it is used to: Locate the availability zone to which the access exception belongs and the set of inaccessible raw data blocks within that availability zone; If the number of inaccessible raw data blocks does not exceed the number of the first checksum blocks in the availability zone, initiate the first decoding repair. Read the accessible raw data blocks and all first check blocks within the availability zone to construct the first decoding bipartite graph; Mark accessible raw data blocks as decoded and perform iterative erasure decoding operations until all inaccessible raw data blocks are recovered or decoding stops. Specifically, at least one of the following operations is performed to write the recovered original data block to the storage node of the availability zone: update the storage location of the affected first check block; regenerate the affected first check block.
[0085] In some embodiments, when the decoding and repair module 550 performs the second decoding and repair, it is used to: Identify the set of availability zones that are not normally accessible and the corresponding complete set of lost original data blocks; Read the original data blocks and all accessible second check blocks within the available area to construct the second decoding bipartite graph; Mark the original data blocks in the availability zone that can be accessed normally as decoded, and perform iterative erasure decoding operations until all lost original data blocks are recovered or decoding stops. The recovered original data blocks are written to storage nodes in the availability zone that can be accessed normally, in order to regenerate the affected first and second check blocks.
[0086] In some embodiments, the fault management module 540 is further configured to: The access status of each availability zone and storage node is monitored through a heartbeat mechanism to detect access anomalies; The fault levels corresponding to access anomalies are classified as either Level 1 faults or Level 2 faults. For first-level faults, a local repair command is issued; for first-level decoding repair failures or second-level faults, a global repair command is issued.
[0087] In some embodiments, the hierarchical erasure coding apparatus 500 further includes an adaptive adjustment module, the adaptive adjustment module being used for: Collect data on network bandwidth, frequency of access anomalies, and node load operation status of the multi-availability zone storage system; Based on the aforementioned operational status data, adjust the number of first and second verification blocks generated. A new verification block is generated based on the adjusted quantity.
[0088] It should be understood that Figure 5 The various modules or units of the apparatus 500 shown can be connected to the reference. Figure 2 The steps in method 200 described correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to apparatus 500 and its included modules and units. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0089] Although specific functions have been discussed with reference to specific modules above, it should be noted that the functions of the units discussed in this article can be divided into multiple units, and / or at least some functions of multiple units can be combined into a single unit.
[0090] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program that, when executed by the at least one processor, implements the method described above.
[0091] According to another aspect of this disclosure, a computer-readable storage medium storing a computer program is also provided, wherein the computer program implements the above-described method when executed by a processor.
[0092] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program implements the above-described method when executed by a processor.
[0093] The following are extensions and alternative implementations of the layered erasure coding method and apparatus, electronic device, computer-readable storage medium, and computer program product described in detail above with reference to the accompanying drawings. Those skilled in the art, based on the foregoing description of this disclosure, can readily understand and implement these extensions without departing from the inventive concept of this disclosure; therefore, these extensions and alternatives also fall within the protection scope of this disclosure.
[0094] Alternatively, an adaptive optimization mechanism can be designed for the degree distribution parameters. The degree distribution parameters used in the aforementioned scheme are pre-set static values. In this extended scheme, real-time statistical data such as fault type distribution, local decoding success rate, and cross-availability zone bandwidth utilization can be collected during system operation. A parameter tuning engine can be built based on reinforcement learning or gradient optimization algorithms to adjust the relevant parameters of local degree distribution and global degree distribution, as well as the number of generated local and global check blocks, online. The encoding blocks under the new parameters are generated progressively in the background to achieve a smooth transition of configuration and better adapt to the dynamically changing operating environment of the system.
[0095] Alternatively, a hybrid coding scheme combining layered fountain coding and traditional erasure coding can be adopted to construct a three-layer coding structure. The first layer uses RS codes to generate local check blocks in each availability zone, providing deterministic and rapid repair capabilities for the most common single-block failures. The second layer uses fountain coding with the aforementioned AZ-OBMD degree distribution to generate additional local redundancy in each availability zone, enhancing fault tolerance in multi-block erasure scenarios. The third layer uses fountain coding with the aforementioned AZ-RLD degree distribution to generate global redundancy, providing backup disaster recovery capabilities for cross-availability zone failures. At the same time, the two-layer or three-layer coding structure can be flexibly selected according to the importance of the data to be stored.
[0096] Alternatively, a non-uniform data placement strategy can be adopted. The aforementioned scheme assumes that data is evenly distributed across availability zones. In this extended scheme, the corresponding weights can be calculated based on indicators such as historical availability, network bandwidth, and storage capacity of each availability zone. Based on the weights, a corresponding number of raw data blocks are allocated to each availability zone. Availability zones with lower weights are allocated fewer raw data blocks and more encoded blocks. At the same time, based on the actual number of data blocks in each availability zone, a locality distribution parameter is customized for each availability zone to better utilize the resources of heterogeneous availability zones and improve the overall system performance.
[0097] Additionally or alternatively, global coding can be further divided into multi-level coding structures to adapt to storage scenarios with wider geographical distribution. First, multiple availability zones are divided into several regional groups according to geographical regions, with availability zones in the same city or geographically adjacent areas grouped into the same regional group. Regional-level fountain coding blocks are generated between availability zones within each regional group to cope with availability zone failures within the region. Cross-regional fountain coding blocks are generated on top of the original data blocks of all regional groups to cope with extreme failures such as city-level network outages or geographical disasters. At the same time, it can reduce the number of coding blocks across long-distance regions, reduce the overhead of long-distance data transmission, and provide a more granular level of fault tolerance.
[0098] Alternatively, an incremental encoding update mechanism can be designed to adapt to data modification scenarios. In the aforementioned scheme, when data is modified, all encoding blocks involving the data block need to be regenerated. In this extended scheme, a dependency graph of all encoding blocks involved in the generation of each original data block can be maintained. When an original data block is modified, only the local encoding blocks and global encoding blocks involving the data block are recalculated. At the same time, a version number is introduced for the encoding blocks to ensure that the version-matching encoding blocks are used during the decoding process, thereby reducing the encoding overhead during data updates and improving write performance in data modification scenarios.
[0099] Alternatively, an encoding scheme combining erasure coding and fountain coding can be adopted. First, RS codes or other MDS codes are used to encode the original data block in the first round to generate k+m MDS coding symbols. Then, the generated MDS coding symbols are used as virtual data blocks, and the aforementioned layered fountain coding scheme is applied on top of the virtual data blocks to combine the superior storage efficiency of MDS codes with the rate-adaptive characteristics of fountain codes, thereby further improving the fault tolerance and operational flexibility of the system.
[0100] Alternatively, an intelligent repair path selection mechanism can be designed to replace the fixed local priority and global fallback repair strategy. First, a repair cost prediction model that considers network bandwidth, transmission latency, and real-time load of storage nodes is constructed. For each detected fault, multiple candidate repair paths are generated, including pure local repair, partial cross-availability zone repair, and global repair. Based on the real-time cost evaluation results, the optimal repair path is selected for the fault to further optimize the system's repair efficiency in complex dynamic operating environments.
[0101] Alternatively, an intelligent caching mechanism for coded blocks can be designed to cache some global coded blocks in client proxies, relay nodes, or other availability zones. First, the repair value of each global coded block is evaluated based on the historical failure probability of the availability zone and the access frequency of data blocks. High-value global coded blocks are cached in multiple preset locations. At the same time, when the original data block is modified, the corresponding cached coded block is invalidated in a timely manner to ensure cache consistency, reduce data transmission latency during the global repair process, and further improve the repair efficiency of cross-availability zone failures.
[0102] See Figure 6 The present invention describes a structural block diagram of an electronic device 600 that can serve as a storage node, client agent, dedicated computing node, or server, etc., as an example of hardware devices applicable to various aspects of the present disclosure. The electronic device can be different types of computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0103] like Figure 6 As shown, the electronic device 600 may include at least one processor 601, working memory 602, input unit 604, display unit 605, speaker 606, storage unit 607, communication unit 608 and other output units 609 that are capable of communicating with each other via system bus 603.
[0104] Processor 601 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 601 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Processor 601 may be configured to acquire and execute computer-readable instructions stored in working memory 602, storage unit 607, or other computer-readable media, such as program code of operating system 602a, program code of application program 602b, etc.
[0105] Working memory 602 and storage unit 607 are examples of computer-readable storage media for storing instructions that are executed by processor 601 to perform the various functions described above. Working memory 602 may include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, storage unit 607 may include hard disk drives, solid-state drives, removable media including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Working memory 602 and storage unit 607 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that can be executed by processor 601 as a specific machine configured to perform the operations and functions described in the examples herein.
[0106] Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone and / or remote control. Output unit can be any type of device capable of presenting information, and can include, but is not limited to, display unit 605, speaker 606 and other output units 609. Other output units 609 can include, but are not limited to, video / audio output terminals, vibrators and / or printers. Communication unit 608 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0107] The application program 602b in the working memory 602 can be loaded to execute the various methods and processes described above, for example... Figure 2Boxes 210 to 260 in the diagram. For example, in some embodiments, the hierarchical erasure coding method for a multi-availability zone storage system may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 607. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 600 via storage unit 607 and / or communication unit 608. When the computer program is loaded and executed by processor 601, one or more steps of the hierarchical erasure coding method for a multi-availability zone storage system described above may be performed. Alternatively, in other embodiments, processor 601 may be configured to perform the hierarchical erasure coding method for a multi-availability zone storage system by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0113] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0114] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0115] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A hierarchical erasure coding method for a multi-availability zone storage system, the method comprising: The raw data to be stored is divided into multiple raw data blocks, which are then allocated to multiple physically isolated availability zones, with each availability zone storing at least one raw data block. Perform the first erasure coding on the raw data blocks in each availability zone to generate the first check block; Perform a second erasure coding on the original data blocks of all availability zones to generate a second check block; Detect access anomalies in raw data blocks, storage nodes, or availability zones in a multi-availability zone storage system to determine the scope and severity of the faults corresponding to the access anomalies. For the first-level fault, the first decoding repair is performed using the first verification block; For scenarios where the first decoding repair fails, or for second-level faults with a wider impact than the first-level fault, the second verification block is used to perform a second decoding repair.
2. The method of claim 1, further comprising: Obtain the mapping table of storage information and dependencies of the original data blocks; After the first and second verification blocks are generated, the mapping table is updated based on the storage information and dependencies of the first and second verification blocks. The mapping table is updated when the storage location or dependency of the original data block, the first check block, or the second check block changes.
3. The method as described in claim 1, wherein, The first erasure coding for the original data blocks in each availability zone includes: After the original data block is allocated to the corresponding availability zone and the write operation is completed, the storage node in that availability zone performs the first erasure coding. The second erasure coding is performed on the original data blocks of all availability zones, including: After all raw data blocks have been allocated to availability zones and written, the second erasure coding is performed by the client agent or a dedicated compute node.
4. The method of claim 1, wherein, The step of dividing the raw data to be stored into multiple raw data blocks includes: The raw data is divided into k equal-length raw data blocks according to a preset fixed block size, and each raw data block is assigned a unique identifier. According to the striped storage logic, k original data blocks are mapped to p physically isolated availability zones, where p ≥ 2. Each availability zone stores at least one original data block. The number of original data blocks stored in each availability zone may be the same or different. When an access anomaly occurs in an availability zone, the original data block is written to another availability zone that can be accessed normally.
5. The method of claim 4, wherein, The step of dividing the raw data to be stored into multiple raw data blocks further includes: The mapping table records the availability zone number, storage node identifier, physical storage location, and unique identifier corresponding to each original data block.
6. The method of claim 1, wherein, The step of performing first erasure coding on the original data blocks within each availability zone to generate a first check block includes: Each availability zone is configured with a preset first degree distribution, which adopts a bimodal structure. For each first check block to be generated, the coding degree value is sampled from the first degree distribution, and the number of original data blocks matching the number of coding degree values is selected from the availability zone. A linear combination operation is performed on the selected original data blocks to generate the first check block, wherein the generated first check block is stored in the storage node within the availability zone.
7. The method of claim 6, wherein, The step of performing first erasure coding on the original data blocks in each availability zone to generate a first check block also includes: Record the original data block dependencies and storage locations corresponding to the first verification block in the mapping table.
8. The method of claim 6, wherein, The step of selecting raw data blocks from the availability zone that match the number of coding degree values includes: Select raw data blocks from all raw data blocks within the availability zone, ensuring the number of raw data blocks matches the number of coding degree values. Record the original data block selection result corresponding to each first check block, as the basis for the decoding process.
9. The method of claim 1, wherein, The second erasure coding is performed on the original data blocks of all availability zones to generate a second check block, including: Configure a preset second-degree distribution; For each second check block to be generated, the coding degree value is sampled from the second degree distribution. Then, original data blocks matching the number of coding degree values are selected from the original data blocks of all available areas. A linear combination operation is performed on the selected original data blocks to generate the second check block. The generated second verification block is stored in the storage nodes of each availability zone according to preset rules.
10. The method of claim 9, wherein, The step of performing second erasure coding on the original data blocks of all availability zones to generate a second check block further includes: The mapping table records the original data block dependencies and storage locations corresponding to the second verification block.
11. The method of claim 9, wherein, The second degree distribution has heavy-tailed characteristics. The step of performing second erasure coding on the original data blocks of all available zones to generate a second check block further includes: A preset ratio of high-order coding symbols is set, and the coding degree value corresponding to the high-order coding symbols covers the original data blocks of at least two availability zones; When selecting the original data blocks, each second check block corresponds to an original data block that comes from at least two different availability zones, and wherein... The generated second verification block is stored in an availability zone that does not contain the original data blocks it depends on, and is distributed to each availability zone in a round-robin manner.
12. The method of claim 9, wherein, The step of selecting raw data blocks from all available zones that match the number of coding degree values includes: Select raw data blocks from all raw data blocks in all availability zones, matching the number of coding degree values. Record the original data block selection result and the availability zone information corresponding to each second check block as the basis for the decoding process.
13. The method of claim 4, wherein, The step of mapping k original data blocks to p physically isolated availability zones according to striped storage logic includes: Create a stripe unit corresponding to the original data, wherein the stripe unit contains k original data blocks and corresponding check blocks; Each stripe unit is assigned a unique stripe identifier, which indicates the association between all raw data blocks and check blocks within the stripe unit; When the original data block within a stripe unit is modified, the version information of the corresponding stripe unit is updated, and the check block associated with the original data block within the stripe unit is also updated.
14. The method of claim 1, wherein, The first decoding repair using the first verification block for the first level fault includes: Locate the availability zone to which the access exception belongs, and the set of inaccessible raw data blocks within that availability zone; If the number of inaccessible raw data blocks does not exceed the number of the first checksum blocks in the availability zone, initiate the first decoding repair. Read all accessible raw data blocks and all first check blocks within the availability zone; Construct the first decoding bipartite graph using the original data block as the variable node and the first check block as the check node; Mark all accessible raw data blocks as decoded and perform iterative erasure decoding operations until all inaccessible raw data blocks are recovered or decoding stops. When decoding stops, the first decoding repair is deemed to have failed. Specifically, at least one of the following operations will be performed to write the recovered original data block to the target storage node of the availability zone: update the storage location of the affected first check block; regenerate the affected first check block.
15. The method of claim 14, wherein, The execution of the iterative erasure decoding operation includes: Repeat the following steps until all inaccessible raw data blocks are marked as decoded, or the absence of a check node with a degree of 1 causes decoding to stall: In the first decoding bipartite graph, a check node with a degree of 1 is found, and the check node with a degree of 1 is connected to only one inaccessible original data block; When a check node with a degree of 1 is found, the inaccessible original data block corresponding to the check node is recovered through linear combination operation. The newly recovered original data block is marked as decoded, and the degree of all check nodes associated with the newly recovered original data block is updated. If no verification node with a degree of 1 is found, perform the pruning or backtracking operation corresponding to the iterative erasure decoding to continue decoding.
16. The method of claim 1, wherein, For scenarios where the first decoding repair fails, or for second-level faults with a wider impact than the first-level fault, the second decoding repair is performed using the second verification block, including: Assess the scope of the access anomaly, identify the set of availability zones that are inaccessible, and all the lost original data blocks corresponding to that set; Read all raw data blocks and all accessible second check blocks within the available zones that can be accessed normally; Construct a second decoding bipartite graph using all original data blocks as variable nodes and accessible second check blocks as check nodes; Mark all accessible raw data blocks in the availability zone as decoded, and perform iterative erasure decoding operations until all lost raw data blocks are recovered or decoding stops. The recovered original data blocks are written to storage nodes in the availability zone that can be accessed normally, in order to regenerate the affected first and second check blocks.
17. The method of claim 16, wherein, The execution of the iterative erasure decoding operation includes: Repeat the following steps until all lost original data blocks are marked as decoded, or the decoding stalls due to the absence of a check node with a degree of 1: In the second decoding bipartite graph, a check node with a degree of 1 is found, and the check node with a degree of 1 is connected to only one lost original data block; When a check node with a degree of 1 is found, the lost original data block corresponding to the check node is recovered through linear combination operation. The newly recovered original data block is marked as decoded, and the degree of all check nodes associated with the newly recovered original data block is updated. If no verification node with a degree of 1 is found, perform the pruning or backtracking operation corresponding to the iterative erasure decoding to continue decoding.
18. The method of claim 1, further comprising: Collect operational status data of the multi-availability zone storage system, including network bandwidth data, access anomaly frequency data, and node load data. Based on the operational status data, adjust the number of the first verification block generated and the number of the second verification block generated; Based on the adjusted number of check blocks, a new check block is generated.
19. The method of claim 1, wherein, The first level of failure includes block-level failure or node-level failure. The block-level failure is a failure in which a single raw data block is inaccessible. The node-level failure is a failure in which a single storage node fails and all inaccessible raw data blocks are located in the same availability zone. The second level of failure is a failure in which all storage nodes in at least one availability zone are inaccessible.
20. The method of claim 1, wherein, The first level of failure is a failure whose impact is limited to a single availability zone, while the second level of failure is a failure whose impact covers all storage nodes in at least two availability zones. The first decoding repair uses only the original data blocks and the first check block that are accessible within the availability zone to which the failure belongs, while the second decoding repair uses the original data blocks and the second check block that are accessible within the availability zone.
21. A hierarchical erasure coding apparatus for a multi-availability zone storage system, the apparatus comprising: The data processing module is used to divide the raw data to be stored into multiple raw data blocks, wherein the raw data blocks will be allocated to multiple physically isolated availability zones, and each availability zone stores at least one raw data block; The first encoding module is deployed within each availability zone and is used to perform first erasure encoding on the original data blocks within that availability zone to generate the first verification block. The second encoding module is used to perform second erasure encoding on the original data blocks of all availability zones to generate a second check block; The fault management module is used to detect access anomalies in raw data blocks, storage nodes, or availability zones in a multi-availability zone storage system, determine the scope of the fault impact and the fault level corresponding to the access anomaly, and issue repair instructions. The decoding and repair module is used to perform first decoding and repair using the first verification block for first-level faults according to the repair instructions; and to perform second decoding and repair using the second verification block for scenarios where the first decoding and repair fails, or for second-level faults whose impact is greater than that of the first-level faults.
22. The apparatus of claim 21, further comprising a metadata management module, configured to: Obtain a mapping table of block storage information and dependencies for the raw data; After the first and second verification blocks are generated, the mapping table is updated based on the storage information and dependencies of the first and second verification blocks. When the storage location or dependency of the original data block, the first check block, or the second check block changes, the mapping table is updated; Maintain the mapping relationship between the original data and the original data blocks; Provides metadata query and update services for other modules.
23. The apparatus of claim 21, further comprising an adaptive adjustment module, configured to: Collect data on network bandwidth, frequency of access anomalies, and node load operation status of the multi-availability zone storage system; Based on the aforementioned operational status data, adjust the number of first and second verification blocks generated. A new verification block is generated based on the adjusted quantity.
24. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores a computer program that, when executed by the at least one processor, implements the method according to any one of claims 1-20.
25. A non-transitory computer-readable storage medium storing a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-20.
26. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-20.