File operations in a distributed storage system

CN117573638BActive Publication Date: 2026-10-09WEKA IO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311660711.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-04
Filing Date
2018-10-05
Publication Date
2026-10-09
Estimated Expiration
2038-10-05

Smart Images

  • Figure CN117573638B_ABST
    Figure CN117573638B_ABST
Patent Text Reader

Abstract

File operations in a distributed storage system are disclosed. A plurality of computing devices are communicatively coupled to one another via a network, and each of the plurality of computing devices is operatively coupled to one or more of a plurality of storage devices. A plurality of failure resilient address spaces are distributed across the plurality of storage devices such that each of the plurality of failure resilient address spaces spans the plurality of storage devices. The plurality of computing devices maintain metadata that maps each failure resilient address space to one of the plurality of computing devices. Each of the plurality of computing devices is operable to read from and write to the plurality of storage blocks while maintaining an extent in the metadata that maps the plurality of storage blocks to the failure resilient address space.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, with its parent application number 201880086360.4, filed on October 5, 2018, and entitled "File Operations in a Distributed Storage System".

[0002] Priority requirements

[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 585,057, filed November 13, 2017, entitled “File Operations in a Distributed Storage System”, and U.S. Patent Application No. 16 / 121,508, filed September 4, 2018, entitled “File Operations in a Distributed Storage System”. Background Technology

[0004] By comparing such a method with some aspects of the present method and system set forth in the remainder of this disclosure with reference to the accompanying drawings, the limitations and disadvantages of conventional methods of data storage will become apparent to those skilled in the art.

[0005] Cross-references to related applications

[0006] The full text of U.S. Patent Application No. 15 / 243,519 entitled “Distributed Erasure Coded Virtual File System” is incorporated herein by way of its entirety. Summary of the Invention

[0007] Methods and systems for file operations in distributed storage systems are provided, which are substantially illustrated and / or described in conjunction with at least one drawing and are set forth more fully in the claims. Attached Figure Description

[0008] Figure 1 Various example configurations of virtual file systems according to aspects of this disclosure are shown.

[0009] Figure 2 An example configuration of a virtual file system node is shown according to aspects of this disclosure.

[0010] Figure 3 This illustrates another representation of a virtual file system according to an example implementation of this disclosure.

[0011] Figure 4 An example of a flash registry with write balancing is shown, according to an example implementation of this disclosure.

[0012] Figure 5 This illustrates an example implementation of the method described in this disclosure. Figure 4 An example of a shadow registry associated with the flash registry.

[0013] Figure 6 This illustrates an example implementation where two of the distributed fault-tolerant address spaces reside on multiple solid-state storage disks.

[0014] Figure 7 A forward error correction scheme according to an example implementation of this disclosure is shown, which can be used to protect data stored in non-volatile memory of a virtual file system. Detailed Implementation

[0015] The system in this disclosure is suitable for small clusters and can also scale to many tens of thousands of nodes. Example embodiments regarding non-volatile memory (NVM) (e.g., flash memory in the form of solid-state drives (SSDs)) are discussed. The NVM can be divided into 4kB “blocks” and 128MB “chunks”. “Extents” can also be stored in volatile memory, such as RAM for fast access, or can be backed up by the NVM storage. The extent can store pointers to blocks, for example, 256 pointers to 1MB of data stored in a block. In other embodiments, larger or smaller memory partitions can also be used. The metadata functionality in this disclosure can be efficiently distributed across many servers. For example, in the case of a large load pointing to a “hotspot” in a specific part of the file system namespace, this load can be distributed across multiple nodes.

[0016] Figure 1 Various example configurations of a virtual file system (VFS) according to aspects of this disclosure are shown. Figure 1 The diagram shows a local area network (LAN) 102, which includes one or more VFS nodes 120 (integer indices from 1 to J, j≥1), and optionally includes (shown in dashed lines): one or more dedicated storage nodes 106 (integer indices from 1 to M, M≥1); one or more compute nodes 104 (integer indices from 1 to N, N≥1); and / or an edge router connecting LAN 102 to a remote network 118. The remote network 118 optionally includes one or more storage services 114 (integer indices from 1 to K, K≥1); and / or one or more dedicated storage nodes 115 (integer indices from 1 to L, for L≥1).

[0017] 120 per VFS node j(j is an integer, where 1 ≤ j ≤ J) is a network computing device (e.g., a server, personal computer, etc.) that includes circuitry for running the VFS process and optional client processes (or directly in device 104). n On the operating system and / or on device 104 n (In one or more virtual machines running in the system).

[0018] Compute node 104 is a networked device that can run a VFS frontend without a VFS backend. Compute node 104 can run the VFS frontend by placing the SR-IOV into the NIC and using a full processor core. Alternatively, compute node 104 can run the VFS frontend by networking via the Linux kernel networking stack and using kernel process scheduling, thus not requiring a full kernel. This is useful if the user does not want to allocate a full kernel for VFS, or if the networking hardware is incompatible with VFS requirements.

[0019] Figure 2 An example configuration of a VFS node according to aspects of this disclosure is shown. The VFS node includes a VFS frontend 202 and drive 208, a VFS memory controller 204, a VFS backend 206, and a VFS SSD agent 214. As used in this disclosure, a “VFS process” is a process that implements one or more of the following: VFS frontend 202, VFS memory controller 204, VFS backend 206, and VFS SSD agent 214. Therefore, in the example implementation, VFS node resources (e.g., processing and memory resources) can be shared between client processes and VFS processes. VFS processes can be configured to require relatively few resources to minimize the impact on client application performance. VFS frontend 202, VFS memory controller 204, and / or VFS backend 206 and / or VFS SSD agent 214 can run on the processor of host 201 or on the processor of network adapter 218. For multi-core processors, different VFS processes can run on different cores and can run different subsets of services. From the perspective of client process 212, the interface with the virtual file system is independent of the specific physical machine running the VFS process. The client process only needs the existence of drive 208 and front end 202 to provide services to them.

[0020] A VFS node can be implemented as a single-tenant server running directly on an operating system (e.g., bare metal), or as a virtual machine (VM) and / or container (e.g., Linux Containers (LXC)) within a bare metal server. A VFS can run as a VM environment within an LXC container. Therefore, within a VM, only an LXC container containing the VFS can run. In a classic bare metal environment, user-space applications exist, and the VFS runs within an LXC container. If the server is running other containerized applications, the VFS may run within an LXC container outside the management scope of a container deployment environment (e.g., Docker).

[0021] A VFS node can be serviced by an operating system and / or a Virtual Machine Monitor (VMM) (e.g., a hypervisor). The VMM can be used to create and run VFS nodes on host 201. Multiple kernels can reside within a single LXC container running VFS, and VFS can run on a single host 201 using a single Linux kernel. Therefore, a single host 201 can include multiple VFS front-ends 202, multiple VFS memory controllers 204, multiple VFS back-ends 206, and / or one or more VFS drives 208. VFS drives 208 can run in kernel space outside the scope of the LXC container.

[0022] A single root input / output virtualization (SR-IOV) PCIe virtualization function can be used to run a network stack 210 in user space 222. SR-IOV allows for the isolation of PCI Express, enabling the sharing of a single physical PCI Express across virtual environments and providing different virtual functions to different virtual components on a single physical server computer. I / O stack 210 enables VFS nodes to bypass the standard TCP / IP stack 220 and communicate directly with network adapter 218. A portable operating system interface for unix (POSIX) VFS functions can be provided to VFS driver 208 via lock-free queuing. SR-IOV or the full PCIe physical function address can also be used to run a non-volatile memory Express (NVMe) driver 214 in user space 222, completely bypassing the Linux I / O stack. NVMe can be used to access non-volatile storage media 216 connected via the PCI Express (PCIe) bus. The non-volatile storage medium 216 may be, for example, flash memory in the form of a solid-state drive (SSD) or a storage class memory (SCM) in the form of a solid-state drive (SSD) or a memory module (DIMM). Other examples may include storage class memory technologies such as 3D-XPoint.

[0023] By coupling the physical SSD 216 with the SSD agent 214 and networking 210, the SSD can be implemented as a networked device. Alternatively, the SSD can be implemented as an NVMe SSD 242 or 244 attached to a network using a networking protocol such as NVMe-oF (structure-based NVMe). NVMe-oF allows access to NVMe devices using redundant network links, providing a higher level of resilience. Network adapters 226, 228, 230, and 232 can include hardware acceleration for connecting to the NVMe SSDs 242 and 244, transforming them into networked NVMe-oF devices without the need for a server. The NVMe SSDs 242 and 244 can each include two physical ports, and all data can be accessed through either of these ports.

[0024] Each client process / application 212 may run directly on the operating system, or it may run in a virtual machine and / or container served by the operating system and / or hypervisor. Client process 212 may read data from and / or write data to memory during the execution of its primary functions. However, the primary function of client process 212 is storage-independent (i.e., the process is only concerned with the reliable storage of its data and its retrieval when needed, regardless of the location, timing, or manner of data storage). Example applications that give rise to such a process include: email servers, web servers, office productivity applications, customer relationship management (CRM), animated video rendering, genomic computing, chip design, software building, and enterprise resource planning (ERP).

[0025] Client application 212 can make system calls to kernel 224, which communicates with VFS driver 208. VFS driver 208 places the corresponding requests on the queue of VFS frontend 202. If several VFS frontends exist, the driver may load balance access to different frontends to ensure that a single file / directory is always accessed through the same frontend. This can be done by “sharding” the frontends based on file or directory IDs. VFS frontend 202 provides an interface for routing filesystem requests to the appropriate VFS backend based on the bucket responsible for the operation. The appropriate VFS backend can be on the same host or on another host.

[0026] The VFS backend 206 hosts multiple buckets, each of which serves the file system requests it receives and performs tasks to additionally manage the virtual file system (e.g., load balancing, logging, maintaining metadata, caching, moving data between layers, deleting outdated data, correcting corrupted data, etc.).

[0027] VFS SSD agent 214 handles interactions with the corresponding storage device 216. This may include, for example, address translation and generating commands to be issued to the storage device (e.g., on SATA, SAS, PCIe, or other suitable buses). Therefore, VFS SSD agent 214 acts as an intermediary between the VFS backend 206 of the virtual file system and the storage device 216. SSD agent 214 can also communicate with standard network storage devices that support standard protocols, such as NVMe-oF (structure-based NVMe).

[0028] Figure 3 This illustrates another representation of a virtual file system according to an example implementation of this disclosure. Figure 3 In this context, element 302 represents the memory resources (e.g., DRAM and / or other short-term memory) and processing resources (e.g., x86 processor, ARM processor, NIC, ASIC) of the various nodes (compute, storage, and / or VFS) where the virtual file system resides, as described above. Figure 2 As described. Element 308 represents one or more physical storage devices 216 that provide long-term storage for a virtual file system.

[0029] like Figure 3 As shown, the physical storage is organized into multiple Distributed Failure Resilience Address Spaces (DFRAS) 518. Each DFRAS comprises multiple blocks 310, and each block 310 comprises multiple blocks 312. Organizing blocks 312 into blocks 310 is only a convenience in some implementations and may not be done in all implementations. Each block 312 stores committed data 316 (which may take various states discussed below) and / or metadata 314 describing or referencing the committed data 316.

[0030] Organizing storage device 308 into multiple DFRAS enables high-performance parallel commits from many (potentially all) nodes of the virtual file system (e.g., Figure 1 All nodes 1041 to 104 N 1061 to 106 M and 1201 to 120 J (Concurrent commits can be executed in parallel). In one example implementation, each node of the virtual file system can own one or more corresponding DFRASs and have exclusive read / commit access to the DFRAS it owns.

[0031] Each bucket has its own DFRAS, and therefore does not require coordination with any other nodes when writing to it. Each bucket may create stripes on many different blocks across many different SSDs, so each bucket and its DFRAS can select the "block stripe" to write to based on many parameters, and once the block is assigned to that bucket, writing can proceed without coordination. All buckets can effectively write to all SSDs without coordination.

[0032] Each DFRAS is owned and accessed only by its owner bucket running on a specific node, allowing each node in the VFS to control a portion of storage device 308 without having to coordinate with any other nodes (except during initialization or after a node failure, during the (re)allocation of buckets holding DFRAS, which can be asynchronous with the actual read / commit operations of storage device 308). Therefore, in such an implementation, each node can read / commit DFRAS to its bucket independently of what other nodes are doing, without needing any consensus when reading and committing to storage device 308. Furthermore, the fact that a particular node owns multiple buckets in the event of a failure allows for a smarter and more efficient redistribution of its workload to other nodes (rather than distributing the entire workload to a single node, which could create "hotspots"). In this respect, in some implementations, the number of buckets may be large relative to the number of nodes in the system, making any single bucket potentially a relatively small load placed on another node. This allows for fine-grained redistribution of the load on the failed node based on the capabilities and capacity of other nodes (e.g., a node with more capabilities and capacity may receive a higher percentage of the failed node's bucket).

[0033] To allow this operation, metadata can be maintained that maps each bucket to the node it currently owns, making it possible to redirect reads and commits to storage device 308 to the appropriate node.

[0034] Load balancing is possible because the entire file system metadata space (e.g., directories, file attributes, ranges of content within files, etc.) can be broken down (e.g., shredded or fragmented) into uniform pieces (e.g., “fragments”). For example, a large system with 30k servers can divide its metadata space into 128k or 256k fragments.

[0035] Each such metadata fragment can be stored in a "bucket". Each VFS node may be responsible for several buckets. When a bucket is serving a metadata fragment on a given backend, that bucket is considered the "active" or "leader" of that bucket. Typically, there are more buckets than VFS nodes. For example, a small system with 6 nodes may have 120 buckets, while a large system with 1000 nodes may have 8k buckets.

[0036] Each bucket can operate on a small group of nodes, typically a 5-tuple consisting of 5 nodes. Cluster configuration ensures that all participating nodes keep up-to-date with the 5-tuple allocation for each bucket.

[0037] Each quintuplet monitors itself. For example, if there are 10,000 servers in the cluster, and each server has 6 buckets, then each server will only need to communicate with 30 different servers to maintain the state of its bucket (6 buckets will have 6 quintuplets, so 6 x 5 = 30). This is far less than a centralized entity having to monitor all nodes and maintain cluster-wide state. Using quintuplets allows performance to scale with larger clusters because nodes don't perform more work as the cluster size increases. This can have a drawback: in "dumb" mode, a small cluster can actually generate more communication than the physical nodes, but this can be overcome by having them share all buckets and sending only a single heartbeat between two servers (this changes to only one bucket as the cluster grows, but if you have a small cluster of 5 servers, it will only include all buckets in all messages, and each server will only communicate with the other 4 servers). Quintuplets can use algorithms similar to the Raft consensus algorithm to make decisions (i.e., reach consensus).

[0038] Each bucket may have a set of compute nodes that can run it. For example, five VFS nodes can run one bucket. However, at any given time, only one node in the group is the controller / leader. Furthermore, for sufficiently large clusters, no two buckets share the same group. If there are only 5 or 6 nodes in the cluster, most buckets can share a backend. In a fairly large cluster, there may be many different groups of nodes. For example, in the case of 26 nodes, there are over 64,000 There are 5 possible five-node pairs (i.e., quintuples).

[0039] All nodes in the group know and agree (i.e., reach consensus) which node is the actual activity controller (i.e., leader) for that bucket. A node accessing a bucket may remember ("caches") the last node among the group's (e.g., five) members that was the leader of that bucket. If it accesses the bucket leader, the bucket leader performs the requested action. If the node it accesses is not the current leader, the node instructs the leader to "redirect" the access. If accessing the cached leader node times out, the contacting node can try using another node from the same five-tuple. All nodes in the cluster share a common cluster "configuration," which allows nodes to know which server can run each bucket.

[0040] Each bucket can have a load / utilization value, which indicates how much the application running on the file system is using the bucket. For example, even if the number of buckets used is unbalanced, a server node with 11 buckets with low utilization can receive another metadata bucket to run before a server with 9 buckets with high utilization. The load value can be determined based on average response latency, the number of concurrent operations, memory consumption, or other metrics.

[0041] Even if a VFS node is not faulty, redistribution can still occur. If the system identifies a node as busier than others based on tracked load metrics, it can move one of its buckets (i.e., "failover") to a less busy server. However, load balancing can be achieved by shifting writes and reads before the bucket is actually relocated to a different host. Since each write may end on a different group of nodes determined by DFRAS, a node with a higher load may not be selected for the stripe where data is written. The system may also choose not to provide reads from high-load nodes. For example, a "degraded-mode read" can be performed, where blocks from a high-load node are reconstructed from blocks in the same stripe. Degraded-mode reads are reads performed by the remaining nodes in the same stripe and the data is reconstructed through failover. Degraded-mode reads may be performed when read latency is too high, as the initiator of the read may assume the node is down. If the load is high enough to create even higher read latency, the cluster can recover by reading the data from other nodes and using degraded-mode reads to reconstruct the required data.

[0042] Each bucket manages its own Distributed Erasure Coding instance (i.e., DFRAS 518) and does not need to cooperate with other buckets to perform read or write operations. There can be thousands of concurrent Distributed Erasure Coding instances working simultaneously, each for a different bucket. This is an integral part of scalability performance because it effectively divides any large file system into independent parts that do not require coordination, thus providing high performance regardless of the scale of the expansion.

[0043] Each bucket handles all file system operations belonging to its fragments. For example, directory structure, file attributes, and file data ranges will fall under the jurisdiction of a specific bucket.

[0044] An operation performed by any frontend begins by determining which bucket owns the operation. Then, the backend leader and nodes for that bucket are determined. This determination can be performed by trying the most recently known leader. If the most recently known leader is not the current leader, the node may know which node is the current leader. If the most recently known leader is no longer part of the bucket's 5-tuple, the backend will inform the frontend that it should go back to the configuration to find the bucket's 5-tuple members. This distribution of operations allows complex operations to be handled by multiple servers instead of a single computer in a standard system.

[0045] Each 4KB block in the file is handled by an object called an extension. The extension ID is based on the inode ID and an offset within the file (e.g., 1MB) and is managed by the bucket. Extensions are stored in the registry of each bucket. Extension ranges are protected by read-write locks. Each extension manages its 1MB content through pointers that manage up to 256 4KB blocks responsible for a contiguous 1MB of data in the file.

[0046] Write operation

[0047] At box 401, the extended region in the VFS backend (“BE”) receives write requests. The VFS frontend (“FE”) calculates the bucket responsible for the 4k write (or aggregate write) of the extended region. The FE then sends the block over the network to the relevant bucket, which begins processing the write. Figure 3The DFRAS request in the code will be used to write the block IDs of these blocks. If the write contains enough 4k blocks to span several consecutive extensions, the first extension receives the entire write request. At box 403, each extension that is part of the write locks its associated 4k blocks, for example, in order from the first block of the first extension to the last block of the last extension. Locking is done on a per-4k-block basis, so even if new data is written for a portion of a 4k block, the entire 4k region will be blocked for consistency purposes. Write locking can be very “expensive” and is partly why other systems can only perform a limited number of write operations when they have only a few metadata servers. Since the current public content allows an unlimited number of BEs, an unlimited number of 4k write operations can be performed simultaneously.

[0048] If the requested data does not fill a complete 4k block, each touched but not fully rewritten 4k block is read first. New data is overwritten on that block, and then the complete 4k block can be written to DFRAS. For example, a user can write 5k data at an offset of 3584. The first 512 bytes are written to the first 4k block (as its last 512 bytes), the second 4k block is fully rewritten, and the third 4k block has its first 512 bytes rewritten. The first and third 4k blocks are read first, new data will overwrite its correct portion, and then the three complete 4k blocks are actually written to DFRAS.

[0049] Each block has its EEDP (End-to-End Data Protection) calculated before being written to DFRAS, and the EEDP information is also provided to DFRAS, thus ensuring that each written block reaches the correct, unchanged final NVM target. If a bit flip or logical error occurs on the network, the end node will receive a 4k block of data with a different EEDP, and DFRAS will know that it needs to retransmit the data.

[0050] In box 405, once the extended region has confirmed the lock, writing can begin. Write locking allows only a single write to each portion of the extended region at a time. Two concurrent writes may occur in different ranges of the extended region. If the extended region is on another bucket, remote procedure calls (RPCs) can be used to make the writes perform in the address space as if they were local writes, for example, without explicitly coding the details of the remote interaction.

[0051] In box 407, the log is saved to DFRAS. If possible, the first block written can be compressed to fit the log. The log (either its own block or within the first block) will be updated using the returned block ID. The first few bytes of the written block can actually be replaced by the return pointer to the extension area, after which (with EEDP) the system can ensure that the data read is complete.

[0052] In box 409, if the data to be written is not 4k aligned, or its size is not a multiple of 4k, the data can be padded to a multiple of 4k at either (or both) of the buffer. If data is padded, existing data is read from the first and / or last block (or from the same block if the write is less than 4k), and the requested write data is padded with the read data, thus allowing a complete 4k block to be written. In this way, small files can be merged into the inode and extension area.

[0053] In box 411, after DFRAS is written, the extended region is updated using the block IDs, EEDPs, and the first few bytes of data for all blocks. The extended region can be stored in the registry and downgraded to DFRAS on the next registry downgrade. In box 413, once all writes are complete, the extended region range is unlocked in reverse order of locking. The data can then be read once the registry is updated.

[0054] Additional operations

[0055] A special form of writing is appending. With appending, the system obtains a binary long object (blob) containing data that may be appended to the end of the current data. Doing this from a host with a single writer is straightforward—simply find the last offset and perform the write operation. However, append operations may occur simultaneously from multiple Front Ends (FEs), and the system must ensure that each append is consistent internally. For example, two large append operations interleaving data must be prevented.

[0056] The inode knows which extension is the last extension of the file, and the extension also knows it is the last extension (there may only be one). Only the last extension will accept append operations. Any other extensions will be rejected.

[0057] In each append operation, the FE searches through the inodes to find the last extended region. If it encounters an extended region that refuses to perform the operation, the FE will search through the inodes again to find the last extended region.

[0058] When an extension performs an append operation, it executes each append in the order it was received. Therefore, when each append begins, the extension finds the last data block and performs the write (after locking that block). If the append ends after the last extension, the extension creates a new extension in a special mode that may only perform one append before being "open for use." The extension can then forward the remaining data fragments to this extension. This extension becomes the last extension and is marked as open for use. Extensions that are no longer the last cannot perform all remaining append operations. Therefore, the remaining append operations return to the FE to find out from the inode which extension is the last and continue normally. If the write is large enough to span multiple extensions, the last extension will create all the extensions needed in that mode and forward the correct data to the created extensions. Once all data has been written to all extensions, the original extension marks the last appended extension as the last and marks all extensions as open for use. Then, the original extension region will search for all pending append operations that have failed and are sent to the original FE to restart the operation and receive the correct extension region.

[0059] The semantics of the operation are unaffected by the order of concurrent appends. However, it must be ensured that no data is interleaved between different writes. If the FE or inode finds the last offset and performs a standard write, a race condition will occur, and the file will end with inconsistent data related to the append operation semantics.

[0060] Read operation

[0061] In box 501, the extension region in the VFS frontend receives read requests. It calculates and determines which bucket is responsible for holding the extension region for the read, and then identifies the backend (node) currently running that bucket. It then sends the read request to that bucket via RPC. If the read operation spans more than one extension region (exceeding a contiguous 1MB), the VFS FE can split it into several distinct read operations, which are then aggregated back by the VFS FE. Each such read operation can request a read from a single extension region. The VFS FE then sends the read request to the node responsible for holding that extension region.

[0062] In box 503, the VFS BE holding the bucket looks up the extended region key in the registry. If the extended region does not exist, a zero buffer is returned in box 505. In box 507, if the extended region does exist, 4k blocks in the extended region are locked, and data is read from DFRAS. In box 509, the registry overwrites the stored data with the corresponding copy from the VM (if present) and returns the actual content pointed to by the extended region. The copy of the extended region includes data changes that have not yet been committed to non-volatile storage (e.g., flash SSDs). If the requested data is not 4k aligned or is not 4k in size, the data is retrieved from the persistent distributed encoding scheme, and then both ends can be "chopped" to meet the requested (offset, size).

[0063] In box 511, the return pointer (pointing to the block ID of the extension area) and EEDP (End-to-End Data Protection) of each received block are checked. If either check fails, the extension area requests a degraded mode from the distributed erasure coding block from the bucket, which is read in box 513, where other blocks of the stripe are read to reconstruct the requested block. If the extension area exists, a list of all stripe blocks stored in the persistent layer is provided. In box 515, the return pointer (pointing to the block ID of the extension area) and EEDP (End-to-End Data Protection) of the reconstructed block are checked. In box 517, if either check fails, an IO error is returned. Once the correct data is confirmed in box 511, it returns to the requested FE after replacing the first few bytes with the original block data stored in the extension area and overwriting the return pointer; a new write to the Distributed Erasure Coding System (DECS) is initiated; and the extension area with the new block ID is also updated.

[0064] Figure 6 This illustrates an example implementation where two of the distributed fault-tolerant address spaces reside on multiple solid-state storage disks. (Tie 510) 1,1 Up to 510 D,C It can be organized into multiple chunk stripes 5201 to 520 S (S is an integer). In the example implementation, forward error correction (e.g., erasure coding) is used to protect each chunk stripe 520 separately. s (s is an integer, where 1 ≤ s ≤ S). Therefore, any particular chunk stripe 520 can be determined based on the desired data protection level. s Block 510 in d,c The quantity.

[0065] For illustrative purposes, assume each block strip is 520. s Includes N = M + K (where each of N, M, and K is an integer) blocks 510 d,c Then there are N blocks 510 d,cThe M bytes can store data codes (typically binary numbers or "bits" for current storage devices) and N blocks 510 d,c K bits in the array can store protection codes (again, typically bits). Then, the virtual file system can allocate 520 bits to each stripe. s Assign N blocks from N different fault domains 510 d,c .

[0066] As used herein, a "fault domain" refers to a set of components, where a failure in any one component (a component losing power, becoming unresponsive, etc.) could cause all components to fail. For example, if a rack has a single top-of-rack switch, a failure of that switch would disconnect all components on that rack (e.g., compute, storage, and / or VFS nodes). Therefore, for the rest of the system, this is equivalent to all components on that rack failing together. The virtual file system according to this disclosure may include fewer fault domains than block 510.

[0067] In the example implementation where each virtual file system node is connected and powered to a single storage device 506 in a fully redundant manner for each such node, the fault domain can be limited to that single storage device 506. Therefore, in the example implementation, each block stripe 520 s Including storage devices 5061 to 506 D Multiple blocks on each of the N ones 510 d,c (Therefore, D is greater than or equal to N). An example of this implementation is... Figure 7 It is displayed in the middle.

[0068] exist Figure 6 In this configuration, D=7, N=5, M=4, K=1, and the storage devices are organized into two DFRAS. These numbers are for illustrative purposes only and are not intended as limitations. The three block stripes 520 of the first DFRAS are arbitrarily selected for illustrative purposes. The first block stripe 5201 consists of block 510. 1,1 510 2,2 510 3,3 510 4,5 and 510 5,6 Composition; the second block strip 5202 is composed of block 510 3,2 510 4,3 510 5,3 510 6,2 and 510 7,3 Composition; the third block strip 5203 is composed of block 510 1,4 510 2,4 510 3,5 510 5,7 and 510 7,5 composition.

[0069] Although in the actual implementation Figure 6 In the simple example, D=7 and N=5, but D can be much larger than N (e.g., an integer multiple of 1, and possibly several orders of magnitude higher), and two values ​​can be chosen so that the probability of any two block stripes 520 of a single DFRAS residing on the same group of N storage devices 506 (or more generally, on the same group of N fault domains) is below a desired threshold. In this way, any single storage device 506 d (Or more generally, a failure in any single fault domain) will result in (expected statistics may be determined based on: selected values ​​of D and N, the size of N storage devices 506, and the arrangement of the fault domain) any particular stripe 520 s At most one block 510 b,c The loss. Furthermore, a double failure will cause the vast majority of stripes to lose at most a single block of 510. b,c And only a small number of bands (determined based on the values ​​of D and N) will be drawn from any given band 520. s Two segments are lost (for example, the number of stripes in two failures may be exponentially reduced compared to the number of stripes in one failure).

[0070] For example, if each storage device 506d is 1TB and each block is 128MB, then storage device 506 d The failure will result in (expected statistics may be determined based on the following: selected values ​​of D and N, the size of N storage devices 506, and the arrangement of the failure domain) 7812 (=1TB / 128MB) block stripes 520 losing one block 510. For each such affected block stripe 520 s Appropriate forward error correction algorithms and chunking stripes 520 can be used. s The other N-1 chunks are used to quickly reconstruct the lost chunks. d,c Furthermore, because the 7812 affected block stripes are evenly distributed across all storage devices 5061 to 506... D Therefore, the 7812 missing blocks were reconstructed. d,c The statistics to be involved (the desired statistics may be determined based on the following: selected values ​​of D and N, the size of N storage devices 506, and the layout of the fault domain) will be from each storage device 5061 to 506. D Reading the same amount of data (i.e., the burden of reconstructing lost data is evenly distributed across all storage devices 5061 to 506) D In order to recover very quickly from a failure.

[0071] Next, we turn to the two storage devices 5061 to 506. DIn the case of concurrent failures (or more generally, concurrent failures of two fault domains), due to the block stripes 5201 to 520 of each DFRAS... S In all storage devices 5061 to 506 D The surface is uniformly distributed, with only a very small number of block bands from 5201 to 520. S Two out of N chunks will be lost. The virtual file system is operable to quickly identify such twice-lost chunk stripes based on metadata indicating chunk stripes 5201 to 520. S With storage devices 5061 to 506 D The mapping between them. Once such two-times-lost-block stripes are identified, the virtual file system can prioritize rebuilding those stripes before initiating a single-time-lost-block striping reconstruction. The remaining block stripes will have only a single lost block, and for them (the vast majority of the affected block stripes), the two storage devices 506 d Concurrent failures with only one storage device 506 d The faults are the same. A similar principle applies to three concurrent faults (in a two-fault scenario, the number of chunk stripes with three fault blocks will be far less than the number with two fault blocks), and so on. In the example implementation, it can be based on chunk stripe 520. s The number of missing items in the control block stripe 520 is used to control the execution of the stripe. s The rate of reconstruction. This can be achieved, for example, by controlling the rate at which reads and commits are performed for reconstruction, the rate at which FEC calculations are performed for reconstruction, the rate at which network messages for reconstruction are transmitted, etc.

[0072] Figure 7 A forward error correction scheme according to an example implementation of this disclosure is shown, which can be used to protect data stored in non-volatile memory of a virtual file system. Storage block 902 of block stripes 5301 to 5304 of DFRAS is shown. 1,1 Up to 902 7,7 .exist Figure 7 In the protection scheme, five blocks in each strip are used to store data codes, and two blocks in each strip are used to store protection codes (i.e., M=5 and K=2). Figure 7 In the middle, the protection code is calculated using the following formulas (1) to (9):

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082] therefore, Figure 7 The four stripes 5301 to 5304 are part of a multi-strip (in this case, four stripes) FEC protection domain, and the loss of any two or fewer blocks in any block stripe 5301 to 5304 is recovered by using various combinations of equations (1) to (9) above. For comparison, an example of a single-strip protection domain is: if only P1 protects D11, D22, D33, D44, D54, and writes D11, D22, D33, D44, D54, and P1 all into stripe 5301 (5301 is a single-strip FEC protection domain).

[0083] According to an example implementation of this disclosure, multiple computing devices are communicatively coupled to each other via a network, and each of the multiple computing devices includes one or more of a plurality of storage devices. Multiple failover address spaces are distributed across the multiple storage devices, such that each of the multiple failover address spaces spans multiple storage devices. Each of the multiple failover address spaces is organized into multiple stripes (e.g., as shown in the example implementation). Figure 6 and Figure 7 The multiple 530s shown). Each or more of the multiple stripes are multiple forward error correction (FEC) protection domains (e.g., such as...). Figure 7 A portion of a corresponding domain in a multi-strip FEC domain. Each of the multiple stripes may include multiple storage blocks (e.g., multiple 512). Each block of a particular stripe in the multiple stripes may reside on different storage devices in multiple storage devices. The first portion of the multiple storage blocks (e.g., by...) Figure 7 The 5301 band of 902 1,2 Up to 902 1,6 The five components) can be used to store data codes, while the second part of multiple storage blocks (e.g., Figure 7 The two 902 stripes of 5301 1,1 and 902 1,7The number of data stripes can be used to store protection codes calculated at least in part based on data codes. Multiple computing devices can be operable to sort multiple stripes. This sorting can be used to select which of the multiple stripes is used for the next commit operation to one of the multiple failover address spaces. The sorting can be based on the number of protected and / or unprotected storage blocks in each of the multiple stripes. For any of the multiple stripes, the sorting can be based on a bitmap stored on multiple storage devices having a particular one of the multiple stripes. The sorting can be based on the number of blocks currently storing data in each of the multiple stripes. The sorting can be based on the read and write overhead for committing to each of the multiple stripes. Each failover address space can be owned by only one of the multiple computing devices at any given time, and each of the multiple failover address spaces can only be read and written by its owner. Each computing device can own multiple failover address spaces. Multiple storage devices can be organized into multiple failover domains. Each of the multiple stripes can span multiple failover domains. Each fault resilience address space can span all multiple fault domains, such that when any particular fault domain fails, the workload for reconstructing lost data is distributed among each of the other fault domains in the multiple fault domains. Multiple stripes can be distributed across multiple fault domains such that, in the event that two fault domains in the multiple fault domains fail simultaneously, the probability of two blocks appearing in any particular stripe of multiple stripes across the multiple fault domains is less than the probability of only one block appearing in any particular stripe of multiple stripes across the multiple fault domains. Multiple computing devices can be used to first reconstruct any one of the multiple stripes with two fault blocks, and then reconstruct any one of the multiple stripes with only one fault block. Multiple computing devices can be used to perform the reconstruction of multiple stripes with two fault blocks at a higher rate than the reconstruction rate of multiple stripes with only one fault block (e.g., a larger percentage of CPU clock cycles dedicated to reconstruction, a larger percentage of network transfer opportunities dedicated to reconstruction, etc.). Multiple computing devices can be used to determine the rate at which to reconstruct any particular lost block based on the number of other blocks in the same stripe across the multiple stripes that are lost, in the event that one or more fault domains fail. One or more of the multiple fault domains may include multiple storage devices. Each of the multiple FEC protection domains may span multiple stripes. The multiple stripes may be organized into multiple groups (e.g., Figure 6 Block strips 5201 to 520 in SThis involves multiple groups, each comprising one or more stripes, and multiple computing devices operable to sort one or more stripes within each group. The multiple computing devices are operable to perform successive commit operations on a selected group of groups until one or more stripes in the group no longer meet a defined criterion, and if one of the selected groups no longer meets the defined criterion, another group is selected. The criterion may be based on the number of blocks available for new data writing.

[0084] Although this method and / or system has been described with reference to certain implementations, those skilled in the art will understand that various changes and substitutions can be made without departing from the scope of this method and / or system. Furthermore, many modifications can be made to adapt particular situations or materials to the teachings of this invention without departing from the scope of the invention. Therefore, it is intended that this method and / or system be limited to the specific implementations disclosed, but rather that this method and / or system will include all implementations falling within the scope of the appended claims.

[0085] As used herein, the terms “circuit” and “circuit system” refer to physical electronic components (i.e., hardware) and any software and / or firmware (“code”) that can configure, be executed by, and / or be associated with the hardware. As used herein, for example, a particular processor and memory may include a first “circuit” when executing one or more lines of code, and may include a second “circuit” when executing a second or more lines of code. As used herein, “and / or” refers to any one or more items in a list connected by “and / or”. As an example, “x and / or y” represents any element in the three-element set {(x), (y), (x, y)}. In other words, “x and / or y” means “one or both of x and y”. As another example, “x, y and / or z” represents the seven-element set {(x), (y), (z), (x, y), (x, z), (y, z), (x, y, z)}. In other words, “x, y and / or z” means “one or more of x, y, and z”. As used herein, the term "exemplary" means used as a non-limiting example, instance, or illustration. As used herein, the terms "for example" and "e.g." introduce a list of one or more non-limiting examples, instances, or illustrations. As utilized herein, a circuit is "operable" to perform a function as long as it contains the necessary hardware and code (if necessary) required to perform that function, regardless of whether the function's performance is disabled (e.g., by user-configurable settings, factory settings, etc.).

Claims

1. A system comprising: Multiple computing devices, wherein the multiple computing devices are communicatively coupled to each other via a network, The plurality of computing devices are operable to maintain metadata that maps the plurality of storage blocks to a fault-tolerant address space. The metadata is divided into multiple storage buckets. Each of the plurality of storage buckets is associated with a unique group of two or more computing devices selected from the plurality of computing devices. The unique group is capable of operating to maintain the state of the storage bucket using a consensus algorithm. The number of buckets in the plurality of buckets is determined based on the number of different blocks into which the metadata is divided. The number of different blocks into which the metadata is divided is greater than the number of computing devices among the plurality of computing devices. Each bucket manages its own distributed erase-encoded instance and does not need to cooperate with other buckets to perform read or write operations. Each of the plurality of computing devices is capable of reading data from the plurality of storage blocks and writing the data to the plurality of storage blocks, and Based on the blocks identified in the extended region, distributed erasure codes are used to check for errors in data read from specific storage blocks among the plurality of storage blocks.

2. The system according to claim 1, wherein, Multiple storage devices include non-volatile memory.

3. The system according to claim 1, wherein, The buffer of data to be written to the first of the plurality of storage blocks is compressed, thereby providing space for the log associated with the data to be written to the plurality of storage blocks.

4. The system according to claim 1, wherein, The buffer for data to be written to one of the plurality of storage blocks is filled with data that was previously written to a portion of the storage block.

5. The system according to claim 1, wherein, In coordination with writing data to one of the plurality of storage blocks, an extension area is updated, wherein the updated extension area includes identifiers of the plurality of blocks associated with protecting the data written to the storage block.

6. The system according to claim 1, wherein, When an error is detected in a specific storage block among the plurality of storage blocks, a degraded data read is performed.

7. The system according to claim 6, wherein, The degraded data read includes regenerating the specific storage block from one or more storage blocks other than the specific storage block among the plurality of storage blocks.

8. A method for accessing a storage medium, the method comprising: Metadata is maintained across multiple computing devices to map multiple storage blocks to a fault-tolerant address space, where... The metadata is divided into multiple storage buckets. Each of the plurality of buckets is associated with a unique group selected from two or more computing devices, the unique group being operable to maintain the state of the bucket using a consensus algorithm. The number of buckets in the plurality of buckets is determined based on the number of different blocks into which the metadata is divided, and The metadata is divided into a greater number of different blocks than the number of computing devices in the plurality of computing devices. Each bucket manages its own distributed erasure coding instance and does not need to cooperate with other buckets to perform read or write operations. Read data from the plurality of storage blocks, and Based on the blocks identified in the extended region, distributed erasure codes are used to check for errors in data read from specific storage blocks among the plurality of storage blocks.

9. The method according to claim 8, wherein, The fault resilience address space includes non-volatile memory.

10. The method according to claim 8, wherein, The method includes writing data into the plurality of storage blocks.

11. The method according to claim 10, wherein, Writing data includes: compressing a buffer of data to be written to a first storage block of the plurality of storage blocks, thereby providing space for a log associated with the data to be written to the plurality of storage blocks.

12. The method according to claim 10, wherein, Writing data includes filling a buffer of data to be written to the plurality of storage blocks with data previously written to the storage blocks.

13. The method according to claim 10, wherein, In coordination with writing data to one of the plurality of storage blocks, the method includes updating an extension area such that the extension area includes identifiers of a plurality of blocks associated with protecting the data written to the storage block.

14. The method according to claim 8, wherein, The method includes: when an error is found in a specific storage block among the plurality of storage blocks, performing a degraded data read.

Citation Information

Patent Citations

  • Mechanism to include hints within compressed data

    US20050144387A1

  • Failure Mapping in a Storage Array

    US20160041878A1

  • Distributed Erasure Coded Virtual File System

    US20170052847A1