Distributed storage data read-write method and device, equipment, medium and program product

CN119806398BActive Publication Date: 2026-09-29CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411781365.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2026-09-29
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

[0006]有鉴于此,本发明提供了一种分布式存储数据读写方法、装置、设备、介质及程序产品,以解决现有技术中软件栈复杂,计算资源占用较高的问题

Benefits of technology

[0008]本发明没有各种分区的复杂概念和多个分区的DHT计算,在视图中抽象出保护组的概念,提前做好映射,通过查表取代运算,在超高性能场景下计算资源占用更少,软件栈时延更低。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806398B_ABST
    Figure CN119806398B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of storage, and discloses a distributed storage data read-write method, device, equipment, medium and program product, the method comprising the following steps: acquiring distributed cluster information; generating a data block distribution view corresponding to each storage pool according to the distributed cluster information, wherein the data block distribution view comprises a plurality of protection groups, each protection group comprises corresponding disk information; querying the data block distribution view according to received service data, determining the protection group of each data block distribution, and performing a data read operation or a write operation. The application does not have the complex concept of various partitions and DHT calculation of multiple partitions, the concept of the protection group is abstracted in the view, mapping is completed in advance, and table lookup is used to replace operation, so that the occupation of computing resources is smaller and the software stack time delay is lower in a super-high-performance scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage technology, specifically to a distributed storage data read / write method, apparatus, device, medium, and program product. Background Technology

[0002] Distributed storage is a data storage technology that distributes data across multiple nodes, each of which can be a computer, server, or other storage device. Data is typically divided into multiple blocks or objects and distributed across different storage nodes, each capable of independently storing and accessing the data.

[0003] Distributed storage technology is now widely used in traditional data centers, public cloud data centers, private cloud data centers, and hyperconverged workstations. Among them, the distributed storage engine provides high-performance, highly reliable, and efficient data access capabilities for distributed storage, and is the core component of distributed storage.

[0004] Currently, there are various implementation schemes for distributed storage engines. For example, Google publicly discloses its distributed storage engine, GFS (Google File System), which provides an append-only file system. GFS uses a unified master to manage the metadata mapping between files and spaces. Chunk information is queried by the master from the specific chunk server at startup. The master manages the chunk information of the entire cluster. As the cluster grows, managing the chunk information of the entire cluster not only consumes master resources but also results in slow response times.

[0005] Another widely used storage solution is Ceph. Ceph is a distributed storage system that uses a distributed algorithm (CRUSH algorithm) to distribute data, distributing data objects according to the weight of each storage device, making the distribution approximately uniform. Ceph calculates a list of OSDs (Object Storage Devices) for reading and writing. Because these OSD lists are calculated and Ceph writes new data by overwriting existing data, PG locks and object locks are used to ensure data consistency and prevent out-of-order writes and inconsistencies, resulting in a complex software stack. When a cluster fails, the PG must make a consistency decision to determine a new OSD master within the PG, which then provides read and write services. In some unstable network scenarios, prolonged suspended I / O can lead to service interruptions. Furthermore, Ceph's in-place overwrite logic is not suitable for scenarios where SSDs are widely used, preventing the full performance of SSDs from being utilized. Summary of the Invention

[0006] In view of this, the present invention provides a distributed storage data reading and writing method, apparatus, device, medium and program product to solve the problems of complex software stacks and high computing resource consumption in the prior art.

[0007] In a first aspect, the present invention provides a distributed storage data read / write method, the method comprising: Obtain distributed cluster information; Based on the distributed cluster information, a data block distribution view is generated for each storage pool. The data block distribution view includes multiple protection groups, and each protection group includes the corresponding disk information. Based on the received business data, query the data block distribution view to determine the protection group for each data block distribution, so as to perform data read or write operations.

[0008] This invention avoids the complex concepts of various partitions and the DHT calculations of multiple partitions. It abstracts the concept of protection groups in the view, performs mapping in advance, and replaces calculations with table lookups. This results in less computational resource consumption and lower software stack latency in ultra-high performance scenarios.

[0009] In one optional implementation, the data block distribution view is queried based on the received service data to determine the protection group for each data block distribution in order to perform data write operations, including: Determine the data block identifier corresponding to the business data; Based on the data block identifier, determine the protection group identifier mapped to the data block identifier; Based on the protection group identifier, query the disks included in the protection group identifier; Write the business data to the corresponding disk.

[0010] In this embodiment, the chunk ID can be directly used after being requested, just like a virtual built-in address. Physical space is mapped at write time, naturally supporting thin allocation. Through the design of simple read / write logic, single-write lock logic, and fault decision logic, it has a simpler software stack, resulting in higher reliability. It is better suited to the write-time redirection characteristics of SSD media, allowing the full performance of the SSD to be utilized with lower-configuration CPUs, thus saving costs.

[0011] In an alternative implementation, in the event of at least one disk failure, the method further includes: Determine if an authoritative set exists in the protection group. The authoritative set is the set of disk fragments whose disk status has always remained normal. If an authoritative set exists, determine whether the authoritative set includes the recovery set and whether the number of shards in the recovery set is greater than or equal to the number of data shards. If the recovery set is the shards that have been successfully written to the business data, then the recovery set is the shards that have been successfully written to the business data. If the authoritative set includes the recovery set and the number of shards in the recovery set is greater than or equal to the number of data shards sent, the recovery set is selected for data recovery.

[0012] In one alternative implementation, in the absence of an authority set, all fragments in the protection group are selected for data recovery.

[0013] The selection of the data consistency consensus set is simple and reliable, and the correctness of the set selection can be easily proven mathematically, making the handling of faulty data consistency more reliable.

[0014] In an alternative implementation, the method further includes: reclaiming the migrated data blocks from the data block distribution view after the data blocks have been migrated from the disk.

[0015] In this embodiment, data block recycling can effectively prevent data block leakage, reduce performance loss caused by data redundancy, and thus improve the overall responsiveness of the system.

[0016] In one optional implementation, before querying the data block distribution view based on the received business data, the method includes: determining whether the client request information is consistent with the view information in the data block distribution view.

[0017] In a second aspect, the present invention provides a distributed storage data read / write device, the device comprising: The collection module is used to obtain information about the distributed cluster. The generation module is used to generate a data block distribution view for each storage pool based on the distributed cluster information. The data block distribution view includes multiple protection groups, and each protection group includes the corresponding disk information. The query module is used to query the data block distribution view based on the received business data, determine the protection group of each data block distribution, and perform data read or write operations.

[0018] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the distributed storage data read / write method of the first aspect or any corresponding embodiment described above.

[0019] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the distributed storage data read / write method of the first aspect or any corresponding embodiment described above.

[0020] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the distributed storage data read / write method of the first aspect or any corresponding embodiment described above.

[0021] It should be noted that the distributed storage data read / write device, computer equipment, and computer-readable storage medium provided by this invention correspond to the aforementioned distributed storage data read / write method. Therefore, for the beneficial effects of the distributed storage data read / write device, computer equipment, and computer-readable storage medium, please refer to the description of the corresponding beneficial effects of the distributed storage data read / write method above, and will not be repeated here. Attached Figure Description

[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating a distributed storage data read / write method according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a storage pool according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the logical topology according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the physical hardware topology according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a data block distribution view according to an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the composition of a chunk id according to an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating the composition of pg id according to an embodiment of the present invention; Figure 8 This is a schematic diagram illustrating that the authority set is a non-empty set according to an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating that the authority set is an empty set according to an embodiment of the present invention; Figure 10 This is a schematic diagram of chunk recycling according to an embodiment of the present invention; Figure 11 This is a schematic diagram of a chunk client according to an embodiment of the present invention; Figure 12This is a structural block diagram of a distributed storage data read / write device according to an embodiment of the present invention; Figure 13 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] The following are explanations of some terms used in this invention: EC stands for Erasure Code, a coding technique that takes n original data sets, adds m data sets, and can restore the original data using any n data sets from the n+m sets. Mathematically, encoding is represented by constructing a multivariate linear equation, and decoding is represented by solving the multivariate linear equation. Further abstraction leads to matrix operations. Current erasure coding techniques are mainly divided into two categories: one is erasure coding based on Galois field operations (such as RS); the other is erasure coding based on the "XOR" operation (LDPC).

[0026] Chunk: meaning data block, uniquely identified by chunk ID, provides a byte stream space that supports appending and writing, similar to virtual address memory write-on-write functionality. It is generally up to 128M in size and provides chunk operation interfaces for reading, appending, closing, deleting, and querying length.

[0027] pool: Used to describe a contiguous virtual logical address space of chunk IDs. The pool ID serves as a unique identifier, and when a business uses a chunk, it needs to request a chunk ID from within the pool.

[0028] PG: Protect Group, meaning logical protection group, uniquely identified by a PG ID. In a pool, to simplify chunk ID management, a PG is abstracted between the chunk and the pool, mapping chunk IDs to fixed PG IDs using a simple and reliable algorithm. A PG describes a group of independent virtual hard disks, organized together using ECs or replicas according to the protection strategy defined by the pool to provide reliability guarantees. It is a logical protection group that provides read / write space externally, simplifying cluster management.

[0029] Strong consistent write: Writing data to all shards returns success, ensuring that data is never lost under any circumstances.

[0030] chunkserver: meaning a standalone storage server. In this invention, it is used to describe a standalone storage server that manages the storage space on the standalone server and manages the standalone byte stream chunk space.

[0031] Chunkmaster: Meaning cluster management center, in this invention it is responsible for a series of cluster management functions, including data distribution, data balancing, fault recovery, cluster scaling, cluster monitoring, virtual space allocation, and virtual space to physical space conversion. Different storage engines use different naming methods, and it is also referred to as master or CM (cluster master), etc.

[0032] chunkclient: meaning cluster read / write SDK (Software Development Kit).

[0033] Current distributed storage data read / write methods often suffer from the following problems: 1. Long fault handling time. For example, in order to ensure data reliability, Ceph needs to determine the authoritative node before reading and writing when encountering a fault in the storage cluster, and this time is uncontrollable.

[0034] 2. Ceph's read and write methods require locking, which prevents the new hardware from performing at its full potential. Compared to leading products in the field, Ceph's latency is many times higher.

[0035] 3. To achieve one-to-many reads, many products use relatively complex control methods, such as Ceph. In the event of a hardware failure, the write is degraded and the background catches up with the data difference, which reduces reliability.

[0036] In view of this, according to an embodiment of the present invention, a distributed storage data read and write method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0037] This embodiment provides a distributed storage data read / write method, which can be executed by devices such as servers, terminals, and mobile terminals. Figure 1 This is a flowchart of a distributed storage data read / write method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain distributed cluster information.

[0038] Step S102: Generate a data block distribution view for each storage pool based on the distributed cluster information. The data block distribution view includes multiple protection groups (PGs), and each protection group includes corresponding disk information.

[0039] Step S103: Query the data block distribution view based on the received service data to determine the protection group of each data block chunk distribution for data read or write operations.

[0040] The above process can be executed by the chunk master in devices such as servers. The chunk master controls the data distribution of the entire cluster. It can be a Raft (a distributed consensus algorithm) cluster or a Paxos (a distributed consensus algorithm) cluster, or it can rely on the cluster management capabilities provided by an external Raft cluster (such as etcd (a highly available distributed key-value database)). As the metadata management center of the entire cluster, the chunk master can receive distributed cluster information reported by the chunkserver, including distributed cluster physical disk information and server information, and report the disk allocation cluster-unique disk identifiers to the chunkserver. Furthermore, the chunk master does not directly manage the chunks of the entire cluster, but manages the abstract protection groups (PGs) above the chunks.

[0041] Based on the physical disk information and server information reported by the chunk server, a logical topology can be constructed, referring to... Figure 2 As shown, the cluster is divided into several storage pools. PG management is performed on each storage pool, including data balancing, fault handling, PG migration, generating a data block distribution view (cdmap), chunk allocation, disk selection, and view analysis. Further, refer to... Figure 3 As shown, the logical topology of the storage pool is as follows: starting from root, there are multiple logical levels including AZ (data center), Row (same group of aggregation switches), Rack (same group of access switches), Host, and Vdisk. Figure 4 This refers to the physical hardware topology of the cluster.

[0042] Multiple protection policies can be created on the logical topology, such as Row fault domain + 3 replicas, Rack fault domain + EC 4 + 2:1, Host fault domain + EC 4 + 2, etc. Additionally, during storage pool creation, a PG is created based on the topology's protection policies, such as Rack fault domain + EC 4 + 2:1. During PG creation, disks are added to the PG's authoritative shard set according to the protection policy. After successful storage pool creation, a cdmap is generated. After cdmap generation, during the keep-alive process for physical disks, the chunk server is notified by the chunk master that the designed storage pool needs to be updated with the cdmap. When hardware failures occur in the cluster, such as disk, node, or network failures, the cluster's fault handling will also update the cdmap according to the hardware failure, thereby ensuring the accuracy of chunk allocation.

[0043] A view of the data block distribution corresponding to one of the storage pools. Figure 5 As shown, the data block distribution view mainly consists of four parts: cdmap_ver, pool information, PG information (i.e., ... Figure 5 Available PG sets in the database) and disk information (i.e. Figure 5 (The cluster disk collection in the cluster).

[0044] Pool information includes: pool_id, pool_ver, available_chunk_cnt, pg_cnt, available_pg_cnt, and flags. `pool_id` represents a contiguous virtual logical space containing 2^56 virtual chunk spaces. `pool_ver`, combined with SDK logic, indicates whether the current chunk space belongs to the pool space at the current view time. Chunk IDs are mapped to chunks through simple calculations, enabling data reading and writing. `avail_chunk_cnt` indicates the number of available chunks, `pg_cnt` indicates the number of PGs, and `avail_pg_cnt` indicates the number of available PGs. Additionally, `flags` records some flags, such as allowing non-authoritative reads.

[0045] PG information includes: pg_id, pg_ver, status, init_set, auth_set, and r_set. pg_id represents a unique protection group space within the cluster under a specific logical clock. This protection group organizes the physical space of multiple disks within the cluster into protection spaces according to a protection policy. These protection spaces may have multiple replicas, such as EC 4+2 or EC 20+4. pg_ver represents the protection group version number, used during cluster fault recovery to ensure consistency of the PG used for recovery across all chunk servers. status indicates the current read / write status of the PG. init_set represents the set of raw disks for the PG; auth_set represents the set of authoritative disks; r_set represents the set of recovery disks; and m_set represents the set of migration disks. The raw disk information included in each protection group PG can be determined during disk balancing.

[0046] The disk information records unique information about all disks in the cluster. The vdisk_id corresponds to the vdisk_id recorded in the PG. That is, the identifier 1 of vdisk1 in the disk information corresponds to 1 in init_set[(1,13,25,37)(49,61)] in the PG information.

[0047] Finally, the data block distribution is queried based on the data block distribution view to determine the final data read and write addresses.

[0048] This invention provides a write-time redirection storage engine read / write method that eliminates the complex concepts of various partitions and the DHT calculations of multiple partitions. It abstracts the concept of protection groups in the view, performs mapping in advance, and replaces calculations with table lookups. This results in lower computational resource consumption and lower software stack latency in ultra-high performance scenarios.

[0049] The distributed storage data read / write method provided by this invention can be used as a data read / write method for public and private cloud storage infrastructures in append-write scenarios. It can also be used as a data read / write method for software-defined storage products in append-write scenarios. Furthermore, it can be used as a persistent data read / write method for clusters in append-write scenarios.

[0050] In some optional implementations, the data block distribution view is queried based on the received service data to determine the protection group for each data block distribution in order to perform data write operations, including: Determine the data block identifier corresponding to the business data, i.e., the chunk id in this embodiment; Based on the data block identifier, determine the protection group identifier mapped to the data block identifier, i.e., the pg id in this embodiment; Based on the protection group identifier, query the disks included in the protection group identifier; Write the business data to the corresponding disk.

[0051] Reference Figure 6 As shown, the chunk ID consists of a 1-byte pool_id and a 7-byte chunk index. The chunk_id represents a virtual logical space, with a maximum value determined by the view (e.g., 256MB). During use, it is mapped to the actual physical space, is 8 bytes long, and is unique within the cluster. The chunk index indicates the protection group the chunk belongs to. (Refer to...) Figure 7 As shown, the pg id consists of a 1-byte pool_id and a 3-byte pg index. pg_index = chunk index & (pg_cnt - 1); pg_cnt is an integer power of 2.

[0052] Business data is written by requesting a chunk ID through the chunk client. The chunk client caches the chunk ID; if no chunk ID is cached, writing is not allowed.

[0053] Business data needs to carry the following information when reading and writing: `cdmap_ver`, `pool_ver`, `pg_ver`, `shard_ver`, and `chunk_id`. Specifically: `cdmap_ver` is the version number of the cdmap, used by the chunk client and chunk server to determine if the local view version is up-to-date, thus achieving view propagation. `pool_ver` is the pool version number, globally unique across the entire cluster, used to check if the chunk client's cdmap and the chunk server's cdmap belong to the same pool. `pg_ver` is the PG version number, used to check if the PG in the chunk client's cdmap and the chunk server's cdmap are consistent. `shard_ver` is the shard version number within the PG, globally unique, used to prevent the same disk from being selected at different times and in different shard locations during recovery, which could lead to indistinguishable chunks. `chunk_id` is the unique identifier for the chunk, containing the most important information: `pool_id` and `pg_id`.

[0054] Based on the protection group identifier (PG ID), the corresponding disk in the CD map is queried, such as: PG (ProtectGroup) pg_id1, pg_ver, status, init_setl[1, 13, 25, 37)(49, 61)], auth_set[(1, 13, 25.37)(49,61)], r setl.m set[], where 1, 13, 25, and 37 are the disks corresponding to pg_id1, and the data is written to the corresponding disks. In this embodiment, the database chunk space is similar to a virtual memory address, which can be used after application. When data is actually written, it is mapped to the actual physical space. Moreover, the write permission of the chunk is controlled by the memory of the client chunkclient SDK. When the pool version changes, the chunk client SDK directly clears the chunk memory, prohibits all chunk writing, and ensures that the chunk is only written by a single thread. In addition, the size of a chunk can be limited by the granularity of disk management by the chunk server, and synchronized to the chunk client SDK by the CDMap generated by the chunk master.

[0055] Reading based on cdmap works as follows: Based on the chunk ID passed in by the business data, the PG is calculated based on the cdmap. Based on the protection policy and sharding information in the PG, the data at the specified location can be read. If the PG is found to be unreadable and writable, it is necessary to wait for 1 second and try to update it to see if there is the latest cdmap. In the event of cluster anomalies, the pause is relatively normal and should not put pressure on the chunkmaster.

[0056] In this embodiment, the chunk ID can be directly used after being requested, just like a virtual built-in address. Physical space is mapped at write time, naturally supporting thin (a dynamic disk allocation method) allocation. Through the design of simple read / write logic, single write lock logic, and fault decision logic, it has a simpler software stack, resulting in higher reliability. It is more adaptable to the write-time redirection characteristics of SSD media, allowing the full performance of the SSD to be utilized with lower-configuration CPUs, thus saving costs.

[0057] In some alternative implementations, in the event of a failure of at least one disk, the method further includes: Determine if an authoritative set exists in the protection group. An authoritative set is a set of fragments whose disk status has always remained normal. An authoritative fragment is a fragment whose disk status has remained normal after the PG has reached full fragmentation and consensus.

[0058] If an authoritative set exists, determine whether the authoritative set includes the recovery set and whether the number of shards in the recovery set is greater than or equal to the number of data shards. If the recovery set is the shards that have been successfully written to the business data, then the recovery set is the shards that have been successfully written to the business data. If the authoritative set includes the recovery set and the number of shards in the recovery set is greater than or equal to the number of data shards sent, the recovery set is selected for data recovery.

[0059] The storage cluster contains multiple firmwares, and the annual disk failure rate is <=5%. However, failures can still occur within the cluster. After a failure, the following principles are followed to select a set of disks that can make PG consistency decisions. This can ensure data availability in the event of a single point of failure and improve the overall system reliability.

[0060] Reference Figure 8 As shown, the total set is: A = {n | n is the initial shard created by PG: data shard + verification shard}; the authoritative set is: B = {x | x is the shard whose state has been normal after PG reaches consensus}; the recoverable set is: C = {y | y is the shard within PG, and the number of sets is >= the number of data shards}.

[0061] When the authoritative shard contains data that was successfully written by the business, if the length consensus set is chosen as C∈D and D∈B, then set D will definitely reach a consensus on the data that was successfully written by the business.

[0062] The above addresses situations where certain disks and nodes within the cluster fail and cannot be recovered.

[0063] In some alternative implementations, in the absence of an authority set, all fragments in the protection group are selected for data recovery.

[0064] Reference Figure 9 As shown, the total set is: A = {n | n is the initial shard created by PG: data shard + check shard}; the authoritative set is: B = {x | x is an empty set}; the recoverable set is: C = {yly is the shard within PG, and the number of sets is >= the number of data shards}.

[0065] When the authoritative shard is an empty set, and the length consensus set is chosen as D=A, which is the entire set of PG, the entire set must contain the data that was successfully written by the business and can be reached through consensus.

[0066] The above addresses scenarios where large-scale cluster failures can be recovered, especially in cases of intermittent network outages.

[0067] In other words, when there are sufficient authoritative shards, the set for consensus decisions is chosen as {x | x is an authoritative shard, and the number of authoritative shards >= the number of data shards}. When there are insufficient authoritative shards, the set for consensus decisions is the entire set. Clearly defining the set of disks for consensus decisions ensures data availability in the event of a single point of failure, improving the overall system reliability. The selection of the data consistency consensus set is simple and reliable, and its correctness can be easily proven mathematically, making the handling of failed data consistency more reliable.

[0068] In some alternative implementations, the method further includes: After a data block is migrated from the disk, the corresponding data block is reclaimed from the data block distribution view.

[0069] There are two ways to reclaim chunk space: one is through application deletion of chunks; the other is through chunk reclamation after changes to the cdmap. Application-based chunk reclamation occurs immediately upon receiving a chunk deletion command; in the event of a failure, consensus negotiation will continue the deletion process.

[0070] This embodiment uses cdmap-based recycling, which includes the distribution of chunks. When the utilization of a disk is too high, the PGs (Platform Groups) on that disk need to be migrated, that is, the chunks mapped to the PGs need to be migrated. After the chunk migration, the cdmap changes, and the migrated chunks can then be recycled through the chunk server. (Refer to...) Figure 10 As shown, when a chunk in the 8th PG in pool1 needs to be reclaimed, PG1.8 can be migrated to another disk, such as disk 4, while PG1.3 is migrated from disk 4. Then, the chunk server reclaims the chunk space belonging to PG1.3 according to the cdmap.

[0071] In this embodiment, data block recycling can effectively prevent data block leakage, reduce performance loss caused by data redundancy, and thus improve the overall responsiveness of the system.

[0072] In some optional implementations, before querying the data block distribution view based on the received business data, the process includes determining whether the client request information is consistent with the view information in the data block distribution view. In this embodiment, the client request information includes the client's chunk client cdmap version number, pool version number, data shard location, etc., as detailed below.

[0073] The chunk master controls all version number changes within the cluster; all version numbers in the chunk server and chunk client components originate from the chunk master. The chunk master can also control whether data on the chunk server needs to be reclaimed in the background via CDMAP. Additionally, the chunk master provides the function of allocating chunk_ids; applications request chunk_ids from the chunk master through the chunk client, and requests are not allowed without the chunk client's permission.

[0074] The chunk server manages the storage media on storage nodes in the cluster and is the single-machine storage engine in this invention. It is responsible for discovering local storage media, adding them to its management, and synchronizing the information to the cluster management chunkmaster. After the chunkmaster generates a cdmap view, it updates the cdmap locally. Upon receiving read / write requests from a client chunk client, it also needs to perform the following operations: Check the chunk client's cdmap version number and the local version number. A higher version number is more authoritative and does not affect read / write operations. If the cdmap version number is too low, fetch a new cdmap from the chunk master.

[0075] Checking the consistency of `pool_ver` in the `cdmap` determines if a pool has been deleted. Inconsistent `pool_ver` will affect read and write operations, requiring updates to the lagging end of the `cdmap`. This check prevents chunk clients from failing to update the `cdmap` after `pool_id` reuse. When `pool_ver` changes, the chunk ID memory set is cleared, making the chunk IDs held by upper-layer applications unwritable.

[0076] Calculate the PG ID based on the chunk ID, obtain the PG status, and check whether the chunk is writable, readable, or not writable.

[0077] When reading a chunk, it's crucial to compare the shard vertex (shard_ver) with the actual shard vertex on the chunk. Checking the consistency of the shard_ver is essential to determine if the client's perceived data shard location matches the local perception. This is critical in EC (Elastic Compute Service) scenarios. Inconsistent shard_ver will impact read and write operations, requiring updates to the lagging end of the cdmap. The shard_ver can also distinguish whether a disk has been repeatedly selected into PGs, but at different shard locations within EC.

[0078] pg_ver primarily checks that during recovery, PG requires all chunk servers within PG to be on the same logical clock.

[0079] When managing chunks, the chunk server needs to persist the pool_ver along with the chunk for easy deletion in the future.

[0080] In addition, regarding the chunk client, as the SDK provided to the business by the read / write method mentioned in this embodiment, the chunk client mainly provides the following functions, as referred to... Figure 11 As shown: Provides an interface to open a pool. When a business application opens a pool, it synchronizes the cdmap of that pool from the chunk master.

[0081] Provides an interface for requesting chunks. When a business requests a chunk ID, the chunk ID is stored in the current worker thread.

[0082] When a business thread writes data to a chunk, it checks whether the chunk ID to be written by the client is in the memory of the current thread. If it does not exist, writing is not allowed, thus ensuring that only a single thread writes to the chunk.

[0083] The chunk client will retry the write operation as needed. If the chunk client is unable to retry the write operation, it will return an error code suggesting that the application switch to writing to a different chunk.

[0084] The chunk client provides an interface for closing chunks. Closing chunks is unrestricted; any client can close a chunk, and the closing operation is idempotent. Once a chunk is closed, it cannot be written to.

[0085] The chunk client provides a query length interface, but the queried length is not guaranteed to be the latest length and may be outdated.

[0086] The chunk client provides a deletion interface. After deletion is called, the chunk becomes untrusted and can no longer be used.

[0087] When a chunk client uses a pool, it first updates the cdmap by opening the pool. Subsequent chunk read / write operations, chunk ID requests, and other operations check if the cdmap version number is outdated, and update the cdmap accordingly.

[0088] The chunk client requests a chunk ID from the pool. Since the chunk master does not directly manage chunks, the chunk master generates the chunk ID from the PG within the pool.

[0089] This invention proposes a view-based data read / write method. All servers install the chunkserver service of this invention, with a unified cluster management chunkmaster. After the chunkserver monitors the hardware and generates a view, services read and write data in the cluster using the chunk client SDK provided by this invention. This invention offers a simpler and more efficient data read / write method; using an RMDA network, read / write latency on NVMe can reach 50µs. Unified view management allows each component to obtain data distribution through a table-like lookup. Adding protection modes only requires changing the table lookup method of the view, making it more flexible. As a self-developed distributed engine read / write method, it can replace the currently used open-source distributed engine Ceph.

[0090] This embodiment also provides a distributed storage data read / write device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0091] This embodiment provides a distributed storage data read / write device, such as... Figure 12 As shown, the device includes: Collection module 201 is used to obtain distributed cluster information; The generation module 202 is used to generate a data block distribution view for each storage pool based on the distributed cluster information. The data block distribution view includes multiple protection groups, and each protection group includes corresponding disk information. The query module 203 is used to query the data block distribution view based on the received business data, determine the protection group of each data block distribution, and perform data read or write operations.

[0092] In some optional implementations, the query module 203 is further configured to determine the data block identifier corresponding to the business data; determine the protection group identifier mapped to the data block identifier based on the data block identifier; query the disks included in the protection group identifier based on the protection group identifier; and write the business data to the corresponding disk.

[0093] In some alternative embodiments, the apparatus further includes: The fault handling module determines whether an authoritative set exists within the protection group. The authoritative set is a collection of disk fragments whose disk status has remained consistently normal. If an authoritative set exists, it checks if the authoritative set includes a recovery set, and if the number of fragments in the recovery set is greater than or equal to the number of data fragments. The recovery set consists of fragments from which business data has been successfully written. If the authoritative set includes the recovery set, and the number of fragments in the recovery set is greater than or equal to the number of data fragments, the recovery set is selected for data recovery. If no authoritative set exists, all fragments in the protection group are selected for data recovery.

[0094] The recycling module is used to reclaim migrated data blocks from the data block distribution view after they have been migrated from the disk.

[0095] The inspection module is used to determine whether the client request information is consistent with the view information in the data block distribution view.

[0096] In this embodiment, the distributed storage data read / write device is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0097] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0098] This invention also provides a computer device having the above-described features. Figure 12 The distributed storage data read / write device shown is shown.

[0099] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 13 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 13 Take a processor 10 as an example.

[0100] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0101] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0102] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0103] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0104] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0105] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0106] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0107] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A distributed storage data read / write method, characterized in that, The method includes: Obtain distributed cluster information; Based on the distributed cluster information, a data block distribution view is generated for each storage pool. The data block distribution view includes multiple protection groups, and each protection group includes corresponding disk information. Determine the data block identifier corresponding to the received business data; Based on the data block identifier, determine the protection group identifier mapped to the data block identifier; Query the data block distribution view and determine the protection group for each data block distribution based on the protection group identifier to perform data read or write operations.

2. The method according to claim 1, characterized in that, The process of querying the data block distribution view and determining the protection group for each data block distribution based on the protection group identifier for data write operations includes: Based on the protection group identifier, query the disks included in the protection group identifier; The business data is written to the corresponding disk.

3. The method according to claim 1, characterized in that, In the event that at least one disk fails, the method further includes: Determine whether an authoritative set exists in the protection group, wherein the authoritative set is a set of disk fragments whose disk status has always remained normal; If the authority set exists, determine whether the authority set includes the recovery set and whether the number of shards in the recovery set is greater than or equal to the number of data shards. If the recovery set is the shard that was successfully written to the business data, then the recovery set is the shard that was successfully written to the data. If the recovery set is included in the authoritative set and the number of shards in the recovery set is greater than or equal to the number of data shards sent, the recovery set is selected for data recovery.

4. The method according to claim 3, characterized in that, In the absence of the aforementioned authority set, all fragments in the protection group are selected for data recovery.

5. The method according to claim 1, characterized in that, The method further includes: After a data block is migrated from the disk, the migrated data block is reclaimed from the data block distribution view.

6. The method according to claim 1, characterized in that, Before querying the data block distribution view based on the received business data, the following is included: Determine whether the client request information is consistent with the view information in the data block distribution view.

7. A distributed storage data read / write device, characterized in that, The device includes: The collection module is used to obtain information about the distributed cluster. The generation module is used to generate a data block distribution view for each storage pool based on the distributed cluster information. The data block distribution view includes multiple protection groups, and each protection group includes corresponding disk information. The query module is used to determine the data block identifier corresponding to the received service data; determine the protection group identifier mapped by the data block identifier based on the data block identifier; query the data block distribution view, and determine the protection group of each data block distribution based on the protection group identifier, so as to perform data read or write operations.

8. A computer device, characterized in that, include: The system includes a memory and a processor, which are interconnected and communicate with each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the distributed storage data read / write method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause the computer to perform the distributed storage data read / write method according to any one of claims 1-6.

10. A computer program product, characterized in that, It includes computer instructions, which are used to cause a computer to perform the distributed storage data read / write method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Data storage method and device, equipment and computer readable storage medium

    CN113031861A

  • Intelligent EC processing method and device

    CN117827097A