A data storage method, system, computer device and storage medium

By using CRUSH algorithm to perform data redundancy and balanced distribution in centralized storage systems, the pressure problem of RAID technology under massive disk data is solved, achieving higher performance and security, and more flexible storage management.

CN116339636BActive Publication Date: 2025-06-24INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310326694.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-06-24
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

When facing massive disk data, the existing centralized storage architecture has too much pressure on RAID technology to flexibly adjust the RAID group, and it is difficult to flexibly adjust the hardware to poor tolerance, and the data reconstruction time is long.

Method used

The CRUSH algorithm is used to perform data redundancy and balance at the data processing level. By chunking and encoding the disk, and the data is equalized and distributed in combination with fault domain information, redundancy policies, placement rules, etc., to form a storage pool and an Extent.

Benefits of technology

Improves the performance and security of storage systems, allowing more disks to fail simultaneously, providing higher redundancy, and supporting flexible disk additions and bad block shielding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116339636B_ABST
    Figure CN116339636B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data storage, and specifically discloses a data storage method, system, computer device, and storage medium. The method includes: fragmenting and encoding a disk into Blocks, and allocating a failure domain for the disk; forming BGs from the Blocks of the disks in multiple failure domains according to data redundancy rules; forming a storage pool based on the BGs, and cutting the BGs into multiple Extents in the storage pool; recording the mapping relationships between Blocks, BGs, Extents, the storage pool, and the physical blocks of the disk; in response to data being written, writing the data to the corresponding physical block of the disk, and recording the data writing information into the mapping relationships. Through the solution of the present invention, a centralized storage method is realized, the redundancy of storage is improved, and in application scenarios with a large amount of disk data, more disks are allowed to fail simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular, to a data storage method, system, computer device, and storage medium. Background Art

[0002] With the rapid development of informatization, the demand for data storage in all walks of life has increased explosively, and various storage systems have emerged. Storage systems with different architectures are carrying huge amounts of data. Currently, the mainstream storage architecture is centralized storage. Centralized storage has centralization, and the entire storage is concentrated on one or more devices in a system. It has mature technology and simple deployment, and usually uses methods such as RAID to avoid single-point disk failures and achieve data redundancy.

[0003] Current centralized architecture storage uses traditional technologies such as RAID and RAID 2.0 to provide fault tolerance for disks in the storage system through operations such as mirroring or parity checking. In scenarios with a large amount of disk data, RAID technology needs to create multiple RAID groups to reduce the calculation and parity checking pressure of each RAID group. As time goes by, the amount of data has increased explosively, and people's demand for large-capacity storage systems is increasing day by day. However, the current RAID technology is already overwhelmed by storage systems with a large number of disks. The existing centralized architecture storage has the following disadvantages:

[0004] 1) It is not suitable for scenarios with a large amount of disk data. When the disk data in the system increases, the calculation and parity checking pressure of RAID also increase exponentially, limiting the disk scale of the storage system and ultimately limiting the performance of the storage system;

[0005] 2) It is not possible to flexibly add or remove disks from the RAID group. It is necessary to plan in advance the level, number of disks, etc. of the RAID group according to the requirements, and it is difficult to change and adjust later;

[0006] 3) After a local bad block appears on the disk, the entire disk needs to be replaced in a timely manner, and the system has poor tolerance for hardware;

[0007] 4) The data reconstruction time is long. After replacing the disk, the new disk needs to synchronize all data blocks according to the rules of the RAID group, and the long reconstruction time reduces the data security of the system;

[0008] 5) The number of disks that can be redundant in each RAID group is small. Summary of the Invention

[0009] To address the deficiencies in existing technical solutions, the present invention proposes a data storage method, system, computer device, and storage medium. At the data processing level, algorithms such as CRUSH are used to achieve data redundancy and balance, replacing traditional RAID technology; by partitioning and encoding disks, and combining fault domain information, redundancy policies, placement rules, etc., data is evenly distributed to improve system performance and security.

[0010] Based on the above objectives, one aspect of the embodiments of the present invention provides a data storage method, which specifically includes the following steps:

[0011] Partition and encode the disk into Blocks, and assign a fault domain to the disk;

[0012] According to the data redundancy rule, form Blocks of disks in multiple fault domains into a BG;

[0013] Based on the BG, form a storage pool, and cut the BG into multiple Extents in the storage pool;

[0014] Record the mapping relationships between Blocks, BGs, Extents, storage pools, and disk physical blocks;

[0015] In response to data being written, write the data to the corresponding disk physical block, and record the data write information into the mapping relationship.

[0016] In some embodiments, the method further includes:

[0017] Mark the level of each Extent in the storage pool, where the levels include the first level, the second level, and the third level;

[0018] In response to data being written, store the data in the Extent of the second level, and manage the historical access records of each Extent based on the I / O manager;

[0019] Judge whether the historical access record of each Extent triggers migration at a preset period, and in response to triggering migration, migrate the data in the Extent to the Extent of the first level or the third level based on the judgment result.

[0020] In some embodiments, writing the data to the corresponding disk physical block and recording the data write information into the mapping relationship includes:

[0021] Divide the data into data blocks of a fixed size, and record the volume information of each data block;

[0022] Select the Extent location, and record the Extent location information and the volume information into the mapping relationship;

[0023] Write the data block to the corresponding physical disk block according to the Extent location information, and record the Extent location information and the physical disk block information into the mapping relationship.

[0024] In some embodiments, writing data to the corresponding physical disk block and recording data writing information into the mapping relationship further includes:

[0025] Divide the data into data blocks of a fixed size;

[0026] Record the volume information of each data block and create a fingerprint index for each data block;

[0027] Determine whether the corresponding data block is duplicate data according to the fingerprint index;

[0028] Based on the judgment result, determine whether to perform deduplication operation or write operation on the corresponding data block, and in response to performing the write operation on the corresponding data block, record the data writing information into the mapping relationship.

[0029] In some embodiments, the method further includes:

[0030] In response to receiving a read data request, obtain the logical address of the read data from the read data request;

[0031] Based on the logical address and the mapping relationship, find the corresponding physical disk block to read the data.

[0032] In some embodiments, the method further includes:

[0033] In response to the occurrence of a disk bad block or failure, obtain the BG where the bad block or the failed disk is located, and perform data reconstruction on the bad block or the failed disk based on the CRUSH algorithm, the BG, the failure domain corresponding to the bad block or the failed disk, and the mapping relationship.

[0034] In some embodiments, the method further includes:

[0035] In response to a disk being added, encode the added disk into multiple blocks and add the multiple blocks to the existing BG;

[0036] Based on the CRUSH algorithm, redistribute the data in the BG with blocks added;

[0037] In response to a disk being removed, redistribute the data in the BG corresponding to the removed disk based on the CRUSH algorithm.

[0038] In some embodiments, the data redundancy rules include replica redundancy rules and erasure code redundancy rules;

[0039] The data writing information includes the disk physical block, Block, BG, Extent, and volume information corresponding to the data.

[0040] On the other hand, an embodiment of the present invention further provides a data storage system, including:

[0041] A data processing module configured to encode disk shards into Blocks and allocate a failure domain for the disk;

[0042] The data processing module is further configured to form BGs from the Blocks of the disks in multiple failure domains according to data redundancy rules;

[0043] The data processing module is further configured to form a storage pool based on the BGs and cut the BGs into multiple Extents in the storage pool;

[0044] The data processing module is further configured to record the mapping relationships among Blocks, BGs, Extents, the storage pool, and disk physical blocks;

[0045] A data writing module configured to, in response to data writing, write the data to the corresponding disk physical block and record the data writing information into the mapping relationship.

[0046] In yet another aspect of the embodiments of the present invention, a computer device is further provided, including: at least one processor; and a memory storing a computer program that can run on the processor, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0047] In still another aspect of the embodiments of the present invention, a computer-readable storage medium is further provided, and the computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the above method.

[0048] The present invention has at least the following beneficial technical effects: A centralized storage method is realized, which can allow more disks to fail simultaneously when there is a large amount of disk data, improving redundancy; and on the basis of this solution, disks can be flexibly added or removed, and bad blocks can be masked; the solution of the present invention evenly distributes data by partitioning and encoding disks and combining failure domain information, redundancy rules, placement rules, etc., improving the performance and security of storage. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.

[0050] Figure 1 Block diagram of an embodiment of the data storage method provided by the present invention;

[0051] Figure 2 Schematic diagram of an embodiment of the disk striping encoding provided by the present invention;

[0052] Figure 3 Schematic diagram of an embodiment of the mapping relationship provided by the present invention;

[0053] Figure 4 Schematic diagram of an embodiment of the data architecture provided by the present invention;

[0054] Figure 5 Schematic diagram of an embodiment of the data writing provided by the present invention;

[0055] Figure 6 Schematic diagram of an embodiment of the data reconstruction provided by the present invention;

[0056] Figure 7 Schematic diagram of an embodiment of the data storage system provided by the present invention;

[0057] Figure 8 Schematic diagram of the structure of an embodiment of the computer device provided by the present invention;

[0058] Figure 9 Schematic diagram of the structure of an embodiment of the computer-readable storage medium provided by the present invention. Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention in detail in conjunction with specific embodiments and with reference to the accompanying drawings.

[0060] To better understand the embodiments of the present invention, the following explains the technical terms that appear in the embodiments of the present invention.

[0061] Block: The hard disk is divided into strips according to a fixed granularity, and each strip is numbered according to a certain rule to form a Block. A Block is the smallest unit that makes up a RAID (Redundant Array of Independent Disks).

[0062] Block Group: A set formed by Blocks according to certain rules, abbreviated as BG. The Blocks in BG are selected from different failure domains according to the replica or erasure code strategy. If it is the replica strategy, the corresponding replicas of the Block are selected from the disks in different failure domains; if it is the erasure code strategy, they are selected according to the K+M rule, where K is the number of data Blocks and M is the number of parity Blocks, and the values of K and M are determined by actual requirements.

[0063] Extent: A block divided from BG with a fixed size, which is the basic unit for storage pool allocation and also the basic unit for forming a LUN (Logical Unit Number, often used to describe a logical device of a disk array).

[0064] Failure domain: A collection of one or more physical disks. Multiple disk failures can be tolerated within the same failure domain without data loss.

[0065] Storage pool: Composed of BGs, it is a container for storing storage space resources, and all the storage space used by LUNs comes from the storage pool.

[0066] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two entities or parameters with the same name but different identities. It can be seen that "first" and "second" are only for the convenience of expression and should not be construed as a limitation on the embodiments of the present invention. This will not be elaborated in subsequent embodiments.

[0067] Based on the above purposes, in the first aspect of the embodiments of the present invention, an embodiment of a data storage method is proposed. As Figure 1 shown, it includes the following steps:

[0068] Step S10: Encode the disk slices into Blocks and assign failure domains to the disks;

[0069] Step S20: Form BGs from the Blocks of the disks in multiple failure domains according to the data redundancy rule;

[0070] Step S30: Form a storage pool based on the BGs and cut the BGs into multiple Extents in the storage pool;

[0071] Step S40: Record the mapping relationships between Blocks, BGs, Extents, the storage pool and the physical disk blocks;

[0072] Step S50: In response to data writing, write the data into the corresponding physical disk block and record the data writing information into the mapping relationship.

[0073] Those skilled in the art can understand that the order of the steps of the methods described above and below is not limited to the listed order and can be adjusted as needed in practical applications. Some steps can also be combined or omitted without departing from the protection scope of the present invention.

[0074] Specifically, based on steps S10, S20, S30, S40

[0075] In step S10, as Figure 2 shown, the disk forms Blocks through fragmentation and encoding. Each Block is represented by X.Y, where X represents the disk number and Y represents the Yth Block inside the disk. Each Block corresponds to the corresponding block (Logical Block Address, abbreviated as LBA) of the disk.

[0076] Set up failure domains for single or multiple disks to ensure that all disks are allocated in multiple failure domains.

[0077] In step S20, form BGs from the Blocks in multiple failure domains according to the data redundancy rules. On the one hand, it ensures that each physical disk contains a part of the Blocks in several BGs; on the other hand, it ensures that different BGs can use different redundancy rules as needed, and can flexibly utilize the storage space.

[0078] In step S30, one or more BGs contribute Extents to form a storage pool; an Extent is the smallest allocation unit of the storage pool, and several Extents in the storage pool are combined into a LUN. There can be multiple storage pools in the storage system, each BG can only belong to one storage pool, and the data redundancy in the storage pool is determined by the BGs that make up it; each physical disk contains the data of one or more storage pools.

[0079] Since the failure domains have been considered when the BG selects Blocks, each Block of the BG comes from different failure domains. When a disk or a disk block in the same failure domain fails, the data on it can be recovered through the Blocks outside the failure domain, and the data will not be lost. Therefore, all disks in the same failure domain are allowed to fail simultaneously without affecting the integrity of the data.

[0080] In step S40, the mapping relationship between each data unit and the physical disk block is maintained by the mapping table.

[0081] The following uses a specific embodiment to illustrate this mapping relationship. It should be understood that the embodiments described here are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0082] As Figure 3 shown, the storage system maintains a mapping table that records the mapping relationship from user information to the data storage location. The mapping relationship can include:

[0083] Volume information and Extent address mapping relationship; Extent and BG mapping relationship; BG and Block mapping relationship; Block and disk physical block mapping relationship.

[0084] Based on the actual application scenario, the volume information and Extent address mapping relationship can also be replaced with the volume information and fingerprint index mapping relationship, and the fingerprint index and Extent address mapping relationship.

[0085] This mapping relationship is the basis for subsequent data processing.

[0086] Through the above solution, a centralized storage method is implemented, which solves problems such as the number of RAID disks in traditional centralized storage and long reconstruction time. On the basis of this solution, disks can be flexibly added or removed, and bad blocks can be masked. In the case of a large number of disks, more disks are allowed to fail simultaneously, improving redundancy.

[0087] In some embodiments, the method further includes the following steps:

[0088] S61. Mark the level of each Extent in the storage pool, where the levels include the first level, the second level, and the third level;

[0089] S62. In response to data writing, store the data in the Extent of the second level, and manage the historical access records of each Extent based on the I / O manager;

[0090] S63. Determine whether the historical access record of each Extent triggers migration at a preset period, and in response to triggering migration, migrate the data in the Extent to the Extent of the first level or the Extent of the third level based on the judgment result.

[0091] Among them, in step S63, the condition for triggering migration is whether the historical access record of the Extent at the second level is greater than the first threshold or less than the second threshold. The first threshold and the second threshold can be customized by the user based on the actual usage scenario, and the first threshold is greater than the second threshold. If the historical access record of the Extent is greater than the first threshold, the data in the Extent is migrated to the Extent of the first level; if the historical access record of the Extent is less than the second threshold, the data in the Extent is migrated to the Extent of the third level, so as to achieve hierarchical management of data and improve storage performance.

[0092] The following uses specific embodiments to illustrate this mapping relationship. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0093] As shown Figure 4 As shown, a storage pool is composed of one or more BG contribution extents. BGs from different disk media are marked with different levels and classified based on the performance of the disk media. The better the performance of the disk media, the higher the level. For example, a certain type of SSD is classified as the first level Tier0, a second type of SSD and SAS disks are classified as the first level Tier1, and a third type of NearLine disk is classified as the first level Tier2. Thus, the storage pool contains extents of different levels, which can provide a data tiering function for LUNs. When new data is written, it is stored in the Tier1 level. The historical access records of each extent are managed by the I / O manager, and hot data analysis is performed at regular time intervals to migrate hot data to faster levels and cold data to lower levels, thereby improving storage performance.

[0094] In some embodiments, writing data to the corresponding disk physical block and recording the data writing information into the mapping relationship specifically includes the following steps:

[0095] Dividing the data into data blocks of a fixed size and recording the volume information of each data block;

[0096] Selecting an extent location and recording the extent location information and the volume information into the mapping relationship;

[0097] Writing the data block to the corresponding disk physical block according to the extent location information and recording the extent location information and the disk physical block information into the mapping relationship.

[0098] Specifically, when writing data to a physical disk, it includes three functional modules: a default functional module, a data compression functional module, and a data deduplication functional module. The data compression functional module and the data deduplication functional module can be selectively enabled to improve the data writing speed.

[0099] This embodiment describes the default functional module, and the specific steps are as follows:

[0100] S11. Dividing the data into data blocks of a fixed size and recording the volume information of each data block;

[0101] S12. Selecting an extent location and recording the extent location information and the volume information into the mapping relationship;

[0102] S13. Writing the data block to the corresponding disk physical block according to the extent location information and recording the extent location information and the disk physical block information into the mapping relationship.

[0103] In some embodiments, in step S50, writing the data into the corresponding physical disk block and recording the data writing information into the mapping relationship further includes:

[0104] S51. Split the data into data blocks of a fixed size;

[0105] S52. Record the volume information of each data block and create a fingerprint index for each data block;

[0106] S53. Determine whether the corresponding data block is duplicate data according to the fingerprint index;

[0107] S54. Based on the judgment result, determine whether to perform a deduplication operation or a writing operation on the corresponding data block, and in response to performing a writing operation on the corresponding data block, record the data writing information into the mapping relationship.

[0108] Among them, in step S54, determining whether to perform a deduplication operation or a writing operation on the corresponding data block based on the judgment result includes: if the corresponding data block is duplicate data, perform a deduplication operation on the corresponding data block; if the corresponding data block is not duplicate data, select an Extent location, and record the volume information and the fingerprint index, the fingerprint index and the Extent location information into the mapping relationship, write the corresponding data block into the corresponding physical disk block according to the Extent location information, and record the Extent location information and the physical disk block information into the mapping relationship.

[0109] The above solution is an embodiment of enabling the data deduplication function module. The following is an example to illustrate the data writing process of enabling the data deduplication function module. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0110] As Figure 5 shown, when the host performs a write I / O operation on the storage, the data stream written into the storage is decomposed into multiple data blocks, and a corresponding fingerprint index is generated for each data block. When new data is written, the storage checks whether there is already storage for the fingerprint index and this data block before, determines whether the data has been stored on the disk through the fingerprint index, and at the same time records information such as the mapping table in the metadata. The specific data writing process is as follows:

[0111] S510. The host writes data to the storage.

[0112] S520. The storage splits the written data into data blocks of a fixed size.

[0113] S530. Record the volume information for each data block and create a fingerprint index.

[0114] S540. Store information for determining whether it is duplicate data based on fingerprint information. If it is duplicate data, perform deduplication and execute step S550. If it is new data, establish a mapping between the fingerprint information and the Extent and execute step S560.

[0115] S550. Increase the fingerprint reference count corresponding to the data block.

[0116] S560. Select the Extent location and write the data to the physical disk according to the Extent location information; and simultaneously record the fingerprint and Extent mapping information.

[0117] S570. The storage returns a write success message to the host.

[0118] In this embodiment, through data deduplication, redundant data blocks in the storage system are deleted, reducing the physical storage capacity occupied by the data and improving the data writing speed and storage performance.

[0119] In the above embodiment, the data can also be compressed before being written to the physical disk, that is, without losing information, by reorganizing the data to reduce the data volume and thus reduce the storage space, thereby improving the transmission, processing, and storage efficiency of the storage system.

[0120] In some embodiments, the method further includes:

[0121] In response to receiving a read data request, obtain the logical address of the read data from the read data request;

[0122] Based on the logical address and the mapping relationship, find the corresponding physical disk block to read the data.

[0123] Specifically, when the host performs a read I / O operation on the storage, the system searches for the logical address of the read data according to the read I / O operation. The system finds the corresponding physical disk block to retrieve the data according to the mapping relationship.

[0124] In this embodiment, the data location can be quickly found through the mapping relationship, improving the read data speed and storage performance.

[0125] In some embodiments, the method further includes:

[0126] In response to the occurrence of a disk bad block or failure, obtain the BG where the bad block or the failed disk is located, and perform data reconstruction on the bad block or the failed disk based on the CRUSH algorithm, the BG, the failure domain corresponding to the bad block or the failed disk, and the mapping relationship.

[0127] Specifically, when a hard disk fails, it is the process of restoring all or the data of the Blocks affected by the failure on the failed hard disk to other available Blocks. During data reconstruction, if the BG uses the replica strategy for redundancy, the replica data on the non-failed disks is read. If the BG uses the erasure code strategy for redundancy, the non-failed data and parity data are read. By performing corresponding processing on the read data, the data is restored to other available Blocks to ensure data integrity.

[0128] The reconstruction granularity is Block. Based on the CRUSH algorithm (Controlled Replication Under Scalable Hashing, a data distribution algorithm), according to the failure domain and mapping relationship, data reconstruction is performed on the failed BG to improve the data recovery speed when a hard disk fails.

[0129] When performing effective data reconstruction, only the actually modified data segments are reconstructed to quickly return to the normal state.

[0130] As Figure 6 shown, the data is in a triple-replica configuration. Taking the data distribution of Block 1.1 as an example, Block 1.1, 2.2, 3.3, 4.4, 5.5 form a BG. The data replicas on Block 1.1 are distributed to 2.2, 3.3, 4.4, 5.5. When Block 1.1 fails, Block 2.2, 3.3, 4.4, 5.5 will all participate in data reconstruction.

[0131] In some embodiments, the method further includes:

[0132] In response to a disk being added, the added disk is sharded and encoded into multiple blocks, and the multiple blocks are added to the existing BG;

[0133] Based on the CRUSH algorithm, the data in the BG with blocks added is redistributed;

[0134] In response to a disk being removed, based on the CRUSH algorithm, the data in the BG corresponding to the removed disk is redistributed.

[0135] In this embodiment, disk expansion and contraction are achieved, supporting dynamic expansion. The minimum granularity of expansion / contraction is Block. Blocks or disks can be flexibly added / removed according to actual needs, thereby abstracting the data from the hardware. As the number of disks increases, the system storage capacity and performance also increase. The CRUSH algorithm realizes dynamic data distribution by calculating and accepting multi-dimensional parameters. Data redundancy and balanced distribution are achieved through rule definition.

[0136] After capacity expansion or contraction, the system will automatically complete data balancing operations through the capacity balancing function. The data is stored globally in a balanced manner, and there is no centralized data hotspot.

[0137] The disk-based capacity contraction method can also achieve shielding of faulty disk cards. It contracts at the granularity of Blocks. When there are local bad blocks on the disk, the corresponding Blocks of the bad blocks can be contracted to shield the bad blocks, thereby improving the system's hardware tolerance and hardware utilization rate, and reducing the pressure and risk brought by the reconstruction of the entire disk's data when replacing the entire disk.

[0138] In some embodiments, the data redundancy rules include replica redundancy rules and erasure code redundancy rules;

[0139] The data write information includes the physical disk block, Block, BG, Extent, and volume information corresponding to the data.

[0140] The data redundancy rules include replica redundancy rules and erasure code redundancy rules.

[0141] Replica redundancy rules: Replicas protect data by writing data into multiple Blocks in the system and support multiple replicas. Users can set different data redundancy policies according to actual needs, and support the coexistence of replicas and erasure codes.

[0142] Erasure code redundancy rules: Erasure code protection is provided at the Block level of data storage. Different protection capabilities are provided according to the data importance level, and users can create BGs and storage pools with different erasure code protection levels. The erasure code level is K+M, where K represents the number of data blocks and M represents the number of parity blocks. Without data loss, the number of Blocks that can fail simultaneously allowed by the system is M. Taking 8+2 as an example, it means that every 8 data Blocks in the BG correspond to 2 parity Blocks, and the effective space is 8 / (8+2), that is, 80%.

[0143] In summary, the present invention provides an efficient storage architecture and storage system that performs physical disk sharding coding, uses the CRUSH algorithm to comprehensively perform data distribution based on the failure domain and placement rules, and combines technologies such as fingerprint indexing, mapping tables, and data layering, and has functions of deduplication and compression, layering, and dynamic adjustment of the number of disks, and supports shielding of bad blocks.

[0144] The following further illustrates the centralized storage method of the present invention through a specific embodiment, which includes physical disk sharding coding, using the CRUSH algorithm to comprehensively perform data distribution based on the failure domain and placement rules, and combining fingerprint indexing, mapping tables, data layering, etc. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0145] S70. Divide physical disks into failure domains according to a predetermined rule;

[0146] S71. Perform sharding encoding on each disk to form Blocks;

[0147] S72. Use the CRUSH algorithm to select Blocks that meet the system placement rules from different failure domains to form BGs, and the BGs are responsible for data redundancy and balance;

[0148] S73. A number of BG sets form a storage pool, and BGs from different disk media are marked as different levels in the storage pool to achieve automatic data tiering;

[0149] S74. Cut out smaller-granularity Extents from the storage pool, and a number of Extents form a LUN, and the LUN is mapped to the front-end host for use;

[0150] S75. The mapping table records the mapping information of LUNs, storage pools, Extents, BGs, Blocks, and disks and records it in the metadata;

[0151] S76. When the host writes data, the system cuts and shards the written data and creates a fingerprint index for all shards, and realizes data deduplication through fingerprint comparison and the mapping table;

[0152] S77. When the host reads data, the system finds the logical address of the data to be read according to the read I / O operation, and finds the corresponding physical disk block to read the data according to the mapping relationship;

[0153] S78. If the storage pool contains disks with different performances, the I / O manager records the historical access records of each Extent, performs hot spot analysis regularly, and migrates hot and cold data between levels;

[0154] S79. Taking Blocks as the reconstruction granularity, the number of disks in the storage pool can be flexibly adjusted, faulty bad blocks can be masked, and faulty disks can be replaced.

[0155] By optimizing the storage data processing architecture, problems such as the number of RAID disks in traditional centralized storage and long reconstruction time are solved; disks can be flexibly added or removed, and bad blocks can be masked. In the case of a large number of disks, more disks are allowed to fail simultaneously, providing higher redundancy.

[0156] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a data storage system, as Figure 7 shown, the data storage system includes:

[0157] A data processing module 110, the data processing module 110 is configured to shard and encode disks into Blocks and assign failure domains to the disks;

[0158] The data processing module 110 is further configured to form BGs by combining Blocks of disks in multiple failure domains according to data redundancy rules;

[0159] The data processing module 110 is further configured to form a storage pool based on the BGs and cut the BGs into multiple Extents in the storage pool;

[0160] The data processing module 110 is further configured to record the mapping relationships among Blocks, BGs, Extents, the storage pool, and physical disk blocks;

[0161] A data writing module 120, which is configured to write data into corresponding physical disk blocks in response to data writing and record data writing information into the mapping relationships.

[0162] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer device, as Figure 8 shown. In this computer device 30, a processor 310 and a memory 320 are included. The memory 320 stores a computer program 321 that can run on the processor. When the processor 310 executes the program, it performs the steps of the above method.

[0163] Among them, the memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the data storage method in the embodiments of the present application. The processor executes various functional applications and data processing of the system by running the non-volatile software programs, instructions, and modules stored in the memory, that is, implements the data storage method in the above method embodiments.

[0164] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the system, etc. In addition, the memory may include a high-speed random access memory, and may further include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the local module through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0165] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, as Figure 9As shown, the computer-readable storage medium 40 stores a computer program 410 that, when executed by a processor, performs the above method.

[0166] Finally, it should be noted that those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be implemented by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium of the program can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. The above embodiments of the computer program can achieve the same or similar effects as the corresponding foregoing method embodiments.

[0167] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, a general description of the functions of various illustrative components, blocks, modules, circuits, and steps has been provided. Whether this function is implemented as software or hardware depends on the specific application and the design constraints imposed on the overall system. The functions that those skilled in the art can implement in various ways for each specific application, but this implementation decision should not be construed as causing a departure from the scope of the disclosure of the embodiments of the present invention.

[0168] The above are the exemplary embodiments disclosed by the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the disclosure of the embodiments of the present invention defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein need not be performed in any particular order. The above serial numbers of the disclosed embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. In addition, although the elements disclosed in the embodiments of the present invention can be described or claimed in individual form, they can also be understood as plural unless clearly limited to the singular.

[0169] It should be understood that, as used herein, unless the context clearly supports an exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations of one or more of the associated listed items.

[0170] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope (including the claims) disclosed by the embodiments of the present invention is limited to these examples; Under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A data storage method, characterized in that, including: encoding disk shards into Blocks and assigning failure domains to the disks; forming BGs from Blocks of disks in multiple failure domains according to data redundancy rules; forming a storage pool based on the BGs and slicing the BGs into multiple Extents in the storage pool; recording the mapping relationships among Blocks, BGs, Extents, the storage pool, and physical disk blocks; in response to data being written, splitting the data into fixed-size data blocks; recording the volume information of each data block and creating a fingerprint index for each data block; judging whether the corresponding data block is duplicate data according to the fingerprint index; determining whether to perform deduplication or write operation on the corresponding data block based on the judgment result, and in response to performing a write operation on the corresponding data block, selecting an Extent location and recording the Extent location information and the volume information into the mapping relationship; writing the data block to the corresponding physical disk block according to the Extent location information, and recording the Extent location information and the physical disk block information into the mapping relationship.

2. The method according to claim 1, characterized in that, further including: marking the level of each Extent in the storage pool, where the levels include the first level, the second level, and the third level; in response to data being written, storing the data in the Extents of the second level and managing the historical access records of each Extent based on the I / O manager; judging whether the historical access records of each Extent trigger migration at a preset period, and in response to triggering migration, migrating the data in the Extent to the Extents of the first level or the third level based on the judgment result.

3. The method according to claim 1, wherein further including: in response to receiving a read data request, obtaining the logical address of the read data from the read data request; finding the corresponding physical disk block based on the logical address and the mapping relationship to read the data.

4. The method according to claim 1, characterized in that, further including: in response to a disk bad block or failure occurring, obtaining the BG where the bad block or the failed disk is located, and performing data reconstruction on the bad block or the failed disk based on the CRUSH algorithm, the BG, the failure domain corresponding to the bad block or the failed disk, and the mapping relationship.

5. The method according to claim 1, wherein further including: in response to a disk joining, encoding the joined disk shards into multiple blocks and adding the multiple blocks to the existing BGs; redistributing the data in the BGs with blocks added based on the CRUSH algorithm; in response to a disk being removed, redistributing the data in the BG corresponding to the removed disk based on the CRUSH algorithm.

6. The method according to claim 1, characterized in that, The data redundancy rules include replica redundancy rules and erasure code redundancy rules; The data write information includes the physical disk block, Block, BG, Extent, and volume information corresponding to the data.

7. A data storage system, characterized in that, including: a data processing module configured to encode disk shards into Blocks and assign failure domains to the disks; the data processing module is further configured to form BGs from Blocks of disks in multiple failure domains according to data redundancy rules; The data processing module is further configured to form a storage pool based on the BGs, and cut the BGs into a plurality of extents in the storage pool; The data processing module is further configured to record the mapping relationships among blocks, BGs, extents, the storage pool, and physical disk blocks; A data writing module, configured to, in response to data being written, split the data into data blocks of a fixed size; record the volume information of each data block and create a fingerprint index for each data block; determine whether the corresponding data block is duplicate data according to the fingerprint index; determine whether to perform a deduplication operation or a writing operation on the corresponding data block based on the determination result, and in response to performing a writing operation on the corresponding data block, select an extent location, and record the extent location information and the volume information into the mapping relationships; write the data block into the corresponding physical disk block according to the extent location information, and record the extent location information and the physical disk block information into the mapping relationships.

8. A computer device, comprising: At least one processor; And A memory storing a computer program that can run on the processor, wherein when the processor executes the program, it executes the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Disk bad block processing method and device

    CN111143116A

  • Data processing method based on erasure codes and related device

    CN114443350A