A storage system and a storage method for a storage volume

By splitting storage volumes into data units and setting up a metadata management cluster, flexible mapping and migration of storage volumes between different storage pools are achieved, solving the problems of expansion and data migration in traditional storage volume architectures, and improving the flexibility and business continuity of the storage system.

CN114995763BActive Publication Date: 2026-05-05HUARUI INDEX CLOUD TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUARUI INDEX CLOUD TECH (SHENZHEN) CO LTD
Filing Date
2022-05-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional storage volume architectures are constrained by the boundaries between logical and physical pools, causing expansion and data migration to affect business continuity. Existing cross-pool volume migration technologies are costly and impact front-end business I/O.

Method used

The storage volume is divided into multiple data units, each with its own metadata set independently of the metadata management cluster. This allows data units to be mapped between different storage pools, enabling flexible migration and mapping at the data unit level.

Benefits of technology

It achieves lightweight data migration, avoids the impact of large-scale data migration on business, meets diversified storage needs, adapts to business changes, and improves the flexibility and reliability of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995763B_ABST
    Figure CN114995763B_ABST
Patent Text Reader

Abstract

This invention relates to a storage system and a storage volume storage method, belonging to the field of data storage technology. The invention divides a storage volume into multiple data units of the same size. Each data unit is equipped with corresponding metadata for managing storage volume attribute information. This metadata is stored in a dedicated metadata management cluster, which stores and manages the metadata corresponding to each data unit. When mapping the storage volume, it can be mapped to different storage pools at the data unit level, so that the storage volume is no longer limited to a single storage pool. Therefore, in the storage system of this invention, the storage volume can be mapped to any logical or physical boundary, making the storage volume space completely independent of any boundary and avoiding the need for storage pool expansion or data migration of the entire storage volume.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a storage system and a storage volume storage method, belonging to the field of data storage technology. Background Technology

[0002] A storage volume specifically refers to a block storage volume, which is a logical space unit used by users in a block storage system, typically corresponding to a virtual disk in the user's system. In current block storage product architectures, a block storage volume belongs to a logical or physical boundary and does not cross this boundary dimension. The logical boundary refers to the scope of the redundancy policy, also called the logical pool. Both traditional and distributed storage employ a set of redundancy policy algorithms to ensure the reliability of multiple copies of data, such as RAID technology and replication technology. The scope of the redundancy policy is the logical space only within the protection space formed by this redundancy policy. Data within this scope is protected by a unique redundancy policy within that scope. The physical boundary refers to the concept of a storage pool or storage cluster, collectively referred to as the physical pool. Both traditional and distributed storage have a physical boundary, which is the storage pool or storage cluster. For traditional block storage systems, this is the collection of all disks corresponding to a single redundancy policy for a block storage volume. For distributed storage, this scope is broader, encompassing not only the collection of disks but also the collection of multiple servers forming a distributed storage cluster.

[0003] In traditional storage architectures (specifically block storage architectures), volumes are subsets of the logical and physical pools. This limits users' access to block storage volumes to the boundaries of these two dimensions. Volumes, however, are logical spaces built on top of the storage system. Volumes have an overselling characteristic (i.e., the effective space of a volume is larger than the actual physical space provided; for example, if the effective space of a volume is defined as 10TB, the physical pool's storage space is 5TB). Therefore, when the logical and physical pools are insufficient, new data cannot be written to the volume due to the limitations imposed by these three dimensions. The only solution is continuous expansion of the underlying space. However, expansion leads to rebalancing of the entire storage cluster, affecting the user experience of other storage volumes within the boundary. If resource constraints prevent timely expansion, services will be unavailable, impacting business continuity. Furthermore, if the redundancy strategy within the boundary has been in a state of fault degradation for a long time and a fault that cannot be repaired in a short time occurs, if any other level of fault occurs, the data will need to be moved as a whole to other logical pools or physical pools. The movement of a large amount of data will inevitably affect the continuity of business.

[0004] To address the aforementioned issues, cross-pool volume migration technology has emerged. This involves adding a software or program outside the storage system to copy and write storage volumes from one logical or physical pool to another, overcoming the previous limitations of not being able to cross pools and thus resolving the two constraints mentioned above to some extent. This solution is implemented out-of-band by the storage system, requiring additional assembly functionality and incurring significant data migration costs. Furthermore, the migration process involves migrating the entire block storage volume. For example, if there is insufficient space in the logical or physical pool, or if a degradation cannot be recovered in a short time, it is necessary to urgently activate the online volume migration function for certain volumes to free up more available space or migrate the affected volumes to the normal cluster. Because the entire volume needs to be migrated during the migration process, a large amount of background I / O is generated. This background I / O affects the foreground business I / O, causing fluctuations in overall business I / O, which can only be restored after the migration is complete.

[0005] Therefore, due to the current architecture of storage volumes, cross-pool volume migration technology does not fundamentally solve the problem at the architectural level, but only addresses it out of the band. It is still limited by the capacity of the storage pool itself, and there are still significant data migration issues. Summary of the Invention

[0006] The purpose of this invention is to provide a storage system and a storage volume storage method to solve the problems of capacity expansion and large-scale data migration caused by the current storage volume storage process being limited by the boundary constraints of the storage pool.

[0007] To achieve the above objectives, the present invention provides a storage system including a storage volume and a storage pool, wherein the storage pool includes a logical pool and / or a physical pool. The storage system further includes a metadata management cluster. The storage volume is divided into at least three data units, each data unit being a virtual space unit of a set size. Each data unit is configured with corresponding metadata, which is used to manage the attribute information of the corresponding data unit in the storage volume. The metadata is stored in the metadata management cluster to achieve metadata storage and management. Each data unit is mapped to multiple storage pools to enable the storage volume to be stored in different storage pools. The metadata management cluster is independent of the storage pools.

[0008] This invention divides a storage volume into multiple data units of equal size. Each data unit has corresponding metadata for managing storage volume attribute information. This metadata is stored in a dedicated metadata management cluster, which stores and manages the metadata for each data unit. When mapping the storage volume, it can be mapped to different storage pools at the data unit level, freeing the storage volume from being confined to a single pool. Therefore, in this invention's storage system, storage volumes can be mapped to any logical or physical boundary, making the storage volume space completely independent of any boundary and avoiding the need for storage pool expansion or entire storage volume data migration.

[0009] Furthermore, the specified size is 512KB-4MB.

[0010] This invention enables the setting of data unit size according to actual business needs, allowing the size of data units in the storage volume to be adjusted according to actual business requirements.

[0011] Furthermore, when a storage pool mapped to a data unit fails, the data in the mapped data unit within that storage pool is migrated to another storage pool.

[0012] When a storage pool fails, the storage system of the present invention only needs to migrate the data of the data unit mapped to the storage pool. Compared with the migration of data in the entire storage volume in the prior art, the amount of data to be migrated is greatly reduced, realizing lightweight data migration without affecting normal data read and write operations.

[0013] Furthermore, when a data unit performs storage pool mapping, the appropriate storage pool is selected for mapping based on the data unit's business data requirements.

[0014] In the storage system of this invention, there are multiple storage pools mapped to storage volumes. Some storage pools may use storage media with faster read and write speeds, while others may use storage media with larger effective capacity. Different data units may have different requirements. For example, some data units are hot data required by the business and require faster storage media, so they can be mapped to storage pools with faster read and write speeds. Or, if the business requires some data units to have larger effective capacity, their data units can be mapped to storage pools with larger effective capacity, thus meeting the diversified data storage needs of the business.

[0015] Furthermore, when the business data requirements of a data unit change, the storage pool mapped to the corresponding data unit is adjusted according to the changed requirements, and the data of that data unit is migrated from the storage pool mapped before the adjustment to the storage pool mapped after the adjustment.

[0016] This invention can also adaptively adjust the mapping relationship of data units in the storage volume in response to dynamic changes in business data. For example, when the business data corresponding to a data unit changes from non-hotspot to hotspot, the mapping relationship of these data units can be adjusted to map them to a storage pool with faster data read and write speeds to facilitate faster data reading and writing. Therefore, this invention can adaptively migrate data according to changes in business data. This migration is more targeted and flexible, and each migration is performed on a data unit basis, resulting in a small amount of data migration.

[0017] The present invention also provides a storage method for a storage volume, the storage method comprising the following steps:

[0018] 1) Divide the storage volume into at least three data units according to a set size, wherein the data unit is a virtual space unit of a set size;

[0019] 2) Set corresponding metadata for each data unit and store each metadata in the metadata management cluster. The metadata is used to manage the attribute information of the corresponding data unit in the storage volume;

[0020] 3) Map each data unit to a different storage pool to achieve mapping of storage volumes at the data unit level. The storage pool includes a logical pool and / or a physical pool. The metadata management cluster is independent of the storage pool.

[0021] This invention divides a storage volume into multiple data units of equal size. Each data unit has corresponding metadata for managing storage volume attribute information. This metadata is stored in a dedicated metadata management cluster, which stores and manages the metadata for each data unit. When mapping the storage volume, it can be mapped to different storage pools at the data unit level, freeing the storage volume from being confined to a single pool. Therefore, in this invention's storage system, storage volumes can be mapped to any logical or physical boundary, making the storage volume space completely independent of any boundary and avoiding the need for storage pool expansion or entire storage volume data migration.

[0022] Furthermore, the specified size is 512KB-4MB.

[0023] This invention enables the setting of data unit size according to actual business needs, allowing the size of data units in the storage volume to be adjusted according to actual business requirements.

[0024] Furthermore, when a storage pool mapped to a data unit fails, the data in the mapped data unit within that storage pool is migrated to another storage pool.

[0025] When a storage pool fails, the storage system of the present invention only needs to migrate the data of the data unit mapped to the storage pool. Compared with the migration of data in the entire storage volume in the prior art, the amount of data to be migrated is greatly reduced, realizing lightweight data migration without affecting normal data read and write operations.

[0026] Furthermore, when a data unit performs storage pool mapping, the appropriate storage pool is selected for mapping based on the data unit's business data requirements.

[0027] In the storage system of this invention, there are multiple storage pools mapped to storage volumes. Some storage pools may use storage media with faster read and write speeds, while others may use storage media with larger effective capacity. Different data units may have different requirements. For example, some data units are hot data required by the business and require faster storage media, so they can be mapped to storage pools with faster read and write speeds. Or, if the business requires some data units to have larger effective capacity, their data units can be mapped to storage pools with larger effective capacity, thus meeting the diversified data storage needs of the business.

[0028] Furthermore, when the business data requirements of a data unit change, the storage pool mapped to the corresponding data unit is adjusted according to the changed requirements, and the data of that data unit is migrated from the storage pool mapped before the adjustment to the storage pool mapped after the adjustment.

[0029] This invention can also adaptively adjust the mapping relationship of data units in the storage volume in response to dynamic changes in business data. For example, when the business data corresponding to a data unit changes from non-hotspot to hotspot, the mapping relationship of these data units can be adjusted to map them to a storage pool with faster data read and write speeds to facilitate faster data reading and writing. Therefore, this invention can adaptively migrate data according to changes in business data. This migration is more targeted and flexible, and each migration is performed on a data unit basis, resulting in a small amount of data migration. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the storage system architecture of the present invention;

[0031] Figure 2 This is a schematic diagram of the data migration process of the storage system of the present invention. Detailed Implementation

[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0033] Storage System Examples

[0034] To address the problem that current storage systems typically confine storage volumes to a single logical or physical boundary, limiting user access to these boundaries, this invention proposes a storage system where storage volumes are divided into multiple data units. Each data unit has corresponding metadata managed by an independent metadata management cluster. Initially, all data units of the storage volume are independent of any physical or logical boundary. This ensures the storage volume's space is completely independent of any boundary, making its data space virtually unlimited.

[0035] Specifically, such as Figure 1 As shown, the storage system of this invention includes a storage volume and a metadata management cluster. The storage volume is divided into multiple data units of the same size. Each data unit is a virtual space unit, and each data unit has corresponding metadata. This metadata is used to manage the attribute information of the corresponding data unit in the storage volume, such as the location and number of the data unit within the entire storage volume. In addition to the metadata corresponding to each data unit, there is also metadata for managing the attributes of the entire storage volume. That is, the metadata here has two layers of management: one layer manages the storage volume information, and the other layer manages the data units within the storage volume. Each metadata is stored in a separately configured metadata management cluster, which stores and manages the metadata corresponding to each data unit. Each data unit has a fixed block size, which can be between 512KB and 4MB. After the storage volume is partitioned, each data unit is not initially associated with any logical or physical pool. Therefore, the data space of the storage volume is unlimited and not limited by any logical or physical pool. The purpose of a storage volume is still for data storage, therefore it needs to be mapped to logical pools or physical pools. Mapping can be done at the data unit level, meaning each data unit is mapped independently. Different data units can be mapped to the same logical pool or physical pool, or to different logical pools or physical pools. For example, assuming the storage volume in this embodiment is divided into 100 data units of 4MB size, these 100 data units are mapped to 10 storage pools (logical pools or physical pools), and each storage pool can contain one or more data units. Therefore, the storage volume in the storage system of this invention can be mapped to multiple logical pools or physical pools, allowing the storage volume's space capacity to be distributed across different boundaries. The metadata management cluster stores the metadata of each data unit. Metadata requires very little storage capacity compared to the data unit, typically around a few KB. Even if the data space of the data source is large, the corresponding metadata's capacity is relatively small, and a single data cluster can meet the requirements.

[0036] According to the storage system of the present invention, all data units of the storage volume have independent metadata management, which can be mapped to any logical boundary or physical boundary, giving them different attributes within different boundary ranges. The storage system of the present invention will be described below from two aspects: redundancy strategy and performance.

[0037] Redundancy strategy dimension: For example, if a distributed storage cluster A adopts a three-replica redundancy strategy, then one or more data units in the storage volume X of this invention can be mapped to the spatial boundary of A, and these data units will all have the three-replica redundancy strategy of A; and if there is another distributed storage cluster B that adopts an EC (8+2) redundancy strategy, then one or more data units in the storage volume X can also be mapped to the spatial boundary of B, and these data units will have the EC (8+2) redundancy strategy of B.

[0038] Performance dimension: Assuming that a distributed storage cluster C uses all SSDs as the physical storage medium, this invention maps several data units of storage volume X to storage cluster C, and these storage units will have the read and write performance of SSD, a high-performance storage medium.

[0039] When the capacity of the storage pool (logical pool or physical pool) mapped to a data unit is full, it is only necessary to stop mapping other data units to that storage pool, without expanding the storage pool or migrating the data. When the storage pool mapped to a data unit is downgraded, the data in that storage pool needs to be migrated. Since the storage volume in this invention is mapped at the granularity of data units, only the data units in the storage pool need to be migrated during the data migration process. These data units are much smaller than the entire space of the storage volume, so the actual amount of data migrated for the data unit can be very small, without affecting the normal writing of data.

[0040] Furthermore, since the storage pool of this invention can be mapped to multiple different storage pools, and the storage media in each storage pool can be different—for example, some storage pools use high-speed SSDs, while others use storage media with moderate read / write speeds—and different services may have different requirements for different data units within a storage volume—for example, some data units may be hot data required by the service and require faster storage media. In this invention, these data units are migrated and mapped to the boundary of the full SSD cluster. Alternatively, if the service requires some data units to have larger effective capacity, these data units are mapped from the three-replica pool to the EC (M+N) pool. Therefore, this invention can adaptively map data units to suitable storage pools based on service needs, making data migration more adaptable to the diverse data requirements of the service.

[0041] This invention enables flexible mapping of data units to meet diverse needs. Therefore, as business operations change, the mapping relationship between data units and storage pools can be adjusted to achieve partial data migration, such as... Figure 2 As shown. For example, previously frequently accessed and written "hot data" may now be replaced by less frequently accessed and written "regular data." In this case, the corresponding data unit can be mapped from a faster storage pool to a slower storage pool. Similarly, if previously regular data becomes hot data, the corresponding data unit can be mapped from a slower storage pool to a faster storage pool. This mapping adjustment allows for partial migration during the data migration process. Because each data unit is an independent management unit, a large-scale data migration over a long period is avoided. This enables migration during periods of low business load and allows for on-demand migration, thus optimizing and ensuring IOPS (read / write) stability during migration.

[0042] Therefore, in addition to addressing the migration requirements of the aforementioned fault scenarios, this invention maximizes its advantages by enabling arbitrary data migration at the data unit level. If a data block needs to be migrated from a three-replica redundancy strategy to an EC (8+2) redundancy strategy, only a simple data migration is required. In summary, the storage system of this invention can mitigate the risks of large-scale data migration caused by online volume migration, achieving lightweight and burden-free data migration, and can also address the risks of boundary capacity exhaustion and degradation scenarios.

[0043] Furthermore, since the storage volume in this invention is completely independent of the boundary range, it can be extended beyond the general description of the boundary range, such as spanning multiple data centers, multiple regions, and multiple public clouds. The storage system based on this invention can easily realize the ability of storage volumes to flow and distribute data across multiple clouds and multiple data centers, enabling richer technical implementations of multi-cloud environments based on this boundaryless storage volume technology.

[0044] Storage method embodiment of storage volume

[0045] This invention first divides a storage volume into at least three data units according to a set size. Each data unit is a virtual space unit of a set size. Then, corresponding metadata is set for each data unit and stored in a management cluster. This metadata is used to manage the attribute information of the storage volume. Finally, each data unit is mapped to a different logical pool or storage pool to achieve storage volume mapping at the data unit level. During the mapping of data units, the mapping relationship can be adaptively adjusted according to specific business needs. The adjustment results enable lightweight data migration without affecting the normal storage of other data. The implementation process of this method has been described in detail in the storage system embodiments and will not be repeated here.

Claims

1. A storage system comprising storage volumes and storage pools, the storage pools comprising logical pools and / or physical pools, characterized in that, The storage system also includes a metadata management cluster. The storage volume is divided into at least three data units, which are virtual space units of a set size. Each data unit is equipped with corresponding metadata, which is used to manage the attribute information of the corresponding data unit in the storage volume. Each metadata is stored in the metadata management cluster to realize the storage and management of metadata. Storage volumes are mapped to different storage pools at the data unit level, enabling storage volumes to be stored in different storage pools. This allows storage volumes to be mapped to any logical or physical boundary. The metadata management cluster is independent of the storage pools; the logical boundary refers to the logical pool, and the physical boundary refers to the physical pool. When a data unit is mapped to a storage pool, the appropriate storage pool is selected based on the data unit's business data requirements. When the business data requirements of a data unit change, the storage pool mapped to the corresponding data unit is adjusted according to the changed requirements, and the data of that data unit is migrated from the storage pool mapped before the adjustment to the storage pool mapped after the adjustment.

2. The storage system according to claim 1, characterized in that, The specified size is 512KB-4MB.

3. The storage system according to claim 1 or 2, characterized in that, When a storage pool mapped to a data unit fails, the data in the mapped data unit within that storage pool is migrated to another storage pool.

4. The storage system according to claim 1, characterized in that, When the storage pool mapped to a data unit is full, other data units will no longer be mapped to that storage pool.

5. The storage system according to claim 1, characterized in that, When the business data of a data unit changes from hot data that requires frequent reading and writing to ordinary data that does not require frequent reading and writing, the data unit is mapped from a storage pool with fast storage speed to a storage pool with average storage speed.

6. A method for storing a storage volume, characterized in that, The storage method includes the following steps: 1) Divide the storage volume into at least three data units according to a set size, wherein the data unit is a virtual space unit of a set size; 2) Set corresponding metadata for each data unit and store each metadata in the metadata management cluster. The metadata is used to manage the attribute information of the corresponding data unit in the storage volume; 3) Map storage volumes to different storage pools at the data unit level to achieve data unit-based mapping of storage volumes, allowing storage volumes to be mapped to any logical or physical boundary. The storage pool includes a logical pool and / or a physical pool. The metadata management cluster is independent of the storage pool. The logical boundary refers to the logical pool, and the physical boundary refers to the physical pool. When a data unit is mapped to a storage pool, the corresponding storage pool is selected for mapping based on the business data requirements of the data unit. When the business data requirements of a data unit change, the storage pool mapped to the corresponding data unit is adjusted according to the changed requirements, and the data of that data unit is migrated from the storage pool mapped before the adjustment to the storage pool mapped after the adjustment.

7. The storage method for a storage volume according to claim 6, characterized in that, The specified size is 512KB-4MB.

8. The storage method for a storage volume according to claim 6 or 7, characterized in that, When a storage pool mapped to a data unit fails, the data in the mapped data unit within that storage pool is migrated to another storage pool.

9. The storage method for a storage volume according to claim 6, characterized in that, When the storage pool mapped to a data unit is full, other data units will no longer be mapped to that storage pool.

10. The storage method for a storage volume according to claim 6, characterized in that, When the business data of a data unit changes from hot data that requires frequent reading and writing to ordinary data that does not require frequent reading and writing, the data unit is mapped from a storage pool with fast storage speed to a storage pool with average storage speed.

Citation Information

Patent Citations

  • Distributed block storage system and data routing method thereof

    CN109327539A

  • Intelligent data placement

    US20160085467A1