A Distributed Storage Data Remapping Method
By introducing vertical and horizontal stripe mapping methods in distributed storage, combining cache disk configuration and data relocation optimization algorithm, the data synchronization problem during hardware replacement is solved, efficient data migration and recovery is achieved, and cloud computing storage is suitable for complex environments.
Patent Information
- Application Number
- CN202411120825.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-08-15
AI Technical Summary
The existing distributed storage technology has too long data resynchronization cycle during hardware replacement, resulting in inefficiency. Especially in complex physical storage environments, data synchronization problems caused by frequent hardware replacement are difficult to effectively solve.
Using a strip mapping method combining vertical and horizontal, through cache disk configuration and data relocation optimization algorithm, data distribution is monitored in real time and flexible reconstruction tasks are performed to shorten data migration and recovery time.
It effectively shortens the data reconstruction time, improves the data synchronization efficiency during hardware replacement, and adapts to cloud computing scenarios with frequent expansion and maintenance, especially in environments with significant altitude changes and large air pressure fluctuations.
Smart Images

Figure CN118819423B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic digital data processing, and relates to a method for redistributing stored data in a distributed manner. Background Art
[0002] Currently, the mainstream storage technologies are divided into three categories, namely file storage, block storage, and object storage. File storage is convenient for access and sharing, but has low read and write efficiency; block storage has high read and write efficiency, but is not convenient for access and is usually bound to a database system; object storage combines the advantages of the former two, but is not flexible enough and has a high cost.
[0003] Currently, mainstream cloud platforms usually adopt the block or object storage CEPH distributed technology (as Figure 1 shown), and store data distributively in the form of BLOCK or OBJECT to provide metadata services. This technology can replace the traditional file system, effectively utilize network bandwidth, give play to the collaborative effect of multiple servers, and improve data read and write efficiency. This technology usually divides files into objects, uses the HASH algorithm to obtain the PG (Placement Group) placement group according to the object name, and then maps multiple copies of the PGID to different OSDs (Object Storage Devices) in the form of a POOL pool group according to the CRUSH algorithm, as Figure 2 shown. This implementation method essentially determines the distribution of data on the data disk by lottery according to the weight of the OSD. Although the storage efficiency is high, it is not flexible enough and lacks means for real-time monitoring, adjustment, and optimization. When hardware is replaced, there is often a problem of too long a data replica resynchronization period. This problem is a probability event and lacks means of manual control. In a scenario with a complex physical environment and frequent hardware maintenance, it is impossible to accurately evaluate the job duration. The fundamental reason is that the operations of data replica resynchronization are all concentrated, resulting in a too long synchronization period and occupying more hardware resources due to the overall data capacity limit or when hardware changes.
[0004] In a physical storage environment with a large altitude range, significant temperature differences, and obvious air pressure fluctuations, on the one hand, the heat dissipation performance of the storage medium decreases when the number of gas molecules per unit volume in the environment becomes small, and on the other hand, the creepage distance of electronic devices decreases due to the lower insulation strength, resulting in a decrease in insulation performance. Therefore, in such an environment, the service life of storage devices becomes shorter, which inevitably leads to frequent hardware replacement and frequent redistribution of stored data. Combining the above, using the existing storage technology for synchronization will inevitably lead to a long time and low efficiency. Summary of the Invention
[0005] To solve the above problems, based on object storage technology, the present invention optimizes and designs a strip mapping method combining vertical and horizontal directions, which can flexibly adjust parameters according to the configuration of data disks and cache disks. The storage service NaBaFs for cloud platforms can effectively shorten data migration and recovery times. The solution of the present invention is applicable to cloud computing scenarios with frequent scaling and maintenance, and is particularly applicable to physical storage environments with a large altitude range, significant temperature differences, and obvious air pressure fluctuations.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A distributed storage data remapping method, comprising the following steps:
[0008] Cache disk configuration: Continuously write to the NaBadev partition in a round-robin manner on each hard disk within a node, and allocate read caches on solid-state disks; Set horizontal striping mode and vertical striping mode according to the configuration of data disks and cache disks, and stripe the data.
[0009] Data relocation: Reconstruct new data replicas from existing replicas, and the object manager places the members on available disks that do not contain member replicas; Add new disks within the cluster to replace faulty disks.
[0010] Data reconstruction: Construct a new tree based on the calculated strip set and the number of replicas; During the process of reconstructing data from the old tree to the new tree, the new tree is marked as the stale tree; When the new tree is completed and reconstructed, it is marked as the latest, and the old tree is deleted.
[0011] Further, the horizontal striping includes: Striping the data, decomposing the data into several strips, each strip containing a unit data volume, several strips are combined to form a strip set, mirroring each strip set to form members, the members have member IDs and storage UUIDs, and writing the members to the disks and nodes of the entire cluster.
[0012] Further, the vertical striping includes: Mirroring the data between nodes, and then striping between disks within each node.
[0013] Further, when a single node contains multiple solid-state disks, both NaBaIntent Log write caches and read caches will create partitions of equal size.
[0014] Further, during data relocation, before adding a new disk, the newly reconstructed data replicas are first placed on other disks, and after adding the new disk, the new data replicas are moved to the new disk.
[0015] Further, the following steps are also included: continuously monitor the capacity usage of nodes and disks, and automatically migrate data when a node or data disk deviates from the average value of resource occupancy of nodes or data disks in the cluster by a sufficient amount, so as to restore the balance of all nodes and disks in the cluster.
[0016] Further, the following steps are also included: when data imbalance is monitored, if the imbalance is between nodes, reconstruct in a horizontal striping mode; when the load is uneven between disks within a node, reconstruct in a vertical striping mode, or reconstruct in a mixed mode of horizontal striping and vertical mode.
[0017] The beneficial effects of the present invention are as follows:
[0018] The present invention is specifically designed for the process of data reconstruction, and performs algorithm optimization design, which can give full play to the read and write performance of the cache SSD hard disk, and shorten the duration of data reconstruction. And it can flexibly set the horizontal striping or vertical striping mode according to the specific configurations of the data disk and the cache disk, make a balance between high availability and performance, and more accurately control the job duration of single hard disk or multi-hard disk replacement. The present invention can monitor the data distribution in real time, and create reconstruction tasks targeted, so as to shorten the data reconstruction duration when the number of hardware changes. Description of the Drawings
[0019] Figure 1 It is a schematic diagram of the CEPH architecture.
[0020] Figure 2 It is a schematic diagram of the object storage mapping logic.
[0021] Figure 3 It is a schematic diagram of data striping.
[0022] Figure 4 It is a schematic diagram of horizontal striping.
[0023] Figure 5 It is a schematic diagram of vertical striping.
[0024] Figure 6 It is a schematic diagram of data relocation.
[0025] Figure 7 It is a schematic diagram of adding a new disk for data relocation.
[0026] Figure 8 It is a schematic diagram of generating a temporary new tree during data reconstruction.
[0027] Figure 9 It is a schematic diagram after the new tree is completely rebuilt during data reconstruction. Detailed Embodiments
[0028] The technical solution provided by the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that the following specific implementation methods are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0029] The present invention Figure 2 Introducing the striping optimization algorithm in step 3 can flexibly control the distribution of data on the hard disk and associate it with the cache disk, so as to give full play to the performance of the SSD during data reconstruction and shorten the data reconstruction cycle.
[0030] NaBaFS (Nature Balance File System) will actively balance the capacity of the NaBaFS cluster. NaBaFS will not wait for nodes to reach a certain capacity limit before rebalancing data, but will continuously monitor the capacity usage of nodes and disks to see if they deviate from the average of other cluster resources. When a node or data disk deviates sufficiently from the average of the resources occupied by nodes or data disks in the cluster, NaBaFS will automatically migrate data to restore the balance of all nodes and disks in the cluster. When data imbalance is detected, it will be reconstructed in horizontal stripe mode if it is imbalanced between nodes, in vertical mode if the load between disks in a node is uneven, or in a mixture of the two modes. When the difference between the hard disk usage between nodes and the usage of different hard disks in the same node reaches a certain proportion of the average usage (the proportion should be pre-set), it indicates data imbalance.
[0031] In NaBaFS, the stripe width is defined as the minimum amount of data that can be placed on disk (1MB by default). Figure 3 As shown in the figure, a stripe group is a logical entity that represents a collection of these stripes. The number of stripe sets (or stripe groups) for a file is determined by the base of the number of disks in the cluster divided by the number of replicas of the file. To reduce the amount of associated metadata, the maximum number of stripe groups for any file is set to 8 by default, but this is a configurable parameter. NaBaFS breaks data into multiple 1MB stripes, and then rotates the 1MB stripes to form the total stripe set based on the previously calculated number of stripe sets. The number of stripe sets is determined by the number of all data disks and replicas of the storage node. Assuming there are 12 data disks and 3 replicas, when the total segment is 36 and the stripe width is 1M, the 1M, 37M, and 73M of the storage object are combined together. This process of adjusting the order of 1, 2, and 3 to 1, 37, and 73 is called rotation. The total data size divided by 36M is the number of stripe sets. Each stripe set is mirrored according to the replication strategy to form members, which have associated member IDs and storage UUIDs (Universally Unique Identifiers). Members will write to disks and nodes across the cluster. When striping, the file tree is built by member.
[0032] Through horizontal striping, NaBaFS can stripe data as widely as possible across the nodes in the cluster. The object manager ensures that the data and its replicas are stored on different nodes within the cluster to provide redundancy in case of node failure. It also stripes the data to prevent different members from coexisting on the same disk. The object manager preferentially places the data on the storage pool (hard disk) with the largest available capacity. The data is first striped and then each stripe set is mirrored across the disks and nodes of the entire cluster.
[0033] Figure 4 Shows a scaled-down environment where each disk contains only one member of a file. In actual deployment, multiple members from multiple files will reside on each disk.
[0034] In a vertical striping environment, the data is first mirrored between nodes and then striped between the disks within each node. The number of members on a node in a vertical striping environment is equal to the number of available disks on that node.
[0035] Figure 5 Shows a scaled-down environment where each disk contains only one member of a file. In actual deployment, multiple members from multiple files will reside on each disk.
[0036] Step 1, Cache disk configuration scenario:
[0037] NaBaFS maximizes performance by using flash media for real-time logs, read caches, and flash-accelerated metadata in a hybrid storage configuration. The NaBaIntent Log (real-time log NxL) is a write cache with a default consumption of 100GB per node, and its data can be evenly distributed across up to 4 solid-state drives on average. Each hard disk within a node has an associated NaBadev (NaBa device partition) partition with a capacity of up to 5% of the hard disk capacity. The NaBadev partitions are written to the solid-state drives and the consecutive NaBadev partitions are written to the solid-state drives within the node in a round-robin (RR) manner. After creating the NaBadev partitions, read caches are allocated on the solid-state disks. For the minimum configuration, at least 100GB of space must be available for the read cache to successfully install NaBaFS (Nature Balance FileSystem, it is recommended that the read cache be 300GB). When a single node contains multiple solid-state drives, both NXL and the read cache will create partitions of equal size, with a total capacity of at least 100GB for each partition. The present invention sets the horizontal striping mode and the vertical striping mode according to the configuration of the data disks and the cache disks.
[0038] The following three tables are the cache configuration tables for single SSD, dual SSD, and triple SSD configurations respectively
[0039] Table 1: Single SSD Configuration
[0040] SSD mirror open SSD mirror closed NXL log Not allocated >100G / SSD NaBadev partition Not allocated All created on 1 SSD Read cache Not allocated >100G / SSD
[0041] Table 2: Dual SSD Configuration
[0042] SSD mirror open SSD mirror closed NXL log The 1st SSD > 100G imaged to the 2nd SSD > 100G Per SSD > 50G NaBadev partition Write the mirror of each partition of each SSD to each SSD in RR round-robin fashion Write to each SSD in RR round-robin fashion Read cache Total sum > 100G Total sum > 100G
[0043] Table 3: Triple SSD Configuration
[0044] SSD mirror open SSD mirror closed NXL log The 1st SSD > 100G imaged to the 2nd SSD > 100G, the 3rd SSD 0G Per SSD > 33G NaBadev partition Write the mirror of each partition of each SSD to each SSD in RR round-robin fashion Write to each SSD in RR round-robin fashion Read cache Total sum > 100G Total sum > 100G
[0045] Step 2, data relocation process:
[0046] In NaBaFS, there are two forms of data re-layout: relocation and reconstruction. The most common data relocation is due to capacity rebalancing, but node or drive failures can also cause relocation. Data relocation only means that the object Manager is moving a single member to a new disk or storage pool. Data reconstruction occurs when the storage administrator changes the policy (such as the number of replicas) or adds disks or nodes to the cluster. This process involves the re-layout of the entire file to meet the new parameters set by the environmental changes.
[0047] The workflow described in the following example shows how data relocation works in the case of disk failure. This example also includes the relocation strategy applicable to capacity rebalancing.
[0048] In Figure 6 the workflow shown:
[0049] (1) There are a total of 12 hard disks and two replicas.
[0050] (2) Six stripe sets are mirrored to create a member and its replica (for example, stripe set A1 has member A1-r1 and replica A1-r2).
[0051] (3) File A has 12 members and replicas (A1-r1, A1-r2, A2-r1, A2-r2, etc.), striped across the hard disks of nodes X, Y, and Z.
[0052] Note: For simplicity, Figure 6 a scaled-down environment is shown where each disk contains only one member of one file. In actual deployment, multiple members from multiple files will reside on each disk.
[0053] The data relocation process for the failed disk is as follows:
[0054] Rebuild a new data copy (A6-r2) from the existing copy (A6-r1). The object Manager places the members on available disks that do not contain member copies.
[0055] In this case, the data is co-located on a disk that already contains another member (A5-r2) in the same file.
[0056] Eventually, the cluster adds a new disk to replace the failed disk. The workflow at this time is the same as capacity rebalancing. The existing disk on node z has two members, the new disk has zero members, and all other disks in the cluster have one member. The Object Manager will see an imbalance in the capacity in the cluster and rebalance the data accordingly. One of the members in the disk with two members (A6-r2 in this example) will be moved to the new disk on node z, as Figure 7 shown.
[0057] Step 3, data reconstruction process:
[0058] Figures 8 - 9 Shows how the object Manager handles a policy change from two data copies to three copies. In the Figure 8 workflow shown:
[0059] There are initially 12 hard disks and two replicated copies, which means there are six stripe groups.
[0060] Each stripe set is mirrored with a member and its copy (for example, stripe set A1 has member A1-r1 and copy A1-r2).
[0061] File A has 12 members and copies (A1-r1, A1-r2, A2-r1, A2-r2, etc.), striped across nodes X, Y, and Z.
[0062] The administrator changes the replication policy from two copies to three copies without adding any hard disks.
[0063] Note: For simplicity, Figure 8 、 Figure 9 shows a scaled-down environment where each disk contains only one member from one file. In an actual deployment, multiple members from multiple files will reside on each disk.
[0064] The data reconstruction process from two copies to three copies is as follows:
[0065] Build a new tree using the newly calculated stripe sets and number of copies. While the data is being rebuilt from the old tree to the new tree, the new tree is marked as a stale tree.
[0066] AsFigure 9 As shown, when the new tree is successfully reconstructed, it will be marked as the latest, while the old tree will be deleted.
[0067] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A distributed storage data remapping method, characterized in that, It includes the following steps: Cache disk configuration: Continuously write to the NaBa dev partition in a round-robin manner on each hard disk within the node, and allocate read caches on the solid-state disks; Set the horizontal striping mode and vertical striping mode according to the configurations of the data disks and cache disks, and stripe the data. The horizontal striping mode means first decomposing the data into several stripes, combining several stripes to form a stripe set, and mirroring each stripe set; The vertical striping mode is to first mirror the data among the nodes, and then stripe the data among the disks within each node; Data relocation: Reconstruct new data replicas from the existing replicas, and the object manager places the members on the available disks that do not contain member replicas; Add new disks within the cluster to replace the faulty disks; Data reconstruction: Build a new tree according to the calculated stripe sets and the number of replicas; During the process of reconstructing the data from the old tree to the new tree, the new tree is marked as the stale tree; When the new tree is completed and reconstructed, it is marked as the latest, and the old tree is deleted; It also includes: When it is monitored that the data is unbalanced, if the imbalance is among the nodes, it is reconstructed in the horizontal striping mode; When the load is unbalanced among the disks within a node, it is reconstructed in the vertical striping mode, or a mixed reconstruction of the horizontal striping and vertical modes is adopted.
2. The distributed storage data remapping method according to claim 1, characterized in that The horizontal striping includes: Striping the data, decomposing the data into several stripes, each stripe containing a unit data volume, combining several stripes to form a stripe set, mirroring each stripe set to form members, and each member has a member ID and a storage UUID, and distributing the members to the disks and nodes of the entire cluster.
3. The distributed storage data remapping method according to claim 1, characterized in that, When a single node contains multiple solid-state disks, both the NaBa Intent Log write cache and the read cache will create partitions of equal size.
Citation Information
Patent Citations
Distributed storage system and storing and reading method for files
CN104639661A
Metadata management and data reconstruction method and device and storage medium
CN109542342A