Data balancing method and device for redirection-on-write distributed storage engine

Through the hierarchical data equalization method, the problem of data imbalance in the distributed storage engine is solved, and the balanced sharing of capacity and performance pressure is achieved, which improves the performance and scalability of the system.

CN119937905AActive Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
CN202411765680.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

There is a problem of data imbalance in the existing distributed storage engine, which causes some nodes to bear too much storage or processing burden, affecting the system's performance, scalability and fault tolerance.

Method used

The hierarchical data balance method is adopted to split the data balance into four levels: disk selection equalization, space allocation equalization, usage equalization and background equalization adjustment. The cluster hardware is monitored through a unified cluster management center, and the data balance within the cluster is achieved by combining the ROW storage engine and a hierarchical storage architecture.

Benefits of technology

Through this method, the capacity pressure and performance pressure equalization can be shared as much as possible on each storage medium and hardware, so as to achieve a balanced distribution of cluster data, and improve the performance and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937905A_ABST
    Figure CN119937905A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed storage, and discloses a data balancing method and device for a redirection-on-write distributed storage engine, and the method comprises the steps: executing a preset disk selection balancing strategy based on a logic topological structure; after the preset disk selection balancing strategy is executed, the used capacity of the PG and the used capacity of the hard disk are obtained; executing a preset space allocation balancing strategy based on the PG used capacity and the hard disk used capacity; after the preset space allocation equalization strategy is executed, executing a preset use equalization strategy based on the star-shaped read-write mode; and after the preset use balancing strategy is executed, under the condition that the cluster is still not balanced, executing the preset background balancing strategy. According to the method, balanced distribution of cluster data can be effectively achieved through disk selection balance, space distribution balance, use balance and background balance adjustment and through a unified cluster management center monitoring cluster hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed storage technology, and in particular to a method and device for balancing data of a distributed storage engine for write-time redirection. Background Art

[0002] Distributed storage is also known as software-defined storage (SDS) in the storage industry. The Storage Networking Industry Association (SNIA) defines software-defined storage as: a virtualized storage with a service management interface. Software-defined storage includes storage pooling functions, and can define the data service characteristics of the storage pool through the service management interface. The most important part of the definition of virtualized storage is the virtualization of storage hardware. Compared with traditional storage, distributed storage no longer relies on proprietary hardware (such as storage controllers) and proprietary networks (such as FC networks), and can provide professional storage services through general hardware and a unified network.

[0003] Distributed storage technology is now widely used in traditional data centers, public cloud data centers, private cloud data centers and hyper-converged workstation products. The distributed storage engine provides high-performance, high-reliability and efficient data access capabilities for distributed storage and is the core component of distributed storage.

[0004] There are many implementation schemes for distributed storage engines, such as the widely used distributed storage system Ceph and distributed file system GFS (Google File System). Both adopt distributed storage architecture and can store data in a dispersed manner on multiple servers, which is suitable for scenarios such as big data analysis, cloud computing and large-scale data storage. However, they all face the problem of data imbalance, that is, data is unevenly distributed among nodes in the distributed system, which will cause some nodes to bear too much storage or processing burden, which will have a significant impact on the performance, scalability and fault tolerance of the system. Summary of the invention

[0005] In view of this, the present invention provides a method and device for balancing data of a distributed storage engine by redirecting the data during writing, so as to solve the problem of data imbalance of the distributed storage engine in the prior art.

[0006] In a first aspect, the present invention provides a method for balancing data in a distributed storage engine for write-time redirection, the method comprising:

[0007] Executing a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing a heterogeneous cluster physical topology structure, and each logical topology structure includes a plurality of logical isolation sub-topologies;

[0008] After executing the preset disk selection balancing strategy, obtain the used capacity of PG and the used capacity of hard disk;

[0009] Based on the used capacity of PG and the used capacity of hard disk, the preset space allocation balance strategy is executed;

[0010] After executing the preset space allocation balancing strategy, based on the star read and write mode, execute the preset usage balancing strategy;

[0011] If the cluster still fails to reach a balance after executing the preset balancing strategy, the preset background balancing strategy will be executed.

[0012] The present invention proposes a hierarchical data balancing method applied to a distributed storage engine with redirection on write, which divides the data balancing of the cluster into four levels, namely disk selection balancing, space allocation balancing, usage balancing and background balancing adjustment, monitors the cluster hardware through a unified cluster management center, and realizes data balancing within the cluster by combining the ROW storage engine and the hierarchical storage architecture. When the method is applied in a distributed storage engine, the capacity pressure on the cluster will be evenly distributed on each storage medium as much as possible, and the performance pressure will be balanced on each network, computing and storage hardware in the cluster, that is, the cluster will be finely controlled by a unified cluster management center to achieve balanced distribution of cluster data.

[0013] In an optional implementation, executing a preset disk selection balancing strategy based on a logical topology structure includes:

[0014] Step a, determining the weight and fault domain of each topological node in the logical topology structure;

[0015] Step b, selecting the Nth layer in the logical topology structure, where N=1;

[0016] Step c, determining the first topological node in the maximum weight set in the Nth layer and the fault domain corresponding to the first topological node; the first topological node is any node in the maximum weight set;

[0017] Step d, when the fault domain corresponding to the first topological node is the set fault domain, determining whether the first topological node meets the fault requirement;

[0018] Step e1, when the fault requirements are met, select the first topological node, set N=N+1, and return to step c until N is the last layer in the logical topology structure, completing the disk selection of the logical topology;

[0019] Step e2, if the fault requirement is not met, return to step b and use the second largest weight set in the Nth layer as the maximum weight set.

[0020] In disk selection balancing, hardware of the same type is organized into a logical topology. Large heterogeneous storage server clusters are managed through logical topology isolation and storage federation management, which improves the balance of heterogeneous clusters to usage balance. Disk selection balancing only considers the balance of hard disks of the same type. The number of PGs available in the cluster is determined by the disk selection within the storage cluster, which can ensure that each hard disk appears the same number of times in each position in the PG, and even in the EC case, the cluster read and write balance can be guaranteed as much as possible.

[0021] In an optional implementation manner, after completing the disk selection of the logical topology, the method further includes:

[0022] Rearrange the check shards and data shards of all PGs in the cluster so that each hard disk in the logical topology is used as an EC check shard and an EC data shard an equal number of times.

[0023] When selecting disks in EC scenarios, all disks are evenly distributed on the EC's check shards and data shards from the perspective of the logical topology cluster, thereby effectively improving the performance balance of the cluster.

[0024] In an optional implementation, based on the used capacity of the PG and the used capacity of the hard disk, a preset space allocation balancing strategy is executed, including:

[0025] Determine whether the difference in used capacity of the PG exceeds a first preset capacity;

[0026] In the case where the difference in the used capacity of the PG exceeds the first preset capacity, determining whether the time for which the first preset capacity is exceeded exceeds the first preset time;

[0027] In the case where the time of exceeding the first preset capacity exceeds the first preset time, establishing a first cycle PG group, the first cycle PG group being used to balance the used capacity of the PGs;

[0028] Determine whether the difference in used capacity of the hard disk exceeds a second preset capacity;

[0029] When the difference in the used capacity of the PG exceeds the second preset capacity, determining whether the time for exceeding the second preset capacity exceeds the second preset time;

[0030] When the time exceeding the second preset capacity exceeds the second preset time, a second cycle PG group is established on the hard disk that exceeds the second preset capacity and exceeds the second preset time, and the second cycle PG group is used to balance the used capacity of the hard disk.

[0031] In an optional implementation, based on the star-shaped reading and writing mode, a preset usage balancing strategy is executed, including:

[0032] Get the IO data of each hard disk in star read and write mode;

[0033] Based on IO data, determine the data types with different heat levels;

[0034] Data types of different temperatures are allocated to corresponding logical topology structures according to the types of logical topology structures.

[0035] In this embodiment, star-shaped reading and writing are adopted, and there is no master responsible for forwarding reading and writing. In the scenario of large-scale reading, the master will not become a reading hotspot. When heterogeneous servers appear in a cluster, the logical topology is used to logically isolate servers of different structures. When using space, different topologies are used according to the characteristics of the business to achieve heterogeneous cluster balance, which can further effectively optimize data balance.

[0036] In an optional implementation, executing a preset background balancing strategy includes:

[0037] Get the statistics of all PGs and all hard disks in the cluster;

[0038] Based on the statistics of all PGs, determine the hot PGs;

[0039] Determine hot disks based on the statistics of all disks;

[0040] When the hotspot PG and the hotspot hard disk are in conflict for a third preset period of time, the hotspot PG is migrated from the hotspot hard disk to the non-hotspot hard disk through PG migration.

[0041] Migrate hot PGs to non-hot hard disks. Since the number of PGs to be migrated and the migration bandwidth can be set, the impact on the entire cluster is small and controllable, and can further effectively improve data balance.

[0042] In a second aspect, the present invention provides a distributed storage engine data balancing device for write-time redirection, the device comprising:

[0043] A balancing strategy execution module, used to execute a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing the physical topology structure of the heterogeneous cluster, and each logical topology structure includes a plurality of logical isolation sub-topologies;

[0044] The space allocation balancing strategy execution module is used to obtain the used capacity of the PG and the used capacity of the hard disk after executing the preset disk selection balancing strategy; based on the used capacity of the PG and the used capacity of the hard disk, execute the preset space allocation balancing strategy;

[0045] A usage balancing strategy execution module is used to execute a preset usage balancing strategy based on a star-shaped reading and writing method after executing a preset space allocation balancing strategy;

[0046] The background balancing strategy execution module is used to execute the preset background balancing strategy when the cluster still fails to achieve balance after executing the preset balancing strategy.

[0047] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the write-time redirection distributed storage engine data balancing method of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0048] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the write-time redirection distributed storage engine data balancing method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0049] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for write-time redirection of distributed storage engine data balancing according to the first aspect or any corresponding embodiment thereof.

[0050] It should be noted that the apparatus for data balancing of a distributed storage engine for redirection on write, the computer device, and the computer-readable storage medium provided by the present invention correspond to the above-mentioned method for data balancing of a distributed storage engine for redirection on write. Therefore, for the beneficial effects of the apparatus for data balancing of a distributed storage engine for redirection on write, the computer device, and the computer-readable storage medium, please refer to the description of the corresponding beneficial effects of the method for data balancing of a distributed storage engine for redirection on write, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 It is a flow chart of a method for data balancing of a distributed storage engine for redirection during writing according to an embodiment of the present invention;

[0053] Figure 2is a schematic diagram of object topology according to an embodiment of the present invention;

[0054] Figure 3 is a schematic diagram of a logical topology structure according to an embodiment of the present invention;

[0055] Figure 4 is a schematic diagram of disk selection according to an embodiment of the present invention;

[0056] Figure 5 is a schematic diagram of readjustment of check slices and data slices according to an embodiment of the present invention;

[0057] Figure 6 is a schematic diagram of adjusted check slices and data slices according to an embodiment of the present invention;

[0058] Figure 7 is a schematic diagram of a PG cycle according to an embodiment of the present invention;

[0059] Figure 8 is a schematic diagram of star-shaped reading and writing according to an embodiment of the present invention;

[0060] Fig. 9 is a schematic diagram of usage balancing according to an embodiment of the present invention;

[0061] Fig.10 is a schematic diagram of PG migration according to an embodiment of the present invention;

[0062] Fig.11 It is a structural block diagram of a distributed storage engine data balancing device for write-time redirection according to an embodiment of the present invention;

[0063] Fig.12 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0065] Currently, Ceph uses relatively coarse-grained management when balancing the cluster, so it cannot fully utilize the hardware capabilities when the cluster is actually running in production. It requires people who have a design and code-level understanding of Ceph to do operation and maintenance to alleviate the data balancing capabilities to a certain extent. In addition, due to problems with the Ceph architecture, some tuning, such as balancing the data sharding and check sharding of EC, cannot be done.

[0066] GFS management is more detailed, and it can balance cluster data from two aspects: space allocation balance and background balance adjustment. However, the master memory resource consumption is huge, and GFS is mainly a replica architecture. It is not designed for EC's data sharding and check sharding balance, so data balance still cannot reach a better level.

[0067] In view of this, according to an embodiment of the present invention, an embodiment of a method for write-time redirection of distributed storage engine data balancing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0068] In this embodiment, a method for balancing data in a distributed storage engine for redirection during write is provided, which can be executed by a cluster management center in a server, terminal, mobile terminal or other device. Figure 1 is a flow chart of a method for balancing data in a distributed storage engine by redirecting the data during write operation according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0069] Step S101, executing a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing a heterogeneous cluster physical topology structure, and each logical topology structure includes a plurality of logical isolation sub-topologies.

[0070] The physical topology reflects the mapping relationship between physical hardware servers and hard disks, and collects all server information in the cluster and the hard disk information managed by the server into the chunkmaster (chunkmaster means cluster management center, which is responsible for a series of cluster management functions such as data distribution, data balancing, fault recovery, cluster expansion and contraction, cluster monitoring, virtual space allocation and virtual space conversion to physical space in the distributed storage engine in this embodiment. Different storage engines have different naming methods, such as master and CM (cluster master)). Among them, the server information includes CPU, memory, number of hard disks, network bandwidth information, number of hard disks and hard disk information. Hard disk information includes: hard disk media (such as: TLC SSD, QLC+SLC SSD, HDD, etc.), hard disk protocol (SAS, SATA, NVMe), hard disk capacity, etc. Refer to the object topology diagram Figure 2As shown in the figure, the logical topology divides the same type of hardware into a logical topology pool according to the physical topology and actual reliability requirements, and a logical topology is a logically isolated fault domain, and computing resources, network resources and storage resources are logically isolated. In other words, servers of the same type are organized into logically isolated logical topologies, and one ultra-high performance logical topology is referred to as Figure 3 As shown, the logical topology structure may also include a high-performance logical topology, a common logical topology, a large IO capacity logical topology, and the like.

[0071] In different logical topologies, balanced disk selection is performed. Since the physical hardware characteristics in a topology are consistent, only the disk weight is needed as the disk selection weight basis to complete the disk selection of the logical topology, where the disk weight = capacity / PG number. After the logical topology disk selection, since the disk is randomly selected, it leads to the EC (Erasure Code, which means erasure code, a coding technology that adds m copies of data to n copies of original data and can restore the original data through any n copies of n+m copies. Mathematically, encoding is expressed as constructing a multivariate linear equation, decoding is solving the solution of the multivariate linear equation, and further abstraction is matrix operation. The current correction coding technology is mainly divided into two categories: one is the correction coding based on Galois field operation (such as RS); the other is the erasure code (LDPC) based on "XOR" operation) protection mode. From the perspective of the cluster, all hard disks are not necessarily balanced as EC check shards and EC data shards. Therefore, further balancing adjustments are required to make each hard disk equal to the number of times it is used as an EC check shard and EC data shard, so as to complete the disk selection balance strategy.

[0072] Step S102, after executing the preset disk selection balancing strategy, obtain the used capacity of the PG and the used capacity of the hard disk.

[0073] Step S103, executing a preset space allocation balancing strategy based on the used capacity of the PG and the used capacity of the hard disk.

[0074] In this embodiment, the cluster management center uniformly allocates virtual address chunks (chunk means large block, which is used to describe a byte stream space for additional writing allocated from PG (Protect Group, which means logical protection group, which is used to describe a logical protection group that is organized by a group of independent virtual hard disks using EC or replicas to provide reliability guarantee and provide read and write space to the outside) and is distinguished by a unique identifier chunkID. The large cycle prioritizes chunk allocation according to the principle of the same number of chunks allocated on each PG. For a cluster in normal use, the number of chunks is the same and the capacity is also the same. When the difference in the used capacity of the PG reported by the monitoring chunkserver (which means a single-machine storage server, which is used to describe a single-machine storage server in this embodiment, manages the storage space on a single-machine server, and manages the single-machine byte stream chunk space) exceeds 10G and lasts for more than 1 day, a "small cycle PG group" is established to catch up. When it is monitored that the difference in the used capacity of the hard disks reported by the stand-alone storage servers exceeds 10G and lasts for more than 2 days, a "small cycle PG group" is established for the PGs on these hard disks to catch up with the capacity, thereby implementing a space allocation balance strategy.

[0075] Step S104, after executing the preset space allocation balancing strategy, execute the preset usage balancing strategy based on the star read and write mode.

[0076] Reading and writing a chunk actually means reading and writing a group of independent virtual hard disks. The virtual hard disk only processes the reading and writing related to itself. The IO data aggregation and distribution are the responsibility of an external client independent of the virtual hard disk. This reading and writing method is called star reading and writing in this embodiment. In this embodiment, after the ROW writes data at the chunk position, the data is not allowed to be changed, and star reading and writing are supported. With the average distribution of the check disk, further balance can be achieved. That is, when heterogeneous servers appear in a cluster, servers with different structures are logically isolated through logical topology. When using space, different logical topologies are used according to the characteristics of the data type to achieve heterogeneous cluster balance.

[0077] Step S105: After executing the preset use balancing strategy, if the cluster still fails to reach a balance, execute the preset background balancing strategy.

[0078] Specifically, the cluster management center collects the statistical information of all PGs and hard disks in the cluster, and then identifies the hot PGs and hot hard disks in the cluster due to burst traffic based on the statistical information. If hot PGs and hot hard disks continue to appear, the hot PGs are migrated from hot hard disks to non-hot hard disks through PG migration, thereby achieving performance and capacity balance within the cluster.

[0079] In the related technology, for example, Google's publicly available distributed storage engine GFS (Google FileSystem) can provide an append-write file system. GFS manages the mapping metadata of files to spaces by a unified management node, and the data block information is queried by the management node on the specific data block server at startup. Data balance is achieved by the management node evenly allocating data blocks and directly migrating data blocks, resulting in a huge number of data blocks that the management node needs to manage when the GFS cluster is large, and because it is provided as a file to the outside, the memory usage of data blocks and file information will be large. Since there is no concept of multi-logical topology, it is complicated to handle when heterogeneous servers appear in the cluster. Another widely used in the industry is ceph, which is a distributed storage that uses a distributed algorithm (CRUSH algorithm) for data distribution. It can distribute data objects according to the weight of each storage device, so that the distribution is close to uniform distribution. Ceph obtains a set of OSD (Object Storage Device) lists for reading and writing through calculation. Since these OSD lists are calculated, it is difficult to ensure that the number of times each OSD appears in each position is consistent. Therefore, when the cluster performs EC reading normally, the OSD used as verification does not participate in the reading, resulting in unbalanced cluster reading. In addition, since the number of PGs (Placement Groups) in Ceph is limited and not as large as the number of objects in Ceph, the number of times each OSD serves as the PG master is different. Since Ceph reads and writes are all handled by the master OSD, the cluster is unbalanced. When the cluster is found to be unbalanced, it can be adjusted by adjusting the weights. However, after the weights are changed, data migration of multiple PGs will occur, making it difficult to achieve fine-grained control like a unified management center cluster.

[0080] The present invention proposes a hierarchical data balancing method applied to a distributed storage engine with redirection on write, which divides the data balancing of the cluster into four levels, namely disk selection balancing, space allocation balancing, usage balancing and background balancing adjustment, monitors the cluster hardware through a unified cluster management center, and realizes data balancing within the cluster by combining the ROW storage engine and the hierarchical storage architecture. When this method is applied in a distributed storage engine, it will evenly share the capacity pressure of the cluster on each storage medium as much as possible, and balance the performance pressure on each network, computing and storage hardware in the cluster, that is, the unified cluster management center finely controls the cluster to achieve balanced distribution of cluster data.

[0081] In some optional implementations, a preset disk selection balancing strategy is executed based on a logical topology structure, including:

[0082] Step a, determining the weight and fault domain of each topological node in the logical topology structure;

[0083] Step b, selecting the Nth layer in the logical topology structure, where N=1;

[0084] Step c, determining the first topological node in the maximum weight set in the Nth layer and the fault domain corresponding to the first topological node; the first topological node is any node in the maximum weight set;

[0085] Step d, when the fault domain corresponding to the first topological node is the set fault domain, determining whether the first topological node meets the fault requirement;

[0086] Step e1, when the fault requirements are met, select the first topological node, set N=N+1, and return to step c until N is the last layer in the logical topology structure, completing the disk selection of the logical topology;

[0087] Step e2, if the fault requirement is not met, return to step b and use the second largest weight set in the Nth layer as the maximum weight set.

[0088] Reference Figure 4 As shown, in this embodiment, a greedy disk selection method is adopted:

[0089] Step 1: Randomly break up the layers from the root layer downwards, and select logical topological points in order of weight;

[0090] Step 2: When selecting the lower layer, if you find that the logical topology point is a fault domain, decide whether to select it according to the fault domain requirements;

[0091] Step 3: After the cleanup fails to meet the fault domain requirements, return to step 1 and select the next weighted logical topology point.

[0092] Step 4: When the fault domain requirements are met, the layer is randomly broken down and the logical topology point with the largest weight is selected until the hard disk is selected.

[0093] In disk selection balancing, hardware of the same type is organized into a logical topology. Large heterogeneous storage server clusters are managed through logical topology isolation and storage federation management, which improves the balance of heterogeneous clusters to usage balance. Disk selection balancing only considers the balance of hard disks of the same type. The number of PGs available in the cluster is determined by the disk selection within the storage cluster, which can ensure that each hard disk appears the same number of times in each position in the PG, and even in the EC case, the cluster read and write balance can be guaranteed as much as possible.

[0094] In some optional implementations, after completing the disk selection of the logical topology, the method further includes:

[0095] Rearrange the check shards and data shards of all PGs in the cluster so that each hard disk in the logical topology is used as an EC check shard and an EC data shard an equal number of times.

[0096] The storage field has found that EC protection groups can provide higher read concurrency and higher hard disk utilization than replicas (three replicas <33%, EC 4+2 can provide >60% utilization). The limitation that writing must be done with EC full stripes has been gradually resolved by the industry, and distributed storage engines using EC protection groups have begun to become the mainstream in the distributed storage industry. EC data shards support cluster reads, while check shards only support repair reads. When the check shards and data shards are unevenly distributed on the hard disks, when the cluster is normal, the hard disk read pressure distribution in the cluster is uneven. The cluster is in a normal state for more than 99.99% of the time, and for example, database OLTP business is generally 70% read, and OLAP is generally 90% read. Read balance can significantly improve cluster throughput. For example, the distributed storage engine used in related technologies, such as ceph, uses random object names and consistent hash calculations to a group of storage resources, and the IO flow is first concentrated on the master of this group of storage resources. Although from the perspective of mathematical probability, in the case of complete randomness and a large enough number of objects, the capacity and performance in the cluster are balanced, but the actual business is not a perfect mathematical model. The EC protection mode widely used in the storage industry is divided into data sharding and checksum sharding. When the cluster is normal, only the data sharding supports reading. When the cluster is abnormal, the checksum sharding is used to repair the reading. The cluster is normal for more than 99.999% of the time, which causes unbalanced reading. The number of PGs in Ceph is not huge, and mathematically it cannot guarantee that each medium is balanced for checksums, so the reading cannot be balanced.

[0097] Therefore, in this embodiment, reference Figure 5 as well as Figure 6 As shown in the figure, the parity shards and data shards of all PGs in the cluster are re-adjusted so that all hard disks appear the same number of times in each position in the PG, thereby ensuring that reads can be balanced to all network, computing, and storage resources in the cluster. In the EC scenario, all disks are evenly distributed on the parity shards and data shards of the EC from the perspective of the logical topology cluster, thereby effectively improving the performance balance of the cluster.

[0098] In this embodiment, the erasure coded hard disks are selected to ensure that the number of times the hard disks in the cluster appear in all erasure code protection groups is the same, thereby achieving balanced reading and writing of erasure codes. This method can effectively prevent uneven hard disk load and improve the performance and reliability of the overall storage system.

[0099] In some optional implementations, based on the used capacity of the PG and the used capacity of the hard disk, a preset space allocation balancing strategy is executed, including:

[0100] Determine whether the difference in used capacity of the PG exceeds a first preset capacity;

[0101] In the case where the difference in the used capacity of the PG exceeds the first preset capacity, determining whether the time for which the first preset capacity is exceeded exceeds the first preset time;

[0102] In the case where the time of exceeding the first preset capacity exceeds the first preset time, establishing a first cycle PG group, the first cycle PG group being used to balance the used capacity of the PGs;

[0103] Determine whether the difference in used capacity of the hard disk exceeds a second preset capacity;

[0104] When the difference in the used capacity of the PG exceeds the second preset capacity, determining whether the time for exceeding the second preset capacity exceeds the second preset time;

[0105] When the time exceeding the second preset capacity exceeds the second preset time, a second cycle PG group is established on the hard disk that exceeds the second preset capacity and exceeds the second preset time, and the second cycle PG group is used to balance the used capacity of the hard disk.

[0106] Specifically, a unified cluster management center allocates virtual address chunks, which are distinguished by a unique identifier chunkID. The cluster management center allocates chunks with adjacent chunkIDs, and the hard disk groups contained in the PGs basically do not overlap. There will be no hard disk hot spots when the chunks are used in a normal order. The large cycle prioritizes chunk allocation based on the principle of the same number of chunks allocated on each PG. For a cluster in normal use, the number of chunks is the same and the capacity is also the same. When the cluster management center monitors that the difference in the used capacity of PG reported by a single storage server exceeds 10G and lasts for more than 1 day, it starts to establish a "small cycle PG group" to catch up. When it is monitored that the difference in the used capacity of the hard disks reported by a single storage server exceeds 10G and lasts for more than 2 days, a "small cycle PG group" is established for the PGs on these hard disks to catch up. For the specific example processing process, please refer to Figure 7 As shown, the effective PG allocates chunks in a "round-robin manner". When the PG capacity is inconsistent, 90% of the requested chunks are used for a large loop, and 10% are used for a small loop to catch up with the PG capacity. The loop range is uniformly controlled by the cluster management center. Furthermore, the cluster management center monitors the available capacity and available percentage of the hard disk. When the hard disk capacity is uneven, the relevant PG is isolated from the large loop and used for a small loop to try to ensure that the chunks allocated by the cluster management center are writable.

[0107] In this embodiment, the cluster space allocation is uniformly balanced by the cluster management center based on the statistical information in the cluster. Chunks are allocated through a "large loop" to make the number of chunks in each PG consistent. When the failure is restored or expanded after a period of time, the number of chunks is added through a small loop while the large loop is allocated. When the number of chunks is consistent but the capacity imbalance lasts for more than 1 day, the PG with low capacity is individually subjected to a small loop to add the number of chunks to achieve balance.

[0108] In some optional implementations, based on the star-shaped reading and writing mode, a preset usage balancing strategy is executed, including:

[0109] Get the IO data of each hard disk in star read and write mode;

[0110] Based on IO data, determine the data types with different heat levels;

[0111] Data types of different temperatures are allocated to corresponding logical topology structures according to the types of logical topology structures.

[0112] Reference Figure 8 As shown, in this embodiment, star write replaces chain write to make the read and write in the cluster more balanced. When heterogeneous servers appear in a cluster, logical topology is used to logically isolate servers of different structures. When using space, IO type analysis is performed and different topologies are used according to the characteristics of the business to achieve heterogeneous cluster balance. For example, after long-term statistics, hot data is placed in high-performance topology through gc (Garbage Collection, an automatic memory management mechanism), and cold data is placed in ordinary performance topology. Fig. 9 shown.

[0113] By organizing the same type of hardware into different topologies, a heterogeneous storage cluster federation is used to manage large storage heterogeneous clusters. Different business types and different hardware characteristics balance data to different federated clusters to achieve cluster performance balance. For example, in the public cloud, the scheduled massive IO model does not require high IO performance, but affects the performance of other cluster businesses. Through the sandbox cluster, the business is gradually guided to the sandbox cluster without being aware of the business. After a long period of hot and cold analysis of IO, the hot and cold quantity can be separated to achieve overall cluster performance balance.

[0114] In this embodiment, star-shaped reading and writing are adopted, and there is no master responsible for forwarding reading and writing. In the scenario of large-scale reading, the master will not become a reading hotspot. When heterogeneous servers appear in a cluster, the logical topology is used to logically isolate servers of different structures. When using space, different topologies are used according to the characteristics of the business to achieve heterogeneous cluster balance, which can further effectively optimize data balance.

[0115] In some optional implementations, executing a preset background balancing strategy includes:

[0116] Get the statistics of all PGs and all hard disks in the cluster;

[0117] Based on the statistics of all PGs, determine the hot PGs;

[0118] Determine hot disks based on the statistics of all disks;

[0119] When the hotspot PG and the hotspot hard disk are in conflict for a third preset period of time, the hotspot PG is migrated from the hotspot hard disk to the non-hotspot hard disk through PG migration.

[0120] If the data balance of the above three layers still fails to achieve cluster balance, the cluster management center will collect cluster information. If the cluster has hard disk hot spots and PG hot spots for a long time, and the hard disk capacity and PG capacity are inconsistent, the cluster management center can perform PG migration. Fig.10 As shown, the hot PG is migrated to the non-hot hard disk. Since the number of PGs to be migrated and the migration bandwidth can be set, the impact on the entire cluster is small and controllable, and data balance can be further effectively improved, where the authoritative shard represents reliable data.

[0121] The cluster management center monitors the information of each component of the cluster. For example, if it is found that some hard disks and some PGs on the hard disks will have performance hotspots at a fixed interval of 1 hour, while other hard disks and PGs will not have this situation, the hotspot PG will be migrated to other hard disks without hotspots, and the non-hotspot PG will be migrated back. This design can also solve the problem of scheduled hotspot PG on the public cloud to a certain extent.

[0122] In this embodiment, a device for data balancing of a distributed storage engine for redirection during writing is also provided, and the device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0123] This embodiment provides a distributed storage engine data balancing device for redirecting writes, such as Fig.11 As shown, the device comprises:

[0124] A balancing strategy execution module 201 is used to execute a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing the physical topology structure of the heterogeneous cluster, and each logical topology structure includes a plurality of logical isolation sub-topologies;

[0125] The space allocation balancing strategy execution module 202 is used to obtain the used capacity of the PG and the used capacity of the hard disk after executing the preset disk selection balancing strategy; based on the used capacity of the PG and the used capacity of the hard disk, execute the preset space allocation balancing strategy;

[0126] A usage balancing strategy execution module 203 is used to execute a preset usage balancing strategy based on a star-shaped reading and writing method after executing the preset space allocation balancing strategy;

[0127] The background balancing strategy execution module 204 is used to execute the preset background balancing strategy when the cluster still fails to reach a balance after executing the preset use balancing strategy.

[0128] The distributed storage engine data balancing device for write-time redirection in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0129] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0130] The embodiment of the present invention also provides a computer device having the above Fig.11 The data balancing device shown is used for redirecting distributed storage engines during write.

[0131] See also Fig.12 , Fig.12 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig.12 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.12 A processor 10 is taken as an example.

[0132] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0133] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0134] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0135] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0136] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0137] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0138] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0139] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for balancing data in a distributed storage engine for redirection during write operation, characterized in that: The method comprises: Executing a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing a heterogeneous cluster physical topology structure, and each of the logical topologies includes a plurality of logical isolation sub-topologies; After executing the preset disk selection balancing strategy, obtaining the used capacity of the PG and the used capacity of the hard disk; Based on the used capacity of the PG and the used capacity of the hard disk, a preset space allocation balancing strategy is executed; After executing the preset space allocation balancing strategy, based on the star-shaped reading and writing mode, executing the preset usage balancing strategy; After executing the preset balancing strategy, if the cluster still fails to reach a balance, the preset background balancing strategy is executed.

2. The method according to claim 1, characterized in that The executing of a preset disk selection balancing strategy based on the logical topology structure includes: Step a, determining the weight and fault domain of each topological node in the logical topological structure; Step b, selecting the Nth layer in the logical topology structure, where N=1; Step c, determining a first topological node in the maximum weight set in the Nth layer and a fault domain corresponding to the first topological node; the first topological node is any node in the maximum weight set; Step d: when the fault domain corresponding to the first topological node is a set fault domain, determining whether the first topological node meets the fault requirement; Step e1, when the fault requirement is met, the first topology node is selected, and N=N+1 is set, and the process returns to step c until N is the last layer in the logical topology structure, and the selection of the logical topology is completed; Step e2, when the fault requirement is not met, return to step b, and use the second largest weight set in the Nth layer as the maximum weight set.

3. The method according to claim 2, characterized in that After completing the disk selection of the logical topology, the following steps are also required: The check shards and data shards of all PGs in the cluster are re-adjusted so that each hard disk in the logical topology is used as an EC check shard and an EC data shard an equal number of times.

4. The method according to claim 1, characterized in that: The executing of a preset space allocation balancing strategy based on the used capacity of the PG and the used capacity of the hard disk includes: Determine whether the difference in used capacity of the PG exceeds a first preset capacity; In the case where the difference in used capacity of the PG exceeds the first preset capacity, determining whether the time for exceeding the first preset capacity exceeds a first preset time; In the case where the time of exceeding the first preset capacity exceeds the first preset time, establishing a first cycle PG group, the first cycle PG group being used to balance the used capacity of the PGs; Determine whether the difference in used capacity of the hard disk exceeds a second preset capacity; When the difference in the used capacity of the PG exceeds the second preset capacity, determining whether the time for exceeding the second preset capacity exceeds a second preset time; When the time exceeding the second preset capacity exceeds the second preset time, a second circular PG group is established on the hard disk that exceeds the second preset capacity and exceeds the second preset time, and the second circular PG group is used to balance the used capacity of the hard disk.

5. The method according to claim 1, characterized in that The star-shaped read-write method is used to execute a preset usage balancing strategy, including: Obtaining IO data of each hard disk in the star-shaped reading and writing mode; Based on the IO data, determine data types with different heat levels; The data types of different temperatures are allocated to the corresponding logical topology structures according to the types of the logical topology structures.

6. The method according to claim 1, characterized in that The execution of the preset background balancing strategy includes: Get the statistics of all PGs and all hard disks in the cluster; Based on the statistical information of all the PGs, determine the hotspot PG; Based on the statistical information of all the hard disks, determine the hot hard disk; When the hotspot PG and the hotspot hard disk are continuously in conflict for a third preset time, the hotspot PG is migrated from the hotspot hard disk to a non-hotspot hard disk through PG migration.

7. A data balancing device for distributed storage engine redirection during write, characterized in that: The device comprises: A balancing strategy execution module, used to execute a preset disk selection balancing strategy based on a logical topology structure, wherein the logical topology structure is one of a plurality of different types of logical topologies determined after reorganizing the physical topology structure of the heterogeneous cluster, and each of the logical topologies includes a plurality of logical isolation sub-topologies; A space allocation balancing strategy execution module, used to obtain the used capacity of the PG and the used capacity of the hard disk after executing the preset disk selection balancing strategy; based on the used capacity of the PG and the used capacity of the hard disk, execute the preset space allocation balancing strategy; A usage balancing strategy execution module is used to execute a preset usage balancing strategy based on a star-shaped reading and writing method after executing the preset space allocation balancing strategy; The background balancing strategy execution module is used to execute the preset background balancing strategy when the cluster still fails to reach balance after executing the preset use balancing strategy.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data balancing method for write-time redirection distributed storage engine as described in any one of claims 1-6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the data balancing method for write-time redirection distributed storage engine as described in any one of claims 1-6.

10. A computer program product, characterized in that It includes computer instructions, which are used to cause a computer to execute the data balancing method for write-time redirection distributed storage engine described in any one of claims 1-6.

Citation Information

Patent Citations

  • Resource allocation method and device, monitor and machine readable storage medium

    CN110515724A

  • Distributed storage system data balancing method and related device

    CN112015708A

  • Distributed storage collocation group PG balancing method and device, equipment and medium

    CN117389465A

  • Optimization method for data balance in distributed storage system

    CN117850680A

Cited By

  • A data storage strategy conversion system and method for Ceph distributed storage system

    CN122672726A

  • Data balancing method and apparatus for redirect-on-write distributed storage engine

    WO2026118920A1