Capacity expansion method and device based on Multi-Raft distributed storage system
By building a multi-Group Multi-Raft storage system and recording Group information when data is written, the problem of limited expansion in the prior art is solved, performance and capacity expansion is achieved, and business performance impact is avoided.
Patent Information
- Application Number
- CN202510031579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The existing distributed storage system based on Multi-Raft is limited by the initial configuration when expanding capacity, and cannot achieve on-demand expansion of performance and capacity. The expansion process is likely to affect business performance and lead to operation and maintenance accidents.
By building a multi-group Multi-Raft storage system, each group of Raft nodes is a group, adding a group to achieve capacity expansion, and recording Group information into metadata when data is written, and obtaining Group information from metadata when reading to read to avoid data migration.
It realizes expansion without affecting existing services, avoids data distribution abnormalities and business performance impacts, and achieves simultaneous expansion of performance and capacity.
Smart Images

Figure CN119440422B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and in particular relates to a capacity expansion method and device based on a Multi-Raft distributed storage system. Background Art
[0002] With the rapid development of the Internet, storage systems have gradually evolved from single entities to distributed systems. Data consistency has become one of the important factors to ensure the reliability and performance of storage systems. The Raft algorithm is widely used to ensure the consistency and high availability of multi-copy data. In distributed storage systems, by sharding data, each shard corresponds to a Raft group, and the Multi-Raft algorithm can effectively improve the performance and scalability of the system.
[0003] The current distributed storage system based on Multi-Raft is limited by the initial system configuration. The number of Raft groups and system resource configuration are usually difficult to modify after the system is initialized. As a result, the distributed storage system based on Multi-Raft can only store a small amount of data persistently, and cannot achieve on-demand expansion of performance and capacity, and cannot be applied to large-scale storage system scenarios.
[0004] Common expansion methods are to migrate Raft replicas to new nodes or vertically expand the capacity of a single Raft replica. However, the above expansion methods have shortcomings, as follows:
[0005] 1. Capacity expansion is achieved by migrating Raft replicas to new nodes. However, due to the number of initially configured Raft groups, unlimited capacity expansion is not possible. Migrating Raft replicas requires changing Raft members. Since Raft performs operations by reaching a consensus among multiple members, changes in Raft members can easily affect existing businesses and cause operation and maintenance accidents.
[0006] 2. There is a capacity limit for expanding the capacity of a single Raft replica by vertically expanding the capacity. All Raft replicas need to be shut down one by one for capacity expansion, which is likely to affect business performance. Summary of the invention
[0007] The object of the present invention is to provide a capacity expansion method and device based on a Multi-Raft distributed storage system, which realizes capacity expansion by combining Raft group partitioning and Raft group scheduling logic. No data migration is required during the capacity expansion process, and business performance is not affected.
[0008] In order to achieve the above object, the technical solution of the present invention is as follows:
[0009] A capacity expansion method based on a Multi-Raft distributed storage system, comprising:
[0010] S1. Build a multi-group Multi-Raft storage system. Each group of Raft nodes is a group. Each group of Raft nodes is divided into multiple sub-groups. Each sub-group includes a leader node and multiple follower nodes. Add Group when capacity expansion is needed.
[0011] S2. When writing data, the corresponding Group information is recorded in the metadata. When reading data, the Group information is first obtained from the metadata information and the data is read from the corresponding Group.
[0012] Furthermore, in step S2, when the number of groups is greater than 1 during data writing, data distribution logic is executed according to data distribution requirements.
[0013] Furthermore, it also includes: when the original Group data is full and a new Group is expanded, the data is migrated to the new Group through data copying.
[0014] Furthermore, it also includes: during the expansion process, the newly added group is not reflected in the scheduling rules. After the new group is ready, data reading and writing are sent to the new group. If the entire group goes offline together, all data in the group is migrated to other groups, and then the group that needs to be offline is disabled to complete the offline of the faulty group.
[0015] Another aspect of the present invention further provides a capacity expansion device based on a Multi-Raft distributed storage system, comprising:
[0016] Multi-Group module: Build a multi-group Multi-Raft storage system. Each group of Raft nodes is a group. Each group of Raft nodes is divided into multiple sub-groups. Each sub-group includes a leader node and multiple follower nodes. Add Group when expansion is needed.
[0017] Metadata module: When writing data, the corresponding Group information is recorded in the metadata. When reading data, the Group information is first obtained from the metadata information and the data is read from the corresponding Group.
[0018] Furthermore, in the metadata module, when the number of groups is greater than 1 during data writing, the data distribution logic is executed according to the data distribution requirements.
[0019] Furthermore, it also includes a migration module: when the original Group data is full and a new Group is expanded, the data is migrated to the new Group through data copying.
[0020] Furthermore, it also includes a guarantee module: during the expansion process, the newly added group is not reflected in the scheduling rules. After the new group is ready, data reading and writing are sent to the new group. If the entire group goes offline together, all data in the group is migrated to other groups, and then the group that needs to be offline is disabled to complete the offline of the faulty group.
[0021] The present invention also provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned capacity expansion method based on the Multi-Raft distributed storage system.
[0022] The present invention also provides a computer program product, including a computer program, which implements the above-mentioned capacity expansion method based on the Multi-Raft distributed storage system when executed by a processor.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] (1) The distributed storage system expansion method based on Multi-Raft proposed in the present invention combines Raft group partitioning and Raft group scheduling logic, and deploys a complete Raft Group to join the existing storage system, so that the expansion does not affect the existing business.
[0025] (2) The implementation method of the present invention is simple and efficient. When writing data, the group information is recorded in the metadata information, and when reading data, the group information is obtained from the metadata information, thereby avoiding data distribution anomalies caused during the expansion process and solving the problem of data migration required for expansion at the lowest cost. In addition, the new group can share the read and write requests from the client, thereby achieving simultaneous expansion of performance and capacity. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 The figure is a schematic diagram of the capacity expansion process of the Multi-Raft distributed storage system according to Embodiment 1 of the present invention.
[0027] Figure 2 This is a schematic diagram of multi-Group deployment in Example 1 of the present invention.
[0028] Figure 3 This is a schematic diagram of the data writing process based on the Multi-Raft distributed storage system according to Example 1 of the present invention.
[0029] Figure 4 This is a schematic diagram of the data reading process of the Multi-Raft distributed storage system according to Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0031] The distributed storage system expansion method based on Multi-Raft proposed in this invention has four main parts:
[0032] 1. Node expansion:
[0033] In the prior art, the deployment of Multi-Raft storage system is as follows: Figure 1 As shown in the structure in the upper part, several Raft nodes (in this embodiment, the number of Raft nodes is 3, as usual, for cost and consistency considerations) are divided into multiple subgroups, each of which includes a leader node (Leader) and multiple follower nodes; if capacity expansion is required, it can only be achieved by migrating Raft copies to new nodes or vertically expanding the capacity of a single Raft copy.
[0034] The present invention regards the above Raft nodes as a group, called Group; when capacity expansion is required, add Group; Figure 1 As shown in the lower part, the original group of Raft nodes is used as Group 0. If expansion is required, another group of Raft nodes is directly added, that is, Figure 1 Group1 in.
[0035] The storage capacity and overall bandwidth of a single group are limited by the node where it is located. After expanding to provide services in the form of multiple groups, multiple groups can not only share the read and write pressure from the client, but also achieve unlimited capacity expansion while ensuring that availability is not reduced.
[0036] 2. Data distribution:
[0037] For the above-mentioned multi-group system, data distribution usually adopts certain rules, such as hashing or averaging to different groups. However, this method will cause the original data distribution method to be abnormal when a new group is added, and the data needs to be redistributed. The use of methods such as consistent hashing cannot ensure that data migration is not performed. Once data migration is required, it will inevitably affect business performance. The present invention records the Group information in the metadata when writing data, obtains the Group information from the metadata information when reading data, and reads the data from the corresponding Group. Since the acquisition of metadata is essential when reading data, the recording and reading of Group information will not affect the data reading and writing performance. This method avoids the problem that when a new group is added, the data distribution rules become invalid, the data cannot be read normally, and the data needs to be migrated.
[0038] A single group does not need to execute data distribution logic, and data reading and writing can use the same group. When the number of groups is greater than 1, different scenarios may have different data distribution requirements, such as:
[0039] (1) When multiple groups have a large amount of remaining capacity, it is appropriate to evenly distribute data to multiple groups.
[0040] (2) When the remaining space of the original group is insufficient and a new group is expanded, it is appropriate to: allocate most of the data to the new group;
[0041] (3) When the performance of the groups is different, such as when the disks are HDDs and SSDs, it is suitable to: distribute data to different groups on demand, such as storing large files in the HDD group and small files in the SSD group;.
[0042] 3. Data Migration:
[0043] The present invention achieves capacity expansion by adding a new group, and reads data to determine the group information through the records in the metadata. In theory, there is no need for data migration; if a new group is expanded when the original group data is full, the data can be re-migrated to the new group through the data copy process. This data migration method has no impact on business performance, and some low-frequency access data can be migrated when the business is relatively idle.
[0044] 4. High availability guarantee:
[0045] The data stored in the group is highly available through Raft replicas, and data reading and writing are not affected by single point failures.
[0046] During the expansion process, since the newly added group has not yet been reflected in the scheduling rules, no data reading and writing will be sent to the new group. After the new group is ready, the client can perceive that the new group is ready and send data reading and writing to the new group.
[0047] When a group fails for some reason and the entire group needs to be taken offline, you can migrate all data in the group to other groups and then disable the group that needs to be taken offline to complete the offline of the failed group without affecting the application.
[0048] Based on the above general idea, the present invention is described in detail below in combination with specific embodiments and drawings.
[0049] Embodiment 1:
[0050] This embodiment proposes specific method steps for capacity expansion based on the Multi-Raft distributed storage system, such as Figure 2 As shown, including:
[0051] Step S101: Prepare a new idle node, and deploy and initialize a new Raft Group (hereinafter referred to as the new Group) on the idle node in the same way as the existing Raft Group. Figure 1 As shown, the existing Raf nodes are regarded as Group 0, and a new Group (Group 1) is deployed and initialized on the idle nodes in the same way as Group 0.
[0052] Step S102: Add new Group configuration information to the metadata for service startup and client identification of Group.
[0053] Step S103: Start the service process on the new Group. At this time, the new Group has started the service according to the configuration, but has not yet started to provide the service.
[0054] Step S104: After the new Group is ready, the new Group enable flag in the metadata is set to True, and the new Group can be perceived by the client.
[0055] Step S105: The client detects that a new group is online, and records the group in the metadata when writing. When reading data, the group is obtained from the metadata information and the data is read from the corresponding group. Since the acquisition of metadata is essential when reading data, the recording and reading of the group will not affect the data reading and writing performance. This method avoids the problem that when a new group is added, the data distribution rules become invalid, the data cannot be read normally, and the data needs to be migrated. Since the user has identified the group information in the metadata when writing data, the group can be read from the metadata information when reading data, so there is no need to migrate data due to changes in data distribution when a new group is online.
[0056] Step S106: The client data read and write is scheduled to the new Group. When the client writes data, Figure 3 As shown in the figure, the group to be written is automatically determined based on the enabled group information using the configured data distribution rule, such as configuring an average distribution rule, for example: targetGroupID = inode % TotalGroupCount. Then the group information is identified in the metadata; when reading data, Figure 4 As shown, according to the Group information recorded in the metadata, read the data from the corresponding Group.
[0057] In this embodiment, a new Raft Group is deployed and initialized on the node to be expanded. The new Group is connected to the existing storage system as an independent service. The expansion of the storage system does not affect the business, and the new Group can share the read and write requests from the client, thereby achieving simultaneous expansion of performance and capacity.
[0058] In this embodiment, the group information is added to the metadata information when writing data, and the data distribution is learned through the group recorded in the metadata information when reading, thereby avoiding the impact of data migration caused by the addition of a new group on business performance.
[0059] Embodiment 2:
[0060] This embodiment proposes a capacity expansion device based on a Multi-Raft distributed storage system, including:
[0061] Multi-Group module: Build a multi-group Multi-Raft storage system. Each group of Raft nodes is a group. Each group of Raft nodes is divided into multiple sub-groups. Each sub-group includes a leader node and multiple follower nodes. Add Group when expansion is needed.
[0062] Metadata module: When writing data, the corresponding Group information is recorded in the metadata. When reading data, the Group information is first obtained from the metadata information and the data is read from the corresponding Group.
[0063] In the metadata module, when data is written and the number of groups is greater than 1, the data distribution logic is executed according to the data distribution requirements.
[0064] In addition, it also includes a migration module: when the original Group data is full and a new Group is expanded, the data is migrated to the new Group through data copying.
[0065] It also includes a guarantee module: during the expansion process, the newly added group is not reflected in the scheduling rules. After the new group is ready, data reading and writing are sent to the new group. If the entire group goes offline together, all data in the group is migrated to other groups, and then the group that needs to be offline is disabled to complete the offline of the faulty group.
[0066] The capacity expansion device based on the Multi-Raft distributed storage system proposed in this embodiment can implement the capacity expansion method based on the Multi-Raft distributed storage system proposed in Example 1, and has the same technical effect as Example 1.
[0067] The above-mentioned embodiments are only preferred implementations of the present invention, and are only used to help understand the method and core ideas of the present application. The protection scope of the present invention is not limited to the above-mentioned embodiments. All technical solutions under the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A capacity expansion method based on Multi-Raft distributed storage system, characterized in that: include: S1. Build a multi-group Multi-Raft storage system. Each group of Raft nodes is a group. Each group of Raft nodes is divided into multiple sub-groups. Each sub-group includes a leader node and multiple follower nodes. Add Group when capacity expansion is needed. S2. When writing data, the corresponding group information is recorded in the metadata. When reading data, the group information is first obtained from the metadata information and the data is read from the corresponding group to avoid the need to migrate data when adding a new group. If the original Group data is full, copy the data to the new Group when expanding the capacity of the new Group.
2. The capacity expansion method based on the Multi-Raft distributed storage system according to claim 1 is characterized in that: In step S2, when the number of groups is greater than 1 during data writing, the data distribution logic is executed according to the data distribution requirements.
3. The capacity expansion method based on the Multi-Raft distributed storage system according to claim 1 is characterized in that: Also includes: During the expansion process, the newly added group is not reflected in the scheduling rules. After the new group is ready, data reading and writing are sent to the new group. If the entire group goes offline together, all data in the group is migrated to other groups, and then the group that needs to be offline is disabled to complete the offline of the faulty group.
4. A capacity expansion device based on a Multi-Raft distributed storage system, characterized in that: include: Multi-Group module: Build a multi-group Multi-Raft storage system. Each group of Raft nodes is a group. Each group of Raft nodes is divided into multiple sub-groups. Each sub-group includes a leader node and multiple follower nodes. Add Group when expansion is needed. Metadata module: When writing data, the corresponding group information is recorded in the metadata. When reading data, the group information is first obtained from the metadata information and the data is read from the corresponding group, avoiding the need to migrate data when adding a new group. If the original Group data is full, copy the data to the new Group when expanding the capacity of the new Group.
5. The capacity expansion device based on the Multi-Raft distributed storage system according to claim 4 is characterized in that: In the metadata module, when data is written and the number of groups is greater than 1, the data distribution logic is executed according to the data distribution requirements.
6. The capacity expansion device based on the Multi-Raft distributed storage system according to claim 4, characterized in that: It also includes a guarantee module: during the expansion process, the newly added group is not reflected in the scheduling rules. After the new group is ready, data reading and writing are sent to the new group. If the entire group goes offline together, all data in the group is migrated to other groups, and then the group that needs to be offline is disabled to complete the offline of the faulty group.
7. A computer-readable storage medium storing a computer program, characterized in that: The computer program is used to execute the capacity expansion method based on the Multi-Raft distributed storage system as described in any one of claims 1 to 3.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the capacity expansion method based on the Multi-Raft distributed storage system is implemented as described in any one of claims 1 to 3.