Database cluster management method and database cluster management system

Through the automated slot migration method, the slot sets are divided into parallel migration according to the data volume and node number, which solves the problems of low resource utilization efficiency and load imbalance caused by manual operations, and realizes efficient database cluster management.

CN120277055APending Publication Date: 2025-07-08SHENZHEN QIANHAI EVOC ASIA-PACIFIC ELECTRONIC EQUIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510412540.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the slot migration process of database clusters relies on manual operations, resulting in low resource utilization efficiency, unbalanced load and consuming a lot of manpower, making it difficult to ensure database performance during capacity expansion or reduction.

Method used

Through the automated slot migration method, the slot sets are divided and migration tasks are performed in parallel according to the data volume of the target node and the number of nodes after shrinking or expanding capacity to ensure the balanced distribution of the data volume and realize the automated migration and load balancing of the slots.

Benefits of technology

It improves the efficiency of slot migration and the performance of database clusters, reduces labor costs, and ensures load balancing and performance of database clusters after shrinking or expanding capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277055A_ABST
    Figure CN120277055A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a database cluster management method and a database cluster management system, and relates to the technical field of information. The method comprises the following steps: determining a target node in M nodes of a target database cluster, wherein the target database cluster is a database cluster to be subjected to capacity reduction; n slot position sets are determined according to the data volume of each slot position in the plurality of slot positions mapped to the target node and the target migration data volume, N = M-1, each slot position set comprises at least one slot position in the plurality of slot positions, and the target migration data volume is equal to the data volume of the target node divided by N, the data volume of each slot set is determined according to the target migration data volume; and migrating the N slot position sets to N first nodes in the target database cluster, wherein the first nodes are nodes except the target node in the M nodes. According to the scheme, slot automatic migration can be achieved, load balance can be guaranteed as much as possible, and the performance of a capacity-reduced database cluster is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular, to a method for managing a database cluster and a database cluster management system. Background Art

[0002] To cope with the explosive growth of enterprise data volume and usage concurrency, a database cluster is introduced for storing specific data to improve concurrent operations and response speed. With the change of business scenarios, the resource usage requirements change accordingly. To utilize resources more efficiently, it is necessary to expand or contract the database cluster. During the process of expansion or contraction, slot migration is required. Usually, slot migration is implemented manually, which consumes a large amount of human resources, has a complex implementation process, and may cause the business to be inaccessible due to unreasonable manual operations. Summary of the Invention

[0003] Embodiments of this application provide a method for managing a database cluster and a database cluster management system, which can realize automatic slot migration, and can ensure load balancing as much as possible and ensure the performance of the database cluster after contraction.

[0004] In a first aspect, a method for managing a database cluster is provided, including: determining a target node among M nodes of a target database cluster, where the target database cluster is a database cluster to be contracted, the target node is a node to be subject to slot migration, and M>3; determining N slot sets according to the data volume of each slot among the multiple slots mapped to the target node and a target migration data volume, where N = M - 1, and each of the slot sets includes at least one of the multiple slots, the target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume; migrating the N slot sets to N first nodes in the target database cluster, where the first nodes are the nodes other than the target node among the M nodes, and the N slot sets correspond to the N first nodes one by one.

[0005] In a possible implementation manner, the migrating the N slot sets to N first nodes in the target database cluster includes: creating N migration tasks, where each migration task is used to migrate one slot set to one of the first nodes; and executing the N migration tasks in parallel.

[0006] In a possible implementation manner, the method further includes: sending a first message to the target database cluster, where the first message instructs the target database cluster to delete the target node, and the first message includes the connection address of the target node and the identifier of the target node.

[0007] In a possible implementation, the method further includes: when the migration of the N slot sets is completed and the target node is deleted, initializing the system of the target node, and marking the status of the target node as resource available after the system initialization is completed.

[0008] In a possible implementation, the target node is the node with the lowest average resource utilization rate within a preset duration among the M nodes.

[0009] In a second aspect, a database cluster management method is provided, including: adding a target node to a target database cluster, where the target database cluster is a database cluster to be expanded; for each first node, determining the target migration data volume of the first node according to the average data volume and the data volume of the first node, where the first node is a node other than the target node in the database cluster, and the average data volume is equal to the data volume of the target database cluster divided by the number of nodes in the target database cluster after expansion; for each of the first nodes, determining at least one slot to be migrated in the slots mapped to the first node according to the target migration data volume of the first node; for each of the first nodes, migrating the at least one slot to be migrated mapped to the first node to the target node.

[0010] In a possible implementation, the step of migrating the at least one slot to be migrated mapped to each of the first nodes to the target node includes: creating M migration tasks, where each migration task is used to migrate the at least one slot to be migrated of one of the first nodes to the target node, and M is the number of the first nodes; and executing the M migration tasks in parallel.

[0011] In a third aspect, an embodiment of the present application provides a database cluster management device, including units for executing each step in the method of the first aspect or any possible implementation manner in the first aspect, or including units for executing each step in the method of the second aspect or any possible implementation manner in the second aspect

[0012] In a fourth aspect, an embodiment of the present application provides a database cluster management device, including a processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the database cluster management system executes the method of the first aspect or any possible implementation manner in the first aspect or executes the method of the second aspect or any possible implementation manner in the second aspect.

[0013] In a fifth aspect, a data center is provided, where the data center includes a database cluster management system and a target database cluster.

[0014] In a possible implementation, the database cluster management system is configured to: determine a target node among M nodes of the target database cluster, where the target database cluster is a database cluster to be scaled down, the target node is a node to perform slot migration, and M > 3; determine N slot sets according to the data volume of each slot among the multiple slots mapped to the target node and the target migration data volume, where N = M - 1, and each of the slot sets includes at least one slot among the multiple slots, the target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume; migrate the N slot sets to N first nodes in the target database cluster, where the first nodes are the nodes other than the target node among the M nodes, and the N slot sets correspond to the N first nodes one by one.

[0015] In another possible implementation, the database cluster management system is configured to: add a target node to a target database cluster, where the target database cluster is a database cluster to be scaled up;

[0016] For each first node, determine the target migration data volume of the first node according to the average data volume and the data volume of the first node, where the first node is a node other than the target node in the database cluster, and the average data volume is equal to the data volume of the target database cluster divided by the number of nodes in the target database cluster after scaling up; for each first node, determine at least one slot to be migrated among the slots mapped to the first node according to the target migration data volume of the first node; for each first node, migrate the at least one slot to be migrated mapped to the first node to the target node.

[0017] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer storage medium stores a computer program. When the computer program is executed, the method in the first aspect or any possible implementation manner in the first aspect is executed, or the method in the second aspect or any possible implementation manner in the second aspect is executed.

[0018] In a seventh aspect, an embodiment of the present application provides a computer program product. The computer program product includes computer program instructions. When the computer program instructions are executed, the method in the first aspect or any possible implementation manner in the first aspect is executed, or the method in the second aspect or any possible implementation manner in the second aspect is executed.

[0019] In an eighth aspect, an embodiment of the present application provides a chip, including a processor and a data interface. The processor reads instructions stored in a memory through the data interface to implement the method in the first aspect or any possible implementation manner in the first aspect, or to implement the method in the second aspect or any possible implementation manner in the second aspect.

[0020] The beneficial effects of the embodiment of the present application compared with the prior art are as follows:

[0021] According to the database cluster management method provided in the first aspect of the present application, by the data volume of each slot mapped to the node to be deleted (i.e., the target node) and the target migration data volume obtained according to the data volume of the node to be deleted (i.e., the target node) and the number of nodes after scale-down, the database cluster to be scaled down can be divided into multiple slot sets, so that the data volume of each slot set is close to the target migration data volume. Since the data volume of each slot set is close to the target migration data volume, after migrating the multiple slot sets to the corresponding nodes, the data volumes of the nodes after scale-down can be made close. The above solution can not only achieve automatic slot migration, but also ensure load balancing as much as possible and ensure the performance of the database cluster after scale-down.

[0022] According to the database cluster management method provided in the second aspect of the present application, by the data volume of each original node in the database cluster and the average data volume obtained according to the total data volume of the original nodes and the number of nodes after expansion, the data volume that each node in the original nodes needs to migrate to the newly added node (i.e., the target node) can be determined, and thus the slot information to be migrated in each node in the original nodes can be determined. According to the above solution for slot migration, not only can automatic slot migration be achieved, but also the data volumes of each node after expansion can be made close, so that load balancing can be ensured as much as possible and the performance of the database cluster after scale-down can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of a database cluster provided by an embodiment of the present application;

[0024] Figure 2 is a schematic flowchart of a database cluster management method provided by an embodiment of the present application;

[0025] Figure 3 is a schematic diagram of slot migration in a scale-down scenario provided by an embodiment of the present application;

[0026] Figure 4 is a schematic flowchart of a database cluster management method provided by an embodiment of the present application;

[0027] Figure 5Schematic diagram of slot migration in the expansion scenario provided by the embodiments of the present application;

[0028] Figure 6 Schematic block diagram of the database cluster management system provided by the embodiments of the present application;

[0029] Figure 7 Schematic flowchart of a specific example of the database cluster management method provided by the embodiments of the present application;

[0030] Figure 8 Schematic block diagram of a data center provided by the embodiments of the present application;

[0031] Figure 9 Schematic block diagram of a data center provided by the embodiments of the present application;

[0032] Figure 10 Schematic block diagram of a data system provided by the embodiments of the present application;

[0033] Figure 11 Schematic block diagram of a data system provided by the embodiments of the present application;

[0034] Figure 12 Schematic block diagram of a data system provided by the embodiments of the present application;

[0035] Figure 13 Structural schematic of a database cluster management system provided by the embodiments of the present application;

[0036] Figure 14 Structural block diagram of a database cluster management system provided by the embodiments of the present application. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0038] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0039] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0040] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0041] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0042] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0043] Exemplarily, the technical solution provided by the present application can be applied to a cache database cluster, for example, RedisCluster (cluster).

[0044] Figure 1 A schematic diagram of a database cluster that can be applied to the present application is shown. Refer to Figure 1 , the database cluster includes multiple nodes, such as node 1 to node N. Each node can map multiple slots, or rather, each node can manage multiple slots. A slot is the smallest unit of logical data sharding, and data is distributed to different nodes through hash mapping. Slots do not directly store data but serve as routing identifiers.

[0045] Exemplarily, one node can correspond to one server, or multiple nodes can correspond to one server. The server can be a physical server or a cloud server. For example, the server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, etc.

[0046] In the related art, during the process of expansion or contraction, the migration of slots is realized by manual operation. For example, during contraction, it is manually determined and configured to migrate the slots mapped to node 1 in the database cluster as shown in Figure 1 to node 2. The above method has a cumbersome manual operation process, and may cause the business access to slow down or become inaccessible due to manual lag and error. Moreover, in the scenario of large-scale applications, there are numerous database clusters, and a large amount of human resources are required for resource expansion or contraction, resulting in a high labor cost.

[0047] In view of this, the present application provides a related database cluster management solution, which can automatically and reasonably perform slot migration during the process of expansion or contraction. The solution provided by the present application will be described below.

[0048] Figure 2 It is a database cluster management method provided by an embodiment of the present application. This method 200 can be executed by a database cluster management system.

[0049] The functions of the database cluster management system can be implemented by one or more nodes, for example, by one or more physical servers. Exemplarily, the database cluster management system and the target database cluster in method 200 can be deployed in the same data center, or the database cluster management system and the target database cluster in method 200 can be deployed in different data centers.

[0050] This method 200 may include S210 to S230, and each step will be described below.

[0051] S210, determine the target node among the M nodes of the target database cluster.

[0052] Among them, the target database cluster is the database cluster to be shrunk, the target node is the node to be subjected to slot migration, and M>3.

[0053] In the embodiment of the present application, the database cluster that needs to be shrunk can be referred to as the database cluster to be shrunk. The target database cluster can be any database cluster that needs to be shrunk. By reducing the number of nodes in the target database cluster from M to M-1 (M-1 = N), that is, deleting a certain node in the target database cluster, the shrinkage of the target database cluster is realized. In the embodiment of the present application, the target node in the target database cluster is used as the node to be deleted (that is, the node to be subjected to slot migration). Before deleting the target node, it is necessary to first migrate the slots mapped to the target node to other nodes in the target database cluster.

[0054] In some embodiments, when the resource utilization rate of at least one node in the target database cluster meets the first preset condition, the target database cluster can be shrunk.

[0055] Exemplarily, the resource utilization rate may include one or more of the following: CPU utilization rate, memory utilization rate, or disk utilization rate.

[0056] For example, CPU utilization can refer to the percentage of the CPU's busy time within a certain duration (e.g., called duration 1). That is, CPU utilization = (working time / duration 1) × 100%. For example, the working time is the time when the CPU executes application program code within duration 1, or the proportion of the time when the CPU executes system kernel operations within duration 1, etc. For example, duration 1 can be 1s, 5s, 100us, etc.

[0057] For example, memory utilization can refer to the ratio of the currently used memory capacity to the total memory capacity.

[0058] For example, disk utilization can refer to the ratio of the currently used disk capacity to the total disk capacity.

[0059] For example, resource utilization can be obtained at a preset period, such as CPU utilization, memory utilization, or disk utilization. The corresponding preset periods for different types of resources can be different or the same.

[0060] In one possible implementation, the first preset condition may include one or more of the following: the duration for which the CPU utilization is less than or equal to the first threshold is greater than or equal to the first preset duration, the duration for which the memory utilization is less than or equal to the second threshold is greater than or equal to the second preset duration, or the duration for which the disk utilization is less than or equal to the third threshold is greater than or equal to the third preset duration.

[0061] Any two of the above first threshold, second threshold, and third threshold can be the same or different. Similarly, any two of the first preset duration, second preset duration, and third preset duration can be the same or different. For example, the above three thresholds can all be 20% or 10%. Another example is that the above three preset durations can all be 30min, 24 hours, or 1 week, etc.

[0062] In one design, taking CPU utilization as an example, if the CPU utilization obtained each time within the first preset duration is less than or equal to the first threshold, it is considered that the duration for which the CPU utilization is less than or equal to the first threshold is equal to the first preset duration. The same applies to memory utilization and disk utilization.

[0063] For example, if the memory utilization of a certain node in the target database cluster has been continuously lower than 10% for 7 days (i.e., the duration for which the memory utilization is less than 10% is equal to 7 days), and the disk utilization has been continuously lower than 20% for 7 days (i.e., the duration for which the disk utilization is less than 20% is equal to 7 days), then it is determined to downsize the target database cluster.

[0064] In another possible implementation, the first preset condition may include one or more of the following: the average CPU utilization rate within a fourth preset duration is less than or equal to a fourth threshold, the average memory utilization rate within a fifth preset duration is less than or equal to a fifth threshold, and the average disk utilization rate within a sixth preset duration is less than or equal to a sixth threshold.

[0065] Any two of the above fourth threshold, fifth threshold, and sixth threshold may be the same or different. Similarly, any two of the fourth preset duration, fifth preset duration, and sixth preset duration may be the same or different. For example, the above three thresholds may all be 20% or 10%. Another example is that the above three preset durations may all be 24 hours or 1 week, etc.

[0066] Taking the CPU as an example, the average CPU utilization rate within the fourth preset duration means the sum of the CPU utilization rates obtained N times within the fourth preset duration divided by N. The corresponding concepts for memory and disk are similar to that of the CPU.

[0067] For example, if the average memory utilization rate of a certain node in the target database cluster within 24 hours is less than or equal to 10%, and the average disk utilization rate of this node within 48 hours is less than or equal to 20%, then it is determined to downsize the target database cluster.

[0068] It should be understood that the above is only an exemplary description of the first preset condition. In specific implementations, the first preset condition may also adopt other designs. For example, the first preset condition may be that the duration during which the memory utilization rate of the target database cluster is less than or equal to the threshold is greater than or equal to the preset duration, where the memory utilization rate of the target database cluster refers to the ratio of the sum of the used memory capacities of each node in the target database cluster to the sum of the memory capacities of each node in the target database cluster.

[0069] It should also be understood that the above only uses CPU utilization rate, memory utilization rate, and / or disk utilization rate, etc. as measurement indicators to determine whether to trigger downsizing. According to actual requirements, other measurement indicators may also be used to determine whether to trigger downsizing. Exemplarily, the measurement indicators (such as CPU utilization rate, etc.) can be directly obtained or calculated by executing general commands on the nodes of the database cluster.

[0070] In the case of determining to downsize the first database cluster, it is also necessary to determine the node to be deleted, that is, to determine the target node.

[0071] In a possible implementation, the target node may be the node with the lowest average resource utilization rate within an eighth preset duration in the target database cluster.

[0072] Exemplarily, for each node of the target database cluster, a weighted sum is calculated for at least two of the average CPU utilization rate, average memory utilization rate, and average disk utilization rate within the eighth preset time period, and the node corresponding to the minimum weighted sum is used as the target node.

[0073] Exemplarily, the target node can be the node with the lowest values in at least two of the average CPU utilization rate, average memory utilization rate, and average disk utilization rate within the eighth preset time period in the target database cluster.

[0074] For example, if the nodes of the target database cluster are Node 1, Node 2, Node 3, and Node 4, and Node 4 has the lowest average CPU utilization rate and average memory utilization rate within the eighth preset time period, then Node 4 is used as the target node.

[0075] It should be understood that the target node can also be determined by other means, and the embodiments of the present application do not limit the method for determining the target node.

[0076] S220. According to the data volume of each slot among the multiple slots mapped to the target node and the target migration data volume, determine N slot sets.

[0077] Wherein, N is the number of the target database clusters after capacity reduction, that is, N = M - 1. Each slot set includes at least one slot among the multiple slots. The target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume.

[0078] For the purpose of achieving load balancing, when performing slot migration, it is desired to evenly divide the data of the target node into N parts. Each part of the data (i.e., the data with a data volume equal to the target migration data volume) can be migrated to a node (i.e., the first node) other than the target node in the target database cluster. However, since all the data in the slot needs to be migrated during slot migration, it may not be possible to evenly divide the data of the target node into N parts. In this case, according to the principle of as even distribution as possible, that is, dividing the data of the target node into N parts with basically the same data volume, which can largely ensure load balancing.

[0079] Dividing the data of the target node into N parts with basically the same data volume means dividing the multiple slots mapped to the target node into N slot sets. The data volume of each slot set is close to the above-mentioned target migration data volume, or in other words, the error between the data volume of each slot set and the target migration data volume is within a certain range.

[0080] Based on the above concept, in a possible implementation, in S220, the range of the migration data volume can be determined first according to the target migration data volume and the preset error. Then, according to the range of the migration data volume and the data volume of each slot among the multiple slots, the set of the N slots is determined.

[0081] Specifically, the range of the migration data volume = the target migration data volume ± the preset error * the target migration data volume. Wherein, the preset error is a percentage, such as ±5%. Or, the range of the migration data volume = the target migration data volume ± the preset error. Wherein, the preset error is a specific value, such as ±3G.

[0082] For example, the target database cluster includes Node 1, Node 2, Node 3, and Node 4, and Node 4 is the target node. The slots mapped to the target node are slots 401 to 500, a total of 100 slots, and the total data volume of the 100 slots is 100G of data. Then, the target migration data volume Qt = 100G / 3 ≈ 33G. For example, if the preset error is ±2G, then the range of the migration data volume is [31, 35]G. For example, if the preset error is ±6%, then the range of the migration data volume is 33G ± 6% * 33G, that is, the range of the migration data volume is [31.02, 34.98]G.

[0083] Continuing with the above example, assume that the range of the migration data volume is [31, 35]G, and assume that the slot data volume situation of Node 4 is shown in Table 1 and Table 2.

[0084] Table 1

[0085] Slots (401 - 500) Data volume (total 100G) 401~419 30G 420~474 29G 475~500 41G

[0086] Table 2

[0087] Slots (401 - 500) Data volume (total 100G) 401~420 35G 421~475 33G 476~500 32G

[0088] It can be seen that the data volume of any slot set shown in Table 1 is not within [31, 35]G, while the data volume of each slot set shown in Table 2 is within [31, 35]G. Therefore, the three slot sets shown in Table 2 can be used as the set of the N slots.

[0089] S230, migrate the set of the N slots to the N first nodes in the target database cluster.

[0090] Wherein, the first node is the node other than the target node among the M nodes of the target database cluster. For example, the M nodes are Node 1 to Node 4, and the target node is Node 4, then the N first nodes are Node 1 to Node 3.

[0091] The set of N slots corresponds to N first nodes one by one. That is to say, one slot set is migrated to only one of the N nodes, and only one of the N nodes receives one slot set.

[0092] Taking the set of N slots as the three slot sets shown in Table 2 as an example, referring to Figure 3 the corresponding relationship between the slot sets and the nodes shown, when slot migration is performed, the three slot sets shown in Table 2 can be migrated to three nodes from Node 1 to Node 3. For example, slots 401 - 420 can be migrated to Node 1, slots 421 - 475 can be migrated to Node 2, and slots 476 - 500 can be migrated to Node 3.

[0093] The embodiments of the present application do not limit the corresponding relationship between the set of N slots and the N first nodes, that is, which slot set is specifically migrated to which first node. For example, this corresponding relationship can be determined randomly or in other ways. The following gives examples to illustrate this corresponding relationship.

[0094] Example 1, the node corresponding to the smaller data represented by the node identifier corresponds to the slot set with a smaller slot identifier, or the node corresponding to the smaller data represented by the node identifier corresponds to the slot set with a larger slot identifier.

[0095] For example, the identifiers of Node 1 to Node 3 are 1, 2, and 3 respectively, and the identifiers of the slots are the numbers shown in Table 2, that is, the identifier of slot 401 is 401, and the others are similar. If the method of the node corresponding to the smaller data represented by the node identifier corresponding to the slot set with a smaller slot identifier is adopted, then, slots 401 - 420 will be migrated to Node 1 (that is, slots 401 - 420 correspond to Node 1), slots 421 - 475 will be migrated to Node 2 (that is, slots 421 - 475 correspond to Node 2), and slots 476 - 500 will be migrated to Node 3 (that is, slots 476 - 500 correspond to Node 3).

[0096] Example 2, the slot set with a larger data volume corresponds to the node with a smaller data volume.

[0097] For example, the data volumes of Node 1 to Node 3 are 500G, 540G, and 532G respectively. Then, slots 401 - 420 with a data volume of 35G will be migrated to Node 1 (that is, slots 401 - 420 correspond to Node 1), slots 421 - 475 with a data volume of 33G will be migrated to Node 3, and slots 476 - 500 with a data volume of 32G will be migrated to Node 2.

[0098] Those skilled in the art can understand that the slot migration operation involved in this application (for example, migrating an N - slot set to N first - level nodes in a target database cluster) means: migrating the data in the slots to the nodes and changing the mapping relationship between the slots and the nodes. The descriptions about the slot migration operation in this article can all refer to this description, and will not be elaborated at other places where the slot migration operation is involved.

[0099] For example, taking the migration of slots 401 - 420 to node 1 as an example, migrating slots 401 - 420 to node 1 means migrating the data in slots 401 - 420 to node 1 and changing the mapping relationship between slots 401 - 420 and the nodes from mapping slots 401 - 420 to node 4 to mapping slots 401 - 420 to node 1.

[0100] According to the database cluster management method provided by the embodiments of this application, by the data volume of each slot mapped to the node to be deleted (i.e., the target node) and the target migration data volume obtained based on the data volume of the node to be deleted (i.e., the target node) and the number of nodes after scale - down, the database cluster to be scaled down can be divided into multiple slot sets, so that the data volume of each slot set is close to the target migration data volume. Since the data volume of each slot set is close to the target migration data volume, after migrating these multiple slot sets to the corresponding nodes, the data volume of each node after scale - down can be made close. The above - mentioned solution can not only realize automatic slot migration, but also can ensure load balancing as much as possible and ensure the performance of the database cluster after scale - down.

[0101] In a possible implementation manner, S230 may specifically include: creating N migration tasks, each migration task is used to migrate a slot set to a first - level node; and executing the N migration tasks in parallel.

[0102] In this solution, by dividing the slots mapped to the target node into N parts, N independent and parallel - executable migration tasks can be created. For example, 3 migration tasks can be created, which are task 1, task 2, and task 3 respectively. Task 1 is used to migrate slots 401 - 420 to node 1, task 2 is used to migrate slots 421 - 475 to node 2, and task 3 is used to migrate slots 476 - 500 to node 3.

[0103] In the related art, for scaling down of a database cluster, the slot migration range and the correspondence between the migration slot and the target node are first manually confirmed, and then the data migration interface is manually called to perform the migration operation. Then, the database cluster starts a built-in task to migrate the slots to be migrated in sequence according to the manually input slot information and the target node for migration. In the above scheme, the manually confirmed slot migration range may not be reasonable. For example, all slots of the target node are mapped to a certain node, which may cause load imbalance and affect the performance of the database cluster. In addition, if the slot to be migrated is split into multiple parts, it is necessary to manually perform the above-mentioned manual call to the data migration interface operation multiple times and the database cluster needs to execute multiple migration tasks in sequence, which will cause the slot migration to take a long time.

[0104] Compared with the above-mentioned related technologies, the solution of the embodiment of the present application can improve the slot migration speed and quickly complete the slot migration by creating multiple migration tasks at the same time and executing the multiple migration tasks in parallel. In addition, the amount of data migrated by the multiple migration tasks is close, so the load balancing can be guaranteed as much as possible, and the performance of the database cluster after the reduction can be guaranteed.

[0105] In a possible implementation, the method 200 may further include:

[0106] S240: Send first information to the target database cluster, wherein the first information instructs the target database cluster to delete the target node, and the first information includes a connection address of the target node and an identifier of the target node.

[0107] After the N slot sets are migrated to the N first nodes in the target database cluster, for example, after the above-mentioned N migration tasks are completed, the database cluster management system can send a first message to the target database cluster, and the first message instructs the target database cluster to delete (or eliminate) the target node. Correspondingly, the target database cluster will delete the target node (that is, delete the corresponding relationship between the target database cluster and the target node) after receiving the first message, and return a command of successful deletion to the database cluster management system. Exemplarily, the first message can be an operation instruction to delete (or eliminate) a node without data.

[0108] In the above solution, the target database cluster deletes the target node so that the target node can be recycled, which is beneficial to improving resource utilization.

[0109] In a possible implementation, after the database cluster management system learns that the target database cluster has successfully deleted the target node, the target node status may be changed from data to be migrated to data to be recycled.

[0110] Based on this solution, by changing the state of the target node from data to be migrated to data to be recycled, the target node can be recycled, which is beneficial to improving resource utilization.

[0111] In a possible implementation, after the N slot sets are migrated to N first nodes in the target database cluster, the resource utilization rate of the scaled-down target database cluster can be analyzed to determine whether to trigger scaling down or scaling up again. For example, it is determined whether to trigger scaling down again based on the foregoing first preset condition, or it is determined whether to trigger scaling up based on the second preset condition described later; alternatively, it can also be determined whether to trigger scaling up or scaling down based on conditions similar to the first preset condition or the second preset condition. In the case of triggering scaling down, the specific scaling-down process and method 200 are similar. In the case of triggering scaling up, the specific scaling-down process is similar to method 400 described later.

[0112] In an example, the minimum number of nodes in the target database cluster is 3, that is, when the target database cluster is scaled down to only 3 nodes, even if the target database cluster meets the conditions for triggering scaling down, such as the foregoing first preset condition, the scaling-down operation will not be triggered again, and the scaling-down task will end directly.

[0113] In a possible implementation, the method may further include:

[0114] S250, perform system initialization on the target node, and mark the status of the target node as available resources after the system initialization is completed.

[0115] After the database cluster management system determines to delete the target node from the target database cluster, or after the database cluster management system changes the status of the target node to to-be-reclaimed, it can trigger the recycling of the node without data (i.e., the target node). That is, the database cluster management system can perform system initialization on the target node, and mark the status of the target node as available resources after the system initialization is completed, so as to realize the recycling of resources. In an example, the database cluster management system can mark the status of the target node as available resources from to-be-reclaimed in its configuration file, so as to complete the recycling of resources.

[0116] According to the above solution, by recycling resources of the node without data, the resource turnover rate can be improved, and unnecessary resource waste can be avoided.

[0117] The above describes the slot migration in the scaling-down scenario. The following combines Figure 4 and Figure 5 to describe the slot migration in the scaling-up scenario.

[0118] Figure 4 is another database cluster management method provided by an embodiment of the present application. This method can be executed by the database cluster management system described above. This method 400 may include S410 to S440. The following describes each step.

[0119] S410, add the target node to the target database cluster, where the target database cluster is the database cluster to be expanded.

[0120] In the embodiments of the present application, the database cluster that needs to be expanded can be referred to as the database cluster to be expanded. The target database cluster can be any database cluster that needs to be expanded. By increasing the number of nodes in the target database cluster from M to M + 1 (M + 1 = N), that is, adding a certain node to the target database cluster, the expansion of the target database cluster is achieved. In the embodiments of the present application, the target node is used as the newly added node. After the target node is added to the target database cluster, some slots in other nodes in the target database cluster need to be migrated to the target node.

[0121] In some embodiments, when the resource utilization rate of at least one node in the target database cluster meets the second preset condition, the target database cluster is expanded. For the relevant content or definition of the resource utilization rate, etc., reference can be made to the description of method 200, which will not be elaborated here.

[0122] In a possible implementation, the second preset condition may include one or more of the following: the duration during which the CPU utilization rate is greater than or equal to the seventh threshold is greater than or equal to the seventh preset duration, the duration during which the memory utilization rate is greater than or equal to the eighth threshold is greater than or equal to the eighth preset duration, or the duration during which the disk utilization rate is greater than or equal to the ninth threshold is greater than or equal to the ninth preset duration.

[0123] Any two of the above seventh threshold, eighth threshold, and ninth threshold can be the same or different. Similarly, any two of the seventh preset duration, eighth preset duration, and ninth preset duration can be the same or different. For example, the above three thresholds can all be 70% or 80%, and the above three preset durations can all be 24 hours or 1 week, etc. Exemplarily, the seventh preset duration can be the same as the first preset duration in method 200, the eighth preset duration can be the same as the second preset duration in method 200, and the ninth preset duration can be the same as the third preset duration in method 200.

[0124] For example, if the memory utilization rate of a certain node in the target database cluster has been greater than 90% for 24 hours (that is, the duration during which the memory utilization rate is greater than 90% is equal to 24 hours), and the disk utilization rate has been greater than 90% for 24 hours, it is determined to expand the target database cluster.

[0125] In another possible implementation, the second preset condition may include one or more of the following: the average CPU utilization rate within the tenth preset duration is greater than or equal to the tenth threshold, the average memory utilization rate within the eleventh preset duration is greater than or equal to the eleventh threshold, and the average disk utilization rate within the twelfth preset duration is greater than or equal to the twelfth threshold.

[0126] Any two of the above-mentioned tenth threshold, eleventh threshold, and twelfth threshold may be the same or different. Similarly, any two of the tenth preset duration, eleventh preset duration, and twelfth preset duration may be the same or different. For example, the above three thresholds may all be 80% or 90%, and the above three preset durations may all be 24 hours or 1 week, etc. Exemplarily, the tenth preset duration may be the same as the fourth preset duration in method 200, the eleventh preset duration may be the same as the fifth preset duration in method 200, and the twelfth preset duration may be the same as the sixth preset duration in method 200.

[0127] For example, if the average memory utilization rate of a certain node in the target database cluster within 24 hours is greater than 90%, and the average disk utilization rate of this node within 48 hours is greater than 80%, then it is determined to expand the target database cluster.

[0128] It should be understood that the above is only an exemplary description of the second preset condition. In specific implementations, the second preset condition may also adopt other designs. For example, the second preset condition may be that the duration for which the memory utilization rate of the target database cluster is greater than or equal to the threshold is greater than or equal to the preset duration.

[0129] It should also be understood that the above only uses CPU utilization rate, memory utilization rate, and / or disk utilization rate, etc. as the measurement indicators for triggering expansion. According to actual requirements, other measurement indicators may also be used to determine whether to trigger expansion. Exemplarily, the measurement indicator (such as CPU utilization rate, etc.) can be directly obtained or calculated by executing a general command on the nodes of the database cluster.

[0130] Exemplarily, in the case of determining to expand the target database cluster, appropriate resources can be selected from the resource pool as the target nodes according to the resource pool information and in combination with its resource allocation algorithm, where the specifications of the target nodes are consistent with the specifications of the existing nodes in the target database cluster.

[0131] For example, if the memory utilization rate of a certain node in the target database cluster has been greater than 90% for 24 consecutive hours, and the disk utilization rate has been greater than 90% for 24 consecutive hours, then the memory and disk can be expanded. For example, if the specifications of the existing nodes in the target database cluster are all 4g50g, then the specifications of the target nodes are also 4g50g, and 4g50g means 4G memory + 50G disk.

[0132] Exemplarily, after the database cluster management system selects a target node from the resource pool, it can mark the target node as used. Then, the database cluster management system can connect to the target database cluster and send an instruction to add a new node to the target database cluster. The instruction may include the address information of the target node. After receiving the instruction, the target database cluster adds the address information of the target node to the target database cluster.

[0133] Furthermore, after the target database cluster successfully adds the target node to the target database cluster, it can send a notification to the database cluster management system. The notification may include the information of the existing nodes and the new node in the target database cluster. The information may be, for example, an identifier and / or an address, etc.

[0134] S420. For each first node, determine the target migration data volume of the first node according to the average data volume and the data volume of the first node.

[0135] S430. For each first node, determine at least one slot to be migrated in the slot mapped to the first node according to the target migration data volume of the first node.

[0136] Herein, the first node is a node in the database cluster other than the target node, and the average data volume is equal to the data volume of the target database cluster divided by the number of nodes in the target database cluster after expansion (i.e., N).

[0137] For the purpose of achieving load balancing, it is desired that after slot migration, the data of all nodes (N nodes) in the target database cluster is equal, that is, the data volume of each of the N nodes is the above-mentioned average data volume. However, since all the data in the slot needs to be migrated during slot migration, it may not be possible to satisfy the condition that the data volume of the N nodes is equal after slot migration. In this case, as long as it is ensured that the data volume of the N nodes is close after migration, load balancing can be guaranteed to a large extent.

[0138] Based on the above concept, in a possible implementation, in S420, for each first node, the difference between the data volume of the first node and the average data volume can be used as the target migration data volume of the first node. Then, in S430, according to the target migration data volume of the first node and a preset error, determine the at least one slot to be migrated.

[0139] In the above solution, the target migration data volume of the first node is a specific value. The target migration data volume of the first node is the expected migration data volume of the first node, but the actual migration data volume may be different from this.

[0140] In method 400, assume that the data volume of the target database cluster (i.e., the sum of the data volumes of the slots mapped to all nodes of the target database cluster) is Nsum, then the average data volume Nav = Nsum / N. Assume that the data volume of the first node is Nnode, then the target migration data volume Nt of the first node = Nnode - Nav. Thus, the data volume actually migrated from the first node to the target node belongs to Nt ± preset error * Nav, that is, the data volume range of the at least one slot to be migrated is (Nt ± preset error * Nav). Here, the preset error is a percentage, such as ±5%. Or, the data volume actually migrated from the first node to the target node belongs to Nt ± preset error, that is, the data volume range of the at least one slot to be migrated is Nt ± preset error. Here, the preset error is a specific value, such as ±5G.

[0141] In another possible implementation, in S420, for each first node, a data volume range can be determined according to the data volume of the first node, the average data volume, and the preset error, and this data volume range is used as the target migration data volume of the first node. Then, in S430, according to the target migration data volume of the first node, the at least one slot to be migrated is determined.

[0142] In this solution, the target migration data volume of the first node is a data volume range, and the data volume actually migrated from the first node to the first node, that is, the data volume of the at least one slot to be migrated, is within this data volume range.

[0143] Exemplarily, the target migration data volume of the first node = the data volume of the first node - (average data volume ± preset error * average data volume). Here, the preset error is a percentage, such as ±5%. Also, for example, the target migration data volume of the first node = the data volume of the first node - (average data volume ± preset error). Here, the preset error is a specific value, such as ±5G.

[0144] In one example, the target database cluster originally had 3 nodes, namely Node 1 to Node 3. Currently, the target database cluster is to be expanded from 3 nodes to 4 nodes, and the newly added node is Node 4. Then, some of the slots mapped to the existing 3 nodes need to be migrated to the newly added node. Suppose the slots mapped to Node 1 are 1 to 100, with a total of 30G of data, the slots mapped to Node 2 are 101 to 200, with a total of 32G of data, and the slots mapped to Node 3 are 201 to 300, with a total of 38G of data. There are a total of 300 slots and 100G of data in the target database cluster. Then, Node 1 needs to migrate approximately 5G (30 - 100 / 4) of data, Node 2 needs to migrate approximately 7G (32 - 100 / 4) of data, and Node 3 needs to migrate approximately 13G (38 - 100 / 4) of data. Assuming the preset error is ±6%, then the slots to be migrated can be filtered out as follows: the slots 1 to 30 mapped to Node 1 (data volume 6G), the slots 101 to 120 mapped to Node 2 (data volume 7G), and the slots 201 to 240 mapped to Node 3 (data volume 14G).

[0145] S440. For each first node, migrate at least one to-be-migrated slot mapped to the first node to the target node.

[0146] For example, referring to Figure 5 , continuing with the above example, migrate the slots 1 to 30 of Node 1, the slots 101 to 120 of Node 2, and the slots 201 to 240 of Node 3 to Node 4.

[0147] According to the database cluster management method provided by the embodiments of the present application, by the data volume of each original node in the database cluster and the average data volume obtained based on the total data volume of the original nodes and the number of nodes after expansion, the data volume that each node in the original nodes needs to migrate to the newly added node (i.e., the target node) can be determined, and thus the slot information that needs to be migrated in each node in the original nodes can be determined. Performing slot migration according to the above solution can not only achieve automatic slot migration, but also make the data volume of each node after expansion close, so as to ensure load balancing as much as possible and ensure the performance of the database cluster after scaling down.

[0148] In a possible implementation, S440 may specifically include: creating M migration tasks, each migration task being used to migrate at least one to-be-migrated slot of a first node to the target node, where M is the number of first nodes; and executing the M migration tasks in parallel.

[0149] For example, 3 migration tasks can be created, namely Task 1 to migrate slots 1 to 30, Task 2 to migrate slots 101 to 120, and Task 3 to migrate slots 201 to 240.

[0150] Based on the above solution, by simultaneously creating multiple migration tasks and executing the multiple migration tasks in parallel, the slot migration speed can be increased, and the slot migration can be quickly completed.

[0151] Exemplarily, when a migration task has a migration exception, the database cluster management system can re-execute the migration task until the migration task is successfully completed.

[0152] In a possible implementation, for each first node, after migrating at least one slot to be migrated mapped to the first node to the target node, for example, after the M migration tasks complete the migration tasks, the database cluster management can analyze the resources of the expanded target database cluster to determine whether to trigger expansion or contraction again. For example, based on the foregoing first preset condition to determine whether to trigger contraction again, or based on the foregoing second preset condition to determine whether to trigger expansion; or, it can also be determined whether to trigger expansion or contraction based on conditions similar to the first preset condition or the second preset condition. In the case of triggering contraction, the specific contraction process and method are similar to 200. In the case of triggering expansion, the expansion can refer to method 400.

[0153] Figure 6 It is a database cluster management system provided by an embodiment of the present application. The database cluster management system can execute any of the foregoing method embodiments, such as method 200 or method 400.

[0154] See Figure 6 , the database cluster management system 600 may include a monitoring module 610, a scaling module 620, and a slot migration module 630. Optionally, the database cluster management system 600 may further include a resource pool 640.

[0155] The monitoring module 610 is used to monitor at least one database cluster. The monitoring module 610 has an independent configuration file, which records the information of the database clusters to be monitored (for example, the data center to which the database cluster belongs, the name of the database cluster, the addresses of each node in the database cluster, etc.), the threshold values corresponding to various indicator information (such as CPU utilization rate, memory utilization rate, and / or disk utilization rate) (such as the first threshold to the twelfth threshold described above), and the triggering conditions (such as the first preset condition and the second preset condition described above). Exemplarily, for different database clusters, the indicator information to be monitored may be the same or different, and can be set according to specific requirements. The monitoring module 610 also periodically or real-time collects various indicator information and compares it with the corresponding threshold values. When the above comparison result meets the triggering condition, the monitoring module 610 sends the corresponding database cluster information to the scaling module 620.

[0156] For example, for the database cluster A monitored by the monitoring module 610, its metric information can be CPU utilization, memory utilization, and disk utilization. The preset thresholds corresponding to the CPU utilization are 70% and 20%, the preset thresholds corresponding to the memory utilization are 85% and 10%, and the preset thresholds corresponding to the disk utilization are 90% and 20%. The trigger condition is that at least two of the following items are satisfied: (1) The CPU utilization is greater than 70%; (2) The duration during which the memory utilization is greater than 85% is greater than or equal to 24 hours; (3) The duration during which the disk utilization is greater than 90% is greater than or equal to 24 hours; (4) The CPU utilization is less than 20%; (5) The duration during which the memory utilization is less than 10% is greater than or equal to 7 days; (6) The duration during which the disk utilization is less than 20% is greater than or equal to 7 days.

[0157] Based on the above example, if the monitoring module 610 monitors that the metric information of each node in the database cluster A satisfies at least two of the trigger conditions, it sends the database cluster information of the database cluster A to the scaling module 620. For example, the database cluster information of the database cluster A includes: Cluster: 008, Node address: 0.0.0.81,..., 0.0.0.84, Information: The duration during which the memory utilization is less than 10% is equal to 7 days (i.e., the memory utilization has been less than 10% for 7 days), and the duration during which the disk utilization is less than 20% is greater than or equal to 7 days.

[0158] The metric information and trigger conditions of the database cluster B are the same as those of the database cluster A. If the monitoring module 610 also monitors that the metric information of each node in the database cluster B satisfies at least two of the trigger conditions, it sends the database cluster information of the database cluster B to the scaling module 620. For example, the database cluster information of the database cluster B includes: Cluster: 001, Node address: 0.0.0.1,..., 0.0.0.4, Information: The duration during which the CPU utilization is greater than 90% is 24 hours, and the duration during which the disk utilization is greater than 90% is equal to 24 hours.

[0159] It should be understood that the database cluster A can be the target database cluster in the foregoing method 200. The database cluster B can be the target database cluster in the foregoing method 400.

[0160] The scaling module 620 receives the database cluster information sent by the monitoring module 610 and triggers scaling up or down.

[0161] For example, based on the database cluster information of database cluster A received by the scaling module 620, it is determined that the resource utilization rate is low for a long time and scaling down is required. Specifically, the operation "node reduction" can be determined, with the current value: 64g500g 4 nodes and the target value: 64g500g 3 nodes. That is, reduce the 64g500g 4 nodes to 64g500g 3 nodes.

[0162] For another example, based on the database cluster information of database cluster B received by the scaling module 620, it is determined that the resource utilization rate is high and scaling up is required. Specifically, the operation "node increase" can be determined, with the current value: 4g50g 3 nodes and the target value: 4g50g 4 nodes. That is, scale up the 4g50g 3 nodes to 4g50g 4 nodes.

[0163] In some embodiments, the scaling module 620 may execute S250 described above.

[0164] The slot migration module 630 is used to migrate slots on certain nodes in the database clusters (such as database cluster A and database cluster B) that need to be scaled down or up as determined by the scaling module 620 to other nodes.

[0165] In some embodiments, the slot migration module 630 may execute S210 to S240 described above.

[0166] In some embodiments, the slot migration module 630 may execute S410 to S440 described above.

[0167] The resource pool 640 has an independent configuration file, which is used to store resources of different specifications and their current states, such as 2c4g50g (i.e., 2 CPU cores + 4GB of memory + 50GB of storage) is unused, 4c8g100g is unused,..., 32g64g500g is unused, etc., for different specifications. The resource specifications can be set according to the actual resources, and the specifications and usage conditions of all resources are recorded in the configuration file. The scaling module 620 can schedule the resource pool 640. The resource pool 640 is planned in the data center to which the database cluster belongs.

[0168] Next, based on Figure 6 the database cluster management system shown, in combination with Figure 7 a schematic specific process of the database cluster management method provided by the embodiments of the present application will be described. The interactions between the modules in the database cluster management system and the interactions between the database cluster management system and the database cluster (such as the target database cluster mentioned above) are described in this process. For the content related to scaling down in this method, reference can also be made to method 200, and for the content related to scaling up, reference can also be made to method 400.

[0169] Step 1: The monitoring module continuously monitors database cluster A, database cluster B, ..., database cluster X.

[0170] Step 2: When the indicator information of the database cluster monitored by the monitoring module meets the corresponding trigger condition, the monitoring module pushes the database cluster information that meets the corresponding trigger condition to the expansion and contraction module.

[0171] As described above Figure 6 In the example listed above, if the indicator information of database cluster A and database cluster B both meet the corresponding trigger conditions, the monitoring module sends the database cluster information of database cluster A and the database cluster information of database cluster B to the scaling module. The following description still takes the example that the indicator information of database cluster A and database cluster B both meet the corresponding trigger conditions.

[0172] Step 3: After receiving the database cluster information of database cluster A or the database cluster information of database cluster B sent by the monitoring module, the capacity expansion and contraction module will make a capacity expansion and contraction judgment, that is, determine whether capacity expansion or contraction is required.

[0173] The following describes the two scenarios of expansion and reduction respectively.

[0174] Scenario 1 (reduction)

[0175] The scaling module receives the database cluster information of database cluster A and can determine the operation: node reduction, the current value: 64g500g 4 nodes, and the target value: 64g500g 3 nodes.

[0176] Step 4a: When capacity reduction is required, the capacity expansion and contraction module may call the slot migration module and notify the slot migration module of the environment information of the database cluster A that needs to be reduced.

[0177] For example, the environment information of database cluster A is: cluster: 008, node address: 0.0.0.81, ..., 0.0.0.84, operation: data migration.

[0178] In step 5a, the slot migration module connects to database cluster A and determines the resource usage of each node in database cluster A (such as CPU utilization, memory utilization and / or disk utilization, etc.) based on the received environmental information of database cluster A, determines the node suitable for slot migration (i.e., the target node in method 200, and node 4 is taken as an example below), and marks the status of node 4 in the resource pool as data to be migrated.

[0179] It should be understood that node 4 is the node to be migrated, and node 4 may be the target node in the above method 200.

[0180] Exemplarily, node 4 can be the node with the lowest resource utilization rate in database cluster A.

[0181] Step 6a: The slot migration module performs slot migration on the nodes marked with data to be migrated.

[0182] Referring to the foregoing method 200, in one example, the slot migration process may include the following steps A to E.

[0183] Step A: Determine N slot sets.

[0184] The slot migration module may split the data on the node to be migrated into N parts according to the principle of as even distribution as possible based on the resource usage of each node in database cluster A analyzed in step 5a. N represents the number of nodes after scale-down, and the slot information involved in each part of the data is recorded. For example, when scaling down from 4 nodes to 3 nodes, the data on the 4th node is split into 3 parts. Suppose the slots on the 4th node are 401 to 500, a total of 100 slots with 100G of data, 35G of data is stored in slots 401 to 420, 33G of data is stored in slots 421 to 475, and 32G of data is stored in slots 476 to 500.

[0185] For more details regarding step 1, reference can be made to the relevant content in method 200.

[0186] Step B: The slot migration module starts 3 migration tasks to simultaneously migrate the slots where the split data on node 4 is located to other nodes in database cluster A.

[0187] Among them, migration task 1 migrates slots 401 to 420, migration task 2 migrates slots 421 to 475, and migration task 3 migrates slots 476 to 500. The 3 migration tasks are independent of each other and can be executed in parallel. In this solution, by simultaneously creating multiple migration tasks to be executed in parallel by the monitoring module, the migration speed can be improved.

[0188] Step C: After each migration task completes the migration task, it notifies the slot migration module. When a certain migration task is abnormally notified, the slot migration module can re-execute the migration task until the migration task is successfully completed.

[0189] Step D: After the slot migration module receives the notification that all tasks have been successfully migrated, it can send an operation instruction to delete the node without data to database cluster A. The operation instruction includes the connection address of the node to be deleted and the ID of the node.

[0190] Step E: Database cluster receives the operation instruction to delete the node without data and executes the deletion action. After the deletion is completed, it returns a command indicating successful deletion to the slot migration module.

[0191] Step 7a, the slot migration module marks the status of the deleted node from having data to be migrated to to-be-recycled, and notifies the scaling module that the slot migration is completed.

[0192] Step 8a, after receiving the notification from the slot migration module, the scaling module automatically triggers the recycling of the node without data (i.e., Node 4). That is, the recycling node (i.e., Node 4) is initialized by the system. After completion, the status of Node 4 in the configuration file is marked from to-be-recycled to available resource, completing the resource recycling operation.

[0193] Step 9a, the scaling module performs an analysis on the resources of the database cluster A after scaling down. It determines whether further scaling down or scaling up is needed based on the resource utilization rate of the nodes in the database cluster A after scaling down.

[0194] For example, when the average value of the resource utilization rates of each node in the database cluster A after scaling down is greater than or equal to the corresponding threshold within a preset duration, the scaling down ends. When the average value of the resource utilization rate of a certain node in the database cluster A after scaling down is less than the corresponding threshold within a preset duration, a second scaling down is triggered, and the above slot migration and node recycling steps are repeated until the average value of the resource utilization rates of each node in the database cluster A is greater than or equal to the corresponding threshold within a preset duration, and the scaling down ends.

[0195] Exemplarily, the minimum number of database cluster nodes is 3. That is, when the number of database cluster nodes is scaled down to only 3 nodes, even if the average value of the resource utilization rate of a certain node in the database cluster A is less than the corresponding threshold within a preset duration, the scaling down operation will not be triggered again, and the scaling down task ends directly.

[0196] Scenario 2 (scaling up):

[0197] When the scaling module receives the database cluster information of database cluster B, it can determine the operation: operation: nodeincrease, current value: 4g50g 3 nodes, target value: 4g50g 4 nodes.

[0198] Step 4b, when scaling up is needed, the scaling module can select appropriate resources from the resource pool according to the resource pool information stored in itself and in combination with its resource allocation algorithm (i.e., determining the target node described in Method 400 or selecting the newly added node), where the selected resource specifications are consistent with the specifications of the existing nodes in database cluster B.

[0199] Step 5b: The scaling module marks the selected resources as used, then connects to database cluster B and sends an instruction to add a new node to database cluster B. The instruction may contain the address information of the selected new node, and expands the node address of this resource into database cluster B. Database cluster B automatically completes the node expansion operation.

[0200] Step 6b: After the scaling module triggers the expansion, after the node successfully joins database cluster B, it sends a notification to the slot migration module. The content of the notification includes the existing node information and the new node information in database cluster B.

[0201] Step 7b: The slot migration module performs slot migration according to the received information in the following steps A to F.

[0202] Step A: The slot migration module connects to database cluster B and analyzes the data distribution of all nodes in database cluster B, and analyzes the resource utilization rate of each node storing data in database cluster B.

[0203] Step B: The slot migration module calculates the total amount of data stored in database cluster B, divides it by the number of nodes after expansion to obtain the average value Nav, that is, it is necessary to migrate data with a volume of Nav to the new node, and subtracts the average value Nav from the actual amount of data stored in each node in database cluster B to obtain the actual amount of data that each node needs to migrate, denoted as Y.

[0204] Step C: The slot migration module filters out data with a volume around Y (a certain error is allowed, and the error value can be set artificially, such as a plus or minus 5% error) on each node and records the slot information involved in each piece of data. For example, if it is expanded from 3 nodes to 4 nodes, that is, part of the data of the existing 3 nodes is migrated to the new node. Suppose the slots on the first node are 1 to 100, with a total of 30G of data, the slots on the second node are 101 to 200, with a total of 32G of data, and the slots on the third node are 201 to 300, with a total of 38G of data. There are a total of 300 slots and 100G of data in the cluster. That is: Node 1 needs to migrate about 5G of data, Node 1 needs to migrate about 7G of data, Node 1 needs to migrate about 13G of data (because a slot can only belong to a certain node, and the entire slot needs to be migrated during migration, so there will be a deviation in the actual amount of data migrated).

[0205] Step D, the slot migration module starts 3 migration tasks to simultaneously migrate the slots where the data screened out on the original 3 nodes is located to the newly added nodes in database cluster B. Assume the slot information to be migrated is as follows: for node 1, slots 1 to 30 (data volume 6G); for node 2, slots 101 to 120 (data volume 7G); for node 3, slots 201 to 240 (data volume 14G). Migration task 1 migrates slots 1 to 30, migration task 2 migrates slots 101 to 120, and migration task 3 migrates slots 201 to 240. The 3 migration tasks are independent of each other and can be executed in parallel to speed up the migration.

[0206] Step E, after each migration task completes the migration task, it can notify the slot migration module. When a node task migration exception is notified, the slot migration module can re-execute the task until the task is completed normally.

[0207] Step F, after the slot migration module receives the notification that all tasks have been migrated successfully, it can return a notification of the completion of the expansion to the scaling module.

[0208] Step 8b, after the scaling module receives the notification of the completion of the expansion of database cluster B, it will analyze the expanded database cluster B once, and determine whether to scale down again or whether to expand based on the resource utilization rate of the nodes in the expanded database cluster B. For example, when the average value of the resource usage rates of each node in the expanded database cluster B within the preset time period is less than the corresponding threshold, the scaling down ends. When the average value of the resource usage rate of a certain node in the scaled-down database cluster A within the preset time period is greater than or equal to the corresponding threshold, the second expansion is triggered, and the above slot migration steps are repeated until the average value of the resource usage rates of each node in database cluster A within the preset time period is less than the corresponding threshold, and the expansion ends.

[0209] In summary, the above process basically does not require manual intervention, and the entire process depends on the intelligent linkage of the monitoring module, the scaling module, and the slot migration module to complete. Moreover, only a relatively low labor cost is required to ensure the normal operation of the monitoring module, the scaling module, and the slot migration module and to maintain the resource scaling of the resource pool.

[0210] It should be noted that the slot A of node X (slot 101 of node 1) described in some places in this article refers to the slot A mapped to node X. In addition, "greater than or equal to" in this article can be replaced with "greater than", and "less than" can be replaced with "less than or equal to".

[0211] Embodiments of the present application also provide various designs of a data center and / or a data system. The database cluster management system in the data center described below may be the database cluster management system 600 described above, and the database cluster may be a cache database cluster. In addition, it should be noted that the following description only takes the resource pool not being deployed in the database cluster management system as an example. Exemplarily, the resource pool may also be deployed in the database cluster management system.

[0212] It should be understood that the database cluster management system described below has a management function for the corresponding database cluster. That is, the database cluster management system can scale out or scale in the corresponding database cluster. For example, the database cluster management system can implement any of the method embodiments described above. It should be understood that the process of the database cluster management system managing the corresponding database cluster involves interaction with the resource pool. For example, selecting a suitable node from the resource pool to join the database cluster to achieve scale out, or recycling the node into the resource pool in the case of scale in.

[0213] Figure 8 is an example of the data center provided by the embodiments of the present application. Refer to Figure 8 , the data center 800 may include a database cluster management system, at least one database cluster (such as database cluster 1 to database cluster N), and a resource pool (the resource pool may be one or more). The database cluster management system can be used to manage the at least one database cluster.

[0214] Figure 9 shows another example of the data center provided by the embodiments of the present application. Refer to Figure 9 , the data center 900 may include multiple database cluster management systems (such as database cluster management system 1 to database cluster management system N), multiple database clusters (such as database cluster 1 to database cluster M), and multiple resource pools (such as resource pool A to resource pool X). Among them, each database cluster management system is independent and manages different database clusters, that is, the same database cluster will only be managed by a certain database cluster management system. For example, the data center 900 may include database cluster management system 1 and database cluster management system 2, database clusters 1 to 100, and resource pool 1 and resource pool 2. Database cluster management system 1 can be used to manage database clusters 1 to 55, database cluster management system 2 can be used to manage database clusters 56 and 100, resource pool 1 corresponds to database cluster management system 1, and resource pool 2 corresponds to database cluster management system 2.

[0215] Figure 10 shows a data system provided by the embodiments of the present application. Refer to Figure 10, the data system 1000 may include multiple data centers, and each data center may include a database cluster management system, multiple database clusters, and one or more resource pools. For example, data center 1010 may include database cluster management system 1, database clusters 1A to 1N, and resource pool 1. Data center 1020 may include database cluster management system 2, database clusters 2A to 2N, and resource pool 2. Among them, the database cluster management system in any one of the data centers is used to manage the multiple database clusters in that data center, and the database clusters and resource pools in different data centers are different.

[0216] Figure 11 shows another data system provided by an embodiment of the present application. Refer to Figure 11 , the data system 1100 may include multiple data centers 1100, such as data center 1 to data center N. Among them, data center 1 includes a database cluster management system, and this database cluster management system can manage the database clusters in data center 1 to data center N. Each of data center 1 to data center N also includes multiple database clusters and one or more resource pools, and the database clusters and resource pools in different data centers are different.

[0217] Exemplarily, data center 1 may include a configuration file, which records information about data center 1 to data center N, the names of the database clusters in data center 1 to data center N, the node addresses in each database cluster in each data center, and the central nodes to which each database cluster belongs.

[0218] For example, the configuration file may record the following information: Cluster: 001, Node Address: 0.0.0.1,..., 0.0.0.4, Belonging: Data Center 1; Cluster: 008, Node Address: 0.0.0.81,..., 0.0.0.84, Belonging: Data Center 2;..., Cluster: 00n, Node Address: 0.0.0.n1,..., 0.0.0.n4, Belonging: Data Center n.

[0219] Figure 12 shows yet another data system provided by an embodiment of the present application. Refer to Figure 12 , the data system 1200 may include multiple data centers, such as data center 1 to data center N. Data center 1 may include database cluster management system 1, and data centers 2 to N include replicas of database cluster management system 1. Database cluster management system 1 can manage the corresponding database clusters in the figure.

[0220] Exemplarily, instead of deploying a replica of the database cluster management system 1 in data centers 2 to N, a replica of the database cluster management system 1 can be deployed in data center 1.

[0221] In the above solution, multiple replicas of the same database cluster management system can be deployed in the same data center or different data centers, thereby ensuring the high availability of the data system and avoiding the unavailability of the entire data system caused by the failure of a single data center.

[0222] Based on Figures 8 to 12 the data centers or data systems shown, the database clusters can be uniformly managed or partitioned according to the number of database clusters and the distribution of data centers according to actual needs, improving the flexibility of managing data.

[0223] Figure 13 FIG. is a schematic structural diagram of another database cluster management system provided by an embodiment of the present application. The database cluster management system 1300 may include a processing unit 1310.

[0224] In a possible implementation, the database cluster management system 1300 may be used to implement the steps shown in method 200.

[0225] The processing unit 1310 is configured to determine a target node among the M nodes of the target database cluster, where the target database cluster is a database cluster to be scaled down, and the target node is a node to be subject to slot migration, M>3; according to the data volume of each slot among the multiple slots mapped to the target node and the target migration data volume, determine N slot sets, N = M - 1, where each slot set includes at least one slot among the multiple slots, the target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume; migrate the N slot sets to N first nodes in the target database cluster, where the first nodes are the nodes among the M nodes other than the target node, and the N slot sets correspond to the N first nodes one by one.

[0226] Optionally, the processing unit 1310 is further configured to create N migration tasks, where each migration task is used to migrate one slot set to one first node; and execute the N migration tasks in parallel.

[0227] Optionally, the database cluster management system 1300 may further include a communication unit 1320, configured to send a first message to the target database cluster, where the first message instructs the target database cluster to delete the target node, and the first message includes the connection address of the target node and the identifier of the target node.

[0228] Optionally, the processing unit 1310 is specifically configured to: after completing the migration of the N slot sets and deleting the target node, perform system initialization on the target node, and mark the state of the target node as resource available after completing the system initialization.

[0229] Optionally, the target node is the node with the lowest average resource utilization rate within a preset duration among the M nodes.

[0230] In a possible implementation manner, the database cluster management system 1300 can be used to implement the steps shown in method 400.

[0231] The processing unit 1310 is configured to: add a target node to a target database cluster, where the target database cluster is a database cluster to be expanded; for each of the first nodes, determine the target migration data volume of the first node according to the average data volume and the data volume of the first node, where the first node is a node other than the target node in the database cluster, and the average data volume is equal to the data volume of the target database cluster divided by the number of nodes in the target database cluster after expansion; for each of the first nodes, determine at least one slot to be migrated in the slots mapped to the first node according to the target migration data volume of the first node; for each first node, migrate the at least one slot to be migrated mapped to the first node to the target node.

[0232] Optionally, the processing unit 1310 is specifically configured to: create M migration tasks, each migration task is used to migrate the at least one slot to be migrated of one of the first nodes to the target node, and M is the number of the first nodes; execute the M migration tasks in parallel.

[0233] It should be noted that the above database cluster management system 1300 is embodied in the form of functional units. The term "unit" here can be implemented in software and / or hardware forms, and no specific limitation is made thereto.

[0234] For example, the "unit" can be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merging logic circuit, and / or other suitable components that support the described functions.

[0235] Therefore, the units of the examples described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0236] Figure 14 FIG. 4 is a schematic structural diagram of a database cluster management system 1400 provided by an embodiment of the present application. As Figure 14 shown, the database cluster management system 1400 of this embodiment includes: at least one processor 1401 ( Figure 14 only one processor is shown in the figure), a memory 1402, and a computer program 1403 stored in the memory 1402 and executable on the at least one processor 1401. When the processor 1401 executes the computer program 1403, the steps of any of the above method embodiments are implemented.

[0237] Those skilled in the art can understand that Figure 14 FIG. 4 is merely an example of the database cluster management system 1400, and does not constitute a limitation on the database cluster management system 1400. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0238] The processor 1401 may be a central processing unit (CPU), and the processor 1401 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0239] In some embodiments, the memory 1402 may be an internal storage unit of the database cluster management system 1400, such as the hard disk or memory of the database cluster management system 1400. In some other embodiments, the memory 1402 may also be an external storage device of the database cluster management system 1400, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the database cluster management system 1400. Further, the memory 1402 may also include both the internal storage unit of the database cluster management system 1400 and external storage devices. The memory 1402 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 1402 may also be used to temporarily store data that has been output or will be output.

[0240] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details will not be elaborated here.

[0241] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.

[0242] The embodiments of the present application also provide a database cluster management system, which includes: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the database cluster management system implements any of the above-mentioned method embodiments.

[0243] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the methods of any of the above-mentioned method embodiments can be implemented.

[0244] The embodiments of the present application provide a computer program product, and when the computer program product runs, the methods of any of the above-mentioned method embodiments can be implemented.

[0245] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the methods of the above-mentioned embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code to the photographing device / terminal device. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.

[0246] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0247] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0248] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0249] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0250] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A database cluster management method, characterized in that Including: Determine a target node among M nodes of a target database cluster, where the target database cluster is a database cluster to be scaled down, the target node is a node to be migrated with slots, and M > 3; According to the data volume of each slot among the multiple slots mapped to the target node and the target migration data volume, determine N slot sets, where N = M - 1. Each of the slot sets includes at least one slot among the multiple slots. The target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume; Migrate the N slot sets to N first nodes in the target database cluster. The first nodes are the nodes among the M nodes other than the target node, and the N slot sets correspond to the N first nodes one by one.

2. The method according to claim 1, wherein The migrating the N slot sets to N first nodes in the target database cluster includes: Create N migration tasks, where each migration task is used to migrate one of the slot sets to one of the first nodes; Execute the N migration tasks in parallel.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Send a first message to the target database cluster, where the first message instructs the target database cluster to delete the target node. The first message includes the connection address and the identifier of the target node.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: In the case where the migration of the N slot sets is completed and the target node is deleted, perform system initialization on the target node, and mark the status of the target node as available resources after the system initialization is completed.

5. The method according to any one of claims 1-4, characterized in that, The target node is the node with the lowest average resource utilization rate within a preset duration among the M nodes.

6. A method for managing a database cluster, characterized in that, Including: Add a target node to a target database cluster, where the target database cluster is a database cluster to be scaled up; For each first node, determine the target migration data volume of the first node according to the average data volume and the data volume of the first node. The first node is a node among the database cluster other than the target node, and the average data volume is equal to the data volume of the target database cluster divided by the number of nodes in the target database cluster after expansion; For each of the first nodes, determine at least one slot to be migrated among the slots mapped to the first node according to the target migration data volume of the first node; For each of the first nodes, migrate the at least one slot to be migrated mapped to the first node to the target node.

7. The method according to claim 6, wherein The migrating the at least one slot to be migrated mapped to each of the first nodes to the target node includes: Create M migration tasks, where each migration task is used to migrate the at least one slot to be migrated of one of the first nodes to the target node, and M is the number of the first nodes; Execute the M migration tasks in parallel.

8. A database cluster management system, characterized in that, It includes a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the database cluster management system executes the method described in any one of claims 1-7.

9. A data center, characterized in that, The data center includes a database cluster management system and a target database cluster; the database cluster management system is used for: determining a target node among the M nodes of the target database cluster, where the target database cluster is the database cluster to be scaled down, the target node is the node to be subject to slot migration, and M>3; determining N slot sets according to the data volume of each slot among the multiple slots mapped to the target node and the target migration data volume, where N = M - 1. Each of the slot sets includes at least one of the multiple slots. The target migration data volume is equal to the data volume of the target node divided by N, and the data volume of each slot set is determined according to the target migration data volume; migrating the N slot sets to N first nodes in the target database cluster, where the first nodes are the nodes among the M nodes other than the target node, and the N slot sets correspond to the N first nodes one by one.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed, the method described in any one of claims 1-7 is executed.