Method, system and equipment for expanding and shrinking capacity of database cluster, storage medium and computer program product
By automatically obtaining the response time information and data access request type of the database cluster, the automatic scaling of the database cluster is realized, solving the problems of cumbersome and inefficient manual operations in the existing technology, and improving the intelligence and efficiency of scaling.
Patent Information
- Application Number
- CN202412000486.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the scaling of the database cluster requires manual operation, and there are problems of human operation errors and cumbersome steps, resulting in inefficient scaling.
By obtaining the response time information of the database cluster for different types of data access requests, the scaling strategy is automatically determined based on the response time information and the type of data access requests, and the scaling processing of the database cluster is realized.
The automatic trigger scaling mechanism of the database cluster is realized, which improves the intelligence and efficiency of scaling, and reduces human operation errors.
Smart Images

Figure CN119961245A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of database expansion and contraction, and in particular to a method, system, device, storage medium and computer program product for expanding and contracting a database cluster. Background Art
[0002] A database cluster can be a distributed database system composed of at least two or more database servers, which is used to provide more intelligent and comprehensive data services to the user end. During the use of the database cluster, it is particularly necessary to expand or shrink its capacity. Expansion can increase data storage capabilities, and shrinking can reduce unnecessary resource consumption.
[0003] In the related art, the solutions for scaling database clusters are all to manually scale them up or down when encountering performance bottlenecks or other scaling requirements. The manual scaling method has the problem of human operation errors and the steps are relatively cumbersome, resulting in low efficiency in scaling database clusters. Therefore, there is an urgent need to propose a mechanism for automatically triggering scaling. Summary of the invention
[0004] Embodiments of the present application provide a method, system, device, storage medium, and computer program product for scaling up or down a database cluster. The method can automatically trigger scaling up or down a database cluster based on business requirements for the database cluster and the needs of database processing business, thereby improving the intelligence of triggering scaling up or down a database cluster and the efficiency of scaling up or down a database cluster.
[0005] The technical solution of the application embodiment is implemented as follows:
[0006] The present application provides a method for scaling a database cluster, including:
[0007] Obtain the response time information of the database cluster to be expanded or reduced for different types of data access requests;
[0008] Determine the scaling strategy for the database cluster based on the response time information and the type of data access request;
[0009] Based on the expansion and contraction strategy, the database cluster is expanded and contracted.
[0010] The present application embodiment provides a system for scaling up and down a database cluster, including:
[0011] The cluster status real-time detection module is used to obtain the response time information of the database cluster for different types of data access requests;
[0012] The cluster status analysis module is used to determine the expansion and contraction strategy for the database cluster based on the response time information and the type of data access request;
[0013] The cluster computing power scheduling module is used to scale the database cluster based on the scaling strategy.
[0014] The present application provides a device for scaling up and down a database cluster, including:
[0015] A memory, used to store instructions for scaling up and down the executable database cluster;
[0016] The processor is used to implement the method for scaling the database cluster provided in the embodiment of the present application when executing the executable database cluster scaling instruction stored in the memory.
[0017] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable database cluster scaling instructions are stored. The computer-executable database cluster scaling instructions are configured to execute the database cluster scaling method provided in the embodiment of the present application.
[0018] An embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for scaling up and down a database cluster provided in the embodiment of the present application is implemented.
[0019] The embodiment of the present application provides a method, system, device, storage medium and computer program product for scaling up and down a database cluster. The method includes: obtaining the response time information of the database cluster to be scaled up and down for different types of data access requests; determining the scaling up and down strategy for the database cluster based on the response time information and the type of data access request; and scaling up and down the database cluster based on the scaling up and down strategy. In this way, the scaling up and down strategy of the database cluster is determined by the response time information of the database cluster for different types of data access requests and the type of data access request, realizing a mechanism for automatically triggering the scaling up and down of the database cluster based on the business needs of the database cluster and the business needs of the database processing, thereby improving the intelligence of triggering the scaling up and down of the database cluster and the efficiency of the scaling up and down of the database cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of a method for scaling up and down a database cluster provided in an embodiment of the present application;
[0021] Figure 2 A flow chart of a K8S-based large-scale host node computing power integrated resource dynamic scheduling management method provided in an embodiment of the present application;
[0022] Figure 3 A flowchart of a method for real-time monitoring of Redis cluster data processing status and resource intelligent Redis node provided in an embodiment of the present application;
[0023] Figure 4 A schematic diagram of the connection status of a proxy cluster and Redis nodes provided in an embodiment of the present application;
[0024] Figure 5 A schematic diagram of the structure of a capacity expansion and contraction system provided in an embodiment of the present application;
[0025] Figure 6 A flowchart of a Redis cluster expansion and contraction method based on elastic data synchronizer computing power scheduling provided in an embodiment of the present application;
[0026] Figure 7 A schematic diagram of computing power scheduling based on a computing power scheduler is provided in an embodiment of the present application.
[0027] Figure 8 A schematic diagram of the structure of a database cluster expansion and contraction system provided in an embodiment of the present application;
[0028] Fig. 9 A schematic diagram of the composition structure of a database cluster expansion and contraction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0030] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0031] In the following description, reference is made to “some embodiments\other embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments\other embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0033] In the related art, when the online expansion solution of the open source Remote Dictionary Server (Redis) cluster (a database cluster) is used to expand the node capacity, it is first necessary to manually start a new Redis instance and add the new Redis instance to the Redis cluster as a new node. This needs to be done manually by using the management command or application programming interface (API) of the Redis cluster. Then data migration is performed to migrate part of the data from the existing node to the new node. The Redis cluster provides a data redistribution function to automatically migrate data from nodes with higher loads to new nodes. After the data migration is completed, the configuration information of the Redis cluster is manually updated, including the node list, slot allocation, etc., to ensure the normal operation of the cluster. Although the online expansion solution of the open source Redis cluster provides great flexibility and convenience, it also has some disadvantages: data migration will have a certain impact on the performance of the cluster, especially in large-scale data migration, which may cause the cluster's response time to increase or throughput to decrease. When processing data migration of large key values, the entire cluster will be stuck and completely unavailable. After the data migration is completed, the client's Internet Protocol (IP) connection information needs to be updated. Otherwise, most clients will not be able to continue to use it and changes will need to be made without business interruption.
[0034] In the related art, a Codis proxy-based online capacity expansion solution for Redis clusters can add nodes to Redis clusters for capacity expansion without affecting the external services of the cluster. First, new Redis server nodes need to be added to the cluster, and these service nodes need to be manually configured and added to the Codis cluster. Next, an expansion command is initiated to Codis through the API, and the data migration function of Codis is used to migrate part of the data from the existing Redis nodes to the new nodes. This process is completed by the Codis management tool, which will determine which data needs to be migrated according to certain strategies (such as load balancing, data distribution, etc.), and ensure the integrity and consistency of the data during the migration process. After the data migration is completed, Codis will update the data routing table to reflect the new nodes and data distribution. In this way, when the Codis proxy receives a request from the client, it can forward the request to the correct Redis node according to the latest routing table. However, this method also has some disadvantages: First, the data migration process may have a certain impact on the performance of the cluster, especially in large-scale data migration, which may cause the cluster response time to increase or the throughput to decrease; second, the nodes need to be deployed manually, and the entire solution is based on Codis. Therefore, the stability and reliability of Codis are crucial to the performance and availability of the entire cluster.
[0035] In the related art, a data processing method obtains the first storage capacity corresponding to the first Redis cluster in real time, wherein the first storage capacity increases as the data written to the first Redis cluster increases; when the first storage capacity is greater than or equal to the first capacity threshold, the first data to be written corresponding to the first Redis cluster is obtained, the writing of the first data to be written to the first Redis cluster is stopped, and the first data to be written is written to the second Redis cluster. In this way, the resources of a set of spare Redis clusters are used to provide the Redis cluster with automatic expansion capabilities, thereby improving the efficiency of expanding the Redis storage capacity. In addition, during the expansion process, there is no need to shut down the Redis cluster, thereby ensuring the stability of the service and improving the user experience. The defects of this method are more obvious. It only considers the expansion of capacity, not the reduction of capacity, and the capacity expansion is carried out by adding a set of Redis clusters, which is very inconvenient.
[0036] Based on the problems in the related art, the embodiment of the present application provides a method for scaling up and down a database cluster. The method can realize a mechanism for automatically triggering scaling up and down a database cluster based on the business needs of the database cluster and the business needs of the database processing, thereby improving the intelligence of triggering scaling up and down a database cluster and the efficiency of scaling up and down a database cluster. The following will describe the method for scaling up and down a database cluster provided by the embodiment of the present application. Figure 1FIG. 1 is a flow chart of a method for scaling up and down a database cluster provided in an embodiment of the present application, and the method comprises the following steps:
[0037] S101. Obtain response time information of a database cluster to be expanded or reduced for different types of data access requests.
[0038] It should be noted that the database cluster to be expanded or reduced can be a distributed database system that needs to be expanded by adding nodes or reduced by reducing nodes. The database cluster can be a Redis cluster, a relational database management system-MySQL cluster, a database system based on distributed file storage-MongoDB cluster, etc. This application uses the database cluster as a Redis cluster for illustration.
[0039] In some embodiments, the data access request may be a request for data reading, writing, querying, etc. sent by a client to a database cluster. The data access request to the database cluster may be triggered by calling a command corresponding to accessing the database cluster.
[0040] In some embodiments, data access requests may include multiple types, such as query type, read data type, and write data type, etc. The type of data access request may be determined by the type of the calling command, for example, if the type of the calling command is a query command, the type of the corresponding data access request may be a query type; if the type of the calling command is a read-write type command, the type of the corresponding data access request may be a read-write type.
[0041] In some embodiments, the response time information may include the time taken by the database cluster to process the data access request, the number of times the database cluster processes query data access requests per second, etc. The response time information of the database cluster to be scaled to historical data access requests and / or the response time information of the database cluster to be scaled to real-time data access requests may be obtained.
[0042] In some embodiments, response time information of the database cluster to be scaled up or down for multiple data access requests can be obtained, and different types of data access requests can be classified to obtain response time information of the database cluster to be scaled up or down for different types of data access requests. The response time information of the same type of data access requests may include one or more types.
[0043] S102: Determine a scaling strategy for the database cluster based on the response time information and the type of data access request.
[0044] In some embodiments, the response time information corresponding to different types of data access requests may be different. The response time information can reflect the processing speed, processing capability, etc. of the database cluster for the data access requests. Therefore, based on the response time information of the database cluster for different types of data access requests, the scaling strategy of the database cluster can be determined with the goal of improving the processing speed, processing capability, etc. of the database cluster.
[0045] In some embodiments, different data access requests can represent different business needs, and the response time information can be represented as the time information of the database cluster processing the business needs. Therefore, the scaling processing of the database cluster is determined according to the type of data access request and the corresponding response time information. This can be understood as determining the scaling strategy for the database cluster according to the type of business need and the time information of the database cluster processing the business need. Since the data access requests to be processed by the database cluster to be scaled and the types of data access requests are dynamically changing, the resource allocation of the database cluster can be adjusted according to different business needs based on the response time information and the type of data access request to adapt to changes in business needs.
[0046] In some embodiments, the expansion trigger conditions and contraction trigger conditions corresponding to the database cluster can be determined based on the response time information and the type of data access request, and when it is determined that the expansion trigger conditions are met, the expansion strategy is determined based on the response time information and the type of data access request, or when it is determined that the contraction trigger conditions are met, the contraction strategy is determined based on the response time information and the type of data access request.
[0047] S103: Based on the expansion and contraction strategy, the database cluster is expanded and contracted.
[0048] It should be noted that the expansion and contraction processing may include expanding the database cluster and contracting the database cluster. The expansion processing may include creating a new node for the database cluster, and migrating part of the data in each node of the database cluster to the new node; the expansion processing may include determining the candidate nodes that need to be deleted or destroyed in the database cluster, and migrating the data in the candidate nodes to other nodes in the database cluster excluding the candidate nodes.
[0049] In some embodiments, the expansion and contraction strategy may include the expansion strategy of the database cluster, and the expansion and contraction strategy may include information such as the storage capacity of the newly added node to be expanded in the database cluster, and the data content and data volume that need to be migrated out of each node in the database cluster. In the process of expanding the database cluster based on the expansion strategy, the data that needs to be migrated out of each node can be migrated to the newly added node according to the information such as the storage capacity of the newly added node, and the data content and data volume that need to be migrated out of each node in the database cluster.
[0050] In other embodiments, the scaling strategy may include a scaling strategy for the database cluster, and the scaling strategy may include identification information of the nodes to be scaled in the database cluster, data content and data volume in the nodes to be scaled in, etc. In the process of scaling down the database cluster based on the scaling strategy, the data content in the nodes to be scaled in can be migrated to other nodes in the database cluster according to the identification information of the nodes to be scaled in the database cluster, data content and data volume in the nodes to be scaled in, etc., and the nodes to be scaled in can be destroyed.
[0051] In an embodiment of the present application, the response time information of the database cluster to be expanded or reduced for different types of data access requests is obtained; based on the response time information and the type of data access request, the expansion and reduction strategy for the database cluster is determined; based on the expansion and reduction strategy, the database cluster is expanded or reduced. In this way, the expansion and reduction strategy of the database cluster is determined by the response time information of the database cluster for different types of data access requests and the type of data access requests, and a mechanism for automatically triggering the expansion and reduction of the database cluster based on the business needs of the database cluster and the business needs of the database processing is realized, thereby improving the intelligence of triggering the expansion and reduction of the database cluster and the efficiency of the expansion and reduction of the database cluster.
[0052] In some embodiments of the present application, the response time information includes the average response time of the database cluster to data access requests. Based on this, the response time information of the database cluster to different types of data access requests is obtained, that is, step S101 can be implemented by the following steps S1011 to S1013, and each step is described below.
[0053] S1011. Obtain a reference-type data access request using a proxy cluster, and forward the reference-type data access request to a first node of a corresponding database cluster.
[0054] It should be noted that the reference type is any type of a data access command. The proxy cluster may be a cluster other than the database cluster, and the proxy cluster may include one or more.
[0055] In some embodiments, when multiple reference-type data access requests are obtained, the data access requests of each reference type may be distributed to corresponding proxy clusters according to the number of reference-type data access requests and the number of proxy clusters to achieve load balancing.
[0056] In some embodiments, a reference type data access request carries slot information of the accessed data, such as a slot index value, and the proxy cluster can perform a hash calculation based on the slot index value, locate the first node corresponding to the data access request according to the calculated hash value, and forward the reference type data access request to the first node. The first node can be one or more nodes in the database cluster.
[0057] S1012: Determine a response duration of the first node according to a first time when the proxy cluster sends a reference type data access request and a second time when the proxy cluster receives feedback data from the first node in response to the reference type data access request.
[0058] In some embodiments, after receiving a reference type data access request sent by a proxy cluster, the first node may perform reading, writing, querying or other processing based on the reference type data access request, and send the processing result as feedback data to the proxy cluster. During this process, the first time when the proxy cluster sends the reference type data access request to the first node and the second time when the proxy cluster receives the feedback data sent by the first node may be recorded, wherein both the first time and the second time may be specific time points or moments.
[0059] In some embodiments, the response time of the first node can be determined based on the difference between the second time and the first time. For example, if the first time is 8:30:16:32 and the second time is 8:30:16:33, the determined response time of the first node is 0.01s.
[0060] S1013: Based on the response time and the access frequency information of the reference type data access request, determine the average response time of the database cluster to the reference type data access request.
[0061] In some embodiments, the access frequency information of the reference type data access request may be the average number of accesses per second for the reference type data access request, and the sum of the response times of the first node to each reference type data access request may be determined as the total response time of the database cluster to the reference type data access request, and the ratio of the total response time to the average number of accesses per second may be determined as the average response time of the database cluster to the reference type data access request.
[0062] It can be understood that the response time of the first node is determined by the first time when the proxy cluster sends the reference type of data access request and the second time when the proxy cluster receives the feedback data of the first node for the reference type of data access request, taking into account factors such as the computing processing time of the first node itself, network transmission time and network congestion, so that the average response time of the database cluster for the reference type of data access request finally determined based on the response time of the first node and the access frequency information of the reference type of data access request is more accurate.
[0063] In some embodiments of the present application, the scaling strategy includes the number of nodes to be expanded in the database cluster, and the type of data access request includes a first type of global access to the database cluster; based on this, based on the response time information and the type of data access request, the scaling strategy for the database cluster is determined, that is, the above-mentioned step S102 can be implemented by the following steps S1021A to S1023A, and each step is described below.
[0064] S1021A. Determine first access frequency information of a first type of data access request.
[0065] It should be noted that the first type of data access request can be a global access type for the data of each node in the database cluster, that is, the first type of data access request can indicate that the user needs to query or read and write all the data in each node; the first type of data access request can also be a data access request corresponding to a cross-shard scan command, and the cross-shard scan command can be keys, scan, etc.
[0066] In some embodiments, the first access frequency information may include the number of accesses of the first type of data access request within a preset time period and the average number of accesses per second. The number of acquired first type of data access requests can be monitored within the preset time period, and the average number of accesses per second of the first type of data access request can be determined based on the length of the preset time period and the number of first type of data access requests.
[0067] S1022A: Determine whether the database cluster meets a first expansion trigger condition based on the first access frequency information and the response time information of the first type of data access request.
[0068] In some embodiments, the response time information of the first type of data access request may include the average time taken by the database cluster to process the first type of data access request and the query rate per second (QPS) of the database cluster, where QPS may represent the number of query requests in the first type of data access request processed by the database cluster per second. Whether the database cluster meets the first expansion trigger condition may be determined based on the average time taken by the database cluster to process the first type of data access request and QPS.
[0069] S1023A: If the database cluster meets the first expansion trigger condition, determine the number of nodes in the database cluster to be expanded according to the first access frequency information, the response time information of the first type of data access request, and the number of nodes in the database cluster.
[0070] In some embodiments, when the database cluster meets the first expansion trigger condition, the number of nodes in the database cluster to be expanded can be calculated based on the first access frequency information, the response time information of the first type of data access request and the number of nodes in the database cluster.
[0071] For example, if the database cluster is a Redis cluster, the first type is the keys type, and the average number of calls per second for data access requests of the keys type is Count. avg1 Unit (times / second), average time consumed is Usec avg1 Unit (seconds), can be obtained by the current query rate per second QPS avg1 The maximum QPS of a single shard (a single node in a database cluster) is set to 90,000 units (times / second). avg1 *Usec avg 1 is the time required to execute key type data access requests per second. It is the percentage of QPS occupied by all current commands per second. Redis can process a maximum of 90,000 commands per second. The current average processing is QPS avg1 The remaining performance time can be The remaining performance time is less than the time required to process a data access request of the keys type to determine whether the first expansion trigger condition is met. When the remaining performance time is less than the time required to process a data access request of the keys type, that is, when: , it is determined that the database cluster meets the first expansion trigger condition.
[0072] Furthermore, since Redis is a single-threaded system, Redis is already too late to process the problem, and the delay will continue to increase. Assuming that the number of shards in the Redis cluster is N, the number of shards is expanded. The expanded number of shards NScal1 can be expressed by formula (1):
[0073]
[0074] It should be noted that the number of shards NScal1 that needs to be expanded obtained from formula (1) can be rounded up. For example, when Nscal1 is 1.6, it is rounded up to 2. Since the maximum size of the database cluster does not exceed 100 nodes, the maximum value of N+Nscal1 is 100.
[0075] It can be understood that, based on the first access frequency information and response time information of the first type of data access request, it can be determined whether the current performance of the database cluster meets the processing requirements of the first type of data access request, that is, whether the database cluster meets the first expansion trigger condition; further, when the database cluster meets the first expansion trigger condition, based on the first access frequency information, the response time information of the first type of data access request, and the number of nodes in the database cluster, the number of nodes to be expanded that meet the processing requirements of the first type of data access request can be determined, thereby realizing dynamic allocation of database cluster resources.
[0076] In some embodiments of the present application, the type of data access request includes a second type of access to key data of a first specific type. Based on this, based on the response time information and the type of data access request, the scaling strategy for the database cluster is determined, that is, step S102 can also be implemented by the following steps S1021B to S1024B, and each step is described below.
[0077] S1021B. Determine a first memory occupancy size of key data in the key-value pair storage structure accessed by the second type of data access request.
[0078] It should be noted that the key data may be a key, and the first specific type may be a type such as String, List, Set, SortedSet, Hash, etc. The first memory occupied size may be a memory size occupied after encoding the key data of a specific type.
[0079] In some embodiments, the data in the database cluster can be stored in a key-value pair storage structure, and the second type of data access request can be an access request for a specific type of key data in the storage structure. The specific type of key data may occupy a large amount of memory, which may affect the response speed of the database cluster to the client. Therefore, analyzing the second type of data access request and determining the expansion strategy of the database cluster can trigger the expansion and contraction strategy based on business needs.
[0080] S1022B: Determine a second memory occupation size that meets the memory occupation condition from the first memory occupation size.
[0081] In some embodiments, the first memory occupancy size may include multiple sizes, and the second memory occupancy size that meets the memory occupancy condition may be the first memory occupancy size whose memory occupancy size is greater than the memory occupancy threshold, or the first memory occupancy size whose memory occupancy size is equal to the memory occupancy threshold. The memory occupancy threshold may be any pre-set value, such as 5 megabytes (MB), 8MB, etc.
[0082] In some embodiments, in the process of determining a second memory occupancy size that satisfies a memory occupancy condition in the first memory occupancy size, each first memory occupancy size and a memory occupancy threshold may be compared in turn to determine the second memory occupancy size, and the second memory size may include one or more.
[0083] S1023B: Determine whether the database cluster meets a second expansion trigger condition based on the response time information of the second type of data access request.
[0084] In some embodiments, the response time information of the second type of data access request may include the average time taken by the database cluster to process the second type of data access request and the QPS of the database cluster for the second type of data access request.
[0085] In some embodiments, the time taken by the database cluster to process the second type of data access request per second can be determined based on the average time taken by the database cluster to process the second type of data access request and the number of accesses per second of the second type of data access request, and then combined with the QPS of the second type of data access request, it is determined whether the database cluster meets the second expansion trigger condition. The second expansion trigger condition is used to determine whether the database cluster needs to be expanded, and the second expansion trigger condition is different from the first expansion trigger condition.
[0086] S1024B: If the database cluster meets the second expansion trigger condition, determine the number of nodes in the database cluster to be expanded according to the number of key data corresponding to the second memory occupation size, the response time information of the second type of data access request, and the number of nodes in the database cluster.
[0087] In some embodiments, the second memory occupied sizes corresponding to different key data may be different or the same, and the number of key data corresponding to the second memory size can be determined based on the number of the second memory occupied sizes.
[0088] In some embodiments, when it is determined that the database cluster meets the second expansion trigger condition, the number of nodes in the database cluster to be expanded can be calculated based on the number of key data corresponding to the second memory occupancy size, the response time information of the second type of data access request, and the number of nodes in the database cluster.
[0089] For example, if the database cluster is a Redis cluster, for key data types such as String, List, Set, Sorted Set, and Hash, if the bytes after calculating the key encoding are greater than 5MB, the key is considered to be a key with large memory usage, and the number of large memory keys is counted as COUNT. BigMem If the average number of calls per second for the command corresponding to each second type access request is Count avg2 , the average time is Usec avg2 , where the average number of calls per second for the command corresponding to the nth type of data access request is CDount avg2,n , the average time is Usec avg2,n , then the second expansion trigger condition can be:
[0090] Furthermore, since the key of a single shard occupies too much memory, querying and writing a single shard takes a long time. Assuming that the number of shards in the Redis cluster is N, the number of shards is expanded. The expanded number of shards NScal2 can be expressed by formula (2):
[0091]
[0092] It should be noted that the number of shards Nscal2 that needs to be expanded obtained from formula (2) can be rounded up. For example, when Nscal2 is 1.6, it can be rounded up to 2. In addition, since the maximum size of the database cluster does not exceed 100 nodes, the maximum value of N+Nscal2 is 100.
[0093] It can be understood that, based on the response time information of the second type of data access request, it can be determined whether the current performance of the database cluster meets the processing requirements of the second type of data access request, that is, whether the database cluster meets the second expansion trigger condition; further, when the database cluster meets the second expansion trigger condition, based on the number of key data corresponding to the second memory occupancy size, the response time information of the second type of data access request, and the number of nodes in the database cluster, the number of nodes to be expanded that meet the processing requirements of the second type of data access request can be determined, thereby realizing dynamic allocation of database cluster resources.
[0094] In some embodiments of the present application, the type of data access request includes a third type of key-value access for a second specific type. Based on the response time information and the type of data access request, the scaling strategy for the database cluster is determined, that is, the above-mentioned step S102 can also be implemented by the following steps S1021C to S1024C. Each step is described below.
[0095] S1021C. Determine a first quantity of key values corresponding to each key data in the key-value pair storage structure accessed by the third type of data access request.
[0096] It should be noted that the key value corresponding to the key data can be value, and the key value of the second specific type can be a value of type String, List, Set, Sorted Set, Hash, etc.
[0097] In some embodiments, the number of key values corresponding to different key data may be different. The number of key values of the key data corresponding to the key value accessed by each third type of data access request can be determined separately to obtain the first number of key values corresponding to each third type of data access request.
[0098] S1022C: Determine a second quantity that meets a key value quantity threshold condition from each first quantity.
[0099] In some embodiments, the second quantity that satisfies the key value quantity threshold condition may be the first quantity whose quantity value is greater than the key value quantity threshold, or the first quantity whose quantity value is equal to the key value quantity threshold. The second quantity may include at least one, and the number of the second quantity may be the same as the number of the first quantity, or may be less than the number of the first quantity.
[0100] S1023C. Determine whether the database cluster meets a third expansion trigger condition based on the response time information of the third type of data access request.
[0101] In some embodiments, the response time information of the third type of data access request may include the average time taken by the database cluster to process the third type of data access request and the QPS of the database cluster for the third type of data access request.
[0102] In some embodiments, the time taken by the database cluster to process the third type of data access request per second can be determined based on the average time taken by the database cluster to process the third type of data access request and the number of accesses per second of the third type of data access request, and then combined with the QPS of the third type of data access request, it is determined whether the database cluster meets the third expansion trigger condition. The third expansion trigger condition is used to determine whether the database cluster needs to be expanded, and the third expansion trigger condition is different from the first expansion trigger condition and the second expansion trigger condition.
[0103] S1024C: If the database cluster meets the third expansion trigger condition, determine the number of nodes in the database cluster to be expanded according to the second quantity, the response time information of the third type of data access request, and the number of nodes in the database cluster.
[0104] In some embodiments, when it is determined that the database cluster meets the third expansion trigger condition, the number of nodes in the database cluster to be expanded can be calculated based on the second quantity, response time information of the third type of data access requests, and the number of nodes in the database cluster.
[0105] For example, if the database cluster is a Redis cluster, there will be a number of secondary data for the third type of value such as List, Set, Sorted Set, Hash, etc. The number of shards is N, and the number of its members is greater than 10,000. The key is considered to be a key with a large amount of data, and the number of secondary data of the value of the large key is counted as COUNT BigCount Get the average number of calls per second for each type of command Count avg3 , the average time is Usec avg3 , where the average number of calls per second for the nth type of command is Count avg3,n , the average time is Usec avg3,n The second expansion trigger condition may be:
[0106] Furthermore, since there are too many values of large keys in a single shard, querying and writing a single shard takes a long time. Assuming that the number of shards in the Redis cluster is N, the number of shards is expanded. The expanded number of shards NScal3 can be expressed by formula (3):
[0107]
[0108] It should be noted that the number of shards NScal3 that needs to be expanded obtained from formula (2) can be rounded up. For example, when Nscal3 is 1.6, it can be rounded up to 2. In addition, since the maximum size of the database cluster does not exceed 100 nodes, the maximum value of N+NScal3 is 100.
[0109] It can be understood that, based on the response time information of the third type of data access request, it can be determined whether the current performance of the database cluster meets the processing requirements of the third type of data access request, that is, whether the database cluster meets the third expansion trigger condition; further, when the database cluster meets the third expansion trigger condition, based on the second quantity, the response time information of the third type of data access request, and the number of nodes in the database cluster, the number of nodes to be expanded that meet the processing requirements of the second type of data access request can be determined, thereby realizing dynamic allocation of database cluster resources.
[0110] In some embodiments of the present application, the scaling strategy includes the number of nodes to be scaled down in the database cluster; the type of data access request includes a fourth type of local access to the database cluster; based on this, based on the response time information and the type of data access request, the scaling strategy for the database cluster is determined, that is, step S102 can also be implemented through the following steps S1021D to S1023D, and each step is described below.
[0111] S1021D. Determine second access frequency information of the fourth type of data access request.
[0112] It should be noted that the fourth type of data access request may be a data access request type for some nodes in a database cluster, for example, it may be a data access request for key data or key values of types such as String, List, Set, Sorted Set, Hash, etc. The amount of data (key data or key values) accessed by the fourth type of data access request is small, so there is no high latency.
[0113] In some embodiments, the second access frequency access information may include the number of accesses of the fourth type of data access request within a preset time period and the average number of accesses per second. The number of acquired fourth type of data access requests can be monitored within the preset time period, and the average number of accesses per second of the fourth type of data access request can be determined based on the length of the preset time period and the number of fourth type of data access requests.
[0114] S1022D. Determine whether the database cluster meets the first shrinking trigger condition based on the second access frequency information and the response time information of the fourth type of data access request.
[0115] In some embodiments, the response time information of the fourth type of data access request may include the average time taken by the database cluster to process the fourth type of data access request and the number of query requests in the first type of data access request processed by the database cluster per second, i.e., QPS. Whether the database cluster meets the first scaling-down trigger condition may be determined based on the average time taken by the database cluster to process the fourth type of data access request and the average number of accesses per second.
[0116] S1023D: If the database cluster meets the first shrinking trigger condition, determine the number of nodes in the database cluster to be shrunk according to the second access frequency information, the response time information of the fourth type of data access request, and the number of nodes in the database cluster.
[0117] In some embodiments, when the database cluster meets the first shrinking trigger condition, the number of nodes in the database cluster to be shrunk can be calculated based on the second access frequency information, the response time information of the fourth type of data access request, and the number of nodes in the database cluster.
[0118] For example, if the database cluster is a Redis cluster, for the fourth type of data access request corresponding to data types such as String, List, Set, Sorted Set, and Hash, the number of shards is N, and the average number of calls per second for each type of command is Count. avg4 , average time Usec avg4 , where the average number of calls per second for the nth type of command is Count avg4,n , the average time is Usec avg4,n The first shrinking trigger condition may be:
[0119] Furthermore, since all commands take less than 0.5 seconds per second, the performance of the Redis cluster is far from being fully utilized, and the shrinking operation is performed. The number of shards to be shrunk Nshrink can be expressed by formula (4):
[0120]
[0121] It should be noted that the number of shards Nshrink obtained from formula (2) that needs to be shrunk can be rounded down. For example, when Nshrink is 1.6, it can be rounded down to 1. In order to minimize the cluster size to no less than 3 nodes, the minimum value of N+Nshrink is 3.
[0122] It can be understood that, based on the second access frequency information and response time information of the fourth type of data access request, it can be determined whether the current performance of the database cluster is fully utilized, that is, whether the database cluster meets the first scaling down trigger condition; further, when the database cluster meets the first scaling down trigger condition, based on the second access frequency information, the response time information of the fourth type of data access request, and the number of nodes in the database cluster, it can be determined that the number of nodes that need to be scaled down in the database cluster while meeting the processing requirements of the fourth type of data access request, thereby achieving reasonable allocation of database cluster resources.
[0123] In some embodiments of the present application, the database cluster is expanded or reduced in capacity based on the expansion or reduction strategy, that is, step S103 can be implemented by the following steps S1031A to S1032A, and each step is described below.
[0124] S1031A. Determine the second nodes corresponding to the number of nodes to be expanded, and create first data synchronization modules corresponding to the third nodes of the database cluster from which data is to be migrated.
[0125] It should be noted that the second nodes may be newly added nodes of the database cluster, the number of the second nodes is the same as the number of nodes to be expanded, and the third nodes may be the original nodes of the database cluster.
[0126] In some embodiments, the first data synchronization module is used to synchronize the slot data in the corresponding third node. Each third node has its own corresponding first data synchronization module. The first data synchronization module can record the changes in the data stored in the corresponding third node and always keep it consistent with the data in the corresponding third node.
[0127] S1032A: Load the slot data to be migrated out in the third node into the corresponding first data synchronization module, and forward the corresponding slot data to be migrated out to the corresponding second node in sequence through the first data synchronization module.
[0128] In some embodiments, the slot data to be migrated out in the third node may be a slot corresponding to a key-value pair, and one key-value pair may correspond to one slot data. The slot data to be migrated out in the third node may be determined according to the number of second nodes, the storage capacity of the second node, etc. The slot data to be migrated out in each third node may be loaded into the corresponding first data synchronization module, and then the obtained slot data to be migrated out may be forwarded to the second node in sequence through each first data synchronization module.
[0129] Exemplarily, the psync command can be used to synchronize the data between the first data synchronization module and the corresponding third node to keep the data consistent; create a child process to save the slot Redis database backup file (RDB) data, and use the slot-key as an index to find the key data belonging to the same slot; create a bio to load data from the RDB file to the temporary database of the elastic data synchronizer. After the load is complete, the key data in the temporary database is moved to the corresponding third node through psync. If the load fails, delete the temporary database. Since bio helps to maintain slot-level atomicity, by loading into a temporary database, it is possible to avoid contamination of the third node (primary database) when loading data, and easily roll back the task in case of failure.
[0130] It can be understood that by creating first data synchronization modules corresponding to the third nodes of the data to be migrated in the database cluster, the slot data to be migrated in the third node is loaded into the corresponding first data synchronization module, and the corresponding slot data to be migrated is forwarded to the corresponding second node in sequence through each first data synchronization module. This will not affect the normal response of the third node to business needs, and avoids the problem of increased response time or decreased throughput of the database cluster, or even stuck or unusable database cluster caused by a single large key data or large key value when migrating in the native migrate method.
[0131] In some embodiments of the present application, in the process of determining the second node corresponding to the number of nodes to be expanded, the state change information of the database cluster can be monitored in real time based on the resource state change notification mechanism; if it is determined that a new node is connected to the database cluster, the resource information of each node in the database cluster is obtained; the resource information is analyzed and processed to obtain resource optimization information and network delay optimization information; based on the resource optimization information and the network delay optimization information, and in combination with the ant colony algorithm, the second node is determined from at least one new node.
[0132] It should be noted that the resource status change notification mechanism can be a watch mechanism provided by the application program interface service of Kubernetes (K8S). The status change information includes the access of a new node to the database cluster, and the new node includes at least one. The status change information can also include the usage status of each node in the database cluster, such as ready state, normal state, abnormal state, unschedulable state, etc.
[0133] In some embodiments, K8S can monitor new nodes added to the database cluster through a watch mechanism, that is, actively discover nodes added to the database cluster, thereby obtaining multiple nodes to which data of the third node can be migrated.
[0134] In some embodiments, the resource information includes resource information of the new node, and may also include resource information of the original nodes in the database cluster. The resource information may include the total memory of each node, the used memory, the total central processing unit (CPU), the used CPU, the total number of Pods in the resource pool, the used number of Pods, the total input / output (I / O) bandwidth, the occupied I / O size, and the basic network delay time.
[0135] In some embodiments, resource optimization information may be information used to minimize resource consumption and maximize resource utilization of database cluster nodes, for example, resource optimization information may be the remaining memory usage, remaining CPU usage, remaining number of available Pods, remaining I / O bandwidth, etc. of database cluster nodes. Network delay optimization information may be information used to minimize the response time of the database cluster, and the network delay optimization information may be determined based on the basic network delay time of the database cluster.
[0136] In some embodiments, after obtaining the resource information of each node in the database cluster, first resource information related to resource optimization and second resource information related to network delay optimization can be determined, and then the first resource information is subjected to a first processing to obtain resource optimization information, and the second resource information is subjected to a second processing to obtain network delay optimization information.
[0137] Exemplarily, the first resource information related to resource optimization may include the total memory, the used memory, the total CPU, the used CPU, the total number of Pods in the resource pool, the used number of Pods, the total I / O bandwidth, and the occupied I / O size. The first processing of the first resource information to obtain the resource optimization information may be to subtract the used memory from the total memory to obtain the remaining memory usage; to subtract the used CPU from the total CPU to obtain the remaining CPU available; to subtract the used number of Pods from the total resource pool to obtain the remaining number of available Pods; to subtract the occupied I / O size from the total I / O bandwidth to obtain the remaining I / O bandwidth, etc.
[0138] Exemplarily, the second resource information related to network delay optimization may be the basic network delay duration of the database cluster. The second processing of the second resource information to obtain the network delay optimization information may be to determine the inverse of the basic network delay duration, and the inverse of the basic network delay duration is determined as the network delay optimization information.
[0139] In some embodiments, based on resource optimization information and network delay optimization information, comprehensive heuristic information for migrating data of a third node in a database cluster to each new node can be determined, and the selection probability corresponding to each new node can be determined according to the comprehensive heuristic information and the pheromone concentration from the third node to the new node, wherein the pheromone concentration can be updated by an ant colony algorithm, and a new node with a selection probability greater than a preset threshold is determined as the second node.
[0140] Exemplarily, the resource information of each node of the database cluster includes M i,total : The total memory of the i-th node; M i,used : The memory usage of the i-th node; C i,total : The total CPU of the i-th node; C i,used : CPU usage of the i-th node; P i,total : The total number of Pods on the i-th node; P i,used : The number of Pods used by the i-th node; I i,total : The total I / O volume of the i-th node; I i,used : The occupied I / O size of the i-th node; D i,net : Basic network delay of the ith node; A i : The abnormal state of the i-th node (0 means normal, 1 means abnormal).
[0141] To minimize resource consumption and maximize resource utilization, resource optimization information includes the remaining resource amount M of each node. i,total -M i,used , C i,total -C i,used , P i,total -P i,used ,I i,total -I i,used ; For minimizing response time, the network delay optimization information can be the basic network delay D between nodes i,net The reciprocal of
[0142] Each resource optimization information and network delay optimization information can be normalized by the following formula (5):
[0143]
[0144] Among them, θ ij Indicates the resource optimization information or network delay optimization information of node j (third node) migrating to node i (new node), max(θ ij ) represents the maximum value corresponding to the resource optimization information or the maximum value corresponding to the network delay optimization information.
[0145] Next, the normalized resource optimization information or network delay optimization information is combined to form a comprehensive heuristic information, which can be determined by the weighted summation method shown in formula (6):
[0146]
[0147] in, Respectively represent the normalized values of the remaining memory usage, remaining CPU usage, remaining number of available Pods, remaining I / O bandwidth, and the reciprocal of the basic network delay. i Indicates the abnormal state of the i-th host. If the host is normal, then A i =0, otherwise A i =1, w6 is the weight of the abnormal state, which is usually set to a negative value to indicate a penalty. i As a penalty factor for comprehensive heuristic information, the probability of selecting abnormal nodes is reduced. The weights corresponding to each resource optimization information can be w1=0.4, w2=0.1, w3=0.3, w4=0.1, w5=0.1, and the abnormal state of the host is w6=1.
[0148] Comprehensive heuristic information It can be substituted into the selection probability formula (7):
[0149]
[0150] Among them, τ ij is the pheromone concentration from node i to node j, N i is the set of neighbor nodes of node i. Pheromone update includes evaporation and deposition. The evaporation process reduces the pheromone on all paths, while the deposition process increases the pheromone according to the path quality of the ant. The pheromone update formula can be expressed by formula (8):
[0151]
[0152] Among them, Δτ ij is the amount of pheromone deposited according to the path quality of ant k.
[0153] Through the above steps and formulas, combined with multiple key factors such as resource utilization, network latency and system stability, the ant colony algorithm can find the best deployment node (second node) in the K8S cluster to schedule the Redis instance to the best node for deployment.
[0154] It can be understood that by analyzing and processing the resource information of each node in the database cluster, resource optimization information and network delay optimization information are obtained, and based on the resource optimization information, network delay optimization information and the ant colony algorithm, a second node that has been optimized through resource optimization and network delay can be determined from new nodes actively discovered based on the resource status change notification mechanism, so that capacity expansion based on the determined second node can achieve the purpose of improving resource utilization and reducing response time.
[0155] In some embodiments of the present application, the scaling strategy includes the number of nodes to be shrunk in the database cluster, and the fourth node to be deleted in the database cluster corresponding to the number of nodes to be shrunk; based on this, based on the scaling strategy, the database cluster is scaled up or down, that is, step S103 can also be implemented by the following steps S1031B to S1032B, and each step is described separately below.
[0156] S1031B: Create a second data synchronization module corresponding to the fifth node in the database cluster where data is to be migrated.
[0157] In some embodiments, the fifth node may be other nodes in the database cluster excluding the fourth node that needs to be deleted or destroyed, the second data synchronization module may be one, the type of the second data synchronization module and the first data synchronization module may be the same, and the second data synchronization module may synchronize the data in each fifth node.
[0158] S1032B. Load the slot data of the fourth node into the second data synchronization module, and forward the slot data of the fourth node to the corresponding fifth node through the second data synchronization module.
[0159] In some embodiments, during the process of scaling down the database cluster, the slot data in the fourth node needs to be transferred to each fifth node. In this process, the slot data in the fourth node can be first loaded into the second data synchronization module, and then the data in the fourth node can be forwarded to each fifth node respectively through the second data synchronization module.
[0160] In some embodiments, the second data synchronization module can allocate slot data that needs to be migrated to each fifth node according to the index value order of the slot data. For example, if the slot data includes 100 and the fifth node includes 20, the slot data corresponding to 5 consecutive index values can be forwarded to the corresponding fifth node in turn, thereby achieving reasonable allocation of the slot data.
[0161] It can be understood that by creating a second data synchronization module corresponding to the fifth node in the database cluster to which data is to be migrated, loading the slot data of the fourth node into the second data synchronization module, and forwarding the slot data of the fourth node to the corresponding fifth node through the second data synchronization module, the impact of the slot data on the business demand response of the fifth node during the process of migrating to the fifth node can be avoided.
[0162] In an embodiment of the present application, the response time information of the database cluster to be expanded or reduced for different types of data access requests is obtained; based on the response time information and the type of data access request, the expansion and reduction strategy for the database cluster is determined; based on the expansion and reduction strategy, the database cluster is expanded or reduced. In this way, the expansion and reduction strategy of the database cluster is determined by the response time information of the database cluster for different types of data access requests and the type of data access requests, and a mechanism for automatically triggering the expansion and reduction of the database cluster based on the business needs of the database cluster and the business needs of the database processing is realized, thereby improving the intelligence of triggering the expansion and reduction of the database cluster and the efficiency of the expansion and reduction of the database cluster.
[0163] The following is an introduction to the implementation process of the embodiments of the present application in actual application scenarios.
[0164] In some embodiments, Figure 2 As shown, it is a flow chart of a method for dynamic scheduling and management of large-scale host node computing power integrated resources based on K8S provided in an embodiment of the present application, and the method includes:
[0165] S201. Use K8S technology to dynamically discover new host nodes (equivalent to the "new nodes" in other embodiments), and monitor the resource status data of the host nodes in the cluster in real time (equivalent to the "resource information" in other embodiments).
[0166] For host node resources, K8S technology is used for unified management to form a large-scale host cluster. The integrated computing power management system uses the watch mechanism provided by the K8S API server to monitor node changes. In the integrated computing power system, nodes are watched to monitor the node resource status in real time. It can discover newly added nodes and status changes of existing nodes. For the resource information of each host node, after the watch obtains the node, the status real-time monitor is used to obtain the indicators that mainly affect the performance and stability of Redis use, such as the total memory of each node host, the amount of memory used, the total CPU, the amount of CPU used, the total Pod, the amount of Pod used, the total I / O bandwidth, the I / O occupied size and the basic network delay.
[0167] S202: Based on the resource status data of the host nodes, determine the best deployed host node (equivalent to the "second node" in other embodiments) from the newly added host nodes.
[0168] The environment of the K8S cluster is modeled as a graph, in which nodes represent host nodes and edges represent connections between host nodes. The heuristic information for migrating each host node to a newly added host node can be determined based on the resource status data of each host node. Based on the heuristic information and the corresponding contents of formulas (6) to (8), the node with the highest selection probability among the newly added host nodes is determined as the best deployed host node.
[0169] The K8S-based large-scale host node computing power integrated resource dynamic scheduling management method provided in this application is particularly suitable for cloud computing and containerized application scenarios. This method normalizes and calculates multiple factors that affect the use of Redis, sets weight configurations, and uses the ant colony algorithm to calculate and obtain the optimal scheduling host node. It can fully utilize the resources of each host node while ensuring the normal and stable operation of the Redis cluster.
[0170] This application provides a method for real-time monitoring of Redis cluster data processing status and intelligent resource expansion and contraction, such as Figure 3 As shown, the method includes:
[0171] S301. The proxy cluster records the Redis fine-grained command processing delay.
[0172] like Figure 4 As shown, all business requests of client 401 will pass through proxy cluster proxy 402, and then be forwarded to Redis cluster 403 by proxy 402. After processing, Redis cluster 403 will return the result to proxy 402, and proxy 402 will return it to client 401.
[0173] Proxy 402 records all command types entered by the client. The recorded information is as follows:
[0174] Number of calls: records the number of times each command type is called.
[0175] Total time consumed: records the cumulative time consumed for each command to be sent from proxy 402 to Redis cluster 403 and then returned from Redis cluster 403 to the proxy, in seconds.
[0176] Average time consumption: Average time consumption = total time consumption / number of calls. In order to ensure the real-time and accuracy of the average time consumption, the total time consumption and number of calls of all commands will be reset every 10 minutes.
[0177] Average number of calls per second: Average number of calls per second = number of calls / call time. The total time and number of calls of all commands are reset every 10 minutes, and the call time is also reset.
[0178] It should be noted that, unlike the average time consumption statistics of Redis itself, which is the average time consumption of calculation and processing inside Redis, the average time consumption counted by the proxy is the processing time of the real data return, which involves the calculation and processing time of Redis itself, network transmission time consumption, network congestion, etc., and the judgment of Redis processing time consumption is more accurate. At the same time, the average time consumption of Redis itself is the average time consumption from the start of Redis, and the calculation result is inaccurate. As time goes by, the data will be averaged, and the sudden increase in time consumption cannot be accurately reflected. The average number of calls per second does not exist in the information of Redis itself.
[0179] S302. The data processing status real-time monitor obtains information collected by the proxy cluster and the Redis cluster in real time.
[0180] like Figure 5 The figure shows a structural diagram of a capacity expansion and contraction system provided by the present application. The capacity expansion and contraction system 500 includes a real-time status detector 501, a data storage device 502, a status analyzer 503 and a computing power scheduler 504. The real-time status detector 501 collects information of the proxy 402 and the Redis cluster 403 into a unified data storage device 502 every 10 seconds for the status analyzer 503 to analyze and process according to the collected information.
[0181] Among them, proxy 402 collects the average time consumption of all command types, Redis cluster 403 collects the average QPS of all shards, and the real-time status monitor 501 will also persist data for all Redis shards every 10 minutes, and calculate the memory occupied and element information of the persisted data in units of key, and sort it, including: parsing RDB files: first read and parse the RDB data file after data persistence. The RDB data file is a binary dump of the Redis database, which contains the serialized form of all keys and values; traverse key-value pairs: traverse all key-value pairs in the RDB file. Since the RDB file contains a snapshot of the database, this step can obtain all keys stored in the Redis instance; calculate the size of the key: for each key, calculate its encoded byte size. For keys of collection types (List, Set, Sorted Set, Hash), calculate the number of members and the total size.
[0182] S303. Generate a scaling strategy for multiple command types based on the information collected by the proxy cluster and the Redis cluster (equivalent to "determining a scaling strategy for the database cluster based on response time information and the type of data access request" in other embodiments).
[0183] Multiple command types include dangerous cross-shard scanning commands, single-key large memory query commands, single-key value large-scale query commands, and no high-latency commands. Among them, the expansion strategy corresponding to the dangerous cross-shard scanning commands can be determined by the corresponding content of formula (1); the expansion strategy corresponding to the single-key large memory query commands can be determined by the corresponding content of formula (2); the expansion strategy corresponding to the single-key value large-scale query commands can be determined by the corresponding content of formula (3); the shrinking strategy corresponding to the no high-latency command can be determined by the corresponding content of formula (4).
[0184] For the expansion strategy corresponding to the single-key large memory query command, after determining the number of nodes to be expanded, the corresponding shard can be obtained for the slot where the large memory key is located, and the large memory key on each shard is recorded. The number of large memory keys in the Nth shard is M. N , then the number of large memory keys migrated from each shard is Finally, through polling allocation, the slots are allocated to new shards one by one in sequence until all shards are allocated the same number of slots.
[0185] For the expansion strategy corresponding to the query command of a large number of values of a single key, after determining the number of nodes to be expanded, the corresponding shards can be obtained for the slots where the keys with large values are located, and the keys with large values on each shard are recorded. If the number of keys with large values on the Nth shard is M N Then, the number of keys with large values migrated from each shard is Finally, through random allocation, a new shard is randomly selected, the slot is assigned to it, and then the process is repeated until all slots are allocated.
[0186] After the number of nodes to be shrunk is determined for the scaling-down strategy corresponding to the command without high latency, the slots of the shrunk shards can be evenly distributed to the remaining shards, and then the remaining keys can be distributed to the remaining shards through the hash algorithm.
[0187] The method for real-time monitoring of Redis cluster data processing status and intelligent resource expansion and contraction strategy provided in this application obtains the status information related to the Redis cluster through a self-developed agent, dynamically obtains the trigger points for resource expansion and contraction, adapts to load changes in different business scenarios, dynamically adjusts the resource allocation of Redis, and achieves intelligent expansion and contraction of resources such as memory. When a single host cannot meet the requirements, it can realize hot expansion and contraction and dynamic scheduling of shards, ensuring that resource utilization is maximized within the performance baseline of the Redis cluster.
[0188] This application provides a Redis cluster expansion and contraction method based on elastic data synchronizer computing power scheduling. Figure 6 As shown, the method includes:
[0189] S601. Create an elastic data synchronizer.
[0190] When the computing power scheduler 504 receives the Redis cluster shard expansion strategy and determines to initiate shard expansion for a Redis cluster, such as Figure 7 As shown, the computing power scheduler 504 will evaluate and calculate the CPU, memory, storage, etc. used by the Redis shard that has been migrated out of the slot, and dynamically create an elastic data synchronizer 701 in the computing power resource pool as the data synchronizer for the Redis that has been migrated out of the slot. At the same time, a new Redis node 702 is created in the computing power resource pool to expand the Redis cluster. After the creation is completed, a migration instruction is sent to the proxy cluster proxy402 and the elastic data synchronizer 701, as shown in FIG. Figure 7 As shown, an elastic data synchronizer 701 is dynamically allocated to each Redis node, and a new Redis node 702 to be expanded is generated at the same time.
[0191] Elastic Data Synchronizer 701 can be created using the self-developed K8S operator. The creation process is as follows:
[0192] (1) Pre-check: Check whether Redis exists. If not, return an error. Verify whether the status of the Redis instance is normal. If not, return an error.
[0193] (2) Build annotations: Create an annotation mapping in K8S orchestration, including network rules, user information, VPC network, and other information;
[0194] (3) Obtain the Pod IP address of the master node of each shard: Traverse and find the node with the corresponding shard group ID and the master role. Obtain the Pod IP address of the node.
[0195] (4) Prepare environment variables: Build an environment variable list, including host name, metadata host, port, user, password and other information;
[0196] (5) Build startup parameters and commands: Format startup parameters and commands to initialize services in the container. Define port mapping: Define the port mapping of the container for monitoring and exposure of performance indicators.
[0197] (6) Prepare volume mounts and volumes: Create volume mounts and volumes for logs, configurations, and status to persist the logs and configurations generated by the container. Attach RDB directory volume: Add volume mounts and volumes for RDB persistence.
[0198] (7) Build a job template: Use the collected information to build a DtsJob Template Info structure, including annotations, tags, parameters, commands, environment variables, resource limits, volume mounts, and volumes.
[0199] (8) Create a data synchronization job: Use the K8S job API to create a K8S job based on the above template, which will start the elastic data synchronizer 701.
[0200] S602: Use the computing power scheduler to perform data migration based on the expansion and contraction strategy.
[0201] The computing power scheduler 504 can implement the basic structure and API of the source Redis node, the target Redis node and the elastic data synchronizer node. The computing power scheduler 504 saves the source node information, including slot, host, port, etc. and the target node information, including slot, host, RDB file name, etc.
[0202] The computing power scheduler 504 can refresh asynchronously, record and indicate which slots are ready to be cleaned, help block all operations on the slots being cleaned, and execute the deletion of slot data in another scheduled task.
[0203] The computing power scheduler 504 can also pause the proxy 402, so that the proxy 402 waits for the pause to achieve imperceptible migration.
[0204] When all data migration is completed, the computing power scheduler 504 can obtain status information in real time and pause proxy 402. Proxy 402 will release the route of the elastic data synchronizer 701 and destroy the elastic data synchronizer 701 at the same time, release resources in real time, and asynchronously refresh and delete the data on the original slot.
[0205] In addition, during the entire slot migration process, the computing power scheduler 504 provides real-time monitoring and logging so that the administrator can track the migration status. If there is a problem during the migration process, the computing power scheduler 504 can be used to troubleshoot and re-trigger the migration task.
[0206] Through the above process, the Redis cluster can achieve online expansion of the Redis cluster shards without downtime, and use the computing power scheduler to ensure the consistency of data migration and the accuracy of the cluster status.
[0207] The Redis cluster expansion and contraction method based on elastic data synchronizer computing power scheduling provided in the present application dynamically schedules slave nodes for data migration through the computing power scheduler. This method performs real-time node switching in the elastic data synchronizer to ensure that the master node is not affected by any data migration, and the user can hardly perceive any impact. It effectively solves the business congestion problem that may be caused by large value migration and improves the availability and business continuity of the cluster.
[0208] This application monitors the state of the Redis cluster in multiple dimensions and combines multi-factor measurement technology to conduct real-time predictive analysis. This method can accurately trigger the elastic expansion and contraction mechanism of the cluster, ensuring that the cluster can automatically adjust resource allocation when facing different load demands to adapt to business changes. Compared with the traditional open source Redis Cluster online expansion solution, the method provided by this application does not require manual startup of new Redis instances and manual data migration, nor does it require business interruption to update the client's IP connection information. The automated expansion process reduces the errors and complexity of manual operations, while ensuring seamless data migration and node expansion, thereby significantly improving operation and maintenance efficiency and user experience.
[0209] The present application also provides a system for scaling up and down a database cluster. Figure 8 A schematic diagram of the structure of a database cluster expansion and contraction system provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the scaling system 800 for a database cluster includes:
[0210] The cluster status real-time detection module 801 is used to obtain the response time information of the database cluster to different types of data access requests;
[0211] A cluster status analysis module 802, configured to determine a scaling strategy for the database cluster based on the response time information and the type of the data access request;
[0212] The cluster computing power scheduling module 803 is used to scale the database cluster based on the scaling strategy.
[0213] In some embodiments, the expansion and contraction strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a first type for global access to the database cluster; the cluster status analysis module 802 includes:
[0214] A first determining submodule, used to determine first access frequency information of the first type of data access request;
[0215] A second determination submodule, configured to determine whether the database cluster meets a first expansion trigger condition according to the first access frequency information and the response time information of the first type of data access request;
[0216] The third determination submodule is used to determine the number of nodes to be expanded in the database cluster according to the first access frequency information, the response time information of the first type of data access request, and the number of nodes in the database cluster if the database cluster meets the first expansion trigger condition.
[0217] In some embodiments, the scaling strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a second type for key data access of the first specific type; the cluster state analysis module 802 further includes:
[0218] a fourth determination submodule, configured to determine a first memory occupation size of key data in a key-value pair storage structure accessed by the second type of data access request; the second type being a data access request for a specific type of key data;
[0219] a fifth determining submodule, configured to determine a second memory occupation size satisfying a memory occupation condition from the first memory occupation size;
[0220] a sixth determination submodule, configured to determine whether the database cluster satisfies a second expansion trigger condition according to the response time information of the second type of data access request;
[0221] The seventh determination submodule is used to determine the number of nodes to be expanded in the database cluster according to the number of key data corresponding to the second memory occupancy size, the response time information of the second type of data access request, and the number of nodes in the database cluster if the database cluster meets the second expansion trigger condition.
[0222] In some embodiments, the expansion and contraction strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a third type of key-value access for a second specific type; the cluster state analysis module 802 further includes:
[0223] an eighth determination submodule, configured to determine a first number of key values corresponding to each key data in the key-value pair storage structure accessed by the third type of data access request; the third type being a data access request for a specific type of key value;
[0224] A ninth determination submodule, configured to determine, from each of the first quantities, a second quantity that satisfies a key value quantity threshold condition;
[0225] a tenth determining submodule, configured to determine whether the database cluster satisfies a third expansion triggering condition according to the response time information of the third type of data access request;
[0226] An eleventh determination submodule is used to determine the number of nodes of the database cluster to be expanded according to the second number, the response time information of the third type of data access request, and the number of nodes of the database cluster if the database cluster meets the third expansion trigger condition.
[0227] In some embodiments, the scaling strategy includes the number of nodes to be scaled down in the database cluster; the type of data access request includes a fourth type of local access to the database cluster; the cluster status analysis module 802 further includes:
[0228] A twelfth determining submodule, configured to determine second access frequency information of the fourth type of data access request;
[0229] A thirteenth determination submodule, configured to determine whether the database cluster satisfies a first shrinking trigger condition according to the second access frequency information and the response time information of the fourth type of data access request;
[0230] A fourteenth determination submodule is used to determine the number of nodes of the database cluster to be shrunk according to the second access frequency information, the response time information of the fourth type of data access request, and the number of nodes of the database cluster if the database cluster meets the first shrinking trigger condition.
[0231] In some embodiments, the cluster computing power scheduling module 803 includes:
[0232] A fifteenth determination submodule is used to determine the second nodes corresponding to the number of nodes to be expanded, and to create first data synchronization modules corresponding to the third nodes in the database cluster to which data is to be migrated; the first data synchronization module is used to synchronize the slot data in the corresponding third nodes;
[0233] The first data transfer submodule is used to load the slot data to be migrated out of the third node into the corresponding first data synchronization module, and forward the corresponding slot data to be migrated out to the corresponding second node in sequence through the first data synchronization module.
[0234] In some embodiments, the fifteenth determining submodule includes:
[0235] A first monitoring unit is used to monitor the state change information of the database cluster in real time based on a resource state change notification mechanism; the state change information includes a new node connected to the database cluster; the new node includes at least one;
[0236] A first acquisition unit is configured to acquire resource information of each node in the database cluster if it is determined that a new node is connected to the database cluster; the resource information includes resource information of the new node;
[0237] An analysis and processing unit, used to analyze and process the resource information to obtain resource optimization information and network delay optimization information;
[0238] The first determination unit is used to determine the second node from at least one of the new nodes based on the resource optimization information and the network delay optimization information and in combination with an ant colony algorithm.
[0239] In some embodiments, the expansion and contraction strategy includes the number of nodes to be contracted in the database cluster, and the fourth node to be deleted in the database cluster corresponding to the number of nodes to be contracted; the cluster computing power scheduling module 803 also includes:
[0240] A first creation submodule, used to create a second data synchronization module corresponding to a fifth node in the database cluster to which data is to be migrated;
[0241] The second data transfer submodule is used to load the slot data of the fourth node into the second data synchronization module, and forward the slot data of the fourth node to the corresponding fifth node through the second data synchronization module.
[0242] In some embodiments, the response time information includes an average response time of the database cluster to the data access request; the cluster status real-time detection module includes:
[0243] A data access request receiving and sending submodule is used to obtain a reference type of data access request using a proxy cluster, and forward the reference type of data access request to the first node of the corresponding database cluster; the reference type is any type of data access command;
[0244] A sixteenth determination submodule, configured to determine a response duration of the first node according to a first time when the proxy cluster sends the reference type of data access request and a second time when the proxy cluster receives feedback data of the first node in response to the reference type of data access request;
[0245] A seventeenth determination submodule is used to determine an average response time of the database cluster to the data access request of the reference type based on the response time and the access frequency information of the data access request of the reference type.
[0246] It should be noted that the description of the expansion and contraction system of the database cluster in the embodiment of the present application is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment, so it is not repeated. For technical details not disclosed in the embodiment of the system, please refer to the description of the method embodiment of the present application for understanding.
[0247] Accordingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for scaling up and down a database cluster provided in the above embodiment is implemented.
[0248] The present application also provides a device for scaling up and down a database cluster. Fig. 9 A schematic diagram of the composition structure of a database cluster expansion and contraction device provided in an embodiment of the present application, such as Fig. 9 As shown, the scaling device 900 for a database cluster includes: a memory 901, a processor 902, a communication interface 903, and a communication bus 904. The memory 901 is used to store scaling instructions for executable database clusters; the processor 902 is used to execute the scaling instructions for executable database clusters stored in the memory to implement the scaling method for a database cluster provided in the above embodiment.
[0249] The embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 902 of the database cluster scaling device 900 to complete the steps of the database cluster scaling method provided in the above embodiment.
[0250] The description of the above database cluster expansion and contraction device, storage medium and computer program product embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the database cluster expansion and contraction device, storage medium and computer program product embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0251] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for scaling up and down a database cluster, characterized in that: include: Obtain the response time information of the database cluster to be expanded or reduced for different types of data access requests; Determining a scaling strategy for the database cluster based on the response time information and the type of the data access request; Based on the expansion and contraction strategy, the database cluster is expanded and contracted.
2. The method according to claim 1, characterized in that The expansion and contraction strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a first type for global access to the database cluster; The determining, based on the response time information and the type of the data access request, a scaling strategy for the database cluster includes: Determining first access frequency information of the first type of data access request; determining, according to the first access frequency information and the response time information of the first type of data access request, whether the database cluster meets a first expansion trigger condition; If the database cluster meets the first expansion trigger condition, the number of nodes of the database cluster to be expanded is determined according to the first access frequency information, the response time information of the first type of data access request, and the number of nodes of the database cluster.
3. The method according to claim 1, characterized in that The expansion and contraction strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a second type for key data access of the first specific type; The determining, based on the response time information and the type of the data access request, a scaling strategy for the database cluster includes: Determine a first memory occupancy size of key data in the key-value pair storage structure accessed by the second type of data access request; Determine a second memory occupation size that meets a memory occupation condition from the first memory occupation size; Determining whether the database cluster meets a second expansion trigger condition according to the response time information of the second type of data access request; If the database cluster meets the second expansion trigger condition, the number of nodes of the database cluster to be expanded is determined according to the number of key data corresponding to the second memory occupancy size, the response time information of the second type of data access request, and the number of nodes of the database cluster.
4. The method according to claim 1, characterized in that: The expansion and contraction strategy includes the number of nodes to be expanded in the database cluster; the type of the data access request includes a third type of key-value access for a second specific type; The determining, based on the response time information and the type of the data access request, a scaling strategy for the database cluster includes: Determine a first quantity of key values corresponding to each key data in the key-value pair storage structure accessed by the third type of data access request; Determine a second quantity that satisfies a key value quantity threshold condition from each first quantity; Determining whether the database cluster meets a third expansion trigger condition according to the response time information of the third type of data access request; If the database cluster meets the third expansion trigger condition, the number of nodes of the database cluster to be expanded is determined according to the second number, the response time information of the third type of data access request, and the number of nodes of the database cluster.
5. The method according to claim 1, characterized in that The scaling strategy includes the number of nodes to be scaled down in the database cluster; the type of data access request includes a fourth type of local access to the database cluster; The determining, based on the response time information and the type of the data access request, a scaling strategy for the database cluster includes: Determining second access frequency information of the fourth type of data access request; determining, according to the second access frequency information and the response time information of the fourth type of data access request, whether the database cluster meets a first shrinking trigger condition; If the database cluster meets the first shrinking trigger condition, the number of nodes of the database cluster to be shrunk is determined according to the second access frequency information, the response time information of the fourth type of data access request, and the number of nodes of the database cluster.
6. The method according to any one of claims 1 to 5, characterized in that The expansion / contraction strategy includes the number of nodes to be expanded in the database cluster; and the expansion / contraction processing of the database cluster based on the expansion / contraction strategy includes: Determine the second nodes corresponding to the number of nodes to be expanded, and create first data synchronization modules corresponding to the third nodes in the database cluster where data is to be migrated; the first data synchronization modules are used to synchronize the slot data in the corresponding third nodes; The slot data to be migrated out in the third node is loaded into the corresponding first data synchronization module, and the corresponding slot data to be migrated out is forwarded to the corresponding second node in sequence through the first data synchronization module.
7. The method according to claim 6, characterized in that The determining the second node corresponding to the number of nodes to be expanded includes: Based on the resource status change notification mechanism, the status change information of the database cluster is monitored in real time; the status change information includes a new node connected to the database cluster; the new node includes at least one; If it is determined that a new node is connected to the database cluster, resource information of each node in the database cluster is obtained; the resource information includes resource information of the new node; Analyzing and processing the resource information to obtain resource optimization information and network delay optimization information; Based on the resource optimization information and the network delay optimization information, and in combination with an ant colony algorithm, the second node is determined from at least one of the new nodes.
8. The method according to any one of claims 1 to 5, characterized in that: The expansion and contraction strategy includes the number of nodes to be contracted in the database cluster, and the fourth node to be deleted in the database cluster corresponding to the number of nodes to be contracted; The step of scaling the database cluster based on the scaling strategy includes: Creating a second data synchronization module corresponding to a fifth node in the database cluster where data is to be migrated; The slot data of the fourth node is loaded into the second data synchronization module, and the slot data of the fourth node is forwarded to the corresponding fifth node through the second data synchronization module.
9. The method according to any one of claims 1 to 5, characterized in that: The response time information includes an average response time of the database cluster to the data access request; The obtaining of the response time information of the database cluster to different types of data access requests includes: Using a proxy cluster to obtain a data access request of a reference type, and forwarding the data access request of the reference type to a first node of the corresponding database cluster; the reference type is any type of data access command; Determine a response duration of the first node according to a first time when the proxy cluster sends the reference type of data access request and a second time when the proxy cluster receives feedback data of the first node in response to the reference type of data access request; Based on the response time and the access frequency information of the data access request of the reference type, an average response time of the database cluster to the data access request of the reference type is determined.
10. A system for scaling up and down a database cluster, characterized in that: include: The cluster status real-time detection module is used to obtain the response time information of the database cluster for different types of data access requests; A cluster status analysis module, used to determine a scaling strategy for the database cluster based on the response time information and the type of the data access request; The cluster computing power scheduling module is used to scale the database cluster based on the scaling strategy.
11. A device for scaling up and down a database cluster, characterized in that: include: A memory, used to store instructions for executing a method for scaling up or down a database cluster; The processor is used to implement the method described in any one of claims 1 to 9 when executing the instructions of the method for scaling up and down the executable database cluster stored in the memory.
12. A computer-readable storage medium, characterized in that: The method for scaling up and down a database cluster is stored, and is used to cause a processor to execute the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 9.
Citation Information
Cited By
DataX distributed data synchronization method and device based on Kubernetes and electronic equipment
CN120470060A
Redis cluster federated data routing method and device, equipment, medium and product
CN120811974A
Redis cluster federation data routing method and device, equipment, medium and product
CN120811974B
Method, system and device for scaling up / down database cluster, storage medium, and computer program product
WO2026145576A1