Method and apparatus for implementing database partitioning and distributed database management system
By dynamically adjusting the hash value of the database partition, the performance bottleneck problem caused by the unbalanced hash key data in traditional database partition technology is solved, efficient database partition management is achieved, and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202010477561.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-05-29
AI Technical Summary
When traditional database partitioning technology faces big data scenarios and complex analysis needs, it is easy to cause partial hot partitioning due to imbalance of hash key data, resulting in database performance bottlenecks.
By obtaining the operation information of the database partition, establish the corresponding relationship between the value of the partition key after the hash function hash and the partition, dynamically adjust the hash value of the partition key corresponding to the partition to minimize it, and dynamic partition storage is performed according to the adjusted relationship.
Dynamic database partitioning is implemented, which reduces operation and maintenance costs in the cloud computing era, improves database query efficiency and performance, and avoids performance bottlenecks caused by hot partitioning.
Smart Images

Figure CN113297315B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, cloud computing technology, and particularly to a method and apparatus for implementing database partitioning and a distributed database management system. Background Art
[0002] As the amount of data in a database grows larger, database partitioning emerges as the times require. Database partitioning is mainly used to solve the problems of low query efficiency and degraded performance caused by the increasing amount of data. Using the database partitioning function, tables, indexes, etc. are further divided into segments, and these database objects are partitions. These partitions can be managed either individually or collectively.
[0003] In traditional database partitioning technologies, the hash algorithm is mostly used to evenly hash data into different partitions, and each partition is only responsible for its own data. However, with the complexity of big data scenarios and the increasingly complex analytical requirements nowadays, in some cases where the data of the hash key itself is unbalanced, it will cause some hot partitions, thus creating a bottleneck in the database. Summary of the Invention
[0004] This application provides a method and apparatus for implementing database partitioning and a distributed database management system, which can achieve dynamic database partitioning and reduce the operation and maintenance costs in the cloud computing era.
[0005] An embodiment of the present invention provides a method for implementing database partitioning, including:
[0006] Obtain the running information of the database partition, where the running information of the database partition includes: the running information of the partition and the running information of the computing node; multiple partitions are included on one computing node;
[0007] Generate a dynamic partition according to the obtained running information of the database partition and the corresponding relationship between the partition and the value obtained by hashing the partition key through a hash function.
[0008] In an exemplary instance, it further includes:
[0009] Write the running information of the database partition into the memory table of the distributed database.
[0010] In an exemplary instance, the generating a dynamic partition according to the obtained running information of the database partition and the corresponding relationship between the partition and the value obtained by hashing the partition key through a hash function includes:
[0011] Establish the corresponding relationship between the value obtained by hashing the partition key through the hash function and the partition;
[0012] From the established correspondence relationship, adjust the value obtained by hashing the partition key corresponding to the current database partition according to the running information of the partition to minimize it;
[0013] Perform dynamic partition storage according to the adjusted correspondence relationship between the partition and the value obtained by hashing the partition key.
[0014] In an exemplary instance, minimizing the value obtained by hashing the partition key corresponding to the adjusted partition and performing dynamic partition storage according to the correspondence relationship between the adjusted partition and the value obtained by hashing the partition key includes:
[0015] Determine the distribution probability of the action according to the partition, the value obtained by hashing the partition key, and the running information of the computing node where the partition is located; and store data in the partition in the corresponding relationship corresponding to the optimal action value according to the determined distribution probability.
[0016] In an exemplary instance, determining the distribution probability of the action and storing data in the partition in the corresponding relationship corresponding to the optimal action value according to the determined distribution probability includes:
[0017] Select, from all actions in the action space, the action with the smallest calculated O value that satisfies the pre-established constraint conditions, and adjust the current state to the state corresponding to the smallest O;
[0018] Perform dynamic partition storage according to the correspondence relationship between the value obtained by hashing the partition key corresponding to the partition in the state corresponding to the smallest O value;
[0019] Among them, the action space A is: A = {a 1 , a 2 ,..., a τ}, and each action a in the action space is used to adjust the combination of the values obtained by hashing the partition key corresponding to a certain partition;
[0020] Among them, represents the overall variance of the partition load; represents how many q i correspond to p i . Among them, o 1 and o 2 are weights; N represents the number of partitions.
[0021] In an exemplary instance, the established correspondence relationship between the value obtained by hashing the partition key and the partition is:
[0022] p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on a Markov decision process, HASH represents a hashing operation, key represents a partition key, and mod represents a modulo operator.
[0023] In an exemplary instance, after the dynamic partitioning, it further includes:
[0024] Recording the information of this generation process, generating a new partitioning strategy and updating the previous partitioning strategy.
[0025] In an exemplary instance, it further includes: writing the new partitioning strategy into a storage table of a distributed database.
[0026] In an exemplary instance, before updating the previous partitioning scheme, it further includes:
[0027] If the overall variance of the load of the partition decreases and the write query performance of the overall database corresponding to the number of reduced values of the partition keys corresponding to all partitions after being hashed by the hash function improves, then give a reward;
[0028] If the write and query performance of the overall database decreases after an action, then give a penalty.
[0029] In an exemplary instance, it further includes:
[0030] Accumulating the reward values of each iteration in multiple iterations, which is used to represent the return of the dynamic partitioning effect after multiple iterations.
[0031] This application also provides a computer-readable storage medium storing computer-executable instructions for executing the method for implementing database partitioning described in any one of the above.
[0032] This application further provides a device for implementing database partitioning, including a memory and a processor, where the memory stores the following instructions executable by the processor: steps for executing the method for implementing database partitioning described in any one of the above.
[0033] This application further provides a device for implementing database partitioning, including: an acquisition module and a processing module; where
[0034] The acquisition module is configured to obtain the running information of the database partitioning, where the running information of the database partitioning includes: the running information of the partition and the running information of the computing node; there are multiple partitions on one computing node;
[0035] The processing module is configured to generate a dynamic partition according to the obtained running information of the database partitioning and the corresponding relationship between the partition and the value of the partition key after being hashed by the hash function.
[0036] In an exemplary instance, the processing module is specifically configured as:
[0037] Establish the correspondence between the value obtained by hashing the partition key through a hash function and the partition; from the established correspondence, adjust the value obtained by hashing the partition key corresponding to the partition according to the running information of the partition of the current database to minimize it, and perform dynamic partition storage according to the adjusted correspondence between the partition and the value obtained by hashing the partition key.
[0038] In an exemplary instance, the establishment of the correspondence between the value obtained by hashing the partition key through a hash function and the partition in the processing module includes:
[0039] p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on the Markov decision process, HASH represents the hashing operation, key represents the partition key, and mod represents the modulo operator.
[0040] In an exemplary instance, the device further includes an iteration module, which is configured as:
[0041] Record the information of the current generation process, generate a new partition policy and update the previous partition policy.
[0042] This application also provides a distributed database management system, including: a plurality of computing nodes, and the device for implementing database partitioning described in any one of the above.
[0043] Through collecting the running state information of the database partition (including partitions and computing nodes) in the embodiments of this application, the advantages and disadvantages of the partition effect are well evaluated. This application not only realizes dynamic partitioning, but also further automatically adjusts the partition according to the data situation of the database, thereby realizing true dynamic partitioning and greatly reducing the operation and maintenance cost in the cloud computing era.
[0044] Furthermore, according to the running data of the computing nodes and partitions of the distributed database collected, the agent in the Markov decision process senses the running situation of the distributed database partition, and stores data in the partitions in the correspondence corresponding to the optimal action value, truly realizing dynamic database partitioning and reducing the operation and maintenance cost in the cloud computing era. Compared with the dynamic partitioning with fixed rules in the related art, this application is more flexible and effective.
[0045] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings are used to provide a further understanding of the technical solution of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application, and do not constitute a limitation to the technical solution of the present application.
[0047] Figure 1 It is a schematic flowchart of the method for implementing database partitioning in the present application;
[0048] Figure 2 It is a schematic diagram of the composition structure of the device for implementing database partitioning in the present application;
[0049] Figure 3 It is a schematic flowchart of the embodiment for implementing database partitioning in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the embodiments of the present application will be described in detail below with reference to the drawings. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily.
[0051] In a typical configuration of the present application, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0052] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0053] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0054] The steps illustrated in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions. And although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0055] In the methods of the related art for implementing database partitioning, data is stored in corresponding partitions according to the partitioning key of the data and fixed data partitioning rules, and partitioning can be customized, so that re-partitioning can be achieved without modifying the application code of the database. In this way, in order to implement customized partitioning, different partitioning strategies need to be customized for different business requirements. However, in the cloud computing scenario, from one field such as the financial field to another field such as the education field and then to another field such as the e-commerce field, etc., there are numerous database instances, numerous requirements, and diverse application methods. Different requirements and different data will all cause problems such as inability to customize partitioning and analyzing and partitioning partitions one by one.
[0056] In database processing, each row in a table is uniquely determined by the primary key (PK). When creating a table, the columns that make up the primary key must be specified, and these columns are called primary key columns. The primary key columns must have values, and it is necessary to ensure that the combination of the values of the primary key columns can uniquely determine a row. In the related art, a memory table divides the table into different data partitions to achieve load balancing of the stored data, and the partitioning granularity of the data partitions is the first column of the primary key, which is the partitioning key of the data.
[0057] Figure 1 The flowchart of the method for implementing database partitioning in this application is as Figure 1 shown, including:
[0058] Step 100: Obtain the running information of the database partition.
[0059] In an exemplary instance, the running information of the database partition includes: the running information of the partition and the running information of the computing node (worker). Among them, the main body of the partition is the computing node. One computing node includes multiple partitions, and a distributed database management system manages numerous partitions on multiple computing nodes. A distributed database is a logically unified database formed by connecting multiple physically dispersed databases with a computer network and is managed by a unified database management system (also called a distributed database management system).
[0060] In an exemplary instance, a timing task can be set through a timer in the distributed database management system to periodically obtain the running information of the database partition in a loop.
[0061] In an exemplary instance, the running information of partitions includes, but is not limited to, for example: the queries per second (QPS) of partitions, the data volume, the data distribution of each partition key, etc. The running information of the same partition can be represented by a data vector. Assuming the number of partitions is n, the running information of partitions can be expressed as: P = {p 1 , p 2 ,..., p n}, where P represents the set of running data of all partitions, and p i represents the running data vector of partition i.
[0062] In an exemplary instance, the running information of computing nodes includes, but is not limited to, for example: CPU occupancy rate, input / output (IO) latency, the number of IO per second transmitted by storage (IOPS), memory occupancy, etc. The running information of the same computing node can be represented by a data vector. Assuming the number of computing nodes is m, the running information of computing nodes can be expressed as: W = w 1 , w 2 ,..., w m , where W represents the set of running data of all computing nodes, and w i represents the running data vector of computing node i.
[0063] In an exemplary instance, this step further includes: writing the obtained running information of the database partitions into the memory table of the distributed database. In this way, low-latency writing and querying are achieved.
[0064] Step 101: Generate dynamic partitions according to the corresponding relationship between the obtained running information of the database partitions and the values after hashing the partitions and partition keys through a hash function.
[0065] To solve the problem of database bottleneck caused by uneven traffic of partition keys. In the embodiments of the present application, the dynamic partition can be expressed as: p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on the Markov decision process, that is, establishing the corresponding relationship between the values after hashing the partition keys and the partitions. Where HASH represents the hashing operation, key represents the partition key, and mod represents the modulo operator.
[0066] The Markov Decision Process (MDP) is constructed based on the interaction objects of a group of agents and the environment. The agent perceives the current system state, implements actions on the environment according to the policy, thereby changing the environment, and can further obtain rewards and punishments according to the changed environment state. Markov in MDP means that the state s at time t t is only related to the state s at time (t - 1) t-1 and the action a t-1 and has nothing to do with the history.
[0067] Several concepts involved in MDP are briefly introduced as follows: State refers to the set of states where the agent is in each step; Action refers to the set of actions that the agent can execute in each step; Transition probability refers to the probability that the agent will transfer to state s' after executing action a when it is in state s; Reward refers to the immediate reward value obtained after the agent transfers to state s' after executing action a when it is in state s; Policy refers to the probability that the agent should execute action a when it is in state s.
[0068] In an exemplary instance, this step may include:
[0069] Establish the correspondence between the value of the partition key after being hashed by the hash function and the partition;
[0070] From the established correspondence, adjust the value of the partition key corresponding to the partition after being hashed by the hash function according to the running information of the partition of the current database to minimize it, and perform dynamic partition storage according to the adjusted correspondence between the partition and the value of the partition key after being hashed by the hash function.
[0071] In an exemplary instance, the generation of dynamic partitions in the step can be established based on, such as, Markov decision, etc.
[0072] In an exemplary instance, from the established correspondence, adjust the value of the partition key corresponding to the partition after being hashed by the hash function according to the running information of the partition of the current database to minimize it, and perform dynamic partition storage according to the adjusted correspondence between the partition and the value of the partition key after being hashed by the hash function, including:
[0073] Determine the distribution probability of the action according to the partition, the value of the partition key after being hashed by the hash function, and the running information of the computing node where the partition is located;
[0074] Store the data in the partition in the correspondence corresponding to the optimal action value according to the determined distribution probability.
[0075] In an exemplary instance, the value q obtained by hashing the partition key through a hash function can be expressed as: q = HASH(key) mod n, where HASH represents the hashing operation, key represents the partition key, and mod represents the modulo operator.
[0076] In an exemplary instance, based on Markov decision-making, the state space (State) of this application is set as: S = {s 1 , s 2 ,..., s τ}, where each s represents various corresponding relationships between the partition p i and the value q i obtained by hashing the partition key through a hash function.
[0077] Taking partition = 3 as an example, then each partition has corresponding cases. Taking the partition p 1 as an example, the various corresponding relationships between the partition p 1 and the value q i obtained by hashing the partition key through a hash function include (the following symbol ":" represents correspondence):
[0078] p 1 : [q 1 , p 1 : [q 2 , p 1 : [q 3 p 1 : [q 1 , q 2 , p 1 : [q 2 , q 3 , p 1 : [q 1 , q 2 , q 3 ;
[0079] That is to say, if it is assumed that the number of partitions is n, then the state space S of the Markov decision-making process is a finite set, with a total of τ = n · (2 n - 1) states.
[0080] In an exemplary instance, based on the partition, the value obtained by hashing the partition key through a hash function, and the running information of the computing node where the partition is located, determine the distribution probability of the action, and store data in the partition corresponding to the optimal action value according to the determined distribution probability, including:
[0081] First, establish constraint conditions in advance, such as those shown in formulas (2) to (3):
[0082]
[0083]
[0084]
[0085] Among them, represents the overall variance of the partition load; represents p i corresponding to how many qs i . Among them, o 1 and o 2 are weights.
[0086] Then, select the action with the smallest O value calculated according to formula (4) from all actions in the action space that satisfies the pre-established constraint conditions, and adjust the current state to the state corresponding to the smallest O;
[0087] Perform dynamic partition storage according to the corresponding relationship between the values obtained by hashing the partition keys corresponding to the partitions in the state corresponding to the smallest O value.
[0088] Among them, set the action space (Action) of this application as: A = {a 1 , a 2 ,..., a τ}, and each action a in the action space is used to adjust the q i corresponding to p i combination.
[0089] In an exemplary instance, the adjustment range may include, for example: each time it can only be increased or decreased. If it is increased, try to increase the q i that has appeared before; minimize the number of qs i corresponding to p i , etc. In this way, while ensuring the data writing performance of the database, the query performance of the database is further optimized, and the partitions involved in the query are reduced.
[0090] The method for implementing database partitioning provided by this application can well evaluate the advantages and disadvantages of partitioning by collecting the running status information of database partitions (including partitions and computing nodes). This application not only realizes dynamic partitioning, but also further automatically adjusts the partitions according to the data situation of the database, thus realizing true dynamic partitioning and greatly reducing the operation and maintenance costs in the cloud computing era. More specifically, according to the running data of the computing nodes and partitions of the distributed database collected, the agent in the Markov decision process senses the running situation of the distributed database partitioning, and stores the data in the partitions corresponding to the optimal action values, truly realizing dynamic database partitioning and reducing the operation and maintenance costs in the cloud computing era. Compared with the dynamic partitioning with fixed rules in the related technology, this application is more flexible and effective.
[0091] In an exemplary instance, this application may further include:
[0092] After this dynamic partitioning (which can also be called the strategy of dynamic load balancing adjustment) is completed, record the information of the process of generating this dynamic partitioning, that is, the Markov decision process, generate a new partitioning strategy and update the previous partitioning strategy to prepare for the next iteration, that is, database partitioning.
[0093] In an exemplary instance, write the new partitioning strategy into the Memory Table of the distributed database. In this way, the distributed database management system will perform data writing and query according to the latest partitioning scheme, truly realizing dynamic database partitioning and reducing the operation and maintenance costs in the cloud computing era.
[0094] Through continuous iterative update process, this application makes the partitioning strategy better and better, so that the cloud database can make better decisions and scheduling.
[0095] In an exemplary instance, before updating the existing partitioning scheme, this application may further include:
[0096] If the overall variance of the load of the calculated partitions decreases, and the write and query performance of the overall database corresponding to all the reduced numbers of p i corresponding q i increases, then give a reward; if the write and query performance of the overall database decreases after the action, then give a penalty, that is, the reward is negative. It is expressed as shown in formula (5):
[0097] R = R(s t , a t , s t+1 ) = ΔO (5)
[0098] As shown in formula (5), R represents the reward and punishment result; after taking an action, assuming that the state space changes from state s t to state s t+1 and if the O corresponding to state s t+1 is less than the O corresponding to state s t , then this action is considered effective and a reward is given.
[0099] This application evaluates the effectiveness of this action based on the latest operating conditions of the distributed database partition, and obtains the intensity of rewards and punishments, thereby better ensuring the improvement and effectiveness of subsequent strategies. In this way, the strategy of the database partition is continuously corrected, making the implementation of the database partition more and more effective.
[0100] In an exemplary example, this application further includes:
[0101] According to the accumulation of the reward value R in each iteration of multiple iterations, determine the return G representing the dynamic partitioning effect after multiple iterations, as shown in formula (6): t
[0102]
[0103] The return G realizes the effect of adopting the database partitioning method of this application within a period of time, providing a relatively intuitive reference for subsequent improvements.
[0104] This application also provides a computer-readable storage medium storing computer-executable instructions for executing the method for implementing database partitioning described in any one of the above.
[0105]
[0105] This application further provides a device for implementing database partitioning, including a memory and a processor, wherein the memory stores the following instructions executable by the processor: steps for executing the method for implementing database partitioning described in any one of the above.
[0106] Figure 2 is a schematic structural diagram of the composition of the device for implementing database partitioning of this application, as Figure 2 shown, including: an acquisition module, a processing module; wherein,
[0107] The acquisition module is configured to obtain the operation information of the database partition;
[0108] The processing module is configured to generate a dynamic partition according to the obtained operation information of the database partition and the corresponding relationship between the partition and the value obtained by hashing the partition key through a hash function.
[0109] In an exemplary instance, the running information of the database partitions includes: the running information of the partitions and the running information of the computing nodes. Assuming the number of partitions is n, the running information of the partitions can be expressed as: P = {p 1 , p 2 ,..., p n}, where P represents the set of all partitions, and p i represents the running data vector of partition i. Assuming the number of computing nodes is m, the running information of the computing nodes can be expressed as: W = w 1 , w 2 ,..., w m , where W represents the set of all computing nodes, and w i represents the running data vector of computing node i.
[0110] In an exemplary instance, the acquisition module is specifically set as:
[0111] Set a timing task through a timer to periodically obtain the running information of the database partitions in a loop.
[0112] In an exemplary instance, the acquisition module is also set as:
[0113] Write the obtained running information of the database partitions into the memory table of the distributed database.
[0114] In an exemplary instance, the processing module is specifically set as:
[0115] Establish the correspondence between the value obtained by hashing the partition key through a hash function and the partition; from the established correspondence, adjust the value obtained by hashing the partition key corresponding to the partition according to the running information of the current database partitions to minimize it, and perform dynamic partition storage according to the adjusted correspondence between the partition and the value obtained by hashing the partition key.
[0116] In an exemplary instance, establishing the correspondence between the value obtained by hashing the partition key through a hash function and the partition in the processing module includes:
[0117] p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on the Markov decision process, that is, establishing the correspondence between the value obtained by hashing the partition key and the partition. Where HASH represents the hash operation, key represents the partition key, mod represents the modulo operator, and q represents the value obtained by hashing the partition key.
[0118] In an exemplary instance, performing dynamic partition storage according to the adjusted correspondence between the partition and the value obtained by hashing the partition key in the processing module includes:
[0119] Select the action with the smallest calculated O value from all actions in the action space that satisfies the pre-established constraint conditions, and adjust the current state to the state corresponding to the smallest O value; perform dynamic partition storage according to the correspondence of the values hashed by the hash function for the partition keys corresponding to the partitions in the state corresponding to the smallest O value. Among them, set the action space (Action) of this application as: A = {a 1 , a 2 ,..., a τ}. Each action a in the action space is used to adjust the q i combination corresponding to p i .
[0120] The device for implementing database partitioning provided by this application collects the operation status information of the database partitioning (including partitions and computing nodes), and well evaluates the pros and cons of the partitioning effect. Specifically, according to the operation data of the computing nodes and partitions of the distributed database collected, the agent in the Markov decision process perceives the operation of the distributed database partitioning, and stores data in the partitions in the correspondence of the optimal action values, truly realizing dynamic database partitioning and reducing the operation and maintenance costs in the cloud computing era. Compared with the dynamic partitioning with fixed rules in the related art, this application is more flexible and effective.
[0121] In an exemplary instance, the processing module is also set to: write the new partitioning policy into the Memory Table of the distributed database. In this way, the distributed database management system will perform data writing and query according to the latest partitioning scheme, truly realizing dynamic database partitioning and reducing the operation and maintenance costs in the cloud computing era.
[0122] In an exemplary instance, the device for implementing database partitioning of this application further includes: an iteration module, which is set to: record the process of generating dynamic partitioning this time, that is, the information of the Markov decision process, generate a new partitioning policy and update the previous partitioning policy.
[0123] Through continuous iterative update process, this application makes the partitioning policy better and better, so that the cloud database can make better decisions and scheduling.
[0124] In an exemplary instance, the processing module can also be set to:[[]]
[0125] If the overall variance of the calculated partition load decreases, and all p i corresponding to q iIf the write query performance of the overall database corresponding to the reduced number improves, then a reward is given; if the write and query performance of the overall database decreases after the action, then a penalty is imposed.
[0126] This application evaluates the effectiveness of this action based on the latest operating conditions of the distributed database partitioning, and determines the intensity of the reward and penalty, thus better ensuring the improvement and effectiveness of subsequent strategies. In this way, the strategy of database partitioning is continuously corrected, making the implementation of database partitioning more and more effective.
[0127] In an exemplary instance, the processing module can also be set as:
[0128] Based on the accumulation of the reward value R in each iteration of multiple iterations t to determine the return G representing the dynamic partitioning effect after multiple iterations. Through the return G, the effect of using the database partitioning method of this application within a period of time is realized, providing a relatively intuitive reference for subsequent improvements.
[0129] This application also provides a distributed database management system, including: multiple computing nodes, and the device for implementing database partitioning described in any of the above.
[0130] The method for implementing database partitioning in this application will be described in detail below in conjunction with an embodiment.
[0131] Figure 3 It is a schematic flowchart of an embodiment for implementing database partitioning in this application. As Figure 3 shown, under the distributed database, in this embodiment, the implementation of dynamic partitioning is deployed in the distributed database management system to perform global management of dynamic partitioning of the distributed database. In this embodiment, scheduling is performed by setting a timing task in the distributed database management system.
[0132] First, start dynamic partitioning balancing (such as Figure 3 step 1 in): Schedule the dynamic partitioning timing task; the dynamic partitioning timing task periodically calls the control task (Control Service) in a loop (such as Figure 3 step 2 in) to perform global adjustment of dynamic partitioning. Perform data collection (such as Figure 3 step 3 in) and storage (such as Figure 3 step 4 in) of the running data of the partition Partition and data collection (such as Figure 3 step 5 in) and storage (such as Figure 3 step 6 in) of the running data of the computing node:
[0133] The Partition Scanner collects the running data of partitions, mainly collecting the TPS, QPS, data volume, data distribution of each partition key, etc. of each partition. The running information of the partition can be expressed as: P = {p 1 , p 2 ,..., p n}, where P is the set of all partitions, and p i represents the running data vector of partition i, and the number of partitions is n. The collected data is written by the Partition Scanner into the Memory Table of the distributed database to achieve low-latency writing and querying;
[0134] The main body of the partition is the computing node. There are multiple partitions on one computing node, and the distributed database management system manages the numerous partitions on multiple computing nodes. The running data of the computing node is also very important for the implementation of database partitioning. Information such as CPU occupancy and memory occupancy is resource-isolated in units of computing nodes. The Worker Scanner collects the running data of the computing nodes, mainly collecting the CPU occupancy, IO latency, IOPS, memory occupancy, etc. of each computing node. The running information of the computing node can be expressed as: W = w 1 , w 2 ,..., w m , where W represents the set of all computing nodes, and w i represents the running data vector of computing node i, and the number of computing nodes is m. The collected data is written by the Worker Scanner into the Memory Table of the distributed database to achieve low-latency writing and querying of statistical information.
[0135] Then, dynamic balancing of partitions is performed based on the above-collected running data, thereby generating new dynamic partitions (such as steps 7, 8, and 9 in Figure 3 ):
[0136] The Control Service calls the Analyzer Processor to perform modeling and calculation of dynamic partitioning. As in Figure 3In step 7, the Analyzer Processor reads the Memory Table of the distributed database to obtain the running data of the current partition, the running data of the computing nodes, and the current dynamic partition situation. A dynamic partition model is established: p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on the Markov decision process, that is, the correspondence between the value obtained by hashing the partition key through the hash function and the partition is established. Here, HASH represents the hashing operation, key represents the partition key, and mod represents the modulo operator. The correspondence between the value obtained by hashing the partition key through the hash function and the partition can be maintained through all the dynamic partition information (All dynamic partition info) in the Memory Table of the distributed database.
[0137] Markov decision process: According to p i , q i and p i where the computing node w i is located (such as Figure 3 in step 8), select the corresponding action according to the conditional probability distribution of the action. That is, select the action with the smallest O value calculated according to formula (4) from all actions in the action space that satisfies the pre-established constraint conditions (such as formula (2) and formula (3) above), and adjust the current state to the state corresponding to the smallest O; according to the correspondence between the value obtained by hashing the partition key corresponding to the partition in the state corresponding to the smallest O, perform dynamic partition storage. Here, set the action space (Action) of this application as: A = {a 1 , a 2 ,..., a τ}. Each action a in the action space represents an adjustment of the q i corresponding to p i combination to generate a new routing partition ( Figure 3 in step 9).
[0138] Iteratively update the Markov decision process ( Figure 3 in steps 10 and 11): After the completion of the current dynamic load balancing adjustment strategy, record the model and data of the current Markov decision process, update the routing of all partitions, and prepare for the next iteration. If the Control Service discovers a new partition scheme, it will call the update service (UpdateProcessor), and the Update Processor will write the new partition policy into the Memory Table of the distributed database. In this way, the distributed database management system will perform data writing and query according to the latest partition policy.
[0139] If the overall variance of the calculated partition load decreases, and for all p i the corresponding q i the number of decreases corresponds to an improvement in the write query performance of the overall database, then it is a reward; if the write and query performance of the overall database decreases after the action, then it is a penalty, that is, the reward is negative. That is to say, after making an action, assuming the state space changes from state s t to state s t+1 after that, if the O corresponding to state s t+1 is less than the O corresponding to state s t then this action is considered valid and a reward is given.
[0140] Finally, the return G representing the dynamic partitioning effect after multiple iterations can also be determined according to the accumulation of the reward values R t for each iteration in multiple iterations.
[0141] Although the embodiments disclosed in this application are as above, the content described is only the embodiments adopted for the convenience of understanding this application and is not used to limit this application. Any person skilled in the art within the scope of this application, without departing from the spirit and scope disclosed in this application, can make any modifications and changes in the form and details of the implementation, but the patent protection scope of this application shall still be subject to the scope defined by the appended claims.
Claims
1. A method for implementing database partitioning, including: Obtaining the running information of the database partitioning, where the running information of the database partitioning includes: the running information of the partitions and the running information of the computing nodes; multiple partitions are included on one computing node; According to the obtained running information of the database partitioning, generating dynamic partitions according to the corresponding relationship between the partitions and the values of the partition keys after being hashed by a hash function, including: establishing the corresponding relationship between the values of the partition keys after being hashed by the hash function and the partitions; From the established corresponding relationship, adjusting the values of the partition keys after being hashed by the hash function corresponding to the partitions according to the running information of the partitions in the current database to minimize them; Performing dynamic partition storage according to the adjusted corresponding relationship between the partitions and the values of the partition keys after being hashed by the hash function.
2. The method according to claim 1, further including: Writing the running information of the database partitioning into the memory table of the distributed database.
3. The method according to claim 2, wherein, The adjusting the values of the partition keys after being hashed by the hash function corresponding to the partitions to minimize them and performing dynamic partition storage according to the adjusted corresponding relationship between the partitions and the values of the partition keys after being hashed by the hash function includes: Determining the distribution probability of the actions according to the partitions, the values of the partition keys after being hashed by the hash function, and the running information of the computing nodes where the partitions are located; and storing data in the partitions in the corresponding relationship corresponding to the optimal action value according to the determined distribution probability.
4. The method according to claim 3, wherein, The determining the distribution probability of the actions and storing data in the partitions in the corresponding relationship corresponding to the optimal action value according to the determined distribution probability includes: Selecting the action with the smallest calculated O value that satisfies the pre-established constraint conditions from all the actions in the action space, and adjusting the current state to the state corresponding to the smallest O; Performing dynamic partition storage according to the corresponding relationship between the values of the partition keys after being hashed by the hash function corresponding to the partitions in the state corresponding to the smallest O value; Among them, the action space A is: A = {a 1 , a 2 ,..., a τ}, and each action a in the action space is used to adjust the combination of the values obtained by hashing the partition key corresponding to a certain partition; Among them, represents the overall variance of the partition load; represents p i corresponding to how many q i ; among them, o 1 and o 2 are weights; N represents the number of partitions.
5. The method according to any one of claims 2 to 4, wherein, The establishing the corresponding relationship between the values of the partition keys after being hashed by the hash function and the partitions is: p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partition matching based on a Markov decision process, HASH represents a hashing operation, key represents a partition key, and mod represents a modulo operator.
6. The method according to claim 1, after the dynamic partitioning, further including: Recording the information of this generation process, generating a new partitioning strategy and updating the previous partitioning strategy.
7. The method according to claim 6, further including: Writing the new partitioning strategy into the storage table of the distributed database.
8. The method according to claim 7, before updating the previous partitioning scheme, further including: If the overall variance of the load of the partitions decreases and the write query performance of the overall database corresponding to the number of decreases in the values of the partition keys after being hashed by the hash function corresponding to all partitions improves, then give a reward; If the write and query performance of the overall database degrades after an operation, a penalty is imposed.
9. The method according to claim 8, further comprising: Accumulating the reward values for each iteration in multiple iterations, which is used to represent the return of the dynamic partitioning effect after multiple iterations.
10. A computer-readable storage medium storing computer-executable instructions for executing the method for implementing database partitioning according to any one of claims 1 to 9.
11. An apparatus for implementing database partitioning, comprising a memory and a processor, wherein, The memory stores the following instructions executable by the processor: steps for executing the method for implementing database partitioning according to any one of claims 1 to 9.
12. An apparatus for implementing database partitioning, comprising: An acquisition module and a processing module; wherein, The acquisition module is configured to obtain the operation information of the database partitioning, where the operation information of the database partitioning includes: the operation information of the partitioning and the operation information of the computing nodes; multiple partitions are included on one computing node; The processing module is configured to generate dynamic partitioning according to the obtained operation information of the database partitioning and the correspondence between the partitioning and the value obtained by hashing the partitioning key through a hash function; wherein, the processing module is specifically configured to: Establish the correspondence between the value obtained by hashing the partitioning key through a hash function and the partitioning; from the established correspondence, adjust the value obtained by hashing the partitioning key corresponding to the partitioning according to the operation information of the current database partitioning to minimize it, and perform dynamic partitioning storage according to the adjusted correspondence between the partitioning and the value obtained by hashing the partitioning key.
13. The apparatus according to claim 12, wherein, The establishment of the correspondence between the value obtained by hashing the partitioning key through a hash function and the partitioning in the processing module includes: p = F(q) = F(HASH(key) mod n), where the function F is a function for calculating dynamic partitioning matching based on a Markov decision process, HASH represents a hashing operation, key represents a partitioning key, and mod represents a modulo operator.
14. The apparatus according to claim 12, the apparatus further comprising an iteration module configured to: Record the information of the current generation process, generate a new partitioning strategy and update the previous partitioning strategy.
15. A distributed database management system, comprising: Multiple computing nodes, and the apparatus for implementing database partitioning according to any one of claims 12 to 14.
Citation Information
Patent Citations
Composite Partition Functions
US20160110391A1