ClickHouse Cluster Data Balancing Method, Device and Equipment
By building temporary tables on the source node of the ClickHouse cluster and copying data from the target table to the target node, the problem of time-consuming and large space-consuming data balance in the expansion of ClickHouse cluster is solved, and efficient data balance and query efficiency are improved.
Patent Information
- Application Number
- CN202210906036.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The existing ClickHouse cluster expansion data balance method takes a long time, takes up a lot of extra space, and the service downtime is proportional to the amount of data.
The temporary table with the same structure as the target table is built on the source node of the ClickHouse cluster, which is used to receive newly inserted data during the equalization process, so that the data in the target table does not change during the equalization process. Then, the data that needs to be equalized is copied and sent to the target node. After receiving successful feedback from the target node, the data in the target table is deleted and merged with the temporary table.
It realizes a data balance solution that does not require stopping nodes, takes up less additional storage space, and adjusts weights without human intervention. Each node consumes less time to achieve a balanced state, improving cluster query efficiency.
Smart Images

Figure CN115455122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and in particular to a method, device and equipment for data balancing in a ClickHouse cluster. Background Art
[0002] ClickHouse is a columnar database management system (DBMS) for online analytical processing (OLAP). A ClickHouse cluster refers to a ClickHouse database cluster. Please refer to Figure 1 , when new nodes need to be added for expansion and data balancing is required, the existing data balancing solutions generally rewrite after emptying. Specifically, the cluster data is backed up and then emptied, and after adding nodes, the data is re-imported. The backup data occupies a large amount of space, the data import takes a long time, and the downtime is proportional to the data volume. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, device and equipment for data balancing in a ClickHouse cluster to solve problems such as long time consumption and large additional space required by the existing data balancing method for ClickHouse cluster expansion.
[0004] According to a first aspect, an embodiment of the present invention provides a method for data balancing in ClickHouse cluster expansion, the method comprising:
[0005] On a source node of the ClickHouse cluster, a first table with the same structure as the target table is constructed, and the first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process;
[0006] Copy out first data that needs to be balanced to a target node from the target table and send it to the target node;
[0007] After receiving a successful feedback from the target node, delete the first data from the target table and merge it with the first table.
[0008] Optionally, before copying out the first data that needs to be balanced to the target node from the target table, it further comprises:
[0009] Determine a balancing index according to the number of source nodes and the number of target nodes;
[0010] Determine the size of the first data according to the balancing index.
[0011] Optionally, when there are multiple target nodes, the balancing index is the ratio of the first data that needs to be balanced to one of the target nodes to the current data size of the target table;
[0012] After receiving the successful feedback from the target node, deleting the first data from the target table and merging it with the first table includes:
[0013] After receiving the successful feedback from all the target nodes, deleting the first data from the target table and merging it with the first table.
[0014] Optionally, before constructing the first table with the same structure as the target table on the source node of the ClickHouse cluster, it further includes:
[0015] For each source node, determining the disk space size required for the balancing process;
[0016] Judging whether the available disk space is greater than or equal to the disk space size.
[0017] Optionally, there are multiple target nodes, and the disk space size is equal to the sum of the first space size and the second space size. The first space size is the size of the first data that needs to be balanced to one of the target nodes, and the second space size is the size of the second table itself for storing the first data that needs to be balanced to one of the target nodes.
[0018] Optionally, the size of the second table itself is equal to the number of rows of the first data that needs to be balanced to one of the target nodes multiplied by the number of bytes required for each row.
[0019] Optionally, before copying the first data that needs to be balanced to the target node from the target table, it further includes:
[0020] Adding a serial number column to the target table and filling in the row numbers.
[0021] According to a second aspect, an embodiment of the present invention provides a device for data balancing in ClickHouse cluster expansion, including:
[0022] A temporary table construction module, configured to construct a first table with the same structure as the target table on the source node of the ClickHouse cluster. The first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process;
[0023] A data transfer module, configured to copy the first data that needs to be balanced to the target node from the target table and send it to the target node;
[0024] A table update module, configured to delete the first data from the target table and merge it with the first table after receiving the successful feedback from the target node.
[0025] According to a third aspect, an embodiment of the present invention provides an electronic device, including:
[0026] A memory and a processor, which are communicatively connected to each other. The memory is used to store a computer program. When the computer program is executed by the processor, the method for data balancing during ClickHouse cluster expansion described in any one of the above first aspects is implemented.
[0027] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium is used to store a computer program. When the computer program is executed by a processor, the method for data balancing during ClickHouse cluster expansion described in any one of the above first aspects is implemented.
[0028] In the embodiments of the present invention, when the ClickHouse cluster is expanded, that is, a new node is added, and data balancing is required, only the source node and the target node need to be selected to start the data balancing task, and finally the data distribution on the nodes tends to be balanced. Moreover, the data balancing solution provided by the embodiments of the present invention does not require stopping the nodes, occupies a small amount of additional storage space, does not require manual intervention to adjust weights, and each node takes less time to reach the balanced state, improving the cluster query efficiency. In addition, the data balancing solution provided by the embodiments of the present invention is a table-level balance, and other tables are not affected during the balancing process, which can give full play to the advantages of cluster resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present invention. In the drawings:
[0030] Figure 1 A schematic diagram for data balancing during ClickHouse cluster expansion;
[0031] Figure 2 A flowchart of a method for data balancing during ClickHouse cluster expansion provided by an embodiment of the present invention;
[0032] Figure 3 A schematic diagram of the data volume before and after data balancing during ClickHouse cluster expansion provided by an embodiment of the present invention;
[0033] Figure 4 A schematic diagram of data transfer during data balancing during ClickHouse cluster expansion provided by an embodiment of the present invention;
[0034] Figure 5 A schematic diagram of the structure of a device for data balancing during ClickHouse cluster expansion provided by an embodiment of the present invention;
[0035] Figure 6 A schematic structural diagram of control and execution in the process of data balancing during the expansion of a ClickHouse cluster provided by an embodiment of the present invention;
[0036] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] It should be noted that the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the element. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. In the description of the following embodiments, "a plurality of" means two or more unless otherwise specifically defined.
[0039] Please refer to Figure 2 , an embodiment of the present invention provides a method for data balancing during the expansion of a ClickHouse cluster, and the method includes:
[0040] S201: On the source node of the ClickHouse cluster, construct a first table with the same structure as the target table, where the first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process;
[0041] S202: Copy the first data that needs to be balanced to the target node from the target table and send it to the target node;
[0042] S203: After receiving the successful feedback from the target node, delete the first data from the target table and merge it with the first table.
[0043] Specifically, the source node is the original node of the ClickHouse cluster, and the target node is the newly added node after expansion. To facilitate the first table to receive newly inserted data during the data balancing process (i.e., data newly stored in the source node), the table name of the first table can be changed to the table name of the target table, and the table name of the target table can be replaced with a new table name. The first table here can avoid cluster query errors. The first data that needs to be balanced to the target node is the data that needs to be transferred to the target node.
[0044] Of course, if the successful feedback from the target node is not received or the failure feedback from the target node is received, the data balancing process of the source node needs to be re-executed.
[0045] In the embodiments of the present invention, when the ClickHouse cluster is expanded, that is, new nodes are added, and data balancing is required, only the source node and the target node need to be selected to start the data balancing task, and finally the data distribution on the nodes tends to be balanced. Moreover, the data balancing method provided by the embodiments of the present invention does not require stopping the nodes, occupies less additional storage space, does not require manual intervention to adjust weights, and each node takes less time to reach the balanced state, improving the cluster query efficiency. In addition, the data balancing method provided by the embodiments of the present invention is a table-level balancing, and does not affect other tables during the balancing process, and can give full play to the advantages of cluster resources.
[0046] In some specific embodiments, before copying the first data that needs to be balanced to the target node from the target table, it further includes:
[0047] Determine the balancing index according to the number of source nodes and the number of target nodes;
[0048] Determine the size of the first data according to the balancing index.
[0049] Here, it can be considered that the current data sizes of the target tables on each source node are equal, that is, before expansion, the data on each original node is in a balanced state.
[0050] Specifically, the balancing index bq can be determined according to the following formula: bq = 1 / (m + n), where m and n are the number of source nodes and the number of target nodes respectively.
[0051] For example, please refer to Figure 3 and Figure 4 , the number of source nodes is 4, and each source node has 90 pieces of data in a certain table of a certain database. After expanding 2 nodes (target nodes), the balancing index = 15 / 90 = 1 / (4 + 2).
[0052] In other optional specific embodiments, the balancing metrics can also be customized. Since the embodiments of the present invention are table-level balancing methods, the data balancing strategy at the database table level can be adjusted to meet different requirements.
[0053] In some specific embodiments, there are multiple target nodes, and the balancing metric is the ratio of the size of the first data to be balanced to one of the target nodes to the current data size of the target table.
[0054] After receiving the successful feedback from the target node, deleting the first data from the target table and merging it with the first table includes:
[0055] After receiving the successful feedback from all the target nodes, deleting the first data from the target table and merging it with the first table.
[0056] Specifically, copying the first data to be balanced to the target node from the target table and sending it to the target node includes:
[0057] Copying the first data to be balanced to the first target node from the target table and sending it to the first target node;
[0058] Copying the first data to be balanced to the second target node from the target table and sending it to the second target node;
[0059] Repeat this process until all the first data to be balanced has been sent.
[0060] In other optional specific embodiments, after receiving the successful feedback from each target node, the first data sent to that target node can be deleted from the target table. After all the first data has been deleted, the target table is then merged with the first table.
[0061] In some specific embodiments, before constructing the first table with the same structure as the target table on the source node of the ClickHouse cluster, it further includes:
[0062] For each source node, determining the size of the disk space required for the balancing process;
[0063] Judging whether the available disk space is greater than or equal to the disk space size.
[0064] Of course, only when the available disk space is greater than or equal to the required disk space size will the data balancing process be carried out.
[0065] In the embodiments of the present invention, the size of the disk space required for the balancing process of each source node can be equal.
[0066] In some specific embodiments, there are multiple target nodes, the disk space size is equal to the sum of the first space size and the second space size, the first space size is the size of the first data that needs to be balanced to one of the target nodes, and the second space size is the size of the second table itself for storing the first data that needs to be balanced to one of the target nodes.
[0067] In the embodiments of the present invention, a second table is also created in the source node. This second table is used to store the first data in the source node that needs to be balanced to one of the target nodes, that is, the first data that needs to be balanced to one of the target nodes is copied from the target table to the second table. After receiving the successful feedback from the target node, the data in the second table is updated, that is, the first data that needs to be balanced to another target node is copied to the second table. Finally, after the data balancing in the source node is successful, for example, after receiving the successful feedback from all target nodes, the second table is deleted.
[0068] Specifically, the second table is created according to the target table and its structure can be the same as that of the target table. After copying the first data that needs to be balanced to one of the target nodes to the second table, the data file of the second table can be exported, and then the data file is compressed and sent to the target node. After receiving the compressed data file, the target node can decompress the data file and import it into a table that is pre-created and has the same structure as the target table.
[0069] In some specific embodiments, the size of the second table itself is equal to the number of rows of the first data that needs to be balanced to one of the target nodes multiplied by the number of bytes required per row, for example, 4 bytes.
[0070] In some specific embodiments, before copying the first data that needs to be balanced to the target node from the target table, it further includes:
[0071] Add a serial number column to the target table and fill in the row numbers.
[0072] In the embodiments of the present invention, in order to facilitate the identification and selection of the first data during the processes such as copying and deleting the first data, row numbers are added to the target table as the identifier for each row of data. ClickHouse is a columnar database, and writing columns is very fast and will not cause an increase in the balancing duration.
[0073] All source nodes of the ClickHouse cluster perform data balancing according to the above process.
[0074] Please refer to Figure 5, the method for data balancing during ClickHouse cluster expansion provided by the embodiments of the present invention can be implemented by a control module in the cluster controller and an agent program module installed on each node.
[0075] Specifically, the control module is used to obtain cluster-related information, including the database that needs to be balanced, the tables in the database that need to be balanced (i.e., target tables), the data size in the tables, the number of original nodes in the cluster (i.e., source nodes) and the number of newly added nodes during expansion (i.e., target nodes), available disk space, balance metrics, etc. The control module is also used to schedule balance tasks according to the balance metrics and send them to the agent program modules on each node for execution without human intervention. After receiving the balance task from the control module, the agent program module on the source node constructs a first table with the same structure as the target table on the source node of the ClickHouse cluster, copies the first data that needs to be balanced to the target node from the target table, and sends it to the target node. After receiving the successful feedback from the target node, it deletes the first data from the target table and merges it with the first table. The agent program module on the source node is also used to send feedback information to the control module after sending the first data to the target node, so that the control module can understand the current task progress. After successfully importing the first data sent by the source node, the agent program module on the target node sends feedback to the agent program module and the control module on the source node.
[0076] Correspondingly, please refer to Figure 6 , the embodiments of the present invention provide a device for data balancing during ClickHouse cluster expansion, including:
[0077] A temporary table construction module 601, configured to construct a first table with the same structure as the target table on the source node of the ClickHouse cluster, where the first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process;
[0078] A data transfer module 602, configured to copy the first data that needs to be balanced to the target node from the target table and send it to the target node;
[0079] A table update module 603, configured to delete the first data from the target table and merge it with the first table after receiving the successful feedback from the target node.
[0080] In the embodiments of the present invention, when the ClickHouse cluster is expanded, that is, new nodes are added and data balancing is required, only the source node and the target node need to be selected to start the data balancing task, and finally the data distribution on the nodes tends to be balanced. Moreover, the data balancing device provided by the embodiments of the present invention does not need to stop the nodes, occupies less additional storage space, does not require manual intervention to adjust weights, and each node takes less time to reach the balanced state, improving the query efficiency of the cluster. In addition, the data balancing device provided by the embodiments of the present invention is a table-level balance, and other tables are not affected during the balancing process, which can give full play to the advantages of cluster resources.
[0081] In some specific embodiments, the device further includes:
[0082] An equilibrium index determination module, configured to determine an equilibrium index according to the number of the source nodes and the number of the target nodes;
[0083] A first data size determination module, configured to determine the size of the first data according to the equilibrium index.
[0084] In some specific embodiments, there are multiple target nodes, and the equilibrium index is the ratio of the first data to be balanced to one of the target nodes to the current data size of the target table;
[0085] The table update module 603 is configured to delete the first data from the target table and merge it with the first table after receiving the successful feedback from all the target nodes.
[0086] In some specific embodiments, the device further includes:
[0087] A space determination module, configured to determine the disk space size required for the balancing process for each of the source nodes;
[0088] A judgment module, configured to judge whether the available disk space is greater than or equal to the disk space size.
[0089] In some specific embodiments, there are multiple target nodes, and the disk space size is equal to the sum of a first space size and a second space size. The first space size is the size of the first data to be balanced to one of the target nodes, and the second space size is the size of the second table itself for storing the first data to be balanced to one of the target nodes.
[0090] In some specific embodiments, the size of the second table itself is equal to the number of rows of the first data to be balanced to one of the target nodes multiplied by the number of bytes required for each row.
[0091] In some specific embodiments, the device further includes:
[0092] A supplementary module for adding a serial number column to the target table and filling in the row numbers.
[0093] The embodiments of the present invention are device embodiments based on the same inventive concept as the above method embodiments. Therefore, for specific technical details and corresponding technical effects, please refer to the above method embodiments and will not be elaborated here.
[0094] The embodiments of the present invention also provide an electronic device, as Figure 7 shown. The electronic device may include a processor 71 and a memory 72, where the processor 71 and the memory 72 may be communicatively connected to each other through a bus or other means. Figure 7 Taking the connection through the bus as an example.
[0095] The processor 71 may be a central processing unit (CPU). The processor 71 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0096] The memory 72, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for data balancing in ClickHouse cluster expansion in the embodiments of the present invention (for example, Figure 6 the temporary table construction module 601, data transfer module 602, and table update module 603 shown). The processor 71 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 72, that is, implements the method for data balancing in ClickHouse cluster expansion in the above method embodiments.
[0097] The memory 72 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor 71 and the like. In addition, the memory 72 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 72 may optionally include a memory remotely disposed relative to the processor 71, and these remote memories may be connected to the processor 71 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0098] The one or more modules are stored in the memory 72 and, when executed by the processor 71, perform the method for data balancing in ClickHouse cluster expansion as described in Figures 2 - 5 the embodiments shown.
[0099] Specific details of the above electronic device may be correspondingly referred to Figures 2 to 5 the corresponding related descriptions and effects in the embodiments shown for understanding, and will not be elaborated here.
[0100] Correspondingly, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, it implements each process of the method embodiment for data balancing in ClickHouse cluster expansion as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0101] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0102] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, electronic device embodiments, and computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.
[0103] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for data balancing during ClickHouse cluster expansion, characterized in that, the method includes: On the source node of the ClickHouse cluster, construct a first table with the same structure as the target table, and the first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process; Copy the first data that needs to be balanced to the target node from the target table, and send it to the target node; After receiving the successful feedback from the target node, delete the first data from the target table and merge it with the first table; Before constructing the first table with the same structure as the target table on the source node of the ClickHouse cluster, it further includes: For each source node, determine the disk space size required for the balancing process; Judge whether the available disk space is greater than or equal to the disk space size; There are multiple target nodes, and the disk space size is equal to the sum of the first space size and the second space size. The first space size is the size of the first data that needs to be balanced to one of the target nodes, and the second space size is the size of the second table itself for storing the first data that needs to be balanced to one of the target nodes; The size of the second table itself is equal to the number of rows of the first data that needs to be balanced to one of the target nodes multiplied by the number of bytes required for each row.
2. The method according to claim 1, characterized in that, Before copying the first data that needs to be balanced to the target node from the target table, it further includes: According to the number of source nodes and the number of target nodes, determine the balancing index; According to the balancing index, determine the size of the first data.
3. The method according to claim 2, characterized in that, There are multiple target nodes, and the balancing index is the ratio of the first data that needs to be balanced to one of the target nodes to the current data size of the target table; After receiving the successful feedback from the target node, deleting the first data from the target table and merging it with the first table includes: After receiving the successful feedback from all the target nodes, delete the first data from the target table and merge it with the first table.
4. The method according to claim 1, characterized in that, Before copying the first data that needs to be balanced to the target node from the target table, it further includes: Add a serial number column to the target table and fill in the row numbers.
5. A device for data balancing during ClickHouse cluster expansion, characterized in that, includes: A space determination module for determining the disk space size required for the balancing process for each source node; A judgment module for judging whether the available disk space is greater than or equal to the disk space size; A temporary table construction module for constructing a first table with the same structure as the target table on the source node of the ClickHouse cluster, and the first table is used to receive newly inserted data during the balancing process, so that the data in the target table does not change during the balancing process; A data transfer module, configured to copy first data that needs to be balanced to a target node from the target table and send it to the target node; A table update module, configured to delete the first data from the target table and merge it with the first table after receiving a successful feedback from the target node; Wherein, there are multiple target nodes, the disk space size is equal to the sum of a first space size and a second space size, the first space size is the size of the first data that needs to be balanced to one of the target nodes, and the second space size is the size of a second table itself for storing the first data that needs to be balanced to one of the target nodes; The size of the second table itself is equal to the number of rows of the first data that needs to be balanced to one of the target nodes multiplied by the number of bytes required for each row.
6. An electronic device, Characterized in that, It includes: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory is used to store a computer program, and when the computer program is executed by the processor, the method for data balancing during ClickHouse cluster expansion according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium, Characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method for data balancing during ClickHouse cluster expansion according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Base station transmission circuit batch scheduling method and system
CN103906101A
Service data processing method, device and equipment
CN112860694A