Partition expansion and contraction method and system

By controlling the sending end to send data to the reserved partitions during the Kafka system, the sending end controls the sending end to send data to the reserved partitions and allows the consumer to receive data from all partitions until the data meets specific conditions, the major network overhead problems caused by the inability to reduce capacity and data migration during the trough period, and the reduction of partition capacity and network consumption is achieved.

CN114936095BActive Publication Date: 2025-06-24CHONGQING UNISINSIGHT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210617236.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-06-24
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The existing Kafka system cannot perform partition reduction during the trough period, and data migration is required when node reduction is reduced, resulting in large network overhead.

Method used

During capacity reduction, the control sender only sends data to the partitions retained after capacity reduction, while the consumer still receives data from all partitions until the data under the partitions that need to be deleted meets certain conditions (such as all consumption or exceeding the retention period), then deletes and completes the capacity reduction.

Benefits of technology

Supports partition shrinkage, reducing partition data migration and network transmission consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936095B_ABST
    Figure CN114936095B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for partition scaling. When triggering node scaling and partition scaling, the data sent by each sender can be sent to each node under the partitions retained after partition scaling. Each consumer receives data from each node under all partitions before partition scaling. In this way, when the partitions to be deleted corresponding to the partition scaling operation meet the deletion conditions, the nodes and partitions to be deleted are deleted to complete the scaling. In this solution, during scaling, the sender is controlled to send data only to the partitions retained after scaling, while the consumer still receives data from all partitions. In this way, the data under the partitions to be deleted is continuously consumed without new addition. When the data under the partitions to be deleted meets certain conditions, the scaling can be completed. This method supports partition scaling and can minimize the migration of partition data and reduce network transmission consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data storage, and more particularly, to a method and system for partitioning scale-out and scale-in. Background Art

[0002] The open-source middleware Kafka, as a message middleware for many companies, can decouple each system. The sending end sends message data, and the consuming end consumes the message data. There is a certain time difference from the sending to the consumption of a message, which has the effect of peak shaving. Businesses have peak and trough periods, and the throughput of messages may vary by several times or even hundreds of times. For example, during the day, the traffic is large and the business operations are intensive, while at night, the traffic is small and the business operations are less. The message throughput of Kafka can be horizontally scaled by increasing the number of nodes and partitions, improving the performance of the system.

[0003] If it is necessary to scale out the nodes and partitions of Kafka during the peak period, the number of partitions can be increased, and the added partitions can be distributed on the expanded nodes.

[0004] However, if it is necessary to scale in the nodes and partitions of Kafka during the trough period, according to the partition design of Kafka, the number of partitions can only be increased and cannot be decreased, because the sending end will continuously send new data to all partitions. At the same time, there is message data in all partitions, and the consuming end needs to consume from all partitions. Moreover, when scaling in nodes, the data on the nodes to be scaled in needs to be copied and migrated to other nodes, and only after the migration is completed can the nodes be scaled in, which incurs a large network overhead. Summary of the Invention

[0005] The objectives of the present invention include, for example, providing a method and system for partitioning scale-out and scale-in, which can support partition scale-in and reduce the migration of partition data points and network transmission consumption.

[0006] The embodiments of the present invention can be implemented as follows:

[0007] In a first aspect, the present invention provides a method for partitioning scale-out and scale-in, which is applied to a partitioning scale-out and scale-in system. The system includes a Kafka cluster, multiple consuming ends, and a sending end. The Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions. The method includes:

[0008] When triggering node scale-in and partition scale-in, send the data sent by each sending end to each node under the partitions remaining after partition scale-in;

[0009] Each consuming end receives data from each node under all partitions before partition scale-in;

[0010] When the partitions to be deleted corresponding to the partition shrinking operation meet the deletion conditions, delete the nodes and partitions to be deleted to complete the shrinking.

[0011] In an alternative embodiment, the partition scaling system further includes a zookeeper cluster. Partition information is recorded in the zookeeper cluster. The partition information for data sent by each sending end and the partition information for data received by each consuming end are recorded in different directories. After the steps of triggering node shrinking and partition shrinking, the method further includes:

[0012] Update the partition information in the directory corresponding to each sending end to the partition information after shrinking, and keep the partition information in the directory corresponding to each consuming end as the original partition information;

[0013] After the step of deleting the nodes and partitions to be deleted to complete the shrinking, the method further includes:

[0014] Update the partition information in the directory corresponding to each consuming end to the partition information after shrinking.

[0015] In an alternative embodiment, the step of, when the partitions to be deleted corresponding to the partition shrinking operation meet the deletion conditions, deleting the nodes and partitions to be deleted includes:

[0016] Judge whether the data in the partitions to be deleted corresponding to the partition shrinking operation needs to be traced back. If it needs to be traced back, after processing the data in the partitions to be deleted according to the requirements of the shrinking time limit, delete the nodes and partitions to be deleted;

[0017] If it does not need to be traced back, when the data in the partitions to be deleted meets the invalidation conditions, delete the nodes and partitions to be deleted.

[0018] In an alternative embodiment, the step of, after processing the data in the partitions to be deleted according to the requirements of the shrinking time limit and then deleting the nodes and partitions to be deleted, includes:

[0019] If the requirements of the shrinking time limit indicate that the shrinking operation needs to be performed immediately, migrate all the data in the partitions to be deleted to the remaining nodes, delete the nodes to be deleted, and wait until the set retention period is reached, then delete the partitions to be deleted;

[0020] If the requirements of the shrinking time limit indicate that the shrinking operation does not need to be performed immediately, wait until the set retention period is reached, then delete the partitions to be deleted and delete the nodes to be deleted.

[0021] In an alternative embodiment, the step of deleting the node and the partition to be deleted when the data in the partition to be deleted meets the invalidation condition without the need for backtracking includes:

[0022] If backtracking is not required, determine whether all the data in the partition to be deleted is invalid. If all the data is invalid, then delete the node and the partition to be deleted;

[0023] If not all the data in the partition to be deleted is invalid, process the non-invalid data according to the shrinkage time limit requirement, and then delete the node and the partition to be deleted.

[0024] In an alternative embodiment, the step of processing the non-invalid data according to the shrinkage time limit requirement and then deleting the node and the partition to be deleted includes:

[0025] When the shrinkage time limit requirement indicates that the shrinkage operation needs to be performed immediately, migrate the non-invalid data in the partition to the remaining nodes, delete the node to be deleted, and wait until all the data on the partition is consumed or reaches the set retention period, and then delete the partition to be deleted;

[0026] When the shrinkage time limit requirement indicates that the shrinkage operation does not need to be performed immediately, wait until the storage time of the last piece of data in the partition reaches the set retention period, determine that all the data in the partition is invalid, and then delete the partition and the node to be deleted.

[0027] In an alternative embodiment, the step of determining whether all the data in the partition to be deleted is invalid includes:

[0028] For each partition to be deleted, obtain the record offset of the last piece of data in the partition;

[0029] Obtain the consumption offset of each consumer for consuming the data in the partition, where the consumption offset represents the maximum record offset of the data in the partition that the consumer has consumed;

[0030] When each consumption offset is equal to the record offset, determine that all the data in the partition is invalid.

[0031] In an alternative embodiment, the method further includes:

[0032] When triggering node expansion and partition expansion, distribute the partitions to be expanded under the nodes to be expanded;

[0033] Send the data sent by each sender to each node under all the partitions after partition expansion;

[0034] Each of the consumer ends receives data from each node under all partitions after partition expansion of the partition.

[0035] In a second aspect, the present invention provides a partition scaling system, the system includes a Kafka cluster, multiple consumer ends and a sender end, the Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions;

[0036] When triggering node scaling down and partition scaling down, each of the sender ends is used to send data to each node under the partitions retained after partition scaling down;

[0037] Each of the consumer ends is used to receive data from each node under all partitions before partition scaling down;

[0038] When the partitions to be deleted corresponding to the partition scaling down operation meet the deletion conditions, the management node among the multiple nodes is used to delete the nodes and partitions to be deleted, and complete the scaling down.

[0039] In an optional implementation manner, the partition scaling system further includes a zookeeper cluster, the partition information is recorded in the zookeeper cluster, and the partition information for the sender ends to send data and the partition information for the consumer ends to receive data are recorded in different directories;

[0040] The zookeeper cluster is used to update the partition information in the directory corresponding to each of the sender ends to the partition information after scaling down when triggering node scaling down and partition scaling down, and keep the partition information in the directory corresponding to each of the consumer ends as the original partition information;

[0041] The zookeeper cluster is further used to update the partition information in the directory corresponding to each of the consumer ends to the partition information after scaling down after completing the scaling down.

[0042] The beneficial effects of the embodiments of the present invention include, for example:

[0043] This application provides a method and system for partition scaling. When triggering node scaling and partition scaling, the data sent by each sender can be sent to each node under the partitions retained after partition scaling. Each consumer receives data from each node under all partitions before partition scaling. In this way, when the partitions that need to be deleted in the partition scaling operation meet the deletion conditions, the nodes and partitions that need to be deleted are then deleted to complete the scaling. In this solution, during scaling, the sender is controlled to only send data to the partitions retained after scaling, while the consumer still receives data from all partitions. In this way, the data under the partitions that need to be deleted is continuously consumed without new addition. When the data under the partitions that need to be deleted meets certain conditions, the scaling can be completed. This method supports partition scaling and can minimize the migration of partition data and reduce network transmission consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 Structural block diagram of the partition scaling system provided by the embodiment of this application;

[0046] Figure 2 Flowchart of the partition scaling method provided by the embodiment of this application;

[0047] Figure 3 For Figure 2 Flowchart of the sub-steps included in step S103 in

[0048] Figure 4 For Figure 3 Flowchart of the sub-steps included in step S1033 in

[0049] Figure 5 Flowchart of the expansion method in the partition scaling method provided by the embodiment of this application;

[0050] Figure 6 Schematic diagram of the relationship between nodes, Topics, and partitions provided by the embodiment of this application;

[0051] Figure 7 Schematic diagram of data sending and receiving between the sender and the consumer and each partition before scaling provided by the embodiment of this application;

[0052] Figure 8Schematic diagram of data transmission and reception between the sender and consumer and each partition after capacity expansion provided by the embodiments of the present application;

[0053] Figure 9 Schematic diagram of partition information update of the sender and consumer during capacity reduction provided by the embodiments of the present application;

[0054] Figure 10 Schematic diagram of data transmission and reception between the sender and consumer and each partition after capacity reduction provided by the embodiments of the present application;

[0055] Figure 11 One of the schematic diagrams of the consumption offsets of each consumer provided by the embodiments of the present application;

[0056] Figure 12 Another schematic diagram of the consumption offsets of each consumer provided by the embodiments of the present application. Detailed implementation manners

[0057] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0058] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0059] It should be noted that: like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0060] It should be noted that, without conflict, the features in the embodiments of the present invention may be combined with each other.

[0061] In this application, we first introduce the principle of Kafka's partitioned data storage, mainly from the partition dimension. Each program running in Kafka is called a Broker (node). Generally, different Brokers run on different physical machines. Multiple Brokers form a Kafka cluster, and a Kafka cluster has multiple Topics. A Topic has multiple Partitions. The sender and consumer of Kafka interact with one or more partitions. The sender sends data to the partition, and the consumer gets data from the partition. Partitions are the basic units for storing messages in Kafka.

[0062] Kafka uses offset to record the number of messages in a partition. Each time a message is written to a partition, an offset number is generated, which is an auto-increment rule. Messages in a partition are written sequentially and read sequentially by the consumer.

[0063] Partitions are the basic unit for Kafka to manage data. A partition can have multiple copies, divided into a primary partition and multiple secondary partitions. The secondary partitions will continuously synchronize data from the primary partition to prevent single node failures from causing partition unavailability. Kafka's sender and consumer end send and consume data in the primary partition. At the same time, the partition will generate many other metadata files, such as partition index files, primary and secondary partition synchronization epoh files, etc. Kafka needs to manage these files.

[0064] Therefore, the number of partitions in a Kafka cluster is not the more the better. More partitions will increase the management overhead of the Kafka cluster. The appropriate number of partitions should be set according to the amount of data generated by the Topic.

[0065] Kafka is a message middleware, and data cannot be stored permanently. Messages in partitions have deletion policies and message retention time settings. For example, if the message retention period is set to 2 hours, data older than 2 hours will be deleted.

[0066] Kafka may have multiple consumers consuming data in the same partition. Each consumer will submit a consumer offset to record which data the consumer has consumed in the partition. Generally, the offset of the last message in the partition is greater than or equal to the consumer offset. At the same time, Kafka supports message backtracking. Consumers can change their consumer offsets and re-consume data that has already been consumed, as long as it is within the message retention period and has not been deleted.

[0067] The number of partitions of a topic T in Kafka is recorded under the directory / brokers / topics / T / partitions in a ZooKeeper cluster. Both the sender and the consumer obtain the number of partitions of topic T from this directory. The sender and the consumer will update the partition information regularly. The sender will send data to all partitions.

[0068] Under the existing mechanism of Kafka, if it is necessary to expand the number of partitions of Kafka's brokers and topics during peak hours, the number of partitions of the topic can be increased, and the added partitions can be distributed on the expanded brokers.

[0069] However, if it is necessary to downsize the Kafka brokers and topic partitions during off-peak hours, according to the partition design of Kafka, the number of partitions can only be increased, not decreased, because the sender will continuously send new data to all partitions. At the same time, there is message data in all partitions. The consumer needs to consume from all partitions.

[0070] At the same time, when the broker is downsized, the broker to be downsized needs to copy and migrate the partition data of the node to the remaining brokers, and only after the migration is completed can the broker be downsized, which incurs a huge network overhead in the middle.

[0071] It can be seen that there is a defect in the existing Kafka mechanism that the number of topic partitions in the cluster cannot be reduced according to the business volume. Moreover, when a node in the Kafka cluster is downsized, the partition data of the downsized node must be migrated, resulting in a large amount of network overhead.

[0072] Based on the above research findings, this application provides a method for expanding and downsizing partitions. When downsizing, the sender is controlled to only send data to the partitions remaining after downsizing, while the consumer still receives data from all partitions. In this way, the data in the partitions to be deleted is continuously consumed without new additions. When the data in the partitions to be deleted meets certain conditions, downsizing can be completed. This method supports the downsizing of partitions and can minimize the migration of partition data and reduce network transmission consumption.

[0073] Please refer to Figure 1 , this application embodiment provides a partition expansion and downsizing system, which includes a Kafka cluster, multiple consumers and senders. Among them, the consumers and senders communicate with the Kafka cluster respectively. Each sender can send data to the Kafka cluster, and each consumer can obtain data from the Kafka cluster. Among them, each consumer and sender can be a client.

[0074] Among them, the Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions. For example, two partitions can be distributed on one node.

[0075] Please refer to Figure 2 , which is the flowchart of the partition scaling method provided by the embodiment of the present application. The method steps defined by the processes related to the partition scaling method can be implemented by the above partition scaling system. The following will Figure 2 elaborate in detail on the specific processes shown in

[0076] S101, when triggering node scaling down and partition scaling down, send the data sent by each of the sending ends to each node under the partitions retained after partition scaling down.

[0077] S102, each of the consuming ends receives data from each node under all partitions before partition scaling down.

[0078] S103, when the partitions to be deleted corresponding to the partition scaling down operation meet the deletion conditions, delete the nodes and partitions to be deleted to complete the scaling down.

[0079] In this embodiment, before performing node scaling down and partition scaling down, each sending end can send data to the partitions of all nodes, and each consuming end can also obtain data from the partitions of all nodes. Assume that there are currently three nodes, which are divided into a total of 6 partitions. Then each sending end can send data to each of the 6 partitions respectively, and each consuming end can also obtain data from each of the 6 partitions.

[0080] When the traffic volume is small, such as when the input traffic in the cluster decreases or the input traffic of the Topic decreases, for example, when it is less than the threshold of the monitoring metrics corresponding to Kafka, or when it passes the peak traffic period of the service, node scaling down and partition scaling down can be triggered.

[0081] In this embodiment, the partition scaling system further includes a zookeeper cluster, and the zookeeper cluster can communicate with each sending end, consuming end, and Kafka cluster respectively. The partition information is recorded in the zookeeper cluster. In the prior art, the partition information corresponding to each sending end and consuming end is recorded in the same directory, such as the / brokers / topics / T / partitions directory, and both the sending end and the consuming end obtain the partition information from this directory.

[0082] In this embodiment, the schematic diagrams of data transmission and reception between the sending end and the consuming end and each partition record the partition information for data transmission by each sending end and the partition information for data reception by each consuming end in different directories. For example, they are respectively in the directories / brokers / topics / T / producer / partitions and / brokers / topics / T / consumer / partitions. In the case where there is no scale-out or scale-in, the information in the two directories is the same.

[0083] When node scale-in and partition scale-in are triggered, assume that the current number of nodes is M and the number of partitions is W. It is necessary to scale the nodes down to N and the number of partitions down to Q. At this time, the nodes and partitions cannot be scaled down immediately. First, the ZooKeeper cluster can update the partition information in the directories corresponding to each sending end to the partition information after scale-in, that is, partitions 1 to Q, while the partition information in the directories corresponding to each consuming end remains the original partition information, that is, partitions 1 to W.

[0084] Each sending end can obtain the updated partition information from the ZooKeeper cluster, so that each sending end will only send new data to partitions 1 to Q, that is, nodes 1 to N, and no new data will be written to the nodes N + 1 to M (partitions Q + 1 to W) that need to be deleted.

[0085] The partition information obtained by each consuming end is still the original partition information. Therefore, each consuming end will continue to receive data from partitions 1 to W (nodes 1 to M) because there are remaining unconsumed messages in the nodes N + 1 to M (partitions Q + 1 to W).

[0086] Each sending end and each consuming end perform data transmission and data acquisition respectively in the above manner. When the partitions that need to be deleted corresponding to the scale-in operation meet the deletion conditions, the nodes and partitions that need to be deleted can be deleted to complete the scale-in. Among them, the so-called meeting the deletion conditions can be, for example, that the data in the partitions to be deleted has been completely consumed, or the data in the partitions to be deleted has exceeded the retention period for deletion, or the remaining unconsumed data has been migrated, etc.

[0087] After the scale-in is completed, in the ZooKeeper cluster, the partition information in the directories corresponding to each consuming end can be updated to the partition information after scale-in, that is, partitions 1 to Q. At the same time, delete partitions Q + 1 to W and nodes N + 1 to M, that is, scale the nodes down to 1 to N and the partitions to 1 to Q. After that, both the sending end and the consuming end send and receive data from partitions 1 to Q.

[0088] In this embodiment, during capacity reduction, the sender is controlled to send data only to the partitions retained after capacity reduction, while the consumer still receives data from all partitions. In this way, the data in the partitions to be deleted is continuously consumed without new addition. When the data in the partitions to be deleted meets certain conditions, capacity reduction can be completed. This method supports capacity reduction of partitions and can minimize the migration of partition data and reduce network transmission consumption.

[0089] In this embodiment, the data in each partition supports message backtracking, that is, in the message backtracking mode, the data that has been consumed by the consumer can be consumed again as long as it is within the set retention period.

[0090] Therefore, determining whether the partitions to be deleted meet the deletion conditions can be judged in two cases. Please refer to Figure 3 , the above step S103 may include the following sub-steps:

[0091] S1031, determine whether the data in the partitions to be deleted corresponding to the partition capacity reduction operation needs to be backtracked. If it needs to be backtracked, execute the following step S1032. If it does not need to be backtracked, execute the following step S1033.

[0092] S1032, after processing the data in the partitions to be deleted according to the capacity reduction time limit requirements, delete the nodes and partitions to be deleted.

[0093] S1033, when the data in the partitions to be deleted meets the invalidation conditions, delete the nodes and partitions to be deleted.

[0094] In this embodiment, whether the data in each partition needs to be backtracked can be set according to requirements. In general stable operation business scenarios, the consumed messages generally do not need to be re-consumed, that is, generally, data backtracking is not required. In the entire Kafka cluster, there are multiple Topics, and each Topic has multiple partitions. Among them, most Topics generally do not need backtracking. During implementation, an identifier indicating whether backtracking is required can be set for each Topic, and then separate processing can be performed.

[0095] For the partitions that need to perform data backtracking, if these partitions are the partitions to be deleted corresponding to the capacity reduction operation, then after processing the data in these partitions according to the capacity reduction time limit requirements, the deletion of nodes and partitions can be performed.

[0096] In addition, for the partitions that do not need to perform data backtracking, when the data in these partitions meets the invalidation conditions, that is, when the data is invalid, the deletion of nodes and partitions can be performed.

[0097] Among them, for the partitions that need to perform data backtracking, they can be processed separately according to whether the downsizing time limit requires immediate downsizing operations.

[0098] In one case, if the downsizing time limit requirement indicates that immediate downsizing operations are required, then all the data in the partitions to be deleted is migrated to the remaining nodes, and the nodes to be deleted are deleted. When the set retention period is reached, the partitions to be deleted are deleted.

[0099] Among them, the set retention period can be unrestricted, such as 2 hours, 3 hours, etc. In this case, a full migration of the data in the partitions to be downsized corresponding to the downsizing operation is performed. Although the network overhead is not reduced, the partitions can finally be downsized, reducing the cluster's management of the partitions.

[0100] In another case, if the downsizing time limit requirement indicates that immediate downsizing operations are not required, then when the set retention period is reached, the partitions to be deleted are deleted, and the nodes to be deleted are deleted.

[0101] In addition, for the partitions that do not need to perform data backtracking, the downsizing operation can be processed according to whether the data in the partitions is invalid.

[0102] In this embodiment, in the case of no need for backtracking, it can first be determined whether all the data in the partitions to be deleted is invalid. If all are invalid, then the nodes and partitions to be deleted are deleted.

[0103] If the data in the partitions to be deleted is not all invalid, then the non-invalid data is processed according to the downsizing time limit requirement, and the nodes and partitions to be deleted are deleted.

[0104] Please refer to Figure 4 , in this embodiment, determining whether all the data in the partitions is invalid can be achieved through the following method:

[0105] S10331, for each partition to be deleted, obtain the record offset of the last piece of data in the partition.

[0106] S10332, obtain the consumption offset of each of the consumer terminals for consuming the data in the partition, where the consumption offset represents the maximum record offset of the data in the partition that the consumer terminal has consumed.

[0107] S10333, when each of the consumption offsets is equal to the record offset, determine that all the data in the partition is invalid.

[0108] In this embodiment, the data in the same partition may be consumed by multiple consumers. Each consumer will submit a consumption offset (consumer offset) when consuming the data in the partition to record which piece of data in the partition it has consumed. For each partition, the record offset (offset) of each piece of data stored in the partition will be recorded, and the record offset of the last piece of data is the largest.

[0109] Suppose partition W is the partition to be scaled down. The offset of the last piece of data recorded above is L (since the sender will no longer send data to partition W, L will no longer increase). There are consumers A to Z on it. When all the consumer offsets submitted by all the consumers on this partition have consumed the data of L, that is, the consumption offsets of each consumer are equal to the record offset, it can be considered that all the data in partition W has become invalid and this partition can be deleted.

[0110] If one of the consumers A to Z in this partition hangs up or consumes very slowly, and the consumer offset submitted by this consumer has not been able to reach the offset L of the last message in this partition and only submits the consumer offset to K (K < L). That is, there is a consumer whose consumption offset is less than the record offset, then it is considered that not all the data in this partition has become invalid.

[0111] In the above way, it is judged whether all the data in the partition to be scaled down has become invalid. If all have become invalid, the node and partition to be scaled down can be directly deleted. If not all have become invalid, the unexpired data can be processed according to the requirements of the scaling-down time limit, and then the node and partition to be deleted can be deleted.

[0112] In this embodiment, when processing the unexpired data according to the requirements of the scaling-down time limit and then performing the deletion process, it can be implemented in the following way:

[0113] When the requirements of the scaling-down time limit indicate that the scaling-down operation needs to be performed immediately, the unexpired data in the partition will be migrated to the remaining nodes, and the node to be deleted will be deleted. Wait until all the data on the partition has been consumed or reaches the set retention period, and then delete the partition to be deleted.

[0114] In this case, taking the above as an example, for a scaled-down partition W, if one consumer only submits the consumer offset to K, and the current record offset of the partition is L (K < L), then the data from K to L in the unexpired partition can be migrated to the remaining nodes. Then, first delete the node, and then wait until all the data on partition W has been consumed by each consumer or reaches the retention period, and then delete the partition.

[0115] In this embodiment, only a small amount of data in a small number of partitions needs to be migrated in the above manner, rather than all of them, so the network overhead is relatively small.

[0116] If the scaling-down time limit requirement indicates that the scaling-down operation does not need to be performed immediately, then when the storage time of the last piece of data in the partition reaches the set retention period, it is determined that all the data in the partition becomes invalid, and the partitions and nodes that need to be deleted are deleted.

[0117] In this embodiment, for the case where the scaling-down operation does not need to be performed immediately, assuming the retention period is S, then when the time of the last piece of data (corresponding to offset L) in the partition exceeds the time S, the partitions and nodes that need to be scaled down can be deleted.

[0118] The above is an introduction to the scaling-down processing method of the Kafka cluster. In this embodiment, the cluster expansion operation is also supported. Please refer to Figure 5 The following is an introduction to the cluster expansion processing method.

[0119] S201, when triggering node expansion and partition expansion, distribute the partitions to be expanded under the nodes to be expanded.

[0120] S202, send the data sent by each of the sending ends to each node under all the partitions after partition expansion.

[0121] S203, each of the consuming ends receives data from each node under all the partitions after partition expansion.

[0122] In this embodiment, when the business volume increases, such as when the input traffic in the cluster increases or the input traffic of a Topic increases, such as being greater than a certain threshold, or when it reaches the peak time period of the business, node expansion and partition expansion can be triggered.

[0123] Suppose there is a Kafka cluster with N nodes currently, a Topic named T, and this Topic T has Q partitions. It is necessary to expand the Kafka nodes to M (M > N) and expand the partitions of Topic T to W (W > Q). Distribute the partitions with numbers greater than Q and less than W (Q + 1 to W) of the expanded partitions on the expanded nodes N + 1 to M.

[0124] The partition information of Topic T at the sender side is in the directory brokers / topics / T / producer / partitions in the zookeeper cluster, changing to 1 to W. The number of partitions of Topic T at the consumer side is also in the directory brokers / topics / T / consumer / partitions, changing to 1 to W. After distributing the partitions to be expanded under the nodes to be expanded, when the sender side and the consumer side update the partition information from the corresponding directories respectively, each sender side sends data to partitions 1 to W, and each consumer side receives data from partitions 1 to W.

[0125] In this embodiment, since both expansion and contraction are time-consuming processes, the triggering of expansion and contraction should be avoided too frequently. The cycle of one expansion and contraction can be set to be greater than the retention period.

[0126] In the solution provided in this embodiment, the partitions in the Kafka cluster can be dynamically increased or decreased according to the business volume. When the business volume is large, the number of partitions is increased, and the sender side and the consumer side can send and consume data from the same number of partitions. When the business volume is small, the sender side only sends data to the reserved partitions. The consumer side consumes all partitions. When all the partition data to be contracted meet the deletion conditions, these partitions are deleted.

[0127] For the contraction of Kafka cluster nodes, the migration of partition data can be minimized as much as possible. By the criteria that all consumer sides have consumed completely and the criteria that the retention period has been exceeded, it is judged whether all the partition data is invalid, and the invalid partitions are deleted without migration.

[0128] To further introduce the partition expansion and contraction solution provided in this embodiment, the solution will be further described below in combination with specific examples.

[0129] Suppose the Kafka cluster has 3 Broker nodes numbered 1 - 3, there is a Topic called TEST with 6 partitions numbered 1 - 6, distributed on 3 Broker nodes, as Figure 6 shown. The message retention period set for Topic TEST is 1 hour. Topic TEST does not require data backtracking (the messages that have been consumed in the partitions do not need to be consumed again).

[0130] Among them, the number of partitions of TEST at the sending end is recorded under the directory / brokers / topics / TEST / producer / partitions in zookeeper, and the number of partitions of TEST at the consuming end is recorded under the directory / brokers / topics / TEST / consumer / partitions. At present, there are 6 partitions numbered 1 - 6. Both the sending end and the consuming end send and receive data from partitions 1 - 6, as Figure 7 shown.

[0131] When the peak business period arrives, the cluster is now expanded to 5 Broker nodes, and Topic TEST is expanded to 10 partitions. Broker 4 - 5 are the expanded nodes, and partitions 7 - 10 are the expanded partitions, which are distributed on the nodes of Broker 4 - 5.

[0132] Update the number of partitions of TEST at the sending end (under the directory / brokers / topics / TEST / producer / partitions) and the number of partitions of TEST at the consuming end (under the directory / brokers / topics / TEST / consumer / partitions) to 1 - 10. After the sending end and the consuming end update the number of partitions of Topic TEST, they will both send and receive data from partitions 1 - 10, as Figure 8 shown.

[0133] When the off - peak business period arrives, 5 Broker nodes need to be restored to 3. The partitions of Topic TEST are restored to 6.

[0134] First, update the number of partitions of TEST at the sending end to 1 - 6, and keep the consuming end unchanged, as Figure 9 shown. After the sending end updates the number of partitions of Topic TEST, new messages will only be sent to partitions 1 - 6. The consuming end still keeps partitions 1 - 10 unchanged and fetches data from partitions 1 - 10, as Figure 10 shown. There are still some messages in partitions 7 - 10 that have not been consumed by the consuming end.

[0135] Suppose it is necessary to delete partitions 7 - 10, and each partition has 3 consumers, namely Consumer A, B, and C. Taking partition 7 as an example, the offset of the last message in this partition is 10000. After a short period of time, the consumer offsets submitted by Consumer A, B, and C all reach 10000, as Figure 11As shown. At this time, all the data in partition 7 can be considered invalid, and partition 7 can be deleted. Partitions 8 and 9 are in the same situation as partition 7, and the messages are all invalid. At this time, the partition count of TEST in the consumer side (under the directory / brokers / topics / TEST / consumer / partitions) is updated to [1-6, 10]. At the same time, all the data in partitions 7, 8, and 9 is deleted. Since all the partitions on Broker5 have been deleted, Broker5 is deleted, and the number of Brokers is scaled down to 1-4.

[0136] On partition 10, the offset of the last message in this partition is 9000. Because one of the consumers A, B, and C consumes slowly or crashes, the consumer offset submitted by one consumer always cannot reach 9000, such as (8000). As a result, on partition 10, this partition cannot meet the state where all messages are invalid, as Figure 12 shown.

[0137] If the Broker is not eager to scale down resources, it can wait for up to 1 hour. If after one hour, the messages in partition 10 are still not fully consumed by all consumers (the consumer offset submitted by one consumer always cannot reach 9000), since the retention period of the messages has been reached, this partition can also be considered invalid. The partition count of TEST in the consumer side (under the directory / brokers / topics / TEST / consumer / partitions) is updated to [1-6], and all the data in partition 10 is deleted. Since all the partitions on Broker4 have been deleted, Broker4 is deleted, and the number of Brokers is scaled down to 1-3.

[0138] If the Broker is eager to scale down and release resources, the messages with offsets from 8000 to 9000 in partition 10 are copied to one of the machines among Brokers 1-3. Broker 4 is deleted, and the number of Brokers is scaled down to 1-3. When the partition 10 on Brokers 1-3 meets the condition that all consumers have consumed up to 9000 or reaches the retention period setting of 1 hour. The partition count of TEST in the consumer side (under the directory / brokers / topics / TEST / consumer / partitions) is updated to [1-6], and all the data in partition 10 is deleted.

[0139] At this time, although partition data migration has been carried out, the amount of migrated data is relatively small, only part of the data in some partitions.

[0140] In addition, assuming that data backtracking is required for Topic TEST, the method of considering a partition invalid after all consumers have consumed it cannot be used. Instead, it is necessary to wait until the last message in the partition meets the retention period setting of 1 hour. Only then can the Broker and the partition be scaled down. Alternatively, the partition data can be migrated first. Then, the Broker is scaled down first, and after waiting for the retention period, the partition is scaled down.

[0141] The embodiment of the present application also provides a partition scaling system. As can be seen from the above, the partition scaling system includes a Kafka cluster, multiple consumer ends, and a sender end. The Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions.

[0142] Among them, when triggering node scaling down and partition scaling down, each sender end is used to send data to each node under the partition retained after the partition is scaled down.

[0143] Each consumer end is used to receive data from each node under all partitions before the partition is scaled down.

[0144] When the partition to be deleted corresponding to the partition scaling down operation meets the deletion condition, the management node among the multiple nodes is used to delete the node and the partition to be deleted, and complete the scaling down.

[0145] In this embodiment, the management node can be any node among the multiple nodes and can implement the management of the Kafka cluster.

[0146] On this basis, the partition scaling system further includes a zookeeper cluster. The partition information is recorded in the zookeeper cluster. The partition information for each sender end to send data and the partition information for each consumer end to receive data are recorded in different directories.

[0147] The zookeeper cluster is used to update the partition information in the directory corresponding to each sender end to the partition information after scaling down when triggering node scaling down and partition scaling down, and keep the partition information in the directory corresponding to each consumer end as the original partition information.

[0148] The zookeeper cluster is further used to update the partition information in the directory corresponding to each consumer end to the partition information after scaling down after the scaling down is completed.

[0149] It should be noted that the partition scaling system provided in this embodiment can be used to implement the partition scaling method in any of the above embodiments. For details not described in this embodiment, reference can be made to the relevant descriptions in the above embodiments, and this embodiment will not be elaborated here.

[0150] In summary, for the partition scaling method and system provided in the embodiments of the present application, when triggering node scaling and partition scaling, the data sent by each sender can be sent to each node under the partitions remaining after the partition scaling. Each consumer receives data from each node under all partitions before the partition scaling. In this way, when the partitions to be deleted corresponding to the partition scaling operation meet the deletion conditions, the nodes and partitions to be deleted are deleted to complete the scaling. In this solution, during scaling, the sender is controlled to only send data to the partitions remaining after the scaling, while the consumer still receives data from all partitions. In this way, the data under the partitions to be deleted is continuously consumed without new addition. When the data under the partitions to be deleted meets certain conditions, the scaling can be completed. This method supports the scaling of partitions and can minimize the migration of partition data and reduce network transmission consumption.

[0151] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for partition scaling, characterized in that, Applied to the partition scaling system, the system includes a Kafka cluster, multiple consumer ends, and a sender end. The Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions. The method includes: When triggering node scaling down and partition scaling down, send the data sent by each sender end to each node under the partitions retained after partition scaling down; Each consumer end receives data from each node under all partitions before partition scaling down; When the partitions to be deleted corresponding to the partition scaling down operation meet the deletion conditions, delete the nodes and partitions to be deleted to complete the scaling down; Among them, the step of deleting the nodes and partitions to be deleted when the partitions to be deleted corresponding to the partition scaling down operation meet the deletion conditions includes: Judge whether the data in the partitions to be deleted corresponding to the partition scaling down operation needs to be traced back. If it needs to be traced back, after processing the data in the partitions to be deleted according to the scaling down time limit requirements, delete the nodes and partitions to be deleted; If it does not need to be traced back, when the data in the partitions to be deleted meets the invalidation conditions, delete the nodes and partitions to be deleted.

2. The partition scaling method according to claim 1, wherein The partition scaling system further includes a zookeeper cluster. The partition information is recorded in the zookeeper cluster. The partition information for sending data by each sender end and the partition information for receiving data by each consumer end are recorded in different directories. After the step of triggering node scaling down and partition scaling down, the method further includes: Update the partition information in the directory corresponding to each sender end to the partition information after scaling down, and keep the partition information in the directory corresponding to each consumer end as the original partition information; After the step of deleting the nodes and partitions to be deleted to complete the scaling down, the method further includes: Update the partition information in the directory corresponding to each consumer end to the partition information after scaling down.

3. The partition scaling method according to claim 1, wherein The step of processing the data in the partitions to be deleted according to the scaling down time limit requirements and then deleting the nodes and partitions to be deleted includes: If the scaling down time limit requirements indicate that the scaling down operation needs to be performed immediately, migrate all the data in the partitions to be deleted to the retained nodes, delete the nodes to be deleted, and wait until the set retention period is reached, then delete the partitions to be deleted; If the scaling down time limit requirements indicate that the scaling down operation does not need to be performed immediately, wait until the set retention period is reached, then delete the partitions to be deleted and delete the nodes to be deleted.

4. The partition scaling method according to claim 1, wherein The step of, if it does not need to be traced back, deleting the nodes and partitions to be deleted when the data in the partitions to be deleted meets the invalidation conditions includes: If it does not need to be traced back, judge whether all the data in the partitions to be deleted is invalid. If all is invalid, then delete the nodes and partitions to be deleted; If not all the data in the partitions to be deleted is invalid, process the non-invalid data according to the scaling down time limit requirements, and delete the nodes and partitions to be deleted.

5. The method for partition scaling according to claim 4, wherein The step of processing the data that has not expired according to the requirements of the capacity reduction time limit and deleting the nodes and partitions to be deleted includes: When the requirements of the capacity reduction time limit indicate that the capacity reduction operation needs to be carried out immediately, the data that has not expired in the partition is migrated to the remaining nodes, and the nodes to be deleted are deleted. When all the data on the partition has been consumed or reaches the set retention period, the partition to be deleted is deleted; When the requirements of the capacity reduction time limit indicate that the capacity reduction operation does not need to be carried out immediately, wait until the storage time of the last piece of data in the partition reaches the set retention period, determine that all the data in the partition has expired, and delete the partition and nodes to be deleted.

6. The partition scaling method according to claim 4, wherein The step of determining whether all the data in the partition to be deleted has expired includes: For each partition to be deleted, obtain the record offset of the last piece of data in the partition; Obtain the consumption offset of each consumer for consuming the data in the partition, where the consumption offset represents the maximum record offset of the data in the partition that the consumer has consumed; When all the consumption offsets are equal to the record offset, determine that all the data in the partition has expired.

7. The method for partitioned scaling according to any one of claims 1-6, characterized in that The method further includes: When triggering node expansion and partition expansion, distribute the partitions to be expanded under the nodes to be expanded; Send the data sent by each sender to each node under all the partitions after partition expansion; Each consumer receives data from each node under all the partitions after partition expansion.

8. A partition scaling system, characterized in that, For implementing the partition expansion and contraction method described in any one of claims 1-7, the system includes a Kafka cluster, multiple consumers and senders. The Kafka cluster includes multiple nodes, and the multiple nodes are divided into multiple partitions; When triggering node contraction and partition contraction, each sender is used to send data to each node under the partitions retained after partition contraction; Each consumer is used to receive data from each node under all the partitions before partition contraction; When the partition to be deleted corresponding to the partition contraction operation meets the deletion condition, the management node in the multiple nodes is used to delete the nodes and partitions to be deleted to complete the contraction.

9. The partition scaling system according to claim 8, wherein The partition expansion and contraction system further includes a zookeeper cluster. The partition information is recorded in the zookeeper cluster. The partition information of the data sent by each sender and the partition information of the data received by each consumer are recorded in different directories; The zookeeper cluster is used to update the partition information in the directory corresponding to each sender to the partition information after contraction when triggering node contraction and partition contraction, and keep the partition information in the directory corresponding to each consumer as the original partition information; The zookeeper cluster is further used to update the partition information in the directory corresponding to each consumer to the partition information after contraction after the contraction is completed.

Citation Information

Patent Citations

  • Sequential message-based capacity expanding and shrinking method and device and electronic equipment

    CN109697187A