Message synchronization method, device, node and readable storage medium
By using high-performance and low-performance storage areas in Kafka's backup partitions, the synchronization of messages to high-performance areas is preferred, which solves the problem of low message synchronization efficiency of Kafka's message system in high throughput and multi-node cluster environments, and improves system performance.
Patent Information
- Application Number
- CN202111506274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The existing Kafka message system is less efficient during message synchronization, especially in high throughput and multi-node cluster environments, resulting in multiple increase in disk capacity and write rates for message storage and backup, affecting system performance.
By using two storage areas in the backup partition, a first storage area with high access performance and a second storage area with low performance, messages synchronized from the primary partition are preferred in the first storage area. The specific steps include obtaining the offset position of the most recent consumption message and the offset position of the most recent synchronization message. If the difference is less than the preset value and there is enough free space for the first storage area, the synchronizes the message to the first storage area.
By prioritizing the synchronization of messages to the first storage area with high access performance, data synchronization response time is reduced, message synchronization efficiency is improved, and system performance bottlenecks in high throughput and multi-node cluster environments are reduced.
Smart Images

Figure CN114253743B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of message processing, and in particular to a message synchronization method, device, node and readable storage medium. Background Art
[0002] Kafka is a high-throughput distributed publish-subscribe messaging system. Producers push messages to Kafka, which manages messages and provides an interface for consumers to pull messages from Kafka. Kafka is usually deployed in a cluster consisting of multiple nodes. In order to ensure the security and reliability of messages, Kafka usually stores messages and their copies on disks of different nodes in the cluster according to partitions. The partition responsible for providing services to the outside world is the primary partition, and the partition responsible for backing up messages is the backup partition. Messages in the primary partition are synchronized to the backup partition according to preset rules. How to improve the efficiency of message synchronization is an urgent problem to be solved by technicians in this field. Summary of the invention
[0003] The present invention provides a message synchronization method, device, node and readable storage medium, which can improve the efficiency of message synchronization.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] In a first aspect, the present invention provides a message synchronization method, which is applied to a first node in a Kafka cluster, wherein the Kafka cluster also includes a second node communicating with the first node, wherein the second node includes a primary partition, wherein there are messages in sequentially increasing order, and the first node includes a backup partition, wherein the backup partition is used to store copies of messages in the primary partition, and the backup partition includes a first storage area and a second storage area, wherein the access performance of the first storage area is greater than that of the second storage area, and the method includes: obtaining a first number of a consumption message consumed most recently, wherein the first number is used to characterize an offset position of the consumption message in the primary partition; obtaining a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to characterize an offset position of the synchronization message in the primary partition; if the difference between the second number and the first number is less than a preset value, and the free area in the first storage area meets a preset condition, then synchronizing the message between the third number of the write message most recently written in the primary partition and the second number to the first storage area.
[0006] In a second aspect, the present invention provides a message synchronization device, which is applied to a first node in a Kafka cluster, wherein the Kafka cluster also includes a second node communicating with the first node, wherein the second node includes a primary partition, wherein there are messages in sequentially increasing order, wherein the first node includes a backup partition, wherein the backup partition is used to store copies of messages in the primary partition, wherein the backup partition includes a first storage area and a second storage area, wherein the access performance of the first storage area is greater than that of the second storage area, wherein the device includes: an acquisition module, which is used to acquire a first number of a consumption message consumed most recently, wherein the first number is used to characterize an offset position of the consumption message in the primary partition; the acquisition module, which is also used to acquire a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to characterize an offset position of the synchronization message in the primary partition; and a synchronization module, which is used to synchronize messages between a third number of a write message most recently written in the primary partition and the second number to the first storage area if the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets a preset condition.
[0007] In a third aspect, the present invention provides a node, comprising a memory and a controller, wherein the controller implements the message synchronization method as described above when executing the computer program.
[0008] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the message synchronization method as described above when executed by a controller.
[0009] Compared with the prior art, the present invention forms a backup partition by using a first storage area and a second storage area, and the access performance of the first storage area is greater than that of the second storage area. When synchronizing messages from the primary backup partition to the backup partition, first, the first number of the most recently consumed consumption message and the second number of the locally pre-stored synchronization message most recently synchronized from the primary partition are obtained, the first number is used to characterize the offset position of the consumption message in the primary partition, and the second number is used to characterize the offset position of the synchronization message in the primary partition. If the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets the preset conditions, the messages between the third number and the second number of the most recently written write message in the primary partition are synchronized to the first storage area. By synchronizing the messages from the primary partition to the first storage area in a priority manner and the access performance of the first storage area is greater than that of the second storage area, the message synchronization efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 An example diagram of a Kafka cluster provided in an embodiment of the present invention.
[0012] Figure 2 An example block diagram of a node provided in an embodiment of the present invention.
[0013] Figure 3 An example diagram of offset numbering of messages provided in an embodiment of the present invention.
[0014] Figure 4 An example flowchart of a message synchronization method provided by an embodiment of the present invention.
[0015] Figure 5 An example diagram of messages in a primary partition and a backup partition provided by an embodiment of the present invention.
[0016] Figure 6 An example flowchart of another message synchronization method provided by an embodiment of the present invention.
[0017] Figure 7 An example diagram of message migration provided by an embodiment of the present invention.
[0018] Figure 8 An example flowchart of another message synchronization method provided by an embodiment of the present invention.
[0019] Fig. 9 An example flowchart of another message synchronization method provided by an embodiment of the present invention.
[0020] Fig.10 A diagram showing a specific application example of the message synchronization method provided in an embodiment of the present invention.
[0021] Fig.11 A block diagram of a message synchronization device 100 provided in an embodiment of the present invention is shown.
[0022] Icons: 10 - node; 20 - producer; 30 - consumer; 11 - controller; 12 - memory; 13 - bus; 14 - communication interface; 100 - message synchronization device; 110 - acquisition module; 120 - synchronization module; 130 - replacement module. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0024] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0026] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear to indicate an orientation or position relationship, they are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the invention is usually placed when used. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0027] In addition, the terms “first”, “second”, etc., if used, are merely used to distinguish between the descriptions and should not be understood as indicating or implying relative importance.
[0028] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.
[0029] Producers push the produced messages to Kafka, which is responsible for managing messages and providing an interface for consumers to pull messages from Kafka. In order to support high-throughput message storage and publishing, Kafka is usually deployed in a cluster consisting of multiple nodes to form a Kafka cluster. The nodes are also called brokers. Generally speaking, the more brokers there are, the higher the cluster throughput. The Kafka cluster can interact with multiple producers and consumers at the same time. Kafka classifies messages according to topics. Each message published to the Kafka cluster needs to specify a topic. A topic can be divided into multiple partitions. The messages in each partition are ordered. A broker can store one or more partitions. In order to ensure the reliability of messages, the Kafka cluster will also save one or more copies for each partition. For each partition, there is a primary partition and at least one backup partition. The primary partition is used to interact with producers or consumers. The backup partition is a copy of the primary partition and is used to back up messages in the primary partition to improve the reliability of messages in the partition. Usually, for any partition, its primary partition and backup partition are distributed on different brokers. For ease of description, the embodiment of the present invention is described for a partition. For this partition, the broker to which its backup partition belongs is called the first node, and the broker to which the primary partition belongs is called the second node. In fact, in actual applications, each broker can include multiple primary partitions and multiple backup partitions. The backup partition and the primary partition on the same broker do not correspond to the same partition.
[0030] It should be noted that in order to configure and manage the Kafka cluster, Zookeeper is introduced to save the metadata information of the cluster, including cluster configuration information and cluster management information. Zookeeper is an open source distributed application coordination service and a software that provides consistency services for distributed applications. The functions provided include: configuration maintenance, domain name service, distributed synchronization, group service, etc. In this embodiment, Zookeeper includes at least the following functions: (1) save the cluster status information of Kafka; (2) be responsible for monitoring the status information of each broker, and once a broker is found to be down, be responsible for selecting a new controller from the remaining brokers; (3) save the offset of the consumer's consumption information so that the consumer can continue to consume according to the offset next time.
[0031] Please refer to Figure 1 , Figure 1This is an example diagram of a Kafka cluster provided in an embodiment of the present invention. The Kafka cluster includes three nodes 10: broker 1 to broker 3, and simultaneously interacts with four producers 20 and four consumers 30 for messages. Figure 1 In the example, the primary partition of partition 0 is distributed on broker 1, and its two backup partitions are distributed on broker 2 and broker 3 respectively. The primary partition of partition 1 is distributed on broker 2, and its two backup partitions are distributed on broker 1 and broker 3 respectively. The primary partition of partition 3 is distributed on broker 3, and its two backup partitions are distributed on broker 1 and broker 2 respectively. For partition 0, broker 1 is the second node, and broker 2 and broker 3 are both the first nodes. For partition 1, broker 2 is the second node, and broker 1 and broker 3 are both the first nodes. For partition 2, broker 3 is the second node, and broker 1 and broker 2 are both the first nodes.
[0032] In this embodiment, the node 10 may be a physical device such as a host or a server, or may be a virtual machine that implements the same functions as the physical device.
[0033] The producer 20 is a client that produces messages, and may be a host, a host group, or a host cluster.
[0034] Consumer 30 is a client that consumes messages, and may be a host, a host group, or a host cluster.
[0035] based on Figure 1 The embodiment of the present invention also provides a block diagram of a node 10, which is used to execute a message synchronization method applied to a first node when the node 10 acts as a first node. Figure 2 , Figure 2 This is a block diagram of an example node provided in an embodiment of the present invention. The node 10 includes a controller 11, a memory 12, a bus 13 and a communication interface 14. The controller 11, the memory 12 and the communication interface 14 are connected via the bus 13.
[0036] The memory 12 is used to store programs, such as the message synchronization device 100 in an embodiment of the present invention. The message synchronization device 100 includes at least one software function module that can be stored in the memory 12 in the form of software or firmware. After receiving the execution instruction, the controller 11 executes the program to implement the message synchronization method disclosed in the embodiment of the present invention.
[0037] The memory 12 may include a high-speed random access memory (RAM) and may also include a non-volatile memory (NVM).
[0038] The controller 11 may be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the controller 11 or an instruction in the form of software. The above controller 11 may be a general processor, including a central processing unit (CPU), a microcontroller unit (MCU), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), an embedded ARM and other chips.
[0039] There may be multiple communication interfaces 14 , and the node 10 may communicate with other devices through different communication interfaces 14 .
[0040] based on Figure 1 In the application scenario, the existing Kafka is used as a message middleware to store messages and decouple various systems. Kafka's message throughput is particularly large. For example, when performing parsing tasks on 10,000 network camera IPC (IP Camera) devices, the inflow and outflow traffic often reaches more than 300M / s. If the inflow is 150M / s, the write speed of the disk storing messages must also reach 150M / s. Kafka stores data in the Topic dimension (partition of the actual Topic).
[0041] At the same time, in order to prevent Kafka node failures, 2-3 copies may be set for each Topic, which is equivalent to storing 2-3 copies of each data. In this way, the disk write rate will be multiplied by 2-3 times to 300M / s or 450M / s. At the same time, the disk capacity requirement will also be multiplied by 2-3 times. If only 1TB of disk space is required, 2TB or 3TB will be used.
[0042] Ceph distributed file system is usually used to store messages. Ceph distributed file system includes hosts with many physical disks. The number of disks on each host can be 3-60. If a 10G network card is configured, the network will basically not have performance bottlenecks when the system is running. However, if the number of disks is small, the performance bottleneck is often caused by disk access performance. Therefore, the disk input and output IO usage rate is particularly high at this time. Even if there are many disks, the distributed file system needs to write all nodes to return success because the message data to be stored is too scattered, resulting in a slow response to the client.
[0043] As a solution, you can use the local disk as a storage system, but many local disks are shared by multiple programs. The excessive disk usage caused by the Kafka program saving messages will affect other programs, or other programs will affect the Kafka program.
[0044] In addition, the principle of synchronizing data among existing Kafka partitions is as follows: taking a partition with a master partition and two followers as an example, the master is responsible for receiving data sent by the sender and writing it to disk, and is also responsible for reading data requests from the consumer. This is equivalent to the sending and consuming clients only interacting with the master, and the follower simply pulls data from the master and writes it to disk for backup.
[0045] In the existing Kafka partition mechanism, if the machine where the master is located crashes or restarts, the follower may be elected as the new master and provide various services. This is the partition disaster recovery feature of Kafka and the reason why the follower synchronizes data from the master.
[0046] Kafka uses offset to record the number of messages in a partition. Each message written will have an offset number generated, which is an auto-increment rule. Please refer to Figure 3 , Figure 3 An example diagram of offset numbering of messages provided in an embodiment of the present invention, Figure 3 There are 338 messages in the partition, and the offset numbers range from 10001 to 10338, and the numbers increase sequentially.
[0047] In the synchronization mechanism, not every message is synchronized immediately after it is received, but it is continuously polled and pulled for synchronization. Therefore, the master's offset will be greater than the follower's offset if it is not synchronized in time.
[0048] In order to ensure the reliability of messages, Kafka sends 3 types of response acks (which can be represented by 3 different numbers: 0, 1, -1). 0 means that after sending, it will return without receiving any response. It is possible that the server has not received it, but the sending client will not process it. 1 means that the master receives and returns a successful response to the sender. If the Kafka node where the master is located is down, the follower is not synchronized, and the follower becomes the master, the message may be lost. -1 means that the master and the follower must receive the message before returning a successful response. In this mode, the message will not be lost unless all the Kafka nodes where the master and the follower are located are down.
[0049] Kafka's message deletion policy is simply to delete messages based on the message retention time or size setting. For example, if the message retention time is set to 12 hours, data older than 12 hours will be deleted.
[0050] The existing Kafka mainly has the following problems:
[0051] 1. The disk capacity used by Kafka increases exponentially as the number of partition (or topic) copies increases.
[0052] 2. The disk write rate used by Kafka increases exponentially as the number of copies of the partition (or topic) increases.
[0053] In view of this, an embodiment of the present invention provides a message synchronization method, device, node and readable storage medium to solve the above problems, which will be described in detail below.
[0054] exist Figure 1 and Figure 2 Based on the present invention, an embodiment of the present invention provides a message synchronization method, which is applied to Figure 1 The first node in Figure 2 When node 10 acts as the first node, please refer to Figure 4 , Figure 4 A flowchart of a message synchronization method provided by an embodiment of the present invention includes the following steps:
[0055] Step S100, obtaining the first number of the most recently consumed consumption message, wherein the first number is used to represent the offset position of the consumption message in the primary partition.
[0056] In this embodiment, the messages in the primary partition are numbered in order. For example, the number of the message received for the first time is 1, and the number of the message received for the second time is 2, and so on. The first number is the number of the most recently consumed message, and the number also represents the offset position of the message in the primary partition. For example, if the first number is 0, it indicates that the offset position of the message in the primary partition is 0, that is, it is the first message in the primary partition.
[0057] In this embodiment, the number of consumers can be one or more. Usually, in a stable business scenario, all consumers of a topic are fixed. When there is one consumer, the first number is the number of the consumer message consumed by the consumer most recently. When there are multiple consumers, the first number is the minimum number of the consumer message consumed by each consumer most recently. For example, there are three consumers: a, b, and c, and the numbers of their most recent consumer messages are 1, 2, and 3, respectively. Then the first number is the minimum value of 1, 2, and 3, that is, 1.
[0058] In this embodiment, the first number will change as the consumer consumes the message, and the first node can obtain the latest first number from zookeeper when synchronizing the message.
[0059] Step S110, obtaining a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to represent an offset position of the synchronization message in the primary partition.
[0060] In this embodiment, the first node can synchronize the messages of the main partition in real time, or periodically synchronize the messages of the main partition. The second number is the number of the synchronization message that the first node most recently synchronized from the main partition. The number also represents the second number of the synchronization message that most recently synchronized from the main partition.
[0061] Step S120: if the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets a preset condition, the message between the third number and the second number of the most recently written message in the primary partition is synchronized to the first storage area.
[0062] In this embodiment, the backup partition includes a first storage area and a second storage area. The access performance of the first storage area is higher than that of the second storage area. As a specific implementation method, the first storage area can be a memory and the second storage area can be a disk. When there is enough free space in the first storage area, the messages of the primary partition are synchronized to the first storage area of the backup partition first.
[0063] In this embodiment, the preset value is used to characterize the maximum number of messages stored in the first storage area in the backup partition. The preset value is related to the memory size of the first node and the write performance of the disk of the first node. Because, in some scenarios, it is necessary to migrate the messages in the first storage area to the second storage area. If the preset value is too large, the message migration takes too long, which eventually leads to a decrease in message processing performance. Usually, the disk performance can be tested in advance, and the preset value can be determined based on the test results. For example, the preset value is set to 100, that is, the first storage area stores a maximum of 100 messages.
[0064] In this embodiment, it takes about a few seconds, such as 1-10 seconds, from the time a message is sent to Kafka to the time when all business consumers process and consume it successfully and submit the offset number. As a specific implementation, the first storage area can be set in the following way: Assuming that the message inflow rate is QM / s, and the message is received from Kafka and consumed in about N seconds, the size of the second storage area of each broker can be set to about Q x N.
[0065] In this embodiment, the free area in the first storage area satisfies the preset condition by, for example, a ratio of the free area to the total available area is greater than or equal to a preset ratio, or a size of the free area is greater than or equal to a preset size.
[0066] In this embodiment, the third number is the number of the most recently written message, and the number also represents the offset position of the written message in the primary partition. The third number can be the same as the second number, in which case the message in the primary partition is the same as the message in the backup partition, and the third number can also be greater than the second number, in which case some messages in the primary partition have not yet been synchronized to the backup partition.
[0067] In this embodiment, under normal circumstances, the speed at which producers produce messages and the speed at which consumers consume messages will be relatively stable, that is, the difference between the second number and the first number will not be too large. Only when any consumer has an abnormality that causes the consumption speed to decrease significantly, the difference between the second number and the first number will gradually increase, and there will be a risk that the first storage area will be insufficient. At this time, some or all of the messages in the first storage area need to be migrated to the second storage area so that the latest synchronized messages are stored in the first storage area.
[0068] For a clearer explanation of the meaning of the first, second and third numbers, please refer to Figure 5 , Figure 5 An example diagram of messages in a primary partition and a backup partition provided by an embodiment of the present invention, Figure 5In the data, there are 8 messages in the primary partition: message 0 to message 7, among which message 7 is the most recently written message, that is, the third number is 7. The messages in the primary partition are consumed by 3 consumers, namely: consumer a, consumer b and consumer c. The consumption messages of the three consumers most recently are: message 3, message 4 and message 5, and the corresponding numbers are 3, 4 and 5 respectively. Then the first number is the minimum value of the three: 3. 7 messages have been synchronized in the backup partition: message 0 to message 6. The most recently synchronized message is message 6, so the second number is 6.
[0069] It should be noted that in order to ensure the reliability of the second and third numbers, as a specific implementation method, the second and third numbers are synchronized to the partition metadata file of the first node each time a message is processed for persistent storage. There is no special essential difference between the two, and both record the latest offset number in the local message. When the primary partition switches, the two will also convert to each other.
[0070] The above method provided in this embodiment can reduce data synchronization response time and improve message synchronization efficiency by synchronizing messages from the primary partition to the first storage area first. Since messages are synchronized to the first storage area with high access performance first, the data synchronization response time can be reduced.
[0071] In this embodiment, Figure 4 This is for the scenario where messages that need to be synchronized can be stored in the first storage area. In fact, when a consumer consumes messages abnormally, the cause of the abnormal consumption may be that the consumer is down or the consumption processing speed is too slow. At this time, the first number will increase slowly or not at all, and the difference between the second number and the first number will gradually increase, which will lead to the risk of insufficient first storage area. The embodiment of the present invention provides a specific implementation method for this scenario, please refer to Figure 6 , Figure 6 Another example flow chart of a message synchronization method provided by an embodiment of the present invention, the method further includes the following steps:
[0072] Step S130: If the difference between the second number and the first number is greater than or equal to a preset value, or the free area in the first storage area does not meet the preset condition, the number of entries to be migrated is determined according to the second number, the third number and the preset value.
[0073] In this embodiment, after each consumer processes the message in the primary partition, it will submit the offset number it has consumed (the smallest offset number among all consumers is the first number) to Kafka. The offset number submitted by each consumer indicates the position of the message it has processed and consumed. It will also be submitted to a directory in a certain format under Zookeeper, such as / offsets / topic(name) / partition(number) / A~Z(consumer name). The lowest value of all offset numbers under this directory is calculated regularly (for example, 1 second or 3 seconds). If the lowest value is updated, it is updated and written to the / offsets / topic(name) / partition(number) / lowest node, which is defined as offset lowest (i.e., the first number). Offset lowest indicates the minimum value of the offset that has been processed by all consumers of this primary partition. Messages below offset lowest are all processed messages and can be deleted. Loss is allowed. If the lowest value is not updated, offset lowest can be not updated in this round of calculation.
[0074] The backup partition follower will listen to the value under the / offsets / topic(name) / partition(number) / lowest node and obtain the lowest offset. According to the characteristics of Zookeeper, if the lowest offset is updated, the follower will be notified. The follower will hold the latest value of the lowest offset in real time. At the same time, since all consumers are constantly consuming messages, the lowest offset will continue to increase.
[0075] In this embodiment, as a specific implementation, the process of calculating the number of entries to be migrated may be:
[0076] First, the number difference between the second number and the third number is calculated.
[0077] In this embodiment, since the messages in the primary partition are increasing in sequence, the number difference can represent the number of messages between the second number and the third number. Normally, the second number is less than or equal to the third number.
[0078] Secondly, the minimum value between the serial number difference and the preset value is used as the number of entries to be migrated.
[0079] In this embodiment, since the maximum number of messages stored in the first storage area is a preset value, when the number difference is less than or equal to the preset value, messages with the number difference can be migrated from the first storage area, and then the latest message to be synchronized can be synchronized to the first storage area. When the number difference is greater than the preset value, all messages in the first storage area need to be migrated to the second storage area first, and then the most recently written preset value messages between the second number and the third number are stored in the first storage area, and other messages between the second number and the third number are stored in the second storage area.
[0080] Step S140: Migrate the earliest number of messages stored in the first storage area to the second storage area.
[0081] Step S150: Synchronize the target messages of the migration number starting from the third number to the first storage area.
[0082] In this embodiment, in order to store the most recently synchronized messages in the first storage area, when migrating the messages in the first storage area, the earliest stored messages are migrated first, and when synchronizing, the most recently written messages are synchronized to the first storage area first, that is, the messages numbered between the second number and the third number and closest to the third number.
[0083] Step S160: Synchronize the messages between the third number and the second number except the target message to the second storage area.
[0084] In this embodiment, to more clearly illustrate the process of message migration, please refer to Figure 7 , Figure 7 This is an example diagram of message migration provided by an embodiment of the present invention. Taking the preset value as 3 as an example, Figure 7 (a) is an example diagram of message migration when the number difference is greater than the preset value. Figure 7 In (a), the third number is 7 and the second number is 3. During migration, 1, 2, and 3 in the first storage area need to be migrated to the second storage area, 7, 6, and 5 need to be synchronized to the first storage area, and 4 needs to be stored in the second storage area. Figure 7 (b) Example diagram of message migration when the number difference is less than the preset value. Figure 7 In (b), the third number is 5 and the second number is 3. It is necessary to migrate 1 and 2 to the second storage area and synchronize 4 and 5 to the first storage area.
[0085] The above method provided in this embodiment can also correctly synchronize messages in the scenario where consumers consume messages abnormally, thereby ensuring the reliability of messages in this abnormal scenario. Since a consumer consumption abnormality only affects the topic or partition where it is located, the embodiment of the present invention can achieve a multiple reduction in disk capacity and disk IO in most cases. Only in extreme cases (such as most consumers have consumption abnormalities) will the disk capacity usage be the same as the prior art.
[0086] It should be noted that when the abnormal consumption returns to normal, the message consumption speed will gradually increase, the increase of the first number will also accelerate, and the difference between the second number and the first number will gradually decrease. When the difference is less than the preset value, the messages in the second storage area will be gradually deleted, and the latest synchronized messages will be stored in the first storage area first.
[0087] In this embodiment, messages smaller than the first number are messages that have been consumed by all consumers. There is no need to store such messages. In order to avoid such messages occupying storage space, the embodiment of the present invention also provides a processing method for messages smaller than the first number. The processing method is based on Figure 4 , please see Figure 8 , Figure 8 Another example flow chart of a message synchronization method provided by an embodiment of the present invention, the method further includes the following steps:
[0088] Step S170: Delete the messages in the backup partition whose numbers are smaller than the first number.
[0089] In this embodiment, if a message has been processed by all consumers that need to process it, it means that the message has no use value and can be deleted. Since the first number is the smallest number of the most recent message consumed by all consumers, messages smaller than the first number must have been processed by all consumers and can be deleted to promptly release the storage space occupied by these messages of no use value.
[0090] In this embodiment, the first number can be in the first storage area in the backup partition or in the second storage area in the backup partition. If a message with a number smaller than the first number is in the first storage area, it will be deleted from the first storage area. If a message with a number smaller than the first number is in the second storage area, it will be deleted from the second storage area. If a message with a number smaller than the first number is in both the first storage area and the second storage area, the message with a number smaller than the first number in each of them will be deleted.
[0091] In this embodiment, although the messages in the first storage area are often subjected to random read and write operations such as deletion, due to the high performance of the first storage area, it will not have much impact on the performance of the entire system. Under normal circumstances, the message consumption speed is quite fast. Once consumed, there is no need to store the message, thereby avoiding performance consumption and storage space consumption caused by storing the message in the second storage area.
[0092] It should be noted that step S170 can also be performed with Figure 6 In order to achieve the corresponding technical effects in both scenarios, the specific implementation method is similar to the above-described method and will not be repeated here.
[0093] In this embodiment, since only the primary partition provides services to the outside, when an exception occurs in the second node to which the primary partition belongs, in order to continue to provide services to the outside, maintain the disaster recovery feature of the Kafka node downtime, ensure that message data is not lost, and improve the reliability of message data, a new primary partition will be selected from the backup partition at this time to improve the reliability of the entire Kafka cluster. Therefore, the embodiment of the present invention also provides a processing method when an exception occurs in the second node, please refer to Fig. 9 , Fig. 9 Another example flow chart of a message synchronization method provided by an embodiment of the present invention, the method comprises the following steps:
[0094] Step S200: When an abnormality is detected in the second node, it is determined whether the backup partition meets the replacement condition.
[0095] In this embodiment, the abnormality of the second node may be abnormal power failure of the second node, or a software error may occur causing the second node to restart, or an abnormality in external communication may cause the second node to become offline, and the like.
[0096] In this embodiment, the replacement condition is used to characterize that the backup partition can replace the primary partition if the demand is met. The scenarios in which the backup partition meets the replacement condition include at least the following three: (1) If there is a message with the second number in the second storage area, it is determined that the backup partition meets the replacement condition; in this scenario, the existence of the second number in the second storage area means that all messages have been stored in the second storage area, there is no possibility of message loss, and the backup partition can replace the primary partition; (2) If there is no message with the second number in the second storage area, and the second number is equal to the first number, it is determined that the backup partition meets the replacement condition; in this scenario, it means that all messages have been processed normally by the consumer, no new messages are written during the abnormality of the second node, there are no unprocessed messages, there is no possibility of message loss, and the backup partition can replace the primary partition. (3) If there is no message with the second number in the second storage area, and the second number is greater than the first number, and partial message loss is allowed, it is determined that the backup partition meets the replacement condition. In this scenario, it means that there are unprocessed messages in the first storage area of the backup partition, and the unprocessed messages are not written to the second storage area. At this time, you can set the same type of ack on the topic according to the message reliability ack type of the kafka cluster. If the ack requires 0 or 1, it means that some messages are allowed to be lost, and the backup partition can replace the primary partition. If the ack requires -1, it means that only the original primary partition has the corresponding data. You can choose to set that the backup partition cannot replace the primary partition and does not participate in the election. It must wait for the original primary partition to recover. During the recovery of the original primary partition, the partition is in an unavailable state and cannot send or receive messages.
[0097] Step S210: If the backup partition meets the replacement condition, the backup partition is used as the new primary partition according to the election mechanism.
[0098] In this embodiment, if there are multiple backup partitions and all of them meet the replacement conditions, one of them can be determined as the new primary partition through an election mechanism. Algorithms that can implement the election mechanism include, but are not limited to, Zookeeper's Zab, Raft, and Viewstamped Replication.
[0099] In this embodiment, once the backup partition is selected as the new primary partition, in order to ensure the reliability of messages while the new primary partition provides services to the outside normally, as a specific implementation method, the implementation steps of selecting the backup partition as the new primary partition according to the election mechanism may be:
[0100] First, the message in the first storage area is written into the second storage area.
[0101] Secondly, writing of messages to the first storage area is prohibited.
[0102] In this embodiment, once the backup partition becomes a new primary partition, the backup partition needs to store messages in the same manner as the primary partition, that is, the written messages are no longer stored in the first storage area but only in the second storage area.
[0103] It should be noted that if the backup partition becomes the new primary partition, the original primary partition becomes the new backup partition after recovery, and the new backup partition synchronizes data from the new primary partition. The synchronization method is: first obtain the latest first number. If the first number is less than or equal to the second number in the second storage area of the backup partition at this time, then start from the second number in the second storage area and synchronize messages from the primary partition. If the first number is greater than the second number in the second storage area of the backup partition at this time, since messages less than the first number are already consumed, there is no need to synchronize them. Therefore, start from the first number and synchronize messages from the primary partition.
[0104] It should also be noted that if the second node to which the primary partition belongs is normal and the first node to which the backup partition belongs needs to be restarted normally, the messages in the first storage area of the backup partition need to be written into the second storage area before the first node is restarted normally, to prevent the restarted backup partition from being elected as the primary partition, which would cause the messages in the first storage area of the backup partition to be lost. In addition, after the first node is restarted, data needs to be synchronized from the primary partition, and the synchronization method is the same as the synchronization method of the new backup partition synchronizing data from the new primary partition described above, which will not be repeated here.
[0105] It should also be noted that if the Kafka cluster restarts normally or abnormally, that is, all nodes in the Kafka cluster are restarted at this time, at this time, both the primary partition and the backup partition can be judged whether the primary partition and the backup partition can be used as the primary partition according to whether the replacement conditions are met in step S200. If all of them can be used as the primary partition, one is selected as the primary partition according to the election mechanism. If some of them can be used as the primary partition, one is selected as the primary partition from those that can be used as the primary partition according to the election mechanism.
[0106] It should also be noted that if all nodes in the kafka cluster restart abnormally, ack is equal to -1, and the disk of the primary partition is damaged again and cannot be restored, the data is not backed up, and data loss may occur. Therefore, for application scenarios with very strict data requirements, configuration items can be added. The first storage area and the second storage area of the backup partition can be configured in the topic dimension, that is, a part of the topics are configured to use or not use the first storage area and the second storage area provided by the embodiment of the present invention. Topics that do not use this solution use the existing technology, that is, both the primary partition and the backup partition are directly written to the second storage area to prevent data loss caused by disk damage.
[0107] In order to more clearly describe the message synchronization method in the above embodiment as a whole, the embodiment of the present invention also provides a specific application example diagram. Taking the topic with the business name VMS in the Kafka cluster as an example, the topic corresponds to 1 partition, and the partition number is 0. The partition has a master partition (hereinafter referred to as the master partition) and two copies, namely two backup partitions (hereinafter referred to as follower partitions). The backup partition includes a memory segment storage area (ie, the first storage area) and a disk segment storage area (ie, the second storage area). The partition has 3 consumers, namely consumer A, consumer B, and consumer C, all of which consume data from the topic.
[0108] If the inflow flow of the Kafka cluster is 150M, it takes about 5-10 seconds for a message to be successfully processed and consumed by all consumers from production. The memory segment storage size of each Kafka node follower partition is set to 1.5G. The index of writing messages to disk is 100 (that is, the default value is 100).
[0109] Please refer to Fig.10 , Fig.10 A specific application example diagram of the message synchronization method provided by an embodiment of the present invention, Fig.10 (a) is a specific application example diagram of message synchronization under normal consumption conditions provided by an embodiment of the present invention, Fig.10 In (a), the follower partition continuously synchronizes data from the master partition. The latest message of the master partition (corresponding to the third-numbered message) has offset master = 10338, and the latest message synchronized to the follower partition (corresponding to the second-numbered message) has offset follower = 10337. The message offset processed by consumers A and B is 10335, and the value 10335 is written to / offsets / VMS / 0 / A and / offsets / VMS / 0 / B of Zookeeper respectively. The message offset processed by consumer C is 10334, and the value 10334 is written to / offsets / VMS / 0 / C of Zookeeper. The minimum value of the message offset processed by consumers A, B, and C under / offsets / VMS / 0 / is calculated regularly, which is 10334, and written to / offsets / VMS / 0 / lowest with a value of 10334. This is the offset lowest of the partition (corresponding to the first-numbered message).
[0110] The follower partition monitors the changes of the / offsets / VMS / 0 / lowest node and obtains the value of the lowest offset. Since the offset of the latest message synchronized is 10337, and the difference with 10334 is less than 100, the message data is only stored in the memory segment, and the message data below the lowest offset in the memory segment is deleted. For example Fig.10 The message 10333 in (a).
[0111] Please refer to Fig.10 (b) Fig.10 (b) is a specific application example diagram of message synchronization under abnormal consumption conditions provided by an embodiment of the present invention. Taking consumer C as an example, / offsets / VMS / 0 / C does not increase or increases slowly, and the corresponding offsetlowest value also does not increase or increases slowly. The processing offset of consumer C stops at 20002, resulting in offsetlowest=20002, and the offset follower synchronized from the master partition to the follower is 20557. At this time, the number of synchronized messages is greater than 100, and messages with offsets below 20458 are flushed to the disk, and the same data in the memory segment is deleted.
[0112] It should be noted that if consumer C returns to normal, the processed offset grows rapidly, and the difference between offsetlowest and offsetfollower returns to less than 100, all data in the disk segment is deleted, and processing is resumed only in the memory segment.
[0113] It should also be noted that if the restart is normal, the follower partition will write the messages in the memory segment to the disk before restarting. After the restart, the partition participating in the master election (formerly the master partition or the follower partition) first determines whether the latest message offset recorded by itself (previously it may be the offset master or the offset follower) exists on the disk, mainly including the following three situations: (1) If it exists, it means that the data has been written to the disk. If it is elected as the master, it will directly become the master partition; (2) If the latest message offset recorded by itself does not exist on the disk, but its value is equal to the lowest offset, it means that no message data has not been processed. If it is elected as the master, it will directly become the master partition; (3) If neither of the above two situations are met and it is elected as the master partition, it means that there will be message loss. If ack is not required to be equal to -1, it can become the master. If it is required to be equal to -1, it will not become the master and exit the election.
[0114] In order to execute the corresponding steps in the above embodiment and each possible implementation method, an implementation method of the message synchronization device 100 is provided below. Fig.11 , Fig.11 It is a block diagram of a message synchronization device 100 provided by an embodiment of the present invention. It should be noted that the basic principle and technical effects of the message synchronization device 100 provided by this embodiment are the same as those of the above embodiments, which are not mentioned in this embodiment for the sake of brief description.
[0115] The message synchronization device 100 includes an acquisition module 110 , a synchronization module 120 and a replacement module 130 .
[0116] The acquisition module 110 is used to acquire the first number of the consumption message consumed most recently, wherein the first number is used to represent the offset position of the consumption message in the primary partition.
[0117] The acquisition module 110 is further used to acquire the second number of the synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to represent the offset position of the synchronization message in the primary partition.
[0118] The synchronization module 120 is used to synchronize the message between the third number and the second number of the most recently written message in the main partition to the first storage area if the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets the preset condition.
[0119] As a specific implementation, the synchronization module 120 is also used for: if the difference between the second number and the first number is greater than or equal to a preset value, or the free area in the first storage area does not meet the preset conditions, then determining the number of items to be migrated based on the second number, the third number and the preset value; migrating the message of the earliest number of migration items stored in the first storage area to the second storage area; synchronizing the target message of the number of migration items starting from the third number to the first storage area; synchronizing the messages between the third number and the second number except the target message to the second storage area.
[0120] As a specific implementation, when the synchronization module 120 is used to determine the number of items to be migrated based on the second number, the third number and the preset value, it is specifically used to: calculate the number difference between the second number and the third number; and take the minimum value between the number difference and the preset value as the number of items to be migrated.
[0121] As a specific implementation, the synchronization module 120 is further configured to: delete messages in the backup partition whose numbers are smaller than the first number.
[0122] As a specific implementation, the replacement module 130 is used to: when an abnormality is detected in the second node, determine whether the backup partition meets the replacement condition; if the backup partition meets the replacement condition, use the backup partition as the new primary partition according to the election mechanism.
[0123] As a specific implementation, the replacement module 130 is specifically used for: if there is a message with a second number in the second storage area, it is determined that the backup partition meets the replacement condition; if there is no message with the second number in the second storage area, and the second number is equal to the first number, it is determined that the backup partition meets the replacement condition; if there is no message with the second number in the second storage area, and the second number is greater than the first number, and the loss of some messages is allowed, it is determined that the backup partition meets the replacement condition.
[0124] As a specific implementation, when the replacement module 130 is specifically used to use the backup partition as the new primary partition according to the election mechanism, it is also specifically used to: write the message in the first storage area to the second storage area; and prohibit writing the message to the first storage area.
[0125] The present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a controller, the message synchronization method as described above is implemented.
[0126] In summary, an embodiment of the present invention provides a message synchronization method, device, node and readable storage medium, which are applied to a first node in a Kafka cluster, wherein the Kafka cluster also includes a second node communicating with the first node, wherein the second node includes a primary partition, wherein there are messages in the primary partition in sequentially increasing order, and the first node includes a backup partition, wherein the backup partition is used to store copies of messages in the primary partition, and the backup partition includes a first storage area and a second storage area, wherein the access performance of the first storage area is greater than that of the second storage area, and the method includes: obtaining a first number of a consumption message consumed most recently, wherein the first number is used to characterize an offset position of the consumption message in the primary partition; obtaining a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to characterize an offset position of the synchronization message in the primary partition; if the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets a preset condition, then synchronizing the message between the third number and the second number of the write message most recently written in the primary partition to the first storage area. Compared with the prior art, the embodiment of the present invention prioritizes synchronizing messages from the primary partition to the first storage area. Since messages are prioritized to be synchronized to the first storage area with high access performance, data synchronization response time can be reduced and message synchronization efficiency can be improved.
[0127] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A message synchronization method, characterized in that: The method is applied to a first node in a Kafka cluster, the Kafka cluster further comprising a second node communicating with the first node, the second node comprising a primary partition, the primary partition containing messages in sequentially increasing order, the first node comprising a backup partition, the backup partition being used to store copies of messages in the primary partition, the backup partition comprising a first storage area and a second storage area, the access performance of the first storage area being greater than that of the second storage area, the method comprising: Obtaining a first number of a consumption message consumed most recently, wherein the first number is used to represent an offset position of the consumption message in the primary partition; Acquire a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to represent an offset position of the synchronization message in the primary partition; If the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets a preset condition, the message between the third number of the most recently written message in the main partition and the second number is synchronized to the first storage area.
2. The message synchronization method according to claim 1, characterized in that: The method further comprises: If the difference between the second number and the first number is greater than or equal to the preset value, or the free area in the first storage area does not meet the preset condition, determine the number of entries to be migrated according to the second number, the third number and the preset value; Migrate the earliest number of messages stored in the first storage area to the second storage area; Synchronize the target messages of the migration number starting from the third number to the first storage area; The messages between the third number and the second number except the target message are synchronized to the second storage area.
3. The message synchronization method according to claim 2, characterized in that: The step of determining the number of entries to be migrated according to the second number, the third number and the preset value comprises: Calculate the difference between the second number and the third number; The minimum value between the serial number difference and the preset value is used as the number of entries to be migrated.
4. The message synchronization method according to any one of claims 1 to 3, characterized in that: The method further comprises: Delete the messages in the backup partition whose numbers are smaller than the first number.
5. The message synchronization method according to claim 1, characterized in that: The method further comprises: When an abnormality is detected in the second node, determining whether the backup partition meets a replacement condition; If the backup partition meets the replacement condition, the backup partition is used as the new primary partition according to the election mechanism.
6. The message synchronization method according to claim 5, characterized in that: The step of judging whether the backup partition meets the replacement condition comprises: If the second storage area contains the message with the second number, determining that the backup partition satisfies the replacement condition; If the second storage area does not contain the message with the second number, and the second number is equal to the first number, it is determined that the backup partition meets the replacement condition; If the second storage area does not contain the message with the second number, and the second number is greater than the first number, and the loss of some messages is allowed, it is determined that the backup partition meets the replacement condition.
7. The message synchronization method according to claim 5, characterized in that: The step of using the backup partition as a new primary partition according to the election mechanism comprises: Writing the message in the first storage area into the second storage area; Writing messages to the first storage area is prohibited.
8. A message synchronization device, characterized in that: The device is applied to a first node in a Kafka cluster, the Kafka cluster further comprising a second node communicating with the first node, the second node comprising a primary partition, the primary partition containing messages in ascending order, the first node comprising a backup partition, the backup partition being used to store copies of messages in the primary partition, the backup partition comprising a first storage area and a second storage area, the access performance of the first storage area being greater than that of the second storage area, the device comprising: An acquisition module, used to acquire a first number of a consumption message consumed most recently, wherein the first number is used to represent an offset position of the consumption message in the primary partition; The acquisition module is further used to acquire a second number of a synchronization message that is pre-stored locally and most recently synchronized from the primary partition, wherein the second number is used to represent an offset position of the synchronization message in the primary partition; A synchronization module is used to synchronize the message between the third number of the most recently written message in the main partition and the second number to the first storage area if the difference between the second number and the first number is less than a preset value and the free area in the first storage area meets a preset condition.
9. A node, comprising a memory and a controller, characterized in that: When the controller executes the computer program, the message synchronization method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a controller, the message synchronization method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Main memory database cluster synchronization method and main memory database host
CN104965862A
Data synchronization method, apparatus, device and storage medium thereof between clusters
CN109388677A