Message processing method, device, electronic device and readable storage medium

By using message offsets to determine the message consumption status in the Kafka message system, the problem of repeated message consumption caused by system downtime is solved, and efficient message processing and system stability are achieved.

CN113779149BActive Publication Date: 2025-05-13BEIJING KNOWNSEC INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111072091.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-05-13
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

In Kafka message system, consumers cannot update information in time due to external factors such as system downtime or restart, and cannot know the consumption status of the message in a timely and correctly, resulting in repeated consumption of messages.

Method used

By utilizing the offset of the message, whether the message is a consumed message, if not, the offset is updated and the message is processed to avoid repeated consumption.

Benefits of technology

It effectively avoids repeated consumption of messages, improves the stability and efficiency of the system, and is easy to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113779149B_ABST
    Figure CN113779149B_ABST
Patent Text Reader

Abstract

The present application proposes a message processing method, device, electronic device and readable storage medium, which relate to the field of computer technology. The method includes: obtaining a first message to be processed from a message queue; judging whether the first message to be processed is a consumed message according to the first current offset and the first target offset in the first message to be processed, and obtaining a first judgment result, wherein the first target offset is the offset of the first message to be processed that was last consumed by the consumer; if yes, discarding the first message to be processed; if no, updating the first target offset according to the first current offset, and processing and saving the first message to be processed. In this way, the offset can be used to determine whether the message is a consumed message, and the message is consumed if it is not, thereby avoiding repeated consumption, and the method is easy to operate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a message processing method, device, electronic device and readable storage medium. Background Art

[0002] The concept of message queue was created to reduce request response time and decoupling, and Kafka is a high-concurrency, low-latency publishing and subscription messaging system with the advantages of high performance, persistence, multi-copy backup, and horizontal scalability.

[0003] The consumption status of Kafka's message data is maintained by the consumer itself, which can remove the pressure on the server to maintain the consumption status and greatly improve the consumer's freedom to store the consumption status. It can even roll back and re-consume when consumption fails.

[0004] However, if consumers fail to update information in time due to external factors such as system downtime or restart, consumers will not be able to know the consumption status of messages in time and correctly, which will lead to wrong judgments. After completing the consumption loop of a message, the message will be consumed again, which means that a message is consumed repeatedly. Repeated consumption will cause various problems, such as increasing the burden on the system. Therefore, how to avoid repeated consumption of messages has become a technical problem that technicians in this field need to solve urgently. Summary of the invention

[0005] The embodiments of the present application provide a message processing method, device, electronic device and readable storage medium, which can use the offset to determine whether a message is a consumed message, and consume the message if it is not, thereby avoiding repeated consumption. At the same time, this method is easy to operate.

[0006] The embodiment of the present application can be implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a message processing method, including:

[0008] Get the first message to be processed from the message queue;

[0009] According to the first current offset and the first target offset in the first message to be processed, determine whether the first message to be processed is a consumed message, and obtain a first judgment result, wherein the first target offset is the offset of the first message to be processed that was last consumed by the consumer;

[0010] If the first judgment result is that the first message to be processed is a consumed message, discarding the first message to be processed;

[0011] When the first judgment result is that the first to-be-processed message is not a consumed message, the first target offset is updated according to the first current offset, and the first to-be-processed message is processed and saved.

[0012] In a second aspect, an embodiment of the present application provides a message processing device, including:

[0013] A message obtaining module, used for obtaining a first message to be processed from a message queue;

[0014] A judgment module, configured to judge whether the first message to be processed is a consumed message according to a first current offset and a first target offset in the first message to be processed, and obtain a first judgment result, wherein the first target offset is an offset of the first message to be processed that was last consumed by the consumer;

[0015] a processing module, configured to discard the first message to be processed if the first judgment result is that the first message to be processed is a consumed message;

[0016] The processing module is further configured to update the first target offset according to the first current offset, and process and save the first message to be processed when the first judgment result is that the first message to be processed is not a consumed message.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the message processing method described in the aforementioned embodiment.

[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the message processing method as described in the foregoing implementation manner is implemented.

[0019] The message processing method, device, electronic device and readable storage medium provided by the embodiments of the present application determine whether the first to-be-processed message is a consumed message based on the first current offset in the first to-be-processed message obtained in the message queue and the stored first target offset. If it is, the first to-be-processed message is discarded. If it is not, the first target offset is updated according to the first current offset, and the first to-be-processed message is processed and saved. Among them, the first target offset is the offset of the first to-be-processed message consumed by the consumer last time. In this way, the offset in the first to-be-processed message and the offset of the first to-be-processed message consumed last time can be used to determine whether the first to-be-processed message has been consumed, and then the message is consumed if it has not been consumed, thereby avoiding repeated consumption of messages, and the method is easy to operate. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 This is a schematic diagram of an existing method for solving repeated consumption;

[0022] Figure 2 A block diagram of an electronic device provided in an embodiment of the present application;

[0023] Figure 3 One of the flowcharts of the message processing method provided in the embodiment of the present application;

[0024] Figure 4 for Figure 3 A schematic flow chart of the sub-steps included in step S120;

[0025] Figure 5 for Figure 3 A schematic flow chart of the sub-steps included in step S140;

[0026] Figure 6 for Figure 5 A schematic flow chart of the sub-steps included in sub-step S141;

[0027] Figure 7 The second flowchart of the message processing method provided in the embodiment of the present application;

[0028] Figure 8 A block diagram of a message processing device provided in an embodiment of the present application.

[0029] Icon: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication unit; 200 - message processing device; 210 - message acquisition module; 220 - judgment module; 230 - processing module. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0031] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0032] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0033] In order to solve the technical problem of repeated message consumption, some developers have specified some solutions based on the specific repeated consumption problem, and try their best to ensure the effective consumption of messages through these solutions. Currently, only the following two solutions are used.

[0034] Solution 1: Add key values ​​to messages to solve duplicate consumption. The specific solution is: the consumer reads messages from the Kafka message queue; when reading a message, the consumer adds a unique key to the message, so that each message becomes a unique key-value pair, where the value in the key-value pair is the message; before consuming a message, the consumer first checks whether the key of the message is in the keys of the saved consumed messages, so as to determine whether the message has been consumed; if the key of the message is found in the keys of the saved consumed messages, it means that the message has been consumed, and the message is directly removed from the message queue; if the key of the message is not found in the keys of the saved consumed messages, the message is consumed normally, and then the unique key of the message is recorded.

[0035] Option 2: If Figure 1 As shown, idempotence is used to solve repeated consumption. Among them, the mathematical definition of idempotence is that if a function f(x) satisfies: f(f(x)) = f(x), then the function f(x) satisfies idempotence. The characteristic of an idempotent operation is that the impact of any multiple executions is the same as the impact of a single execution. Solution 2 is as follows: Step 1. The consumer reads the message from the Kafka message queue; Step 2. The consumed message is stored in the database, and according to the association of the record, a transfer table is established. The table will splice these associated primary keys (for example, primary key 1, primary key 2...primary key N) into the unique primary key of this information (that is, the joint primary key), and form a record with the message that really needs to be consumed and save it in the transfer table; Step 3. When the message is transferred from the message queue again, first check whether there is already a record in the transfer table. If there is already a record, update the data directly, otherwise return to step 2.

[0036] However, the above solutions are relatively complicated in actual implementation.

[0037] For example, the key-value pair technology used in Solution 1 is not easy to use due to its complexity and limitations. The complexity refers to the fact that Kafka processes a large amount of data. If a unique key is added to each message to be consumed at the beginning, adding a unique key will bring a huge workload and reduce work efficiency. The limitation is that adding a unique key before data processing needs to be based on business logic. If the business logic is too complex, this method cannot be implemented.

[0038] For another example, Solution 2 uses idempotency technology, but its limitations and business restrictions make it inconvenient to use. The limitation here refers to the fact that idempotent producers cannot achieve idempotency across multiple partitions and sessions. To achieve atomicity across multiple partitions, transactions need to be introduced, and even if the same producer crashes and restarts, it cannot guarantee that the message is processed only once, that is, the Exactly Once semantics of the message. Business restrictions mean that designers need to start with business logic, but not all businesses can be designed to be naturally idempotent. Often, some other methods or techniques are needed to assist in achieving idempotency, making the implementation process more cumbersome and prone to logical defects, resulting in excessively high maintenance costs in the later stages.

[0039] The defects existing in the above schemes are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in the embodiments of this application below for the above problems should be the contributions made by the inventor to this application during the application process.

[0040] Based on the above situation, the embodiments of the present application provide a message processing method, device, electronic device and readable storage medium, which can use the offset to determine whether the message is a consumed message, and consume the message if it is not, thereby avoiding repeated consumption. At the same time, this method is easy to operate.

[0041] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0042] Please refer to Figure 2 , Figure 2 A block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may be, but is not limited to, a computer, a server, etc. The server may be a separate server or a clustered server. The electronic device 100 may include a Kafka system.

[0043] Kafka is a high-throughput, low-latency distributed publish-subscribe messaging system. The messages stored in Kafka come from multiple processes called "producers". Messages can be assigned to different "partitions" under different "topics" under different "brokers" (instances). Within a partition, these messages are indexed and stored together with timestamps. Other processes called "consumers" can query messages from partitions. Kafka runs on a cluster consisting of one or more servers, and partitions can be distributed across cluster nodes.

[0044] Among them, Broker is a server in the Kafka cluster. There are one or more servers in the Kafka cluster, and the server is called Broker. Each Broker in the Kafka cluster has a unique number. Topic represents the subject of the message, which can be understood as the classification of the message. Kafka's data (i.e., messages) is stored in the Topic. Multiple Topics can be created on each Broker, and each message entering Kafka will be placed under a Topic. Partition represents the partition of the Topic. Each Topic can have multiple Partitions. The role of the partition is to do load, improve Kafka's throughput, and improve the efficiency of message processing; the data of the same Topic in different partitions is not repeated, and the Partition is expressed as folders one by one.

[0045] The electronic device 100 includes a memory 110, a processor 120, and a communication unit 130. The memory 110, the processor 120, and the communication unit 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0046] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0047] The processor 120 is used to read / write data or programs stored in the memory 110 and perform corresponding functions. For example, the memory 110 stores a message processing device 200, and the message processing device 200 includes at least one software function module that can be stored in the memory 110 in the form of software or firmware. The processor 120 executes various functional applications and data processing by running software programs and modules stored in the memory 110, such as the message processing device 200 in the embodiment of the present application, that is, implementing the message processing method in the embodiment of the present application.

[0048] The communication unit 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through a network, and to send and receive data through the network.

[0049] It should be understood that Figure 2 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may also include Figure 2 More or fewer components as shown, or with Figure 2 Different configurations shown. Figure 2 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0050] Please refer to Figure 3 , Figure 3 This is one of the flow diagrams of the message processing method provided in the embodiment of the present application. The method can be applied to the above-mentioned electronic device 100. The specific flow of the message processing method is described in detail below. The method may include steps S110 to S140.

[0051] Step S110, obtaining a first message to be processed from a message queue.

[0052] In this embodiment, messages from Kafka are continuously put into the message queue, that is, the messages in the message queue are all messages generated by the producer. The consumer can obtain a message from the consumer queue as the first message to be processed. The position of the message in the message queue can be represented by an offset, which is determined by the producer. For example, if the producer writes messages 1, 2, and 3 in sequence, the offsets of messages 1, 2, and 3 can be 0, 1, and 2 respectively. Each message can carry the offset of the message.

[0053] Step S120: judging whether the first message to be processed is a consumed message according to the first current offset and the first target offset in the first message to be processed, and obtaining a first judgment result.

[0054] When the first message to be processed is obtained, the offset in the first message to be processed can be obtained as the first current offset of the first message to be processed. Then, based on the first current offset and the saved first target offset, it is determined whether the first message to be processed is a consumed message, and a first judgment result is obtained. Among them, the first target offset is the offset of the first message to be processed that the consumer consumed last time, which can be used to indicate the position that the consumer has consumed in the message queue. Among them, the judgment method based on the first current offset and the first target offset can be specifically determined in combination with the change method of the offset of the messages consumed in sequence.

[0055] Step S130: When the first judgment result is that the first message to be processed is a consumed message, discard the first message to be processed.

[0056] The first judgment result is that the first pending message is a consumed message, which means that the first pending message has been consumed before. If the first pending message is consumed this time, the first pending message will be consumed repeatedly. To avoid the first pending message being consumed repeatedly, the consumer can discard the first pending message.

[0057] Step S140: When the first judgment result is that the first to-be-processed message is not a consumed message, the first target offset is updated according to the first current offset, and the first to-be-processed message is processed and saved.

[0058] The first judgment result is that the first pending message is not a consumed message, which means that the first pending message is currently determined to have not been consumed, that is, the consumption performed this time is the first consumption and is normal consumption. Therefore, when the first judgment result is that the first pending message is not a consumed message, the first pending message can be processed and saved to complete the consumption of the first pending message. In addition, the first target offset is also updated according to the first current offset so as to perform the above judgment on the next message obtained from the message queue. Among them, the specific updating method can be to assign the first current offset to the first target offset, that is, after the update, the first target offset and the first current offset are the same value.

[0059] In this way, the offset in the first message to be processed and the offset of the first message to be processed that was saved last time consumed can be used to determine whether the first message to be processed has been consumed, and then the message can be consumed if it has not been consumed, thereby avoiding repeated consumption of the message. At the same time, this method is easy to operate.

[0060] Optionally, as a possible implementation, the offsets of the messages consumed sequentially increase in sequence. For example, the offset of the message consumed first is 1, the offset of the message consumed next is 2, and the offset of the message consumed next is 3. In this case, Figure 4 The first judgment result is obtained in the manner shown. Figure 4 , Figure 4 for Figure 3 Schematic diagram of the flow of sub-steps included in step S120. In this embodiment, step S120 may include sub-steps S121 to S123.

[0061] Sub-step S121: comparing the first current offset with the first target offset.

[0062] Sub-step S122: When the first current offset is not greater than the first target offset, determining that the first to-be-processed message is a consumed message.

[0063] Sub-step S123: When the first current offset is greater than the first target offset, determining that the first to-be-processed message is not a consumed message.

[0064] In this embodiment, the first current offset may be compared with the first target offset to determine whether the first current offset is greater than the first target offset. If the first current offset is less than or equal to the first target offset, it means that the position originally consumed by the consumer is greater than or equal to the position of the current first message to be processed, and the current first message to be processed is a message that has been consumed before. In this case, the first judgment result obtained is that the first message to be processed is a consumed message.

[0065] If the first current offset is greater than the first target offset, it means that the position originally consumed by the consumer is smaller than the current position of the first message to be processed, and this is normal consumption. Therefore, the first judgment result obtained in this case is that the first message to be processed is not a consumed message.

[0066] It is understandable that when Kafka is running for the first time, the first message may not be judged based on the first current offset as above, and the first message may be directly determined to be not a consumed message, and then consumed accordingly. In the case where the message has been consumed, the consumer may perform the above judgment after each message is obtained to determine whether to consume the obtained message.

[0067] If the first judgment result is that the first message to be processed is a consumed message, the first message to be processed may be discarded. Message consumption is ongoing, so after discarding the first message to be processed, the consumer may obtain another message from the message queue as a new first message to be processed, that is, after discarding the first message to be processed, the process starts again from step S110.

[0068] Optionally, in the case of discarding the first message to be processed, in order to prevent the consumer from obtaining the consumed message again, a specified offset may be determined according to the first target offset, and then the message at the position represented by the specified offset may be obtained from the message queue as a new first message to be processed. The specified offset may be the result of adding 1 to the first target offset.

[0069] When the first judgment result is that the first to-be-processed message is not a consumed message, the first target offset can be updated based on the first current offset to ensure that the first target offset corresponds to the actual situation and avoid duplicate consumption due to inaccurate first target offset.

[0070] If the first judgment result is that the first message to be processed is not a consumed message, the first message to be processed may be processed and saved to complete the consumption of the first message to be processed. Optionally, the specific processing process may be determined by actual needs.

[0071] Optionally, the consumer can process the first message to be processed and save the processed first message to be processed in a database. The first message to be processed can be processed and stored in a database in a synchronous manner, which can reduce the probability of repeated consumption of messages.

[0072] Optionally, in a possible implementation manner, Figure 5 The first pending message is processed and saved in the manner shown. Figure 5 , Figure 5 for Figure 3 Schematic diagram of the flow of sub-steps included in step S140. In this embodiment, step S140 may include sub-steps S141 to S145.

[0073] Sub-step S141, performing preset processing on the first message to be processed, and storing the processed message in a first buffer.

[0074] In this embodiment, the first message to be processed can be saved in the first buffer, and then the first message to be processed is processed in the first buffer, and the processed first message to be processed is saved in the first buffer. The specific processing operation of the preset processing can be determined by the actual situation and is not specifically limited here.

[0075] Among them, if the first pending message after processing is saved in the first buffer, then judging whether the first pending message is a consumed message in step S120 can be understood as judging whether the first pending message has been read from the Kafka message queue and stored in the first buffer.

[0076] Sub-step S142, obtaining a second message to be processed from the first buffer.

[0077] The first buffer includes a first message to be processed that has undergone a preset process, and a first message to be processed that has undergone a preset process can be obtained from the first buffer as a second message to be processed. Saving the message to the first buffer and obtaining the second message to be processed from the first buffer can be synchronous or asynchronous, which can be determined according to actual needs.

[0078] If it is asynchronous, it means, for example, that one program continuously obtains the first message to be processed from the message queue and brings the first message to be processed to the first buffer; another program continuously obtains the second message to be processed from the first buffer and brings it to the second buffer.

[0079] Optionally, in a possible implementation, saving the message to the first buffer and obtaining the second message to be processed from the first buffer are asynchronous, so as to improve processing efficiency through asynchrony.

[0080] Sub-step S143, judging whether the second message to be processed is a consumed message according to the second current offset and the second target offset in the second message to be processed, and obtaining a second judgment result.

[0081] In this embodiment, the offset carried in the second message to be processed is the second current offset of the second message to be processed. The second target offset is the offset of the message last stored in the second buffer. It can be determined whether the second message to be processed has been stored in the second buffer based on the second current offset and the second target offset. That is, based on the position of the message last stored in the second buffer and the position of the second message to be processed, it is determined whether the second message to be processed has been placed in the second buffer before. This judgment can avoid repeated storage. Among them, the specific judgment method based on the second current offset can be determined in combination with the change method of the offset of the messages consumed sequentially.

[0082] Optionally, the offsets of the messages consumed sequentially are in a sequentially increasing state. In this case, Figure 6 The second judgment result is obtained in the manner shown. Figure 6 , Figure 6 for Figure 5 The schematic flow chart of the sub-steps included in sub-step S143. Sub-step S143 may include sub-steps S1431 to S1433.

[0083] Sub-step S1431: compare the second current offset with the second target offset.

[0084] Sub-step S1432: When the second current offset is not greater than the second target offset, determine that the second to-be-processed message is a consumed message.

[0085] Sub-step S1433: When the second current offset is greater than the second target offset, determine that the second to-be-processed message is not a consumed message.

[0086] In this embodiment, the second current offset may be compared with the second target offset to determine whether the second current offset is greater than the second target offset, that is, to determine whether the following formula holds: meg.offset>lastoffset, where meg.offset represents the second current offset and lastoffset represents the second target offset.

[0087] If the second current offset is less than or equal to the second target offset, it means that the maximum position of the message originally saved in the second buffer is greater than or equal to the current position of the second message to be processed, and the second message to be processed has been consumed, that is, it has been read from the first buffer and stored in the second buffer. Therefore, the second judgment result obtained in this case is that the second message to be processed is a consumed message.

[0088] If the second current offset is greater than the second target offset, it means that the maximum position of the message originally stored in the second buffer is smaller than the current position of the second message to be processed, and this is normal consumption. Therefore, the second judgment result obtained in this case is that the second message to be processed is not a consumed message.

[0089] It is understandable that, when Kafka is running for the first time, the first message may not be judged based on the second current offset as above, and the first message may be directly determined to be not a consumed message, and then consumed accordingly. In the case where the message has been consumed, after each message is obtained, the above judgment based on the second current offset may be performed to determine whether to save the message to the second buffer.

[0090] Sub-step S144: when the second judgment result is that the second message to be processed is a consumed message, discard the second message to be processed.

[0091] If the second judgment result is that the second pending message is a consumed message, and the second pending message is still saved in the second buffer, duplicate consumption will occur. Therefore, when the second judgment result is that the second pending message is a consumed message, the second pending message can be directly discarded. Message consumption is ongoing, so after discarding the second pending message, while asynchronously storing the messages in the message queue into the first buffer and the messages in the first buffer into the second buffer, the consumer can perform the above-mentioned judgment based on the second current offset for the new second pending message.

[0092] Sub-step S145, when the second judgment result is that the second to-be-processed message is not a consumed message, the second target offset is updated according to the second current offset, and the second to-be-processed message is stored in the second buffer, and the second to-be-processed message stored in the second buffer is saved in the database.

[0093] Optionally, when the second judgment result is that the second to-be-processed message is not a consumed message, the second target offset can be updated based on the second current offset to ensure that the second target offset corresponds to the message position actually saved in the second buffer, thereby avoiding duplicate consumption due to inaccurate second target offset.

[0094] If the second judgment result is that the second pending message is not a consumed message, the second pending message can also be saved in the second buffer. Message consumption is carried out uninterruptedly, so after the second pending message is saved in the second buffer, when asynchronously storing the messages in the message queue in the first buffer and the messages in the first buffer in the second buffer, the consumer can perform the above judgment based on the second current offset for the new second pending message. After the second pending message is saved in the second buffer, the second pending message stored in the second buffer can also be saved in the database.

[0095] Optionally, as an optional method, after the second message to be processed is saved in the second buffer, the second message to be processed can be directly processed to obtain the processing result of the second message to be processed, and the processing result can be written into the database. The specific strategy of the data processing can be determined according to actual needs. In this way, the processing and storage of a second message to be processed can be completed in a synchronous manner.

[0096] In this manner, the offset of the processing result can also be saved as the third target offset. During initialization, the initial value of the second target offset is the third target offset.

[0097] Optionally, as another optional method, after the second message to be processed is saved in the second buffer, the second message to be processed can be directly processed in the second buffer to obtain the processing result of the second message to be processed. Afterwards, it is determined whether the processing result currently obtained but not stored in the database meets the warehousing conditions. Among them, the warehousing conditions can be set in combination with actual needs. For example, the warehousing conditions include a preset data volume. If the data volume of the processing result currently obtained but not stored in the database is greater than or equal to the preset data volume, it can be determined that the current processing result meets the warehousing conditions; conversely, if the data volume of the processing result currently obtained but not stored in the database is less than the preset data volume, it can be determined that the current processing result does not meet the warehousing conditions. It can be understood that other rules can also be set in the warehousing conditions, which can be set specifically according to the situation.

[0098] If the storage conditions are not met, the process continues to wait until the processing results not stored in the database meet the storage conditions.

[0099] If the storage conditions are met, the currently obtained processing results that are not stored in the database can be stored in batches in the database. In this way, the asynchronous batch writing method can reduce the number of interactions with the hard disk where the database is located and improve storage efficiency.

[0100] In the case of batch writing, the maximum offset in the processing result of the database of this batch writing can also be saved as the third target offset. Wherein, during initialization, the initial value of the second target offset is the third target offset. For example, when the Kafka system is restarted, the second target offset can be initialized based on the third target offset.

[0101] Please refer to Figure 7 , Figure 7 The second flowchart of the message processing method provided by the embodiment of the present application. After step S140, the method may further include at least one of step S150 and step S160-step S170.

[0102] Step S150: Save node information of at least one processing node into a log document.

[0103] The processing node represents the processing stage. In this embodiment, in the above steps S110 to S140, relevant information of different stages can be obtained, and the relevant information is used as the node information of the processing stage, and then the node information is saved in a log document. Among them, the log document can be located in a hard disk. Optionally, in order to ensure that the log document is consistent with the actual situation, the log document can be updated accordingly when the relevant information changes.

[0104] The specific content of the node information of each processing node can be determined according to actual needs, as long as fault recovery and / or abnormality inspection can be performed later according to the log document. Optionally, the node information of at least one processing node includes at least one of the first target offset, the second target offset and the third target offset.

[0105] In one possible implementation, the at least one processing node includes: reading messages from a message queue and storing them in a first buffer, reading messages from a first buffer and storing them in a second buffer, storing processing results in a database, etc. Correspondingly, the node information of the at least one processing node includes the first target offset, the second target offset, and the third target offset, etc.

[0106] Step S160: in case of a failure, perform failure recovery processing according to the log document.

[0107] Optionally, when a system failure occurs, such as a program being killed, a runtime error, a restart, an uncommitted offset that has been consumed, a processing result not being stored, or a processing result storage interruption, the data before the failure can be found based on the log document, and then the failure recovery process can be performed based on the data. In this way, the log document can be used to maintain the data integrity of the entire consumption process as much as possible.

[0108] Step S170, periodically checking the log document and handling the exception when an exception is found.

[0109] In this embodiment, it is also possible to regularly detect whether the functions of certain key points are correct and real-time based on the log document, so as to find abnormalities early and handle the abnormalities.

[0110] For example, the offset of each Partition in the Topic can be obtained regularly and compared with the first target offset saved by the consumer. If the offset of the message in the Partition is greater than the first target offset recorded by the consumer, it is normal consumption; otherwise, it means that the message corresponding to the offset of the current Partition has been consumed, and the message can be discarded directly. Alternatively, it is possible to regularly check whether the locally saved data file is empty. If it is not empty, it will be imported into the corresponding database and data table. In particular, when a failure occurs, the processing results that are not stored in the database can be stored locally as the above-mentioned data files.

[0111] Optionally, in order to ensure efficient information consumption and effective data storage, expired logs can also be destroyed regularly to reduce storage pressure.

[0112] In this embodiment, the consumer can detect the location of the message each time after consuming data. If the number of topics subscribed by the consumer changes, or the number of partitions subscribed to the topic changes, or the number of consumers changes, a rebalance will be triggered. During rebalance, Kafka will reallocate consumption tasks based on changes in the number of consumers, the number of topics, and the number of partitions. After the rebalance is completed, the location of the message can be reset to start consumption based on the consumption position of each partition under each topic recorded in the log document (i.e., the first target offset).

[0113] The embodiment of the present application carefully analyzes the process of Kafka information consumption, and provides a detailed solution setting for each possible situation of repeated message consumption, which can solve the problems of repeated message consumption and data anomalies as much as possible. In addition, the embodiment of the present application is based on the traditional consumption processing mode, conforms to most of the design logic, has strong operability, can reduce the difficulty of implementation, can adapt to more business environments, and enhance user experience. At the same time, the detection of key point data processing in a cycle can allow developers to find problems in time, and the log document in the embodiment of the present application includes the node information of at least one processing node, which can ensure the security of data. And by regularly deleting expired log files, it is guaranteed that resource utilization is maximized.

[0114] In order to execute the corresponding steps in the above embodiments and various possible methods, a method for implementing a message processing device 200 is provided below. Optionally, the message processing device 200 may adopt the above Figure 2 The device structure of the electronic device 100 is shown. Figure 8 , Figure 8The block diagram of the message processing device 200 provided in the embodiment of the present application. It should be noted that the basic principle and technical effect of the message processing device 200 provided in this embodiment are the same as those of the above embodiment. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding contents in the above embodiment. The message processing device 200 may include: a message obtaining module 210, a judgment module 220 and a processing module 230.

[0115] The message obtaining module 210 is used to obtain a first message to be processed from a message queue.

[0116] The judgment module 220 is used to judge whether the first to-be-processed message is a consumed message according to the first current offset and the first target offset in the first to-be-processed message, and obtain a first judgment result, wherein the first target offset is the offset of the first to-be-processed message consumed by the consumer last time.

[0117] The processing module 230 is configured to discard the first message to be processed if the first judgment result is that the first message to be processed is a consumed message.

[0118] The processing module 230 is further configured to update the first target offset according to the first current offset, and process and save the first message to be processed when the first judgment result is that the first message to be processed is not a consumed message.

[0119] Optionally, in this embodiment, the offsets of messages consumed sequentially increase sequentially, and the judgment module 220 is specifically used to: compare the first current offset with the first target offset; when the first current offset is not greater than the first target offset, determine that the first to-be-processed message is a consumed message; when the first current offset is greater than the first target offset, determine that the first to-be-processed message is not a consumed message.

[0120] Optionally, in this embodiment, the processing module 230 is specifically used to: perform preset processing on the first message to be processed, and save the processed message in a first buffer; obtain a second message to be processed from the first buffer; determine whether the second message to be processed is a consumed message based on a second current offset and a second target offset in the second message to be processed, and obtain a second judgment result, wherein the second target offset is the offset of the message last stored in the second buffer; if the second judgment result is that the second message to be processed is a consumed message, discard the second message to be processed; if the second judgment result is that the second message to be processed is not a consumed message, update the second target offset according to the second current offset, store the second message to be processed in the second buffer, and save the second message to be processed stored in the second buffer to a database.

[0121] Optionally, in this embodiment, the offsets of the messages consumed sequentially increase sequentially, and the processing module 230 is specifically used to: compare the second current offset with the second target offset; when the second current offset is not greater than the second target offset, determine that the second message to be processed is a consumed message; when the second current offset is greater than the second target offset, determine that the second message to be processed is not a consumed message.

[0122] Optionally, in this embodiment, the processing module 230 is specifically used to: perform data processing on the second message to be processed to obtain processing results; and store the obtained processing results into the database in batches.

[0123] Optionally, in this embodiment, after storing the obtained processing results in batches into the database, the processing module 230 is further used to: save the maximum offset among the processing results stored in batches into the database as the third target offset. During initialization, the initial value of the second target offset is the third target offset.

[0124] Optionally, in this embodiment, the processing module 230 is also used to: save node information of at least one processing node into a log document, wherein the node information of the at least one processing node includes at least one of the first target offset, the second target offset and the third target offset; in the event of a fault, perform fault recovery processing according to the log document; and / or periodically check according to the log document and process the exception when an exception is detected.

[0125] Optionally, the above modules can be stored in the form of software or firmware. Figure 2The memory 110 shown in the figure may be fixed in the operating system (OS) of the electronic device 100 and may be Figure 1 Meanwhile, the data and program codes required for executing the above modules may be stored in the memory 110.

[0126] An embodiment of the present application also provides a readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the message processing method is implemented.

[0127] In summary, the embodiments of the present application provide a message processing method, device, electronic device and readable storage medium, which determines whether the first to-be-processed message is a consumed message based on the first current offset in the first to-be-processed message obtained in the message queue and the stored first target offset. If it is, the first to-be-processed message is discarded. If it is not, the first target offset is updated according to the first current offset, and the first to-be-processed message is processed and saved. Among them, the first target offset is the offset of the first to-be-processed message consumed by the consumer last time. In this way, the offset in the first to-be-processed message and the offset of the first to-be-processed message consumed last time can be used to determine whether the first to-be-processed message has been consumed, and then the message is consumed if it has not been consumed, thereby avoiding repeated consumption of messages, and the method is easy to operate.

[0128] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0129] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0130] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0131] The above description is only an optional embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A message processing method, characterized in that: include: Get the first message to be processed from the message queue; According to the first current offset and the first target offset in the first message to be processed, determine whether the first message to be processed is a consumed message, and obtain a first judgment result, wherein the first target offset is the offset of the first message to be processed that was last consumed by the consumer; If the first judgment result is that the first message to be processed is a consumed message, discarding the first message to be processed; When the first judgment result is that the first to-be-processed message is not a consumed message, the first target offset is updated according to the first current offset, and the first to-be-processed message is processed and saved, wherein the updated first target offset is the first current offset; The processing of the first message to be processed includes: Performing preset processing on the first message to be processed, and saving the processed message in a first buffer, wherein the first judgment result is used to indicate whether the first message to be processed has been performed preset processing and saved in the first buffer; Obtain a second message to be processed from the first buffer; According to the second current offset and the second target offset in the second message to be processed, determine whether the second message to be processed is a consumed message, and obtain a second determination result, wherein the second target offset is the offset of the message last stored in the second buffer; If the second judgment result is that the second message to be processed is a consumed message, discarding the second message to be processed; When the second judgment result is that the second pending message is not a consumed message, the second target offset is updated according to the second current offset, and the second pending message is stored in the second buffer, and the processing result of the second pending message stored in the second buffer is saved in the database, wherein the saving of the message in the first buffer and the obtaining of the second pending message from the first buffer are processed asynchronously.

2. The method according to claim 1, characterized in that The offsets of the messages consumed in sequence increase in sequence. According to the first current offset and the first target offset in the first message to be processed, determining whether the first message to be processed is a consumed message, and obtaining a first determination result, includes: comparing the first current offset with the first target offset; When the first current offset is not greater than the first target offset, determining that the first to-be-processed message is a consumed message; When the first current offset is greater than the first target offset, it is determined that the first to-be-processed message is not a consumed message.

3. The method according to claim 1, characterized in that The offsets of the messages consumed in sequence increase in sequence, and judging whether the second message to be processed is a consumed message according to the second current offset and the second target offset in the second message to be processed, and obtaining a second judgment result, includes: comparing the second current offset with the second target offset; When the second current offset is not greater than the second target offset, determining that the second to-be-processed message is a consumed message; When the second current offset is greater than the second target offset, it is determined that the second to-be-processed message is not a consumed message.

4. The method according to claim 1, characterized in that: The step of saving the second to-be-processed message stored in the second buffer into the database includes: Performing data processing on the second message to be processed to obtain a processing result; The obtained processing results are stored in the database in batches.

5. The method according to claim 4, characterized in that After storing the obtained processing results in the database in batches, the method further includes: The maximum offset in the processing results stored in the database in batches is saved as the third target offset, wherein, during initialization, the initial value of the second target offset is the third target offset.

6. The method according to claim 5, characterized in that The method further comprises: Saving node information of at least one processing node in a log document, wherein the node information of the at least one processing node includes at least one of the first target offset, the second target offset, and the third target offset; In the event of a failure, performing a failure recovery process according to the log document; and / or periodically performing a check according to the log document and processing the exception when an exception is detected.

7. A message processing device, characterized in that: include: A message obtaining module, used for obtaining a first message to be processed from a message queue; A judgment module, configured to judge whether the first message to be processed is a consumed message according to a first current offset and a first target offset in the first message to be processed, and obtain a first judgment result, wherein the first target offset is an offset of the first message to be processed that was last consumed by the consumer; a processing module, configured to discard the first message to be processed if the first judgment result is that the first message to be processed is a consumed message; The processing module is further configured to update the first target offset according to the first current offset, and process and save the first message to be processed, when the first judgment result is that the first message to be processed is not a consumed message, wherein the updated first target offset is the first current offset; The processing module is specifically used to: perform preset processing on the first message to be processed, and save the processed message in a first buffer, wherein the first judgment result is used to indicate whether the first message to be processed has been preset processed and saved in the first buffer; obtain a second message to be processed from the first buffer; determine whether the second message to be processed is a consumed message according to a second current offset and a second target offset in the second message to be processed, and obtain a second judgment result, wherein the second target offset is the offset of the message last stored in the second buffer; if the second judgment result is that the second message to be processed is a consumed message, discard the first message to be processed; if the second judgment result is that the second message to be processed is not a consumed message, update the second target offset according to the second current offset, and store the second message to be processed in the second buffer, and save the processing result of the second message to be processed stored in the second buffer to a database, wherein saving the message in the first buffer and obtaining the second message to be processed from the first buffer are processed asynchronously.

8. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the message processing method described in any one of claims 1-6.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the message processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Message processing method and device

    CN112822091A

  • Message transmission method and device, equipment and storage medium

    CN112954007A