Data Processing Method, Device, Equipment and Storage Medium Based on Distributed System

By introducing intermediate modules and state cache mechanisms in distributed systems, the data consistency problem in the streaming system is solved, end-to-end data consistency is achieved, data duplication is avoided, and data accuracy and processing efficiency are improved.

CN115470205BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211161771.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-24
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

It is difficult for existing streaming systems to achieve end-to-end data consistency in the field of big data, resulting in data duplication when consuming data downstream, especially in the financial, banking and insurance industries that require strict data reliability.

Method used

By introducing intermediate modules in the distributed system, the pre-submitted message data of the production module is received and cached, and the preset transaction IDs and preset identifiers of each distributed node are obtained. When all preset identifiers are successfully received, the data is converted into target message data and the data offset state is stored, so as to achieve end-to-end data consistency between the production module and the intermediate module.

Benefits of technology

It achieves end-to-end data consistency, avoids the problem of data duplication in downstream consumption, improves data accuracy and processing efficiency, and meets the high requirements for data reliability in industries such as finance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470205B_ABST
    Figure CN115470205B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and discloses a data processing method, device, equipment and storage medium based on a distributed system. The method includes: receiving pre-submission message data sent by each first distributed node in the production module and caching the pre-submission message data in the intermediate module; obtaining the preset transaction IDs of each first distributed node and receiving in real time the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs; when the preset identifiers sent by all the first distributed nodes corresponding to the preset transaction IDs are successfully received, converting the pre-submission message data into target message data for consumption by the consumption module, and reading the data offset status of each first distributed node, and storing the data offset status of each first distributed node in the status cache module. By the above method, the present invention can achieve end-to-end data consistency and improve data accuracy and processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data processing method, device, equipment and storage medium based on a distributed system. Background Art

[0002] At present, there has always been a technical problem in the industry regarding stream system middleware in the field of big data, that is, the end-to-end data consistency problem. In a stream system, a piece of data can be accurately processed by the downstream system once. Currently, the reliability of stream systems in the field of big data is divided into three levels: at most once (At Most Once, data is likely to be lost, but there will be no repeated transmission), at least once (At Least Once, data cannot be lost, but there may be repeated transmission), and exactly once (Exactly Once, data cannot be lost and cannot be repeatedly transmitted). However, currently, the vast majority of stream systems in the industry can only meet the reliability of "at least once", that is, it will cause the phenomenon of duplicate data when the downstream consumes data.

[0003] The problems caused by duplicate data are very serious, especially in the financial, banking, and insurance industries, where the requirements for data reliability are very strict, requiring absolute consistency of data, and the reliability of "at least once" cannot meet the industry's needs. The existing method is generally to deduplicate the data according to the unique primary key of the data after the downstream consumes the data, and then perform further processing and display of the data; this method of data deduplication will result in low data processing efficiency, and at the same time, some specific data analysis systems in the downstream of the stream system do not support data deduplication, such as Druid. Summary of the Invention

[0004] The present invention provides a data processing method, device, equipment and storage medium based on a distributed system, which can achieve end-to-end data consistency and improve data accuracy and processing efficiency.

[0005] To solve the above technical problems, a technical solution adopted by the present invention is: to provide a data processing method based on a distributed system, wherein the distributed system includes a production module, an intermediate module connected to the production module, a consumption module connected to the intermediate module, and a status cache module respectively connected to the production module and the consumption module, and the data processing method includes:

[0006] Receiving pre-submission message data sent by each first distributed node in the production module and caching the pre-submission message data in the intermediate module;

[0007] Obtaining the preset transaction ID of each first distributed node and real-time receiving the preset identifier sent by the first distributed node corresponding to each preset transaction ID;

[0008] When all the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs are successfully received, convert the pre-submission message data into target message data for consumption by the consumption module, read the data offset statuses of the first distributed nodes, and store the data offset statuses of the first distributed nodes in the status cache module.

[0009] According to an embodiment of the present invention, before receiving the pre-submission message data sent by the first distributed nodes in the production module and caching the pre-submission message data in the intermediate module, it further includes:

[0010] Obtain the processing and sending events of the pre-submission message data by the first distributed nodes in the production module;

[0011] Configure the preset transaction IDs for the first distributed nodes based on the processing and sending events, and calculate the data offset statuses of the first distributed nodes through a preset calculation engine;

[0012] Cache the data offset statuses of the first distributed nodes in the memory of the production module.

[0013] According to an embodiment of the present invention, after obtaining the preset transaction IDs of the first distributed nodes and receiving in real time the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs, it further includes:

[0014] When all the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs are not successfully received, discard the pre-submission message data, re-read the previous data offset statuses of the first distributed nodes from the status cache module, and re-process the pre-submission message data based on the production module according to the previous data offset statuses of the first distributed nodes.

[0015] According to an embodiment of the present invention, after converting the pre-submission message data into target message data for consumption by the consumption module, reading the data offset statuses of the first distributed nodes, and storing the data offset statuses of the first distributed nodes in the status cache module when all the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs are successfully received, it further includes:

[0016] Send the target data to the consumption module and obtain in real time the consumption and successful processing events of the target data by the second distributed nodes in the consumption module;

[0017] Calculate the data offset status of each of the second distributed nodes through a pre-designed calculation engine based on the consumption and processing success events, and store the data offset status of each of the second distributed nodes in the memory of the consumption module.

[0018] According to an embodiment of the present invention, after sending the target data to the consumption module and real-time obtaining the consumption and processing success events of each second distributed node in the consumption module for the target data, it further includes:

[0019] When obtaining the consumption and processing success events of all the target data, read the data offset status of each of the second distributed nodes, and cache the data offset status of each of the second distributed nodes in the status cache module.

[0020] According to an embodiment of the present invention, after sending the target data to the consumption module and real-time obtaining the consumption and processing success events of each second distributed node in the consumption module for the target data, it further includes:

[0021] When not obtaining the consumption and processing success events of all the target data, abandon the consumption and processing of the target data, and re-read the previous data offset status of each of the second distributed nodes from the status cache module, and perform re-consumption and processing of the target data according to the previous data offset status of each of the second distributed nodes.

[0022] According to an embodiment of the present invention, the reading the data offset status of each of the first distributed nodes and storing the data offset status of each of the first distributed nodes in the status cache module includes:

[0023] Read the data offset status of each of the first distributed nodes, associate and bind each of the first distributed nodes with the corresponding data offset status to obtain a first association and binding result, and store the first association and binding result in the status cache module;

[0024] The reading the data offset status of each of the second distributed nodes and caching the data offset status of each of the second distributed nodes in the status cache module includes:

[0025] Read the data offset status of each of the second distributed nodes, associate and bind each of the second distributed nodes with the corresponding data offset status to obtain a second association and binding result, and store the second association and binding result in the status cache module.

[0026] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a data processing device, including:

[0027] A first receiving module, configured to receive pre-commit message data sent by each first distributed node in the production module and cache the pre-commit message data in the intermediate module;

[0028] A second receiving module, configured to obtain preset transaction IDs of each of the first distributed nodes and receive in real time preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs;

[0029] An execution module, configured to, when successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, convert the pre-commit message data into target message data for consumption by the consumption module, read the data offset statuses of each of the first distributed nodes, and store the data offset statuses of each of the first distributed nodes in the status cache module.

[0030] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the data processing method based on a distributed system as described above is implemented.

[0031] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data processing method based on a distributed system as described above is implemented.

[0032] The beneficial effects of the present invention are: by caching the pre-commit message data of the production module in the intermediate module, when successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, converting the pre-commit message data into target message data for consumption by the consumption module, reading the data offset statuses of each of the first distributed nodes, and storing the data offset statuses of each of the first distributed nodes in the status cache module, end-to-end data consistency between the production module and the intermediate module is achieved through a status cache mechanism and a pre-commit mechanism, and it is ensured that there is only one correct sending action during the process of the production module sending the same batch of data to the intermediate module through the preset transaction ID, avoiding the problem of duplicate data in downstream consumption. Compared with the traditional downstream data deduplication solution, data consistency control is performed from upstream production, improving data accuracy and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the architecture of a distributed system according to an embodiment of the present invention;

[0034] Figure 2It is a schematic flowchart of a data processing method of a chrominance block prediction mode acquisition method according to a first embodiment of the present invention based on a distributed system;

[0035] Figure 3 It is a schematic flowchart of a data processing method based on a distributed system according to a second embodiment of the present invention;

[0036] Figure 4 It is a schematic flowchart of a data processing method of a chrominance block prediction mode acquisition method according to a third embodiment of the present invention based on a distributed system;

[0037] Figure 5 It is a schematic structural diagram of a data processing device according to an embodiment of the present invention;

[0038] Figure 6 It is a schematic structural diagram of a computer device according to an embodiment of the present invention;

[0039] Figure 7 It is a schematic structural diagram of a computer storage medium according to an embodiment of the present invention. Detailed Embodiments

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.

[0041] The terms "first", "second", and "third" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0042] References to "embodiments" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0043] Figure 1 is a schematic diagram of the architecture of a distributed system according to an embodiment of the present invention. Please refer to Figure 1 , the distributed system 10 includes a production module 11, an intermediate module 12 connected to the production module 11, a consumption module 13 connected to the intermediate module 12, and a status cache module 14 respectively connected to the production module 11 and the consumption module 13. Among them, the production module 11 includes a plurality of distributed nodes. The production module 11 is used to produce, process, and send pre-submission message data to the intermediate module 12. When the production module 11 sends the pre-submission message data, it calculates the data offset status through a pre-designed computing engine and caches the data offset status in the memory of the production module 11. After the intermediate module 12 confirms that all the pre-submission message data in the same batch has been successfully pre-submitted, the production module 11 stores the data offset status in the status cache module 14; the intermediate module 12 includes a plurality of partitions. The intermediate module 12 is used to receive and pre-store the pre-submission message data. The pre-submission message data may carry the data offset status. After confirming that all the pre-submission message data in the same batch has been successfully pre-submitted, it converts the pre-submission message data into target data (i.e., the formal data available for the downstream consumption module); the consumption module 13 includes a plurality of distributed nodes. The consumption module 13 is used to obtain the target data of the intermediate module 12 and perform consumption and processing. When the target data is successfully consumed and processed, it calculates the data offset status through a pre-designed computing engine and caches the data offset status in the memory of the consumption module 13. After all the target data has been successfully consumed and processed, the consumption module 13 stores the data offset status in the status cache module 14; the status cache module 14 is used to store the data offset status. The status cache module 14 includes a storage system such as Zookeeper or HDFS that can store distributed locks and statuses and supports persistent storage. During the operation of the distributed system, even if the memory of the production module 11 or the consumption module 13 crashes, the previous stored data offset status can still be read through the status cache module 14. The previous data offset status can be the data offset status of the previous batch of message data or the initialized data offset status, and the initial value of the data offset status is 0. Further, in the status cache module 14, the storage method is (distributed node name = data offset status).

[0044] This embodiment is a distributed node that distinguishes the production module 11 from the consumption module 13. The distributed nodes (P1, P2, P3) of the production module 11 are defined as the first distributed nodes, and the distributed nodes (C1, C2, C3) of the consumption module 13 are defined as the second distributed nodes. Taking 3 as an example for the number of the first distributed nodes, the number of partitions, and the number of the second distributed nodes, please refer to Figure 1 . In other embodiments, there may be other deployment quantity relationships for the number of the first distributed nodes, the number of partitions, and the number of the second distributed nodes, which are not specifically limited herein. This embodiment uses Apache Spark as the pre-designed computing engine for the production module 11 and the consumption module 13. The intermediate module 12 can be a streaming system, and Apache Kafka is used as the streaming system medium.

[0045] Figure 2 It is a schematic flowchart of the data processing method based on a distributed system according to the first embodiment of the present invention. It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 2 the process sequence shown. As Figure 2 shown, the method includes the steps:

[0046] Step S201: Receive the pre-submission message data sent by each first distributed node in the production module and cache the pre-submission message data in the intermediate module.

[0047] In step S201, the production module produces a batch of message data, and evenly scatters and distributes this batch of message data to each first distributed node. Each first distributed node processes and pre-submits the message data to the intermediate module. The action of each first distributed node processing and pre-submitting the message data to the intermediate module is defined as a processing and sending event. Further, before step S201, obtain the processing and sending events of each first distributed node in the production module for the pre-submission message data; configure a preset transaction ID for each first distributed node based on the processing and sending events, and calculate the data offset status of each first distributed node through the pre-designed computing engine; cache the data offset status of each first distributed node in the memory of the production module. In this embodiment, different first distributed nodes have different transaction IDs for processing each batch of message data. The first distributed node can send the pre-submission message data to the intermediate module with the transaction ID and, after sending the pre-submission message data, send a preset identifier to the intermediate module with the transaction ID.

[0048] Step S202: Obtain the preset transaction IDs of each first distributed node and receive in real time the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs.

[0049] In step S202, the preset identifier can be an identifier for confirming the successful sending of the pre-submission message data. After the production module waits for all the first distributed nodes to send the pre-submission message data to the intermediate module, each first distributed node sends the preset identifier to the intermediate module again to confirm the successful sending of the pre-submission message data. However, due to an exception in the intermediate module, or due to network fluctuations, an exception in the first distributed node of the production module, etc., it may result in the unsuccessful reception of the preset identifiers from all the first distributed nodes. Therefore, after step S202, it is necessary to determine whether the preset identifiers sent by all the first distributed nodes corresponding to all the preset transaction IDs have been successfully received.

[0050] Step S203: When the preset identifiers sent by all the first distributed nodes corresponding to all the preset transaction IDs are successfully received, convert the pre-submission message data into target message data for consumption by the consumption module, and read the data offset status of each first distributed node, and store the data offset status of each first distributed node in the status cache module.

[0051] In step S203, when the preset identifiers sent by all the first distributed nodes corresponding to all the preset transaction IDs are successfully received, it is determined that the action of submitting the batch of message data is successful. Then, convert the pre-submission message data into target message data for consumption by the consumption module, and read the data offset status of each first distributed node, associate and bind each first distributed node with the corresponding data offset status to obtain a first association and binding result, and store the first association and binding result in the status cache module to record the processing intermediate state of each first distributed node. The first association and binding result can be: (distributed node name = data offset status). The data offset status in this embodiment can be understood as the processing quantity of the message data, and the data offset can take any integer such as 0, 1, 2, 3, etc. For example, the initial value of the data offset status of the first distributed node is 0. After one processing of the message data, the data offset status is updated to 10. Then, the next processing of the message data by this first distributed node starts from 11.

[0052] The data processing method based on a distributed system in the first embodiment of the present invention caches the pre-submission message data of the production module in the intermediate module. When the preset identifiers sent by all the first distributed nodes corresponding to the preset transaction IDs are successfully received, the pre-submission message data is converted into target message data for consumption by the consumption module, and the data offset statuses of each first distributed node are read and stored in the status cache module. Through the status cache mechanism and the pre-submission mechanism, end-to-end data consistency between the production module and the intermediate module is achieved, and the preset transaction ID is used to ensure that there is only one correct sending action when the production module sends the same batch of data to the intermediate module, avoiding the problem of duplicate data in downstream consumption. Compared with the traditional downstream data deduplication scheme, data consistency control is performed from upstream production, improving data accuracy and processing efficiency.

[0053] Figure 3 It is a flowchart of the data processing method based on a distributed system in the second embodiment of the present invention. It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 3 the process sequence shown. As Figure 3 shown, the method includes the steps:

[0054] Step S301: Receive the pre-submission message data sent by each first distributed node in the production module and cache the pre-submission message data in the intermediate module.

[0055] In this embodiment, Figure 3 step S301 in Figure 2 is similar to step S201 in

[0056] For the sake of simplicity, it will not be elaborated here.

[0057] In this embodiment, Figure 3 step S302 in Figure 2 is similar to step S202 in

[0058] For the sake of simplicity, it will not be elaborated here.

[0059] In this embodiment, Figure 3 step S303 in Figure 2is similar to step S203 in [reference], and for the sake of simplicity, it will not be elaborated here.

[0060] Step S304: When the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs are not successfully received, discard the pre-commit message data, and re-read the last data offset status of each first distributed node from the status cache module, and reprocess the pre-commit message data based on the production module according to the last data offset status of each first distributed node.

[0061] In step S304, due to an exception in the intermediate module, or due to network fluctuations, an exception in the first distributed node of the production module, etc., it may cause the preset identifiers sent by all the first distributed nodes not to be successfully received. When the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs are not successfully received, it is determined that the action of submitting this batch of message data fails, then discard this batch of pre-commit message data, including the message data with successful submission and the message data with failed submission, and will not store the data offset status of this processing. At the same time, re-read the last data offset status of each first distributed node from the status cache module, and reprocess this batch of message data.

[0062] Based on the data processing method of the distributed system in the second embodiment of the present invention, on the basis of the first embodiment, when the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs are not successfully received, determine the message data to be reprocessed by obtaining the last data offset status, ensure that the same batch of message data is only processed and processed once, and only successfully submitted once, avoid the phenomenon of duplicate data, and thus achieve end-to-end data consistency.

[0063] Figure 4 is a schematic flowchart of the data processing method of the distributed system in the second embodiment of the present invention. It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 4 the process sequence shown. As Figure 4 shown, the method includes the steps:

[0064] Step S401: Receive the pre-commit message data sent by each first distributed node in the production module and cache the pre-commit message data in the intermediate module.

[0065] In this embodiment, Figure 4 step S401 in [[reference]] and Figure 2 step S201 in [[reference]] are similar, and for the sake of simplicity, it will not be elaborated here.

[0066] Step S402: Obtain the preset transaction IDs of each first distributed node and receive in real time the preset identifiers sent by the first distributed nodes corresponding to each preset transaction ID.

[0067] In this embodiment, Figure 4 Step S402 in Figure 2 is similar to Step S202 in , and for the sake of simplicity, it will not be elaborated here.

[0068] Step S403: When successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, convert the pre-commit message data into target message data for consumption by the consumption module, and read the data offset statuses of the first distributed nodes, and store the data offset statuses of the first distributed nodes into the status cache module.

[0069] In this embodiment, Figure 4 Step S403 in Figure 2 is similar to Step S203 in , and for the sake of simplicity, it will not be elaborated here.

[0070] Step S404: Send the target data to the consumption module and obtain in real time the consumption and processing success events of the target data by each second distributed node in the consumption module.

[0071] In Step S404, the action of the consumption module obtaining the target data and the successful consumption and processing of the target data is defined as the consumption and processing success event.

[0072] Step S405: Calculate the data offset statuses of the second distributed nodes based on the consumption and processing success events through a preset calculation engine, and store the data offset statuses of the second distributed nodes into the memory of the consumption module.

[0073] In Step S405, when the consumption and processing success events occur, the data offset status changes. Calculate the data offset statuses of the second distributed nodes through a preset calculation engine, and store the data offset statuses of the second distributed nodes into the memory of the consumption module. Considering that the memory of the consumption module is prone to crashing, resulting in the problem of being unable to obtain the memory, it is necessary to cache the data offset status in a module that supports persistent storage. Therefore, after Step S405, it is necessary to determine when to cache the data offset status into the status cache module. Please refer to Step S406.

[0074] Step S406: When obtaining all the consumption and processing success events of the target data, read the data offset statuses of the second distributed nodes, and cache the data offset statuses of the second distributed nodes into the status cache module.

[0075] In step S406, when the consumption and processing success events of all target data are obtained, it is considered that all the message data of this batch has been consumed and processed successfully. Then, the data offset status of each second distributed node is read, and each second distributed node is associated and bound with the corresponding data offset status to obtain a second association and binding result. The second association and binding result is stored in the status cache module. The second association and binding result in this embodiment is similar to the above-mentioned first association and binding result. The difference is that the object of the second association and binding result is the data offset status corresponding to the second distributed node, while the object of the first association and binding result is the data offset status corresponding to the first distributed node.

[0076] Step S407: When the consumption and processing success events of all target data are not obtained, the consumption and processing of the target data are abandoned, and the last data offset status of each second distributed node is read again from the status cache module. The target data is re-consumed and processed according to the last data offset status of each second distributed node.

[0077] In step S407, when the consumption and processing success events of all target data are not obtained, it is considered that the consumption and processing of the message data of this batch have failed. Then, all the message data of this batch is abandoned, including the message data with successful consumption and processing and the message data with failed consumption and processing. And the data offset status of this processing will not be stored. At the same time, the consumption module reads the last data offset status of each second distributed node again from the status cache module, re-obtains the target data of this batch according to the last data offset status, and then performs re-consumption and processing.

[0078] Based on the data processing method of the first embodiment, the data processing method of the third embodiment of the present invention realizes the end-to-end data consistency between the consumption module and the intermediate module through the status cache mechanism and the pre-submission mechanism, improves the data accuracy and processing efficiency. Even if any system fails or crashes, it can be restored again from the status cache module, improving the data disaster tolerance ability.

[0079] Figure 5 It is a schematic structural diagram of the data processing device according to the embodiment of the present invention. As Figure 5 shown, the device 50 includes a first receiving module 51, a second receiving module 52, and an execution module 53.

[0080] The first receiving module 51 is configured to receive the pre-submitted message data sent by each first distributed node in the production module and cache the pre-submitted message data in the intermediate module;

[0081] The second receiving module 52 is used to obtain the preset transaction IDs of each first distributed node and receive in real time the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs.

[0082] The execution module 53 is used to, when successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, convert the pre-submission message data into target message data for consumption by the consumption module, and read the data offset statuses of each first distributed node and store the data offset statuses of each first distributed node in the status cache module.

[0083] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the computer device according to an embodiment of the present invention. As Figure 6 shown, the computer device 60 includes a processor 61 and a memory 62 coupled to the processor 61.

[0084] The memory 62 stores program instructions for implementing the data processing based on the distributed system described in any of the above embodiments.

[0085] The processor 61 is used to execute the program instructions stored in the memory 62 to process data.

[0086] Among them, the processor 61 may also be referred to as a CPU (Central Processing Unit). The processor 61 may be an integrated circuit chip with the ability to process signals. The processor 61 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0087] Refer to Figure 7 , Figure 7Schematic diagram of the structure of the computer storage medium according to an embodiment of the present invention. The computer storage medium according to the embodiment of the present invention stores a program file 71 that can implement all the above methods. Among them, the program file 71 can be stored in the above computer storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned computer storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0088] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0089] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0090] The above is only the embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A data processing method based on a distributed system, characterized in that, The distributed system includes a production module, an intermediate module connected to the production module, a consumption module connected to the intermediate module, and a status cache module respectively connected to the production module and the consumption module. The data processing method includes: Receiving pre-submission message data sent by each first distributed node in the production module and caching the pre-submission message data in the intermediate module; Obtaining the preset transaction IDs of each of the first distributed nodes and receiving in real time the preset identifiers sent by the first distributed nodes corresponding to each of the preset transaction IDs; When successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, converting the pre-submission message data into target message data for consumption by the consumption module, reading the data offset statuses of each of the first distributed nodes, and storing the data offset statuses of each of the first distributed nodes in the status cache module; After obtaining the preset transaction IDs of each of the first distributed nodes and receiving in real time the preset identifiers sent by the first distributed nodes corresponding to each of the preset transaction IDs, it further includes: When not successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, discarding the pre-submission message data, re-reading the previous data offset statuses of each of the first distributed nodes from the status cache module, and re-processing the pre-submission message data by the production module based on the previous data offset statuses of each of the first distributed nodes; After when successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, converting the pre-submission message data into target message data for consumption by the consumption module, reading the data offset statuses of each of the first distributed nodes, and storing the data offset statuses of each of the first distributed nodes in the status cache module, it further includes: Sending target data to the consumption module and obtaining in real time the consumption and processing success events of the target data by each second distributed node in the consumption module; Calculating the data offset statuses of each of the second distributed nodes through a preset computing engine based on the consumption and processing success events, and storing the data offset statuses of each of the second distributed nodes in the memory of the consumption module.

2. The data processing method according to claim 1, wherein Before receiving the pre-submission message data sent by each first distributed node in the production module and caching the pre-submission message data in the intermediate module, it further includes: Obtaining the processing and sending events of the pre-submission message data by each first distributed node in the production module; Configuring the preset transaction IDs for each of the first distributed nodes based on the processing and sending events, and calculating the data offset statuses of each of the first distributed nodes through a preset computing engine; Caching the data offset statuses of each of the first distributed nodes in the memory of the production module.

3. The data processing method according to claim 1, wherein After sending the target data to the consumption module and obtaining in real time the consumption and successful processing events of the target data by each second distributed node in the consumption module, the following steps are further included: When obtaining the consumption and successful processing events of all the target data, read the data offset statuses of each of the second distributed nodes, and cache the data offset statuses of each of the second distributed nodes in the status cache module.

4. The data processing method according to claim 1, wherein After sending the target data to the consumption module and obtaining in real time the consumption and successful processing events of the target data by each second distributed node in the consumption module, the following steps are further included: When not obtaining the consumption and successful processing events of all the target data, abandon the consumption and processing of the target data, and re-read the previous data offset statuses of each of the second distributed nodes from the status cache module, and perform re-consumption and re-processing of the target data according to the previous data offset statuses of each of the second distributed nodes.

5. The data processing method according to claim 3, wherein The step of reading the data offset statuses of each of the first distributed nodes and storing the data offset statuses of each of the first distributed nodes in the status cache module includes: Read the data offset statuses of each of the first distributed nodes, associate and bind each of the first distributed nodes with the corresponding data offset status to obtain a first association and binding result, and store the first association and binding result in the status cache module; The step of reading the data offset statuses of each of the second distributed nodes and caching the data offset statuses of each of the second distributed nodes in the status cache module includes: Read the data offset statuses of each of the second distributed nodes, associate and bind each of the second distributed nodes with the corresponding data offset status to obtain a second association and binding result, and store the second association and binding result in the status cache module.

6. A data processing device, characterized in that, For the data processing method according to any one of claims 1-5, the data processing device includes: A first receiving module, configured to receive the pre-submission message data sent by each first distributed node in the production module and cache the pre-submission message data in the intermediate module; A second receiving module, configured to obtain the preset transaction IDs of each of the first distributed nodes and receive in real time the preset identifiers sent by the first distributed nodes corresponding to the preset transaction IDs; An execution module, configured to, when successfully receiving the preset identifiers sent by the first distributed nodes corresponding to all the preset transaction IDs, convert the pre-submission message data into target message data for consumption by the consumption module, and read the data offset statuses of each of the first distributed nodes, and store the data offset statuses of each of the first distributed nodes in the status cache module.

7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the data processing method based on a distributed system according to any one of claims 1-5 is implemented.

8. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the data processing method based on a distributed system as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Distributed transaction processing system and method based on micro-service

    CN112527472A

  • Message processing method, device and equipment and storage medium

    CN113064742A