MongoDB bidirectional data synchronization method and system based on logic clock
Through a logical clock-based method, the two-way data synchronization of MongoDB cluster is realized, solving the problem that existing tools cannot achieve reliable two-way synchronization, and achieving data consistency, reliability and high availability.
Patent Information
- Application Number
- CN202411928339.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing MongoDB bidirectional data synchronization tools have not yet achieved reliable bidirectional data synchronization, and existing tools such as Mongoshake and Flink CDC can only support one-way synchronization, and the components dependencies are numerous and deployment is heavy.
Using the MongoDB bidirectional data synchronization method based on logical clock, data is received from MongoDB's change data stream through Mongo synchronization service and sent to the Kafka cluster to achieve orderly delivery of messages. At the same time, the data deduplication and reliability are ensured through the automatic increment of id and checkpoint mechanisms, and the write conflict is detected through logical clocks to solve the problem of double write conflict.
The two-way data synchronization of MongoDB cluster is realized, ensuring the consistency and reliability of data synchronization, solving the problem of double write conflict, providing high availability and real-time, and supporting dual-active and high-availability construction in remote areas.
Smart Images

Figure CN120011445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, in particular, to the field of data synchronization technology; specifically, to a MongoDB bidirectional data synchronization method and system based on a logical clock. Background Art
[0002] As corporate business develops, the requirements for high availability and high reliability of services are becoming increasingly higher, and the demand for data synchronization across regions and data centers in database systems will also increase.
[0003] MongoDB is one of the most widely used open source NoSQL databases. It stores data in the form of documents. Compared with traditional relational databases, MongoDB has higher scalability and flexibility. However, MongoDB's support for data synchronization is not perfect enough. It only supports one-way data synchronization from the source database to the target database. At present, there is no reliable out-of-the-box MongoDB two-way data synchronization tool.
[0004] There are two main types of existing MongoDB data synchronization tools:
[0005] 1. Mongoshake is the most commonly used MongoDB data synchronization tool. However, Mongoshake only supports two-way synchronization of Alibaba's deeply customized version of MongoDB, and only one-way data synchronization for the open source version of MongoDB.
[0006] 2. Flink CDC, the full name of which is Flink Change Data Capture, is a real-time data synchronization technology implemented using the Apache Flink stream processing framework. It can capture and synchronize incremental changes in the source database in real time, and transmit the changed data to the target system in real time, thereby achieving data synchronization; however, Flink CDC also only supports one-way synchronization, and has multiple component dependencies, making deployment more cumbersome. Summary of the invention
[0007] In view of this, the purpose of the present invention is to develop a MongoDB bidirectional data synchronization method and system based on logical clock, which can support bidirectional data synchronization of MongoDB clusters and realize the ability of dual active in different locations; the data synchronization platform automatically solves the problem of double write conflicts, while ensuring the consistency of data source data and the high availability and reliability of the synchronization platform.
[0008] The present invention provides a MongoDB bidirectional data synchronization method based on a logical clock, comprising:
[0009] S1. Receive data from the mongodb change data stream (ChangeStream) through the Mongo synchronization service mongo-sync-service. <database> . <collection>As the key of Kafka cluster (Kafka message queue) messages, the message is partitioned by key, so that the change data of the same library table are written to the same topic partition of Kafka to ensure the orderly delivery of messages;
[0010] Specifically, because the data received from the mongodb change data stream through the Mongo synchronization service mongo-sync-service is itself ordered, as long as the order of writing messages to Kafka and consuming messages from Kafka is guaranteed, the ordered delivery of messages can be achieved by only controlling the order of change data in the same library table.
[0011] S2. The source-side Mongo synchronization service maintains an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is increased by 1 (+1) for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka; the source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka; the target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id; after the target-side Mongo synchronization service completes writing a batch of data, it commits the offset;
[0012] S3. Add a dc column to the Mongo table, and leave the dc column to the application for maintenance; the Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column;
[0013] S4. Detect write conflicts through the logical clock, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column is increasing; determine whether there may be a conflict by comparing the column values of the createTime column, and handle it according to the conflict resolution strategy;
[0014] Specifically, there are two conditions that need to be met for a write conflict to exist:
[0015] 1. The key values written at both ends are the same;
[0016] 2. Both ends write at the same time, that is, when DC1 writes data a and it has not yet been synchronized to DC2, DC2 also writes data b.
[0017] The first condition is easy to judge, but the second condition is more difficult to judge. As long as we can determine whether b is written before or after a is synchronized to dc2, we can determine whether there is a write conflict.
[0018] Therefore, when data synchronization is performed, the target end synchronizes the message with this message and writes it to the source end, so that the source end can receive the latest createTime of the target end. In this way, when this message is synchronized to the target end, it can be compared with createTime to determine whether there is a possible conflict;
[0019] S5. Use akka cluster to build Mongo synchronization service. Use akka cluster singleton to select an instance from the akka cluster as the master node. The master node listens to Mongo's change data stream and pushes it to kafka. At the same time, the master node also regularly executes checkpoints and sends the automatically incremented id, resumeToken and the latest received timestamp information to a specified topic in kafka. All other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance.
[0020] Specifically, on both the source and target ends, the master node instance is used to consume data and checkpoints are performed regularly. The automatically incremented ID values of each partition received are sent to the specified topic of Kafka. Other instances besides the master node instance continue to consume the specified topic to obtain the latest status of the master node instance.
[0021] The Mongo synchronization service supports high availability and can still run when a single node fails.
[0022] Furthermore, the method of detecting write conflicts by using a logical clock in step S4 includes:
[0023] Assume that a piece of data a is written to the source MongoDB cluster dc1, and the createTime value of the data a is recorded as Ta. After the data Ta is synchronized to the target MongoDB cluster dc2, dc2 writes data b and synchronizes it to dc1 to overwrite data c. If Tc > Ta, there is no conflict. If Tc ≤ Ta, there may be a conflict.
[0024] Furthermore, the source Mongo synchronization service of the S2 step periodically performs checkpoints, and sends the latest auto-increment id and resumeToken of the changed data stream to the specified topic partition of Kafka, including:
[0025] If the source Mongo synchronization service is restarted, the state before the restart is obtained, the latest auto-increment id and resumeToken of the changed data stream are obtained from the specified topic partition, and sent to the specified topic partition of Kafka.
[0026] Furthermore, the deduplication according to the automatically incremented ID in step S2 includes:
[0027] If the id does not increase, it means that it is duplicate data and the id that does not increase is ignored.
[0028] Furthermore, the step S3 of handing over the dc column to the application for maintenance includes:
[0029] When the application writes data in the dc column, it writes a fixed value based on the application's sequence number. For example, the value in the dc1 environment is 1, and the value in the dc2 environment is 2.
[0030] Furthermore, all other instances except the master node in step S5 continuously consume the specified topic, and obtaining the latest status of the master node instance includes:
[0031] When the master node instance fails, an instance other than the master node instance is elected as the new master node instance to continue providing services according to the latest checkpoint.
[0032] The present invention also provides a MongoDB bidirectional data synchronization system based on a logical clock, which is used to implement the MongoDB bidirectional data synchronization method based on a logical clock as described above, including:
[0033] Message delivery order module: used to receive data from the mongodb change data stream through the Mongo synchronization service mongo-sync-service <database> . <collection>As the key of Kafka cluster messages, the messages are partitioned by key, so that the changed data of the same library table are written to the same topic partition of Kafka, ensuring the orderly delivery of messages;
[0034] Message transmission reliability module: used for the source-side Mongo synchronization service to maintain an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is incremented by 1 for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka; the source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka; the target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id; after the target-side Mongo synchronization service completes the writing of a batch of data, it commits the offset;
[0035] Service storm processing module: used to add a new dc column to the Mongo table and hand over the dc column to the application for maintenance; the Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column;
[0036] Write conflict handling module: used to detect write conflicts through logical clocks, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column increases; judge whether there may be a conflict by comparing the column value of the createTime column, and handle it according to the conflict resolution strategy;
[0037] High availability module: used to build Mongo synchronization service with Akka cluster. Akka clustersingleton is used to select an instance from Akka cluster as the master node. The master node monitors Mongo's change data stream and pushes it to Kafka. At the same time, the master node also performs checkpoints regularly and sends the automatically incremented ID, resumeToken and the latest received timestamp information to a specified topic in Kafka. All other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance.
[0038] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the MongoDB bidirectional data synchronization method based on a logical clock as described above are implemented.
[0039] The present invention also provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the MongoDB bidirectional data synchronization method based on logical clock as described above are implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The logical clock-based MongoDB bidirectional data synchronization method and system provided by the present invention realize bidirectional data synchronization of MongoDB clusters, can ensure the consistency of data synchronization, and effectively solve the problem of double write conflicts; can ensure the reliability of data synchronization, and can ensure that data is not lost regardless of network problems or node failure restarts; can ensure the real-time performance of data synchronization, and the synchronization delay is at the second level; can ensure high availability, and the failure of a single or a few node instances will not affect the normal synchronization of data, providing strong support for the construction of dual-active and high-availability business services in different locations. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiment.The drawings are only for the purpose of illustrating the preferred embodiments and are not to be construed as limiting the invention.
[0043] In the attached picture:
[0044] Figure 1 Schematic diagram of the overall architecture of a MongoDB bidirectional data synchronization system based on a logical clock according to an embodiment of the present invention;
[0045] Figure 2 A schematic diagram of the data flow process of the Mongo synchronization service of dc1 in an embodiment of the present invention monitoring the Mongo change data stream, processing and filtering the data, and then sending it to Kafka;
[0046] Figure 3 A processing logic diagram for realizing high availability of data synchronization according to an embodiment of the present invention;
[0047] Figure 4 It is a flow chart of the MongoDB bidirectional data synchronization method based on logical clock of the present invention;
[0048] Figure 5 The figure is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and products consistent with some aspects of the present disclosure as detailed in the appended claims.
[0050] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "the" used in this disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0051] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0052] The embodiments of the present invention are described in further detail below.
[0053] The embodiment of the present invention provides a MongoDB bidirectional data synchronization method based on a logical clock, see Figure 4 As shown, including:
[0054] S1. Receive data from the mongodb change data stream (ChangeStream) through the Mongo synchronization service mongo-sync-service. <database> . <collection>As the key of Kafka cluster (Kafka database) messages, the messages are partitioned by key, so that the changed data of the same library table are written to the same topic partition of Kafka to ensure the orderly delivery of messages;
[0055] S2. The source-side Mongo synchronization service maintains an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is increased by 1 (+1) for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka; the source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka; the target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id; after the target-side Mongo synchronization service completes writing a batch of data, it commits the offset;
[0056] The source Mongo synchronization service periodically performs checkpoints and sends the latest auto-increment id and resumeToken of the changed data stream to the specified topic partition of Kafka, including:
[0057] If the source Mongo synchronization service is restarted, the state before the restart is obtained, the latest auto-increment id and resumeToken of the changed data stream are obtained from the specified topic partition, and sent to the specified topic partition of Kafka.
[0058] Deduplication according to the automatic increment id includes:
[0059] If the id does not increase, it means that it is duplicate data and the id that does not increase is ignored.
[0060] S3. Add a dc column to the Mongo table, and leave the dc column to the application for maintenance; the Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column;
[0061] Handing over the DC column to the application for maintenance includes:
[0062] When the application writes the dc column data, it writes a fixed value according to the arrangement sequence number of the application. In this embodiment, the value is 1 in the dc1 environment and 2 in the dc2 environment.
[0063] S4. Detect write conflicts through the logical clock, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column is increasing; determine whether there may be a conflict by comparing the column values of the createTime column, and handle it according to the conflict resolution strategy;
[0064] Methods for detecting write conflicts through logical clocks include:
[0065] Assume that a piece of data a is written to the source MongoDB cluster dc1, and the createTime value of the data a is recorded as Ta. After the data Ta is synchronized to the target MongoDB cluster dc2, dc2 writes data b and synchronizes it to dc1 to overwrite data c. If Tc > Ta, there is no conflict. If Tc ≤ Ta, there may be a conflict.
[0066] S5. Use akka cluster to build Mongo synchronization service. Use akka cluster singleton to select an instance from the akka cluster as the master node. The master node listens to Mongo's change data stream and pushes it to Kafka (such as Figure 3 Indicated by the solid line in the figure), the master node also periodically executes checkpoints, sending the information of the automatically incremented id, resumeToken, and the latest received timestamp to a specified topic in Kafka; all other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance (such as Figure 3 denoted by the dotted line in ).
[0067] When the master node instance fails, an instance other than the master node instance is elected as the new master node instance to continue providing services according to the latest checkpoint.
[0068] On both the source and target ends, the master node instance is used to consume data and checkpoints are performed regularly. The automatically incremented ID values of each partition received are sent to the specified topic of Kafka. Other instances besides the master node instance continue to consume the specified topic to obtain the latest status of the master node instance.
[0069] The Mongo synchronization service supports high availability and can still run when a single node fails.
[0070] The embodiment of the present invention further provides a MongoDB bidirectional data synchronization system based on a logical clock, which is used to implement the MongoDB bidirectional data synchronization method based on a logical clock as described above, including:
[0071] Message delivery order module: used to receive data from the mongodb change data stream through the Mongo synchronization service mongo-sync-service <database> . <collection>As the key of Kafka cluster messages, the messages are partitioned by key, so that the changed data of the same library table are written to the same topic partition of Kafka, ensuring the orderly delivery of messages;
[0072] Message transmission reliability module: used for the source-side Mongo synchronization service to maintain an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is incremented by 1 for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka; the source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka; the target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id; after the target-side Mongo synchronization service completes the writing of a batch of data, it commits the offset;
[0073] Service storm processing module: used to add a new dc column to the Mongo table and hand over the dc column to the application for maintenance; the Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column;
[0074] Write conflict handling module: used to detect write conflicts through logical clocks, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column increases; judge whether there may be a conflict by comparing the column value of the createTime column, and handle it according to the conflict resolution strategy;
[0075] High availability module: used to build Mongo synchronization service with Akka cluster. Akka clustersingleton is used to select an instance from Akka cluster as the master node. The master node monitors Mongo's change data stream and pushes it to Kafka. At the same time, the master node also performs checkpoints regularly and sends the automatically incremented ID, resumeToken and the latest received timestamp information to a specified topic in Kafka. All other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance.
[0076] The overall architecture of the embodiment of the present invention in a practical application is as follows: Figure 1 As shown in the figure, dc1 and dc2 represent two data centers, and a MongoDB cluster is deployed in each data center. The solid arrows represent the flow of data from dc1mongo to dc2mongo, and the dotted arrows represent the flow of data from dc2mongo to dc1mongo. The components included are:
[0077] 1. MongoDB cluster, responsible for storing all business data. This component provides the change stream function, which can monitor the change data of the mongo table in real time through the client API, and can start monitoring from the specified position of the stream through the resumeToken function.
[0078] 2. Kafka cluster, responsible for caching and synchronizing data to ensure the reliability and traceability of data transmission;
[0079] 3. MirrorMaker is the data synchronization component of Kafka, responsible for synchronizing Kafka data across data centers;
[0080] 4. Mongo Sync Service, the core component of Mongo Sync Service, has two functions:
[0081] (a) Monitor the change stream in MongoDB, process and filter the data in the data stream according to certain rules, and send it to Kafka;
[0082] (b) Listen to Kafka, receive data sent from the external environment, synchronize it back to MongoDB, and automatically handle write conflicts in the process;
[0083] The data flow process is as follows:
[0084] 1. The application service writes data to the MongoDB cluster, and the MongoDB persistent data changes;
[0085] 2. The Mongo sync service captures the changed data by listening to the change stream and sends the data to the Kafka cluster;
[0086] 3. Kafka Mirror Maker synchronizes the changed data from the local Kafka cluster to the Kafka cluster in the target data center;
[0087] 4. The Mongo synchronization service of the target cluster pulls the new synchronized data from Kafka, executes the conflict resolution strategy, and synchronizes it to the MongoDB cluster of the target cluster.
[0088] Figure 2 The figure shows the data flow process in which the Mongo synchronization service of dc1 listens to the Mongo change data stream, processes and filters the data, and then sends it to Kafka. Figure 2 In the mongo change data flow, there are five OPQRS messages, of which OQS is written by the client and PR is data synchronized from dc2. The Mongo synchronization service maintains an incremental value ( Figure 2 The service receives the maximum createTime value of the synchronization message from dc2 obtained from the change data stream.
[0089] When the Mongo sync service receives a message written by the client (that is, dc=1), it adds an incremented inc value and T2 to the message and sends the message to Kafka. When a synchronization message is received from dc2 (that is, dc=2), the internally maintained T2 value is refreshed and the data is filtered out.
[0090] After receiving the synchronized message from the source, if there is data with the same primary key in mongo, dc2's Mongo synchronization service will compare the message's T2 with the createTime of the data already inserted into mongo to determine whether there is a conflict. The following table describes the centralized scenarios such as conflict, normal, or abnormal, as well as the processing logic.
[0091] "Normal" means that there is no write conflict in data synchronization; "Conflict" means that there is a write conflict and a conflict resolution strategy needs to be executed. The commonly used conflict resolution strategy is to define a priority value for dc1 and dc2 respectively. The data from the cluster with a higher priority can overwrite the data with a lower priority; "Abnormal" means that there is an unexpected problem, which means that there may be a configuration error or a serious system failure, and technical personnel intervention is required.
[0092]
[0093] In the solution of this embodiment, the createTime value in the mongo table is required to be strictly increasing, that is, the createTime of the newly written data must be greater than or equal to the old data. However, in actual applications, due to the influence of time synchronization, gc and other issues, it is impossible to be strictly increasing, and there are more or less disordered situations. Therefore, a watermark mechanism can be introduced, that is, when the synchronization service monitors the change data stream and updates the internal timestamp T value, it subtracts n seconds from the latest createTime. This can effectively avoid the impact of createTime disorder, but it will increase the possibility of triggering a write conflict alarm to a certain extent.
[0094] The logical clock-based MongoDB bidirectional data synchronization method and system of this embodiment realizes bidirectional data synchronization of MongoDB clusters, ensures the consistency of data synchronization, and effectively solves the problem of double write conflicts; ensures the reliability of data synchronization, and can ensure that data is not lost regardless of network problems or node failure restarts; ensures the real-time nature of data synchronization, with synchronization delays at the second level; and ensures high availability, so that failures of a single or a few node instances will not affect normal data synchronization.
[0095] An embodiment of the present invention further provides a computer device, Figure 5 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention; see the accompanying drawings Figure 5 As shown, the computer device includes: an input system 23, an output system 24, a memory 22 and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the MongoDB bidirectional data synchronization method based on logical clocks as provided in the above embodiment; wherein the input system 23, the output system 24, the memory 22 and the processor 21 can be connected via a bus or other means, Figure 5 The example of connecting through bus is taken in the following.
[0096] The memory 22 is a readable and writable storage medium of a computing device, which can be used to store software programs and computer executable programs, such as the program instructions corresponding to the MongoDB bidirectional data synchronization method based on the logical clock described in the embodiment of the present invention; the memory 22 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created according to the use of the device, etc.; in addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device; in some instances, the memory 22 can further include a memory remotely arranged relative to the processor 21, and these remote memories can be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0097] The input system 23 may be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the device; the output system 24 may include display devices such as display screens.
[0098] The processor 21 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 22, that is, implements the above-mentioned MongoDB bidirectional data synchronization method based on logical clock.
[0099] The computer device provided above can be used to execute the MongoDB bidirectional data synchronization method based on logical clock provided in the above embodiment, and has corresponding functions and beneficial effects.
[0100] The embodiment of the present invention also provides a storage medium containing computer executable instructions, which are used to execute the MongoDB bidirectional data synchronization method based on the logical clock as provided in the above embodiment when executed by a computer processor. The storage medium is any of various types of memory devices or storage devices, and the storage medium includes: installation media, such as CD-ROM, floppy disk or tape system; computer system memory or random access memory, such as DRAM, DDRRAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disk or optical storage); registers or other similar types of memory elements, etc.; the storage medium may also include other types of memory or a combination thereof; in addition, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system, which is connected to the first computer system via a network (such as the Internet); the second computer system may provide program instructions to the first computer for execution. The storage medium includes two or more storage media that may reside in different locations (for example, in different computer systems connected via a network). The storage medium may store program instructions (for example, specifically implemented as a computer program) that can be executed by one or more processors.
[0101] Of course, the computer executable instructions of a storage medium containing computer executable instructions provided in an embodiment of the present invention are not limited to the MongoDB bidirectional data synchronization method based on logical clocks as described in the above embodiment, and can also execute related operations in the MongoDB bidirectional data synchronization method based on logical clocks provided in any embodiment of the present invention.
[0102] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments, but it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / collection> < / database> < / collection> < / database> < / collection> < / database> < / collection> < / database>
Claims
1. A MongoDB bidirectional data synchronization method based on logical clock, characterized in that: include: S1. Receive data from the change data stream of mongodb through the Mongo synchronization service mongo-sync-service. <database> . <collection> As the key of Kafka cluster messages, the messages are partitioned by key, so that the changed data of the same library table are written to the same topic partition of Kafka, ensuring the orderly delivery of messages;< / collection> < / database> S2. The source-side Mongo synchronization service maintains an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is incremented by 1 for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka. The source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka. The target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id. After the target-side Mongo synchronization service completes writing a batch of data, it commits the offset. S3 and Mongo tables add a dc column, and leave the dc column to the application for maintenance; The Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column; S4. Detect write conflicts through the logical clock, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column is increasing; determine whether there may be a conflict by comparing the column values of the createTime column, and handle it according to the conflict resolution strategy; S5. Use akka cluster to build Mongo synchronization service. Use akka cluster singleton to select an instance from the akka cluster as the master node. The master node listens to Mongo's change data stream and pushes it to kafka. At the same time, the master node also regularly executes checkpoints and sends the automatically incremented id, resumeToken and the latest received timestamp information to a specified topic in kafka. All other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance.
2. The MongoDB bidirectional data synchronization method based on logical clock according to claim 1 is characterized in that: The method for detecting write conflicts by using a logical clock in step S4 includes: Assume that a piece of data a is written to the source MongoDB cluster dc1, and the createTime value of the data a is recorded as Ta. After the data Ta is synchronized to the target MongoDB cluster dc2, dc2 writes data b and synchronizes it to dc1 to overwrite data c. If Tc> Ta, there is no conflict. If Tc ≤ Ta, there may be a conflict.
3. The MongoDB bidirectional data synchronization method based on logical clock according to claim 1 is characterized in that: The source Mongo synchronization service in step S2 periodically performs checkpoints and sends the latest auto-increment id and resumeToken of the changed data stream to the specified topic partition of Kafka, including: If the source Mongo synchronization service is restarted, the state before the restart is obtained, the latest auto-increment id and resumeToken of the changed data stream are obtained from the specified topic partition, and sent to the specified topic partition of Kafka.
4. The MongoDB bidirectional data synchronization method based on logical clock according to claim 1 is characterized in that: The deduplication according to the automatically incremented ID in step S2 includes: If the id does not increase, it means that it is duplicate data and the id that does not increase is ignored.
5. The MongoDB bidirectional data synchronization method based on logical clock according to claim 1 is characterized in that: The step S3 of handing over the DC column to the application end for maintenance includes: When the application writes the dc column data, a fixed value is written according to the arrangement sequence number of the application.
6. The MongoDB bidirectional data synchronization method based on logical clock according to claim 1 is characterized in that: All other instances except the master node in step S5 continue to consume the specified topic, and obtaining the latest status of the master node instance includes: When the master node instance fails, an instance other than the master node instance is elected as the new master node instance to continue providing services according to the latest checkpoint.
7. The MongoDB bidirectional data synchronization system based on logical clock is characterized by: The method for implementing a MongoDB bidirectional data synchronization method based on a logical clock as claimed in any one of claims 1 to 6 comprises: Message delivery order module: used to receive data from the mongodb change data stream through the Mongo synchronization service mongo-sync-service <database> . <collection> As the key of Kafka cluster messages, the messages are partitioned by key, so that the changed data of the same library table are written to the same topic partition of Kafka, ensuring the orderly delivery of messages;< / collection> < / database> Message transmission reliability module: used for the source-side Mongo synchronization service to maintain an auto-incrementing id internally. When monitoring the change data stream, the auto-incrementing id is incremented by 1 for each data monitored, and the auto-incrementing id is assigned to the change data, and the change data is sent to Kafka; the source-side Mongo synchronization service periodically performs checkpoints, and sends the latest auto-incrementing id and resumeToken of the change data stream to the specified topic partition of Kafka; the target-side Mongo synchronization service consumes the Kafka data synchronized by the source-side Mongo synchronization service, and deduplicates according to the auto-incrementing id; after the target-side Mongo synchronization service completes the writing of a batch of data, it commits the offset; Service storm processing module: used to add a new dc column to the Mongo table and hand over the dc column to the application for maintenance; the Mongo synchronization service determines whether the changed data is written by the client or synchronized from the external environment based on the value of the dc column; Write conflict handling module: used to detect write conflicts through logical clocks, set a createTime column in the Mongo table, use the createTime column to indicate the writing time of a message, and the trend of the column value of the createTime column increases; judge whether there may be a conflict by comparing the column value of the createTime column, and handle it according to the conflict resolution strategy; High availability module: used to build Mongo synchronization service with Akka cluster. Akka cluster singleton is used to select an instance from Akka cluster as the master node. The master node monitors Mongo's change data stream and pushes it to Kafka. At the same time, the master node also performs checkpoints regularly and sends the information of automatically incremented ID, resumeToken and the latest received timestamp to a specified topic of Kafka. All other instances except the master node continuously consume the specified topic to obtain the latest status of the master node instance.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the MongoDB bidirectional data synchronization method based on a logical clock according to any one of claims 1 to 6 are implemented.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the MongoDB bidirectional data synchronization method based on logical clock as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Sensor arrays, method for operating a sensor array and a computer program for performing a method for operating a sensor array
CN111936044A
NoSQL database synchronization method and device, equipment and storage medium
CN115455113A
Timestamp allocation method, equipment, storage medium and system
CN116388916A
Systems and methods for managing distributed database deployments
US20170286516A1
Cited By
A MongoDB dual-master synchronization method and system supporting field-level concurrent writes
CN122507712A