Method, system and data synchronization device for real-time synchronization of distributed data

The method ensures real-time, consistent data synchronization across diverse databases by identifying operation types and using message queues to handle database-specific changes, addressing the challenge of varying database processing methods and reducing maintenance costs.

CN114428820BActive Publication Date: 2025-07-15TP-LINK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210096601.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-07-15
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

In a distributed cluster system, how to achieve real-time synchronization and consistency of data without being affected by database types.

Method used

By obtaining database operation requests, determining the change operation type, and sending the change statement to the message queue according to the type, using the Kafka queue for batch summary and cross-region transmission, combining partitioning and multi-threading to ensure the order and consistency of the change statements in the message queue and database.

Benefits of technology

Real-time synchronization and consistency of data under different database types is achieved, maintenance costs are reduced, data synchronization is improved stability and efficiency, and is suitable for cross-regional data synchronization and data migration of large cloud services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114428820B_ABST
    Figure CN114428820B_ABST
Patent Text Reader

Abstract

This application relates to the field of data synchronization technology, and in particular, to a method, system, and data synchronization device for real-time distributed data synchronization. The method includes: obtaining a database operation request for indicating an operation on a data record in a first database; when the database operation request includes a change statement, obtaining a change operation type of a second database to be synchronized with the first database; sending the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database; and synchronously writing the change statement into the second database based on the message queue. When using this method for distributed data synchronization, it is not affected by the database type and location, can effectively achieve real-time distributed data synchronization, and ensure the consistency of the synchronized data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data synchronization, and in particular, to a method, a system, and a data synchronization device for real-time synchronization of distributed data. Background Art

[0002] Currently, distributed cluster systems are widely used. A complete distributed cluster system is formed by connecting many clusters in different locations through a network, and a large amount of data is distributed in different clusters of the entire system. Among them, data synchronization is required between databases of different clusters. In the related art, various databases basically have their own data synchronization mechanisms, but generally rely on the core technologies of the databases themselves. However, with different types of databases, the database processing methods are also different. How to ensure real-time synchronization and consistency of data in a distributed cluster database is a problem that needs to be considered currently. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, a system, and a data synchronization device for real-time synchronization of distributed data, which can synchronize distributed data in real time without being affected by the type of database and ensure the consistency of the synchronized data.

[0004] The first aspect of the embodiments of this application provides a method for real-time synchronization of distributed data, including:

[0005] Obtain a database operation request, where the database operation request is used to indicate an operation on a data record in a first database;

[0006] When the database operation request includes a change statement, obtain the change operation type of a second database to be synchronized with the first database;

[0007] Send the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database;

[0008] Based on the message queue, synchronously write the change statement into the second database.

[0009] In a possible implementation manner of the first aspect, the step of sending the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database includes:

[0010] When the change operation type of the second database is idempotent, send the change statement to the message queue in the form of a log.

[0011] In a possible implementation of the first aspect, the change statement includes a specified field for identifying a data record to be operated on, and sending the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database includes:

[0012] When the change operation type of the second database is non-idempotent, determine the partition information corresponding to the change statement in the message queue according to the specified field;

[0013] Send the change statement to the partition identified by the partition information.

[0014] In a possible implementation of the first aspect, the partition information includes a partition key, and determining the partition information corresponding to the change statement in the message queue according to the specified field includes:

[0015] Obtain the number of partitions of the message queue;

[0016] Determine the partition key corresponding to the change statement according to the specified field and the number of partitions.

[0017] In a possible implementation of the first aspect, the change statement includes a specified field for identifying a data record to be operated on, and synchronously writing the change statement to the second database based on the message queue includes:

[0018] Obtain the current number of concurrent threads;

[0019] Determine the target thread corresponding to the change statement according to the specified field and the number of concurrent threads;

[0020] Write the change statement to the second database based on the target thread.

[0021] In a possible implementation of the first aspect, synchronously writing the change statement to the second database includes:

[0022] Obtain the database types of the first database and the second database;

[0023] Determine whether the first database and the second database are heterogeneous databases according to the database types;

[0024] If they are heterogeneous databases, process the change statement into a statement form adapted to the second database and then write it to the second database.

[0025] In a possible implementation of the first aspect, obtaining the database operation request includes:

[0026] Obtain the database operation request from the data access layer through aspect-oriented programming technology.

[0027] The second aspect of the embodiments of the present application provides a distributed data real-time synchronization system, and the system includes:

[0028] An operation request acquisition unit, configured to obtain a database operation request, where the database operation request is used to indicate an operation on data records in a first database;

[0029] An operation type determination unit, configured to, when the database operation request includes a change statement, obtain a change operation type of a second database to be synchronized with the first database;

[0030] A statement sending unit, configured to send the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database;

[0031] A data synchronization unit, configured to synchronously write the change statement into the second database based on the message queue.

[0032] The third aspect of the embodiments of the present application provides a data synchronization device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the distributed data real-time synchronization method provided in the first aspect of the embodiments of the present application are implemented.

[0033] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the distributed data real-time synchronization method provided in the first aspect of the embodiments of the present application are implemented.

[0034] The fifth aspect of the embodiments of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is enabled to execute the steps of the distributed data real-time synchronization method described in the first aspect of the embodiments of the present application.

[0035] In the embodiments of the present application, by obtaining a database operation request, when the database operation request includes a change statement, obtaining a change operation type of a second database to be synchronized with the first database, then sending the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database, and synchronously writing the change statement into the second database based on the message queue, the second database performs the same change operation on the data records to be synchronized. This solution is not affected by the database type and location during distributed data synchronization, can effectively implement distributed data real-time synchronization, and ensure the consistency of the synchronized data. Brief Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a flowchart for implementing the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0038] Figure 1.1 It is a schematic diagram of a scenario for obtaining a data operation request from the data access layer in the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0039] Figure 2 It is a specific flowchart for implementing the method of sending the change statement to the message queue in the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0040] Figure 3 It is a specific flowchart for implementing the method of determining the partition information corresponding to the change statement in the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0041] Figure 4 It is a specific flowchart for implementing the method of synchronously writing the change statement into the second database in the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0042] Figure 5 It is another specific flowchart for implementing the method of synchronously writing the change statement into the second database in the method for real-time distributed data synchronization provided by the embodiment of the present application;

[0043] Figure 6 It is a block diagram of the structure of the system for real-time distributed data synchronization provided by the embodiment of the present application;

[0044] Figure 7 It is a schematic diagram of a data synchronization device provided by the embodiment of the present application. Detailed Embodiments

[0045] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, systems, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0046] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0047] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0048] It should be understood that each method embodiment of the present application provides a method for real-time synchronization of distributed data, which is applicable to various types of data synchronization devices that need to perform data synchronization, specifically, it can be an intermediate device, a cloud server, etc. that are communicatively connected to multiple server clusters.

[0049] The method for real-time synchronization of distributed data provided by the present application will be exemplarily described below in conjunction with specific embodiments.

[0050] Figure 1 The implementation process of the method for real-time synchronization of distributed data provided by the embodiments of the present application is shown. The execution end of the embodiments of the present application can be a data synchronization device. The method process may include the following steps S101 to S104.

[0051] S101: Obtain a database operation request, where the database operation request is used to indicate an operation on the data records in the first database.

[0052] Database operations include queries and changes. In this embodiment, the above database operation request refers to a request for querying or changing the data records in the first database. The above first database is applied to any one of the service clusters in the distributed cluster.

[0053] Generally, all business requests, after being accessed from the upper layer and before being written into the database, will converge to the layer of the database write client. It is the database client that writes the received data into the database. In this embodiment, interception and acquisition are performed according to the feature that all data is written by the database client. For example Figure 1.1As shown in the figure, when a normal service request is received, the service request is processed and parsed by the request access layer and the service processing layer, and then becomes a database operation request when reaching the data access layer. Among them, the data access layer may be implemented as an independent service or embedded in the service layer. The first database directly operates on the data records according to the database operation request. At the same time, the database operation request is obtained at the data access layer, processed and sent to the message queue, and then synchronously written to the second database.

[0054] In the embodiment of the present application, the database operation request can be obtained by using Aspect Oriented Programming (AOP) technology. Specifically, the database operation request is obtained from the data access layer through aspect-oriented programming technology. Among them, AOP technology is a technology that realizes the unified maintenance of program functions through pre-compilation and dynamic proxy during runtime. It can dynamically and uniformly add functions to the program without modifying the source code through pre-compilation and dynamic proxy during runtime. Using AOP technology can isolate each part of the business logic, thereby reducing the coupling degree between parts of the business logic, improving the reusability of the program, and at the same time improving the development efficiency.

[0055] In the embodiment of the present application, by obtaining the database operation request at the data access layer, it is basically unnecessary to perceive the service. When the service adds / modifies the interface, it will not affect the data synchronization model. The change statement is obtained at the data access layer for data synchronization, and service desensitization can be achieved, thereby reducing the subsequent manual maintenance cost.

[0056] S102: When the database operation request includes a change statement, obtain the change operation type of the second database to be synchronized with the first database.

[0057] The above-mentioned second database is the database that needs to be synchronized with the first database. The above-mentioned second database can be applied to another service cluster different from the first database in a distributed cluster.

[0058] The database operation request includes a query statement and / or a change statement. The change includes addition, deletion, and update. In the embodiment of the present application, in step S101, the database operation request is directly obtained at the data access layer first, and then the statements in the database operation request are filtered to determine whether the data operation request includes a change statement, and the non-change statements are filtered out. When the database operation request includes a change statement, the change operation type of the second database to be synchronized with the first database is obtained.

[0059] For example, in the data access layer, the database operation request is obtained through the AOP technology, and then the statements in the database operation request are filtered according to the statement type. If it is a query statement, the data records corresponding to the query statement in the first database are directly read. If it is a change statement, the change operation type of the second database is obtained.

[0060] For databases of different types, the definition of changing the same data record is different. For example, for a redis database, the records corresponding to the same key value are regarded as one data record, and its change statement is to perform operations such as set, del, expire, etc. on this key value. For a MySQL database, the records corresponding to the valid fields of the same unique index are regarded as one data record, and its change statement is to perform operations such as adding, deleting, and modifying on this data record. For example, the commonly used auto-incrementing primary key id is a unique index but not a valid field, and its record cannot be regarded as a valid data record.

[0061] The change operation types of the database include idempotent and non-idempotent. In one implementation, each database includes an operation type identifier, and the operation type identifier is used to identify whether the change operation type of the database is idempotent or non-idempotent. In this embodiment, the change operation type of the second database can be determined by obtaining the operation type identifier of the second database.

[0062] A database with an idempotent change operation type means that for multiple change statements such as add statements, delete statements, and update statements for the same data record, without changing the content of the statements, arbitrarily adjusting the execution order of the statements, the final result of the database is the same. Typically, the Cassandra database belongs to this type. For all adds, deletes, and updates in this type of database, the timestamp used in the statement during execution is used as the criterion. The largest timestamp represents the latest record, which has nothing to do with the execution order of the database.

[0063] For a database with a non-idempotent change operation type, for multiple change statements such as add statements, delete statements, and update statements for the same data record, the operation order will affect the final result of the database. The final result is related to the number and order of execution of the change statements, and the results after each execution operation are different. Most database update operation types are non-idempotent. Commonly used non-idempotent databases include mysql and redis.

[0064] S103: Send the change statement to the message queue according to the preset processing method corresponding to the change operation type of the second database.

[0065] The above message queue is used to batch summarize the change statements of the database, improve the security and stability of long-latency cross-region transmission, and ensure that the change statements can be completely sent to the second database. In practical applications, an appropriate type of message queue can be selected according to the queue usage and the throughput of the change statements of the first database.

[0066] The above message queue can be a Kafka queue. Kafka is a high-throughput distributed publish-subscribe messaging system. The Kafka queue is mainly targeted at scenarios with high cross-region latency, and supports both offline data processing and real-time data processing. Using the Kafka queue can ensure the stability and security of cross-region data transmission. It should be noted that in the embodiments of the present application, the above message queue can also be other queues such as RabbitMQ queue or Rocketmq queue.

[0067] In the embodiments of the present application, different processing methods are selected according to whether the change operation type of the second database is idempotent or non-idempotent.

[0068] In a possible implementation manner, when the change operation type of the second database is idempotent, the change statement is sent to the message queue in the form of a log. Specifically, the log printing framework of log4j / logback can be borrowed to submit the change statement to the Kafka queue.

[0069] In this embodiment, for an idempotent database, there is no need to use a special method to ensure the order of multiple change statements of the same record in the message queue. Therefore, the change of all statements of the idempotent database can be directly submitted to the message queue by borrowing the log printing framework, which not only saves the sending device, but also saves the maintenance cost.

[0070] In a possible implementation manner, when the change operation type of the second database is non-idempotent, the change statement is sent to the message queue in an orderly manner. That is, it is necessary to ensure the order of the change statements in the message queue. In this embodiment, the order of the change statements means that the multiple change statements for the same data record are in the same order as the change statements sent to the first database.

[0071] When operating a non-idempotent database, it is necessary to ensure that the write order of the second database to be synchronized is consistent with the write order of the data access layer. This requires ensuring the order of multiple change statements of the same record in Kafka, so as to ensure the consistency of the data records in the second database and the data records in the first database, and achieve accurate and effective synchronization.

[0072] As a possible implementation manner of the present application, the change statement includes a specified field for identifying the data record to be operated. Figure 2It shows a specific implementation process of sending the change statement to the message queue according to the preset processing method corresponding to the change operation type of the second database in the method provided by the embodiment of the present application, which is described in detail as follows:

[0073] A1: When the change operation type of the second database is non-idempotent, determine the partition information corresponding to the change statement in the message queue according to the specified field.

[0074] The above-mentioned specified field is a field that can index data records, which can be a key value or a primary key value, and can be specifically determined according to the database type. For example, for a redis database, the above-mentioned specified field is a key value; for a MySQL database, the above-mentioned specified field is a valid field that can uniquely index data records, such as a primary key value.

[0075] A2: Send the change statement to the partition identified by the partition information.

[0076] In this embodiment, the message queue is partitioned in advance. After determining the partition information according to the above fields, the change statement is sent to the partition identified by the partition information.

[0077] If there is no partition information corresponding to the specified field, it means that the data record corresponding to the change statement is operated for the first time. At this time, the change statement can be sent to any partition in the message queue in a certain way. For example, the partition information is determined by using a random algorithm, and the change statement is randomly sent to the partition in the message queue.

[0078] The processing of multiple partitions can improve the throughput of change statements in the message queue. If the change statements of the same data record are in different partitions of the distributed message queue, the order of the change statements cannot be guaranteed. Exemplarily, for a key of token:{xiaoming} in the redis database, its stored value is modified twice. The first time, the value of this key is set to account, and the second time, the value of this key is set to accountId. The final record in database 1 is accountId. The time interval between the two changes is very small, just a few milliseconds. The first change statement is sent to partition 1 of Kafka, and the second change statement is sent to partition 2 of Kafka. However, during consumption, due to some reasons, the consumption of partition 1 is 10 ms slower than that of partition 2 of Kafka. That is to say, the second set statement is written to database 2 first, and the first set statement is written to database 2 later. Then the final record of token:{xiaoming} in database 2 is the account changed for the first time, which causes a data inconsistency problem between data 2 and database 1. Ensuring that the same record is on a fixed partition in Kafka is to avoid the performance problem of a single partition affecting the order in which the change statements of the same data record are written to database 2, thereby affecting the data consistency of the two databases to be synchronized.

[0079] In the embodiment of the present application, sending all change statements of the same record to a fixed partition of the message queue can ensure the orderliness of the change statements in the message queue, thereby ensuring the consistency of the synchronized data.

[0080] As a possible implementation manner of the present application, the partition information includes a partition key, the partition key identifies a partition in the message queue, records the mapping relationship between the specified field of the data record and the partition key, determines the partition key according to the specified field, and sends the change statement to the partition identified by the partition key.

[0081] As a possible implementation manner of the present application, Figure 3 It shows the specific implementation process of determining the partition information corresponding to the change statement in the message queue according to the specified field in the method embodiment provided by the embodiment of the present application, which is described in detail as follows:

[0082] B1: Obtain the number of partitions of the message queue.

[0083] B2: Determine the partition key corresponding to the change statement according to the specified field and the number of partitions.

[0084] In the embodiments of the present application, a hash can be specifically used to determine partitions, so as to ensure the order of multiple change statements of the same data record in the first database in the message queue.

[0085] In an application scenario, taking the redis database as an example, the redis database is of the key-value type. The specified field key of the data record to be operated in redis can be directly hashed as shown in the following formula (1) to determine the partition key corresponding to the change statement:

[0086] hash(key) = CRC16(key) % partitionNum (1)

[0087] Where partitionNum is the number of partitions of Kafka, and CRC16 is an algorithm used by redis to calculate hash values. After calculating the key of redis with CRC16 and taking the remainder of the number of partitions partitionNum in Kafka, different keys can be evenly distributed on the partitionNum partitions of Kafka, so as to achieve the purpose of load balancing. The topic of Kafka can have multiple partitions for storage. By sending the change statement corresponding to this key to the hash(key)th partition of the topic, it can be ensured that the change statements of the same key will always exist in the same fixed partition of the Kafka queue, thus ensuring the order of the change statements in Kafka.

[0088] As a possible implementation manner of the present application, different hashing methods can be adopted according to different change operations in the change statement. For example, for the MySQL database, different hash methods can be used for different specific change operations. For the insert statement, the primary key (or unique key) can be filtered out for hashing; for the update and delete statements, the primary key (unique key) needs to be filtered out from the where clause for hashing. The specific hashing method can refer to the above formula (1) and will not be elaborated here. In a small number of scenarios, there may be a situation where changes are made using an index. In this case, the statement of index change needs to be changed into a statement of primary key (or unique key) change. Then filter out the primary key (or unique key) for hashing (generally in actual business, there will not be a situation of large-scale changes through the index. In this case, the latency of obtaining the primary key through index query is not high); of course, if the request pressure of each table in the database is not very different, the table name can also be hashed, but this will cause the pressure distribution of each partition (or each thread writing to the second database) to be not very uniform and needs to be used according to the actual throughput of the changes.

[0089] In the scenario of large-scale distributed cloud services, the write pressure is high. It is certain that single-threadedly sending change statements to the message queue and single-threadedly pulling messages from the message queue to write to the database to be synchronized cannot meet the requirements of real-time synchronization. However, multi-threaded operations will have concurrency problems, affecting the order of writing change statements of the first database to the second database, and ultimately affecting the consistency of data in the second database and the first database.

[0090] S104: Based on the message queue, synchronously write the change statement to the second database.

[0091] In the embodiment of the present application, pull the change statement from the message queue, write the pulled change statement to the second database, and the second database operates on the corresponding data record in real time according to the change statement, so as to achieve real-time synchronization with the first database and ensure the consistency of data in the first database.

[0092] For one queue and multiple consumers, there will be concurrency problems, which will in turn affect the sequentiality of writing change statements to the second database. As a possible implementation manner of the present application, the change statement includes a specified field for identifying the data record to be operated, such as Figure 4 shown, the step of synchronously writing the change statement to the second database based on the message queue further includes:

[0093] C1: Obtain the current number of concurrent threads. The concurrent thread specifically refers to the number of threads that concurrently write to the second database.

[0094] C2: Determine the target thread corresponding to the change statement according to the specified field and the number of concurrent threads. The target thread refers to the thread that fixedly processes the change statement of the data record identified by the specified field.

[0095] C3: Based on the target thread, write the change statement to the second database.

[0096] In the embodiment of the present application, when batch pulling messages from the message queue and then multi-threadedly processing and writing to the non-idempotent second database, the changes to the same data record are operated in the same fixed thread to ensure the order of writing its change statements to the non-idempotent second database.

[0097] In a possible implementation manner, after pulling the change statement from the message queue and multi-threadedly writing to the second database, hash the specified field that uniquely identifies the data record to be operated in the change statement, and then allocate the change statement to a fixed thread for processing according to the hash value, which ensures that all change statements of the same record are always processed by the same thread.

[0098] In an application scenario, taking the hash calculation method with Redis as an example, the hash of the Redis key is calculated as follows:

[0099] hash(key) = CRC16(key) % threadNum

[0100] Where threadNum is the number of concurrent threads writing to the second database. CRC16 is an algorithm used by Redis to calculate hash values.

[0101] In this embodiment, the idea of using hash to determine the number of threads is adopted to ensure that multiple changes to the same data record are processed in the same thread, avoiding the possible concurrency problems that may occur when multiple changes to the same data record are written to the second database through multiple threads.

[0102] For non-idempotent databases, the embodiments of the present application can ensure that the order of writing the change statements to the second database is the same as the order of writing to the first database, thus ensuring the ultimate consistency of the data in the second database and the first database.

[0103] Generally, for steps such as data pulling, parsing, storage, and consumption, the execution has sequentiality. Each step needs to be executed after the previous step is completed, and each step runs based on a single thread, resulting in low overall operation efficiency and long data synchronization time. Multi-threaded and multi-concurrent processing will also affect the sequentiality of multiple change statements of the same data record when writing to the second database. In the example of the present application, after ensuring the sequentiality of all change statements of the same data record in the message queue, it also ensures that all change statements are in order from being pulled from the message queue to being written to the second database.

[0104] A characteristic of Kafka is that only one consumer can exist in a partition at the same time, that is, the data in the same partition will only be pulled by the same processing device for service. This ensures that all change statements of the same data record are processed in the same processing device, avoiding the concurrency problem that all change statements of the same data record are consumed by multiple consumers at the same time, thereby affecting the sequentiality of all change statements of the same data record. Similarly, when multiple messages pulled in batches from Kafka are written to the second database in multi-threaded processing later, it is necessary to ensure that all change statements of the same record are processed in the same thread to ensure their sequential writing to the second database.

[0105] As a possible implementation manner of the present application, Figure 5 The following shows a specific implementation process of synchronously writing the change statements to the second database in the method embodiment provided by the embodiment of the present application, which is described in detail as follows:

[0106] D1: Obtain the database types of the first database and the second database. The database types include Cassandra, MySQL, redis, etc.

[0107] D2: Determine whether the first database and the second database are heterogeneous databases according to the database types. That is, determine whether the types of the first database and the second database are the same.

[0108] D3: If they are heterogeneous databases, after processing the change statement into a statement form adapted to the second database, write it into the second database.

[0109] In the embodiment of the present application, before writing the change statement into the second database, by processing the change statement into a statement form adapted to the second database, the stable progress of the second database synchronization can be guaranteed, the delay or even synchronization exception during data synchronization can be avoided, and the stability and accuracy of distributed data synchronization are further guaranteed.

[0110] The distributed data real-time synchronization proposed in the present application can be widely applied to cross-region data synchronization and data migration of large cloud services. A general processing solution is proposed for both databases with idempotent change operations and databases with non-idempotent change operations. This synchronization solution is universal, does not target specific types of databases, the difference in database types does not affect the distributed data synchronization, and it can achieve real-time synchronization with guaranteed cross-region synchronization security. In practice, it performs well not only for MySQL, but also for the synchronization of two typical databases, redis and Cassandra. When using the solution of the present application for data synchronization and data migration, the synchronization and migration of incremental data can basically be done with zero pause, avoiding the impact on the database during the data migration process. Moreover, the idea of data synchronization in this solution is mainly implemented at the data access layer, and basically does not need to perceive the business. The addition / modification of business interfaces will not affect the data synchronization model and also reduce the subsequent human maintenance cost.

[0111] It should be understood that the magnitudes of the sequence numbers of the steps in the above respective embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0112] Corresponding to the method for distributed data real-time synchronization described in the above embodiments, Figure 6 The structural block diagram of the system for distributed data real-time synchronization provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0113] Refer to Figure 6, this system is applied to a data synchronization device, and the system includes: an operation request obtaining unit 61, an operation type determining unit 62, a statement sending unit 63, and a data synchronization unit 64, where:

[0114] The operation request obtaining unit 61 is configured to obtain a database operation request, and the database operation request is used to indicate an operation on data records in a first database;

[0115] The operation type determining unit 62 is configured to, when the database operation request includes a change statement, obtain a change operation type of a second database to be synchronized with the first database;

[0116] The statement sending unit 63 is configured to send the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database;

[0117] The data synchronization unit 64 is configured to synchronously write the change statement into the second database based on the message queue.

[0118] As a possible implementation manner of this application, the statement sending unit 63 includes:

[0119] The first sending module is configured to, when the change operation type of the second database is idempotent, send the change statement to the message queue in the form of a log.

[0120] As a possible implementation manner of this application, the change statement includes a specified field for identifying a data record to be operated, and the statement sending unit 63 includes:

[0121] The partition information determining module is configured to, when the change operation type of the second database is non-idempotent, determine partition information corresponding to the change statement in the message queue according to the specified field;

[0122] The statement sending module is configured to send the change statement to the partition identified by the partition information.

[0123] As a possible implementation manner of this application, the partition information includes a partition key, and the partition information determining module includes:

[0124] The partition number obtaining sub-module is configured to obtain the number of partitions of the message queue;

[0125] The partition key determining sub-module is configured to determine a partition key corresponding to the change statement according to the specified field and the number of partitions.

[0126] As a possible implementation manner of this application, the change statement includes a specified field for identifying a data record to be operated, and the data synchronization unit 64 includes:

[0127] A thread number acquisition module, configured to acquire the current number of concurrent threads;

[0128] A target thread determination module, configured to determine a target thread corresponding to the change statement according to the specified field and the number of concurrent threads;

[0129] A first data writing module, configured to write the change statement into the second database based on the target thread.

[0130] As a possible implementation manner of this application, the data synchronization unit 64 further includes:

[0131] A database type acquisition module, configured to acquire the database types of the first database and the second database;

[0132] A heterogeneous determination module, configured to determine whether the first database and the second database are heterogeneous databases according to the database types;

[0133] A second data writing module, configured to, if they are heterogeneous databases, process the change statement into a statement form adapted to the second database and then write it into the second database.

[0134] As a possible implementation manner of this application, the operation request acquisition unit 61 includes:

[0135] A request acquisition module, configured to acquire the database operation request from the data access layer through aspect-oriented programming technology.

[0136] In the embodiment of this application, by acquiring a database operation request, when the database operation request includes a change statement, the change operation type of the second database to be synchronized with the first database is acquired, and then the change statement is sent to a message queue according to a preset processing method corresponding to the change operation type of the second database. Based on the message queue, the change statement is synchronously written into the second database, and the second database performs the same change operation on the data records to be synchronized. This solution is not affected by the database type and location during distributed data synchronization, can effectively achieve real-time distributed data synchronization, and ensure the consistency of the synchronized data.

[0137] The embodiment of this application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the distributed data real-time synchronization methods as Figures 1 to 5 represented.

[0138] The embodiment of this application further provides a computer program product, when the computer program product runs on a terminal device, it causes the terminal device to execute and implement asFigures 1 to 5 Steps of any of the methods for real-time synchronization of distributed data represented.

[0139] Figure 7 It is a schematic diagram of a data synchronization device provided by an embodiment of the present application. As Figure 7 shown, the data synchronization device 7 of this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. When the processor 70 executes the computer program 72, the steps in the embodiments of the above-mentioned various methods for real-time synchronization of distributed data are implemented, such as Figure 1 the steps S101 to S104 shown. Alternatively, when the processor 70 executes the computer program 72, the functions of each module / unit in the above-mentioned system embodiments are implemented, such as Figure 6 the functions of the units 61 to 64 shown.

[0140] The computer program 72 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 71 and executed by the processor 70 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 72 in the data synchronization device 7.

[0141] The so-called processor 70 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0142] The memory 71 may be an internal storage unit of the data synchronization device 7, such as a hard disk or a memory of the data synchronization device 7. The memory 71 may also be an external storage device of the data synchronization device 7, such as a plug-in hard disk equipped on the data synchronization device 7, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 71 may also include both an internal storage unit and an external storage device of the data synchronization device 7. The memory 71 is used to store the computer program and other programs and data required by the data synchronization device. The memory 71 may also be used to temporarily store the data that has been output or will be output.

[0143] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be described herein again.

[0144] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working processes of the system, system and unit described above can refer to the corresponding processes in the foregoing method embodiment and will not be described herein again.

[0145] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0146] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0147] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0151] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A method for real-time synchronization of distributed data, characterized in that, The method includes: Obtaining a database operation request for indicating an operation on data records in a first database; When the database operation request includes a change statement, obtaining a change operation type of a second database to be synchronized with the first database; the change operation type includes idempotent and non-idempotent; Sending the change statement to a message queue according to a preset processing method corresponding to the change operation type of the second database; Based on the message queue, synchronously writing the change statement into the second database; The change statement includes a specified field for identifying the data record to be operated. When the change operation type of the second database is non-idempotent, the step of sending the change statement to the message queue according to the preset processing method corresponding to the change operation type of the second database includes: partitioning the message queue in advance. If there is no partition information corresponding to the specified field in the message queue, it indicates that the operation on the data record corresponding to the change statement is performed for the first time. At this time, the change statement is sent to any partition in the message queue. If there is partition information corresponding to the specified field in the message queue, the change statement is sent to the partition identified by the partition information, so as to send all change statements of the same data record to the same fixed partition of the message queue; When the change operation type of the second database is non-idempotent, the step of synchronously writing the change statement into the second database based on the message queue includes: obtaining the current number of concurrent threads; determining the target thread corresponding to the change statement according to the specified field and the number of concurrent threads; writing the change statement into the second database based on the target thread; wherein the target thread is a thread that fixedly processes the change statements of the data records identified by the specified field, so that the changes of the same data record are operated in the same fixed thread, ensuring that the change statements are written into the non-idempotent second database in an orderly manner.

2. The method according to claim 1, wherein The step of sending the change statement to the message queue according to a preset processing method corresponding to the change operation type of the second database includes: When the change operation type of the second database is idempotent, sending the change statement to the message queue in the form of a log.

3. The method according to claim 1, characterized in that, The partition information includes a partition key. The step of determining the partition information corresponding to the change statement in the message queue according to the specified field includes: Obtaining the number of partitions of the message queue; Determining the partition key corresponding to the change statement according to the specified field and the number of partitions.

4. The method according to any one of claims 1 to 3, characterized in that The step of synchronously writing the change statement into the second database includes: Obtaining the database types of the first database and the second database; Determining whether the first database and the second database are heterogeneous databases according to the database types; If they are heterogeneous databases, the change statement is processed into a statement form adapted to the second database and then written into the second database.

5. The method according to any one of claims 1 to 3, characterized in that, The step of obtaining the database operation request includes: Obtaining the database operation request from the data access layer through aspect-oriented programming technology.

6. A system for real-time synchronization of distributed data, characterized in that, The system includes: An operation request acquisition unit for acquiring a database operation request for indicating an operation on data records in a first database; An operation type determination unit for, when the database operation request includes an update statement, acquiring an update operation type of a second database to be synchronized with the first database; the update operation type includes idempotent and non-idempotent; A statement sending unit for sending the update statement to a message queue according to a preset processing method corresponding to the update operation type of the second database; A data synchronization unit for synchronously writing the update statement into the second database based on the message queue; The update statement includes a specified field for identifying the data record to be operated. When the update operation type of the second database is non-idempotent, the step of sending the update statement to the message queue according to the preset processing method corresponding to the update operation type of the second database includes: partitioning the message queue in advance. If there is no partition information corresponding to the specified field in the message queue, it indicates that the data record corresponding to the update statement is being operated for the first time. At this time, the update statement is sent to any partition in the message queue; if there is partition information corresponding to the specified field in the message queue, the update statement is sent to the partition identified by the partition information, so that all update statements of the same data record are sent to the same fixed partition of the message queue; When the update operation type of the second database is non-idempotent, the step of synchronously writing the update statement into the second database based on the message queue includes: acquiring the current number of concurrent threads; determining the target thread corresponding to the update statement according to the specified field and the number of concurrent threads; writing the update statement into the second database based on the target thread; wherein the target thread is a thread that fixedly processes the update statements of the data record identified by the specified field, so that the updates of the same data record are operated in the same fixed thread, ensuring the orderliness of writing the update statement into the non-idempotent second database.

7. A data synchronization device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the steps of the method for real-time synchronization of distributed data according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program implements the steps of the method for real-time synchronization of distributed data according to any one of claims 1 to 5 when executed by the processor.

Citation Information

Patent Citations

  • Data synchronization method and device

    CN110674213A

  • Database synchronization method, system and device, electronic equipment and medium

    CN113094434A