Real-time multi-channel replication method for changed data based on message queue
By using the real-time multi-replication method of changing data based on message queue during data synchronization, dynamically judge the replication type and synchronization conditions, and update the consumption checkpoint, the problems of low efficiency and data loss in cross-database data synchronization in the existing technology are solved, and efficient, secure and consistent data synchronization is achieved.
Patent Information
- Application Number
- CN202510076418.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-24
AI Technical Summary
Existing data synchronization technologies have problems with inefficiency and data loss or excessive differences when accessing across databases, especially when links are interrupted.
The real-time multi-replication method of changing data based on the message queue is adopted. By judging the replication type and synchronization conditions of the changed data in the message queue, the data processing logic is dynamically adjusted, the change data is synchronized to the target data system, and the consumption checkpoint is updated through the consumption status table to achieve flexible, secure and consistent synchronization of the data.
It significantly improves data synchronization efficiency, ensures data stability and consistency, enhances the system's fault tolerance and data processing flexibility, and avoids data loss or repeated synchronization.
Smart Images

Figure CN120201033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data communication, and specifically to a real-time multi-channel replication method for changed data based on a message queue. Background Art
[0002] With the current development of big data and the Internet, people have an increasingly high demand for the synchronization of large volumes of data. However, in the prior art, data synchronization is usually optimized for data synchronization between nodes within a single data system, and there are significant limitations in the face of the increasingly diverse types of databases today, unable to effectively meet the data synchronization requirements across multiple data systems; moreover, in the conventional data synchronization process, if a link interruption occurs, it will lead to missing synchronized data or a large difference in data between different target nodes, ultimately resulting in synchronization failure.
[0003] The patent "Data Synchronization Method and Electronic Device Based on Kafka", publication number: CN118193635A, publication date: June 14, 2024, discloses: comparing the unique identifiers of the data to be synchronized in the hive library with the historical data to determine the synchronization type of the data to be synchronized; according to the synchronization type, sending the data to be synchronized to the corresponding kafka topic; deploying multiple node consumer programs, and determining the target kafka topic to be consumed by each node consumer program according to the kafka topic; creating threads according to the number of target kafka topics to be consumed by each node consumer program, and determining the target kafka topic to be consumed by the threads; reading the data to be synchronized into the mongo table and the elasticsearch table through the threads according to the offset corresponding to the target kafka topic and the synchronization type. However, this solution is still affected by the Kafka storage capacity and the processing capacity of the consumer programs and the backlogged data in solving the link interruption problem, and there are still problems such as missing synchronized data or a large difference in data between different target nodes, ultimately resulting in synchronization failure. Summary of the Invention
[0004] The object of the present invention is to address the problem of low data synchronization efficiency caused by the limitations of conventional data synchronization technologies in cross-database access. A real-time multi-channel replication method for changed data based on a message queue is proposed. By judging the replication type of the changed data in the message queue, the type of the changed data is further dynamically judged to obtain the synchronization conditions for the changed data. Based on the synchronization conditions of the changed data, the corresponding data processing logic is executed to synchronize the changed data to the target data system. The consumption checkpoint for each data synchronization is updated through the consumption status table, enabling flexible full or partial synchronization of data according to business requirements. Moreover, the synchronization data processing logics are independent of each other, ensuring stable data synchronization while guaranteeing the flexibility, security, and consistency of data replication, and significantly improving the data synchronization efficiency.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: A real-time multi-channel replication method for changed data based on a message queue, comprising the following steps: Judge the replication type of the changed data according to the data change message of the source system in the message queue; Dynamically judge the type of the changed data according to the replication type, and obtain the synchronization conditions for the changed data; execute the corresponding data processing logic based on the synchronization conditions of the changed data, and synchronize the changed data to the target data system; Update the consumption checkpoint for each data synchronization through the consumption status table according to the synchronization operation of the changed data.
[0006] In this solution, by judging the replication type and synchronization conditions of the changed data before data replication, the data can be diverted to different processing paths to achieve more refined data replication operations, adjust the replicated data range according to different business scenarios, improve the pertinence and effectiveness of data replication, and avoid unnecessary data transmission and processing. Different data processing logics are used to handle different synchronization conditions, ensuring the correctness, consistency of data operations, and the coherence of business logic, ensuring that data changes conform to business requirements and system specifications, and improving the reliability of the system and data quality. These processing logics not only achieve the synchronization of data between the source and target ends, but also guarantee the rationality and standardization of data operations, while taking into account the system performance and data security. By updating the consumption checkpoint in real time when the synchronization data is completed, it provides a basis for the target data system to resume from interruption. According to the information of the checkpoint, the data synchronization operation can be restarted from the position where the last interruption occurred, avoiding re-synchronizing all data, ensuring the consistency and integrity of data synchronization, improving the fault tolerance of the system, data synchronization efficiency, and data synchronization stability. The checkpoint reduces repeated operations, optimizes resource utilization, and enables efficient incremental data synchronization, overcoming the problem of low data synchronization efficiency caused by the limitations of conventional data synchronization technologies in cross-database access.
[0007] Preferably, determining the distributed replication type of the changed data includes: Reading the data change message in the message queue through the target data system, parsing the message to obtain key change information, including at least the data change type, data change content, and metadata; obtaining the replication type of the changed data in the source system according to the key change information.
[0008] Preferably, dynamically determining the type of the changed data according to the replication type to obtain the synchronization condition of the changed data includes: If the distributed replication type is full replication, directly retrieve the changed data for data synchronization; if the distributed replication type is conditional replication, determine the specific synchronization condition of the changed data based on the data change type.
[0009] Preferably, executing the corresponding data processing logic based on the synchronization condition of the changed data includes: When the changed data is determined to be conditional replication and the synchronization condition is insert, the target data system starts the insert function to execute the data insertion processing logic; when the changed data is determined to be conditional replication and the synchronization condition is update, the target data system starts the update function to execute the data update processing logic; when the changed data is determined to be conditional replication and the synchronization condition is delete, the target data system starts the delete function to execute the data deletion processing logic.
[0010] In this solution, by accurately judging the synchronization condition of the changed data, it is ensured that only the qualified data can be correctly synchronized to the target data system, avoiding the interference of unqualified data on the data consistency of the target side, and ensuring the integrity of the data and the correctness of the business logic. At the same time, by judging the synchronization condition, the data synchronization method can be adjusted flexibly according to the business requirements, so that the target data system can better adapt to the changes of different businesses. For the data that does not meet the replication conditions, no transmission and processing are performed, reducing the occupancy of network bandwidth and the resource consumption of the target system, and improving the data processing efficiency of the target data system.
[0011] Preferably, the target data system starts the insert function to execute the data insertion processing logic, including: Judging whether the newly added data meets the replication conditions of the target data system; if the newly added data meets the conditions, write the newly added data to the master node of the target data system database through the insert function and record the data insertion log; copy the data to all slave nodes of the database based on the data insertion log; the slave nodes write the newly added data to the target database in an asynchronous replication manner through the insert statement in the log.
[0012] Preferably, the target data system starts an update function to execute data update processing logic, including: Determine whether the changed data to be updated meets the replication conditions of the target data system; if the changed data meets the replication conditions, determine whether the corresponding old data value in the target data system meets the replication conditions, and execute corresponding data processing logic according to the judgment result of the old data value; if the changed data does not meet the replication conditions, ignore the current data change message; if the old data value does not meet the replication conditions, write the changed data into the target database based on the insertion function.
[0013] Preferably, the execution of corresponding data processing logic according to the judgment result of the old data value includes: based on the data change content, query whether there is a corresponding old data value in the target data system, if there is no old data value, execute the data insertion processing logic to write the changed data to be updated into the target database; if there is an old data value, write the changed data to the master node of the target data system database through the update function and record the data update log; copy the changed data to all slave nodes of the database based on the data update log; the slave nodes write the changed data into the target database according to the update statement in the log.
[0014] Preferably, the target data system starts a deletion function to execute data deletion processing logic, including: Based on the data change content, determine whether the old data to be deleted corresponding in the target data system meets the deletion conditions, if it meets the deletion conditions, execute the deletion statement through the deletion function to delete the corresponding data in the target data system.
[0015] Preferably, according to the synchronization operation of the changed data, update the consumption checkpoint for each data synchronization through the consumption status table, including: Establish a consumption status table for the target data system, and store the consumption identifier and consumption location information of the data through the consumption status table; when the target data system completes a synchronization of the changed data, record the location information of the data consumption into the consumption status table to update the consumption checkpoint; among them, when the replication type of the changed data is full replication, update the consumption checkpoint after the changed data is successfully applied to the target data system; when the replication type of the changed data is conditional replication, update the consumption checkpoint after the insertion function is successfully executed or the update function is successfully executed or the deletion function is successfully executed.
[0016] Preferably, the method further includes: During the process of judging the synchronization conditions for insertion, update, and deletion, if the current replication condition is not met, the processing logic corresponding to the current synchronization condition is ignored, and the logic processing of other replication conditions is performed.
[0017] Advantages of the present invention: While this application realizes multi-channel replication, it can flexibly perform full or partial synchronization of data according to business requirements. Moreover, the synchronization data processing logics are independent of each other, and the data systems at the synchronization target end do not interfere with each other. While ensuring data stability, it can also flexibly configure the processing logic to synchronize data to multiple different data systems in real time, significantly improving the flexibility, stability, and consistency of data replication. By updating the consumption checkpoint, the position of the last consumption in the message queue can be quickly found, and the unsynchronized data can be continuously synchronized, avoiding data loss or repeated application of synchronized data, and ensuring the accuracy and correctness of data synchronization. Description of the drawings
[0018] Other features, objectives, and advantages of the present invention will become more obvious by reading the detailed description of the non-limiting embodiments with reference to the following drawings. The drawings are only for the purpose of showing the preferred embodiments and are not considered as limiting the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0019] Figure 1 It is a flowchart of a method for real-time multi-channel replication of changed data based on a message queue according to an embodiment of the present invention.
[0020] Figure 2 It is a flowchart of a data insertion processing logic according to an embodiment of the present invention.
[0021] Figure 3 It is a flowchart of a data update processing logic according to an embodiment of the present invention.
[0022] Figure 4 It is a flowchart of a data deletion processing logic according to an embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of a consumption checkpoint update process according to an embodiment of the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific implementation manners described herein are only the best embodiments of the present invention and are only used to explain the present invention, without limiting the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0025] Example 1: As Figure 1 shown, a real-time multi-channel replication method for changed data based on a message queue includes steps S1 - S4, where: S1. According to the data change message of the source system in the message queue, determine the replication type of the changed data; specifically including: reading the data change message in the message queue through the target data system, performing message parsing on it, and obtaining key change information, including at least the data change type, data change content, and metadata; Obtain the replication type of the changed data of the source system according to the key change information.
[0026] In this embodiment, the target data system is used to read the data change message from the queue, the read message is parsed, and the key change information in the message, including the change type, data content, and metadata, is extracted. According to the parsed change type, it is determined which type of replication operation should be performed.
[0027] In some embodiments, the determination of the distributed replication type can be completed through a configuration file or database settings. For example, store replication type information in the system configuration table, and retrieve the corresponding replication type information by querying this table; or set the corresponding relationship of each replication type in the configuration file, match the corresponding relationship of the replication type according to the data change information in the message queue, and query the matching result to obtain the replication type of this data change; for example, use the key-value pair replication_type:full to represent full replication, and replication_type:conditional to represent conditional replication.
[0028] In some other embodiments, different business rules can be set according to business requirements to determine the replication type. During system initialization or data backup, it may be necessary to synchronize all data changes at the source end to the target data system without discrimination, corresponding to the full - volume replication type; while scenarios where only data that meets specific business logics is synchronized, such as only synchronizing order data after a specific date or only synchronizing data of a specific user group, correspond to distributed conditional replication. For example, the corresponding replication type is determined according to data update frequency, time range, etc. For instance, for data with frequent updates, conditional replication may be more suitable to avoid unnecessary data synchronization and only synchronize data that meets specific conditions to the target end to reduce the burden of data transmission and processing; for data with infrequent updates, such as configuration data, full - volume replication may be more suitable because the data volume is relatively small and full - volume replication can ensure data integrity and consistency. For example, when only synchronizing data in the most recent month, conditional replication is required at this time. Data can be filtered by setting time conditions; if the business requires synchronizing all - time - range data at the source end to the target end, such as during data migration, full - volume replication is needed.
[0029] In this embodiment, the replication - type judgment logic is the starting point of the entire data - replication process. It determines whether to perform full - volume replication or conditional replication according to configurations or business rules. The target data system reads and parses the stored replication - type information and then executes the corresponding replication logic respectively, providing the basis and direction for subsequent data - synchronization operations. Among them, full - volume replication focuses on synchronizing data without discrimination, while conditional replication pays more attention to screening data according to business requirements to ensure that only data changes that meet the conditions are synchronized to the target data system. This judgment logic provides flexibility and customizability for the system to meet different business scenarios and data - synchronization requirements.
[0030] S2. Dynamically judge the type of the changed data according to the replication type, and obtain the synchronization conditions of the changed data; specifically, it includes: If the distributed replication type is full - volume replication, directly retrieve the changed data for data synchronization; If the distributed replication type is conditional replication, determine the specific synchronization conditions of the changed data based on the data - change type.
[0031] In this embodiment, if it is a full - volume replication table, the changed data is directly applied; if it is a distributed replication table, it is necessary to continue to analyze according to the type of the changed data to determine how to synchronize the data to the target data system according to different conditions.
[0032] S3. Execute the corresponding data - processing logic based on the synchronization conditions of the changed data, and synchronize the changed data to the target data system; specifically, it includes: When the changed data is determined to be conditional replication and the synchronization condition is insertion, the target data system starts the insertion function to execute the data insertion processing logic; When the changed data is determined to be conditional replication and the synchronization condition is update, the target data system starts the update function to execute the data update processing logic; When the changed data is determined to be conditional replication and the synchronization condition is deletion, the target data system starts the deletion function to execute the data deletion processing logic.
[0033] Furthermore, the target data system can dynamically select the synchronization of changed data according to business data requirements, and flexibly synchronize the changed data by flexibly calling each processing logic function, achieving the adaptability and flexibility of data synchronization, and improving the data synchronization efficiency of the distributed data system.
[0034] In this embodiment, by accurately judging the synchronization conditions of the changed data, it is ensured that only the data meeting the conditions can be correctly synchronized to the target data system, avoiding the interference of the data not meeting the conditions on the data consistency of the target side, and ensuring the integrity of the data and the correctness of the business logic. At the same time, through the judgment of the synchronization conditions, the data synchronization method can be adjusted flexibly according to business requirements, so that the target data system can better adapt to the changes of different businesses. For the data that does not meet the replication conditions, no transmission and processing are performed, reducing the occupation of network bandwidth and the resource consumption of the target-side system, and improving the data processing efficiency of the target data system.
[0035] Specifically, as Figure 2 shown, when the target data system starts the insertion function to execute the data insertion processing logic, it includes: Judging whether the newly added data meets the replication conditions of the target data system; If the newly added data meets the conditions, the newly added data is written into the master node of the target data system database through the insertion function and the data insertion log is recorded; Based on the data insertion log, the data is replicated to all slave nodes of the database; The slave nodes write the newly added data into the target database in an asynchronous replication manner through the insertion statement in the log.
[0036] Furthermore, judging whether the newly added data meets the replication conditions of the target data system can be set according to the key data fields to ensure that the inserted data contains all the required fields. For example, checking whether the newly added data contains the fields necessary for the target data system, and the values of these fields need to meet the corresponding format, or a certain value range, or be associated with other business data in the target data system to ensure the consistency of the business logic.
[0037] In this embodiment, by checking whether the newly added value meets the conditions, illegal or incomplete data is prevented from entering the system, thus maintaining the integrity and consistency of the data. For example, it prevents user information lacking key information or data that does not conform to business logic from being inserted into the system, thereby ensuring the data quality of the entire system.
[0038] Specifically, as Figure 3 shown, the target data system starts an update function to execute data update processing logic, including: Determining whether the change data to be updated meets the replication conditions of the target data system; If the change data meets the replication conditions, then determining whether the corresponding old data value in the target data system meets the replication conditions, and executing corresponding data processing logic according to the judgment result of the old data value; If the change data does not meet the replication conditions, then ignoring the current data change message; If the old data value does not meet the replication conditions, then writing the change data into the target database based on the insert function.
[0039] Specifically, the execution of corresponding data processing logic according to the judgment result of the old data value includes: based on the data change content, querying whether there is a corresponding old data value in the target data system. If there is no old data value, then executing the data insertion processing logic to write the change data to be updated into the target database; If there is an old data value, then writing the change data into the master node of the target data system database through the update function and recording a data update log; Copying the change data to all slave nodes of the database based on the data update log; The slave nodes write the change data into the target database according to the update statement in the log.
[0040] Furthermore, the replication conditions that the change data to be updated needs to meet can be set according to the validity of the updated value, data consistency, and business logic relevance. For example, the updated value needs to meet the conditions set by the business rules, the data cannot be negative or cannot exceed a certain upper limit, etc., the updated data cannot damage data consistency, and the updated value needs to conform to the association relationship of the current data system's business logic. In the update logic, judging the old data value can ensure that the update operation is based on existing data records. If there is no corresponding old data value, the change data is regarded as new data for replication synchronization to ensure data consistency between the target data system and the source data system.
[0041] In this embodiment, the information of the existing data in the target data system is updated through the update processing logic to reflect the latest status of the source-side data, ensuring the timeliness and accuracy of the data. By checking whether the new value and the old value meet the replication conditions, it is ensured that the update operation will not damage the business logic and data consistency, and the coherence of the business process is ensured. Among them, in the conditional judgment of the new value and the old value, when the old value does not meet the replication conditions while the new value meets, it means that there is no historical data corresponding to the changed data in the target data system, then the new value is inserted as a new record to make the data state of the target system consistent with the source side, so as to meet the business requirements for data consistency, avoid business problems caused by data inconsistency, and at the same time ensure the fault tolerance of the system, that is, the target data system can flexibly process update operations according to the actual situation of the data.
[0042] Specifically, as Figure 4 shown, the target data system starts a deletion function to execute the data deletion processing logic, including: Based on the data change content, it is judged whether the corresponding old data to be deleted in the target data system meets the deletion conditions. If the deletion conditions are met, the execution deletion statement will be passed through the deletion function to delete the corresponding data in the target data system.
[0043] Furthermore, by judging whether the corresponding old data to be deleted in the target data system meets the deletion conditions, that is, it is necessary to confirm whether the data record to be deleted exists. If it does not exist, it does not meet the deletion conditions. In addition, the permissions of the old data to be deleted can be checked and whether the data has associated data can be checked. Only when certain permission conditions are met and the data has no associated data can the deletion operation be performed.
[0044] In this embodiment, the data deletion operation on the source side is synchronized to the target system through the deletion logic to ensure the consistency of the data on the source side and the target side, and avoid business problems caused by data redundancy and inconsistency; by checking whether the old value meets the replication conditions, accidental deletion operations are prevented. For example, only data that meets specific conditions can be deleted, avoiding accidental loss of data, and ensuring the security of the data and the reliability of the system. At the same time, it is judged whether the replication conditions are met to ensure that the deletion operation conforms to the business rules and data associations, and avoid damaging the business logic due to the deletion operation.
[0045] S4. According to the synchronization operation of the changed data, update the consumption checkpoint for each data synchronization through the consumption status table; as Figure 5 shown, it specifically includes: Establish a consumption status table for the target data system, and store the consumption identifier and consumption location information of the data through the consumption status table; When the target data system completes one synchronization of the changed data, record the position information of data consumption in the consumption status table to update the consumption checkpoint; Among them, when the replication type of the changed data is full replication, update the consumption checkpoint after the changed data is successfully applied to the target data system; When the replication type of the changed data is conditional replication, update the consumption checkpoint after the insert function is successfully executed or the update function is successfully executed or the delete function is successfully executed.
[0046] Furthermore, the database transaction function of the target data system can be utilized to achieve the characteristic that batch consumption is either all successful or all failed to implement the update of the checkpoint, ensuring that messages are not consumed repeatedly, and at the same time improving the replication performance through the batch mode. Specifically as follows: In the target data system, store the consumption identifier and consumption position information through the consumption status table, and utilize the database transaction function to ensure the consistency of batch consumption, that is, after each batch of messages is consumed, the position information of this consumption will be recorded in the consumption status table; among them, for full replication, update the checkpoint after the data is successfully applied to the target data system; for conditional replication, update the checkpoint after the insert statement is successfully executed, the update statement is successfully executed, or the delete statement is successfully executed. Ensure that in the target data system, each batch of data operations such as insert, update, and delete is either all successful or all failed, guaranteeing the reliability and consistency of data synchronization.
[0047] In this embodiment, the consumption checkpoint provides a starting point identifier for each data replication of the target data system, indicating that this data replication will start from the checkpoint position, and the data before the checkpoint has been successfully synchronized, avoiding repeated data synchronization; for example, when the data synchronization program suddenly interrupts, the checkpoint provides a recovery point, and the program can restart the data synchronization operation from the position where it was interrupted last time according to the consumption position information, avoiding resynchronizing all the data, saving time and resources, and at the same time ensuring the integrity of the data. At the same time, the consistency of data synchronization and the fault tolerance of the target data system are ensured by updating the checkpoint. For example, when synchronizing a batch of data to the target system, multiple insert, update, or delete operations may need to be executed. Updating the checkpoint can be used as the end flag of a transaction to ensure that these operations are either all successful or all failed, preventing data inconsistency problems caused by partial data updates while some data is not updated; for example, in a complex distributed environment, network, hardware, or software failures are inevitable. The checkpoint mechanism allows the system to recover quickly after a failure, reducing the risk of data loss and data inconsistency caused by the failure.
[0048] Specifically, the method further includes: During the process of judging the synchronization conditions for insertion, update, and deletion, if the current replication condition is not met, the processing logic corresponding to the current synchronization condition is ignored, and the logic processing of other replication conditions is performed.
[0049] In this embodiment, the synchronization conditions are independent of each other, that is, the insertion, update, and deletion logic functions do not interfere with each other, and the corresponding programs run independently; when one of the logic functions is running, it will not affect the execution of other data processing logics, and the synchronized data will not be affected either. At the same time, when the replication condition of a logic function is not met, this change data information record is automatically ignored, and the next information record is judged to ensure data consistency in the target data system during the entire data synchronization period.
[0050] The above specific implementation manners are the preferred implementation manners of the present invention, and do not limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. All equivalent changes made according to the shape, structure, and method of the present invention are within the protection scope of the present invention.
Claims
1. A method for real-time multi-path replication of changed data based on a message queue, characterized in that: The steps include: Determine the replication type of the changed data based on the data change message of the source system in the message queue; Dynamically determine the type of the changed data according to the replication type to obtain a synchronization condition for the changed data; Execute corresponding data processing logic based on the synchronization condition of the changed data to synchronize the changed data to the target data system; According to the synchronization operation of the changed data, the consumption checkpoint of each data synchronization is updated through the consumption status table.
2. The method for real-time multi-path replication of changed data based on message queue according to claim 1, characterized in that: The determining of the distributed replication type of the changed data includes: The target data system reads the data change message in the message queue, parses the message, and obtains key change information, including at least the data change type, data change content, and metadata; The replication type of the source end system change data is obtained according to the key change information.
3. The method for real-time multi-path replication of changed data based on message queue according to claim 1, characterized in that: The dynamically judging the type of the changed data according to the replication type to obtain the synchronization condition of the changed data includes: if the distributed replication type is full replication, directly retrieving the changed data for data synchronization; If the distributed replication type is conditional replication, the specific synchronization condition of the changed data is determined based on the data change type.
4. The method for real-time multi-path replication of changed data based on message queue according to claim 3 is characterized in that: The executing corresponding data processing logic based on the synchronization condition of the changed data includes: When the changed data is determined to be the conditional replication and the synchronization condition is insertion, the target data system starts an insertion function to execute data insertion processing logic; When the changed data is determined to be the conditional replication and the synchronization condition is update, the target data system starts an update function to execute data update processing logic; When the changed data is determined to be the conditional replication and the synchronization condition is deletion, the target data system starts a deletion function to execute data deletion processing logic.
5. The method for real-time multi-path replication of changed data based on message queue according to claim 4 is characterized in that: The target data system starts an insert function to execute data insertion processing logic, including: Determine whether the newly added data meets the replication conditions of the target data system; If the newly added data meets the conditions, the newly added data is written into the master node of the target data system database through the insertion function and the data insertion log is recorded; Copying the data to all slave nodes of the database based on the data insertion log; The slave node writes the newly added data into the target database using an asynchronous replication method through the insert statement in the log.
6. The method for real-time multi-path replication of changed data based on message queue according to claim 4, characterized in that: The target data system starts an update function to execute data update processing logic, including: Determining whether the changed data to be updated meets the replication condition of the target data system; If the changed data meets the replication condition, then determine whether the corresponding old data value in the target data system meets the replication condition, and execute corresponding data processing logic according to the determination result of the old data value; If the changed data does not meet the replication condition, the current data change message is ignored; If the old data value does not satisfy the replication condition, the changed data is written into the target database based on the insert function.
7. The method for real-time multi-path replication of changed data based on message queue according to claim 6, characterized in that: The executing corresponding data processing logic according to the judgment result of the old data value includes: Based on the data change content, query whether there is a corresponding old data value in the target data system, and if there is no old data value, execute the data insertion processing logic to write the changed data to be updated into the target database; If there is an old data value, the changed data is written to the master node of the target data system database through the update function and the data update log is recorded; Based on the data update log, copy the changed data to all slave nodes of the database; The slave node writes the changed data into the target database according to the update statement in the log.
8. The method for real-time multi-path replication of changed data based on message queue according to claim 4, characterized in that: The target data system starts a deletion function to execute data deletion processing logic, including: Based on the data change content, it is determined whether the corresponding old data to be deleted in the target data system meets the deletion conditions. If the deletion conditions are met, the deletion statement is executed through the deletion function to delete the corresponding data in the target data system.
9. The method for real-time multi-path replication of changed data based on message queue according to claim 4, characterized in that: The synchronization operation according to the changed data, updating the consumption checkpoint of each data synchronization through the consumption status table, comprises: Establishing a consumption status table of the target data system, and storing the consumption identification and consumption location information of the data through the consumption status table; When the target data system completes synchronization of the changed data once, the location information of the data consumption is recorded in the consumption status table to update the consumption checkpoint; Wherein, when the replication type of the changed data is full replication, the consumption checkpoint is updated after the changed data is successfully applied to the target data system; When the replication type of the changed data is conditional on replication, the consumption checkpoint is updated after the insert function is successfully executed, the update function is successfully executed, or the delete function is successfully executed.
10. The method for real-time multi-path replication of changed data based on message queue according to claim 9, characterized in that: The method further comprises: During the synchronization condition judgment process of inserting, updating and deleting, if the current replication condition is not met, the processing logic corresponding to the current synchronization condition is ignored, and the logic processing of other replication conditions is performed.
Citation Information
Patent Citations
Data synchronization method based on kafka and electronic equipment
CN118193635A