Automatic complement method and device for solving data synchronization difference data
By capturing change data, performing real-time comparison, and classifying differential data, combined with automated complement strategies, the problem of rapidly repairing differential data during data synchronization is solved, enabling efficient and accurate data synchronization and complementation, and improving system efficiency and flexibility.
Patent Information
- Application Number
- CN202510782412.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the data synchronization process, data loss, duplication or conflict may occur due to network delays, system failures or human operational errors. Existing technologies make it difficult to quickly and accurately detect discrepancies and automatically supplement them. Existing methods are inefficient and cannot meet the needs of large data volumes.
Adopt change data capture, real-time comparison, difference data classification and automated complement strategy, use database CDC function or log parsing technology to capture data changes, perform efficient comparison through XORMap collection, process difference data in batches, formulate complement strategy based on difference type, and automatically execute complement operation through triggering scripts or API.
It can quickly and accurately find and repair data differences under limited hardware resources, improve system efficiency and accuracy, support multiple complement strategies, adapt to large data volume requirements, reduce manual intervention, and support system horizontal expansion.
Smart Images

Figure CN120631979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular provides a method and device for automatically complementing data for solving data synchronization discrepancy. Background Art
[0002] During the data synchronization process, network delays, system failures or human errors may cause data loss, duplication or conflict, resulting in data differences.
[0003] Manual comparison and repair are inefficient, error-prone, and cannot meet the needs of large data volumes.
[0004] Simple script processing lacks flexibility and scalability and is difficult to handle complex difference scenarios.
[0005] Although big data platforms can handle large amounts of data, they require high hardware resources and have high deployment and maintenance costs.
[0006] After data synchronization is completed, how to quickly find specific difference data within limited hardware resources, promptly discover inconsistent data, display the reconciliation results of data synchronization, and complete the complement operation through the automatic complement function to improve data consistency and system efficiency is an urgent problem that needs to be solved by technical personnel in this field. Summary of the Invention
[0007] The present invention aims to address the deficiencies of the above-mentioned prior art and provides a highly practical method for automatically complementing data for resolving data synchronization differences.
[0008] A further technical task of the present invention is to provide a rationally designed, safe and applicable automatic complement device for solving data synchronization difference data.
[0009] The technical solution adopted by the present invention to solve its technical problem is:
[0010] An automatic complement method for resolving data synchronization discrepancies has the following steps:
[0011] S1, change data capture;
[0012] S2, real-time data comparison;
[0013] S3, data difference classification;
[0014] S4, step strategy development;
[0015] S5. Perform automatic complementation.
[0016] Furthermore, in step S1, based on the CDC function or log parsing technology of the database, the insert, update and delete operations of the source and target database tables are captured in real time, the changed data is encapsulated as events, and sent to the message queue through Kafka.
[0017] Furthermore, the CDC function is a built-in function provided by the database, which is used to capture the change operation of the data table. For databases that do not support the CDC function, the data changes are captured through log parsing technology.
[0018] Log parsing technology reads the database operation log, parses the data change operations, parses the change events of the source table and the target table, encapsulates them into message queues, and uses Kafka to send them to the difference task in real time for comparison.
[0019] Furthermore, in step S2, a special XORMap set is constructed based on the hashMap in JDK. The set is similar to an exclusive OR operation: if the set does not contain the key, or the value of the corresponding key is different, the data will be stored in the set normally; if the set already contains the key, the data in the key will be deleted.
[0020] Furthermore, the change event message queue sent is monitored, and the XORMap efficient data structure is used. The batch processing mode is adopted to divide the large data set into multiple small data sets for processing, and a double-layer loop is constructed to process the message queues of the source and target change data respectively. The table with large data volume is placed in the outer layer to ensure that after the outer layer is traversed, all the data in the inner layer are traversed, and the difference data between the source table and the target table is obtained.
[0021] Furthermore, in step S3, the difference data obtained in step S2 is divided into missing data, redundant data and conflicting data according to the difference type of the data, and primary keys for the corresponding data are created for collection storage respectively;
[0022] For missing data, the source table exists but the target table lacks data. By creating a missingData collection to store the primary key,
[0023] For redundant data, which does not exist in the source table but exists in the target table, create a redundantData collection to store the primary key.
[0024] For data that exists in both the source table and the target table but has inconsistent content, create a conflictData collection to store the primary key.
[0025] Furthermore, in step S4, according to the difference type in step S3, the corresponding complement strategy is analyzed and customized:
[0026] (1) Missing data: For missing data in missingData, traverse to obtain the key from the source table and insert it into the target table;
[0027] (2) Redundant data: For redundant data in redundantData, traverse to get the key and delete it from the target table;
[0028] (3) For the conflicting data of conflictData, traverse to obtain the key, automatically determine the latest data according to the preset rules, and perform the update operation.
[0029] Furthermore, in step S5, for the obtained data set of difference type, a corresponding complement strategy is adopted to automatically perform the complement operation by triggering a script or API;
[0030] For automatic insertion of missing data, read the missing data from the source table and insert the data into the target table;
[0031] For automatic deletion of redundant data, redundant data is read from the target table and the data is deleted from the target table;
[0032] For automated updates of conflicting data, read conflicting data from the source and target tables, compare them according to preset rules, select the data to retain, and update the data in the target table with the data in the source table.
[0033] An automatic complement device for resolving data synchronization discrepancy data includes: at least one memory and at least one processor;
[0034] The at least one memory is configured to store a machine-readable program;
[0035] The at least one processor is configured to call the machine-readable program to execute an automated complement method for resolving data synchronization difference data.
[0036] Compared with the prior art, the method and device for automatically complementing data synchronization discrepancies of the present invention have the following outstanding beneficial effects:
[0037] This invention achieves automated repair of data synchronization discrepancies through real-time comparison, differential data classification, automated complement strategies, and result verification. It efficiently completes large-scale comparison and complement operations with limited resources, reducing manual intervention and improving system efficiency and accuracy. It supports multiple complement strategies and can be flexibly configured according to business needs. It detects data differences in real time and provides rapid feedback on the results. Through message queues and event-driven mechanisms, it supports horizontal expansion of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 It is a flow chart of an automated method for complementing data to solve data synchronization discrepancies;
[0040] Figure 2 The invention is a schematic diagram of traversing and searching for difference data in an automated complement method for solving data synchronization difference data. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0042] A best embodiment is given below:
[0043] like Figure 1 As shown, an automatic complement method for resolving data synchronization difference data in this embodiment has the following steps:
[0044] S1, change data capture;
[0045] Capture insert, update, and delete operations on source and target database tables in real time using the database's CDC functionality or log parsing technology. Encapsulate the changed data as events and send them to a message queue via Kafka.
[0046] CDC is a built-in database feature that captures changes to data tables. Common databases (such as MySQL, PostgreSQL, Oracle, and SQL Server) support CDC. CDC works by monitoring the database's transaction log (such as MySQL's binlog and PostgreSQL's WAL log) to capture data changes in real time.
[0047] For databases that don't support CDC, data changes can be captured using log parsing technology. Log parsing technology reads the database's operation log (such as MySQL's binlog or MongoDB's oplog) and analyzes the data change operations.
[0048] Parse the change events of the source table and the target table, encapsulate them into a message queue, and use Kafka to send them to the difference task in real time for comparison.
[0049] S2, real-time data comparison;
[0050] Based on the hashMap in JDK, a special XORMap collection structure is constructed as follows:
[0051] public class XORMap<K,V> extends HashMap<K,V> {+
[0052] @Override.
[0053] public V put(K key,V value){+
[0054] if(super.containsKey(key)&&value==super.get(key)){+
[0055] return remove(key);+
[0056] }else{+
[0057] }+
[0058] return super.put(key,value);+
[0059] This set is similar to an XOR operation: if the set does not contain this key, or the value of the corresponding key is different, the data will be stored in the set normally; if the set already contains this key, the data in the key will be deleted.
[0060] like Figure 2 As shown, the incoming change event message queue is monitored. Leveraging the efficient XORMap data structure and adopting batch processing, large datasets are divided into multiple smaller ones for processing, reducing memory requirements. We construct a two-layer loop to process the message queues for source and target change data separately. Tables with large data volumes are placed in the outer layer to ensure that after the outer layer is traversed, all data in the inner layer is traversed, allowing the difference data between the source and target tables to be obtained.
[0061] S3, data difference classification;
[0062] The difference data obtained in step S2 is divided into missing data, redundant data, and conflicting data according to the difference type. A primary key key is created for each set to store the corresponding data.
[0063] Missing data: Data that exists in the source table but is missing in the target table. Create a missingData collection to store the primary key.
[0064] Example: Source table: id = 1, name = John; Target table: No record with id = 1.
[0065] Redundant data: redundant data that does not exist in the source table but exists in the target table. Create a redundantData collection to store the primary key.
[0066] Example: Source table: No record with id=2, target table: id=2, name=Alice.
[0067] Conflicting data: Data that exists in both the source and target tables but has inconsistent content. Create a conflictData collection to store the primary key.
[0068] Example: Source table: id=3, name=Bo, target table: id=3, name=Bob.
[0069] S4, step strategy formulation;
[0070] According to the difference type in step S3, the corresponding complement strategy is analyzed and customized:
[0071] Missing data: For missing data in missingData, traverse to obtain the key, obtain the data from the source table, and insert it into the target table.
[0072] Redundant data: For redundant data in redundantData, traverse to obtain the key and delete it from the target table.
[0073] Conflicting data: For conflictData, traverse to obtain the key, automatically determine the latest data according to the preset rules (timestamp), and perform the update operation.
[0074] S5, perform automatic complementation;
[0075] For the resulting datasets with different types, we apply the corresponding complement strategy and automatically execute complement operations through triggering scripts or APIs. We leverage parallel processing modes using multithreading or distributed computing frameworks to increase processing speed. We also support transaction mechanisms during processing to ensure the atomicity of complement operations.
[0076] For automatic insertion of missing data, read the missing data from the source table and insert the data into the target table;
[0077] For automatic deletion of redundant data, redundant data is read from the target table and the data is deleted from the target table.
[0078] For the automated update of conflicting data, the conflicting data is read from the source table and the target table, compared according to the preset rules (timestamp), and the data of one side is selected to be retained, and the data in the target table is updated with the data of the source table.
[0079] After the complement is completed, a secondary comparison is automatically triggered to verify data consistency and the verification results are fed back to the system.
[0080] Secondary comparison:
[0081] Re-compare the data in the source table and the target table to check whether there are any discrepancies.
[0082] Verification result processing:
[0083] If the data consistency verification passes, the log is recorded and the process ends. If the data consistency verification fails, an alarm is triggered and the complement operation is executed again.
[0084] Through the above steps, you can quickly find the difference data between the source table and the target table under limited resources, automatically complete the complement processing, and ensure the data synchronization function.
[0085] Based on the above method, an automatic complement device for solving data synchronization discrepancy data in this embodiment includes: at least one memory and at least one processor;
[0086] The at least one memory is configured to store a machine-readable program;
[0087] The at least one processor is configured to call the machine-readable program to execute an automated complement method for resolving data synchronization difference data.
[0088] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.
[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An automated method for solving data synchronization discrepancy data, characterized in that: The steps are as follows: S1, change data capture; S2, real-time data comparison; S3, data difference classification; S4, step strategy development; S5. Perform automatic complementation.
2. The method for automatically complementing data to solve data synchronization discrepancy according to claim 1, characterized in that: In step S1, based on the database's CDC function or log parsing technology, the insert, update, and delete operations of the source and target database tables are captured in real time, the changed data is encapsulated as events, and sent to the message queue through Kafka.
3. The method for automatically complementing data synchronization differences according to claim 2, characterized in that: CDC is a built-in function provided by the database, which is used to capture changes to data tables. For databases that do not support CDC, data changes are captured through log parsing technology. Log parsing technology reads the database operation log, parses the data change operations, parses the change events of the source table and the target table, encapsulates them into message queues, and uses Kafka to send them to the difference task in real time for comparison.
4. The method for automatically complementing data synchronization differences according to claim 3, characterized in that: In step S2, a special XORMap set is constructed based on the hashMap in JDK. The set is similar to an exclusive OR operation: if the set does not contain the key, or the value of the corresponding key is different, the data will be stored in the set normally; if the set already contains the key, the data in the key will be deleted.
5. The method for automatically complementing data to solve data synchronization discrepancy according to claim 4, characterized in that: Listen to the incoming change event message queue, use the XORMap efficient data structure, adopt the batch processing mode, divide the large data set into multiple small data sets for processing, build a double-layer loop, process the message queues of the source and target change data respectively, and place the table with large data volume in the outer layer to ensure that after the outer layer is traversed, all the data in the inner layer are traversed, and the difference data between the source table and the target table is obtained.
6. The method for automatically complementing data to solve data synchronization discrepancy according to claim 5, characterized in that: In step S3, the difference data obtained in step S2 is divided into missing data, redundant data and conflicting data according to the difference type of the data, and a primary key key for the corresponding data is created for each set; For missing data, the source table exists but the target table lacks data. By creating a missingData collection to store the primary key, For redundant data, which does not exist in the source table but exists in the target table, create a redundantData collection to store the primary key. For data that exists in both the source table and the target table but has inconsistent content, create a conflictData collection to store the primary key.
7. The method for automatically complementing data to solve data synchronization discrepancy according to claim 6, characterized in that: In step S4, according to the difference type in step S3, the corresponding complement strategy is analyzed and customized: (1) Missing data: For missing data in missingData, traverse to obtain the key from the source table and insert it into the target table; (2) Redundant data: For redundant data in redundantData, traverse to get the key and delete it from the target table; (3) For the conflicting data of conflictData, traverse to obtain the key, automatically determine the latest data according to the preset rules, and perform the update operation.
8. The method for automatically complementing data to solve data synchronization discrepancy according to claim 7, characterized in that: In step S5, for the obtained data set of difference type, a corresponding complement strategy is adopted to automatically perform the complement operation by triggering a script or API; For automatic insertion of missing data, read the missing data from the source table and insert the data into the target table; For automatic deletion of redundant data, redundant data is read from the target table and the data is deleted from the target table; For automated updates of conflicting data, read conflicting data from the source and target tables, compare them according to preset rules, select the data to retain, and update the data in the target table with the data in the source table.
9. An automatic complement device for solving data synchronization difference data, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.