Remote database synchronization system and method based on incremental data extraction and verification
Through a remote database synchronization system based on incremental data extraction and verification, the problems of low transmission efficiency and great impact on the performance of the source database in the existing technology are solved, and efficient and accurate data synchronization is achieved, which is suitable for complex distributed environments.
Patent Information
- Application Number
- CN202510108561.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing remote database synchronization method has low transmission efficiency in scenarios with large data volume and high real-time requirements, and has a great impact on the performance of the source database, making it difficult to cope with the synchronization requirements in complex distributed environments.
A remote database synchronization system based on incremental data extraction and verification is adopted, including a variable detection module, a data transmission module, a data verification module and an exception processing module. By monitoring the changes of the source database in real time, an incremental data change set is generated, and data compression, encryption and verification are carried out to achieve efficient data transmission and synchronization.
It improves data transmission and synchronization efficiency, adapts to high-frequency data changes scenarios, ensures the integrity and accuracy of data synchronization, and supports complex distributed environments and cross-regional data synchronization.
Smart Images

Figure CN120050291A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data transmission and database synchronization, and particularly relates to a remote database synchronization system and method based on incremental data extraction and verification. Background Art
[0002] With the rapid development of big data and distributed systems, the remote synchronization of databases has become an important technical requirement for enterprise informatization. The existing remote database synchronization methods mainly include two ways: full - volume transmission and trigger - based synchronization:
[0003] 1. Full - volume transmission: By regularly copying the complete data set in the source database to the target database. This method is simple and easy to use, but when the data volume is large, the transmission efficiency is low, the network load is high, and it is difficult to meet the scenarios with high real - time requirements.
[0004] 2. Trigger - based synchronization: By capturing the changes in the source database through triggers and transmitting them to the target database in real - time. However, this method has a greater impact on the performance of the source database, and there are certain limitations in data conflict detection, transmission error handling, etc., and it is difficult to meet the synchronization requirements in complex distributed environments.
[0005] In addition, in a large - scale data environment, the data transmission process often faces the following challenges: Conflict between real - time and performance: High - frequency data changes may lead to network and database performance bottlenecks, affecting the timeliness of synchronization. Insufficient guarantee of data accuracy: During the transmission process, data may be lost or damaged due to network fluctuations or transmission errors, and there is a lack of effective verification and recovery mechanisms. Imperfect exception - handling mechanism: The existing technologies lack a sound automated processing ability in dealing with exceptions such as network interruptions and data loss, resulting in data inconsistency or synchronization interruption. Summary of the Invention
[0006] The problem to be solved by the present invention is to improve the synchronization efficiency and data consistency, and adapt to complex and changeable distributed database environments, and a remote database synchronization system and method based on incremental data extraction and verification are proposed.
[0007] To achieve the above object, the present invention is realized through the following technical solutions:
[0008] A remote database synchronization system based on incremental data extraction and verification includes a variable detection module, a data application module, a data transmission module, a data verification module, and an exception - handling module;
[0009] The variable detection module and the data application module are respectively connected to the data transmission module, the data transmission module is respectively connected to the data verification module and the exception - handling module, and the exception - handling module is also connected to the data verification module;
[0010] The source database connection variable detection module is respectively connected to the data application module and the data transmission module for the target database;
[0011] The variable detection module is used to monitor the new, updated or deleted operations in the source database in real time and generate an incremental data change set;
[0012] The data transmission module is used to transmit the incremental data change set generated by the variable detection module to the target database in batches through the network;
[0013] The data verification module is used to verify the integrity and accuracy of the incremental data in the incremental data change set during and after the transmission process;
[0014] The data application module is used to parse, apply the incremental data in the incremental data change set transmitted to the target database, and resolve data conflicts;
[0015] The exception handling module is connected to the data transmission module and the data verification module, and is used to trigger a retry mechanism or an alarm mechanism when network exceptions, data loss or verification failures are detected.
[0016] Furthermore, the data transmission module includes a data compression unit and a data encryption unit.
[0017] Furthermore, the exception handling module includes an alarm unit and a recovery unit.
[0018] A remote database synchronization method based on incremental data extraction and verification is realized relying on the described remote database synchronization system based on incremental data extraction and verification, and includes the following steps:
[0019] S1. The source database connection change detection module, the change detection module captures the new, updated and deleted operations of the source database in real time, extracts the table name, primary key, field content, operation type and timestamp of the changed data by parsing the transaction log or using the trigger mechanism, and the changed data is encapsulated into an incremental data change set DCS and stored in the queue for buffering;
[0020] S2. The data transmission module divides the incremental data in the incremental data change set obtained in step S1 into blocks. Before transmission, the data compression unit uses the compression algorithm Gzip+ to compress the incremental data, and then the data encryption unit encrypts the incremental data through the encryption algorithm, and then transmits the encrypted incremental data to the target database;
[0021] The data verification module uses the national secret SM3 hash algorithm to generate a verification code for the transmitted data block for data integrity verification during the data transmission process, and uses the redundancy verification algorithm to perform redundancy verification on all incremental data after the data transmission is completed;
[0022] After the target database receives and verifies the data, the data application module parses the incremental data and synchronizes it to the target database. It calls database instructions according to the operation type in the incremental data to complete the synchronization, and constructs a conflict detection mechanism and a conflict resolution mechanism to handle data conflicts;
[0023] S4. During the data transmission and data verification processes, the exception handling module processes network interruptions, data loss, or verification failures. The alarm unit of the exception handling module sends alarm prompts via email, text message, or real-time notification. When the target database is abnormal or the update fails, the recovery unit of the exception handling module rolls back or reapplies the incremental data according to the transmission log.
[0024] Furthermore, the specific implementation method of step S2 includes the following steps:
[0025] S2.1. The data compression unit compresses the incremental data using the optimized compression algorithm Gzip+. The compression ratio CR is calculated using the following formula:
[0026]
[0027] where H(d i ) and H(d' j ) are the entropy values of the original data block and the compressed data block respectively, f(α,β,W i ) is an adaptive compression function that depends on the data block size and data type weight. The first parameter α is used to adjust the sliding window size, the second parameter β is used to select different coding strategies, and W i is the weight of the i-th data block. During the compression process, the weight of the i-th data block is automatically set according to the size and type of the data block, indicating the relative importance of each data block in the compression ratio calculation, so as to ensure that each data block is given an appropriate proportion in the compression algorithm. R is the amount of redundant data, which is expressed as the redundant information added during the compression process, and λ is the weight coefficient of the redundant data, which is used to balance the compression effect and the redundant capacity;
[0028] By adaptively adjusting the sliding window size and data type weight, the compression algorithm dynamically selects the most suitable compression strategy according to the characteristics of the input data;
[0029] S2.2. The data encryption unit uses the national cryptographic algorithm, including the SM4 block encryption algorithm, to encrypt the data;
[0030] S2.3. The data verification module combines a verification mechanism with a redundant verification algorithm to verify the transmission data packets and the full incremental data set respectively. The checksum mechanism uses the national cryptographic SM3 hash algorithm to generate the hash value of each data packet for data integrity verification;
[0031] The redundancy check algorithm uses cyclic redundancy check to perform redundancy check on the full - volume incremental data set. The expression is:
[0032] CRC(X) = (X n ·D(X)+H(X)) mod f(G(X),W)
[0033] Among them, CRC(X) is the target redundancy check result, H(X) is the hash value of D(X), which is used to increase the unpredictability of the data; f(G(X),W) is a dynamically generated non - linear generation function, and D(X) is the polynomial corresponding to the data.
[0034] Furthermore, the data transmission module in step S2 supports multi - protocol transmission, including HTTP, HTTPS, MQTT and custom transmission protocols.
[0035] Furthermore, the specific implementation method of step S3 includes the following steps:
[0036] S3.1. Build a conflict detection mechanism for detecting conflicts between incremental data and existing data in the target database. The specific method includes:
[0037] First, according to the primary key value of each record in the incremental data, check whether there is a corresponding record in the target database;
[0038] If there is a record with the same primary key in the target database, calculate the timestamp difference. If the timestamp difference is greater than the preset threshold, then continue to compare the version numbers. If the version number in the incremental data is greater than the version number of the record in the target database, then there is a conflict;
[0039] For the comparison of field contents, an adaptive feature difference measurement algorithm is used to automatically analyze the conflicts of data field contents. The formula is as follows:
[0040]
[0041] Among them, Diff(D 1 ,D 2 ) represents the difference degree between field contents D 1 and D 2 . w i is the dynamic weight of field i, d i (D 1 ,D 2 ) is the content difference measurement of field i between data records D 1 and D 2 . n is the total number of fields;
[0042] By integrating the field content difference measurement and weights, the system judges whether a field has a conflict according to the difference degree score. The formula is as follows:
[0043]
[0044] Among them, ConflictScore i is the difference degree score of field i, Diff j (D 1 , D 2 ) is the difference degree of the j-th feature in field i, w j is the weight of the j-th feature. If ConflictScore i exceeds the set threshold, it is determined as a conflict;
[0045] S3.2. Build a conflict resolution mechanism: Automatically or manually resolve data conflicts through predefined priority rules or manual intervention mechanisms. Among them, the priority rules determine the processing order of conflict records according to timestamps, version numbers, or the update frequency of data. When automatic resolution is not possible, manual intervention is performed.
[0046] Furthermore, the specific implementation method of step S4 includes the following steps:
[0047] S4.1. Build a data packet transmission confirmation mechanism. The receiving end returns a confirmation response after successfully receiving the data packet by using the confirmation response mechanism. After receiving the confirmation response, the sending end considers that the data packet transmission is successful;
[0048] S4.2. Based on the data packet transmission confirmation mechanism in step S4.1, the exception handling module performs transmission failure detection. If the confirmation response is not received within the specified time window, or the confirmation response is incorrect, it is considered that the data packet transmission fails, and the retransmission mechanism is triggered;
[0049] S4.3. According to the detected data packet loss or checksum failure, use a dynamic adaptive retransmission algorithm for data recovery, specifically including:
[0050] First, the system records the index set U of unconfirmed data packets = {u 1 , u 2 ... u n}, and initializes the retry count R k = 0 for each data packet, where k is the data packet number;
[0051] Build a dynamic retransmission decision function:
[0052] T k = α·RTT + β·ER k + γ·L
[0053] Among them, T k is the retransmission priority score of the k-th data packet; RTT is the round-trip delay of the current network; ER kis the historical error rate of the k-th data packet; L is the current network load; α, β, and γ are weighting factors that are dynamically adjusted according to the network state and satisfy α + β + γ = 1.
[0054] Then, sort the unacknowledged data packets in U in descending order according to the T k value, and preferentially retransmit the packet with the largest T k ; if the number of retransmissions R k reaches the maximum retry threshold R max , stop retransmitting, record the unrecovered data packets as the failure set F, and send a notification to the administrator at the same time;
[0055] S4.4. During the data transmission process, if the alarm unit detects transmission failures, excessive retransmission times, or data consistency anomalies, it automatically sends an alarm prompt via SMS, email, or real-time notification.
[0056] S4.5. When the target database update fails, automatically identify and locate the failed operation by querying the incremental data change operations recorded in the transmission log; if the target database update fails, the recovery unit will roll back or reapply the incremental data based on the data records in the log to restore the database to a consistent state and ensure the integrity and accuracy of the data.
[0057] Advantages of the present invention:
[0058] A remote database synchronization system based on incremental data extraction and verification according to the present invention greatly improves the data transmission and synchronization efficiency and adapts to high-frequency data change scenarios through incremental extraction, compressed transmission, and batch writing technologies.
[0059] A remote database synchronization system based on incremental data extraction and verification according to the present invention guarantees the integrity and accuracy of data synchronization through a multi-level data verification mechanism and an exception recovery strategy.
[0060] A remote database synchronization system based on incremental data extraction and verification according to the present invention meets the requirements for low-latency transmission through real-time capture and synchronization of incremental data.
[0061] A remote database synchronization system based on incremental data extraction and verification according to the present invention supports multi-protocol adaptation and dynamic adjustment strategies and is applicable to complex distributed environments and cross-regional data synchronization scenarios.
[0062] A remote database synchronization system based on incremental data extraction and verification according to the present invention adopts domestic encryption transmission technology to ensure the confidentiality and anti-tampering ability of data during transmission.
[0063] A remote database synchronization system based on incremental data extraction and verification according to the present invention can effectively solve the core problems in the existing database synchronization technology, is applicable to application fields with high requirements for data synchronization such as financial transactions, logistics tracking, and industrial Internet of Things, and has broad practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic structural diagram of a remote database synchronization system based on incremental data extraction and verification according to the present invention;
[0065] Figure 2 It is a flowchart of a remote database synchronization method based on incremental data extraction and verification according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0067] Therefore, the detailed description of the specific embodiments of the present invention provided in the drawings below is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0068] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and are accompanied by the attached Figure 1 - Attached Figure 2 The details are as follows:
[0069] Embodiment 1:
[0070] A remote database synchronization system based on incremental data extraction and verification includes a variable detection module 1, a data application module 2, a data transmission module 3, a data verification module 4, and an exception handling module 5;
[0071] The variable detection module 1 and the data application module 2 are respectively connected to the data transmission module 3, the data transmission module 3 is respectively connected to the data verification module 4 and the exception handling module 5, and the exception handling module 5 is also connected to the data verification module 4;
[0072] The source database 6 is connected to the variable detection module 1, and the target database 7 is respectively connected to the data application module 2 and the data transmission module 3;
[0073] The variable detection module 1 is used to monitor the new, updated or deleted operations in the source database in real time and generate an incremental data change set;
[0074] The data transmission module 3 is used to transmit the incremental data change set generated by the variable detection module 1 to the target database 7 in batches through the network;
[0075] The data verification module 4 is used to verify the integrity and accuracy of the incremental data in the incremental data change set during and after the transmission;
[0076] The data application module 2 is used to parse, apply the incremental data in the incremental data change set transmitted to the target database 7, and resolve data conflicts;
[0077] The exception handling module 5 is connected to the data transmission module 3 and the data verification module 4, and is used to trigger a retry mechanism or an alarm mechanism when network exceptions, data loss or verification failures are detected.
[0078] Furthermore, the data transmission module 3 includes a data compression unit and a data encryption unit.
[0079] Furthermore, the exception handling module 5 includes an alarm unit and a recovery unit.
[0080] Embodiment 2:
[0081] A remote database synchronization method based on incremental data extraction and verification is realized relying on the remote database synchronization system based on incremental data extraction and verification described in Embodiment 1, and is characterized by including the following steps:
[0082] S1. Source database connection change detection module, the change detection module captures the new, updated and deleted operations of the source database in real time, extracts the table name, primary key, field content, operation type and timestamp of the changed data by parsing the transaction log or using the trigger mechanism, and the changed data is encapsulated into an incremental data change set DCS and stored in a queue for buffering;
[0083] S2. The data transmission module divides the incremental data in the incremental data change set obtained in step S1 into blocks. Before transmission, the data compression unit uses the compression algorithm Gzip+ to compress the incremental data, and then the data encryption unit encrypts the incremental data through the encryption algorithm, and then transmits the encrypted incremental data to the target database;
[0084] During the data transmission process, the data verification module uses the national cryptography SM3 hash algorithm to generate a verification code for the transmitted data block for data integrity verification, and uses a redundancy verification algorithm to perform redundancy verification on all incremental data after the data transmission is completed;
[0085] Further, the specific implementation method of step S2 includes the following steps:
[0086] S2.1. The data compression unit uses the optimized compression algorithm Gzip+ to compress the incremental data. The compression ratio CR is calculated using the following formula:
[0087]
[0088] where, H(d i ) and H(d' j ) are the entropy values of the original data block and the compressed data block respectively, f(α,β,W i ) is an adaptive compression function, which depends on the data block size and data type weight. The first parameter α is used to adjust the sliding window size, the second parameter β is used to select different coding strategies, and W i is the weight of the i-th data block. During the compression process, the weight of the i-th data block is automatically set according to the size and type of the data block, indicating the relative importance of each data block in the compression ratio calculation, so as to ensure that each data block is given an appropriate proportion in the compression algorithm. R is the amount of redundant data, which is expressed as the redundant information added during the compression process, and λ is the weight coefficient of the redundant data, which is used to balance the compression effect and the redundant capacity;
[0089] By adaptively adjusting the sliding window size and data type weight, the compression algorithm dynamically selects the most suitable compression strategy according to the characteristics of the input data;
[0090] S2.2. The data encryption unit uses national cryptography algorithms, including the SM4 block encryption algorithm, to encrypt the data;
[0091] S2.3. The data verification module combines a verification mechanism with a redundancy verification algorithm to verify the transmitted data packets and the full incremental data set respectively; the checksum mechanism uses the national cryptography SM3 hash algorithm to generate the hash value of each data packet for data integrity verification;
[0092] The redundancy verification algorithm uses cyclic redundancy check to perform redundancy verification on the full incremental data set. The expression is:
[0093] CRC(X)=(X n ·D(X)+H(X))modf(G(X),W)
[0094] Among them, CRC(X) is the redundancy check result of the target, H(X) is the hash value of D(X), which is used to increase the unpredictability of the data; f(G(X), W) is a dynamically generated non-linear generation function, and D(X) is the polynomial corresponding to the data.
[0095] Furthermore, the data transmission module in step S2 supports multi-protocol transmission, including HTTP, HTTPS, MQTT, and custom transmission protocols.
[0096] S3. After the target database receives and verifies, the data application module parses the incremental data and synchronizes it to the target database, calls the database instruction according to the operation type in the incremental data to complete the synchronization, and constructs a conflict detection mechanism and a conflict resolution mechanism to handle data conflicts;
[0097] Furthermore, the specific implementation method of step S3 includes the following steps:
[0098] S3.1. Construct a conflict detection mechanism for detecting conflicts between the incremental data and the existing data in the target database. The specific method includes:
[0099] First, according to the primary key value of each record in the incremental data, check whether there is a corresponding record in the target database;
[0100] If there is a record with the same primary key in the target database, calculate the timestamp difference. If the timestamp difference is greater than the preset threshold, continue to compare the version numbers. If the version number in the incremental data is greater than the version number of the record in the target database, there is a conflict;
[0101] For the comparison of field contents, an adaptive feature difference measurement algorithm is used to automatically analyze the conflicts of data field contents. The formula is as follows:
[0102]
[0103] Among them, Diff(D 1 ,D 2 ) represents the difference degree between the field contents D 1 and D 2 , w i is the dynamic weight of field i, d i (D 1 ,D 2 ) is the content difference measurement of field i between the data records D 1 and D 2 , and n is the total number of fields;
[0104] By integrating the field content difference measurement and the weight, the system judges whether the field has a conflict according to the difference degree score. The formula is as follows:
[0105]
[0106] Among them, ConflictScore i is the difference score of field i, Diff j (D 1 , D 2 ) is the difference of the j-th feature in field i, w j is the weight of the j-th feature. If ConflictScore i exceeds the set threshold, it is considered a conflict;
[0107] S3.2. Build a conflict resolution mechanism: Automatically or manually resolve data conflicts through predefined priority rules or manual intervention mechanisms. Among them, the priority rules determine the processing order of conflict records according to timestamps, version numbers, or the update frequency of data. When automatic resolution is not possible, manual intervention is performed.
[0108] S4. During data transmission and data verification, the exception handling module processes network interruptions, data loss, or verification failures. And the alarm unit of the exception handling module sends alarm prompts through emails, text messages, or real-time notifications. And the recovery unit of the exception handling module rolls back or reapplies incremental data according to the transmission log when the target database is abnormal or the update fails.
[0109] Furthermore, to improve synchronization performance, the data application module supports batch processing operations, writes incremental data into the database in batches, and dynamically adjusts the batch size according to the I / O performance of the target database.
[0110] Furthermore, the specific implementation method of step S4 includes the following steps:
[0111] S4.1. Build a data packet transmission confirmation mechanism. By using the confirmation response mechanism, the receiving end returns a confirmation response after successfully receiving the data packet. After the sending end receives the confirmation response, it considers that the data packet transmission is successful;
[0112] S4.2. Based on the data packet transmission confirmation mechanism in step S4.1, the exception handling module performs transmission failure detection. If no confirmation response is received within the specified time window, or the confirmation response is incorrect, it is considered that the data packet transmission fails, and the retransmission mechanism is triggered;
[0113] S4.3. According to the detected data packet loss or verification failure, use a dynamic adaptive retransmission algorithm for data recovery, specifically including:
[0114] First, the system records the index set U = {u 1 , u 2 ... u n} of unacknowledged data packets, and initializes the retry count R of each data packetk = 0, where k is the data packet number;
[0115] Construct a dynamic retransmission decision function:
[0116] T k = α·RTT + β·ER k + γ·L
[0117] where, T k is the retransmission priority score of the k-th data packet; RTT is the round-trip delay of the current network; ER k is the historical error rate of the k-th data packet; L is the current network load; α, β, γ are weight factors, dynamically adjusted according to the network state, and satisfy α + β + γ = 1.
[0118] Then, sort the unacknowledged data packets in U according to the T k value from high to low, and preferentially retransmit the packet with the largest T k ; if the number of retransmissions R k reaches the maximum retry threshold R max , stop retransmitting, record the unrecovered data packets as the failure set F, and send a notification to the administrator at the same time;
[0119] Furthermore, by dynamically adjusting the weight factors, this algorithm flexibly allocates retransmission resources according to the real-time network conditions, significantly improving the stability and efficiency of data transmission.
[0120] Furthermore, the retransmission start based on the conflict detection mechanism: When it is detected that the data in the target database fails to be successfully synchronized or an exception occurs (such as inconsistent timestamps, changes in field content), the retransmission mechanism is triggered, the incremental data is resent, and the data consistency check is performed again. Retransmission strategy optimization: Adopt a bandwidth adaptation strategy and transmission rate control, dynamically adjust the retransmission frequency and the priority of retransmitted data according to the network conditions (for example, preferentially retransmit unacknowledged critical data packets or incremental data packets), and optimize the bandwidth utilization and transmission efficiency of the system.
[0121] S4.4. During the data transmission process, if the alarm unit detects transmission failures, excessive retransmission times, or data consistency anomalies, it automatically sends alarm prompts through text messages, emails, or real-time notifications;
[0122] S4.5. When the update of the target database fails, the failed operations are automatically identified and located by querying the incremental data change operations recorded in the transmission log; if the update of the target database fails, the recovery unit will roll back or reapply the incremental data according to the data records in the log, restoring the database to a consistent state to ensure the integrity and accuracy of the data.
[0123] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0124] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any manner, and the fact that the combinations thereof are not exhaustively described in this specification is only for the consideration of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A remote database synchronization system based on incremental data extraction and verification, characterized in that: It comprises a variable detection module (1), a data application module (2), a data transmission module (3), a data verification module (4), and an exception handling module (5); The variable detection module (1) and the data application module (2) are respectively connected to the data transmission module (3); the data transmission module (3) is respectively connected to the data verification module (4) and the exception processing module (5); the exception processing module (5) is also connected to the data verification module (4); The source database (6) is connected to the variable detection module (1), and the target database (7) is connected to the data application module (2) and the data transmission module (3) respectively; The variable detection module (1) is used to monitor the addition, update or deletion operations in the source database in real time and generate an incremental data change set; The data transmission module (3) is used to transmit the incremental data change set generated by the variable detection module (1) to the target database (7) via the network in batches; The data verification module (4) is used to verify the integrity and accuracy of the incremental data of the incremental data change set during and after the transmission; The data application module (2) is used to parse and apply the incremental data of the incremental data change set transmitted to the target database (7), and resolve data conflicts; The exception handling module (5) is connected to the data transmission module (3) and the data verification module (4) and is used to trigger a retry mechanism or an alarm mechanism when a network anomaly, data loss or verification failure is detected.
2. A remote database synchronization system based on incremental data extraction and verification according to claim 1, characterized in that: The data transmission module (3) comprises a data compression unit and a data encryption unit.
3. A remote database synchronization system based on incremental data extraction and verification according to claim 2, characterized in that: The exception handling module (5) comprises an alarm unit and a recovery unit.
4. A remote database synchronization method based on incremental data extraction and verification, implemented by a remote database synchronization system based on incremental data extraction and verification as claimed in any one of claims 1 to 3, characterized in that: The steps include: S1. The source database is connected to the change detection module. The change detection module captures the addition, update and deletion operations of the source database in real time. By parsing the transaction log or using the trigger mechanism, the table name, primary key, field content, operation type and timestamp of the changed data are extracted. The changed data is encapsulated into an incremental data change set DCS and stored in the queue for buffering; S2. The data transmission module divides the incremental data in the incremental data change set obtained in step S1 into blocks, and the data compression unit compresses the incremental data using the compression algorithm Gzip+ before transmission, and then the data encryption unit encrypts the incremental data using the encryption algorithm, and then transmits the encrypted incremental data to the target database; The data verification module uses the national secret SM3 hash algorithm to generate a check code for the transmission data block to verify data integrity during the data transmission process, and uses a redundant check algorithm to perform redundant check on all incremental data after the data transmission is completed; S3. After the target database receives and verifies the data, the data application module parses the incremental data and synchronizes it to the target database. It calls the database command to complete the synchronization according to the operation type in the incremental data, and builds a conflict detection mechanism and a conflict resolution mechanism to handle data conflicts. S4. During the data transmission and data verification process, the exception handling module handles network interruption, data loss or verification failure, and the alarm unit of the exception handling module sends alarm prompts via email, SMS or real-time notification, and the recovery unit of the exception handling module rolls back or reapplies incremental data according to the transmission log when the target database is abnormal or the update fails.
5. A remote database synchronization method based on incremental data extraction and verification according to claim 4, characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. The data compression unit uses the optimized compression algorithm Gzip+ to compress the incremental data. The compression ratio CR is calculated using the following formula: Among them, H(d i ) and H(d' j ) are the entropy values of the original data block and the compressed data block, respectively, f(α, β, W i ) is an adaptive compression function that depends on the data block size and data type weight. The first parameter α is used to adjust the sliding window size, and the second parameter β is used to select different encoding strategies. i is the weight of the ith data block. During the compression process, the weight of the ith data block is automatically set according to the size and type of the data block, indicating the relative importance of each data block in the compression ratio calculation, thereby ensuring that each data block is given an appropriate weight in the compression algorithm. R is the amount of redundant data, which is expressed as the redundant information added during the compression process. λ is the weight coefficient of redundant data, which is used to balance the compression effect and redundant capacity. By adaptively adjusting the sliding window size and data type weight, the compression algorithm dynamically selects the most suitable compression strategy based on the characteristics of the input data; S2.
2. The data encryption unit uses national secret algorithms, including the SM4 block encryption algorithm, to encrypt data; S2.
3. The data verification module uses a combination of a verification mechanism and a redundant verification algorithm to verify the transmission data packet and the full incremental data set respectively; the checksum mechanism uses the national secret SM3 hash algorithm to generate a hash value for each data packet for data integrity verification; The redundancy check algorithm uses cyclic redundancy check to perform redundancy check on the full incremental data set. The expression is: CRC(X)=(X n ·D(X)+H(X))modf(G(X),W) Among them, CRC(X) is the redundant check result of the target, H(X) is the hash value of D(X), which is used to increase the unpredictability of the data; f(G(X), W) is a dynamically generated nonlinear generating function, and D(X) is the polynomial corresponding to the data.
6. A remote database synchronization method based on incremental data extraction and verification according to claim 5, characterized in that: The data transmission module of step S2 supports multi-protocol transmission, including HTTP, HTTPS, MQTT and custom transmission protocols.
7. A remote database synchronization method based on incremental data extraction and verification according to claim 6, characterized in that: The specific implementation method of step S3 includes the following steps: S3.
1. Construct a conflict detection mechanism to detect conflicts between incremental data and existing data in the target database. The specific methods include: First, based on the primary key value of each record in the incremental data, find out whether there is a corresponding record in the target database; If there are records with the same primary key in the target database, the timestamp difference is calculated. If the timestamp difference is greater than the preset threshold, the version numbers are compared. If the version number in the incremental data is greater than the version number recorded in the target database, there is a conflict. For the comparison of field contents, an adaptive feature difference measurement algorithm is used to automatically analyze the conflict of data field contents. The formula is as follows: Among them, Diff(D1,D2) represents the difference between the field contents D1 and D2, w i is the dynamic weight of field i, d i (D1, D2) is the content difference measure of field i between data records D1 and D2, and n is the total number of fields; By combining the difference measurement and weight of field contents, the system determines whether a field conflicts based on the difference score. The formula is as follows: Among them, ConflictScore i Score the difference of field i, Diff j (D1, D2) is the difference of the jth feature in field i, w j is the weight of the jth feature, if ConflictScore i If the set threshold is exceeded, it is considered a conflict; S3.
2. Build a conflict resolution mechanism: automatically or manually resolve data conflicts through predefined priority rules or manual intervention mechanisms. The priority rules determine the processing order of conflicting records based on timestamps, version numbers, or data update frequency. When conflicts cannot be resolved automatically, manual intervention is performed.
8. A remote database synchronization method based on incremental data extraction and verification according to claim 7, characterized in that: The specific implementation method of step S4 includes the following steps: S4.
1. Construct a data packet transmission confirmation mechanism. By using the confirmation response mechanism, the receiving end returns a confirmation response after successfully receiving the data packet. After receiving the confirmation response, the sending end considers that the data packet has been successfully transmitted; S4.
2. Based on the data packet transmission confirmation mechanism in step S4.1, the exception handling module performs transmission failure detection. If no confirmation response is received within the specified time window, or the confirmation response is incorrect, the data packet transmission is considered to have failed, triggering the retransmission mechanism; S4.
3. Based on the detected data packet loss or verification failure, a dynamic adaptive retransmission algorithm is used to recover data, including: First, the system records the index set U of the unconfirmed data packets = {u1,u2...u n }, and initialize the number of retries R for each data packet k =0, where k is the data packet number; Construct a dynamic retransmission decision function: T k =α·RTT+β·ER k +γ·L in, T k is the retransmission priority score of the kth data packet; RTT is the round-trip delay of the current network; ER k is the historical error rate of the kth data packet; L is the current network load; α, β, γ are weight factors, which are dynamically adjusted according to the network status to satisfy α+β+γ=1. Then, press T for the unacknowledged packets in U k The values are sorted from high to low and retransmitted first. k The largest packet; if the number of retransmissions is R k Reached the maximum retry threshold R max , stop retransmission, and record the unrecovered data packets as the failure set F, and send a notification to the administrator; S4.
4. During the data transmission process, if the alarm unit detects a transmission failure, excessive retransmission times, or abnormal data consistency, it will automatically send an alarm prompt via SMS, email or real-time notification; S4.
5. When the target database fails to be updated, the failed operation is automatically identified and located by querying the incremental data change operations recorded in the transmission log; if the target database fails to be updated, the recovery unit will roll back or reapply the incremental data based on the data records in the log, restore the database to a consistent state, and ensure the integrity and accuracy of the data.
Citation Information
Patent Citations
Extensive makeup language (XML)-based method for synchronously updating increment of spatial data
CN102508886A
Data synchronization increment tracking method and system
CN105938492A
Detection data platform data synchronization method and system
CN118708652A
Distributed database data synchronization method, system, device and medium
CN118861160A
Method, system and device for retransmission and error correction of communication packet loss and storage medium
CN119109559A
Cited By
Heterogeneous database structure dynamic synchronization and fault-tolerant migration method, system and equipment
CN121029719A
Data integrity transmission system based on USB high-batch reading and writing
CN122240542A
A data integrity transmission system based on high-volume USB read / write
CN122240542B