A distributed database data synchronization system and data consistency detection method
Through multiple data synchronization tests and version sequence analysis, combined with dynamic threshold adjustment and modular design, the problem of data inconsistency in distributed databases is solved, fast and accurate data consistency detection and automated management are achieved, and the operation and maintenance efficiency and reliability of the system are improved.
Patent Information
- Application Number
- CN202510757942.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing distributed databases are susceptible to network delays and node failures during data synchronization, resulting in data inconsistency. They lack automated and intelligent cross-cluster data consistency management solutions, making it difficult to meet the efficient operation and maintenance needs of large-scale distributed database systems.
Through multiple data synchronization tests, the data version sequence collection is generated, the version change range is calculated to determine the node synchronization consistency, abnormal nodes are located and isolated, dynamic threshold adjustment and multiple calculation methods are used to realize cross-cluster data consistency calibration, and automated management is carried out through a modularly designed data synchronization system.
Quickly and accurately detect data consistency problems, automatically locate and isolate abnormal nodes, improve operation and maintenance efficiency, reduce business interruption time, realize cross-cluster data consistency management, and improve system reliability and availability.
Smart Images

Figure CN120276999B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of databases, and in particular to a distributed database data synchronization system and a data consistency detection method. Background Art
[0002] With the rapid development of information technology, data volumes are exploding. Traditional centralized databases are no longer able to meet the demands of large-scale data storage and processing. Distributed databases, by distributing data across multiple nodes, can effectively improve storage and processing capabilities, meeting enterprises' requirements for high-concurrency data access and massive storage.
[0003] In a distributed database environment, data synchronization is a critical step in ensuring data consistency. Currently, common data synchronization technologies include log-based synchronization and message queue-based synchronization. However, these technologies still face numerous challenges in practical application. First, in complex network environments, the data synchronization process is susceptible to factors such as network latency and packet loss, leading to delayed or failed data synchronization and, in turn, data inconsistencies. For example, in a distributed database cluster across multiple locations, data synchronization delays can reach hundreds of milliseconds or even seconds due to long network distances, resulting in significant discrepancies between data versions on different nodes. Furthermore, when some nodes fail, existing data consistency detection methods often struggle to quickly and accurately locate the faulty node. Furthermore, when addressing data inconsistencies caused by node failures, repair efficiency is low, which can easily lead to data loss or business interruption. Furthermore, for distributed databases with multi-cluster architectures, coordinating data consistency across multiple clusters remains a major challenge. Existing methods for cross-cluster data synchronization and consistency detection often require manual intervention and lack automated and intelligent solutions, making them incapable of meeting the requirements for efficient operation and maintenance of large-scale distributed database systems. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting data consistency in a distributed database, comprising the following steps:
[0005] Step 1: Perform multiple data synchronization tests on multiple data nodes in a distributed database to generate a set of data version sequences corresponding to each set of tests; the data nodes belong to a data cluster, each data cluster contains N data nodes, and the data version sequence includes the data version number and synchronization timestamp of each data node;
[0006] Step 2: Based on multiple sets of data version sequence sets, extract the independent version sequence of each data node and calculate the version change range of each data node; determine the synchronization consistency of the data nodes based on the version change range. If all data nodes are synchronized, proceed to step 3; otherwise, isolate the inconsistent data nodes and proceed to step 3;
[0007] Step 3: Sort the synchronized data nodes by the preset identifier to generate a sorted version sequence set; calculate the overall cluster version consistency based on this set. If the consistency meets the threshold, proceed to step 6; otherwise, proceed to step 4.
[0008] Step 4: Locate and isolate the data node with the abnormal version based on the version difference of each node in the sorted version sequence set; if the database has a single-cluster architecture, proceed to step 6; if it has a multi-cluster architecture, proceed to step 5;
[0009] Step 5: Count the number of isolated nodes in each cluster and calculate the storage capacity loss rate of the corresponding cluster. If the loss rate of all clusters does not exceed the set threshold, perform cross-cluster data consistency calibration and proceed to step 6. Otherwise, trigger a data integrity warning.
[0010] Step 6: Based on the calibrated effective storage capacity of each cluster, generate a distributed database global consistency report to complete data synchronization management.
[0011] Furthermore, the data cluster includes multiple logical shards, each logical shard consists of multiple data copies; the data synchronization test includes cross-shard transaction submission and copy synchronization verification.
[0012] Furthermore, the version change range is:
[0013] Extract the minimum version number V_min and the maximum version number V_max in the data node version sequence. If (V_max - V_min) ≤ the preset version tolerance threshold, the node is considered synchronized; otherwise, it is considered inconsistent.
[0014] Furthermore, the overall version consistency of the computing cluster is:
[0015] Calculate the maximum Δ_max and minimum Δ_min of the version change range of all data nodes. If (Δ_max - Δ_min) is less than the cluster difference threshold, the cluster consistency is determined to be met.
[0016] Furthermore, the positioning of the version abnormal node includes:
[0017] Calculate the average Δ_avg of the version differences between each pair of data nodes in the cluster. If the absolute deviation between the version change range of a node and Δ_avg exceeds a dynamic floating threshold, the node is determined to be an abnormal node; the dynamic floating threshold is dynamically adjusted according to the historical synchronization volatility.
[0018] Furthermore, the storage capacity loss rate is:
[0019] For each cluster, the ratio of the number of isolated nodes to the total number of initial nodes is used as the capacity loss rate, and the replica loss weight coefficient is superimposed for correction.
[0020] Furthermore, the cross-cluster data consistency calibration includes:
[0021] The cluster with the largest capacity loss rate is selected as the benchmark cluster, and the effective storage capacity of the remaining clusters is adjusted to match the benchmark cluster through dynamic shard migration or replica expansion and contraction.
[0022] Furthermore, after the data integrity warning is triggered, the following steps are performed:
[0023] a) Automatically initiate the data recovery process and rebuild lost data based on geo-redundant copies;
[0024] b) Mark the abnormal cluster as read-only and generate an operation and maintenance work order;
[0025] c) Trigger data redistribution across availability zones based on preset policies.
[0026] A distributed database data synchronization system, applying the distributed database data consistency detection method, comprises: a data synchronization management module, a consistency detection engine, an exception handling module, an early warning decision module, a metadata repository and a communication module;
[0027] The data synchronization management module, consistency detection engine, exception handling module, early warning decision module, and metadata repository are respectively connected to the communication module for communication;
[0028] The data synchronization management module is used to control multi-node concurrent reading and writing and version number generation;
[0029] The consistency detection engine is used to analyze the version sequence in real time and calculate the consistency index;
[0030] The exception handling module: performs node isolation, data repair and capacity calibration operations;
[0031] The early warning decision module: dynamically triggers a multi-level early warning strategy according to the capacity loss rate;
[0032] The metadata repository records node topology relationships, version history, and calibration logs.
[0033] The beneficial effects of this invention include: Through multiple data synchronization tests and comprehensive version sequence analysis, data consistency issues in data nodes and clusters can be quickly and accurately detected. Dynamic threshold adjustment and multiple calculation methods are used to improve detection accuracy and adaptability, meeting the needs of different business scenarios and system scales.
[0034] The system can automatically locate and isolate abnormal data nodes and perform corresponding exception handling operations based on different situations, such as data repair and capacity calibration. Without manual intervention, it greatly improves the system's operation and maintenance efficiency and reduces business interruption time caused by data inconsistencies.
[0035] A comprehensive cross-cluster data consistency management solution is proposed for distributed databases with multi-cluster architectures. By calculating storage capacity loss rates and performing cross-cluster data consistency calibration, it effectively balances storage loads across clusters, improving data consistency and availability across the entire distributed database system.
[0036] The early warning decision module dynamically triggers a multi-level early warning strategy based on the capacity loss rate and automatically executes emergency response measures in severe cases, such as data repair and marking the cluster as read-only. This enables timely detection and resolution of data integrity risks, ensuring stable system operation and data security.
[0037] The distributed database data synchronization system adopts a modular design, with each module clearly divided into different responsibilities and working in collaboration. The data synchronization management module, consistency detection engine, and exception handling module work together to achieve automated management of the entire process, including data synchronization, consistency detection, and exception handling, improving the system's scalability and maintainability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The figure is a flowchart of a distributed database data consistency detection method. DETAILED DESCRIPTION
[0039] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0040] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0041] like Figure 1 As shown, a distributed database data consistency detection method includes the following steps:
[0042] Step 1: Data synchronization test and data version sequence set generation
[0043] In a distributed database, data nodes belong to data clusters. Each data cluster contains N data nodes, and each data cluster contains multiple logical shards, each of which consists of multiple data replicas. Data synchronization testing is the foundation of the entire testing method. Its core purpose is to obtain data version changes on each data node under different test scenarios.
[0044] Data synchronization testing involves cross-shard transaction submission and replica synchronization verification. During cross-shard transaction submission, the system simulates the complex transaction operations of real-world business scenarios, submitting data modifications across multiple logical shards as a single transaction. For example, in an e-commerce order processing system, a single order might involve modifications to data across multiple logical shards, including product inventory information (stored on one logical shard) and user account information (stored on another). Using a distributed transaction processing mechanism, the system ensures that all cross-shard data modifications are either successfully submitted or rolled back, guaranteeing data atomicity and consistency.
[0045] During the replica synchronization verification phase, the system monitors data synchronization between replicas within each logical shard in real time. Whenever data on the master replica is updated, the system immediately synchronizes the update to each slave replica and records the synchronization timestamp. To ensure synchronization accuracy and reliability, the system employs various verification mechanisms. For example, by comparing data checksums (such as those generated by hash algorithms like MD5 and SHA-256) between the master and slave replicas, the system determines whether data errors or loss occurred during synchronization. Furthermore, the system monitors synchronization delays between replicas and triggers alerts when these delays exceed preset thresholds.
[0046] By executing the above data synchronization tests multiple times, the system generates a corresponding set of data version sequences for each test. The data version sequence contains the data version number and synchronization timestamp for each data node. The data version number uniquely identifies the data update status and increments according to a specific rule with each data update. The synchronization timestamp records the specific time when the data update occurred, accurate to the millisecond level. For example, a data version sequence can be represented as {(node 1, version number 1, timestamp 1), (node 2, version number 2, timestamp 2), ..., (node N, version number N, timestamp N)}. These data version sequence sets provide rich raw data for subsequent data consistency analysis.
[0047] (2) Step 2: Data Node Synchronization Consistency Judgment
[0048] Based on the multiple sets of data version sequences generated in step 1, the system extracts an independent version sequence for each data node. An independent version sequence integrates the data version numbers and synchronization timestamps of the same data node in different test groups to form a complete data version change record for that node.
[0049] To determine the synchronization consistency of data nodes, the system calculates the version range of each data node. The version range is calculated by extracting the minimum version number V_min and the maximum version number V_max in the data node version sequence. If (V_max - V_min) ≤ the preset version tolerance threshold, the node is considered synchronized; otherwise, it is considered inconsistent. The preset version tolerance threshold is a parameter pre-set based on actual business needs and system performance requirements. It reflects the degree of data version discrepancy that the system can tolerate. For example, in financial trading systems with extremely high data consistency requirements, the preset version tolerance threshold may be set to 0 or 1, meaning that version discrepancies between data nodes are not allowed. In contrast, in some logging systems with relatively low real-time requirements, the preset version tolerance threshold can be appropriately relaxed.
[0050] When an inconsistent data node is detected, the system immediately isolates it. This isolation modifies the data node's network configuration or access permissions, temporarily preventing it from participating in data synchronization and business data read and write operations. This prevents the inconsistent data node from impacting other healthy nodes and prevents the spread of data inconsistencies. Only when all data nodes are synchronized and consistent will the system proceed to step three.
[0051] (III) Step 3: Calculating the overall cluster version consistency
[0052] Consistently synchronized data nodes are sorted by a preset identifier, which can be a unique identifier such as the node's IP address or node number. The purpose of sorting is to unify the processing order of data nodes, facilitating subsequent calculations and analysis. After generating a sorted set of version sequences, the system calculates the overall version consistency of the cluster based on this set.
[0053] The method for calculating overall cluster version consistency is to calculate the maximum Δ_max and minimum Δ_min of the version change range of all data nodes. If (Δ_max - Δ_min) is less than the cluster difference threshold, the cluster is considered consistent. The cluster difference threshold is a parameter set based on factors such as cluster scale and business needs. It measures the degree of variation in the version change range of data nodes within the cluster. For example, for small clusters with infrequent data updates, the cluster difference threshold can be set relatively loosely. However, for large-scale, highly concurrent clusters, a stricter cluster difference threshold is required to ensure high data consistency.
[0054] If the cluster consistency meets the threshold requirements, proceed to step 6; otherwise, it means that there are nodes with large data version differences in the cluster, and the system needs to proceed to step 4 to locate and isolate these abnormal nodes.
[0055] (IV) Step 4: Locating and isolating nodes with version anomalies
[0056] Based on the version differences between each node in the sorted version sequence set, the system locates and isolates data nodes with version anomalies. This is done by calculating the average Δ_avg of the version differences between each pair of data nodes in the cluster. If the absolute deviation between a node's version range and Δ_avg exceeds a dynamic floating threshold, the node is identified as an anomaly.
[0057] The dynamic floating threshold is adjusted dynamically based on the historical synchronization volatility. The historical synchronization volatility is derived from a statistical analysis of data node synchronization over a period of time and reflects the degree of fluctuation in the data synchronization process. When the historical synchronization volatility is high, indicating an unstable data synchronization process, the dynamic floating threshold is increased accordingly to avoid misidentifying normal nodes as abnormal. Conversely, when the historical synchronization volatility is low, the dynamic floating threshold is decreased to improve the accuracy of abnormal node detection.
[0058] Once the node with the abnormal version is located, the system immediately performs an isolation operation, similar to the isolation operation in step 2. If the database has a single-cluster architecture, after isolating the abnormal node, the system directly proceeds to step 6; if it has a multi-cluster architecture, it proceeds to step 5 to further address cross-cluster data consistency issues.
[0059] (V) Step 5: Cross-cluster data consistency calibration and early warning
[0060] In a distributed database with a multi-cluster architecture, the system counts the number of isolated nodes in each cluster and calculates the storage capacity loss rate for that cluster. This capacity loss rate is calculated by taking the ratio of the number of isolated nodes to the initial total number of nodes for each cluster and applying a correction factor based on the loss of replicas. The loss of replicas is a parameter set based on the importance and redundancy of data replicas to more accurately reflect actual storage capacity loss. For example, a higher loss of replicas is set for replicas of critical business data to emphasize the impact of their loss on storage capacity.
[0061] If the loss rate for all clusters does not exceed the set threshold, indicating that the system still has a certain degree of fault tolerance, a cross-cluster data consistency calibration operation is performed. Cross-cluster data consistency calibration involves selecting the cluster with the highest capacity loss rate as the baseline cluster and adjusting the effective storage capacity of the remaining clusters to match the baseline cluster through dynamic shard migration or replica scaling.
[0062] Dynamic shard migration migrates some logical shards from clusters with low storage capacity utilization to the baseline cluster to balance the storage load across clusters. During the migration process, the system ensures data integrity and consistency, and uses a distributed transaction processing mechanism to guarantee the atomicity of the shard migration operation. Replica scaling increases or decreases data replicas in other clusters based on the storage capacity of the baseline cluster. For example, when the baseline cluster has sufficient storage capacity, the number of data replicas in other clusters can be appropriately increased to improve data redundancy and availability. Conversely, when the baseline cluster's storage capacity is limited, the number of replicas in other clusters can be reduced to free up storage resources.
[0063] If the loss rate of a cluster exceeds the set threshold, it means that the system's data integrity is at great risk, and a data integrity warning is triggered.
[0064] (6) Step 6: Generate a global consistency report for the distributed database
[0065] Based on the calibrated effective storage capacity of each cluster, the system generates a global consistency report for the distributed database. This report details the data consistency status of each data cluster, information about isolated nodes, storage capacity loss, and the data consistency calibration process and results. This report not only provides system administrators with comprehensive data consistency information, facilitating system maintenance and troubleshooting, but also serves as an important reference for subsequent system optimization and improvement. This completes the entire data synchronization management process.
[0066] A distributed database data synchronization system applies the above-mentioned distributed database data consistency detection method. The system mainly includes a data synchronization management module, a consistency detection engine, an exception handling module, an early warning decision module, a metadata repository and a communication module.
[0067] The data synchronization management module controls multi-node concurrent read and write operations and version number generation. To manage multi-node concurrent read and write operations, the module employs advanced concurrency control algorithms, such as the two-phase locking protocol (2PL) and optimistic locking protocol, to ensure that data conflicts and inconsistencies do not occur when multiple nodes simultaneously read and write data. For example, when using the two-phase locking protocol, the data synchronization management module locks the relevant data at the beginning of a transaction to prevent other transactions from modifying the data simultaneously. The corresponding lock is released after the transaction is committed or rolled back.
[0068] The data synchronization management module uses an efficient version number generation algorithm. This algorithm combines information such as timestamps and node identifiers to generate unique and sequential data version numbers. For example, a version number can consist of a timestamp (accurate to the millisecond), a node ID, and an increasing sequence number. This ensures that each data update operation in a distributed environment generates a unique version number, facilitating subsequent data consistency testing and tracking.
[0069] The consistency check engine analyzes version sequences and calculates consistency metrics in real time. It communicates with the data synchronization management module and the metadata repository to obtain data version sequences and related metadata. During this real-time version sequence analysis, the consistency check engine utilizes efficient data processing algorithms, such as sliding window algorithms and incremental calculation algorithms, to rapidly process and analyze large amounts of version sequence data.
[0070] For example, the sliding window algorithm can count and analyze version sequences within a certain timeframe, promptly identifying trends and anomalies in data version changes. The incremental calculation algorithm, when a data version changes, only calculates consistency metrics for the changed portion, significantly improving computational efficiency. Through these algorithms, the consistency check engine can accurately calculate metrics such as the version change range of data nodes and overall cluster version consistency, providing a reliable basis for determining data consistency.
[0071] The exception handling module performs node isolation, data repair, and capacity calibration. To isolate nodes, it collaborates with the system's network management module to isolate abnormal data nodes by modifying network routing rules and shutting down the node's network ports. To ensure the safety and reliability of isolation operations, the exception handling module performs a series of checks and verifications before executing the isolation operation, such as confirming node status and backing up node data.
[0072] When data loss or corruption is detected, the exception handling module reconstructs the lost data based on geo-redundant replicas. The system pre-configures multiple redundant replicas in different geographic locations to improve data reliability and availability. The exception handling module retrieves the latest and most complete data from the geo-redundant replica based on the data's version information and timestamp, and restores it to the corresponding data node.
[0073] During capacity calibration, the exception handling module performs dynamic shard migration and replica scaling based on the calibration strategy in step 5. During this process, the exception handling module monitors the progress of data migration and replica adjustments, promptly addressing any anomalies, such as data transmission errors and insufficient resources, to ensure the successful completion of capacity calibration.
[0074] The early warning decision module dynamically triggers a multi-level early warning strategy based on the capacity loss rate. The module pre-sets multiple warning thresholds, such as minor and major. When the cluster's capacity loss rate exceeds the minor warning threshold, the module issues a minor warning, alerting system administrators to data consistency issues and enabling them to take appropriate preventative measures, such as increasing data backup frequency and optimizing data synchronization strategies.
[0075] When the capacity loss rate exceeds the critical warning threshold, the warning decision module issues a critical warning and automatically executes a series of emergency response measures. For example, when a data integrity warning is triggered, the following steps are executed: a) Automatically initiate a data repair process to reconstruct lost data using geo-redundant replicas; b) Mark the abnormal cluster as read-only and generate an operations and maintenance work order to prevent further data corruption and loss, notifying operations and maintenance personnel to address the issue; c) Trigger data redistribution across availability zones according to pre-set policies to improve system fault tolerance and data availability.
[0076] The metadata repository records node topology, version history, and calibration logs. Node topology describes the physical and logical location of data nodes within a distributed database, including information such as the cluster and logical shard to which the node belongs. This information is crucial for data synchronization management and consistency testing. The system can rationally schedule data synchronization tasks and resource allocation based on the node topology.
[0077] The version history records data version changes for each data node, including detailed information such as version number, update time, and update operation. By analyzing the version history, system administrators can trace the data update process and locate the root cause of data inconsistencies. The calibration log records all operations and events during the data consistency calibration process, such as node isolation time, data repair operations, and capacity calibration policies. This log information provides important information for system auditing and troubleshooting.
[0078] The data synchronization management module, consistency detection engine, exception handling module, early warning decision module, and metadata repository are each connected to the communication module. The communication module uses efficient and reliable communication protocols such as TCP / IP and HTTP / 2 to ensure fast and accurate data and command transmission between modules.
[0079] During data transmission, the communication module uses data encryption technologies, such as SSL / TLS, to encrypt sensitive data and prevent it from being stolen or tampered with during transmission. Furthermore, the communication module features fault tolerance and retransmission mechanisms. When errors or loss occur during data transmission, the module automatically resends the data to ensure data integrity and reliability.
[0080] Example 1: E-commerce order processing system with a single cluster architecture
[0081] An e-commerce platform uses a distributed database with a single cluster architecture to store order-related data. The cluster contains 10 data nodes, and the data is divided into multiple logical shards, such as order information, product inventory, and user accounts. Each logical shard has three data copies.
[0082] During daily operations, the system regularly performs data synchronization tests according to preset rules. For example, when a user submits an order, the system triggers a cross-shard transaction commit, modifying data in the order information shard (recording order details), the product inventory shard (deducting the corresponding product inventory), and the user account shard (deducting the user's account balance). The distributed transaction processing mechanism uses a two-phase commit protocol to ensure that all data modifications to these three shards are either successfully committed or rolled back. Simultaneously, replica synchronization verification monitors the data synchronization status between the primary and secondary replicas of each shard in real time, compares data checksums, and records synchronization timestamps to generate a data version sequence set.
[0083] After multiple tests, the system extracted the independent version sequence for each data node. Calculations revealed that, within the version range of node 5, V_max - V_min = 3, while the preset version tolerance threshold is 2. Therefore, node 5 was deemed inconsistent and the system immediately isolated it by modifying its network access permissions. The remaining nine nodes were consistent, and the process proceeded to step three.
[0084] The nine synchronized nodes are sorted by IP address and the overall cluster version consistency is calculated. The result is Δ_max - Δ_min = 1, which is less than the cluster difference threshold of 2. The cluster consistency is considered met and the process proceeds directly to step 6. Based on the current effective storage capacity of each node, the system generates a global consistency report that includes the data consistency status of each node and information about isolated nodes. This report provides a basis for subsequent system optimization and maintenance.
[0085] Example 2: Financial transaction system with multi-cluster architecture
[0086] A large financial institution uses a distributed database with a multi-cluster architecture to process transaction data. It has a total of 5 data clusters, each containing 15 data nodes. The data is divided into different logical shards based on region, and the replica loss weight coefficient of key business data is set to 0.8.
[0087] During a data synchronization check, the system discovered inconsistent data synchronization in four nodes in Cluster 2 and three nodes in Cluster 4. First, the storage capacity loss rate for each cluster was calculated. Cluster 2's capacity loss rate was 4 / 15 × 0.8 = 0.213, and Cluster 4's capacity loss rate was 3 / 15 × 0.8 = 0.16. The loss rate threshold was set at 0.2, and the loss rate for Cluster 2 exceeded the threshold, triggering a data integrity alert.
[0088] Cluster 2, due to its highest capacity loss rate, was selected as the baseline cluster. For Clusters 1, 3, and 5, the system performed dynamic shard migration and replica scaling. For example, some logical shards containing non-critical business data in Cluster 1 were migrated to Cluster 2, and the number of data replicas in Cluster 1 was appropriately reduced. For Clusters 3 and 5, the number of replicas for some critical business data was increased based on the storage capacity of the baseline cluster to balance storage loads across clusters and perform cross-cluster data consistency calibration.
[0089] After the calibration is completed, the system generates a global consistency report based on the effective storage capacity of each cluster, recording in detail the consistency calibration process of each cluster, information on isolated nodes, etc., to help operation and maintenance personnel fully understand the system data consistency status and ensure the accuracy and security of financial transaction data.
Claims
1. A distributed database data consistency detection method, characterized in that: The following steps are involved: Step 1: Perform multiple data synchronization tests on multiple data nodes in the distributed database to generate a data version sequence set corresponding to each test; The data nodes belong to a data cluster, each data cluster contains N data nodes, and the data version sequence includes the data version number and synchronization timestamp of each data node; Step 2: Based on multiple sets of data version sequence sets, extract the independent version sequence of each data node and calculate the version change range of each data node; determine the synchronization consistency of the data nodes based on the version change range. If all data nodes are synchronized, proceed to step 3; otherwise, isolate the inconsistent data nodes and proceed to step 3; Step 3: Sort the synchronized data nodes by the preset identifier to generate a sorted version sequence set; calculate the overall cluster version consistency based on this set. If the consistency meets the threshold, proceed to step 6; otherwise, proceed to step 4. Step 4: Locate and isolate the data node with the abnormal version based on the version difference of each node in the sorted version sequence set; if the database has a single-cluster architecture, proceed to step 6; if it has a multi-cluster architecture, proceed to step 5; Step 5: Count the number of isolated nodes in each cluster and calculate the storage capacity loss rate of the corresponding cluster. If the loss rate of all clusters does not exceed the set threshold, perform cross-cluster data consistency calibration and proceed to step 6. Otherwise, trigger a data integrity warning. Step 6: Based on the calibrated effective storage capacity of each cluster, generate a distributed database global consistency report to complete data synchronization management.
2. A distributed database data consistency detection method according to claim 1, characterized in that: The data cluster includes multiple logical shards, each logical shard consists of multiple data replicas; the data synchronization test includes cross-shard transaction submission and replica synchronization verification.
3. A distributed database data consistency detection method according to claim 2, characterized in that: The version change range is: Extract the minimum version number V_min and the maximum version number V_max in the data node version sequence. If (V_max - V_min) ≤ the preset version tolerance threshold, the node is considered synchronized; otherwise, it is considered inconsistent.
4. A distributed database data consistency detection method according to claim 3, characterized in that: The overall version consistency of the computing cluster is: Calculate the maximum Δ_max and minimum Δ_min of the version change range of all data nodes. If (Δ_max - Δ_min) < the cluster difference threshold, the cluster consistency is determined to be met.
5. A distributed database data consistency detection method according to claim 4, characterized in that: Nodes for locating version anomalies include: Calculate the average Δ_avg of the version differences between each pair of data nodes in the cluster. If the absolute deviation between the version change range of a node and Δ_avg exceeds a dynamic floating threshold, the node is determined to be an abnormal node; the dynamic floating threshold is dynamically adjusted according to the historical synchronization volatility.
6. A distributed database data consistency detection method according to claim 5, characterized in that: The storage capacity loss rate is: For each cluster, the ratio of the number of isolated nodes to the total number of initial nodes is used as the capacity loss rate, and the replica loss weight coefficient is superimposed for correction.
7. A distributed database data consistency detection method according to claim 6, characterized in that: The cross-cluster data consistency calibration includes: The cluster with the largest capacity loss rate is selected as the benchmark cluster, and the effective storage capacity of the remaining clusters is adjusted to match the benchmark cluster through dynamic shard migration or replica expansion and contraction.
8. A distributed database data consistency detection method according to claim 7, characterized in that: When the data integrity warning is triggered, perform at least one of the following operations: a) Automatically initiate the data recovery process and rebuild lost data based on geo-redundant copies; b) Mark the abnormal cluster as read-only and generate an operation and maintenance work order; c) Trigger data redistribution across availability zones based on preset policies.
9. A distributed database data synchronization system, characterized in that: A distributed database data consistency detection method according to any one of claims 1 to 8 is applied, comprising: a data synchronization management module, a consistency detection engine, an exception handling module, an early warning decision module, a metadata repository, and a communication module; The data synchronization management module, consistency detection engine, exception handling module, early warning decision module, and metadata repository are respectively connected to the communication module for communication; The data synchronization management module is used to control multi-node concurrent reading and writing and version number generation; The consistency detection engine is used to analyze the version sequence in real time and calculate the consistency index; The exception handling module: performs node isolation, data repair and capacity calibration operations; The early warning decision module: dynamically triggers a multi-level early warning strategy according to the capacity loss rate; The metadata repository records node topology relationships, version history, and calibration logs.
Citation Information
Patent Citations
Database version processing method and system
CN112579568A
Multiple version database concurrency control system
US5280612A