Database cluster partition running method, device and medium in network partition scenario
Through multi-master partition architecture and structured parsing technology, combined with the MapReduce framework, the data synchronization and conflict handling problems of distributed databases under network partitions are solved, efficient and automatic data synchronization and consistency recovery are achieved, and the system availability and performance are improved.
Patent Information
- Application Number
- CN202511128973.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-08-13
AI Technical Summary
In a network partition scenario, existing distributed database systems cannot continue to operate independently after a partition failure. Data synchronization is subject to conflicts and performance bottlenecks, and the recovery process relies heavily on manual intervention, affecting system availability and consistency.
It adopts a multi-master partition architecture, achieves local business continuity by recording transaction update logs, uses structured parsing technology to process data, combines the MapReduce framework for incremental synchronization and conflict arbitration, optimizes data transmission and storage, and realizes automated fault recovery and data consistency.
Ensure that the database cluster runs independently in the event of a network partition, achieve efficient data synchronization and low latency, automatically handle conflicts, improve system high availability and data consistency, and support millisecond-level conflict identification and data consistency convergence.
Smart Images

Figure CN120631637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed data storage, in particular to a database cluster partition running method and device in a network partition scenario and a medium, and is suitable for the fields of database cluster partition running, automatic fault recovery, eventual consistency, conflict detection and resolution, and high availability system design of a distributed database in an extreme scenario such as network partition and brain split. BACKGROUND
[0002] Under the background of rapid development of big data and artificial intelligence technology, distributed databases are widely used in key fields such as financial transactions and Internet of Things real-time monitoring due to their high scalability and elastic architecture. The mainstream distributed database adopts a Shared-Nothing (no sharing) architecture, the core principle of which is to store data shards in independent nodes and achieve data distribution, load balancing and fault isolation through coordination mechanisms. However, in scenarios where network structure is simple or deployment in different places leads to partitioning (such as fields with extremely high requirements for network stability such as finance), the existing technology has the following defects:
[0003] First, the system cannot continue to run independently after partition failure. When there is a network partition (such as a communication interruption between nodes), the Shared-Nothing architecture may not be able to provide independent services due to the lack of shared resources, or after the master node is disconnected from some nodes, concurrent writes to the same data by multiple partitions may cause conflicts. The CAP theorem states that in a distributed system, consistency, availability, and partition tolerance are incompatible.
[0004] Second, there are conflicts in data synchronization after recovery. After the network is restored, a large number of conflicts may occur due to semantic conflicts (such as field format mismatches or business rule violations) or operation timing differences when merging data versions in different partitions. Relying on high-cost distributed transactions or application layer logic to resolve conflicts requires a lot of manual intervention.
[0005] Third, there are performance and delay defects in data synchronization. Existing synchronization schemes (synchronous / asynchronous replication) have significant delays in large-scale clusters. Synchronous replication is limited by network round-trip delay and node processing capacity, while asynchronous replication may result in data loss due to node failure. When deployed across regions, network transmission delay and storage load further exacerbate performance loss. Large-scale data synchronization also increases network and storage load, affecting system performance. SUMMARY
[0006] In view of the defects of the distributed database in consistency, availability, data synchronization efficiency and conflict processing in the network partition scene, the application proposes a database cluster partition running method, equipment and medium in the network partition scene, which solves the core problem through the following innovations:
[0007] First, availability guarantee in the partition scene. A multi-master partition architecture is constructed, so that each partition can still run independently in the separated state, and local business continuity is realized through recording transaction update logs (such as MySQL Binlog). This architecture avoids cross-partition transactions and strong data dependence, ensures that each partition can still handle business requests independently when the network fails or the partition is isolated, thereby breaking the strong dependence of traditional distributed systems on the master node and realizing cross-partition high availability.
[0008] Second, automatic resolution of data merging conflicts. During data synchronization, the data is standardized in the data preprocessing stage, and the data in a unified paradigm is used to cooperate with the pre-defined conflict detection rules for data merging, so as to minimize human intervention; the abnormal data that has not been resolved is marked and isolated to avoid polluting the global copy, and finally the data convergence and strong consistency are realized through arbitration strategies.
[0009] Third, performance optimization and delay control. The application is based on the incremental data synchronization of the database transaction logs (such as MySQL Binlog) recorded by the multi-master partition, which avoids incremental data transmission; the MapReduce framework is used for distributed sharding and parallel processing of incremental data to improve throughput and computing efficiency; transaction order management technology is used to reduce cross-partition data verification overhead; and compression and priority scheduling are used to reduce network transmission delay, and data storage index is optimized to reduce disk IO consumption.
[0010] The technical scheme adopted by the application is as follows:
[0011] A database cluster partition running method in a network partition scene, comprising:
[0012] A multi-master distributed database cluster is constructed, and each node continuously captures the database native transaction log; change event data is obtained, and unified data storage is performed;
[0013] The original log is parsed into structured data, and uniform sharding and parallel processing are performed;
[0014] The change record of the block log is processed, and the change event of the primary key is arranged according to the rules;
[0015] Conflict processing and data synchronization are performed, and all nodes are synchronized after the conflict is resolved;
[0016] Through exception handling and fault tolerance, self-healing and data consistency are realized.
[0017] Further, the construction of the multi-master distributed database cluster, each node continuously capturing database native transaction log, comprising: the cluster adopts multi-master partition architecture, in normal state, the database cluster performs full+incremental backup, and the backup is used as the starting point of recovery check; in abnormal state, each master partition runs independently and records the native transaction log.
[0018] Further, the acquisition of the change event data and the unified data storage, comprising: synchronizing the incremental change event of all nodes to the distributed storage through a preset protocol, and transmitting the data during the network interruption.
[0019] Further, the parsing of the original log into structured data, comprising: filtering non-pre-set events, and parsing the statement-level log into row-level structured log; the pre-set events include update, insert and delete operations.
[0020] Further, the execution of uniform sharding, comprising: dividing the log into input shards in time sequence and data volume, so as to balance the load of Map tasks.
[0021] Further, the processing of the block log change record, arranging the change event of the primary key according to the rules, comprising:
[0022] Hash grouping: distributing events to the same Reduce task through primary key hashing, ensuring that the same primary key change is processed by a single node;
[0023] Time sequence verification: sorting events according to timestamp or global transaction identifier GTID, triggering fault tolerance process if there is timestamp reverse order; through the shuffle and sort mechanism of the MapReduce programming model, grouping and arranging the change events of the same table and primary key in time sequence.
[0024] Further, the execution of conflict processing and data synchronization, synchronizing all nodes after solving the conflict, comprising:
[0025] Priority setting: setting unified priority rules, and automatically arbitrating timestamp conflicts according to preset rules;
[0026] Automatic resolution: grouping change data of the same primary key in the Reduce phase, automatically processing to generate the final data state, marking abnormal conflict data for isolation and providing a manual review interface;
[0027] Data synchronization: writing the merged result to the distributed storage through transaction order management, and synchronizing to all nodes.
[0028] Further, the implementation of self-healing and data consistency through abnormal processing and fault tolerance, comprising:
[0029] Backup and rollback: regular snapshot backup and incremental log storage backup, supporting fast rollback;
[0030] Task retry: MapReduce task automatic retry, failed shard marking does not affect the overall process;
[0031] Consistency check: after synchronization, differences are found by MD5 algorithm comparison, triggering repair process.
[0032] A computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the database cluster partition running method under the network partition scenario.
[0033] A computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the database cluster partition running method under the network partition scenario.
[0034] The beneficial effects of the present application are:
[0035] 1. Partition autonomous operation and horizontal expansion of multi-master architecture. The multi-master partition architecture ensures that each sub-cluster can operate independently and autonomously when the cluster is network partitioned, eliminates cross-partition transaction dependence, and ensures local business continuity in the network partition state; supports dynamically expanding master nodes on demand to horizontally expand capacity, and realizes resource elastic scaling with intelligent load balancing and adaptive routing strategy, thereby improving system high availability and disaster recovery capability without service interruption.
[0036] 2. Real-time data synchronization of multiple distributed databases based on structured parsing technology. Based on structured parsing technology, by designing a unified row-level structured data format and semantic parsing engine, the native statement-level log (Statement-Based) of multiple distributed databases such as MySQL (Binlog), Oracle (Redo Log) is converted into a standardized key-value pair form of row-level (Row-Level) log in real time, thereby eliminating the log semantic differences of heterogeneous databases in cross-partition data synchronization scenarios, improving transmission efficiency and ensuring data consistency.
[0037] 3. High-availability self-healing mechanism based on conflict arbitration and automatic repair. A distributed detection engine is used to automatically determine the priority and conflict logic of conflict data merging in combination with pre-defined data conflict rules (such as timestamp priority, field format error, and business rule violation), mark and isolate abnormal data, and realize low-human-intervention fault recovery and state reconstruction; during synchronization, the system can automatically detect network partition and isolate faulty nodes to ensure that the remaining nodes continue to operate stably during partition.
[0038] 4. Low-latency high-throughput architecture based on incremental synchronization and distributed computing. Incremental synchronization only transmits change events instead of full data, reducing network bandwidth occupancy by more than 90%, and using compression and priority scheduling techniques to optimize transmission efficiency; using the distributed processing capability of MapReduce, through dynamic sharding, transaction sequence management and parallel processing methods, eliminating single-point performance bottlenecks, supporting million-level data throughput per second, and realizing efficient synchronization of large-scale clusters.
[0039] In summary, the present application provides a high-throughput, low-latency, strongly consistent data synchronization solution for large-scale cross-network partition scenario database clusters through the triple innovation of multi-master architecture + structured log synchronization, automated conflict arbitration engine and incremental distributed computing framework, effectively overcoming the limitations of traditional solutions in aspects such as the inability to continue running independently after partition failure, the inability to quickly and automatically recover, and data synchronization performance bottlenecks. The present application can achieve millisecond-level conflict identification and data consistency convergence after network disconnection, ensuring the continuous and stable operation of the cluster in a partitioned state, and realizing self-healing and continuous availability of the system. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 is a network partition scenario database cluster partition running method flowchart of the present application embodiment 1.
[0041] Figure 2 is a network partition database synchronization principle diagram based on MapReduce in the present application embodiment 2. DETAILED DESCRIPTION
[0042] In order to have a more clear understanding of the technical features, purposes and effects of the present application, the specific embodiments of the present application will be described. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, that is, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0043] Embodiment 1
[0044] As shown in Figure 1 , the present embodiment provides a network partition scenario database cluster partition running method, comprising the following steps:
[0045] S10: Construct a multi-master distributed database cluster, each node continuously captures database native transaction log; obtain change event data and perform unified data storage;
[0046] S20: Analyze the original log into structured data, perform uniform sharding and parallel processing;
[0047] S30: processing block log change records, arranging change events of primary keys according to rules;
[0048] S40: performing conflict processing and data synchronization, synchronizing all nodes after solving conflicts;
[0049] S50: realizing self-healing and maintaining data consistency through exception processing and fault tolerance.
[0050] Preferably, step S10 comprises the following sub-steps:
[0051] S101 (architecture design): the cluster adopts a multi-primary partition architecture, in a normal state, the database cluster performs full+incremental backup, and the backup is used as a starting point of a checkpoint for recovery; in an abnormal state, each primary partition independently runs and records a native transaction log (such as a MySQL Binlog).
[0052] S102 (incremental storage): all incremental change events of the nodes are synchronized to distributed storage through a reliable protocol (such as TCP / MQ) to synchronize incremental data, only data during network interruption is transmitted, and full synchronization is avoided.
[0053] Preferably, step S20 comprises the following sub-steps:
[0054] S201 (filtering and standardization): filtering non-INSERT / UPDATE / DELETE (insert / update / delete) events, and parsing a statement-level log (statement) into a row-level structured log (row).
[0055] S202 (sharding strategy): using Input Split (input split) to evenly divide the log into input splits in chronological order and data volume, and ensuring load balancing of Map (mapping) tasks.
[0056] Preferably, step S30 comprises the following sub-steps:
[0057] S301 (hash grouping): distributing events to the same Reduce (reduction) task through primary key hashing, and ensuring that changes of the same primary key are processed by a single node.
[0058] S302 (time sequence verification): sorting events according to timestamp / GTID (global transaction identifier), triggering a fault tolerance process (step S50) if there is reverse timestamp order; and grouping change events of the same table and primary key and arranging them in chronological order through a Shuffle&Sort (shuffle and sort) mechanism of MapReduce (a programming model).
[0059] Preferably, step S40 comprises the following sub-steps:
[0060] S401 (Priority Rules): Establish unified rules, such as DELETE operations taking precedence over INSERT / UPDATE; timestamp conflicts are automatically arbitrated according to preset rules (such as timestamp priority and business rules).
[0061] S402 (Automated Resolution): Enter the Reduce phase to aggregate the changed data with the same primary key. Automated processing is performed based on Table 1 (Data Conflict Merging Rules Table) to generate the final data state. Abnormal conflicting data is marked and isolated, and a manual review interface is provided.
[0062] S403 (Data Synchronization): The merge results are written to distributed storage through transaction order management and synchronized to all nodes.
[0063] Table 1 - Data conflict merging rules table
[0064]
[0065] Preferably, step S50 includes the following sub-steps:
[0066] S501 (Backup and Rollback): Regular snapshot backup and incremental log storage backup, supporting fast rollback.
[0067] S502 (Task Retry): MapReduce failed tasks are automatically retried, and marking the failed shards does not affect the overall process.
[0068] S503 (consistency check): After synchronization is completed, differences are found through MD5 comparison, triggering the repair process (such as re-triggering steps S30-S40).
[0069] Example 2
[0070] This embodiment is based on embodiment 1:
[0071] like Figure 2 As shown, this embodiment provides a database cluster partition operation method in a network partition scenario, simulating the complete synchronization process from data capture to final aggregation using Binlog logs.
[0072] In this example, a MySQL multi-master architecture is configured. Each master node is assigned a unique server-id. Binlog logging is enabled (log-bin=mysql-bin), and gtid_mode=ON is configured to ensure global transaction ID uniqueness. A pre-configured directory structure is used for distributed storage to store sharding and merging results.
[0073] Specifically, this implementation method can be implemented by the following steps:
[0074] Phase S10: Capture and store Binlog change events.
[0075] S101: Real-time parsing of Binlog events through MySQL native tools (such as mysqlbinlog), capturing INSERT / UPDATE / DELETE operations, and filtering non-business change events (such as transaction BEGIN / COMMIT, DDL statements, etc.).
[0076] S102: Incremental synchronization under normal network conditions.
[0077] Sharding strategy: Binlog is divided into shards according to time window (such as every 5 minutes) or file size (such as 100MB), with naming rules binlog_YYYYMMDD_HHMM (such as binlog_20231001_0900).
[0078] Transmission and storage: Asynchronous transmission of shards to distributed storage through Kafka, using ACK mechanism to ensure atomic data writing.
[0079] S103: Network interruption state.
[0080] Local staging: Change events are temporarily stored in a local path (such as / data / mysql_binlog / ), and written to files (such as binlog_offline_20231001_093000.log) in timestamp order.
[0081] Recovery strategy: After network recovery, temporarily stored shards are uploaded in timestamp order, avoiding duplication or omission through version number or timestamp comparison.
[0082] Stage S20: Log parsing and sharding.
[0083] S201: Structured parsing.
[0084] Filter invalid events: Remove non-business change events.
[0085] Log conversion: Convert Statement logs of Binlog to Row format, generating standardized key-value pairs (JSON).
[0086] Unified format: DELETE operation retains old_data field, no need for new_data; INSERT operation only retains new_data.
[0087] S202: Input split (InputSplit).
[0088] Divide Binlog into multiple InputSplit according to time sequence and data volume (such as single shard ≤500MB), assign independent Map tasks to each shard to ensure load balancing.
[0089] Stage S30: Shuffle & Sort.
[0090] S301: Primary key hash grouping.
[0091] Generate hash value using (table name, primary key) joint key, and assign to corresponding Reduce task (e.g. primary key 100 of table users is handled by Reduce Task 3).
[0092] S302: Time sorting and conflict handling.
[0093] Sorting basis: Sort change events in ascending order of timestamp.
[0094] Time sequence conflict handling: If reverse order events are detected (e.g. event with timestamp 1696123000 comes after event with timestamp 1696123500), mark as abnormal data and trigger S50 fault tolerance process.
[0095] S303: Time sorting: Sort events with id=100 and 102 according to event timestamp (timestamp) or MySQL's global transaction ID (GTID).
[0096] S304: If timestamp reverse order is detected (e.g. event disorder caused by network delay), trigger fault tolerance process (mark abnormal events and manually review).
[0097] Stage S40: Reduce phase merges conflicts.
[0098] S401: Conflict handling rules.
[0099] Based on Table 1 (data conflict merge rules table), strictly time sequence priority executes events in time order, DELETE operation is prior to INSERT / UPDATE (e.g. delete first and then insert, final state is new data after deletion).
[0100] S402: Conflict marking and resolution.
[0101] Automatically trigger predefined rules (e.g. last operation result of id=100), complex conflicts reserve manual review interface.
[0102] S403: Result output.
[0103] Merged final state data is written to distributed storage in key-value pair form.
[0104] Stage S50: Fault tolerance and exception handling.
[0105] S501: Data backup, generate full data snapshot regularly (such as every day at 00:00), and keep the latest 7 days of Binlog shards.
[0106] S502: Partial automatic recovery, when the MapReduce task fails, the system automatically retries the failed shards (up to 3 times), without affecting the progress of other shards.
[0107] S503: Consistency check.
[0108] Example scenario: id=102 data check failed, mark the data and return to stage S30 to re-shuffle & sort (shuffle and sort), and synchronize to all nodes after re-checking.
[0109] Final check compares the original MySQL data with the incremental data stored in the distributed storage. If the MD5 check fails, the repair process is triggered: suspend the current task → re-execute the MapReduce process for the failed shard → merge the results after the check passes.
[0110] Embodiment 3
[0111] This embodiment is based on embodiment 1:
[0112] This embodiment provides a computer device, including a memory and a processor, the memory stores a computer program, and the processor implements the network partition scenario database cluster partition running method of embodiment 1 when executing the computer program. The computer program can be in source code form, object code form, executable file or some intermediate form, etc.
[0113] Embodiment 4
[0114] This embodiment is based on embodiment 1:
[0115] This embodiment provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the network partition scenario database cluster partition running method of embodiment 1. The computer program can be in source code form, object code form, executable file or some intermediate form, etc. The storage medium includes any entity or device that can carry computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content of the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0116] It is apparent that, for the method embodiments described previously, the steps of the methods have been described as being arranged in a particular order. However, it is to be appreciated that this is merely one example, and that the steps of the methods can be arranged in other orders or performed contemporaneously. Furthermore, it is to be appreciated that the embodiments described in the specification are merely preferred embodiments, and that the steps of the methods need not be performed in the order described.
Claims
1. A method for running a database cluster partition in a network partition scenario, characterized in that, The application relates to a method for running a database cluster in a network partitioning scenario. A multi-master distributed database cluster is constructed, and each node continuously captures a database native transaction log; Change event data is acquired and unified data storage is performed; Original logs are parsed into structured data, even sharding and parallel processing are executed; Block log change records are processed, and change events of primary keys are arranged according to rules; Conflict processing and data synchronization are executed, and all nodes are synchronized after conflicts are solved; Through exception processing and fault tolerance, self-healing is realized and data consistency is maintained; The method for processing block log change records and arranging change events of primary keys according to rules comprises the following steps: Hash grouping: events are distributed to the same Reduce task through primary key hashing, so that changes of the same primary key are processed by a single node; Time sequence verification: events are sorted according to timestamps or global transaction identifiers GTIDs, and if there is timestamp reverse order, a fault tolerance process is triggered; through the shuffling and sorting mechanism of the MapReduce programming model, change events of the same table and primary key are grouped and arranged in chronological order; The method for executing conflict processing and data synchronization and synchronizing all nodes after conflicts are solved comprises the following steps: Priority setting: a unified priority rule is set, and timestamp conflicts are automatically arbitrated according to the preset rule; Automatic resolution: change data of the same primary key is aggregated in the Reduce stage, and the final data state is automatically processed, and abnormal conflict data is marked for isolation and a manual review interface is provided; Data synchronization: the combined result is written into distributed storage through transaction sequence management, and is synchronized to all nodes. 2.The method of claim 1, wherein, The method for constructing a multi-master distributed database cluster and continuously capturing a database native transaction log of each node comprises the following steps: the cluster adopts a multi-master partitioning architecture, in a normal state, a database cluster performs full+incremental backup, and the backup is used as a recovery check starting point; in an abnormal state, each master partition independently runs and records a native transaction log. 3.The method of claim 1, wherein, The method for acquiring change event data and performing unified data storage comprises the following steps: incremental change events of all nodes are synchronized to distributed storage through a preset protocol to synchronize incremental data, and data during network interruption is transmitted.
4. The database cluster partition running method in a network partition scenario according to claim 1, characterized in that, The method for parsing original logs into structured data comprises the following steps: non-preset events are filtered, and statement-level logs are parsed into row-level structured logs; the preset events comprise update, insert and delete operations.
5. The database cluster partition running method in a network partition scenario according to claim 1, characterized in that, The method for executing even sharding comprises the following steps: input shards are evenly divided according to time sequence and data volume, so that Map task loads are balanced.
6. The database cluster partition running method in a network partition scenario according to claim 1, characterized in that, The method for realizing self-healing and maintaining data consistency through exception processing and fault tolerance comprises the following steps: Backup and rollback: regular snapshot backup and incremental log storage backup are performed, and quick rollback is supported; Task retry: MapReduce failed tasks are automatically retried, and the failed shards are marked without affecting the overall process; Consistency check: after synchronization is completed, differences are found through MD5 algorithm comparison, and a repair process is triggered. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to realize the database cluster partition running method in the network partitioning scenario according to any one of claims 1-5.
8. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to realize the database cluster partition running method in the network partitioning scenario according to any one of claims 1-5.
Citation Information
Patent Citations
Partition-based concurrency control method in multi-main cloud database scene
CN113535742A
Heterogeneous database real-time synchronization method supporting full increment integration
CN119166717A