Distributed multi-master database consistency processing method, device and equipment
By constructing a hierarchical replication topology and a dynamic adjustment mechanism, the problem of data inconsistency caused by network conditions between nodes in a distributed multi-master database is solved, ensuring data consistency and efficiency.
Patent Information
- Application Number
- CN202510873457.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
In a multi-master replication architecture, network conditions between different nodes in a distributed multi-master database cause replication latency fluctuations, resulting in inconsistent read and write data versions on each node within a short period of time, which affects data consistency.
By constructing a hierarchical replication topology, identifying core and edge nodes, establishing semi-synchronous links and asynchronous replication channels, scheduling write request processing, and combining Bloom filters, multi-strategy merging mechanisms, and automatic rollback mechanisms, node status indicators are monitored, and traffic and sharding schemes are dynamically adjusted to ensure data consistency.
It achieves data consistency in a distributed multi-master database during write request processing, balancing cross-regional availability with strong local consistency, and optimizes network interaction and node operating efficiency.
Smart Images

Figure CN120994742A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of database processing, and particularly relates to a distributed multi-master database consistency processing method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] Under the modern Internet and cloud computing environment, distributed databases have become the core infrastructure supporting large-scale and highly available business systems. With the growing demand for microservice architecture and cross-regional deployment, the multi-master replication mode has gradually attracted widespread attention and application due to its read-write load balancing, fault domain isolation and cross-machine room switching capability.
[0003] However, under the multi-master replication architecture, the network conditions between different nodes in the distributed multi-master database (three network operators interconnection, cross-regional link) can cause replication delay fluctuations, making the read-write data versions of each node not uniform in a short period of time, which is not conducive to ensuring the data consistency of the internal nodes of the distributed multi-master database. SUMMARY
[0004] Therefore, it is necessary to provide a distributed multi-master database consistency processing method, device, computer equipment, computer readable storage medium and computer program product capable of ensuring the data consistency of the distributed multi-master database to solve the above technical problems.
[0005] In a first aspect, the present application provides a distributed multi-master database consistency processing method, comprising:
[0006] receiving a write request triggered for any write node in the distributed multi-master database;
[0007] scheduling a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database maintains data consistency during processing of the write request;
[0008] The hierarchical replication topology structure is constructed by: determining a plurality of core nodes and a plurality of edge nodes from the plurality of write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; establishing a semi-synchronous link between the plurality of core nodes, and for each edge node, establishing an asynchronous replication channel between the edge node and each core node.
[0009] In one embodiment, scheduling a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database comprises:
[0010] taking the edge node receiving the write request as a first edge node;
[0011] In the process of the first edge node processing the write transaction corresponding to the write request, the write transaction is pushed to the replication queue of each core node through the asynchronous replication channel established in the hierarchical replication topology for the first edge node;
[0012] The proportion of the core nodes that have processed the write transaction is determined through the semi-synchronous link established in the hierarchical replication topology, and in the case where the proportion reaches a preset proportion, the confirmation character is fed back to the first edge node in batches through the core nodes that have processed the write transaction.
[0013] In one embodiment, the distributed multi-master database consistency processing method further comprises:
[0014] In the process of processing the write request, the evaluation indicators for the plurality of core nodes are monitored;
[0015] If it is determined according to the evaluation indicators that there is an abnormal core node, the semi-synchronous link associated with the abnormal core node is changed to an asynchronous replication channel until it is monitored that the abnormal core node returns to normal, and the asynchronous replication channel obtained by the change is restored to a semi-synchronous link.
[0016] In one embodiment, the distributed multi-master database is configured with a Bloom filter, a multi-strategy merging mechanism and an automatic rollback mechanism. The distributed multi-master database consistency processing method further comprises:
[0017] The write transaction processing frequencies of the plurality of write nodes are monitored, and the vector clocks of the plurality of write nodes are adjusted according to the write transaction processing frequencies;
[0018] The running conflicts of the plurality of write nodes are processed according to the vector clocks of the plurality of write nodes and the configured Bloom filter, multi-strategy merging mechanism and automatic rollback mechanism.
[0019] In one embodiment, the distributed multi-master database consistency processing method further comprises:
[0020] The state indicators of the plurality of write nodes are monitored, and the traffic distribution scheme between the plurality of write nodes is determined according to the state indicators;
[0021] The pressure indicators of the plurality of write nodes are monitored, and the hot spot shard scheme between the plurality of write nodes is determined according to the pressure indicators;
[0022] The delay indicators of the plurality of write nodes are monitored, and the batch adjustment scheme of the plurality of write nodes processing the write request is determined according to the delay indicators;
[0023] The plurality of write nodes processing the write request are dynamically scheduled according to the traffic distribution scheme, the hot spot shard scheme and the batch adjustment scheme.
[0024] In one of the embodiments, the distributed multi-master database consistency processing method further comprises:
[0025] If a new write node exists in the distributed multi-master database, the data in the distributed multi-master database is synchronized to the new write node through full data snapshot and incremental log playback.
[0026] In the case that the full data snapshot has been completed and the incremental log playback has been played back to the preset log, the online switching processing is performed on the new write node, so that the new write node and other write nodes in the distributed multi-master database keep data alignment.
[0027] In one of the embodiments, the distributed multi-master database consistency processing method further comprises:
[0028] For the distributed multi-master database, a visual monitoring platform is constructed, and an automatic partition repair mechanism and a distributed rollback mechanism are configured.
[0029] Based on the visual monitoring platform, the automatic partition repair mechanism and the distributed rollback mechanism, the operation of the distributed multi-master database is controlled.
[0030] In a second aspect, the present application also provides a distributed multi-master database consistency processing device, comprising:
[0031] The request receiving module is configured to receive a write request triggered for any write node in the distributed multi-master database.
[0032] The request processing module is configured to schedule multiple write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database keeps data consistency in the processing of the write request; the hierarchical replication topology structure is constructed by: determining multiple core nodes and multiple edge nodes from the multiple write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; establishing a semi-synchronous link between the multiple core nodes, and establishing an asynchronous replication channel between each edge node and each core node for each edge node.
[0033] In a third aspect, the present application also provides a computer device comprising a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned distributed multi-master database consistency processing method when executing the computer program.
[0034] In a fourth aspect, the present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the above-mentioned distributed multi-master database consistency processing method.
[0035] In a fifth aspect, the present application also provides a computer program product comprising a computer program which, when executed by a processor, implements the distributed multi-master database consistency processing method described above.
[0036] The distributed multi-master database consistency processing method, device, computer device, computer readable storage medium and computer program product described above can receive a write request triggered for any write node in the distributed multi-master database, and then schedule multiple write nodes in the distributed multi-master database to process the write request through the hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database maintains data consistency during processing of the write request. The hierarchical replication topology structure is constructed in the following manner: multiple core nodes and multiple edge nodes are determined from the multiple write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; a semi-synchronous link is established between the multiple core nodes, and for each edge node, an asynchronous replication channel is established between the edge node and each core node. Based on this, the distributed multi-master database consistency processing method described above can ensure data consistency within the distributed multi-master database when processing the write request. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0038] Figure 1 An application environment diagram of the distributed multi-master database consistency processing method in an embodiment;
[0039] Figure 2 A flowchart of the distributed multi-master database consistency processing method in an embodiment;
[0040] Figure 3 A flowchart of switching a core node link in an embodiment;
[0041] Figure 4 A flowchart of fast conflict judgment based on vector clocks and Bloom filters in an embodiment;
[0042] Figure 5 A flowchart of batch transaction submission in an embodiment;
[0043] Figure 6 A flowchart of node expansion by a new write node in an embodiment;
[0044] Figure 7 A distributed consistency rollback timing diagram in an embodiment;
[0045] Figure 8 A structural block diagram of a distributed multi-master database consistency processing apparatus in an embodiment;
[0046] Figure 9 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0047] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0048] The distributed multi-master database consistency processing method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . The terminal 102 communicates with the server 104 through a network. The data storage system can store data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on a cloud or other network server. The server 104 can be deployed with a write node (write master node) of a distributed multi-master database. The server 104 can receive a write request triggered by a user using the terminal 102 for any write node in the distributed multi-master database through communication with the terminal 102, and then schedule multiple write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database maintains data consistency during the processing of the write request. The hierarchical replication topology structure is constructed by the following method: determining a plurality of core nodes and a plurality of edge nodes from the plurality of write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; establishing a semi-synchronous link between the plurality of core nodes, and for each edge node, establishing an asynchronous replication channel between the edge node and each core node. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and Internet of Things devices. The server 104 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0049] In an exemplary embodiment, as shown in Figure 2 , a distributed multi-master database consistency processing method is provided. Taking the server 104 in Figure 1 as an example, the method includes the following steps 202 to 204. Wherein:
[0050] Step 202, receiving a write request triggered by a terminal for any write node in the distributed multi-master database.
[0051] Wherein, the node in the database refers to the basic unit in the database for storing and managing data. The node can be physical or logical, depending on the type and architecture of the database. In a distributed database, a node usually refers to a different server or computer, each node stores a part of data and is responsible for a part of function. The distributed multi-master database is a database architecture that stores data on multiple nodes and allows any node to directly handle read and write requests, especially write requests. The distributed multi-master database combines the core concepts of "distributed" and "multi-master", aiming to provide high availability, high write throughput, low latency and geographical distribution capability. The write node is a node in the database for processing write requests.
[0052] Optionally, the server can be deployed with a write node of the distributed multi-master database, so as to receive a write request triggered by a terminal for any write node in the distributed multi-master database by communicating with the terminal. The server can be a plurality of physical servers or a plurality of cloud servers, each physical server or each cloud server can be a write node, and any physical server or cloud server can communicate with the terminal to receive the write request triggered by the terminal.
[0053] Step 204, scheduling a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database keeps data consistent during processing of the write request.
[0054] Wherein, the hierarchical replication topology structure is constructed by: determining a plurality of core nodes and a plurality of edge nodes from the plurality of write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; establishing a semi-synchronous link between the plurality of core nodes, and for each edge node, establishing an asynchronous replication channel between the edge node and each core node.
[0055] Optionally, the server can schedule a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, and synchronize data between nodes through the semi-synchronous link and the asynchronous replication channel established in the hierarchical replication topology structure during processing of the write request, so that the distributed multi-master database keeps data consistent during processing of the write request.
[0056] Exemplarily, the construction process of the hierarchical replication topology is described as an example: (1) First, from the multiple write nodes of the distributed multi-master database, multiple core nodes (Core Masters) and multiple edge nodes (Edge Masters) are determined. The core nodes and edge nodes meet the preset hierarchical deployment conditions, which can refer to: 1. The core nodes are deployed in the same city or the same machine room in a high-bandwidth and low-latency network environment. The number of core nodes can be configured according to the throughput demand (usually 3-5 nodes). 2. The edge nodes are deployed in cross-regional or relatively unstable link machine rooms, mainly facing local users for nearby writing to reduce access delay. (2) A full Internet-like topology is established between the core nodes, so that semi-synchronous links are established between multiple core nodes. For each edge node, an asynchronous replication channel is established between the edge node and each core node. Among them, Async is used between the edge node and the core node. When the edge node processes a write transaction, it only needs to be locally submitted successfully and receive an acknowledgement character (ACK), and then immediately feedback success to the client, and asynchronously push the write transaction to the core node in the background. Semi-Sync is used between the core nodes. When the core node processes a write transaction, it needs to receive an acknowledgement character (ACK) after most of the core nodes are locally submitted successfully, and then feedback success to the client. By using the above hierarchical replication topology, the pressure of strong consistency can be concentrated on the core cluster composed of core nodes, and the edge cluster composed of edge nodes provides low-latency writing.
[0057] In the above distributed multi-master database consistency processing method, a write request triggered for any write node in the distributed multi-master database can be received, and then the hierarchical replication topology constructed for the distributed multi-master database is used to schedule multiple write nodes in the distributed multi-master database to process the write request, so that the distributed multi-master database maintains data consistency during the processing of the write request. Among them, the hierarchical replication topology is constructed by: determining multiple core nodes and multiple edge nodes from multiple write nodes; the core nodes and edge nodes meet the preset hierarchical deployment conditions; a semi-synchronous link is established between the multiple core nodes, and for each edge node, an asynchronous replication channel is established between the edge node and each core node. Based on this, by using the above distributed multi-master database consistency processing method, the data consistency within the distributed multi-master database can be ensured when processing the write request.
[0058] It should be noted that the user, the client or the middleware can trigger a write request according to the geographical location and the real-time delay, and select any core node or any edge node in the distributed multi-master database. Taking the client triggering a write request through an edge node as an example, in one embodiment, the layered replication topology structure constructed for the distributed multi-master database is used to schedule multiple write nodes in the distributed multi-master database to process the write request, including:
[0059] The edge node receiving the write request is taken as a first edge node;
[0060] In the process of processing the write transaction corresponding to the write request in the first edge node, the write transaction is asynchronously pushed to the replication queue of each core node through the asynchronous replication channel established in the layered replication topology structure for the first edge node;
[0061] Through the semi-synchronous link established in the layered replication topology structure, the proportion of the core nodes that have processed the write transaction is determined, and in the case that the proportion reaches a preset proportion, the confirmation character ACK is fed back to the first edge node in batches through the core nodes that have processed the write transaction.
[0062] Among them, the core node that has processed the write transaction can refer to the core node that has successfully committed the write transaction locally (synchronized to the local WAL) and received the confirmation character ACK. The preset proportion can be three-fifths, three-fourths, etc., which can be flexibly configured according to actual needs.
[0063] Optionally, the server can take the edge node receiving the write request as the first edge node, and in the process of processing the write transaction corresponding to the write request in the first edge node, the first edge node can determine the write transaction by analyzing the write request, synchronize the write transaction to the local WAL (commit the write transaction locally), and immediately return "write success" to the client after successful submission and receiving the confirmation character ACK. At the same time, the write transaction can be asynchronously pushed to the replication queue of each core node through the asynchronous replication channel established in the layered replication topology structure for the first edge node. Further, after receiving the write transaction pushed by the first edge node, the multiple core nodes can determine the proportion of the core nodes that have processed the write transaction through the semi-synchronous link established in the layered replication topology structure, and in the case that the proportion reaches a preset proportion, it is confirmed that most of the core nodes successfully process the write transaction, at which time the confirmation character ACK is fed back to the first edge node in batches through the core nodes that have processed the write transaction.
[0064] Exemplarily, to optimize the ACK batch feedback efficiency, the core node that has processed the write transaction can aggregate the ACKs corresponding to multiple write transactions into one batch message, thereby reducing the number of network interactions. After receiving the batch ACK, the first edge node can apply it to the local state machine to complete the eventual consistency.
[0065] Exemplarily, before the first edge node receives the batch ACK from the core node, the local latest write of the first edge node is only visible to the node (the first edge node) itself, and is visible to other nodes depending on the replication lag. After the first edge node receives the batch ACK from the core, the whole system enters the "strong consistency" state.
[0066] In the write and replication process between the edge node and the core node, a semi-synchronous + asynchronous hybrid can be used to balance the cross-regional availability and local strong consistency of the distributed multi-master database.
[0067] In an exemplary embodiment, the distributed multi-master database consistency processing method further comprises:
[0068] In the processing of the write request, the evaluation indicators for the multiple core nodes are monitored;
[0069] If it is determined according to the evaluation indicators that there is an abnormal core node, the semi-synchronous link associated with the abnormal core node is changed to an asynchronous replication channel until it is monitored that the abnormal core node has returned to normal, and the asynchronous replication channel obtained by the change is restored to a semi-synchronous link.
[0070] The evaluation indicators, i.e., the health monitoring indicators, can include: round trip time (RTT), representing the average RTT between core nodes; packet loss rate, representing the packet loss percentage of the link between core nodes; node pressure, related to CPU, IO, replication queue length, etc.
[0071] Optionally, in the processing of the write request, the evaluation indicators for the multiple core nodes can be monitored in real time. If it is determined according to the evaluation indicators that there is an abnormal core node, the semi-synchronous link associated with the abnormal core node is changed to an asynchronous replication channel (triggering degradation), until it is monitored that the abnormal core node has returned to normal, and the asynchronous replication channel obtained by the change is restored to a semi-synchronous link (triggering recovery). It should be noted that during the link change, all pending write transactions continue to be pushed according to the original strategy to avoid data loss, and after the link change is completed, the pending transactions of the edge node and the core node are automatically aligned to ensure that there is no omission.
[0072] Specifically, the RTT of the core node exceeds 100 ms or the packet loss rate exceeds 2% → it is determined that the core node is abnormal, triggering link degradation; the RTT of the core node falls below 50 ms and the packet loss rate is <1% → it is determined that the core node is normal, triggering recovery. In order to avoid frequent switching caused by link jitter, a minimum mode retention time, such as 5 minutes, can also be set, and the link obtained after the change needs to be maintained for at least 5 minutes before the next change. The degradation step includes: detecting health degree anomaly in the control plane; switching the semi-synchronous link associated with the affected core node to an asynchronous replication channel; broadcasting a notification to the middleware to update the routing strategy. The recovery step includes: after the health degree returns to normal and the hysteresis period is over, re-establishing the semi-synchronous link and performing heartbeat verification; notifying the middleware to restore the semi-synchronous preferred write route.
[0073] Based on this, as shown in Figure 3 a flowchart for switching core node links is provided, first, evaluation indicators (health monitoring indicators) are collected, and the health of the core node is evaluated according to the evaluation indicators. If the result of the health evaluation determines that a certain core node is abnormal, then trigger degradation, change the semi-synchronous link associated with the abnormal core node in the hierarchical replication topology to an asynchronous replication channel, and notify all nodes of this change. If the result of the health evaluation determines that a certain core node is normal, then control it to continue to maintain the original semi-synchronous link, and if the core node that is evaluated as normal is an abnormal core node that has undergone link change before, then restore the asynchronous replication channel obtained by the change to a semi-synchronous link.
[0074] In this embodiment, the dynamic switching mechanism can be combined to automatically balance consistency and performance according to real-time network and load conditions, ensuring that the distributed multi-master database can run smoothly under any network state.
[0075] In some embodiments, a Bloom filter, a multi-strategy merging mechanism, and an automatic rollback mechanism are configured in the distributed multi-master database; the distributed multi-master database consistency processing method further includes:
[0076] Monitoring the respective write transaction processing frequencies of the plurality of write nodes, and adjusting the respective vector clocks of the plurality of write nodes according to the write transaction processing frequencies;
[0077] Processing the running conflicts of the plurality of write nodes according to the respective vector clocks of the plurality of write nodes, and the configured Bloom filter, multi-strategy merging mechanism, and automatic rollback mechanism.
[0078] Among them, the vector clock (Vector Clock) is a logical clock mechanism used to track the causal order of events in a distributed system. It solves the partial ordering problem of events in a distributed environment by maintaining a vector (array) for each node. It is particularly good at detecting concurrent conflicts (such as multiple nodes modifying the same data simultaneously). The Bloom filter (Bloom Filter) is a high-efficiency probabilistic data structure used to quickly determine whether an element is not in a set (there may be false positives, but there will be no false negatives). Its core feature is high space efficiency, with a query time complexity of O(k) (k is the number of hash functions). It sacrifices some accuracy to gain storage and speed advantages.
[0079] Optionally, the write transaction processing frequency of each of the plurality of write nodes can be monitored periodically (e.g., every minute), such as the number of write transactions processed in the past 10 minutes (how many times the write is written), so as to adjust the vector clock of each of the plurality of write nodes according to the write transaction processing frequency. For example, the write node with a write frequency lower than a threshold (e.g., less than 100 writes in the past 10 minutes) is marked as inactive, and the vector clock component is no longer maintained for it, until the inactive node processes a write transaction again, the inactive node is added to the active list and marked as an active node, and its clock stamp is initialized. In addition, the vector clock design can also be compressed in the following ways: 1. Sparse storage: only maintain the mapping NodeID→Counter for active nodes; the clock of all inactive nodes does not occupy space. 2. Incremental update: when a write transaction is submitted, only the Counter of the current write node is incremented; when the copy receiving end is merged, only the components that have changed are updated to avoid full table scanning.
[0080] Further, to prevent the infinite growth of the sparse map, the clock components of nodes that have not appeared in write operations for more than N days (e.g., 7 days) can also be periodically cleaned up to realize garbage collection of clock stamps.
[0081] Optionally, according to the vector clock of each of the plurality of write nodes and the configured Bloom filter, multi-strategy merging mechanism and automatic rollback mechanism, the running conflicts of the plurality of write nodes can be processed, which can specifically include:
[0082] (1) Based on the vector clock of each of the plurality of write nodes and the Bloom filter configured for each of them, fast conflict judgment is performed. Specifically, a Bloom filter is maintained locally in each write node, with a capacity expected to support M active transactions (e.g., M = 10 6 ), and the false positive rate is set to P (e.g., 1%). Before replicating the write transaction to other write nodes, the primary key and modification column fingerprint of the write transaction are hashed together and inserted into the Bloom filter, and fast conflict judgment is performed according to the vector clock and the Bloom filter. For example, Figure 4As shown, a flowchart for fast conflict judgment based on vector clock and Bloom filter is provided. After receiving a write transaction triggered by a write request, the primary key and modified column fingerprint of the write transaction are hashed together and inserted into the Bloom filter for detection. If the Bloom filter detects a miss, it can be almost certain that there is no conflict with the primary key, and the expensive vector clock comparison is bypassed and directly submitted. If the Bloom filter detects a hit, it is determined that there is a conflict possibility, and the next step of vector clock accurate comparison is entered to detect vector clock conflict. If the vector clock is not in conflict, it is directly submitted, and if the vector clock is in conflict, it enters the conflict merging process.
[0083] (2) In the case of vector clock conflict, multi-strategy conflict merging is performed based on a multi-strategy merging mechanism. First, a plug-in framework architecture for multi-strategy conflict merging is built: a merging interface merge(old_row, new_row)→merged_row is defined in the kernel layer of the database, and a plug-in registration point is opened. Different merging strategies are implemented in the form of scripts (such as JavaScript or Lua), and are loaded according to the table name + column through a configuration file or a management API. Common merging strategies are as follows: numerical accumulation: merged.value = old.value + new.value; taking the maximum / minimum value: merged.value = max(old.value, new.value); JSON merging: performing deep merging on the JSON object in the field, and appending and deduplicating the array field. Based on this, the multi-strategy conflict merging process can include: 1. The conflict detection module synchronizes old_row and new_row to the merging engine. 2. According to the table column mapping, the merge function of the corresponding plug-in is called. 3. If the script execution fails, the default LWW (Last-Write-Wins) strategy is used and an alarm is recorded.
[0084] (3) Introduce CRDT automatic rollback mechanism, taking into account the versatility and flexibility, both can let the simple scene "fast merge", but also for complex scenarios to provide "strong convergence" guarantee. First, can be directed at collaborative documents, user session logs, like like count, which has natural tolerance for unordered merge scenarios, pre-defined as CRDT field. Automatic rollback mechanism can include the following processes: the default process first tries "light merge" (LWW + plug-in). If the conflict complexity exceeds the preset threshold (such as the same row has > 5 column conflicts), the system automatically switches to CRDT structure: 1, the row data is serialized into CRDT state and synchronized. 2, the CRDT convergence algorithm is executed asynchronously in the background, and the final global consistent state is formed. 3, developers can set which tables / fields support automatic rollback in the configuration. Among them, the automatic rollback mechanism of CRDT (Conflict-free Replicated Data Type, Conflict-free Replicated Data Type) is a high-level strategy for handling operation undo or state rollback in distributed systems, which uses the mathematical properties of CRDT itself (commutative law, associative law, idempotency) to implement undo operations without central coordination, ensuring that all replicas eventually converge to a consistent state.
[0085] In the embodiment, on the one hand, the lightweight vector clock is used to optimize and compress the vector clock, and the clock is maintained only for active nodes, and combined with sparse storage and garbage collection, accurate conflict detection with low overhead can be realized. On the other hand, since most transactions do not need accurate comparison, the Bloom filter can be used to quickly determine "no conflict", greatly reducing the CPU and memory burden. On the other hand, based on multi-strategy merge + CRDT automatic rollback mechanism, versatility and flexibility are taken into account, both can let the simple scene "fast merge", but also for complex scenarios to provide "strong convergence" guarantee. Based on this, the combination of compressed vector clock and Bloom filter can be used to realize "fast no conflict" judgment in most scenarios, and when needed, multi-strategy merge plug-in and CRDT automatic rollback mechanism can be used to complete efficient and customizable conflict processing for the daily operation process of distributed multi-master database.
[0086] In one possible implementation, the distributed multi-master database consistency processing method further includes:
[0087] Monitoring the state indicators of the plurality of write nodes respectively, and determining a traffic distribution scheme between the plurality of write nodes according to the state indicators;
[0088] Monitoring the pressure indicators of the plurality of write nodes respectively, and determining a hot spot shard scheme between the plurality of write nodes according to the pressure indicators;
[0089] Monitoring the delay indicators of the plurality of write nodes respectively, and determining a batch adjustment scheme for processing write requests of the plurality of write nodes according to the delay indicators;
[0090] According to the traffic distribution scheme, the hotspot sharding scheme and the batch adjustment scheme, the multiple write nodes are dynamically scheduled to handle the write requests.
[0091] The state indicators can include: average round-trip delay L i , utilization U i , recent conflict rate C i . The utilization U i refers to CPU / IO utilization, and the recent conflict rate C i refers to the proportion of transactions triggering merging or rollback per unit time. The pressure indicators can include: the number of writes of each primary key in a certain window (such as 10 seconds) counted by the middleware. The delay indicators can include: network round-trip delay L and transaction processing time P proc .
[0092] Optionally, the real-time state indicators of the multiple write nodes are monitored, and a global write routing table is maintained at the client or middleware node, recording the state indicators of each write node, and the global write routing table is periodically (such as every second) synchronized with the heartbeat channel of each write node. Further, the average round-trip delay L i in the state indicators is smoothed by using exponential weighted moving average to balance real-time and stability, then for each write node, the comprehensive score S i of the write node is calculated according to the utilization U i , the recent conflict rate C i , and the smoothed average round-trip delay L in the state indicators. Still further, candidate write nodes can be selected from the multiple write nodes according to the sizes of the comprehensive scores of the multiple write nodes, and weighted random or polling is performed among the candidate write nodes to smooth the traffic distribution and obtain the traffic distribution scheme among the multiple write nodes.
[0093] Exemplarily, the average round-trip delay L i in the state indicators is smoothed by using exponential weighted moving average, which can be specifically shown in formula (1):
[0094]
[0095] In formula (1), L i collected is smoothed by using exponential weighted moving average (EWMA) to obtain the smoothed average round-trip delay L The value of a is usually 0.2-0.3 to balance real-time and stability.
[0096] According to the utilization U i , the recent conflict rate C i , and the smoothed average round-trip delay L Calculate the comprehensive score S of the write node i , which can be specifically shown in formula (2):
[0097]
[0098] In formula (2), L max is the acceptable maximum delay threshold, and all indicators after normalization are in [0, 1]. W L , W U , and W C are weights flexibly configured according to service importance, satisfying W L +W U +W C = 1.
[0099] Based on this, the traffic distribution scheme between the plurality of write nodes can be determined in the following manner: obtaining the latest routing table from the client / intermediary, and calculating the S i of each write node; sorting the S i of each write node in ascending order, and selecting the first k nodes with the lowest scores as candidate write masters (usually k = 1 or k = 2); if k > 1, weighted random or polling can be performed among the candidates to smooth the traffic distribution.
[0100] Optionally, the pressure indicators of the plurality of write nodes are monitored, such as the write times of each primary key in a window (such as 10 seconds) counted by the intermediary, and if the write times of a key exceed a threshold H (such as 1000 times), the key is determined to be a hot key. According to the pressure indicators, the hot shard scheme between the plurality of write nodes can be determined specifically as follows: 1, hash the hot key to m sub-keys, for example,
[0101] key j = hash (original key || j), j = 0,..., m-1. 2, the client or intermediary writes the sub-keys to different write nodes respectively to disperse the write of the hot key. 3, automatically combine the sub-key results and return the original key view when reading. Based on this, the hot shard scheme is determined through the hot key local partitioning operation.
[0102] Optionally, the delay indicators (network round-trip delay L and transaction processing time P proc ) of the plurality of write nodes are monitored, and according to the network round-trip delay L and the transaction processing time P proc , the batch transaction delay T batch (n) is calculated based on the batch transaction delay T batch(n), to determine the throughput per unit time. Further, in the trade-off between throughput and delay, an adaptive adjustment algorithm is introduced to solve, to determine the optimal batch size, and the throughput per unit time and the optimal batch size are taken as the batch adjustment scheme for multiple write nodes to process write requests. Taking the packaging of n transactions as an example, according to the network round-trip delay L and the transaction processing time P proc , the batch transaction delay T batch (n) can be calculated as shown in formula (3):
[0103]
[0104] In formula (3), λ is the transaction arrival rate (unit: transaction / second), is the queuing waiting time of assembling batches. The throughput per unit time is about However, for the batch value, if the batch is too large, the throughput improvement is limited and will also cause queuing delay, and if the batch is too small, it will increase network interaction, therefore, according to the network round-trip delay L and the transaction processing time P proc , the optimal batch size n * can be solved to realize the trade-off between throughput and delay:
[0105]
[0106] Further, in the process of solving n * , an adaptive adjustment algorithm can be introduced, the minimum batch n min , the maximum batch n man , and the initial batch size n = n min are set. At the same time, the current λ, network delay L and processing time P proc are monitored in real time, and the following operations are performed every Δt (such as 1 second): 1. Calculate the theoretical optimal n*; 2. If n*>n and n<n max , then n←min(n+1,n * )n * ; 3. If n*<n and n>n min , then n←max(n-1,n * ). Based on this, the optimal batch size n* is solved.
[0107] In summary, the flow distribution scheme, the hot spot sharding scheme and the batch adjustment scheme can be used to dynamically schedule multiple write nodes to process write requests.
[0108] Exemplarily, Figure 5As shown, a flowchart of a batch transaction submission process is provided. After a client initiates a write transaction, in a distributed multi-master database, a batch submission, batch write, batch return acknowledgement character (ACK), and the like can be triggered in sequence, and finally, the client can obtain a transaction-by-transaction result under batch ACK.
[0109] In this embodiment, on the one hand, multi-dimensional evaluation indicators and weighted scores can be used to dynamically select a write node traffic distribution scheme, to ensure optimal delay and least conflicts. On the other hand, hot key local partitioning operations can be used to disperse sub-keys and distribute write pressure to multiple write nodes, to avoid an "avalanche" effect. On the other hand, based on queuing theory and delay models, an optimal batch size can be adjusted in real time, to achieve the best balance between throughput and response. Based on this, the throughput can be maximized, the write delay can be minimized, and the conflict probability can be reduced.
[0110] In one of the embodiments, the distributed multi-master database consistency processing method further includes:
[0111] If a new write node is started in the distributed multi-master database, full data snapshots and incremental log replay are used to synchronize data in the distributed multi-master database to the new write node.
[0112] After the full data snapshot is completed and the incremental log replay is replayed to a preset log, online switching processing is performed on the new write node, to keep the new write node and other write nodes in the distributed multi-master database in data alignment.
[0113] The full data snapshot is a complete data copy at a certain time point, like a photo that freezes all data states. The incremental log can record the data change operation sequence between two snapshots, and the replay refers to a process of re-executing the recorded data change operations in sequence. The full data snapshot provides a basic state and a recovery anchor point of data, and the incremental log replay provides the ability of efficient synchronization change and accurate time point recovery. The combination of the two ensures data security and recoverability, and greatly improves the efficiency of data replication and backup. In this embodiment, the full data snapshot and the incremental log replay are combined, and only one full data snapshot is obtained in the initial stage, and subsequent data synchronization can be maintained through incremental log replay, greatly reducing the bandwidth and I / O overhead of full replication.
[0114] Specifically, the data in the distributed multi-master database is synchronized to the new write node through full data snapshot and incremental log playback, which can specifically include: (1) Snapshot generation. Trigger a consistent read-write blocking point on any core node or dedicated snapshot node to generate a full data snapshot file, and divide it into multiple data slices of balanced size, which are stored in a distributed file system (such as HDFS, Ceph). (2) Parallel distribution: After the new write node is started, each full data snapshot slice is pulled in parallel, and each slice is downloaded in parallel through multi-threading or multi-processing while being verified. (3) Parallel decompression and import. Local parallel decompression (e.g., 4 concurrent processes) and import are used to quickly write slice data to local storage and indexing, and the import is completed to enter the "snapshot ready" state. (4) Incremental log streaming playback. At the same time as the snapshot begins to be generated, the core node continuously pushes the incremental logs generated thereafter to the new write node, and the new write node can play back the incremental logs in parallel during the snapshot import process to apply new transactions in real time. (5) Data consistency. After the full data snapshot is imported, the last piece of residual incremental log is played back to ensure that the data version of the new write node is completely consistent with that of other write nodes. In the full data replication stage, the network and disk peak value decreases by about 60%; parallelization and streaming playback reduce the single-node cold start time to 30%-35% of the original.
[0115] Further, two-stage pipeline synchronization can be adopted to separate full data snapshot import and incremental log playback into "parallel loading period" and "synchronization sprint period", complete node switching without interrupting online writing, and not produce service jitter within the synchronization window. Specifically: (1) In the parallel loading period, after the new write node is started, the incremental snapshot synchronization is started immediately, the online write node continues to serve externally, and the incremental log is pushed to the new write node in real time. (2) In the synchronization sprint period, when the full data snapshot import is completed and the playback is close to the latest transaction position, that is, when the full data snapshot has been completed and the incremental log playback has been played to the preset log, the sprint stage is entered, and the online switching processing of the new write node is performed. The online switching processing is implemented by the following steps: 1. Write routing switching. Temporarily add the new write node to the write routing pool, and at the same time, reduce the synchronization confirmation proportion threshold of the core node to the write node (for example, from three-fifths to one-fifth), to speed up the confirmation speed of the latest write transaction. 2. Short-time consistency. In a short window of 1 to 2 seconds, for critical write transactions, the new write node is required to confirm before returning "write success" to the client, to ensure that the new write node is synchronized with the latest state of the node cluster in the distributed multi-master database, and to keep the new write node in data alignment with other write nodes in the distributed multi-master database. 3. Restore normality. After the synchronization sprint is completed and the version consistency is verified, the write routing and the replication proportion are restored to the standard configuration; the new write node formally assumes the same semi-synchronous or asynchronous replication role as other core nodes, the new write node successfully joins the distributed multi-master database, and the node expansion of the distributed multi-master database is realized.
[0116] Based on this, as shown in Figure 6 , a flowchart of the new write node realizing node expansion is provided. After the new write node is started, it goes through the processes of parallel pulling full data snapshot slices, parallel decompression and import, incremental log playback, and then goes through the processes of synchronization sprint period, temporary adjustment of replication proportion and write routing, revisit of participating in incremental log, version verification, restoration of standard replication proportion and write routing, etc. The new write node can formally assume the same semi-synchronous or asynchronous replication role as other core nodes in the distributed multi-master database, and realize the node expansion of the distributed multi-master database.
[0117] In this embodiment, through the two technical means of "incremental snapshot synchronization" and "smooth online switching", the cold start time is effectively shortened and the online writing is stabilized. Among them, the incremental snapshot synchronization combines full data snapshot and incremental log playback, fully utilizes parallel distribution and import, greatly compresses the cold start time and reduces the peak consumption of resources. The smooth online switching is completed through two-stage pipeline synchronization, which does not block online writing, and completes the strong consistency alignment of the new node and the cluster in a very short time, realizes smooth and imperceptible expansion and recovery.
[0118] In one exemplary embodiment, the distributed multi-master database consistency processing method further comprises:
[0119] For the distributed multi-master database, a visual monitoring platform is constructed, and an automatic partition repair mechanism and a distributed rollback mechanism are configured;
[0120] Based on the visual monitoring platform, the automatic partition repair mechanism and the distributed rollback mechanism, the operation of the distributed multi-master database is controlled.
[0121] Optionally, the visual monitoring platform can be constructed in the following ways: (1) data collection and index aggregation. The key indicators collected are: 1, replication delay: average and maximum delay of edge to core, core to core. 2, conflict rate: the proportion of transactions triggering conflict merging or CRDT rollback per unit time. 3, resource utilization: CPU, memory, disk I / O, network bandwidth. 4, queue depth: replication queue, batch commit queue, WAL playback queue length. The collection method is: a lightweight monitoring agent is deployed on each write node, which reports in time through Prometheus protocol or gRPC Push. The central control platform uniformly pulls or receives the data of each node and writes it into a time series database (such as Prometheus TSDB, InfluxDB). (2) Establish dynamic topology graph, dashboard and report. 1, in the dynamic topology graph: nodes are displayed in a graphical node, and edges and cores are arranged in layers. The state of the node and the link is reflected in real time by color and line width: red indicates high delay / high conflict, yellow indicates moderate risk, and green indicates health. 2, in the dashboard and report: display replication delay trend chart, conflict rate heat map, resource utilization stack chart, etc., support filtering by time range, machine room, etc. (3) Set automatic alarm. 1, set multi-dimensional alarm. Threshold alarm: such as replication delay > 200ms, conflict rate > 5%, queue depth > 1000. Trend alarm: the rising rate of delay or conflict rate exceeds the threshold (such as 10% / min) within a short time. Compound alarm: compound multiple indicators (such as high delay + high conflict) trigger more serious level alarm. 2, set automatic execution. Retry replication: when single-link replication fails, automatically call API to re-establish replication channel. Write routing switching: if a core node has high delay for a long time, it is automatically excluded from the write routing table. Container restart: detect node process exception or resource exhaustion, call container orchestration platform (Kubernetes) for rolling restart. 3, set notification and closed loop. Alarm is pushed to email, Slack, PagerDuty; at the same time, the list of pending events is displayed on the console. The script execution result is written back to the monitoring platform to form an event processing closed loop to avoid repeated triggering.
[0122] Optionally, the automatic partition repair mechanism and the distributed rollback mechanism can be configured in the following ways: (1) Partition write conflict isolation and repair. 1. Detection and isolation. When the conflict detection module determines that there is a multi-node write conflict in the same row and the plugin merge fails, it automatically initiates "row-level isolation locking" - marks the row as "to be repaired", and suspends the subsequent write of the row. 2. Background repair task. Schedule a background repair job, read the locked data by row, and call the business-aware merge plugin script for deep merge or CRDT rollback. After repair, record the merge strategy and result in the monitoring platform, release the isolation lock, and restore the write of the row. (2) Distributed consistency rollback. 1. Failed transaction identification. For partial commit failed transactions that occur during write route switching, node expansion or fault recovery, the system summarizes the affected write masters and transaction IDs. 2. Two-phase commit + compensating transaction. Preparation phase (Prepare): Each write master reserves resources and marks the affected transactions; Commit / rollback phase: if all nodes are prepared successfully, commit normally; otherwise, execute rollback; Compensation mechanism: For inconsistencies caused by business side side effects (such as external system calls) during rollback, introduce compensating transactions (Compensating Transaction) to process reversely step by step according to business definition. (3) Idempotency and atomicity guarantee. Each compensating transaction is designed as an idempotent operation, and a global transaction log is recorded to ensure that multiple executions or interrupted retries do not produce secondary side effects. After rollback, the monitoring platform confirms that all nodes have recovered to consistency, and cleans up the temporary markers.
[0123] In summary, based on the visual monitoring platform, automatic partition repair mechanism and distributed rollback mechanism, the operation of the distributed multi-master database can be controlled to achieve the following effects: (1) Shorten the fault detection and response time. The visual monitoring platform can display the real-time topology, and the average detection time of fault link or node anomaly is shortened from 5-10 minutes to <30 seconds. The multi-dimensional alarm strategy cooperates with the automatic script to make the average time from "alarm triggering" to "automatic disposal" decrease from 3 minutes to within 15 seconds. (2) Greatly reduce the cost of manual operation and maintenance. Automatic retry replication, write routing switching, container restart and other operations, the average number of daily manual interventions is reduced from about 10 times per week to <1 time. The operation and maintenance team can process >80% of common faults by automatic scripts, and can invest more energy into architecture optimization and capacity planning. (3) Improve the stability and availability of the cluster. Row-level conflict isolation and background automatic repair reduce the conflict row write suspension time from several hours to several seconds, and improve the overall write availability of the system by ≥99.99%. Distributed rollback combined with 2PC and compensation transactions ensures that after a cross-node failure or partial commit failure, the global data consistency recovery time is reduced from days to minutes. (4) Enhance system observability and decision support. The unified console aggregates all key indicators and event history, supports cross-time and cross-machine room backtracking analysis, and shortens the average fault cause positioning time from several hours to <30 minutes. Rich dashboard and report capabilities make capacity planning, hotspot analysis, and risk prediction more accurate, helping the team to prepare for capacity expansion and performance tuning in advance. (5) Improve business continuity and user experience. Business interruptions caused by human errors or replication link abnormalities are basically eliminated, and the annual unplanned downtime of the system is reduced by ≥90%. Through automatic operation and maintenance and rapid recovery, the business-side perceived availability SLA is improved from 9.9% to 99.99%, significantly improving the access reliability of end users.
[0124] Exemplarily, as shown in Figure 7 , a distributed consistency rollback timing diagram is provided. The monitoring platform can pull indicators from the monitoring agent, the monitoring agent can report data, the monitoring platform can execute the retry replication script by evaluating the alarm rule, and the returned execution result is obtained. Further, the write master node can also report conflict events to the monitoring platform, and the monitoring platform can perform row-level isolation and repair, initiate 2PC rollback through automatic scripts, and report to the monitoring platform after the rollback is completed.
[0125] In this embodiment, on the one hand, the real-time perception of the running state of the distributed multi-master database, the rapid alarm triggering, and the automatic repair and consistency rollback of the partition write conflict and failed transaction can be realized through the visual monitoring platform and the automatic script, so that the operation and maintenance cost is greatly reduced and the system reliability is improved. The visual monitoring platform can provide global situation awareness across nodes and across machine rooms, and realize "early discovery, quick positioning and automatic disposal" of problems through multi-dimensional alarm strategies. On the other hand, the automatic partition repair mechanism can use row-level locking and background merging tasks to automatically troubleshoot and repair conflict writes, avoiding global write stop. In another aspect, the distributed rollback mechanism can combine 2PC and business compensation transactions to ensure that the consistency of cluster data and the idempotency of the system are thoroughly guaranteed in any failure scenario.
[0126] The distributed multi-master database consistency processing method provided in the embodiment of the application has the following mechanisms: hierarchical mixed synchronous / asynchronous replication topology, dynamic replication mode switching mechanism, conflict detection of compressed vector clock combined with Bloom filter, multi-strategy business-aware conflict merging framework, intelligent write routing and local partitioning of hot keys, adaptive batch submission algorithm, incremental snapshot parallel synchronization and streaming incremental log playback, smooth online expansion and write routing sprint alignment, centralized topology monitoring and multi-dimensional alarm automation, row-level isolation repair and distributed consistency rollback. Specifically:
[0127] (1) Hierarchical mixed synchronous / asynchronous replication topology. The core nodes and the edge nodes are deployed hierarchically, and semi-synchronous and asynchronous replication modes are respectively adopted; the edge nodes are pushed asynchronously to the core write master after local fast confirmation, and then the core nodes feedback in batches, realizing the "eventual consistency + controllable lag" model.
[0128] (2) Dynamic replication mode switching mechanism. Based on real-time health indicators such as network RTT, packet loss rate and node pressure, the mode is automatically switched between semi-synchronous and asynchronous replication; the shortest mode retention time is introduced to avoid the jitter protection strategy of frequent switching.
[0129] (3) Conflict detection of compressed vector clock combined with Bloom filter. Only the active write master maintains a sparse compressed vector clock, and the inactive clock components are recycled regularly; Bloom filter is used to quickly determine the "no conflict" scenario, and only when the filter hits, the clock is accurately compared.
[0130] (4) Multi-strategy business-aware conflict merging framework. A plug-in merging interface is provided in the database kernel layer, supporting multiple strategies such as numerical accumulation, maximum / minimum value, JSON depth merging, etc.; CRDT automatic rollback mechanism is provided for complex data types, and the merging process is automatically switched according to the conflict complexity.
[0131] (5) Intelligent write routing and hotspot key local sharding. The client / middleware layer maintains a multi-dimensional routing score model based on EWMA smoothing, enabling dynamic write master selection; for write hot keys, sharding by sub-key and dispersing to multiple write masters avoids single-point overload.
[0132] (6) Adaptive batch commit algorithm. Based on queuing theory and delay model, the optimal batch size is calculated in real time, and dynamically increased or decreased within a given time window; supports "queue length trigger" and "timing trigger" two batch commit strategies, balancing throughput and response delay.
[0133] (7) Incremental snapshot parallel synchronization and streaming WAL replay. Snapshot slices are pulled and imported in parallel, and incremental WAL streaming replay is pushed as soon as it is generated; two-stage pipeline synchronization process shortens the cold start time to 30-35% of the original.
[0134] (8) Smooth online expansion and write routing sprint alignment. After the snapshot import is completed, temporarily adjust the replication confirmation ratio and write routing strategy to ensure strong consistency alignment for a short time; after the sprint is completed, restore the standard replication configuration to ensure smooth and non-perceptual expansion.
[0135] (9) Centralized topological monitoring and multi-dimensional alarm automation. Collect key indicators such as replication delay, conflict rate, and resource utilization, and display them in a dynamic topology graph; support threshold, trend, and composite alarms, and automatically trigger retry replication, write routing switching, and node restart scripts.
[0136] (10) Row-level isolation repair and distributed consistency rollback. When conflict merging fails, lock the specified row, and automatically repair and unlock it in the background combined with business plugins; combined with the distributed rollback mechanism of two-phase commit and compensation transactions, ensure atomicity and idempotency.
[0137] It should be understood that although each step in the flowchart involved in each of the above embodiments is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0138] Based on the same inventive concept, the embodiments of the present application also provide a distributed multi-master database consistency processing apparatus for implementing the distributed multi-master database consistency processing method described above. The apparatus provides a solution to the problem in a similar manner to the implementation described in the method above, and therefore the specific limitations in one or more distributed multi-master database consistency processing apparatus embodiments provided below can refer to the limitations of the distributed multi-master database consistency processing method described above, and will not be described here again.
[0139] In one exemplary embodiment, as shown in Figure 8 A distributed multi-master database consistency processing apparatus is provided, comprising: a request receiving module 802 and a request processing module 804, wherein:
[0140] The request receiving module is configured to receive a write request triggered for any write node in the distributed multi-master database.
[0141] The request processing module is configured to schedule a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database maintains data consistency during the processing of the write request. The hierarchical replication topology structure is constructed in the following manner: a plurality of core nodes and a plurality of edge nodes are determined from the plurality of write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; a semi-synchronous link is established between the plurality of core nodes, and for each edge node, an asynchronous replication channel is established between the edge node and each core node.
[0142] The distributed multi-master database consistency processing apparatus described above can receive a write request triggered for any write node in the distributed multi-master database, and then schedule a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database maintains data consistency during the processing of the write request. The hierarchical replication topology structure is constructed in the following manner: a plurality of core nodes and a plurality of edge nodes are determined from the plurality of write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; a semi-synchronous link is established between the plurality of core nodes, and for each edge node, an asynchronous replication channel is established between the edge node and each core node. Based on this, the distributed multi-master database consistency processing method described above can ensure the data consistency within the distributed multi-master database when processing the write request.
[0143] In one of the embodiments, the request processing module is specifically configured to: take the edge node receiving the write request as a first edge node; in the process of processing a write transaction corresponding to the write request by the first edge node, push the write transaction to a replication queue of each core node through an asynchronous replication channel established in the hierarchical replication topology for the first edge node; determine a proportion of core nodes that have processed the write transaction through a semi-synchronous link established in the hierarchical replication topology, and in the case that the proportion reaches a preset proportion, batch feedback an acknowledgement character to the first edge node through the core nodes that have processed the write transaction.
[0144] In one of the embodiments, the distributed multi-master database consistency processing apparatus further comprises a link switching module, which is specifically configured to: monitor evaluation indexes of the plurality of core nodes in the process of processing the write request; if it is determined according to the evaluation indexes that there is an abnormal core node, change the semi-synchronous link associated with the abnormal core node to an asynchronous replication channel until it is monitored that the abnormal core node returns to normal, and change the asynchronous replication channel obtained by the change to the semi-synchronous link.
[0145] In one of the embodiments, the distributed multi-master database is configured with a Bloom filter, a multi-strategy merging mechanism and an automatic rollback mechanism. The distributed multi-master database consistency processing apparatus further comprises a conflict processing module, which is specifically configured to: monitor write transaction processing frequencies of the plurality of write nodes respectively, adjust vector clocks of the plurality of write nodes respectively according to the write transaction processing frequencies; and process running conflicts of the plurality of write nodes according to the vector clocks of the plurality of write nodes respectively and the Bloom filter, the multi-strategy merging mechanism and the automatic rollback mechanism configured.
[0146] In one of the embodiments, the distributed multi-master database consistency processing apparatus further comprises a dynamic adjustment module, which is specifically configured to: monitor state indexes of the plurality of write nodes respectively, determine a traffic distribution scheme between the plurality of write nodes according to the state indexes; monitor pressure indexes of the plurality of write nodes respectively, determine a hot spot shard scheme between the plurality of write nodes according to the pressure indexes; monitor delay indexes of the plurality of write nodes respectively, determine a batch adjustment scheme of processing write requests of the plurality of write nodes according to the delay indexes; and dynamically schedule the plurality of write nodes to process write requests according to the traffic distribution scheme, the hot spot shard scheme and the batch adjustment scheme.
[0147] In one of the embodiments, the distributed multi-master database consistency processing apparatus further comprises a node expansion module, which is specifically configured to: if a new write-in node is started in the distributed multi-master database, synchronizing data in the distributed multi-master database to the new write-in node through full data snapshot and incremental log replay; and performing online switching processing on the new write-in node in the case that the full data snapshot has been completed and the incremental log replay has been replayed to a preset log, so as to keep the new write-in node and other write-in nodes in the distributed multi-master database in data alignment.
[0148] In one of the embodiments, the distributed multi-master database consistency processing apparatus further comprises a centralized regulation module, which is specifically configured to: constructing a visual monitoring platform for the distributed multi-master database, and configuring an automatic partition repair mechanism and a distributed rollback mechanism; and controlling running of the distributed multi-master database based on the visual monitoring platform, the automatic partition repair mechanism and the distributed rollback mechanism.
[0149] The modules in the distributed multi-master database consistency processing apparatus described above can be all or partially implemented by software, hardware and combinations thereof. The modules described above can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.
[0150] In one exemplary embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in Figure 9 The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store distributed multi-master database consistency processing data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a distributed multi-master database consistency processing method.
[0151] Those skilled in the art can understand that, Figure 9The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0152] In an exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the above embodiments when executing the computer program.
[0153] In an embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program is executed by a processor to implement the above embodiments.
[0154] In an embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the above embodiments.
[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0156] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0157] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.
[0158] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the patent scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for handling consistency in a distributed multi-master database, characterized in that, The method comprises: receiving a write request triggered for any write node in a distributed multi-master database; scheduling a plurality of write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database remains data consistent during processing of the write request; The hierarchical replication topology structure is constructed by: determining a plurality of core nodes and a plurality of edge nodes from the plurality of write nodes; the core nodes and the edge nodes meet a preset hierarchical deployment condition; establishing a semi-synchronous link between the plurality of core nodes, and for each edge node, establishing an asynchronous replication channel between the edge node and each core node.
2. The method of claim 1, wherein, The scheduling of the plurality of write nodes in the distributed multi-master database to process the write request through the hierarchical replication topology structure constructed for the distributed multi-master database comprises: taking an edge node receiving the write request as a first edge node; in the process of the first edge node processing a write transaction corresponding to the write request, asynchronously pushing the write transaction to a replication queue of each core node through the asynchronous replication channel established for the first edge node in the hierarchical replication topology structure; determining the proportion of core nodes that have processed the write transaction through the semi-synchronous link established in the hierarchical replication topology structure, and in the case that the proportion reaches a preset proportion, feeding back a confirmation character to the first edge node in batches through the core nodes that have processed the write transaction.
3. The method of claim 1, wherein, The method further comprises: monitoring evaluation indicators for the plurality of core nodes during processing of the write request; if there is an abnormal core node according to the evaluation indicators, changing the semi-synchronous link associated with the abnormal core node to an asynchronous replication channel until the abnormal core node is monitored to be normal, and restoring the asynchronous replication channel changed to a semi-synchronous link.
4. The method of claim 1, wherein, The distributed multi-master database is configured with a Bloom filter, a multi-strategy merging mechanism and an automatic rollback mechanism; the method further comprises: monitoring the write transaction processing frequency of each of the plurality of write nodes, and adjusting the vector clock of each of the plurality of write nodes according to the write transaction processing frequency; processing the running conflicts of the plurality of write nodes according to the vector clock of each of the plurality of write nodes, and the Bloom filter, the multi-strategy merging mechanism and the automatic rollback mechanism configured.
5. The method of claim 1, wherein, The method further comprises: monitoring state indicators of each of the plurality of write nodes, and determining a traffic distribution scheme between the plurality of write nodes according to the state indicators; monitoring pressure indicators of each of the plurality of write nodes, and determining a hot spot shard scheme between the plurality of write nodes according to the pressure indicators; monitoring delay indicators of each of the plurality of write nodes, and determining a batch adjustment scheme for the plurality of write nodes to process the write request according to the delay indicators; According to the traffic distribution scheme, the hotspot sharding scheme and the batch adjustment scheme, the multiple write nodes are dynamically scheduled to process the write request.
6. The method of claim 1, wherein, The method further comprises: If a new write node is started in the distributed multi-master database, the data in the distributed multi-master database is synchronized to the new write node through full data snapshot and incremental log replay; In a case where the full data snapshot has been completed and the incremental log replay has been replayed to a preset log, online switching processing is performed on the new write node, so that the new write node and other write nodes in the distributed multi-master database keep data alignment.
7. The method of claim 1, wherein, The method further comprises: A visual monitoring platform is constructed for the distributed multi-master database, and an automatic partition repair mechanism and a distributed rollback mechanism are configured; Based on the visual monitoring platform, the automatic partition repair mechanism and the distributed rollback mechanism, the operation of the distributed multi-master database is controlled.
8. A distributed multi-master database consistency processing apparatus, comprising: The apparatus comprises: The request receiving module is configured to receive a write request triggered for any write node in a distributed multi-master database; The request processing module is configured to schedule multiple write nodes in the distributed multi-master database to process the write request through a hierarchical replication topology structure constructed for the distributed multi-master database, so that the distributed multi-master database keeps data consistency in the processing of the write request; the hierarchical replication topology structure is constructed by: determining multiple core nodes and multiple edge nodes from the multiple write nodes; the core nodes and the edge nodes satisfy a preset hierarchical deployment condition; establishing a semi-synchronous link between the multiple core nodes, and establishing an asynchronous replication channel between each edge node and each core node for each edge node. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.