Data synchronization method for distributed system, device, storage medium, and system

By configuring different node weights and primary selection upper limits in a distributed system, decoupling the consensus conditions of the main selection stage and the data synchronization stage, the problems of low efficiency and poor availability caused by the coupling of consensus conditions in the existing technology are solved, and efficient data synchronization and service availability and quality improvement in cross-AZ disaster situations are achieved.

WO2025149829A1PCT designated stage expired Publication Date: 2025-07-17CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/063214
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-12-27
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the distributed system, the consensus condition coupling between the main selection stage and the data synchronization stage leads to low data synchronization efficiency and it is difficult to ensure high availability and data consistency in the case of cross-AZ disasters.

Method used

By configuring different node weights and primary selection upper limits in the availability zones of different regions, we decouple the consensus conditions of the main selection stage and the data synchronization stage, and use the strict consensus conditions of the main selection stage and the flexible consensus conditions of the data synchronization stage to ensure that the consensus conditions of the main selection stage are more stringent, the consensus conditions of the data synchronization stage are more relaxed, and the consensus conditions of the node set intersection are realized.

Benefits of technology

Improve data synchronization efficiency, enhance the availability and service quality of distributed systems in the case of cross-AZ disasters, reduce network latency, and improve the overall availability and response speed of services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024063214_17072025_PF_FP_ABST
    Figure IB2024063214_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a data synchronization method for a distributed system, a device, a storage medium, and a system. The distributed system comprises three AZs located in different regions, and different node weights are configured for nodes in the three AZs. When a target node meets a main node selection triggering condition, vote weights of a plurality of acquired votes are determined by means of the node weights of the nodes and an actual main node selection round; whether the target node can be a main node is determined on the basis of the vote weights of the plurality of votes and a preset first consensus condition; when the target node becomes the main node, the target node sends written data received from a client to other nodes; and if it is determined that a node that feeds back a confirmation message among the other nodes meets a second consensus condition, it is determined that the written data is successfully synchronized, wherein a node set meeting the first consensus condition and a node set meeting the second consensus condition have an intersection. By decoupling consensus conditions in the stages of main node selection and data synchronization, the data synchronization efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TECHNICAL FIELD The present disclosure relates to the field of cloud computing technology, and more particularly to a data synchronization method, device, storage medium, and system for a distributed system. Background: An availability zone (AZ) generally refers to a physical area within a region with independent power and network resources, typically a computer room where multiple servers and other physical devices are deployed. To ensure the disaster recovery capabilities of a distributed system, it is necessary to deploy a distributed system for a particular service across multiple AZs, each containing multiple servers. This ensures that, if a catastrophic event occurs in one AZ, services can continue to be provided normally in other AZs, thereby ensuring the availability of the distributed system. While ensuring cross-AZ disaster recovery, distributed systems must also guarantee data consistency. This means that all data replicas in the distributed system should have the same value at the same time (equivalent to all clients accessing the same, up-to-date copy of the data), ensuring that clients can obtain the latest, consistent data in the distributed system. Data consistency is typically achieved through consensus protocols. Traditional consensus protocols go through two phases when reaching consensus on a proposal: the Prepare phase and the Accept phase. The Prepare phase is used to identify nodes with proposal rights, while the Accept phase is used to reach consensus on the specific proposal. Reaching consensus on such a proposal requires at least two network hops. To improve consensus efficiency, a node is typically designated as having exclusive proposal rights. This node can then continuously organize the Accept phase to reach consensus on the proposal. This type of consensus protocol is also known as a strong-leader consensus protocol. In the original protocol, the Prepare and Accept phases are also referred to as the leader election phase and data synchronization phase. The leader election phase of the strong-leader consensus protocol is similar to the Prepare phase, but the goal shifts from determining proposal rights to determining the leader node. Subsequent proposals can be directly assigned proposal numbers by the master node, eliminating the need for inter-node interaction. The data synchronization phase is consistent with the Accept phase, meaning consensus is reached when a majority of nodes accept the proposal. In traditional strong-leader consensus protocols, the quorum consensus model of the master election and data synchronization phases is coupled, requiring acceptance by a majority of members (i.e., more than half of the total number of server nodes) for the phase to be considered successful. In this article, quorum refers to the majority condition of the server node set, meaning that a majority of more than half of the server nodes must meet this condition.For example, assume a distributed system includes five nodes. During the leader election phase, a node can only become the leader if it receives at least three votes. During the data synchronization phase, data synchronization is considered successful only if at least three nodes have successfully synchronized data. The quorum consensus model in the leader election and data synchronization phases has a node intersection condition: if the majority number in the leader election phase + the majority number in the data synchronization phase > the total number of nodes, then this intersection condition is met. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a distributed system data synchronization method, device, storage medium, and system to improve the high availability of distributed lock services. In a first aspect, an embodiment of the present disclosure provides a distributed system, comprising: at least three availability zones located in different regions, nodes in different availability zones being configured with different node weights; a target node, configured to, when a master election trigger condition is met, determine voting weights of multiple votes obtained from multiple nodes based on the node weights corresponding to the at least three availability zones and an actual master election round of each node, and determine whether the target node can become a master node based on the voting weights of the multiple votes and a preset first consensus condition, the target node being any node in the at least three availability zones; the target node being further configured to, after becoming the master node, send write data received from a client to other nodes, and determine that synchronization of the write data is successful if it is determined that a node among the other nodes that feeds back a confirmation message corresponding to the write data meets a second consensus condition; wherein the set of nodes meeting the first consensus condition and the set of nodes meeting the second consensus condition have an intersection.In a second aspect, an embodiment of the present disclosure provides a distributed system, comprising: R availability zones located in different regions, each availability zone including N nodes, where R is greater than or equal to 3 and N is greater than or equal to 3; nodes in a first availability zone are configured with a first master election round upper limit and a first node weight L, wherein the first master election round upper limit prevents nodes in the first availability zone from being elected as master nodes before the first master election round upper limit is reached; nodes in a second availability zone are configured with a second master election round upper limit and a second node weight, wherein the second node weight is determined based on the second master election round upper limit, the first node weight, a total number of nodes N in the availability zone, and a majority number of nodes in the availability zone, and the second node weight is less than the first node weight; the second availability zone is any availability zone among the R availability zones except the first availability zone, the second master election round upper limit is greater than or equal to the first master election round upper limit, and the second master election round upper limit prevents nodes in the second availability zone from being elected as master nodes before the second master election round upper limit is reached; target nodes in the R availability zones are configured to, when a master election trigger condition is met, combine the first node weight, The second node weight and the actual master election round of each node are used to determine the voting weights of the multiple votes obtained, a voting weight statistical value in each availability zone is determined based on the voting weights of the multiple votes, and whether the target node can become the master node is determined based on the voting weight statistical value in each availability zone and a preset first consensus condition; the target node is configured to send write data received from the client to other nodes after becoming the master node, and determine that the write data synchronization is successful if a node among the other nodes that feeds back a confirmation message corresponding to the write data satisfies the second consensus condition; wherein the set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition have an intersection.In a third aspect, an embodiment of the present disclosure provides a data synchronization method for a distributed system, which is applied to a target node in a target availability zone, where the target availability zone is any one of at least three availability zones located in different regions, and nodes in different availability zones are configured to have different node weights; the method includes: when a master election trigger condition is met, determining the voting weights of multiple votes obtained from multiple nodes in combination with the node weights corresponding to the at least three availability zones and the actual master election round of each node; determining whether the target node can become a master node based on the voting weights of the multiple votes and a preset first consensus condition; if the target node becomes the master node, sending write data received from a client to other nodes; if it is determined that a node among the other nodes that feeds back a confirmation message corresponding to the write data meets a second consensus condition, determining that the write data synchronization is successful; wherein the set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition have an intersection. In a fourth aspect, an embodiment of the present disclosure provides a data synchronization device for a distributed system, applied to a target node in a target availability zone, where the target availability zone is any one of at least three availability zones located in different regions, and nodes in different availability zones are configured to have different node weights; the device includes: a voting acquisition module, configured to determine, when a master election trigger condition is met, the voting weights of multiple votes obtained from multiple nodes in combination with the node weights corresponding to the at least three availability zones and the actual master election round of each node; a role determination module, configured to determine whether the master node can become the master node based on the voting weights of the multiple votes and a preset first consensus condition; a data synchronization module, configured to send write data received from a client to other nodes if the target node becomes the master node; and, if it is determined that a node among the other nodes that feeds back a confirmation message corresponding to the write data meets a second consensus condition, determining that the write data synchronization is successful; wherein the set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition have an intersection. In a fifth aspect, embodiments of the present disclosure provide an electronic device comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the processor executes the executable code, the processor is enabled to implement at least the distributed system data synchronization method described in the third aspect. In a sixth aspect, embodiments of the present disclosure provide a non-transitory machine-readable storage medium, wherein the non-transitory machine-readable storage medium stores executable code, and when the processor of the electronic device executes the executable code, the processor is enabled to implement at least the distributed system data synchronization method described in the third aspect.In embodiments of the present disclosure, a distributed system providing a service may include multiple availability zones located in different regions. Each availability zone may include multiple nodes providing the service. For example, a first availability zone and a second availability zone located in a first region, and a third availability zone located in a second region, with at least three nodes deployed in each availability zone. Weights may be assigned to the nodes in different availability zones based on environmental factors such as their location. For example, nodes in the first availability zone are assigned a first node weight, nodes in the second availability zone are assigned a second node weight, and nodes in the third availability zone are assigned a third node weight. In this deployment scenario, nodes in the first and second availability zones may be nodes providing services under normal circumstances, while the third availability zone serves as a disaster recovery mechanism and is only used in special circumstances, such as when disasters occur in other availability zones. Therefore, the third node weight may be configured to be smaller than the first and second node weights, while the first node weight may be greater than or equal to the second node weight. During the leader election phase, when any target node in multiple availability zones meets the leader election trigger conditions, the voting weights of multiple votes received from multiple nodes are determined based on the node weights corresponding to the different availability zones and the current actual leader election round of each node. Based on these voting weights and the preset first consensus condition, the target node is determined to be a leader. Therefore, the first consensus condition used in the leader election phase is a more stringent condition related to the voting weights determined based on the node weights. After becoming the leader, the target node sends write data received from the client to other nodes. If the node that responds with a confirmation message corresponding to the write data meets the second consensus condition, the write data synchronization is considered successful. The set of nodes meeting the first consensus condition and the set of nodes meeting the second consensus condition intersect. Therefore, the second consensus condition used in the data synchronization phase is related to the number of nodes that successfully synchronized the write data, thus decoupling the consensus conditions of the data synchronization phase from those of the leader election phase. Decoupling the consensus conditions of these two phases facilitates flexible control of the leader election process and results during the leader election phase. This allows for stricter consensus conditions during the leader election phase, thereby relaxing the consensus conditions during the data synchronization phase, enabling more efficient data synchronization. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the figures used in describing the embodiments. Obviously, the figures described below represent some embodiments of the present disclosure. Persons skilled in the art can derive other figures based on these figures without inventive effort.Figure 1 is a schematic diagram of the composition of a distributed system provided by an embodiment of the present disclosure; Figure 2 is a schematic diagram of a consensus condition in the leader election phase provided by an embodiment of the present disclosure; Figure 3 is a schematic diagram of another consensus condition in the leader election phase provided by an embodiment of the present disclosure; Figure 4 is a schematic diagram of a consensus condition in the leader election and data synchronization phases provided by an embodiment of the present disclosure; Figure 5 is a schematic diagram of a consensus condition in the leader election and data synchronization phases in an AZ disaster scenario provided by an embodiment of the present disclosure; Figure 6 is a flow chart of a leader election phase provided by an embodiment of the present disclosure; Figure 7 is a schematic diagram of the structure of a data synchronization device in a distributed system provided by an embodiment of the present disclosure; and Figure 8 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS To further clarify the objectives, technical solutions, and advantages of the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings. It should be understood that the described embodiments represent only a portion of the embodiments of the present disclosure, but not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in the embodiments of this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. The following detailed description of some embodiments of this disclosure is provided in conjunction with the accompanying drawings. The following embodiments and features can be combined unless there is a conflict between them. Furthermore, the sequence of steps in the following method embodiments is merely illustrative and not a strict limitation. The following terms and concepts used in the embodiments of this disclosure are explained: Consensus protocols, such as Paxos and Raft, require a quorum to reach consensus on a proposal and ensure data consistency among different nodes. A quorum refers to a collection of server nodes. Generally, the quorum must be greater than half the number of server nodes, which is the so-called majority. Multi-AZ disaster recovery: To ensure the disaster recovery capabilities of a distributed system, a distributed system is deployed across multiple AZs. If a disaster occurs in one AZ, systems in other AZs can continue to provide services, thus ensuring the availability of the distributed system. This refers to a distributed system that provides a specific service, such as storage services or audio and video services.Decoupling leader election and synchronization: Strong leader consensus protocols are generally designed to couple the quorum consensus requirements of the leader election phase with those of the data synchronization phase. This means that a majority (i.e., a set of nodes greater than half the number of server nodes) must accept the consensus for both phases to succeed. Decoupling leader election and synchronization means that the leader election and data synchronization phases use different consensus criteria. They do not necessarily need to adhere to the quorum requirement. Instead, they only need to satisfy the node intersection requirement between the leader election and data synchronization consensus requirements. Specifically, the set of nodes meeting the leader election consensus requirement and the set of nodes meeting the data synchronization consensus requirement must intersect. To achieve zone-based disaster tolerance, distributed systems providing a service are typically deployed across at least three zones. Furthermore, to ensure that in the event of an AZ-level disaster (all nodes in one of the three AZs fail), the remaining AZs can still tolerate the failure of one node. A common approach is to deploy three nodes (i.e., the server nodes providing the service) in each AZ. In this way, if three AZs each contain three nodes, even if three nodes in one AZ fail and a node in another AZ also fails, the remaining five nodes still meet the quorum (more than half of the total number of nodes) requirement, and the service remains available. Due to limitations in data center construction, the simultaneous provision of three independent AZs in the same region is not guaranteed. Often, distributed systems must cope with situations where only two AZs exist in the same region, with the remaining AZ deployed in another region. Therefore, the disclosed embodiments propose a multi-AZ disaster recovery solution based on a strong leader consensus protocol that decouples leader election from synchronization. This solution allows the distributed system to be flexibly configured based on factors such as the actual AZ's geographic distribution, power supply, and network, while also reducing the impact of cross-AZ network latency on service quality. Specifically, nodes in each AZ are tagged with a weight related to their AZ environment. During the leader election and data synchronization phases, different quorum consensus modes are used, taking into account the node weights associated with the node's AZ. This decouples the consensus conditions during the leader election and data synchronization phases. The strong leader consensus protocol, which decouples leader election from data synchronization, allows for targeted control over the leader election process and results, accelerating data synchronization and ensuring improved service quality. Furthermore, by incorporating richer weight information, consensus conditions during leader election can be more flexibly configured based on actual needs and deployment environments, adapting to cross-region, multi-AZ deployments and daily changes.The following describes a distributed system and a data synchronization method within the distributed system provided by an embodiment of the present disclosure. The distributed system provided by an embodiment of the present disclosure includes at least three availability zones located in different regions, and nodes in different availability zones are configured with different node weights. When any target node in the at least three availability zones meets the master election trigger condition, it initiates a voting request, thereby receiving votes from multiple nodes (including itself). The target node determines the voting weights of multiple votes received from multiple nodes based on the node weights corresponding to the at least three availability zones and the actual master election round of each node. It then determines whether the target node can become the master node based on the voting weights of the multiple votes and a preset first consensus condition. After determining that it has become the master node, the target node sends the write data received from the client to other nodes. If it determines that the node that has responded with a confirmation message corresponding to the write data meets the second consensus condition, the write data synchronization is successful. The set of nodes that meet the first consensus condition (i.e., the multiple nodes that have provided votes) and the set of nodes that meet the second consensus condition (i.e., the nodes that have provided confirmation messages) intersect. As can be seen, the first consensus condition is a more stringent condition related to voting weight, while the second consensus condition is a traditional "majority" condition, achieving decoupling between the two. The above-described distributed system is illustrated in detail with reference to the accompanying figures. Figure 1 is a schematic diagram of the composition of a distributed system provided in an embodiment of the present disclosure. As shown in Figure 1, the distributed system providing a service includes a first AZ (AZ1) and a second AZ (AZ2) located in a first region, as well as a third AZ (AZ3) located in a second region. As shown in Figure 1, assume that AZ1 includes the following three nodes: node a1, node a2, and node a3; AZ2 includes the following three nodes: node b1, node b2, and node b3; and AZ3 includes the following three nodes: node c1, node c2, and node c3. Each of these nine nodes has a service deployed, forming a distributed system located in different AZs for client access. It is understandable that, based on traditional consensus protocols, this distributed system can tolerate the failure of four nodes. That is, as long as five of the nodes can provide services normally, the availability of the distributed system is guaranteed. In the disclosed embodiment, node weights can be set for each of the three AZs. In an optional embodiment, nodes in the same AZ have the same node weight. Here, it is assumed that nodes in AZ1 are configured with a first node weight X1, nodes in AZ2 are configured with a second node weight X2, and nodes in AZ3 are configured with a third node weight Y.In another optional embodiment, the nodes in the same AZ can also be configured with different node weights based on differences in their computing power performance. In actual applications, the node weights of different AZs can be set according to which AZ is more likely to generate a master node (leader). In the case where AZ1 and AZ2 are located in a certain region and AZ3 is located in another region, AZ3 is often used as a backup AZ, that is, when AZ1 or AZ2 fails, AZ3 is used to provide services to the client. Therefore, it is hoped that when all three AZs are normal, the master node is generated in AZ1 and AZ2, but not in AZ3. Based on this, the third node weight Y of the node in AZ3 can be set smaller than the first node weight X1 and the second node weight X2 corresponding to the nodes in AZ1 and AZ2, that is, Y <X1 , Y<X2。 O In actual applications, AZ1 and AZ2 are often located closer to clients than AZ3. Therefore, master nodes are generated in AZ1 and AZ2, and client write data is written through the master node, which allows clients to experience better network latency. In an optional embodiment, the basis for determining node weights in an AZ can be based on the region to which the AZ belongs. In this case, nodes in each AZ in the same region have the same node weight. That is, the first node weight X1 and the second node weight X2 corresponding to the nodes in AZ1 and AZ2 are equal, expressed as X, Y <X O In another optional embodiment, the basis for determining the weight of nodes in an AZ may include other environmental information, such as power information and network information, in addition to the region to which the AZ belongs. In this case, nodes in different AZs in the same region may have different node weights, that is, the first node weight X1 and the second node weight X2 corresponding to nodes in AZ1 and AZ2 are not necessarily equal. For example, when considering power information, if the power system in AZ1 is more stable, then X1 can be set to > X2. For another example, when considering network information, if the network quality of the network environment in AZ1 is better, then X1 can be set to > X2. OBased on the above description, it is assumed in the embodiment of the present disclosure that: the first node weight X1 corresponding to the node in AZ1 is greater than or equal to the second node weight X2 corresponding to the node in AZ2, and both X1 and X2 are greater than the third node weight Y corresponding to the node in AZ3. In an optional embodiment, in addition to configuring the node weights differently for nodes in different AZs, different configurations can be made for the upper limit of the number of master rounds for nodes in different AZs to control the difficulty of nodes in different AZs becoming master nodes. For example, in the case of X1 = X2, the upper limit of the number of master rounds for nodes in AZ1 and AZ2 can be set to a first value, which can be 1, and the upper limit of the number of master rounds for nodes in AZ3 can be set to a second value, which is smaller than the second value, such as the second value is 3. OThe upper limit on the number of consecutive votes a node can initiate. Furthermore, the upper limit also prevents a node from being elected as a master node until the number of votes initiated reaches this upper limit. Therefore, by configuring the upper limit on the number of votes initiated and the node weight, the difficulty of nodes becoming master nodes in different AZs can be controlled. For example, in AZs where a master node is more likely to be generated, nodes in these AZs can be configured with higher node weights and lower upper limits on the number of rounds required for master node generation. Conversely, in AZs where a master node is less likely to be generated, nodes in these AZs can be configured with lower node weights and higher upper limits on the number of rounds required for master node generation. This allows nodes in these AZs to become master nodes only after multiple rounds of master node generation. In practical applications, relevant personnel can set the highest node weight for the AZ where the master node is most likely to be located and set the upper limit on the number of rounds required for master node generation to 1. This reduces the probability of nodes in these AZs being successfully elected as master nodes in fewer rounds. Regarding the node weights and leader election round caps for nodes in other AZs, it is optional to first set a leader election round cap for nodes in other AZs (the value should be greater than or equal to 1). The corresponding node weights of nodes in other AZs are then determined based on the maximum node weight and the leader election round caps set for nodes in other AZs. For example, the third node weight Y is determined based on the first node weight X1 and a second value, assuming that the first node weight X1 is the maximum node weight. Alternatively, when there is no need to distinguish between the node weights and leader election round caps for some AZs (e.g., two AZs), the node weights and leader election round caps for these AZs can be set to the same value. It should be noted that if a node in a particular AZ is configured with the maximum node weight and a leader election round cap of 1, even if the leader election round caps for nodes in other AZs are also set to 1, the node weights of the nodes in these other AZs may not necessarily be the maximum node weights. This weight needs to be flexibly determined based on the first consensus condition used in the leader election phase. Based on the above configuration of node weights and the upper limit of node leader election rounds, the leader election process for the three AZs can be summarized as follows: When the target node meets the leader election trigger conditions, it combines the first node weight X1, the second node weight X2, the third node weight Y, and the actual leader election rounds of each node to determine the voting weights of multiple votes obtained. Based on the voting weights of multiple votes, it determines the voting weight statistics within each AZ. Based on the voting weight statistics within each AZ and the preset first consensus condition, it determines whether the target node can become the leader node.After the target node becomes the master node, the data synchronization process can be summarized as follows: the target node sends the write data received from the client to other nodes. If it is determined that the node among the other nodes that has responded with a confirmation message corresponding to the write data meets the second consensus condition, the write data synchronization is considered successful. The set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition intersect. The target node is any node in the three aforementioned AZs. In the case where each of the three AZs contains three nodes, and the first node weight X1 corresponding to the node in AZ1 is the maximum node weight, the first consensus condition can be set as: the voting weight statistics in at least two AZs are at least twice the first node weight X1; the second consensus condition is: at least two nodes in each of the at least two AZs respond with the confirmation message. Therefore, the first consensus condition for a node to become the master node actually includes two dimensions: the number of AZs and the voting weight within the AZs. Regarding the number of AZs, a majority condition must be met: a majority in at least two AZs is a majority in three AZs. Voting weight within an AZ can be determined based on the maximum node weight (X1 in this example) and the majority of nodes within the AZ. A majority of three nodes is at least two. Therefore, a 2X1 condition can be set for voting weight within the AZ. Consequently, the first consensus condition is: within at least two AZs, the voting weight statistics (i.e., the cumulative sum) must be at least twice the first node weight, X1. The second consensus condition is independent of node weight and is based solely on the majority of AZs and the number of nodes within them: at least two AZs, each with at least two nodes. In the disclosed embodiments, the concept of voting weight refers to the value a node assigns to a vote received from another node. Voting weight is related to the node weight and the node's current actual leader election round. In one optional embodiment, the voting weight of a vote can be determined according to the following formula: Voting weight = max(node ​​weight of the DST node * actual leader election round of the D node, node weight of the SRC node * actual leader election round of the D node). The DST node is the node that initiates the vote, and the SRC node is the node that casts the vote, i.e., the source node of the vote. In another optional embodiment, the voting weight of a vote can be determined according to the following formula: Voting weight = node weight of the D node * actual leader election round of the D node.When any target node in any of the three AZs meets the master election trigger conditions, it can initiate a vote with the remaining nodes (i.e., send a voting request to each node). The voting weight of each received vote is calculated according to the above formula. The voting weight statistics for each AZ are then tallied to determine whether the corresponding voting weight statistics for each AZ meet the aforementioned 2x1 condition. Furthermore, the number of AZs that meet this 2x1 condition is determined to be at least two. If so, the target node is elected as the master node. In practical applications, the voting weight statistics within an AZ can be the cumulative sum of the voting weights of the votes given by the nodes within that AZ. After the target node becomes the master node, client write data will be sent to the target node, which must synchronize this write data with other nodes to complete the data synchronization phase. Under traditional consensus protocols, synchronization is considered successful only after the target node has synchronized the write data with a majority of nodes. For example, if each of the three AZs contains three nodes, the target node needs to receive confirmation messages from at least five nodes confirming the successful synchronization of the write data to confirm that data synchronization is successful. In the disclosed embodiment, based on the aforementioned second consensus condition, the target node only needs to receive confirmation messages from at least four nodes confirming successful data synchronization to determine that data synchronization has been successful. These four nodes are located in at least two AZs, with at least two nodes in each AZ. Therefore, based on the data synchronization solution provided by the disclosed embodiment, the distributed system can tolerate the downtime of five nodes (compared to the traditional tolerance of only four nodes), thereby enhancing the availability of the distributed system. As can be seen from the aforementioned first and second consensus conditions, the disclosed embodiment employs two different consensus conditions in the leader election and data synchronization phases: one based on node weights and the other based on a majority consensus condition, thereby decoupling the consensus conditions of the two phases. Furthermore, leveraging the lower probability of leader election events compared to data synchronization events, the first consensus condition in the leader election phase, which is more stringent than the majority condition in traditional consensus protocols, can reduce the impact of AZ disasters on service availability while ensuring service quality under normal conditions (when all AZs are available). The lower probability of a master election event compared to a data synchronization event means that after a node becomes the master, it often takes some time for the conditions for another master election to be met. During this period, a large amount of data is often written. The impact of AZ disasters on service availability is reduced primarily because more node failures can be tolerated during the data synchronization phase without affecting service availability.Ensuring service quality under normal conditions primarily involves using the first consensus condition to generate the primary node in AZ1 and AZ2, which are closer to the client, thereby reducing network response latency. Because the nodes in AZ1 and AZ2 have greater weights, according to the aforementioned voting weight calculation formula, nodes in AZ1 and AZ2 are more likely to meet the 2x1 condition than nodes in AZ3. For ease of understanding, the following illustrates the execution process of the primary election and data synchronization phases in the three-AZ scenario described above, using Figures 2 and 3. In Figures 2 and 3, it is assumed that nodes in AZ1 and AZ2 have the same node weight X, and nodes in AZ3 have a node weight Y, with X > Y. Furthermore, it is assumed that the maximum number of primary election rounds for nodes in AZ1 and AZ2 is 1, and the maximum number of primary election rounds for nodes in AZ3 is Tmax, with Tmax being greater than or equal to 1. That is, each time a node in AZ3 initiates a vote, its actual primary election round increases by 1 until the election is complete and Tmax is reached, at which point its actual primary election round is reset to 0. In Figures 2 and 3, the first consensus condition in the leader election phase is that the total voting weight in each of at least two AZs is no less than 2X. The second consensus condition in the data synchronization phase is that at least two nodes in each of at least two AZs provide confirmation messages indicating that data synchronization has been successful. In addition, in Figures 2 and 3, the voting weight formula is assumed to be: Voting Weight = max (DST node's node weight * D node's actual leader election round, SRC node's node weight * SRC node's actual leader election round). For example, in the first leader election round, if a node in AZ2 receives a vote from a node in AZ3, the voting weight of this vote is max (X, Y) = X. If a node in AZ3 receives a vote from a node in AZ1 or AZ2, the corresponding voting weight is also max (Y, X) = X. If a node in AZ3 receives a vote from another node in AZ3, the corresponding voting weight is max (Y, Y) = Y. Under the aforementioned upper limit on leader election rounds, as the actual leader election round increases, the voting weight of votes received by a node in AZ3 from other nodes in AZ3 will also increase. For example, if a node receives a vote from another node in AZ3 during the second leader election round, the corresponding voting weight will be updated to max (2*Y, 2*Y) = 2Y. OOf course, when the node's actual master election rounds reach the upper limit of Tmax, its voting weight is updated to max (Tmax*Y, Tmax*Y) = Tmax*Y. Furthermore, Tmax*Y is less than or equal to X. That is, when Tmax is reached, the voting weight cannot exceed the maximum node weight X. Therefore, TmaxWX / Y is rounded up. Figure 2 illustrates an example where the master node falls in AZ1 or AZ2. If a node in AZ1 or AZ2 receives a vote in the master election box, it accumulates the voting weight in each AZ for which it voted, and thus obtains the AZ views shown in the figure (AZ1 View and AZ2 View). Consensus is achieved in both AZ1 and AZ2, where the cumulative voting weight reaches 2X. Specifically, assume that node a1 in AZ1 sends a voting request to nodes in all AZs, and the actual leader election round is updated to 1. Also assume that, as shown in Figure 2, node a1 receives votes from multiple nodes in the leader election box (nodes a1, a2, b1, and b2, each with an actual leader election round updated to 1). The voting weight of each vote is calculated using the above formula. Taking the vote of node a2 for node a1 as an example: Voting weight = (^X(X*1, X*1) = X, that is, the actual master election round of nodes a1 and a2 at this time is 1. The voting weight of the votes given by other nodes is also X. It can be understood that node a1 will vote for itself. Therefore, as shown in Figure 2, the cumulative sum of the voting weights of node a1 in each AZ for its own votes is obtained as AZ1 View: The total voting weight in AZ1 is 2X, and the total voting weight in AZ2 is 2X, which meets the first consensus condition. Therefore, node a1 can become the master node in the first round of master election initiated by it. Similarly, if a node in AZ2 (such as node b1) initiates the vote, based on the votes in the selection box in the figure, we can obtain AZ2 View: The total voting weight in AZ1 is 2X, and the total voting weight in AZ2 is 2X, which meets the first consensus condition. Therefore, node b1 can become the master node in the first round of master election initiated by it. It should be noted that After receiving the above four votes, node a1 can determine that the first consensus condition has been met. At this time, it does not need to wait for votes from other nodes and can directly perform subsequent actions, such as sending heartbeat packets to all other nodes.In addition, since the nodes in AZ1 and AZ2 have the same node weight and upper limit for the number of master election rounds, any node selected in the master selection box in Figure 2 will have the AZ View shown in Figure 2 if it receives votes from all the nodes selected in the box. Figure 3 illustrates an example where the master node falls in AZ2 or AZ3. AZ2 has a node weight of X and a master election round upper limit of 1, and AZ3 has a node weight of Y and a master election round upper limit of Tmax. If a node in AZ2 (node ​​b2 or node b3) receives votes from all the nodes in the master selection box (including nodes b2, b3, c1, c2, and c3) during the first round of master election it initiates (as will be explained below, the actual master election rounds of other nodes are also updated to 1 at this time), the AZ2 View shown in Figure 3 can be obtained: the sum of the voting weights in AZ2 is 2X, and the sum of the voting weights in AZ3 is 3X. This satisfies the first consensus condition, and the node becomes the master node. If a node in AZ3 (node ​​c1, node c2, or node c3) receives votes from all nodes in the leader election frame, the AZ3 View shown in Figure 3 can be obtained: the sum of the voting weights in AZ2 is 2X, and the sum of the voting weights in AZ3 is 3YT, where T is the current actual leader election round of the node that currently triggers the leader election in AZ3, and does not exceed the upper limit of the leader election round Tmax. Therefore, by setting YG [2X / 3, X), the node in AZ3 must receive all three votes in AZ3 in the first leader election round (Tmax=1) to become the leader node in this round. Furthermore, if we want to further increase the difficulty of electing a node in AZ3 as the master, for example, by requiring it to be elected only in the pth round (where p>1), we can set [2X / (3*p), 2X / (3*p-1)]. This ensures that during the first pT rounds of master election, nodes in AZ3 fail to meet the first consensus condition and are therefore not elected as the master. However, nodes in other AZs that trigger a master election during this period will generally succeed, thus making it difficult for AZ3 to be elected as the master. This property is automatically maintained regardless of whether all three AZs are operating normally or in the event of an AZ-level failure. Figure 4 illustrates the consensus conditions during the master election and data synchronization phases, assuming all three AZs are operating normally.As shown in Figure 4, AZ1 has a node weight of X, AZ2 has a node weight of X, and AZ3 has a node weight of Y. Assume that node b2 in AZ2 receives votes from all nodes in the leader selection box (including nodes b2, b3, c1, c2, and c3) during the first round of leader election it initiates. By accumulating the vote weights in each AZ and determining that node b2 meets the first consensus condition, it becomes the leader. Subsequently, after receiving the client's write data, node b2 synchronizes the write data to the other nodes in the three AZs. Assuming that node b2 receives confirmation messages for the write data from all nodes in the synchronization box (including nodes a1, a2, b1, and b2), node b2 determines that the data synchronization phase is complete because it has received confirmation messages from two nodes in each of the two AZs, meeting the second consensus condition. The set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition intersect with node b2, which also satisfies the condition that the nodes participating in the leader election and data synchronization phases intersect. Figure 5 illustrates the disaster recovery mechanism for leader election and data synchronization between AZ1 and AZ3, assuming a disaster in AZ2 causes all nodes to become unavailable. AZ1 has a node weight of X, AZ2 has a node weight of X, and AZ3 has a node weight of Y. Continuing from Figure 4, the failure of node b2 (old leader) in AZ2 triggers nodes in AZ1 and AZ3 to re-initiate leader election. Referring to the leader election process for nodes in AZ2 and AZ3 illustrated in Figure 3 and the data synchronization process illustrated in Figure 4 , the consensus conditions for leader election and data synchronization can also be achieved in this AZ disaster scenario. See Figure 5 for the leader election dialog (including nodes a2, a3, c1, c2, and c3) and the synchronization dialog (including nodes a1, a2, c1, and c2), which are not further described here. Figure 6 is a flowchart of the leader election phase provided in an embodiment of the present disclosure. As shown in Figure 6 , the leader election phase may include the following steps:

[0002] 51. When the target node meets the master election trigger condition, the target node increments its actual master election round (turn) and first term (term) identifier. The actual master election round is incremented up to the target node's corresponding master election round upper limit. The target node's actual master election round is incremented up to the target node's corresponding master election round upper limit. That is, subsequent steps are executed only when the incremented actual master election round does not exceed the target node's corresponding master election round upper limit. Otherwise, the process ends and the target node cannot become the master node. The master election trigger condition can include two scenarios: If a master node has been previously elected, the target node determines that the master election trigger condition has been met if it times out without receiving a heartbeat packet from the master node. Alternatively, if a master node has not been previously elected, the target node is configured with a set term. After the term expires, if no voting request is received from other nodes, the target node determines that the master election trigger condition has been met. In this embodiment, the term can be understood as the node's master election period. Its function and meaning can be further understood with reference to existing related technologies.

[0003] 52. The target node obtains a first data identifier corresponding to the target node's latest data. In practical applications, after receiving a write receipt from the client, the node first writes the written data to a local log file, creating a corresponding data identifier, or log index, in the log file. In this embodiment, the first data identifier may be the data identifier corresponding to the last written data entry in the target node's log file.

[0004] 53. The target node sends a voting request including the first term identifier and the first data identifier to other nodes. The other nodes are nodes in the multiple AZs except the target node.

[0005] 54. The other nodes increase their actual leader election rounds (not exceeding their own leader election round upper limit), confirm that their second term identifiers are not greater than their first term identifiers, and that their second data identifiers are earlier than their first data identifiers. After receiving the target node's voting request, the other nodes first increase their actual leader election rounds. When they confirm that their own term identifiers are not greater than the target node's term identifiers, they update their own term identifiers to match the target node's term identifiers. Furthermore, they determine whether the second data identifier corresponding to the last write data in their local log files is earlier than the first data identifier. If the second data identifier is earlier than the first data identifier, the write data corresponding to the first data identifier is the latest write data and requires synchronization. Otherwise, if the second data identifier is later than the first data identifier, the write data corresponding to the first data identifier is old data and may already be stored in the other nodes, requiring no synchronization. Therefore, the other nodes will only execute step S5 if they determine that the current write data is the latest write data requiring synchronization and that they have voting rights (the term identifiers and actual leader election rounds meet the requirements). Otherwise, the process ends.

[0006] 55. Other nodes send a voting notification message to the target node. The voting notification message includes the current actual leader election round and node weight of other nodes.

[0007] 56. The target node determines the voting weight for the corresponding vote based on the voting notification message, sums the voting weights for each AZ, and determines whether it can become the master node. Determining whether to become the master node, i.e., determining whether the first consensus condition is met, is described in the previous embodiment.

[0008] 57. After the target node becomes the master node, the actual master election round of the target node is reset. The actual master election round is reset to Oo

[0009] 58. After becoming the master node, the target node sends heartbeat packets to other nodes. To let other nodes know the status of the master node, the master node will periodically send heartbeat packets to other nodes.

[0010] 59. Other nodes determine that the target node has become the master node and reset their actual master election rounds. The actual master election rounds are reset to 0.

[0011] S10. After receiving the client's write data, the target node, becoming the master node, performs the data synchronization phase. The following briefly describes the data synchronization process between the target node and other nodes: After receiving the write data, the target node writes it to the local log file. At this point, the write data is in an uncommitted state. The target node then sends an append request containing the write data to other nodes. If the other nodes determine that the write data has not been written to their local log files, meaning there is no data conflict, they write the write data to their local log files. At this point, the write data is also in an uncommitted state on the other nodes. After writing the write data to the local log files, the other nodes send confirmation messages back to the target node. Based on the number of confirmation messages received and the zone to which it belongs, the target node changes the state of the write data to committed. At this point, the target node can return the write result to the client. The target node then sends another append request to the other nodes. Based on this request, the other nodes modify the status of the written data to the confirmed state. This completes the log replication process, or data synchronization, between the target node (the master node) and the other nodes. In summary, the data synchronization solution for a distributed system provided by the embodiments of this disclosure:

[0012] (1) By decoupling the consensus conditions during the leader election and data synchronization phases, the size of the majority required to reach consensus during the synchronization phase is smaller than that of the majority required in traditional non-decoupled strong-leader consensus protocols such as Raft. This allows the data synchronization phase to be completed in a shorter time, improving efficiency. For example, in a distributed system with a total of 9 nodes, the number of majorities required to reach consensus during the data synchronization phase in traditional consensus protocols is at least 5. However, in the disclosed embodiment, the number of majorities can be 4 (2 AZs, 2 nodes in each AZ).

[0013] (2) By introducing a strong leader consensus protocol that decouples the consensus conditions of the leader election and data synchronization phases into a multi-AZ disaster recovery scenario, we can take advantage of the fact that the probability of leader election events is much lower than that of data synchronization events. By using stricter consensus conditions in the leader election phase, we can reduce the impact of AZ disaster events on service availability while ensuring the service quality under normal conditions. Although the strict consensus conditions in the leader election phase increase the difficulty of reaching consensus in the leader election phase, the overall overhead of the leader election phase is not high because the probability of leader election events is lower. However, under the condition that the nodes that reach consensus in the leader election phase and the data synchronization phase need to have an intersection, the consensus conditions in the data synchronization phase can be relaxed - making the majority number in the data synchronization phase smaller. This means that the overall distributed system can tolerate more node failures while still ensuring availability (as long as the majority number of nodes used in the data synchronization phase is valid). For example, in the aforementioned embodiment, the availability of the service can be guaranteed as long as the nodes in the data synchronization selection box are valid. Furthermore, strict consensus conditions during the leader election phase make it more difficult for the leader node to be located in AZ3 for disaster recovery. Instead, it is more likely to be located in AZ1 or AZ2, which is located in the same region as the client. This allows client requests to be responded to with shorter delays, thereby improving service quality.

[0014] (3) By introducing the weight information of different AZs during the master election phase, the views of the nodes initiating the master election in each AZ are different from each other, thereby achieving the effect of flexibly controlling the process and results of the master election phase and improving the efficiency of the subsequent data synchronization phase. Specifically, by setting the node weights and the upper limit of the node master election rounds in different AZs, the difficulty of nodes in different AZs becoming master nodes can be controlled, making it easier for master nodes to be generated in AZs that are close to the client and have good quality, and more difficult to be generated in AZs that are disaster-tolerant and far away. The above embodiments are explained using the case of three AZs, each containing three nodes. The following describes an implementation plan for the master election and data synchronization phase in a more generalized distributed system deployment scenario. In a distributed system that provides a certain service, there are R AZs located in different regions, each containing N nodes, R is greater than or equal to 3, and N is greater than or equal to 3. In actual applications, two or more AZs can be deployed in one region. Here, the R AZs are represented as AZ_1, AZ_2, ..., AZ_R respectively. ONodes in the first AZ are configured with a first master election round upper limit and a first node weight, L. The first master election round upper limit prevents nodes in the first AZ from being elected as master nodes before the first master election round upper limit is reached. Nodes in the second AZ are configured with a second master election round upper limit and a second node weight. The second node weight is determined based on the second master election round upper limit, the first node weight, L, the total number of nodes N in the availability zone, and the number of majority nodes in the availability zone. The second node weight is less than the first node weight. The second AZ is any AZ among the R AZs other than the first AZ. The second master election round upper limit is greater than or equal to the first master election round upper limit. The second master election round upper limit prevents nodes in the second AZ from being elected as master nodes before the second master election round upper limit is reached. Based on the above configuration, when any target node in R AZs meets the master election trigger condition, the node determines the voting weights of multiple votes obtained based on the first node weight L, the second node weight, and the actual master election round of each node. Based on the voting weights of the multiple votes, it determines the voting weight statistics within each AZ. Based on the voting weight statistics within each AZ and the preset first consensus condition, it determines whether the target node can become the master node. After becoming the master node, the target node sends the write data received from the client to other nodes. If it is determined that the node among the other nodes that responds with a confirmation message corresponding to the write data meets the second consensus condition, the write data synchronization is determined to be successful. The set of nodes meeting the first consensus condition and the set of nodes meeting the second consensus condition intersect. Optionally, the first consensus condition is that in at least a first number of AZs, the voting weight statistics in each AZ are at least a second number, where the first number is the result of rounding down R / 2+1, and the second number is the result of rounding down (N / 2+1)*L. The second consensus condition is: In at least the first number of AZs, at least a third number of nodes in each AZ feedback the confirmation message, where the third number is N / 2+1 rounded down. As described above, voting weight = [(node ​​weight of D node * actual leader election round of D node)](node ​​weight of SRC node * actual leader election round of D node). Alternatively, voting weight = max(node ​​weight of DST node * actual leader election round of D node, node weight of SRC node * actual leader election round of D node).In the above embodiment, for a service deployed in R AZs, each with N nodes, and a maximum voting weight of L per vote, the first consensus condition that must be met during the leader election phase is: in at least [R / 2 + 1J AZs, the voting weight generated in each AZ must be at least [(N / 2 + l) * Lj]. The second consensus condition that must be met during the data synchronization phase is: in at least [R / 2 + 1J AZs, the number of nodes reaching consensus in each AZ must be at least L(N / 2 + l)J, where L represents the round-down operator. The maximum voting weight that can be achieved by a single vote is L, which limits the node weights within an AZ and the upper limit of the leader election round. For example, if AZ_1 is determined to be superior based on the environmental information of each AZ, the first AZ is determined to be AZ_1, the upper limit of the first leader election round is set to 1: Tmax(AZ_1)=1, and the first node weight is L: W(AZ_1)=L. In this case, AZ_1 The maximum voting weight that a node can achieve is L. In practical applications, the number of first AZs can be one or more. For example, several AZs within the same region can serve as first AZs, with the upper limit for the first election round set to 1 and the first node weights set to Lo Tmax (AZ_i) representing the upper limit for the election round of AZ_i, and W (AZ_i) representing the node weight of the node in AZ_i. AZs other than the first AZ are collectively referred to as second AZs. Therefore, the second AZs can be any of the remaining R AZs excluding the first AZ. The second node weight of the nodes in each second AZ is determined based on the upper limit for the second election round, the first node weight L, the total number of nodes N in an AZ, and the number of majority nodes in an AZ. The number of majority nodes in an AZ is L(N / 2 + 1)J. oFor example, if you want nodes in AZ_1 to become the master node with the minimum number of votes in the first round of master election initiated by it, you can set the node weight of the nodes in AZ_1 to W (AZ_1) = L, and the upper limit of the master election round Tmax (AZ_1) to 1. If you want nodes in AZ_2 to become the master node in the first round of master election initiated by it, you can set the node weight of the nodes in AZ_2 to W (AZ_2) = [(N / 2 + 1)J*L / N, and the upper limit of the master election round Tmax (AZ_2) to 1. If you want nodes in AZ_3 to become the master node in the second round of master election initiated by it, you can set the node weight of the nodes in AZ_3 to W (AZ_3) = [(N / 2 + 1)J*L / 2*N, and the upper limit of the master election round Tmax (AZ_3) to 2. Based on the above example: If you want a node in AZ_j to have a chance of becoming the master node in the pth round of master election initiated by it, you can set the node weight of the node in AZ_j to W (AZ_j) = [(N / 2 + 1)J*L / p*N], and set the upper limit of master election rounds Tmax (AZ_j) to p, where j is not equal to 1, meaning that AZ_j is not included in the first AZ. In the above example, the "minimum number of votes" for AZ_1 means that [(N / 2 + 1)J votes in each of [R / 2 + 1J AZs] can satisfy the first consensus condition. The "possible" for AZ_j means that the first consensus condition must be satisfied by obtaining votes from all N nodes in each of at least [R / 2 + 1J AZs]. This controls the possible rounds in which nodes in each AZ can be elected as the master, increasing the probability of nodes in AZs with a smaller upper limit of master election rounds being elected as the master. The execution process of the node master election phase and the data synchronization phase can refer to the relevant description in the aforementioned embodiments and will not be repeated here. The data synchronization device of the distributed system of one or more embodiments of the present disclosure will be described in detail below. Those skilled in the art will understand that these devices can all be configured using commercially available hardware components through the steps taught in this solution. Figure 7 is a structural schematic diagram of a data synchronization device of a distributed system provided by an embodiment of the present disclosure, which is applied to a target node in a target availability zone, and the target availability zone is any one of at least three availability zones located in different regions, and the nodes in different availability zones are configured to have different node weights. As shown in Figure 7, the device includes: a voting acquisition module 11, a role determination module 12, and a data synchronization module 13.The vote acquisition module 11 is configured to determine the voting weights of multiple votes obtained from multiple nodes when a master election trigger condition is met, combining the node weights corresponding to the at least three availability zones and the actual master election round of each node. The role determination module 12 is configured to determine whether the master node can become the master node based on the voting weights of the multiple votes and a preset first consensus condition. The data synchronization module 13 is configured to, if the target node becomes the master node, send write data received from the client to other nodes; and, if it is determined that a node among the other nodes that has fed back a confirmation message corresponding to the write data meets a second consensus condition, determine that the write data synchronization is successful; wherein the set of nodes meeting the first consensus condition and the set of nodes meeting the second consensus condition intersect. Optionally, the role determination module 12 is specifically configured to: when a master election trigger condition is met, increase the actual master election round and the first term identifier of the target node; if the actual master election round of the target node does not reach the master election round upper limit corresponding to the target node, obtain the first data identifier of the target node; send a voting request including the first term identifier and the first data identifier to other nodes, so that when the other nodes determine that their second term identifiers are not greater than the first term identifier and their second data identifiers are earlier than the first data identifier, they send a voting notification message to the target node, where the voting notification message includes the increased actual master election round and node weight of the other nodes; wherein the first data identifier and the second data identifier correspond to the latest data of the target node and the other nodes, respectively; and determine the voting weight of the corresponding vote based on the voting notification message, wherein the master election round upper limit of the nodes in the first availability zone and the second availability zone is configured as a first value, and the master election round upper limit of the nodes in the third availability zone is a second value, the first value is less than the second value, and the master election round upper limit is used to limit the number of consecutive votes initiated by the corresponding node. Optionally, the role determination module 12 is further configured to: if becoming the master node, reset the actual master election round of the target node, and send heartbeat packets to other nodes to enable the other nodes to reset their actual master election rounds.Optionally, the at least three availability zones include: a first availability zone and a second availability zone located in a first region, and a third availability zone located in a second region; nodes in the first availability zone are configured to have a first node weight, nodes in the second availability zone are configured to have a second node weight, and nodes in the third availability zone are configured to have a third node weight, wherein the third node weight is less than the first node weight and the second node weight, and the first node weight is greater than or equal to the second node weight; each availability zone includes three nodes, the first node weight is the same as the second node weight, and the first consensus condition is: the voting weight statistics in at least two availability zones are at least twice the first node weight; and the second consensus condition is: at least two nodes in each of the at least two availability zones feedback the confirmation message. The apparatus shown in FIG. 7 can execute the steps performed by the target node in the aforementioned embodiment. The detailed execution process and technical effects are described in the aforementioned embodiment and are not further elaborated here. In one possible design, the structure of the data synchronization apparatus of the distributed system shown in FIG. 7 can be implemented as an electronic device. As shown in FIG. 8 , the electronic device may include: a processor 21, a memory 22, and a communication interface 23. Memory 22 stores executable code. When executed by processor 21, processor 21 can implement at least the distributed system data synchronization method provided in the aforementioned embodiments. Furthermore, embodiments of the present disclosure provide a non-transitory machine-readable storage medium. The non-transitory machine-readable storage medium stores executable code. When executed by a processor of an electronic device, the processor can implement at least the distributed system data synchronization method provided in the aforementioned embodiments. The apparatus embodiments described above are merely illustrative. The network elements described as separate components may or may not be physically separate. Some or all of these modules can be selected based on practical needs to achieve the objectives of the present embodiments. Persons of ordinary skill in the art can understand and implement these embodiments without inventive effort. Through the above description of the embodiments, persons of ordinary skill in the art can clearly understand that each embodiment can be implemented using a necessary general-purpose hardware platform, or alternatively, through a combination of hardware and software. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a computer product. The present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present disclosure, and are not intended to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art will appreciate that modifications may be made to the technical solutions described in the aforementioned embodiments, or that some of the technical features therein may be replaced with equivalents. Such modifications or replacements do not deviate from the essence of the corresponding technical solutions within the spirit and scope of the various embodiments of the present disclosure. Industrial Applicability: During the leader election phase, when any target node in multiple availability zones meets the leader election trigger conditions, the node weights corresponding to the different availability zones and the current actual leader election round of each node are combined to determine the voting weights of multiple votes obtained from multiple nodes. Based on these voting weights and the preset first consensus condition, the target node is determined to be eligible for becoming the leader node. Therefore, the first consensus condition used during the leader election phase is a more stringent condition related to the voting weights determined based on node weights. After the target node becomes the master, it sends the write data received from the client to other nodes. If it determines that the node that responds with a confirmation message corresponding to the write data satisfies the second consensus condition, the write data synchronization is considered successful. The set of nodes that meet the first consensus condition intersects the set of nodes that meet the second consensus condition. Therefore, the second consensus condition used during the data synchronization phase is related to the number of nodes that successfully synchronized the write data. This decouples the consensus conditions of the data synchronization phase from those of the master election phase. This decoupling of the consensus conditions of these two phases facilitates flexible control over the master election process and results during the master election phase. This allows for stricter consensus conditions during the master election phase, thereby relaxing the consensus conditions during the data synchronization phase for more efficient completion.

Claims

Claims 1. A distributed system, comprising: At least three availability zones in different regions, and nodes in different availability zones are configured to have different node weights; A target node, which is used to determine the voting weights of multiple votes obtained from multiple nodes by combining the node weights corresponding to the at least three availability zones and the actual election rounds of each node when the election trigger condition is met, and determine whether the target node can become the master node according to the voting weights of the multiple votes and a preset first consensus condition. The target node is any node in the at least three availability zones; the target node is further used to send the write data received from the client to other nodes after becoming the master node, and if it is determined that the nodes that feedback the confirmation message corresponding to the write data among the other nodes meet the second consensus condition, determine that the write data synchronization is successful; wherein, there is an intersection between the set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition.

2. The system according to claim 1, wherein, The at least three availability zones include: a first availability zone and a second availability zone in a first region, and a third availability zone in a second region; the nodes in the first availability zone are configured to have a first node weight, the nodes in the second availability zone are configured to have a second node weight, the nodes in the third availability zone are configured to have a third node weight, the third node weight is less than the first node weight and the second node weight, and the first node weight is greater than or equal to the second node weight.

3. The system according to claim 2, wherein, The first node weight is the same as the second node weight.

4. The system according to claim 2, wherein The target node is specifically used to: determine the voting weight statistical value in each availability zone according to the voting weights of the multiple votes, and determine whether the target node can become the master node according to the voting weight statistical value in each availability zone and a preset first consensus condition.

5. The system according to claim 4, wherein Each availability zone includes three nodes, and the first consensus condition is: the voting weight statistical values in at least two availability zones are at least twice the first node weight.

6. The system according to claim 1, wherein Each availability zone includes three nodes, and the second consensus condition is: in each of at least two availability zones, there are at least two nodes that feedback the confirmation message.

7. The system according to claim 3, wherein The upper limit of the election rounds of the nodes in the first availability zone and the second availability zone is configured as a first value, the upper limit of the election rounds of the nodes in the third availability zone is a second value, the first value is less than the second value, the upper limit of the election rounds is used to limit the number of consecutive votes initiated by the corresponding nodes, and the third node weight is determined according to the first node weight and the second value.

8. The system according to claim 7, wherein In the process of determining the voting weight, the target node is configured to: when the master election trigger condition is met, increase the actual master election round and the first term identifier of the target node. If the actual master election round of the target node does not reach the upper limit of the master election round corresponding to the target node, obtain the first data identifier of the target node, and send a voting request including the first term identifier and the first data identifier to other nodes, so that when the other nodes determine that their second term identifier is not greater than the first term identifier and their second data identifier is earlier than the first data identifier, they send a voting notification message to the target node, and the voting notification message includes the increased actual master election round and node weight of the other node; and determine the voting weight of the corresponding vote according to the voting notification message; wherein, the first data identifier and the second data identifier respectively correspond to the last written data of the target node and the other node.

9. The system according to claim 8, wherein, The target node is further configured to reset the actual master election round of the target node when it is determined to be the master node, and send a heartbeat packet to other nodes to cause the other nodes to reset their actual master election rounds.

10. The system according to claim 1, wherein, The node weights corresponding to different availability zones are determined according to the environmental information of the corresponding availability zones, and the environmental information includes at least one of the following: the region to which it belongs, power information, and network information.

11. A data synchronization method for a distributed system, applied to a target node in a target availability zone, where the target availability zone is any one of at least three availability zones located in different regions, and nodes in different availability zones are configured with different node weights; the method includes: When the master election trigger condition is met, combine the node weights corresponding to the at least three availability zones and the actual master election rounds of each node to determine the voting weights of multiple votes obtained from multiple nodes; Determine whether the target node can become the master node according to the voting weights of the multiple votes and a preset first consensus condition; If it becomes the master node, send the write data received from the client to other nodes; if it is determined that the nodes that feedback the confirmation message corresponding to the write data among the other nodes meet the second consensus condition, determine that the write data synchronization is successful; wherein, there is an intersection between the set of nodes that meet the first consensus condition and the set of nodes that meet the second consensus condition.

12. The method according to claim 11, wherein The at least three availability zones include: a first availability zone and a second availability zone located in a first region, and a third availability zone located in a second region; nodes in the first availability zone are configured to have a first node weight, nodes in the second availability zone are configured to have a second node weight, nodes in the third availability zone are configured to have a third node weight, the third node weight is less than the first node weight and the second node weight, and the first node weight is greater than or equal to the second node weight; each availability zone includes three nodes, the first node weight is the same as the second node weight, the first consensus condition is that the vote weight statistic in at least two availability zones is at least twice the first node weight; the second consensus condition is that in each of at least two availability zones, there are at least two nodes that feedback the confirmation message.

13. An electronic device, comprising: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the data synchronization method of the distributed system according to claim 11 or 12.

14. A non-transitory machine-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the data synchronization method of the distributed system according to claim 11 or 12.

15. A distributed system, comprising: R availability zones located in different regions, each availability zone contains N nodes, R is greater than or equal to 3, and N is greater than or equal to 3; 19 Nodes in the first availability zone are configured to have a first master election round upper limit and a first node weight L, and the first master election round upper limit causes the nodes in the first availability zone not to be elected as the master node before reaching the first master election round upper limit; The nodes in the second available zone are configured to have a second master election round upper limit and a second node weight, and the second node weight is determined according to the second master election round upper limit, the first node weight L, the total number N of nodes in the available zone, and the number of majority nodes in the available zone. The second node weight is less than the first node weight L. The second available zone is any available zone among the R available zones except the first available zone. The second master election round upper limit is greater than or equal to the first master election round upper limit, and the second master election round upper limit is such that the nodes in the second available zone cannot be elected as the master node before reaching the second master election round upper limit. The target nodes in the R available zones are used to, when the master election trigger condition is satisfied, determine the voting weights of the multiple votes obtained in combination with the first node weight L, the second node weight, and the actual master election rounds of each node, determine the voting weight statistical value in each available zone according to the voting weights of the multiple votes, and determine whether the target node can become the master node according to the voting weight statistical value in each available zone and a preset first consensus condition. The target node is used to, after becoming the master node, send the write data received from the client to other nodes. If it is determined that the nodes among the other nodes that feedback the confirmation message corresponding to the write data satisfy the second consensus condition, then it is determined that the write data synchronization is successful. Among them, there is an intersection between the set of nodes that satisfy the first consensus condition and the set of nodes that satisfy the second consensus condition.

16. The system according to claim 15, wherein The first consensus condition is: in at least the first number of available zones, the voting weight statistical value in each available zone is at least the second number, where the first number is the ceiling result of R / 2 + 1, and the second number is the floor result of (N / 2 + 1) * L. The second consensus condition is: in at least the first number of available zones, at least the third number of nodes in each available zone feedback the confirmation message, and the third number is the floor result of N / 2 + 1.

Citation Information

Patent Citations

  • Distributed disaster recovery system, server node processing method, device and equipment

    CN114490158A

  • Leader election in a distributed system based on node weight and leadership priority based on network performance

    US20230222034A1