Optimization Method for Model Checking of Distributed Consensus Protocol Based on Semantic Information
Through state space reduction and exploration strategies based on semantic information, state space is compressed, and the state space explosion problem of the distributed consensus protocol model is solved, efficient model verification and error discovery is achieved, and the reliability of the system is improved.
Patent Information
- Application Number
- CN202310324016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-30
AI Technical Summary
The state space explosion problem of the distributed consensus protocol model results in high cost and low efficiency in model verification of computing resources, making it difficult to effectively verify the correctness of the model and find errors.
Through state space reduction and exploration strategies based on semantic information, state space is compressed, semantic information is used to focus on key logic, explore equivalence and independence, prune ineffective exploration, limit the number of key events, optimize the initial state, and improve model verification efficiency.
Effectively alleviate the explosion of model state space, reduce computing resource costs, improve the efficiency and accuracy of model verification, and improve the reliability of protocol systems.
Smart Images

Figure CN116389325B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an optimization method for model checking of a distributed consensus protocol based on semantic information, and belongs to the technical field of electronic digital data processing. Background Art
[0002] With the rapid growth in the number of Internet applications, various types of data have shown explosive growth, and the data volume of major companies has reached the EB / ZB level. To avoid single-point storage from becoming a bottleneck for system high availability and high scalability, distributed data systems usually adopt a distributed consensus protocol to store multiple copies of the same data on multiple physical nodes. The consensus protocol avoids system crashes caused by single-point failures under high-concurrency access, and improves the availability and fault tolerance of the system.
[0003] The model of the distributed consensus protocol can be used to perform related work such as verifying the correctness of the protocol or finding errors in the protocol, and has gradually become a mainstream protocol verification method. Through model checking, the correctness of the model can be verified, or the correctness of the model can be demonstrated, or invariant violations in the model can be discovered, which is beneficial to helping the protocol locate and modify errors. Due to the high complexity of the distributed consensus protocol, its model is often very complex, which means that the state space of the model is extremely large. Summary of the Invention
[0004] Object of the Invention: To solve the problem of state space explosion of the model, the present invention provides an optimization method for model checking of a distributed consensus protocol based on semantic information. This method is based on the Zab consensus protocol in ZooKeeper and the TLA+ specification language to form a model of the Zab protocol, relieve the state space explosion of the model, improve the efficiency of verifying the correctness of the model, and achieve the effect of reducing costs and increasing efficiency.
[0005] Technical Solution: An optimization method for model checking of a distributed consensus protocol based on semantic information reduces the computational resource cost during model checking by compressing the state space, and improves the efficiency of verifying the correctness of the model through an efficient exploration strategy of heuristic pruning, thereby enhancing the ability and efficiency to discover errors in the distributed protocol; mainly including:
[0006] 1) A state space reduction process based on semantic information;
[0007] 2) A state space exploration strategy based on semantic information.
[0008] State space reduction process based on semantic information:
[0009] This process aims to reduce the state space before exploring it without compromising the correctness of the verification model and the ability to find invariant violations. The main processing methods include: using semantic information to focus on the key logic in the state space, using semantic information to discover the equivalence between states in the state space, and using semantic information to discover the independence between events in the state space.
[0010] By using the principle of focusing on key logic with semantic information, we can manually mark the importance of each "module" to data synchronization, retain those modules that have a greater correlation with the interaction of "key data" between nodes and pose a greater threat to maintaining the consistency of "key data" between nodes, and abstract or discard modules that pose a smaller threat or are more peripheral to maintaining the consistency of key data between nodes. The event set in the model is divided into multiple "modules" according to the functional type. Each module consists of one or more events to complete a specific function of the protocol. This division method is mainly derived from the natural modularization in the design of the consensus protocol, which makes the module highly cohesive and the modules lowly coupled. For example, the Discovery module of the Zab protocol is responsible for determining the term of office in the current round of consensus process in the cluster, and the Sync module is responsible for completing the log recovery between the leader and each follower in the cluster. In addition, for distributed systems using consensus protocols, the "key data" between nodes is considered to be data that affects the security and activity between nodes, mainly including the term of office, log, and submission information of the node. According to the roles of the server nodes, the roles retained are the leader and follower in the log synchronization phase, and the looker in the master election phase; the roles discarded are the observers. According to the semantics of message transmission between nodes, key messages related to log recovery and synchronization are retained, and non-semantic messages such as heartbeat interaction between master and slave nodes are discarded. According to the importance of affecting the protocol logic, key data such as logs, terms, and submission information are retained, and data related to peripheral details such as node state machine data, network status, log file system, etc. are discarded. The logic of the network module is abstracted and simplified. The network module is modeled as a queue in the model. The sender puts the message at the end of the queue, and the receiver takes the message at the head of the queue for processing. The logic of the server cluster receiving client requests is abstracted and simplified. When the client initiates a write request, the Leader node directly receives the request for broadcast and processing, and the modeling of the client-initiated read request is discarded, which reduces the additional message forwarding caused by the Follower node receiving the request.
[0011] Using the principle of discovering the equivalence between states in the state space of the semantic information discovery model, when the states reached by different successor events are equivalent, consider selecting a representative state from the equivalent state class for exploration and discarding other non-representative states. That is, among the execution event classes that reach the same equivalent state class, select a representative event for execution and reject the execution of other non-representative events. For state S, it has multiple successor states {S1, S2, …, S n}, define the equivalence for the successor states {S1, S2, …, S n}, and thus partition the set of successor states into equivalent classes Select a representative state for each equivalent class for exploration and discard other non-representative states, thereby exponentially compressing the state space. In the model state space of Zab, the state equivalence we consider is the state equivalence of nodes. When event A involves two nodes (s, v) and both s and v are in the Looking role, then A(s, v) ≡ A(v, s), representing that the states reached by executing event A(s, v) and executing event A(v, s) are equivalent.
[0012] Using the principle of discovering the independence between events in the state space by semantic information, adjust the execution order between mutually independent events and further merge related events to compress the model state space without compromising the ability to verify the correctness of the model. For event A, after executing event A, the successor events include event B and event C. Assume that the set of variables affecting the behavior of executing event B is b1, the set of variables modified by executing event B is b2, the set of variables affecting the behavior of executing event C is c1, and the set of variables modified by executing event C is c2. Events B and C are mutually independent, which means that is, executing event C does not change the behavior of executing event B, and that is, executing event B does not change the behavior of executing event C, and that is, the execution order of executing event B and executing event C does not affect the behavior of the successor actions. When events B and C are mutually independent, the execution sequence ABC ≡ ACB, representing that the sequences ABC and ACB are equivalent. Based on this understanding, select the representative sequence ABC for execution and discard the sequence ACB. Further merge the events, let A ′= AB, simplify the above sequence to sequence A'C, thus exponentially compressing the state space. In the model state space of Zab, taking the node failure event NodeCrash(s) acting on node s as an example, its subsequent events include the node startup event NodeStart(s) and events acting on other nodes. Because the objects of action are different, the event NodeStart(s) is independent of other subsequent events. In the model, let the events NodeCrash(s) and NodeStart(s) be executed continuously in sequence. To further compress the state space, these two events are merged into the node restart event Restart(s).
[0013] State space exploration strategy based on semantic information:
[0014] What this process pursues is to discard the space with low exploration value after the model state space is generated on the premise that model errors may be lost, and retain the space with high exploration value that poses a greater threat to the correctness of the model. The main processing methods include: using semantic information to limit the scale of key events in the state space, using semantic information for the heuristic exploration strategy of the state space, and using semantic information to set the initial state of the more critical state space.
[0015] Based on our experience and understanding of the consensus protocol and the ZooKeeper system, repeating the events of a single module for a single execution sequence has very low quality. Limiting the number of executions of key events can effectively control the scale of the state space, and can also make the events between each module appear in combination in the execution sequence, improving the quality of the execution sequence, so as to minimize the loss of model errors and improve the reliability of model correctness verification. For the Zab model, we limit parameters such as the number of node failures, the number of network partitions, the maximum term, and the maximum length of the node log in the event, check and judge before each event is executed, and count and update the number of executions of key events after each event is executed.
[0016] Based on our experience and understanding of the consensus protocol and the ZooKeeper system, limit the continuous execution times of environmental error injection events in each round of log synchronization process to further promote the combined appearance of events between each module. Explore the meaning of each environmental error event in the current execution sequence, record each "meaningless" environmental error event, and limit the execution of these meaningless events in subsequent actions. Taking the environmental error event of node failure as an example, considering the meaning of subsequent actions, assume that in this round of log synchronization process, the node failure event NodeCrash(s) acting on node s has been executed, then the meaning discussion of subsequent environmental error events related to node failure and node startup is as follows:
[0017] 1) If the role of node s before failure is Looking, then it is meaningless to execute the node start event NodwStart(s) subsequently. Because node s does not participate in the current round of log synchronization, its state before failure is the same as that after startup. If the start event of node s is executed, it is equivalent to not executing the event NodeCrash(s).
[0018] 2) If the role of node s before failure is Leader or Follower, then it is meaningful to execute the node start event NodeStart(s) subsequently. Because the role of node s becomes Looking after restart, and the state of node s before and after is inconsistent.
[0019] 3) Regardless of the role of node s, it is meaningless to execute the node failure event NodeCrash(s) subsequently. Because when node s executes NodeCrash(s) for the second time, it means that a node start event NodeStart(s) of node s is executed between the two events NodeCrash(s). Then the part of node s in the state before executing the event NodeStart(s) is the same as the part of node s in the state after the second execution of the event NodeCrash(s).
[0020] 4) When a new round of log synchronization process is restarted due to the execution of the event ElectionAndDiscovery(l,F), where the parameter l is the Leader node and the parameter F is the set of Follower nodes. If s ∈ {l} ∪ F, then the node failure event NodeCrash(s) acting on node s is meaningful, otherwise it is meaningless; regardless of whether s ∈ {l} ∪ F holds, the node start event NodeStart(s) acting on node s is meaningful.
[0021] In the model, the event ElectionAndDiscovery(l,F) represents the process of electing a leader among nodes and reaching a consensus on the current term within the cluster in order to start a new round of log synchronization process. The processing flow of the environmental error event of network partition is similar to that of the environmental error event of node failure. By removing the meaningless events in the execution sequence and retaining the execution of meaningful events, the length of a single execution sequence is greatly shortened, meeting the requirements of compressing the state space and ensuring the quality of the exploration model.
[0022] Let the Zab model state be {s i :(r,p,e,h,c),i = 1…N} ∪ {net:[[s m →s n ,msg],…]}, which is divided into two parts: node state and network state, where s in represents a node, r represents a role, p represents the stage, e represents the current term, h represents the log, c represents the committed log, N represents the total number of nodes in the server cluster, net represents the network, and msg represents a single message. Generally speaking, the default initial state is that each node has not participated in consensus and no message has been transmitted in the network, that is, the initial state S0 = {s i :(Looking, Election, 0, [], []), i = 1…N} ∪ {net: []}. Based on the understanding of ZooKeeper and the investigation of historical errors, a large number of execution sequences that discovered bugs have similar prefix substrings of events. Extract this sequence of event substrings and set the initial state to the state reached after sequentially executing this event sequence. In the model, the initial state S0′ =
[0023] {s1: (Leader, Broadcast, 1, [[(1, 1), v1], [(1, 2), v2]], [[(1, 1), v1]])} ∪
[0024] {s i :(Follower, Broadcast, 1, [[(1, 1), v1]], [[(1, 1), v1]]), i = 2…N} ∪
[0025] {net: [{[s1 → s i , (1, 2), v2], i = 2…N}]}. This initial state is: node s1 is the Leader, other nodes are Followers, all nodes enter the Broadcast stage and can provide services externally, all nodes have persisted and committed the log entry with zxid (1, 1), the Leader node s1 has persisted the log entry with zxid (1, 2), and now broadcasts the proposal of the log entry with zxid (1, 1) to all Follower nodes. The successor state of this initial state has a higher possibility of violating the triggering invariant and has higher exploration value. At the same time, state S0′ can reach a new state equivalent to the original initial state S0 after executing several events, so the ability to find model errors is not lost.
[0026] Beneficial effects: Compared with the prior art, the method for optimizing the model checking of the distributed consensus protocol based on semantic information provided by the present invention effectively alleviates the problem of model state explosion in the distributed system. While compressing the model state space to reduce the computational resource cost of model checking, it extracts the remaining state space with higher exploration value, so as to improve the efficiency of verifying the correctness of the model or finding model errors and improve the reliability of the protocol or system corresponding to the model. Description of the Drawings
[0027] Figure 1A schematic diagram of a state space reduction process framework of an embodiment of the present invention;
[0028] Figure 2 Schematic diagram of the state space exploration strategy framework of an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The present invention is further explained below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0030] The distributed consensus protocol model verification optimization method based on semantic information mainly includes:
[0031] 1) State space reduction process based on semantic information;
[0032] 2) State space exploration strategy based on semantic information.
[0033] State space reduction process based on semantic information:
[0034] like Figure 1 As shown in (a), by using the principle of focusing on key logic with semantic information, each module can be manually marked according to its importance to data synchronization, and the modules set by the user that are highly correlated with the interaction of key data between nodes and pose a greater threat to maintaining the consistency of key data between nodes are retained for proportional modeling, while the modules set by the user that pose a smaller threat to maintaining the consistency of key data between nodes or are more peripheral are abstracted or discarded. According to the role of the server node, the retained roles are the leader and follower in the log synchronization phase, and the finder in the master election phase; the discarded role is the Observer. According to the semantics of message transmission between nodes, key messages related to log recovery and synchronization are retained, and non-semantic messages such as heartbeat interaction between master and slave nodes are discarded. According to the importance of affecting the system logic, key data such as logs, terms, and submission information are retained, and data related to peripheral details such as node state machine data, network status, log file system, etc. are discarded. The logic of the network module is abstracted and simplified. The network module is modeled as a queue in the model. The sender puts the message at the end of the queue, and the receiver takes the message at the head of the queue for processing. The abstraction simplifies the logic of the server cluster receiving client requests. When the client initiates a write request, the Leader node directly receives the request for broadcast and processing, while discarding the modeling of the client-initiated read request. This reduces the additional message forwarding caused by the Follower node receiving the request.
[0035] like Figure 1(As shown in (b)), according to the principle of exploiting semantic information to discover the equivalence between states, when the states reached by the execution of different successor events are equivalent, we consider selecting a representative state from the equivalent state class for exploration and discarding other non-representative states. That is, among the execution event classes that reach the same equivalent state class, we select a representative event for execution and reject the execution of other non-representative events. For state S, which has multiple successor states {S1, S2, …, S n}, we define the equivalence for these states, thereby partitioning the set of successor states into equivalent classes Select a representative state for each equivalent class for exploration and discard other non-representative states, thus exponentially compressing the state space. In the state space of the protocol model, the state equivalence we consider is the state equivalence of nodes. When event A involves two nodes (s, v) and both s and v are in the Looking role, then A(s, v) ≡ A(v, s), indicating that the states reached by executing event A(s, v) and executing event A(v, s) are equivalent.
[0036] For example Figure 1 (As shown in (c)), according to the principle of exploiting semantic information to discover the independence between events, we adjust the execution order between independent events and further merge related events to compress the model state space without compromising the ability to verify the correctness of the model. For event A, after executing event A, the successor events include event B and event C. Suppose the set of variables that affect the behavior of executing event B is b1, the set of variables modified by executing event B is b2, the set of variables that affect the behavior of executing event C is c1, and the set of variables modified by executing event C is c2. The independence between event B and event C means that is, executing event C does not change the behavior of executing event B, and that is, executing event B does not change the behavior of executing event C, and that is, the execution order of executing event B and executing event C does not affect the behavior of the successor actions. When event B and event C are independent, the execution sequence ABC ≡ ACB, indicating that the sequences ABC and ACB are equivalent. Based on this understanding, we select the representative sequence ABC for execution and discard the sequence ACB. Further merge the events, let A ′= AB, the above sequence is simplified to sequence A'C, thus exponentially compressing the state space. In the Zab model state space, taking the node failure event NodeCrash(s) acting on node s as an example, its subsequent events include the node start event NodeStart(s) and events acting on other nodes. Since the objects of action are different, it is easy to prove that the event NodeStart(s) is independent of other subsequent events. In the model, let the events NodeCrash(s) and NodeStart(s) be executed continuously in sequence. To further compress the state space, these two events are merged into the node restart event Restart(s).
[0037] State space exploration strategy based on semantic information:
[0038] What this process pursues is to discard the space with low exploration value after the model state space is generated while possibly losing model errors, and retain the space that poses a greater threat to the model's correctness and has high exploration value. The main processing methods include: restricting the scale of key events in the state space using semantic information, using semantic information for heuristic exploration strategies in the state space, and using semantic information to set more critical initial states of the state space.
[0039] For a single execution sequence, repeatedly executing events of a single module has very low quality. Limiting the number of executions of key events can effectively control the scale of the state space, and at the same time allow events between each module to appear in combination in the execution sequence, improving the quality of the execution sequence, so as to minimize the loss of model errors and improve the reliability of model correctness verification. As Figure 2 (a) shows, in the model, we limit parameters such as the number of node failures, the number of network partitions, the maximum term, and the maximum length of node logs in the event, check and judge before each event is executed, and count and update the number of executions of key events after each event is executed.
[0040] Limit the continuous execution times of environmental error injection events in each round of log synchronization process to further promote the combination of events between each module. Explore the meaning of each environmental error event in the current execution sequence, record each "meaningless" environmental error event, and limit the execution of these meaningless events in subsequent actions. As Figure 2 (b) shows, taking the environmental error event of node failure as an example, considering the meaning of subsequent actions, assume that in the current round of log synchronization process, the node failure event NodeCrash(s) acting on node s has been executed, then the discussion on the meaning of subsequent environmental error events regarding node failure and node start is as follows:
[0041] 1) If the role of node s was Looking before failure, then it is meaningless to execute the node start event NodeStart(s) for the subsequent execution node. Because node s does not participate in the current round of log synchronization, its state before failure is the same as its state after startup. If the start event of node s is executed, it is equivalent to not executing the event NodeCrash(s).
[0042] 2) If the role of node s was Leader or Follower before failure, then it is meaningful to execute the node start event NodeStart(s) for the subsequent execution node. Because the role of node s becomes Looking after restart, and the state of node s before and after is inconsistent.
[0043] 3) Regardless of the role of node s, it is meaningless to execute the node crash event NodeCrash(s) for the subsequent execution node. Because when node s executes NodeCrash(s) for the second time, it means that a node start event NodeStart(s) of node s is executed between the two events NodeCrash(s). Then the part of node s in the state before executing the event NodeStart(s) is the same as the part of node s in the state after executing the event NodeCrash(s) for the second time.
[0044] 4) When a new round of log synchronization process is restarted due to the execution of the event ElectionAndDiscovery(l,F), where the parameter l is the Leader node and the parameter F is the set of Follower nodes. If s ∈ {l} ∪ F, then the node crash event NodeCrash(s) acting on node s is meaningful, otherwise it is meaningless; regardless of whether s ∈ {l} ∪ F holds, the node start event NodeStart(s) acting on node s is meaningful.
[0045] The processing flow of the environmental error event in the network partition is similar to that of the environmental error event of node failure. By removing the meaningless events in the execution sequence, the length of a single execution sequence is greatly shortened, meeting the requirements of compressing the state space and ensuring the quality of the exploration model.
[0046] Generally speaking, the default initial state S0 is that each node has not participated in consensus yet, and no message has been transmitted in the network. Based on the understanding of ZooKeeper and the investigation of historical errors, most of the execution sequences that discovered bugs have a similar prefix substring of events. Extract this event substring sequence and set the initial state to the state reached after sequentially executing this event sequence, such as Figure 2As shown in (c), in the model, the initial state S0′ is: node s1 is the leader, other nodes are followers, all nodes enter the Broadcast phase and can provide services to the outside world, all nodes persist and submit log items with zxid (1,1), leader node s1 persists the log item with zxid (1,2), and now broadcasts the proposal of log items with zxid (1,1) to all follower nodes. The successor state of this initial state has a higher probability of triggering invariant violations and has a higher exploration value. At the same time, state S0′ can reach a new state equivalent to the original initial state S0 by executing several events, so the ability to find model errors is not lost.
[0047] The usage of this method is as follows:
[0048] 1) Deploy the model of the distributed consensus protocol. The typical deployment method is to configure all constants contained in the model and add invariants in the model to verify the correctness of the model. The constants include the number of key events, and the number threshold needs to be manually set by humans.
[0049] 2) Use the simulation model checking mode in the TLC model checker, which starts from the initial state of the state space and randomly executes a complete sequence each time until the execution sequence length limit is reached or there is no next reachable state. Through the simulation model checking mode, multiple execution sequences are obtained, and all generated execution sequences are saved locally for data analysis in subsequent processes.
[0050] 3) Summarize the execution sequences generated by all old models before using this method and new models after using this method, and classify them according to the different parameters of the number of key events. Compare the length of execution sequences in different classes, the number of events of some key modules, the proportion of violations of discovered model invariants, and other data to analyze the effectiveness of this method in optimizing the state space of the model and verifying correctness.
[0051] The technical solution of the present invention is described in detail below through a specific example. The Zab protocol model is selected as the running object of the model verification, and the execution sequence generated by each model verification is collected in the local running framework by reusing the TLC model checker, and the execution sequence is analyzed by the local script.
[0052] 1) Hardware environment:
[0053] Start the running system based on the TLC model checker on the local workstation to obtain the execution sequence, and start the script for analyzing the execution sequence on the local workstation to make statistics on the key data of the execution sequence.
[0054] 2) Running process:
[0055] Before applying this method to the application of the Zab protocol, start the running system based on the TLC model checker on the local workstation. Stop after obtaining 10,000 execution sequences each time. Among them, the maximum length of a single execution sequence is limited to 100. Run the local script to count data such as the length of the execution sequence, the number of occurrences of events in some key modules, and the number of violations of the discovered model invariants. Repeat this process 20 times to eliminate errors.
[0056] After applying this method to the application of the Zab protocol, start the running system based on the TLC model checker on the local workstation. Stop after obtaining 10,000 execution sequences each time. Among them, the maximum length of a single execution sequence is limited to 100. Run the local script to count data such as the length of the execution sequence, the number of occurrences of events in some key modules, and the number of violations of the discovered model invariants. Repeat this process 20 times to eliminate errors.
[0057] Compare the results of the two experiments. The experimental results are shown in Table 2.
[0058] The experimental parameters and default values are shown in Table 1.
[0059] 3) Running results:
[0060] Table 1 Experimental parameters and default values
[0061]
[0062] Table 2 Experimental results
[0063]
Claims
1. A method for optimizing the model checking of a distributed consensus protocol based on semantic information, characterized in that For a distributed consensus protocol model, to compress the state space of the model and improve the verification of the model's correctness, it includes: 1) The state space reduction process based on semantic information; this process reduces the state space before exploring the state space without compromising the ability to verify the correctness of the model and find invariant violations. The processing methods included are: using semantic information to focus on the key logic in the state space, using semantic information to discover the equivalence between states in the state space, and using semantic information to discover the independence between events in the state space; 2) The state space exploration strategy based on semantic information; after the model state space is generated, discard the spaces with low exploration value and retain the spaces that pose a greater threat to the correctness of the model and have high exploration value. The processing methods included are: using semantic information to limit the scale of key events in the state space, using semantic information for the heuristic exploration strategy of the state space, and using semantic information to set the initial state of the more critical state space; In the principle of using semantic information to focus on key logic, mark according to the importance of each module for data synchronization, and retain those modules that have a greater correlation with the key data interaction between nodes and pose a greater threat to maintaining the key data consistency between nodes for proportional modeling, while abstracting or discarding the modules that pose a smaller threat to maintaining the key data consistency between nodes or are more peripheral; In the principle of using semantic information to focus on key logic, according to the roles of server nodes, the retained roles are the Leader, Follower in the log synchronization stage, and the Looking in the leader election stage; the discarded role is the Observer; according to the semantics of the messages transmitted between nodes, retain the key messages related to log recovery and synchronization and discard the meaningless messages of the heartbeat interaction between the master and slave nodes; according to the importance of affecting the protocol logic, retain the key data, and the key data includes logs, terms, and commit information; discard the data related to peripheral details, and the data related to peripheral details includes the state machine data of nodes, network status, and log file system; abstract and simplify the logic of the network module, and in the model, the network module is modeled as a queue, the sender puts the message at the end of the queue, and the receiver takes the message at the head of the queue for processing; Abstract and simplify the logic of the server cluster receiving client requests. When the client initiates a write request, the Leader node directly receives the request for broadcasting and processing, and at the same time discard the modeling of the client initiating a read request.
2. The method for optimizing the model checking of the distributed consensus protocol based on semantic information according to claim 1, characterized in that In the state space reduction process described above, the principle of using semantic information to discover the equivalence between states in the state space; when the states reached by the execution of different successor events are equivalent, consider selecting a representative state from the equivalent state class for exploration and discard other non-representative states, that is, among the execution event classes that reach the same equivalent state class, select a representative event for execution and reject the execution of other non-representative events.
3. The method for optimizing the verification of the distributed consensus protocol model based on semantic information according to claim 2, wherein In the principle of exploiting the equivalence between states in the state space using semantic information, for a state S that has multiple successor states {S1, S2, …, S n}, the equivalence is defined for the successor states {S1, S2, …, S n}, thereby partitioning the set of successor states into equivalence classes Select a representative state for each equivalence class for exploration, discard other non-representative states, thereby exponentially compressing the state space; in the model state space of Zab, state equivalence is the state equivalence of nodes. When an event A involves two nodes (s, v), and s and v are both in the Looking role, then A(s, v) ≡ A(v, s), indicating that the states reached by executing the event A(s, v) and executing the event A(v, s) are equivalent.
4. The method for optimizing the model checking of the distributed consensus protocol based on semantic information according to claim 1, wherein In the state space reduction process described above, the principle of using semantic information to discover the independence between events in the state space, adjust the execution order between mutually independent events, and further merge related events to compress the model state space without compromising the ability to verify the correctness of the model.
5. The method for optimizing the verification of a distributed consensus protocol model based on semantic information according to claim 4, wherein In the principle of exploring the independence between events in the state space using semantic information, for event A, after executing event A, the subsequent events include event B and event C. Let The set of variables that affect the behavior of execution event B is b1, the set of variables modified by execution event B is b2, the set of variables that affect the behavior of execution event C is c1, and the set of variables modified by execution event C is c2; events B and C are independent of each other, which means that is, execution event C does not change the behavior of execution event B, and that is, execution event B does not change the behavior of execution event C, and that is, the order of execution events B and C does not affect the behavior of subsequent actions; when events B and C are independent of each other, the execution sequence ABC ≡ ACB, representing that the sequences ABC and ACB are equivalent; Select the representative sequence ABC for execution and discard the sequence ACB. Further merge the events, let A′ = AB, and simplify the above sequence to the sequence A′C, thus exponentially compressing the state space. In the model state space of Zab, for the node failure event NodeCrash(s) acting on node s, its subsequent events include the node startup event NodeStart(s) and events acting on other nodes. Because the objects of action are different, the event NodeStart(s) is independent of other subsequent events. In the model, let the events NodeCrash(s) and NodeStart(s) be executed continuously in sequence. To further compress the state space, these two events are merged into the node restart event Restart(s).
6. The method for optimizing the verification of a distributed consensus protocol model based on semantic information according to claim 1, characterized in that In the state space exploration strategy based on semantic information, use semantic information to adopt a heuristic exploration strategy for the state space, limit the continuous execution times of the environmental error injection event in each round of log synchronization process, and promote the combination of events among modules; explore the meaning of each environmental error event in the current execution sequence, record each "meaningless" environmental error event, and limit the execution of these meaningless events in subsequent actions; for the environmental error event of node failure, consider the meaning of subsequent actions. Suppose in the current round of log synchronization process, the node failure event NodeCrash(s) acting on node s has been executed. Then the discussion on the meaning of subsequent environmental error events regarding node failure and node startup is as follows: 1) If the role of node s before failure is Looking, then the subsequent execution of the node startup event NodeStart(s) is meaningless; because node s does not participate in the current round of log synchronization, its state before failure is the same as its state after startup. If the startup event of node s is executed, it is equivalent to not executing the event NodeCrash(s); 2) If the role of node s before failure is Leader or Follower, then the subsequent execution of the node startup event NodeStart(s) is meaningful; because the role of node s becomes Looking after restart, and the state of node s before and after is inconsistent; 3) Regardless of the role of node s, the subsequent execution of the node failure event NodeCrash(s) is meaningless; because when node s executes NodeCrash(s) for the second time, it means that there is an execution of the node startup event NodeStart(s) of node s between the two events NodeCrash(s). Then the part of node s in the state before executing the event NodeStart(s) is the same as the part of node s in the state after the second execution of the event NodeCrash(s); 4) When a new round of log synchronization process is restarted due to the execution of the event ElectionAndDiscovery(l,F), where the parameter l is the Leader node and the parameter F is the set of Follower nodes. If s ∈ {l} ∪ F, the node failure event NodeCrash(s) acting on the node s is meaningful; otherwise, it is meaningless. Regardless of whether s ∈ {l} ∪ F holds, the node startup event NodeStart(s) acting on the node s is meaningful; The processing flow of the environmental error event of network partitioning is similar to that of the environmental error event of node failure. Eliminate the meaningless events in the execution sequence and retain the execution of meaningful events.
7. The method for optimizing the verification of the distributed consensus protocol model based on semantic information according to claim 1, wherein In the state space exploration strategy based on semantic information, the initial state of a more critical state space is set using semantic information; let the state of the simplified protocol model be {s i :(r, p, e, h, c), i = 1…N} ∪ {net: [[s m →s n , msg],…]}, which is divided into two parts: node state and network state, where s i represents a node, r represents a role, p represents the stage, e represents the current term, h represents the log, c represents the committed log, N represents the total number of nodes in the server cluster, net represents the network, and msg represents a single message; The default initial state is that each node has not participated in consensus and no messages have been transmitted in the network, that is, the initial state S0 = {s i :(Looking, Election, 0, [], []), i = 1…N} ∪ {net: []}; Based on the understanding of ZooKeeper and the investigation of historical errors, most of the execution sequences that discovered bugs have similar prefix substrings of events. Extract the event substring sequence and set the initial state to the state reached after sequentially executing this event sequence. In the model, the initial state S0' = {s1: (Leader, Broadcast, 1, [[(1,1), v1], [(1,2), v2]], [[(1,1), v1]])} ∪ {s i :(Follower, Broadcast, 1, [[(1,1), v1]], [[(1,1), v1]]), i = 2…N} ∪ {net: [{[s1→s i ,(1,2), v2], i = 2…N}]}; This initial state is: node s1 is the Leader, other nodes are Followers, all nodes enter the Broadcast phase and can provide services externally, all nodes have persisted and committed the log entry with zxid (1,1), the Leader node s1 has persisted the log entry with zxid (1,2), and now broadcasts the proposal of the log entry with zxid (1,1) to all Follower nodes; The state S0' reaches a new state equivalent to the original initial state S0 after executing several events.