Distributed System Testing Method and Device Based on TLA+ Formal Specification Model Checking

By mapping the elements in the TLA+ formal specification to the code implementation of the distributed system and generating test cases using the TLC model inspector, the problem of difficulty in detecting the implementation defects of the distributed system in the existing technology is solved, and the effect of efficient and comprehensive testing is achieved.

CN115309654BActive Publication Date: 2025-06-10INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211041081.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-06-10
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect defects in distributed system implementation, especially due to state space explosion problems and efficiency limitations of relying on user-defined test assertions.

Method used

By mapping the features in the TLA+ formal specification into the code implementation of the distributed system, and verifying the abstract state space graph using the TLC model inspector, generating test cases for controlled testing, comparing run data with controlled test data to find potential flaws.

Benefits of technology

Systematized and complete testing of the huge state space of distributed systems has been achieved, greatly improving the testing efficiency and helping developers to promptly discover defects in the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309654B_ABST
    Figure CN115309654B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed system testing method and device based on TLA+ formal specification model checking. The method includes: mapping the elements of the TLA+ formal specification of the target distributed system to the implementation of the target distributed system, and collecting the operation data of the target distributed system; using the TLC model checker to verify the TLA+ formal specification to obtain an abstract state space graph with verified correctness; performing path traversal on the abstract state space graph, and regarding each obtained path as a test case; performing controlled testing on the target distributed system based on the test case to obtain controlled test data; comparing the operation data with the controlled test data to generate a test report. The present invention can discover potential defects in the distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software, and particularly relates to a distributed system testing method and device based on TLA+ formal specification model checking. Background Art

[0002] With the development and popularization of cloud computing technology, distributed systems have begun to play an important role in various industries. Different types of distributed systems are widely used in various fields such as finance, transportation, and government affairs. Compared with traditional stand-alone software, distributed systems run on multiple nodes and have better reliability. However, defects in the design and implementation of distributed systems can affect their correctness and thus threaten the reliability of the system. Distributed systems have a huge state space and must face various uncertain events, such as user requests, unstable networks, and various types of faults. Therefore, it is very difficult to find deeply hidden defects within the system by exploring all possible scenarios. However, these defects can cause serious consequences such as service outages and huge economic losses. Therefore, designing an efficient testing method for distributed systems that can conduct comprehensive and systematic testing is of great significance for ensuring the reliability of distributed systems.

[0003] Currently, the technologies used to discover defects in distributed systems include formal methods and practical model checking. Formal methods usually formalize the key protocols of distributed systems and then use formal verification tools to check and verify the correctness of the models. For example, Zave modeled and verified the distributed protocol Chord using the Alloy formal language as early as 2013 and found several counterexamples that violated the predetermined properties in the system design. Program developers at Amazon used the formal language TLA+ to verify their systems and successfully discovered several protocol design defects. However, there is a gap between the abstract formal specification and the specific implementation of distributed systems. Even if formal verification proves the correctness of the distributed system in design, it cannot ensure that there are no defects in the system implementation.

[0004] Compared with formal methods, implementation-level model checking technology can effectively identify defects in the implementation of distributed systems. Implementation-level model checking intercepts messages during the runtime of a distributed system, controls the order of message sending and receiving, injects node downtime and restart failures, enabling the system to enter different states, and then detects whether the system's behavior is normal in these states. However, despite the significant amount of work researchers have done in state space reduction, implementation-level model checking is still troubled by the state space explosion problem. In addition, implementation-level model checking relies on user-defined test assertions to determine whether potential defects are triggered in test scenarios, which greatly limits its efficiency in defect detection.

[0005] In summary, distributed systems are widely used, and it is very important to ensure their reliability. Hidden defects in distributed systems can affect the correctness of the system and thus threaten system reliability. However, existing technologies cannot effectively detect defects in the implementation of distributed systems. Therefore, designing and developing a test method for implementation-level distributed systems can efficiently and comprehensively test distributed systems, which is of great value and significance for improving the reliability of distributed systems. Summary of the Invention

[0006] The objective of the present invention is to bridge the gap between the abstract formal specification and the specific code implementation of distributed systems, and propose a test method and device for distributed systems based on model checking of TLA+ formal specifications. The method maps key state and behavior elements in the formal specification and code implementation of the distributed system, and controls various uncertainties during system runtime, so that the system can be tested according to the state transition process verified by model checking during runtime, and then discover code implementations that are inconsistent with the formal specification, that is, potential distributed system defects.

[0007] The technical solution of the present invention includes:

[0008] A test method for distributed systems based on model checking of TLA+ formal specifications, the steps of which include:

[0009] For the TLA+ formal specification of the target distributed system, map the elements of the TLA+ formal specification to the implementation of the target distributed system, and collect the runtime data of the target distributed system; wherein, the runtime data includes: the values of node state variables during runtime and the occurrence of behaviors during runtime;

[0010] Use the TLC model checker to verify the TLA+ formal specification and obtain an abstract state space graph with verified correctness;

[0011] Perform path traversal on the abstract state space graph, and regard each obtained path as a test case;

[0012] Perform a controlled test on the target distributed system based on the test cases to obtain controlled test data; wherein, the controlled test data includes: the variable values of the state nodes in the test path and the occurrence of behaviors in the test path.

[0013] Compare the running data with the controlled test data to generate a test report.

[0014] Further, the elements include: variable elements used to describe the states within the distributed system and behavior elements that cause state updates within the distributed system. The variable elements include: state-related variables and message-related variables. The behavior elements include: internal node behaviors, message-related behaviors, user requests, and faults. The faults include: node downtime, node restart, message loss, and message duplication. Among them, the state-related variables are used to define the critical state values in the distributed system, the message-related variables are used to simulate the information exchange in the message communication of the target distributed system in the form of shared variables, the internal node behaviors are used to describe the behaviors that update the node state values, the message-related behaviors are used to describe the behaviors that update the node state values through message passing, the user requests are used to describe the external behaviors of the system initiated by users, and the faults are used to describe external uncertain behaviors.

[0015] Further, in the implementation of mapping the state-related variables to the target distributed system, it includes:

[0016] Annotate the state-related variables with @Variable($VNAME$); where $VNAME$ represents the naming of the state-related variable in the TLA+ formal specification.

[0017] For each of the state-related variables, instrument the target distributed system based on the annotation of the state-related variable and add a static shadow field to the mapped variable; where the static shadow field is used to assign values based on the initialization or update of the node state variable values.

[0018] Further, in the implementation of mapping the internal node behaviors to the target distributed system, it includes:

[0019] Annotate the internal node behaviors with @Behavior($BNAME$); where $BNAME$ represents the naming of the internal node behavior in the TLA+ formal specification.

[0020] For each of the internal behaviors of the nodes, instrument the target distributed system based on the annotation of the internal behaviors of the nodes, and insert a line of code at the entry and return of each mapped behavior method; wherein, the code at the entry is used to collect the real-time values of the behavior parameters to block the method from continuing to run and wait for the scheduling instruction to return, and the code at the return is used to collect the real-time values of all state-related variables.

[0021] Further, the implementation of mapping the message-related behaviors to the target distributed system includes:

[0022] Annotate the message-related behaviors;

[0023] For each of the internal behaviors of the nodes, instrument the target distributed system based on the annotation of the message-related behaviors to collect the message values transmitted in the target distributed system.

[0024] Further, the implementation of mapping the user requests to the target distributed system includes:

[0025] Use the user script of the target distributed system to establish the correspondence between the user request behaviors in the TLA+ formal specification and the user script;

[0026] Based on the correspondence, match the user requests and the user script.

[0027] Further, the implementation of mapping the faults to the target distributed system includes:

[0028] In the case where the fault is the node downtime or the node restart, kill the process corresponding to the node and restart it through the saved node configuration;

[0029] In the case where the fault is the message loss fault, instrument the message and network-related code, and ignore the internal logic of the message processing function and directly return;

[0030] In the case where the fault is the message duplication fault, instrument the message and network-related code, and perform multiple repeated processing of the same message in the message processing function.

[0031] Further, the path traversal of the abstract state space graph, regarding each obtained path as a test case, includes:

[0032] Start from the initial state node of the abstract state space;

[0033] Judge whether the current state node is the test scenario termination state:

[0034] If not, then sequentially visit each unvisited successor behavior edge of the current state node, add the successor behavior edge and the corresponding successor node to the traversal result, mark the successor edge as visited, use the corresponding successor node as the current state node, and return to determine whether the current state node is the termination state of the test scenario;

[0035] If so, then generate a test case based on the traversal result;

[0036] Further, the comparing the running data with the controlled test data to generate a test report for the target distributed system includes:

[0037] Compare the values of the node state variables during runtime with the variable values of the state nodes in the test path to obtain a test report based on state determination;

[0038] Compare the occurrence of behaviors during runtime with the occurrence of behaviors in the test path to obtain a test report based on behavior determination;

[0039] Integrate the test report based on state determination and the test report based on behavior determination to obtain the test report for the target distributed system; wherein, the test report for the target distributed system includes: whether the implementation of the target distributed system is consistent or inconsistent with the TLA+ formal specification.

[0040] A storage medium stores a computer program, wherein the computer program is configured to execute any of the above methods when running.

[0041] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute any of the above methods.

[0042] Compared with the prior art, the present invention has the following technical advantages:

[0043] 1. The method of the present invention uses the abstract state space generated by model checking of the TLA+ formal specification to guide the test process of the distributed system, and uses the specific state values and behavior transitions in the abstract state space as oracles for testing the distributed system. This enables systematic and complete testing of the huge state space of the distributed system, greatly improving the test efficiency and helping distributed system developers to timely discover defects existing in the system.

[0044] 2. The present invention uses a depth-first algorithm based on edge coverage to traverse the abstract state space of the distributed system and generate test cases, which greatly reduces the number of test cases that need to be tested and improves the test efficiency.

[0045] 3. The present invention utilizes code instrumentation on the target distributed system and uses various scripts to control the uncertainties in the distributed system, such as message order, user requests, various faults, etc. This enables the distributed system to execute completely according to the behavior sequence specified in the test case during the test process, ensuring the accuracy of the test. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Example diagram of the overall process flow of the method of the present invention.

[0047] Figure 2 Example diagram of elements in the TLA+ formal specification.

[0048] Figure 3 Example diagram of the abstract state space obtained by model checking the TLA+ formal specification.

[0049] Figure 4 Example diagram of the controlled test process. DETAILED DESCRIPTION OF THE INVENTION

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only specific embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The technical solution of the present invention includes an automated test method for a distributed system based on model checking of a TLA+ formal specification. In this automated test method, the abstract state space formed by the TLA+ formal specification is used to guide the test of the implementation of the distributed system. By comparing the states and behaviors during the operation of the system during the test with the states and behaviors in the abstract state space, the inconsistencies between the system implementation and the formal specification are found, that is, the potential defects in the distributed system. The specific steps of this automated test method include:

[0052] (1) Map various elements in the TLA+ formal specification, including state-related variables, message-related variables, internal node behaviors, message-related behaviors, user requests, and faults, to the code implementation of the distributed system, so that the state and behavior information of the distributed system can be obtained and controlled in real time during the test.

[0053] (2) Perform TLC model checking on the TLA+ formal specification to obtain an abstract state space diagram.

[0054] (3) Based on the abstract state space diagram, perform graph traversal to generate a path set as the test case set.

[0055] (4) Test the implementation of the distributed system based on test cases and control the uncertainties during the testing process. Specifically, it is necessary to control the message order, user requests, and occurrence of faults in the distributed system.

[0056] (5) Compare the behavioral states of the distributed system during the testing process with the behavioral states in the abstract state space. If there are inconsistencies, report them as potential defects in the distributed system.

[0057] The following Figure 1 illustrates the method of the present invention in detail with reference to the specific implementation process example diagram in

[0058] First, map various elements in the TLA+ formal specification. As Figure 2 shown, the TLA+ formal specification mainly consists of variables and behaviors. All variables together define the states in the distributed system, and different variable values represent different states in the distributed system. Different behaviors define the change relationships between states. Specifically, the variables and behaviors in the distributed system include: state-related variables, message-related variables, internal node behaviors, message-related behaviors, user requests, and various types of faults. The following separately describes the mapping methods for different variables and behaviors.

[0059] State-related variables define the key state values in the distributed system. Message-related variables simulate the information exchange in the distributed system message communication in the form of shared variables. Internal node behaviors occur on a single node and only update the state of a single node. Message-related behaviors act between different nodes and update the state values through message passing. User requests are external behaviors initiated by users. Faults refer to external uncertain behaviors such as node downtime and restart, message loss, and message resending.

[0060] For the state-related variable elements in the TLA+ formal specification, map them to the corresponding class member variables in the distributed system source code. The user needs to annotate these class member variables with @Variable($VNAME$), where $VNAME$ is the name of the variable in the TLA+ formal specification. During the operation of the distributed system, the testing method will perform automated instrumentation according to the annotation and add a static shadow field for each mapped variable. When the variable value is initialized or updated, the shadow field will also be assigned the same value. During subsequent controlled testing, the real-time values of the shadow fields will be collected for state judgment without affecting the original variables.

[0061] For the message-related variables in the TLA+ formal specification, they essentially simulate the message sending and receiving process in a distributed system through shared variables, and it is impossible to directly map the message-related variables to any corresponding source code. In this regard, an independent third-party module is used in the testing method to obtain information and control the message communication process in the distributed system, so as to map the message-related variables.

[0062] For the internal behavior of a node, map it to the corresponding method in the distributed system source code. The user needs to annotate this method with @Behavior($BNAME$), where $BNAME$ is the name of this behavior in the TLA+ formal specification. In addition, the user needs to manually write a line of code at the beginning of this method using the interface provided by the tool to obtain the runtime parameter values of the method, so as to match the specific behavior in the TLA+ specification with the runtime behavior implemented in the distributed system. During the operation of the distributed system, the testing method will perform automatic instrumentation according to the annotation, and insert a line of code at the entry and return of each mapped behavior method. At the entry of the method, automatically collect the real-time values of the behavior parameters and send them to the central module of the testing method, block the method from continuing to run and wait for the scheduling instruction returned by the central module. At the return of the method, automatically collect the real-time values of all state-related variables and send them to the central module of the testing method for state consistency judgment.

[0063] For message-related behaviors, in addition to the user's manual annotation and the tool's automatic instrumentation for the internal behavior of the node, the user also needs to use the interface provided by the testing method tool to collect the information values passed by the message inside the corresponding method, and arrange them in the value order defined in the TLA+ specification, so as to control specific message-related behaviors during testing.

[0064] For user requests, use the user scripts provided by the distributed system itself to establish the correspondence between the user request behaviors in the TLA+ specification and the corresponding script running commands, and match and manage the available request parameters in the TLA+ specification and the corresponding script running parameters.

[0065] For different types of faults, use different methods to simulate them to map different fault behaviors defined in the TLA+ specification. For node downtime and node restart faults, the testing method uses the corresponding script to directly kill the process corresponding to the node and restart it through the saved node configuration. For network-related faults, the testing method instruments the message and network-related code to simulate and inject different types of network faults. Specifically, for message loss faults, ignore the internal logic of the message processing function and directly return. For message duplication faults, process the same message multiple times in the message processing function.

[0066] After mapping the elements of the TLA+ formal specification, the TLA+ specification is verified using the TLC model checker, and a state space diagram is output, as Figure 3 shown. This diagram has an initial state as the entry point. Each node in it represents a unique distributed system state, and the edges between nodes represent the specific behaviors that cause state updates.

[0067] Then, the testing method traverses the state space diagram to obtain a set of paths and generates a set of test cases. Each path is a test case. In this path, each node represents a program point where the values of all state variables are checked, and each edge represents a specific behavior that must occur during the distributed system testing.

[0068] Specifically, the testing method uses a depth-first algorithm based on edge coverage to traverse the state space. Starting from the initial state, it iteratively traverses its successor nodes. The traversal of a single path terminates at a state node in two cases: all the edges from this state node to its successor nodes have been visited before, or a predefined end state of this test scenario is reached. The end state of a specific test scenario is determined by the user. When any termination condition is met, this path is added to the final test case set; otherwise, the traversal process continues. For each edge from this state node to its successor node, if it has not been visited, the edge and its successor node are added to the test path, and the node is given to continue the traversal process. Otherwise, it is directly skipped, and other successor nodes of the current node are continued to be visited.

[0069] Next, the testing method uses the set of test cases to test the target distributed system. The user configures the root directory of the target distributed system, cluster configuration, startup script location and commands, client request script location, etc. The testing method conducts a round of testing for each generated test case according to the user configuration. Specifically, as Figure 4 shown, the testing method first reads the test case to obtain the parameter values for this round of testing, including the number of cluster nodes, client request types and quantities, fault types and quantities, specific behaviors and state values. Then, the testing method initializes and starts the system cluster with a specific number of cluster nodes. Since the behaviors were mapped to the corresponding distributed system code in the previous steps, the system can collect relevant information about the code during operation, including whether it is running, behavior names, behavior parameters, etc. When the system runs to a specific behavior, the testing method suspends the current process and notifies the central module of the testing method. According to the behavior sequence in the test path, the testing method decides whether to resume this process or wait for other behaviors to execute first. Faults and client requests are initiated by the central module of the testing method, and their occurrence order is the same as that of other behaviors. Before a behavior finishes running, the testing method collects and checks the real-time values of all state-related variables of the target system.

[0070] Finally, the test method determines whether the actual execution of the target system is consistent with the definition in the TLA+ formal specification during behavior scheduling and state checking. The determination methods are divided into two types: state-based determination and behavior-based determination. If any of the determination conditions fails, it means that the test fails, that is, there is an inconsistency between the system implementation and the formal specification. For the state, if any state variable value collected during the operation of the target system is different from the value of the corresponding state node in the test path, that is, the variable value used in the formal specification to describe the state within the distributed system, it is reported as an inconsistency between the system implementation and the formal specification. For the behavior, if the next behavior in the test path never occurs during the operation of the target system, or the next behavior that occurs during the operation of the target system does not exist in the test path, it is reported as an inconsistency between the system implementation and the formal specification.

[0071] Although the specific implementation process and example drawings of the present invention are disclosed for illustrative purposes, and the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the shown implementation process and example drawings.

Claims

1. A testing method for a distributed system based on model checking of TLA+ formal specifications, the steps of which include: For the TLA+ formal specification of the target distributed system, map the elements of the TLA+ formal specification to the implementation of the target distributed system, and collect the running data of the target distributed system; wherein, the running data includes: the values of node state variables during operation and the occurrence of behaviors during operation; Use the TLC model checker to verify the TLA+ formal specification and obtain an abstract state space graph with verified correctness; Perform path traversal on the abstract state space graph, and regard each obtained path as a test case; Based on the test case, conduct a controlled test on the target distributed system to obtain controlled test data; wherein, the controlled test data includes: the variable values of state nodes in the test path and the occurrence of behaviors in the test path; Compare the running data with the controlled test data to generate a test report.

2. The method according to claim 1, characterized in that the elements include: variable elements used to describe the states within the distributed system and behavior elements that cause state updates within the distributed system. The variable elements include: state-related variables and message-related variables. The behavior elements include: internal node behaviors, message-related behaviors, user requests, and faults. The faults include: node downtime, node restart, message loss, and message duplication. Among them, the state-related variables are used to define the key state values in the distributed system, the message-related variables are used to simulate the information exchange in the message communication of the target distributed system in the form of shared variables, the internal node behaviors are used to describe the behaviors that update the node state values, the message-related behaviors are used to describe the behaviors that update the node state values through message passing, the user requests are used to describe the external behaviors of the system initiated by users, and the faults are used to describe external uncertain behaviors.

3. The method according to claim 2, characterized in that the mapping of the state-related variables to the implementation of the target distributed system includes: Annotate the state-related variables with @Variable($VNAME$); where $VNAME$ represents the naming of the state-related variable in the TLA+ formal specification; For each state-related variable, based on the annotation of the state-related variable, insert a stake in the target distributed system and add a static shadow field to the mapped variable; wherein, the static shadow field is used to assign values based on the initialization or update of the node state variable values.

4. The method according to claim 2, characterized in that the mapping of the internal node behaviors to the implementation of the target distributed system includes: Annotate the internal node behaviors with @Behavior($BNAME$); where $BNAME$ represents the naming of the internal node behavior in the TLA+ formal specification; For each of the internal behaviors of the nodes, instrument the target distributed system based on the annotation of the internal behaviors of the nodes, and insert a line of code at the entry and return of each mapped behavior method; wherein, the code at the entry is used to collect the real-time values of the behavior parameters to block the method from continuing to run and wait for the scheduling instruction to return, and the code at the return is used to collect the real-time values of all state-related variables.

5. The method according to claim 2, wherein, the implementation of mapping the message-related behaviors to the target distributed system includes: annotating the message-related behaviors; for each of the internal behaviors of the nodes, instrument the target distributed system based on the annotation of the message-related behaviors to collect the message values transmitted in the target distributed system.

6. The method according to claim 2, wherein, the implementation of mapping the user requests to the target distributed system includes: using the user scripts of the target distributed system to establish the correspondence between the user request behaviors in the TLA+ formal specification and the user scripts; based on the correspondence, match the user requests and the user scripts.

7. The method according to claim 2, wherein, the implementation of mapping the faults to the target distributed system includes: in the case where the fault is the node downtime or the node restart, kill the process corresponding to the node and restart it through the saved node configuration; in the case where the fault is the message loss fault, instrument the message and network-related code, and ignore the internal logic of the message processing function and directly return; in the case where the fault is the message duplication fault, instrument the message and network-related code, and perform multiple repeated processing of the same message in the message processing function.

8. The method according to claim 1, wherein, the path traversal of the abstract state space graph, and regarding each obtained path as a test case, includes: starting from the initial state node of the abstract state space; judging whether the current state node is the termination state of the test scenario: if not, then sequentially visit each unvisited successor behavior edge of the current state node, add the successor behavior edge and the corresponding successor node to the traversal result, mark the successor edge as visited, the corresponding successor node as the current state node, and return to judge whether the current state node is the termination state of the test scenario; if so, generate a test case based on the traversal result.

9. The method according to claim 1, wherein, the comparison of the running data with the controlled test data to generate the test report of the target distributed system includes: comparing the values of the node state variables during runtime with the values of the variables of the state nodes in the test path to obtain a test report based on state determination; comparing the occurrence of behaviors during runtime with the occurrence of behaviors in the test path to obtain a test report based on behavior determination; Combining the test report based on state determination and the test report based on behavior determination to obtain the test report of the target distributed system; wherein, the test report of the target distributed system includes: whether the implementation of the target distributed system is consistent or inconsistent with the TLA+ formal specification.

10. An electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute any one of the methods recited in claims 1-9.

Citation Information

Patent Citations

  • Operation system specification formal verification and test method

    CN108509336A

  • Apparatus and method for generating model reference tests

    US6148277A