General concurrent defect detection method and system based on partial order relation compatible with control flow change

CN116383076BActive Publication Date: 2026-10-09INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310379726.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2026-10-09
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

除以上检测技术外,还存在针对特定并发缺陷的检测技术,但其局限性一般较高,难以扩展

Benefits of technology

[0047] 1. By utilizing partial order relations, the detection of multiple concurrent defects is transformed into the verification of the sequence to be tested, providing a general detection technology for multiple concurrent defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116383076B_ABST
    Figure CN116383076B_ABST
Patent Text Reader

Abstract

The application provides a general concurrent defect detection method and system based on a partial order relation compatible with control flow changes, which comprises the following steps: an original concept of fixed point event, equivalent sequence and suffix boundary is proposed; a target program to be tested is instrumented and original execution sequences are collected; fixed point events and suffix boundary information are calculated according to static information recorded in the execution sequences and the target program; part of the events are removed to obtain a simplified execution sequence according to the sharing condition of variables among multiple threads; suspicious event pairs are found and suspicious sequences are constructed according to the target defect type; equivalent sequences are constructed for the suspicious sequences; finally, the sequence to be tested is verified through a partial order closure algorithm, and loop checking is performed to confirm whether the suspicious event pairs constitute concurrent defects. The application is compatible with control flow changes, covers concurrent defects such as data race, deadlock, atomicity violation, reuse after release, null pointer dereference and multiple pointer release, and can obtain a detection result in polynomial time without false positives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of software testing and software reliability, and specifically relates to a method and system for detecting concurrent defects in multithreaded programs. Background Technology

[0002] With the development of hardware technology, the improvement of single-core CPU processing power has slowed down, while multi-core technology has rapidly matured and gradually become the market mainstream. Against this backdrop, to cope with increasingly complex application demands, the software field is accelerating concurrency, fully utilizing hardware resources and improving software performance through multi-threaded programs. However, the widespread use of concurrent programs has brought new security problems. In concurrent programs, multiple threads share a large number of resources. If errors occur in the synchronization operations between threads, it will lead to competition for shared resources, resulting in concurrency defects, causing data errors, system crashes, or even malicious attacks, resulting in huge losses. Worse still, the scheduling of multi-threaded programs is unpredictable, and their scheduling space can explode, making the detection of concurrency defects a huge challenge.

[0003] Common concurrency defects include data races, deadlocks, and atomicity violations. A data race occurs when two or more threads simultaneously operate on the same shared variable, with at least one of these operations being a write operation. A deadlock occurs when each thread in a group is blocked because it holds a portion of resources and requests resources held by other threads in the group. An atomicity violation occurs when, during the execution of an atomic instruction sequence, other threads share access to data involved in that sequence. Besides these three types of concurrency defects, other defects derived from data races exist, including reuse after deallocation, null pointer dereferencing, and multiple pointer deallocations.

[0004] To address the aforementioned concurrency defects, existing detection techniques can be categorized into two types: static detection methods and dynamic detection methods. The former detects concurrency defects based on the program's source code, utilizing techniques such as data flow analysis, symbolic execution, and formal verification. The latter utilizes runtime information for detection, and can be further divided into three categories based on the model employed: lock-set based, runtime detection analysis, and partial order based. Lock-set based techniques treat access to unprotected shared variables as concurrency defects, generating numerous false positives. Runtime detection analysis techniques detect defects by repeatedly running the program and exhaustively exploring all legal thread interleaving scenarios. This type of technique boasts high accuracy but suffers from the state space explosion problem. Partial order based techniques collect runtime information into execution sequences, recording the operations of each thread and their order of occurrence; each operation is called an event. Then, based on read-write event relationships and lock information, partial order relationships are established between these events. Events without a partial order relationship are considered concurrent events, and further detection is performed on these concurrent events using algorithms such as closures. Partially ordered detection techniques rely on the partially ordered model and sequence verification algorithm for their detection effectiveness and efficiency. A reasonable model can ensure the reliability of the results, while efficient sequence verification algorithms can achieve polynomial time complexity. Besides the above detection techniques, there are also detection techniques for specific concurrent defects, but these generally have significant limitations and are difficult to scale.

[0005] Partial order-based techniques excel in both efficiency and accuracy, meeting the needs of practical applications, but they suffer from false negatives. The fundamental reason lies in the fact that these techniques infer other feasible execution sequences based on existing execution sequences, only covering the state space adjacent to the existing sequence. Their coverage depends heavily on the partial order model. Because control flow transformations involve chain reactions, existing partial order models require the inferred execution sequence to maintain a consistent control flow with the original sequence to ensure reliability. Furthermore, since it's impossible to determine which events affect the control flow, existing partial order models employ a strong assumption: any read / write event can potentially affect the control flow of the execution sequence. This becomes a bottleneck for these techniques, severely limiting their detection effectiveness. Summary of the Invention

[0006] The purpose of this invention is to provide a general concurrency defect detection method and system based on partial order relations that is compatible with control flow variations. This method allows for different control flows between the inferred execution sequence and the original execution sequence, while ensuring the reliability of the results. This relaxes the restrictions of partial order relations, covers a wider state space, and accurately detects more concurrency defects. The method is not limited to a specific type of concurrency defect, but encompasses data races, deadlocks, atomicity violations, and various concurrency defects derived from data races. The method can complete the detection in polynomial time and guarantees the reliability of the results, i.e., no false positives.

[0007] To simplify the description of this method, the following notation is introduced:

[0008] Given an event e, its corresponding thread number is tid(e), and its corresponding static code instruction is I(e).

[0009] To achieve the above objectives, this method introduces the following original concepts:

[0010] Observing a read event: Given a read event r(x), its corresponding write event is w(x), that is, the value obtained by r(x) is written by w(x). If r(x) and w(x) belong to different threads, then r(x) is called an observed read event (Obr event).

[0011] Fixed-point event: Given the original execution sequence σ, an event e, and a read event e in the same thread. r , where e r Before e, in thread tid(e), if for what happens at e r Any observation read event e' between e and e r , I(e) versus I(e' r There are no data dependencies or control dependencies, and I(e) r If no lock-related events can be executed between I(e) and I(e), then e is called e. r A fixed-point event, e r The anchor event for e. The observation read event e. r The set of fixed-point events is represented as FpS(e r ).

[0012] Equivalent sequence: Given a sequence<e1,e2,...> Where e2 is the observed read event e r A sequence is a fixed-point event. <e1,e r ...> is a sequence<e1,e2,...> The equivalent sequence.

[0013] Postfix boundary: Given the original execution sequence σ, an observed read event e r Let e′ be the value of e in the same thread. r After that, with e r The closest and target address is affected by e r The affected read / write events cause e br For the same thread e r The first branch event after that, if e′ precedes e br If it occurs, then e r The suffix boundary is e′, otherwise it is e. br .

[0014] Based on the above concepts, the present invention further adopts the following technical solutions:

[0015] A general concurrent defect detection method based on partial order relation that is compatible with control flow variations includes the following steps:

[0016] Instrument the target program to record memory read / write events, lock-related events, and branch events to obtain the original execution sequence;

[0017] Analyze the observed read events in the original execution sequence to obtain their corresponding fixed-point event sets and suffix boundaries;

[0018] Filter the events in the original execution sequence to obtain a simplified execution sequence;

[0019] Based on the simplified execution sequence and the type of defect to be detected, identify suspicious event pairs;

[0020] Based on the type of defect to be detected, generate a suspicious sequence with potential defects for suspicious event pairs;

[0021] Based on the set of fixed-point events, construct equivalent sequences for the suspicious sequences. The suspicious sequences and equivalent sequences are collectively referred to as the sequences to be tested.

[0022] Based on the simplified execution sequence and the sequence to be tested, combined with the suffix boundary, determine the event set and the partial order set;

[0023] Construct a directed graph for the event set and the partial order set, where vertices and edges correspond to events and partial orders, respectively, and solve the partial order closure based on this directed graph. If no loops are generated in the directed graph, the sequence to be tested can actually occur, and the events in the original suspicious sequence or the original suspicious sequence corresponding to the equivalent sequence have concurrency defects. Otherwise, it is determined that the sequence to be tested cannot actually occur and there are no concurrency defects.

[0024] Furthermore, during the instrumentation process, static information is added to each event that needs to be recorded, and this static information can be used to map to the static program instruction corresponding to that event.

[0025] Furthermore, branch events include explicit conditional branch events and implicit branch events in object-oriented programming.

[0026] Furthermore, when calculating fixed-point events and suffix boundaries, a program dependency graph is constructed for the program. Events are mapped to nodes in the graph using static information, and the dependencies between instructions (including data dependencies and control dependencies) are obtained through reachability analysis, thereby calculating fixed-point events and suffix boundaries.

[0027] Furthermore, the nodes in the program dependency graph represent program instructions (and also temporary variables) and global variables, and the edges represent the direct dependencies between variables, which are divided into control dependencies and data dependencies. Data dependencies are further divided into explicit data dependencies and implicit data dependencies. The former is calculated through use-def relationships, and the latter is analyzed through pointer aliases.

[0028] Furthermore, pointer aliasing analysis can employ various existing methods, but to ensure that there are no false positives in the results of fixed-point events and suffix boundaries, if two pointers are likely to be aliased, they should be treated as aliases.

[0029] Furthermore, by removing the events corresponding to single-threaded exclusive variables and read-only variables from the original execution sequence, a simplified execution sequence is obtained.

[0030] Furthermore, the types of defects to be detected include, but are not limited to, data races, deadlocks, atomicity violations, reuse after free, null pointer dereferences, and pointers being freed multiple times.

[0031] Furthermore, regarding data contention, the suspicious event pairs are defined as: access events to the same memory object from two different threads, at least one of which is a write event, and the lock sets held by the two threads do not conflict. The two events in the suspicious event pair are arranged sequentially, ensuring that there are no other events between them, thus forming a suspicious sequence.

[0032] Furthermore, regarding deadlock, the suspicious event pairs are defined as follows: two lock request events within different threads, whose lock sets do not conflict, and which are mutually requesting the lock that the other thread is currently holding. The events in the suspicious event pair are then queued after the most recent lock acquisition event in the other thread corresponding to the currently requested lock, forming a suspicious sequence.

[0033] Furthermore, regarding atomicity violations, the suspicious event pairs are: events performed by the two atoms, and other events that operate on the same shared variable as the two atoms. Placing the third event between the two atomic events forms a suspicious sequence.

[0034] Furthermore, regarding reuse after release, the suspicious event pairs are: pointer release events and usage events related to that pointer. Placing pointer usage events after release events, and requiring that there are no other related events for the same pointer between them, constitutes a suspicious sequence.

[0035] Furthermore, regarding the dereferencing of a null pointer, the suspicious event pairs are: the pointer being set to null and the usage event associated with that pointer. Placing the pointer usage event after the nulling event, and requiring that there are no other related events for the same pointer between them, constitutes a suspicious sequence.

[0036] Furthermore, for multiple pointer releases, the suspicious event pairs are: two release events targeting the same pointer. Arranging these two events sequentially, and ensuring there are no other related events targeting the same pointer between them, constitutes a suspicious sequence.

[0037] Furthermore, replacing the events in the suspicious sequence with the corresponding anchor events yields the equivalent sequence. Not all suspicious sequences have an equivalent sequence; if one exists, the equivalent sequence is used as the test sequence; otherwise, the original suspicious sequence is used as the test sequence.

[0038] Furthermore, the partial order set includes program partial order, observation partial order, lock-escaping partial order, and partial order in the sequence to be tested. Among them, program partial order means that events within the same thread must occur in the order of the original execution sequence; observation partial order means that if a memory read event and the memory write event read by that event are in different threads, the read-write relationship must always be maintained; and lock-escaping partial order means that a single lock object cannot be held by multiple threads at the same time.

[0039] Furthermore, if events in the event set hold the same lock, then the corresponding release event needs to be added to the event set as well; at the same time, a partial order is added to the partial order set to make the two critical sections mutually exclusive; based on the event set, the observation read events of each thread are found, and if the corresponding suffix boundary is not in the event set, then the observation partial order corresponding to the observation read event is not added to the partial order set.

[0040] Furthermore, in the process of solving the partial order closure, closure calculation is performed for the observed partial order and the locked partial order, and further closure calculation is performed based on the transitivity of the partial order relation; for the additional requirement that there cannot be any adjacency relationship between events, the adjacency closure calculation is also required.

[0041] A general concurrent defect detection system based on partial order relations that is compatible with control flow variations includes the following modules:

[0042] Instrumentation and Sequence Collection Module: Instruments the target program under test. The instrumentation code is responsible for recording memory access events, lock-related events, and branch events, generating the original execution sequence, and adding static information to each event to map the event to the corresponding static program instruction.

[0043] Static information calculation module: Load the target program, construct the program dependency graph, map the events to static program instructions based on the static information in the events, and further map them to nodes in the graph. Then, calculate the set of fixed point events and suffix boundaries for the observed read events based on the reachability relationships in the graph.

[0044] Sequence generation module: Filters the original execution sequence, removes events corresponding to single-threaded exclusive variables and read-only variables in the original execution sequence to obtain a simplified execution sequence. Based on this, according to the type of defect to be tested, it finds suspicious event pairs, constructs suspicious sequences with potential defects, and further constructs equivalent sequences to obtain all sequences to be tested.

[0045] Sequence Analysis Module: Based on the simplified execution sequence and the sequence to be tested, the module determines the event set and the partial order set, and maps them one by one to nodes and edges in a directed graph. It then solves for the partial order closure on the directed graph. If no loops are generated after the closure calculation, the sequence to be tested can actually occur, and the events in the original suspected sequence or the original suspected sequence corresponding to the equivalent sequence have concurrency defects. Otherwise, it is determined that the sequence to be tested cannot actually occur, and there are no concurrency defects.

[0046] Compared with the prior art, the present invention has the following advantages:

[0047] 1. By utilizing partial order relations, the detection of multiple concurrent defects is transformed into the verification of the sequence to be tested, providing a general detection technology for multiple concurrent defects.

[0048] 2. By utilizing fixed-point events and equivalent sequences, the verification of the original suspicious sequence is transformed into the verification of the equivalent sequence. If the equivalent sequence can occur, the original suspicious sequence must also occur. Therefore, during the verification process, it is unnecessary to calculate the partial order relationship between anchor events and fixed-point events. While ensuring no false positives, this relaxes the partial order restriction, allowing the control flow between anchor events and fixed-point events in the inferred execution sequence to differ from the original execution sequence. This improves the coverage of the scheduling state space, thereby enabling the detection of more concurrent defects. The reliability of this invention (i.e., no false positives) can be proven through rigorous theoretical reasoning.

[0049] 3. By utilizing the concept of suffix boundaries, the partial order restriction is further relaxed while ensuring reliability. During closure computation, it is not necessary to calculate the partial order of observations that are not in the event set due to the suffix boundary, thereby detecting more concurrency defects. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating the implementation of a general concurrent defect detection method based on partial order relations that is compatible with control flow changes, as described in this invention. Specific Implementation

[0051] To make the technical solution of the present invention clear and easy to understand, a detailed description is provided in conjunction with the accompanying drawings, disclosing an embodiment of a concurrent defect detection system based on partial order that is compatible with control flow variations.

[0052] The processing steps of this embodiment are divided into four stages: instrumentation and sequence collection, static information calculation, sequence generation, and sequence analysis. These stages can be executed by the instrumentation and sequence collection module, static information calculation module, sequence generation module, and sequence analysis module in a partial-order-based concurrency defect detection system compatible with control flow changes. The sequence generation module includes a set of sequence generators targeting various concurrency defects such as data races, deadlocks, atomicity violations, reuse after release, null pointer dereferences, and multiple pointer releases. These sequence generators can be further modularly extended to accommodate more types of concurrency defects. The specific implementation methods for the above four stages are as follows:

[0053] I. Instrumentation and Sequence Collection Phase

[0054] To instrument the target program under test, firstly, each static program instruction that needs to be recorded is assigned a globally unique static ID. Then, instrumentation code is inserted to record memory access events, lock-related events, and branch events. The aforementioned static ID is recorded in the event information during event recording. Specifically, for memory access events, the operation type (read or write) and the memory object operated on need to be recorded. The object can be identified by the target address of the read / write event, using r(x) / w(x) to represent read and write events on the shared variable x, respectively. For lock-related events, the operation type and the lock object operated on are also recorded, using req(l) / acq(l) / rel(l) to represent the acquisition, release, and request events on lock l, respectively. For branch events, unconditional branches do not need to be recorded; only conditional branches and function calls in object-oriented programming need to be recorded. No additional information such as branch conditions needs to be recorded; branch events can be represented by br. Furthermore, for all types of events, the globally unique ID of the static instruction corresponding to the current event needs to be recorded. After instrumentation is complete, the program can be executed, and the recorded execution sequence is written to a file.

[0055] II. Static Information Calculation Stage

[0056] This stage is responsible for calculating fixed-point events and suffix boundary information based on the original execution sequence and static program information. The specific implementation method is as follows:

[0057] First, information is extracted from the program's source code (or intermediate code) and stored in a graph. Global variables and instructions (also representing temporary variables) are mapped to nodes in the graph, and control flow and data flow dependencies are mapped to directed edges. The endpoint of a directed edge has a direct dependency on its origin. Directed edges can be calculated as follows: First, construct the program's control flow graph; then, determine the control dependencies between nodes based on the control flow information in the graph. Next, construct use-def chains between instructions, where adjacent instructions in the chain have direct data flow dependencies. Finally, perform pointer aliasing analysis. Aliasing analysis can use any existing technique. To ensure the final result is free of false positives, if there is a possibility that two pointers are aliased, they are considered aliased, and a bidirectional edge is added between them in the graph.

[0058] Then, the original execution sequence is traversed, and events are mapped to static program instructions based on the static ID recorded in each event, and further mapped to nodes in the program dependency graph. Then, all observed read events are traversed, and the reachability relationship between nodes in the graph is used to determine whether there is a dependency relationship between two instructions, and fixed-point events and suffix boundaries are calculated according to the definition.

[0059] III. Sequence Generation Stage

[0060] This stage is responsible for identifying suspicious event pairs and constructing corresponding test sequences. The specific implementation method is as follows:

[0061] First, the original execution sequence is filtered for events. Events corresponding to single-threaded exclusive variables and read-only variables are unlikely to cause concurrency defects. These events are removed to obtain a simplified execution sequence.

[0062] Then, the simplified execution sequence is traversed, and suspicious event pairs are searched for based on the target defect type to construct a suspicious sequence. For data races, a suspicious event pair is defined as: access events to the same memory object from two different threads, at least one of which is a write event, and the lock sets held by the two threads do not conflict. The two events are arranged sequentially, and no other events are required between them to form a suspicious sequence. For deadlocks, a suspicious event pair is defined as: two lock request events from different threads, whose lock sets do not conflict, and which are requesting locks held by each other. The two events are arranged after the most recent lock acquisition event in the other thread corresponding to the requested lock, forming a suspicious sequence. For atomicity violations, a suspicious event pair is defined as: two atomically executed events and other events that operate on the same shared variable. The third event is arranged between the two atomic events, forming a suspicious sequence. For reuse after free, a suspicious event pair is defined as: a pointer release event and a usage event related to that pointer. The pointer usage event is arranged after the release event, and no other related events for the same pointer are required between them, forming a suspicious sequence. For dereferencing a null pointer, a suspicious event pair is: a pointer-to-null event and a related usage event. The pointer usage event is placed after the pointer-to-null event, and there must be no other related events for the same pointer between the two events, forming a suspicious sequence. For multiple pointer releases, a suspicious event pair is: two release events targeting the same pointer. The two events are arranged sequentially, and there must be no other related events for the same pointer between the two events, forming a suspicious sequence.

[0063] Finally, equivalent sequences are constructed for suspicious sequences by replacing fixed-point events in the suspicious sequence with corresponding anchor events. If no equivalent sequence exists for a suspicious sequence, the suspicious sequence is used as the test sequence; otherwise, its equivalent sequence is used as the test sequence. A single suspicious sequence may correspond to multiple equivalent sequences; to avoid missed detections, all these equivalent sequences are tested as test sequences.

[0064] IV. Sequence Analysis Stage

[0065] This stage is responsible for verifying whether the test sequence can actually occur. The specific implementation method is as follows:

[0066] First, based on the simplified execution sequence and the sequence under test, determine the event set S and the partial order set P. The initial event set is the set of events in the sequence under test, with further requirements: if an event is in S, then all other events preceding that event in the same thread should also be in S; if an observed read event is in S, then its corresponding write event should also be in S; if two events in S hold the same lock, then the corresponding release event should also be in S. The partial order set includes program partial order, observed partial order, lock estrangement partial order, and partial order in the sequence under test, with the requirement that: if the start and end events of a partial order relationship are both in S, then that partial order relationship should be in P. Furthermore, it is necessary to find the observed read events of each thread based on the event set. If the corresponding suffix boundary is not in the event set, then the observed partial order corresponding to that observed read event is not added to the partial order set P.

[0067] Then, based on the event set S and the partial order set P, the partial order closure is calculated. The event set S is mapped to nodes in a directed graph, and the partial order relations in the partial order set P are mapped to directed edges between nodes. The closure is then further calculated according to the following rules:

[0068] (1) Transitive closure: Add edge<u,v> After that, transitive edges will be generated, which are the edges between u′, the predecessor of u in each thread (assuming there are k threads in total), and v′, the successor of v in each thread.

[0069] (2) Observe the closure: Check all read and write edges in the graph in turn.<w,r> For all other write events w′ that occur at the same memory location as the w operation, w′ must not occur between w and r; that is, an edge needs to be added as needed.<w′,w> or<r,w′> .

[0070] (3) Lock closure: For all critical sections in the graph, it is required that critical sections with the same lock cannot intersect. It is necessary to add a partial order between critical sections as needed, which can be achieved by adding the corresponding edges between rel and acq events.

[0071] (4) Adjacency Closure: For the additional requirement that events cannot have adjacency relationships with other related events, the adjacency closure is performed. For the pair of adjacent events (e1, e2) and any other event e, it is required that there is at least one edge in the graph.<e,e1> and<e2,e> One of them.

[0072] The closure is calculated on the directed graph according to the above rules. If no loop is generated during the calculation, the sequence under test can actually occur, and the events in the sequence under test (or its corresponding original suspected sequence) have concurrency defects. Otherwise, it is determined that the sequence under test cannot actually occur and there are no concurrency defects.

[0073] The four modules mentioned above and their corresponding functions are as follows:

[0074] Instrumentation and Sequence Collection Module: Instruments the target program under test. The instrumentation code is responsible for recording memory access events, lock-related events, and branch events, generating the original execution sequence, and adding static information to each event to map the event to the corresponding static program instruction.

[0075] Static information calculation module: Load the target program, construct the program dependency graph, map the events to static program instructions based on the static information in the events, and further map them to nodes in the graph. Then, calculate the set of fixed point events and suffix boundaries for the observed read events based on the reachability relationships in the graph.

[0076] Sequence generation module: Filters the original execution sequence, removes events corresponding to single-threaded exclusive variables and read-only variables in the original execution sequence to obtain a simplified execution sequence. Based on this, according to the type of defect to be tested, it finds suspicious event pairs, constructs suspicious sequences with potential defects, and further constructs equivalent sequences to obtain all sequences to be tested.

[0077] Sequence Analysis Module: Based on the simplified execution sequence and the sequence to be tested, the module determines the event set and the partial order set, and maps them one by one to nodes and edges in a directed graph. It then solves for the partial order closure on the directed graph. If no loops are generated after the closure calculation, the sequence to be tested can actually occur, and the events in the original suspected sequence or the original suspected sequence corresponding to the equivalent sequence have concurrency defects. Otherwise, it is determined that the sequence to be tested cannot actually occur, and there are no concurrency defects.

[0078] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.

[0079] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) storing a computer program that, when executed by a computer, implements the various steps of the method of the present invention.

[0080] The above embodiments are provided merely for the purpose of describing the present invention and are not intended to limit the scope of the invention. The scope of the invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the invention should be covered within the scope of the invention.

Claims

1. A general concurrent defect detection method based on partial order relations that is compatible with control flow variations, characterized in that, Includes the following steps: Instrument the program under test to record memory read / write events, lock-related events, and branch events to obtain the original execution sequence. Analyze the observed read events in the original execution sequence to obtain their corresponding fixed-point event sets and suffix boundaries; Filter the events in the original execution sequence to obtain a simplified execution sequence; Based on the simplified execution sequence and the type of defect to be detected, identify suspicious event pairs; Based on the type of defect to be detected, generate a suspicious sequence with potential defects for suspicious event pairs; Based on the set of fixed-point events, construct equivalent sequences for suspicious sequences, and use both suspicious sequences and equivalent sequences as sequences to be tested. Based on the simplified execution sequence and the sequence to be tested, combined with the suffix boundary, determine the event set and the partial order set; Construct a directed graph for the event set and the partial order set, where vertices and edges correspond to events and partial orders, respectively, and solve the partial order closure based on this directed graph; If no loops are generated in the directed graph, then the sequence under test can actually occur, and the events in the suspected sequence or the suspected sequence corresponding to the equivalent sequence have concurrency defects; otherwise, it is determined that the sequence under test cannot actually occur and there are no concurrency defects. The observed read event is: given a read event The corresponding write event is ,Right now The value obtained is from Written, if and If they belong to different threads, then they are called... For an observation, read the event; Given event Its corresponding thread number is Its corresponding static code instruction is The fixed-point event is: given the original execution sequence An event And a read event observed in the same thread ,in What happened Previously, in threads In the middle, if for what happens and Any observation read event between , right There are no data dependencies or control dependencies, and and If no lock-related events can be executed between them, then it is called... for A fixed-point event, for Anchor point events; The equivalent sequence is: given a sequence ,in To observe the event A sequence is a fixed-point event. For sequence Equivalent sequences; The suffix boundary is defined as follows: given the original execution sequence An observation of an event ,make In the same thread After that and The closest and target address is affected The affected read / write events make In the same thread The first branch event after that, if Prior to If it happens, then The suffix boundary is Otherwise ; The analysis of observed read events in the original execution sequence to obtain their corresponding fixed-point event set and suffix boundary includes: constructing a program dependency graph for the program under test, mapping events to nodes in the program dependency graph through static information, obtaining the dependency relationships between program instructions through reachability analysis, and then calculating the fixed-point event set and suffix boundary; the static information is the static ID recorded in each event; The simplified execution sequence is obtained by removing the events corresponding to single-threaded exclusive variables and read-only variables from the original execution sequence; The partial order set includes program partial order, observation partial order, lock escaping partial order, and partial order in the sequence to be tested. When determining the event set and the partial order set, if events in the event set hold the same lock, the corresponding release event needs to be added to the event set as well. At the same time, the partial order is added to the partial order set to make the two critical sections mutually exclusive. The observation read events of each thread are found according to the event set. If the corresponding suffix boundary is not in the event set, the observation partial order corresponding to the observation read event is not added to the partial order set.

2. The method as described in claim 1, characterized in that, During the instrumentation process, static information is added to each event that needs to be recorded, and this static information is used to map the event to a static program instruction; the branch events include explicit conditional branch events and implicit branch events in object-oriented programming.

3. The method as described in claim 1, characterized in that, The nodes in the program dependency graph represent program instructions and global variables, and the edges represent direct dependencies between global variables, which are divided into control dependencies and data dependencies. Data dependencies are further divided into explicit data dependencies and implicit data dependencies. The former is calculated through use-def relationships, and the latter is obtained through pointer alias analysis. The pointer alias analysis adopts an arbitrary alias analysis method, but in order to ensure that there are no false positives in the results of fixed point events and suffix boundaries, if two pointers have the possibility of being aliased, they are treated as aliases.

4. The method as described in claim 1, characterized in that, The types of defects to be detected include data races, deadlocks, atomicity violations, reuse after free, null pointer dereferences, and multiple pointer frees; the methods for constructing suspicious sequences for each type of defect include: For data races, the suspicious event pairs are: access events to the same memory object in two different threads, at least one of which is a write event, and the lock sets held by the two threads do not conflict; the two events in the suspicious event pair are arranged in order, and there are no other events between them, forming a suspicious sequence; For deadlock, the suspicious event pairs are: two lock request events in different threads, whose lock sets do not conflict, and which request each other's locks; the events in the suspicious event pairs are arranged after the most recent lock acquisition event of the corresponding currently requested lock in the other thread, forming a suspicious sequence; For atomicity violations, the suspicious event pairs are: events performed by two atoms and other events that operate on the same shared variable as the two atoms; placing the third event between the two atomic events forms a suspicious sequence; For reuse after freeing, the suspicious event pairs are: pointer freeing event and pointer usage event related to that pointer; the pointer usage event is placed after the pointer freeing event, and there should be no other related events for the same pointer between the two, which constitutes a suspicious sequence; For dereferencing a null pointer, the suspicious event pairs are: the pointer being set to null and the pointer being used in relation to that pointer. The pointer being used is placed after the pointer being set to null, and there must be no other related events for the same pointer between the two, forming a suspicious sequence. For multiple pointer releases, the suspicious event pairs are: two pointer release events targeting the same pointer; the two pointer release events are arranged in order, and there are no other related events targeting the same pointer between the two pointer release events, which constitute a suspicious sequence.

5. The method as described in claim 1, characterized in that, The equivalent sequence is obtained by replacing the events in the suspicious sequence with the corresponding anchor events. Not all suspicious sequences have an equivalent sequence. If an equivalent sequence exists, the equivalent sequence is used as the sequence to be tested. If no equivalent sequence exists, the suspicious sequence is used as the sequence to be tested.

6. The method as described in claim 1, characterized in that, In solving the partial order closure, closure calculation is performed for observed partial order and locked partial order, and further closure calculation is performed based on the transitivity of the partial order relation; for the additional requirement that there cannot be adjacency relations between events, adjacency closure calculation is also required.

7. A general concurrent defect detection system based on partial order relations that is compatible with control flow variations and employs the method described in any one of claims 1 to 6, characterized in that, Includes the following modules: Instrumentation and Sequence Collection Module: Responsible for instrumenting the target program under test. The instrumentation code is responsible for recording memory access events, lock-related events, and branch events, generating the original execution sequence, and adding static information to each event to map the event to the corresponding static program instruction. Static information calculation module: responsible for loading the target program under test, constructing the program dependency graph, mapping the events to static program instructions based on the static information in the events, and further mapping them to nodes in the graph, and then calculating the set of fixed point events and suffix boundaries for the observed read events based on the reachability relationship in the program dependency graph; The sequence generation module is responsible for filtering the original execution sequence, removing events corresponding to single-threaded exclusive variables and read-only variables in the original execution sequence to obtain a simplified execution sequence. Based on this, according to the type of defect to be tested, suspicious event pairs are found, suspicious sequences with potential defects are constructed, and equivalent sequences are further constructed to obtain all sequences to be tested. The sequence analysis module is responsible for determining the event set and partial order set based on the simplified execution sequence and the sequence to be tested, and mapping them one by one to the nodes and edges in the directed graph. It then solves the partial order closure on the directed graph. If no loop is generated after the closure calculation, the sequence to be tested can actually occur, and the events in the suspected sequence or the suspected sequence corresponding to the equivalent sequence have concurrency defects. Otherwise, it is determined that the sequence to be tested cannot actually occur and there are no concurrency defects.

Citation Information

Patent Citations

  • Concurrent program defect detection method based on memory access mode

    CN114911695A

  • Universal concurrent defect detection method and system based on partial order relation

    CN115080374A