Template-based Concurrent Defect Repair System and Method
By converting concurrent related defect reports into change tree representations and extracting repair templates, the problem of inaccurate and easy introduction of deadlocks in the prior art is solved, and more efficient and accurate concurrent defect repair is achieved.
Patent Information
- Application Number
- CN202111507043.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art is prone to introducing new deadlocks when repairing concurrent defects, and repair does not take into account the defect context, resulting in low accuracy of repairs.
By crawling the concurrent related defect reports in the open source defect tracking system, extracting the patch file and converting it into a special change tree representation, comparing the similarity of the change tree to extract the repair template, and finally using the AST context information of the repair template and the file to be repaired for template matching and source code reconstruction to repair concurrent defects.
It achieves higher repair accuracy and accuracy, avoids the introduction of new deadlocks, and has stronger universality and universality of repairs, and can repair more types of concurrent defects.
Smart Images

Figure CN114327575B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software security, and particularly relates to a template-based concurrent defect repair system and method. Background Art
[0002] With the development of computer software technology, concurrent programs are widely used in the development process of software systems. Due to their inherent concurrency, the uncertainty of thread scheduling, and the large code scale, concurrent programs are prone to encounter concurrent defects. Concurrent defects usually cause the program or system to fall into an uncertain running state, and even lead to serious problems such as system crashes, resulting in huge economic losses. However, concurrent defects often occur in rare and specific interleaved executions, making it difficult for developers to detect and repair them effectively and efficiently. Therefore, the research on concurrent defect repair has become a hot issue in the academic and industrial fields in recent years. When developers face a large number of concurrent defects, if they can mine relevant repair templates based on the defective code and the repair code, the efficiency of developers in repairing concurrent defects can be greatly improved.
[0003] Currently, most of the existing concurrent defect repair work inserts door locks dynamically or statically to serialize the execution of all threads involved in concurrent defects. For example, the literature "G. Jin, L. H, Song, W. Zhang, S. Lu, B. Liblit. Automated atomicity-violation fixing. In Proc. PLDI, 389–400, 2011." repairs common types of concurrent defects: atomicity violations by inserting new locks. However, this repair is likely to introduce new deadlocks and result in high runtime overhead. And this repair does not consider the defect context, that is, the aliasing configuration of shared variables and locks that cause the defect. Therefore, the accuracy of the repair cannot be guaranteed. Some work has also proposed improvements on this basis to avoid introducing new deadlocks in the repair. For example, the literature "P. Liu, O. Tripp, and C. Zhang. Grail: context-aware fixing of concurrency bugs. In Proc. FSE, 318–329, 2014." uses Petri-net analysis to eliminate the introduced deadlocks, but it can only repair deadlocks caused by two threads, and its generality is not high. Summary of the Invention
[0004] Object of the Invention: Aiming at the problems of the above-mentioned prior art, the object of the present invention is to provide a template-based concurrent defect repair system and method with characteristics such as a wider application field, higher precision, and more accurate positioning.
[0005] Technical solution: To achieve the above invention objectives, the present invention specifically adopts the following technical solutions:
[0006] The present invention provides a template-based concurrent defect repair system, including: a dataset collection module, used to retrieve and screen all concurrency-related defects by crawling defect reports of open-source projects in an open-source defect tracking system, and extract the corresponding patch files; a rich edit script calculation module, used to convert the patch files into a special change tree representation; a repair template extraction module, used to extract repair templates by comparing the similarity of change trees; a concurrent defect repair module, used to perform template matching with the repair templates and the files to be repaired with known defect line numbers as inputs, and complete the repair of concurrent defects by refactoring the source code after successful matching.
[0007] Further, in the dataset collection module:
[0008] Crawl defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "defect", and the status being "resolved", "fixed", and "closed"; retrieve defect reports with keywords such as "concurrent", "concurrency", "deadlock", "atomicity violation", "data race", "resource race", "order violation", "thread safety", "multi-threading", and "synchronization", screen all concurrency-related defects, and extract the corresponding patch files.
[0009] Further, in the rich edit script calculation module:
[0010] Use GumTree to calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file and remap it to the relevant nodes in the program AST; starting from the nodes of the GumTree edit operations, traverse the parsed program AST from bottom to top until reaching a predefined root node, and the predefined root node types are as follows: type declaration, field declaration, method declaration, branch statement, catch clause, constructor call, parent constructor call, or any statement node; for the reached predefined root node, extract the AST subtree between the root node and the leaf nodes mapped to the GumTree edit operations; for the extracted AST subtree, create an ordered sequence and store the ordered sequence as a rich edit script; the rich edit script describes the operations performed on a given AST of the program before repair to convert the AST of the program before repair into the AST of the program after repair, and each node in the tree is an AST node affected by the changed code, and each node in the rich edit script contains three different types of information: shape, operation, and token.
[0011] Further, in the repair template extraction module:
[0012] Construct a search index, that is, a set of comparison subspaces. Define a search index for each iteration, and sequentially obtain a shape index, an operation index, and a token index. Among them, the types of iterations include: shape, operation, and token. The construction method of the search index is to group rich edit scripts according to standards, where these standards depend on the embedding information represented by the change tree used in different iterations.
[0013] The construction of the shape index takes the shape tree representation of the rich edit script as input and groups it according to the AST node type based on the structure of the shape tree. Rich edit scripts with the same root node and the same depth are grouped into one group. For each group, a comparison space is created by enumerating pairwise combinations of group members. The shape index is constructed by storing the identifier of each group, represented as root node / depth.
[0014] The construction of the operation index follows the same principle as the shape index. It re-groups the clustering output based on the shape tree. The input is composed of the operation tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated by clustering based on the shape tree.
[0015] The construction of the token index follows the same principle as the operation index. It re-groups the clustering output based on the operation tree. The input is composed of the token tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated by clustering based on the shape tree / id generated by clustering based on the operation tree.
[0016] Calculate the tree edit distance of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree to another tree. For the first two iterations, that is, shape and operation, use the edit script algorithm of GumTree, take two trees as input, and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same. Two identical trees are called an identical tree pair.
[0017] To extract the repair template, use the clusters of the special change tree representation of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representation according to the iteration. Use a clustering process based on the theory of connected component identification in the graph to find a set of identical trees, that is, clusters. Create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent associated trees, that is, identical tree pairs. In the graph, a cluster is defined as a subgraph. Each subgraph contains a set of trees that are identical to each other and not related to each other. A cluster contains a set of rich edit scripts. The rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template.
[0018] Furthermore, in the concurrent defect repair module:
[0019] The open-source tool ANTLR is used to parse the file to be repaired into an AST, and the spaces and comments in the source code are also parsed by ANTLR into nodes of corresponding types. The matching of the repair template is based on the AST context information of the defective code. Each node of the AST of the suspicious statement is traversed in turn, from the first child node to the last leaf node of the AST of the suspicious statement, and each node is tried to be matched with the context AST of the repair template. If a node can match any of the given defective contexts, the corresponding repair template is successfully matched. If the node is not a leaf node, its leaf nodes are continued to be traversed. After successfully matching the repair template, insert, delete, update, and move operations are performed on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of concurrent defects.
[0020] In addition, the present invention provides a template-based concurrent defect repair method, including: Step 1, by crawling the defect reports of open-source projects in the open-source defect tracking system, retrieving and filtering all defects related to concurrency, and extracting the corresponding patch files; Step 2, converting the patch files into a special change tree representation; Step 3, extracting the repair template by comparing the similarity of the change trees; Step 4, using the repair template and the file to be repaired with known defective line numbers as inputs for template matching, and after successful matching, completing the repair of concurrent defects by reconstructing the source code.
[0021] Further, in Step 1:
[0022] Crawl the defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "defect" and the status being "resolved", "fixed", and "closed"; retrieve the defect reports with keywords such as "concurrent", "concurrency", "deadlock", "atomicity violation", "data race", "resource race", "order violation", "thread safety", "multi-threading", and "synchronization", filter all defects related to concurrency, and extract the corresponding patch files.
[0023] Further, in Step 2:
[0024] Use GumTree to calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file and remap it to the relevant nodes in the program AST; starting from the nodes of the GumTree edit operations, traverse the parsed program AST from bottom to top until reaching a predefined root node. The types of the predefined root nodes are as follows: type declarations, field declarations, method declarations, branch statements, catch clauses, constructor calls, super constructor calls, or any statement nodes; for the reached predefined root node, extract the AST subtree between the root node and the leaf nodes mapped to the GumTree edit operations; for the extracted AST subtree, create an ordered sequence and store the ordered sequence as a rich edit script; the rich edit script describes the operations performed on a given AST of the program before repair to convert the AST of the program before repair into the AST of the program after repair. Each node in the tree is an AST node affected by the changed code, and each node in the rich edit script contains three different types of information: shape, operation, and token.
[0025] Furthermore, in step 3:
[0026] Build search indexes, that is, a set of comparison subspaces. Define a search index for each iteration respectively, and obtain a shape index, an operation index, and a token index in sequence; where the types of iterations include: shape, operation, and token; the method for building the search indexes is to group the rich edit scripts according to standards, where these standards depend on the embedding information of the change tree representation used in different iterations;
[0027] The construction of the shape index takes the shape tree representation of the rich edit script as the input and groups it according to the AST node types based on the structure of the shape tree. Rich edit scripts with the same root node and the same depth are grouped into one group. For each group, create a comparison space by enumerating the pairwise combinations of the group members; the shape index is built by storing the identifier of each group, represented as root node / depth;
[0028] The construction of the operation index follows the same principle as the shape index. The regrouping is based on the clustering output of the shape tree, and the input is composed of the operation tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated based on the clustering of the shape tree;
[0029] The construction of the token index follows the same principle as the operation index. The regrouping is based on the clustering output of the operation tree, and the input is composed of the token tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated based on the clustering of the shape tree / id generated based on the clustering of the operation tree;
[0030] Calculate the tree edit distances of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree to another tree. For the first two iterations, namely shape and operation, use the edit script algorithm of GumTree. Take two trees as input and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same, and two identical trees are called an identical tree pair;
[0031] To extract the repair template, use the clusters of the special change tree representations of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representations according to the iteration. Use the clustering process based on the theory of connected component recognition in the graph to find a set of identical trees, that is, clusters; create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent associated trees, that is, identical tree pairs; in the graph, define a cluster as a subgraph. Each subgraph contains a set of trees that are identical to each other and unassociated. A cluster contains a set of rich edit scripts, and the rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template.
[0032] Furthermore, in step 4:
[0033] Use the open-source tool ANTLR to parse the file to be repaired into an AST. The spaces and comment contents in the source code are also parsed by ANTLR into corresponding types of nodes; the matching of the repair template is based on the AST context information of the defective code. Traverse each node of the AST of the suspicious statement in turn, from the first child node to the last leaf node of the AST of the suspicious statement, and try to match each node with the context AST of the repair template. If a node can match any given defective context, the corresponding repair template is successfully matched. If the node is not a leaf node, continue to traverse its leaf nodes; after successfully matching the repair template, perform insert, delete, update, and move operations on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of concurrent defects.
[0034] Beneficial effects: By converting the patch file into a special change tree representation, capturing the code change context, generating a more accurate and fine-grained edit script, and by comparing the similarity between change trees, the present invention can dig out accurate repair templates, can better utilize the syntax information of the defective code, fully explore its relationship with the context, repair more types of concurrent defects, has stronger universality and generality, and the repair is more accurate. Brief Description of the Drawings
[0035] Figure 1 is a flowchart of the template-based concurrent defect repair method of the present invention. Detailed Embodiments
[0036] The following specifically elaborates on the specific implementation manners of the present invention in combination with embodiments and the accompanying drawings.
[0037] Embodiment 1
[0038] This embodiment discloses a template-based concurrent defect repair system, as Figure 1 shown. The template-based concurrent defect repair system includes: a dataset collection module, which is used to retrieve and screen all concurrency-related defects by crawling defect reports of open-source projects in the open-source defect tracking system Bugzilla, and extract the corresponding patch files; a rich edit script calculation module, which is used to convert the patch files into a special change tree representation; a repair template extraction module, which is used to extract repair templates by comparing the similarity of the change trees; and a concurrent defect repair module, which is used to perform template matching with the repair templates and the files to be repaired with known defect line numbers as inputs, and complete the repair of concurrent defects by reconstructing the source code after successful matching.
[0039] Further, in the dataset collection module:
[0040] Crawl defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "bug" (defect), the status being "resolved" (resolved), "fixed" (fixed), and "closed" (closed); use "concurrent" (concurrent), "concurrency" (concurrency), "deadlock" (deadlock), "atomicity violations" (atomicity violation), "data race" (data race), "race condition" (resource race), "order violations" (order violation), "thread-safety" (thread safety), "multi-thread" (multi-thread), and "synchronization" (synchronization) as keywords to retrieve defect reports, screen all concurrency-related defects, and extract the corresponding patch files.
[0041] Further, in the rich edit script calculation module:
[0042] The patch file is used to calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file by GumTree and remap it to the relevant nodes in the program AST. Starting from the nodes of the GumTree edit operations, the parsed program AST is traversed from bottom to top until a predefined root node is reached. The types of the predefined root nodes are as follows: TypeDeclaration (type declaration), FieldDeclaration (field declaration), MethodDeclaration (method declaration), SwitchCase (branch statement), CatchClause (catch clause), ConstructorInvocation (constructor invocation), SuperConstructorInvocation (parent constructor invocation), or any statement node. For the reached predefined root node, the AST subtree is extracted between the root node and the leaf node mapped to the GumTree edit operation. For the extracted AST subtree, an ordered sequence is created and stored as a rich edit script. The rich edit script describes the operations performed on a given AST of the program before repair to convert the AST of the program before repair into the AST of the program after repair. Each node in the tree is an AST node affected by the changed code, and each node in the rich edit script contains three different types of information: shape, operation, and token.
[0043] Furthermore, in the repair template extraction module:
[0044] To address the problem of the combinatorial explosion of the comparison space caused by the direct pairwise comparison of rich edit scripts, a search index is constructed, that is, a set of comparison subspaces. A search index is defined for each iteration, and the shape index, operation index, and token index are obtained in sequence. The types of iterations include: shape, operation, and token. The construction method of the search index is to group the rich edit scripts according to standards, where these standards depend on the embedding information of the change tree representation used in different iterations.
[0045] The construction of the shape index takes the shape tree representation of the rich edit script as the input and groups it according to the AST node type based on the structure of the shape tree. The rich edit scripts with the same root node and the same depth are grouped into one group. For each group, a comparison space is created by enumerating the pairwise combinations of the group members. The shape index is constructed by storing the identifier of each group, represented as root node / depth.
[0046] The construction of the operation index follows the same principle as the shape index. The regrouping is based on the clustering output of the shape tree. The input is composed of the operation tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated based on the clustering of the shape tree.
[0047] The construction of the token index follows the same principle as the operation index. It regroupsthe clustering output based on the operation tree. The input is composed of the token tree representation of the rich edit script. The group identifier for each comparison space is generated as the root node / depth / id generated by the shape tree-based clustering / id generated by the operation tree-based clustering;
[0048] Calculate the tree edit distance of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree into another tree. For the first two iterations, namely shape and operation, use the edit script algorithm of GumTree. Take two trees as input and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same. Two identical trees are called an identical tree pair;
[0049] To extract the repair template, use the clusters of the special change tree representation of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representation according to the iteration. Use a clustering process based on the theory of connected component recognition in the graph to find a set of identical trees, that is, clusters; create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent associated trees, that is, identical tree pairs; in the graph, a cluster is defined as a subgraph. Each subgraph contains a set of trees that are identical to each other and unassociated. A cluster contains a set of rich edit scripts. The rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template.
[0050] Furthermore, in the concurrent defect repair module:
[0051] Use the open-source tool ANTLR to parse the file to be repaired into an AST. The spaces and comments in the source code are also parsed by ANTLR into corresponding types of nodes; the matching of the repair template is based on the AST context information of the defective code. Traverse each node of the AST of the suspicious statement in turn, from the first child node to the last leaf node of the AST of the suspicious statement, and try to match each node with the context AST of the repair template. If a node can match any of the given defect contexts, the corresponding repair template is successfully matched. If the node is not a leaf node, continue to traverse its leaf nodes; after successfully matching the repair template, perform insert, delete, update, and move operations on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of the concurrent defect.
[0052] Example 2
[0053] This example provides a template-based concurrent defect repair method, as Figure 1As shown below. The template-based concurrent defect repair method includes: Step 1, by crawling the defect reports of open-source projects in the open-source defect tracking system, retrieving and filtering all concurrency-related defects, and extracting the corresponding patch files; Step 2, converting the patch files into a special change tree representation; Step 3, extracting the repair template by comparing the similarity of the change trees; Step 4, using the repair template and the file to be repaired with known defect line numbers as input for template matching, and after successful matching, completing the repair of concurrent defects by refactoring the source code.
[0054] Further, in Step 1:
[0055] Crawl the defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "bug" (defect), the status being "resolved" (resolved), "fixed" (fixed), and "closed" (closed); use "concurrent" (concurrent), "concurrency" (concurrency), "deadlock" (deadlock), "atomicity violations" (atomicity violation), "data race" (data race), "race condition" (resource race), "order violations" (order violation), "thread-safety" (thread safety), "multi-thread" (multi-thread), and "synchronization" (synchronization) as keywords to retrieve the defect reports, filter all concurrency-related defects, and extract the corresponding patch files.
[0056] Further, in Step 2:
[0057] Use GumTree to calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file and remap it to the relevant nodes in the program AST; starting from the nodes of the GumTree edit operations, traverse the parsed program AST from bottom to top until reaching a predefined root node. The types of the predefined root nodes are as follows: TypeDeclaration (type declaration), FieldDeclaration (field declaration), MethodDeclaration (method declaration), SwitchCase (branch statement), CatchClause (catch clause), ConstructorInvocation (constructor invocation), SuperConstructorInvocation (parent constructor invocation), or any statement node; for the reached predefined root node, extract the AST subtree between the root node and the leaf node mapped to the GumTree edit operation; for the extracted AST subtree, create an ordered sequence and store the ordered sequence as a rich edit script; the rich edit script describes the operations performed on a given AST of the program before repair to convert the AST of the program before repair into the AST of the program after repair. Each node in the tree is an AST node affected by the changed code, and each node in the rich edit script contains three different types of information: shape, operation, and token.
[0058] Further, in step 3:
[0059] To address the problem of the combinatorial explosion of the comparison space caused by the direct pairwise comparison of rich edit scripts, construct search indexes, that is, a set of comparison subspaces. Define a search index for each iteration, and obtain the shape index, operation index, and token index in sequence; where the types of iterations include: shape, operation, and token; the construction method of the search index is to group the rich edit scripts according to standards, where these standards depend on the embedding information of the change tree representation used in different iterations;
[0060] The construction of the shape index takes the shape tree representation of the rich edit script as the input and groups it according to the AST node type based on the structure of the shape tree. Rich edit scripts with the same root node and the same depth are grouped into one group. For each group, create a comparison space by enumerating the pairwise combinations of the group members; the shape index is constructed by storing the identifier of each group, represented as root node / depth;
[0061] The construction of the operation index follows the same principle as the shape index. The regrouping is based on the clustering output of the shape tree, and the input is composed of the operation tree representation of the rich edit script. The group identifier of each comparison space is generated as root node / depth / id generated based on the clustering of the shape tree;
[0062] The construction of the token index follows the same principle as the operation index. It regroupsthe clustering output based on the operation tree. The input is composed of the token tree representation of the rich edit script. The group identifier for each comparison space is generated as the root node / depth / id generated by the shape tree-based clustering / id generated by the operation tree-based clustering;
[0063] Calculate the tree edit distance of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree to another. For the first two iterations, namely shape and operation, use the edit script algorithm of GumTree. Take two trees as input and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same, and two identical trees are called an identical tree pair;
[0064] To extract the repair template, use the clusters of the special change tree representation of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representation according to the iteration. Use the clustering process based on the theory of connected component recognition in the graph to find a set of identical trees, that is, clusters; create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent associated trees, that is, identical tree pairs; in the graph, define a cluster as a subgraph. Each subgraph contains a set of trees that are identical to each other and unassociated. A cluster contains a set of rich edit scripts, and the rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template.
[0065] Furthermore, in step 4:
[0066] Use the open-source tool ANTLR to parse the file to be repaired into an AST. The spaces and comment contents in the source code are also parsed by ANTLR into nodes of corresponding types; the matching of the repair template is based on the AST context information of the defective code. Traverse each node of the AST of the suspicious statement in turn, from the first child node to the last leaf node of the AST of the suspicious statement, and try to match each node with the context AST of the repair template. If a node can match any of the given defective contexts, the corresponding repair template is successfully matched. If the node is not a leaf node, continue to traverse its leaf nodes; after successfully matching the repair template, perform insert, delete, update, and move operations on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of concurrent defects.
[0067] There are many ways and means to specifically implement this technical solution in the present invention. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A template-based concurrent defect repair system, characterized in that Including: A data set collection module, which is used to crawl defect reports of open-source projects in an open-source defect tracking system, retrieve the defect reports using keywords such as "concurrent", "concurrency", "deadlock", "atomicity violation", "data race", "resource race", "order violation", "thread safety", "multi-threading", and "synchronization", filter all concurrency-related defects, and extract the corresponding patch files; A rich edit script calculation module, which is used to calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file using GumTree and remap it to the relevant nodes in the program AST; Each node in the rich edit script contains three different types of information: shape, operation, and token; A repair template extraction module, which is used to construct search indexes, that is, a set of comparison subspaces, define a search index for each iteration respectively, and obtain a shape index, an operation index, and a token index in sequence; where the types of iteration include: shape, operation, and token; Calculate the tree edit distance of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree to another tree. For the first two iterations, that is, shape and operation, use the edit script algorithm of GumTree, take two trees as input, and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same, and two identical trees are called identical tree pairs; To extract the repair template, use the cluster of the special change tree representation of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representation according to the iteration, and use the clustering process based on the theory of connected component recognition in the graph to find a set of identical trees, that is, a cluster; create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent associated trees, that is, identical tree pairs; in the graph, define a cluster as a subgraph. Each subgraph contains a set of trees that are identical to each other and not associated. A cluster contains a set of rich edit scripts, and the rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template; A concurrent defect repair module, which is used to perform template matching with the repair template and the file to be repaired with known defect line numbers as input, and complete the repair of concurrent defects by refactoring the source code after successful matching.
2. The template-based concurrent defect repair system according to claim 1, wherein: In the data set collection module: Crawl defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "defect" and the status being "resolved", "fixed", and "closed".
3. The template-based concurrent defect repair system according to claim 2, wherein: In the rich edit script calculation module: Starting from the nodes of the GumTree edit operation, traverse the parsed program AST from bottom to top until reaching a predefined root node. The types of the predefined root nodes are as follows: type declaration, field declaration, method declaration, branch statement, catch clause, constructor call, parent constructor call, or any statement node; For the reached predefined root node, extract the AST subtree between the root node and the leaf node mapped to the GumTree edit operation; For the extracted AST sub - trees, create an ordered sequence and store this ordered sequence as a rich edit script; the rich edit script describes the operations performed on a given pre - repair program AST to transform the pre - repair program AST into a post - repair program AST, and each node in the tree is an AST node affected by the changed code.
4. The template-based concurrent defect repair system according to claim 3, wherein: In the repair template extraction module: The construction method of the search index is to group the rich edit scripts according to standards, where these standards depend on the embedding information of the change tree representations used in different iterations; The construction of the shape index takes the shape tree representation of the rich edit script as input and groups it according to the AST node types based on the structure of the shape tree. Rich edit scripts with the same root node and the same depth are grouped into one group. For each group, a comparison space is created by enumerating the pairwise combinations of the group members; the shape index is constructed by storing the identifier of each group, represented as root node / depth; The construction of the operation index follows the same principle as the shape index, regrouping the clustering output based on the shape tree. The input is composed of the operation tree representation of the rich edit script, and the group identifier for each comparison space is generated as root node / depth / id generated by clustering based on the shape tree; The construction of the token index follows the same principle as the operation index, regrouping the clustering output based on the operation tree. The input is composed of the token tree representation of the rich edit script, and the group identifier for each comparison space is generated as root node / depth / id generated by clustering based on the shape tree / id generated by clustering based on the operation tree.
5. The template-based concurrent defect repair system according to claim 4, characterized in that: In the concurrent defect repair module: Use the open - source tool ANTLR to parse the file to be repaired into an AST, and the spaces and comment contents in the source code are also parsed by ANTLR into corresponding types of nodes; The matching of the repair template is based on the AST context information of the defective code. Traverse each node of the suspicious statement AST in sequence, from the first child node to the last leaf node of the suspicious statement AST, and try to match each node with the context AST of the repair template. If a node can match any of the given defective contexts, the corresponding repair template is successfully matched. If the node is not a leaf node, continue to traverse its leaf nodes; After successfully matching the repair template, perform insert, delete, update, and move operations on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of concurrent defects.
6. A template - based concurrent defect repair method, including: Step 1, by crawling the defect reports of open - source projects in the open - source defect tracking system, retrieve the defect reports using keywords such as "concurrent", "concurrency", "deadlock", "atomicity violation", "data race", "resource race", "order violation", "thread safety", "multi - thread", and "synchronization", filter all concurrency - related defects, and extract the corresponding patch files; Step 2, calculate the edit operation sequence of the abstract syntax tree (AST) of the patch file using GumTree and remap it to the relevant nodes in the program AST; each node in the rich edit script contains three different types of information: shape, operation, and token; Step 3, construct search indexes, that is, a set of comparison subspaces, define a search index for each iteration respectively, and obtain a shape index, an operation index, and a token index in sequence; among them, the types of iterations include: shape, operation, and token; Calculate the tree edit distances of the first two representations of the rich edit script respectively, that is, the edit operation sequence for converting one tree to another tree. For the first two iterations, that is, shape and operation, use the edit script algorithm of GumTree, take two trees as input, and generate an edit script. The size of the edit script represents the tree edit distance between the two trees. When the tree edit distance is zero, the two input trees are considered the same, and two identical trees are called an identical tree pair; To extract the repair template, use the clusters of the special change tree representation of the rich edit script. Starting from the identical tree pairs generated for each iteration, extract the corresponding change tree representation according to the iteration, and use the clustering process based on the theory of connected component recognition in the graph to find a set of identical trees, that is, clusters; create an undirected graph from the set of tree pairs. The nodes in the graph represent trees, and the edges represent the associated trees, that is, identical tree pairs; in the graph, define a cluster as a subgraph. Each subgraph contains a set of trees that are identical to each other and not associated with each other. A cluster contains a set of rich edit scripts, and the rich edit scripts in the same set share a common special tree representation. When a cluster has at least two members, the cluster is defined as a template; Step 4, use the repair template and the file to be repaired with known defect line numbers as input for template matching. After successful matching, complete the repair of concurrent defects by reconstructing the source code.
7. The template-based concurrent defect repair method according to claim 6, characterized in that: In Step 1: Crawl the defect reports of open-source projects in the open-source defect tracking system Bugzilla with the type being "defect", and the status being "resolved", "fixed", and "closed".
8. The template-based concurrent defect repair method according to claim 7, wherein: In Step 2: Starting from the nodes of the GumTree edit operation, traverse the parsed program AST from bottom to top until reaching a predefined root node. The types of the predefined root nodes are as follows: type declaration, field declaration, method declaration, branch statement, catch clause, constructor call, parent constructor call, or any statement node; For the reached predefined root node, extract the AST subtree between the root node and the leaf node mapped to the GumTree edit operation; For the extracted AST subtree, create an ordered sequence and store the ordered sequence as a rich edit script; the rich edit script describes the operations performed on a given AST of the program before repair to convert the AST of the program before repair into the AST of the program after repair. Each node in the tree is an AST node affected by the changed code.
9. The template-based concurrent defect repair method according to claim 8, wherein: In Step 3: The method for constructing the search index is to group the rich edit scripts according to standards, where these standards depend on the embedding information of the change tree representation used in different iterations; The construction of the shape index takes the shape tree representation of the rich edit script as input, groups it according to the AST node type based on the structure of the shape tree, and rich edit scripts with the same root node and the same depth are grouped into one group. For each group, a comparison space is created by enumerating the pairwise combinations of the group members; the shape index is constructed by storing the identifier of each group, represented as the root node / depth; The construction of the operation index follows the same principle as the shape index, regrouping the clustering output based on the shape tree. The input is composed of the operation tree representation of the rich edit script, and the group identifier of each comparison space is generated as the root node / depth / id generated by the clustering based on the shape tree; The construction of the token index follows the same principle as the operation index, regrouping the clustering output based on the operation tree. The input is composed of the token tree representation of the rich edit script, and the group identifier of each comparison space is generated as the root node / depth / id generated by the clustering based on the shape tree / id generated by the clustering based on the operation tree.
10. The template-based concurrent defect repair method according to claim 9, wherein: In step 4: The file to be repaired is parsed into an AST using the open-source tool ANTLR, and the spaces and comment contents in the source code are also parsed by ANTLR into corresponding types of nodes; The matching of the repair template is based on the AST context information of the defective code. Each node of the AST of the suspicious statement is traversed in sequence, from the first child node to the last leaf node of the AST of the suspicious statement, and each node is tried to be matched with the context AST of the repair template. If a node can match any of the given defective contexts, the corresponding repair template is successfully matched. If the node is not a leaf node, its leaf nodes are continued to be traversed; After successfully matching the repair template, insert, delete, update, and move operations are performed on the corresponding nodes on the AST of the defective code to reconstruct the source code and complete the repair of concurrent defects.
Citation Information
Patent Citations
A software defect repair template extraction method based on clustering analysis
CN109165155A
Method for realizing defect repair recommendation based on learning algorithm
CN110442514A