Defect code tracing method based on program dependency graph
By constructing abstract syntax tree and program dependency graph similarity calculation, the problem of misidentifying defective code in traditional methods is solved, and more accurate traceability of defective code is achieved.
Patent Information
- Application Number
- CN202510564204.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-01
AI Technical Summary
Traditional defect code traceability methods based on program dependency graphs are difficult to accurately identify changes that are not related to defect repair when processing large-scale code, and it is easy to misidentify form modifications as defective codes, resulting in waste of resources.
By building an abstract syntax tree, using the depth-first search algorithm to obtain the modified node collection, and based on the similarity calculation of the program dependency graph, defect code is identified, and defect code traceability is used to reflect logical semantic information.
It improves the accuracy of defect code identification, reduces the situation where formal modifications are identified as defect codes, and improves the accuracy of traceability.
Smart Images

Figure CN120407379A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of source code program analysis, and particularly relates to a method for tracing defective code based on a program dependence graph. Background Art
[0002] In the modern software development process, the existence of software defects seriously affects software quality and reliability, and tracing defective code aims to trace the source of software defect introduction, which is crucial for improving the software development process and enhancing software quality.
[0003] There are still many deficiencies in traditional methods for tracing defective code based on a program dependence graph. For example, when dealing with large-scale code, traditional methods for tracing defective code based on a program dependence graph are difficult to accurately identify changes that are irrelevant to defect repair (such as code style adjustment, irrelevant logic changes, etc.), and it is easy to identify changes that are irrelevant to defect repair as defective code, resulting in waste of resources caused by tracing irrelevant modifications. Summary of the Invention
[0004] The present disclosure provides a method for tracing defective code based on a program dependence graph, which can reduce the situation of identifying changes that are irrelevant to defect repair as defective code and improve the accuracy of defective code identification. The technical solution at least includes the following: In a first aspect, a method for tracing defective code based on a program dependence graph is provided, including: obtaining a first code file and a second code file, where the first code file and the second code file are code files of adjacent versions, at least one defect exists in the first code file, and at least one defect in the first code file is repaired in the second code file; obtaining a set of changed nodes based on the first code file and the second code file, where the set of changed nodes includes multiple arrays, each array is used to indicate a changed node, and the multiple arrays of the set of changed nodes include first-type arrays, and the changed nodes indicated by the first-type arrays exist in both the first code file and the second code file; obtaining a first program dependence graph and a second program dependence graph corresponding to the i-th first-type array based on the first code file and the second code file; calculating the similarity between the first program dependence graph and the second program dependence graph corresponding to each first-type array; and performing defective code tracing based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first-type array, where the defective code tracing includes determining the location of the defective code in the first code file and the commit version of each defective code line.
[0005] Optionally, obtaining the set of modified nodes based on the first code file and the second code file includes: constructing a first abstract syntax tree based on the first code file; constructing a second abstract syntax tree based on the second code file; traversing the first abstract syntax tree and the second abstract syntax tree using a depth-first search algorithm to construct a first hash map and a second hash map, where the first hash map is used to store each node in the first abstract syntax tree, and the second hash map is used to store each node in the second abstract syntax tree; traversing the first hash map and the second hash map, and storing the first modified node and the second modified node in the first type of array of the set of modified nodes, where the first modified node is a node in the first abstract syntax tree, the second modified node is a node in the second abstract syntax tree, the node positions of the first modified node and the second modified node in the first abstract syntax tree and the second abstract syntax tree are the same, and the value of the first modified node is different from the value of the second modified node.
[0006] Optionally, obtaining the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array based on the first code file and the second code file includes: constructing the first program dependence graph corresponding to the i-th third modified node based on the control dependence and data dependence of the i-th third modified node in the first code file, where the i-th third modified node is the modified node corresponding to the i-th first type of array; constructing the second program dependence graph corresponding to the i-th third modified node based on the control dependence and data dependence of the i-th third modified node in the second code file.
[0007] Optionally, calculating the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array includes: converting the first program dependence graph corresponding to the i-th first type of array into a first feature vector based on graph embedding technology; converting the second program dependence graph corresponding to the i-th first type of array into a second feature vector based on graph embedding technology; calculating the cosine similarity between the first feature vector and the second feature vector to obtain the similarity between the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array.
[0008] Optionally, performing defect code tracing based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array includes: in the case where the cosine similarity between the first feature vector and the second feature vector is less than the similarity threshold, determining the modified node corresponding to the i-th first type of array as a defect repair node; determining the defective code lines in each of the defect repair nodes to obtain a set of defective code lines, where the defective code lines are the code lines where the defective code in the first code file is located.
[0009] Optionally, defect code tracing based on the similarity between the first program dependency graph and the second program dependency graph corresponding to each first type of array further includes: obtaining code files of all versions corresponding to the first code file before the version of the first code file to obtain a code file set; based on the code file set, determining the commit version of each defective code line in the defective code line set.
[0010] In a second aspect, there is also provided a defect code tracing device based on a program dependency graph, including: a first obtaining module, configured to obtain a first code file and a second code file, where the first code file and the second code file are code files of adjacent versions, at least one defect exists in the first code file, and at least one defect in the first code file is fixed in the second code file; a second obtaining module, configured to obtain a set of changed nodes based on the first code file and the second code file, where the set of changed nodes includes multiple arrays, each array is used to indicate a changed node, and the multiple arrays in the set of changed nodes include first type of arrays, and the changed nodes indicated by the first type of arrays exist in both the first code file and the second code file; a third obtaining module, configured to obtain a first program dependency graph and a second program dependency graph corresponding to the i-th first type of array based on the first code file and the second code file; a similarity calculation module, configured to calculate the similarity between the first program dependency graph and the second program dependency graph corresponding to each first type of array; a defect code tracing module, configured to perform defect code tracing based on the similarity between the first program dependency graph and the second program dependency graph corresponding to each first type of array, and the defect code tracing includes determining the position of the defect code in the first code file and the commit version of each defective code line.
[0011] Optionally, the second obtaining module is further configured to construct a first abstract syntax tree based on the first code file; construct a second abstract syntax tree based on the second code file; traverse the first abstract syntax tree and the second abstract syntax tree using a depth-first search algorithm to construct a first hash map and a second hash map, where the first hash map is used to store each node in the first abstract syntax tree, and the second hash map is used to store each node in the second abstract syntax tree; traverse the first hash map and the second hash map, and store a first changed node and a second changed node into the first type of array in the set of changed nodes, where the first changed node is a node in the first abstract syntax tree, the second changed node is a node in the second abstract syntax tree, the node positions of the first changed node and the second changed node in the first abstract syntax tree and the second abstract syntax tree are the same, and the value of the first changed node is different from the value of the second changed node.
[0012] Optionally, the third acquisition module is further configured to construct a first program dependence graph corresponding to the i-th third change node based on the control dependence and data dependence of the i-th third change node in the first code file, where the i-th third change node is the change node corresponding to the i-th first type of array; construct a second program dependence graph corresponding to the i-th third change node based on the control dependence and data dependence of the i-th third change node in the second code file.
[0013] Optionally, the similarity calculation module is further configured to convert the first program dependence graph corresponding to the i-th first type of array into a first feature vector based on a graph embedding technique; convert the second program dependence graph corresponding to the i-th first type of array into a second feature vector based on a graph embedding technique; calculate the cosine similarity between the first feature vector and the second feature vector to obtain the similarity between the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array.
[0014] Optionally, the defective code tracing module is further configured to determine that the change node corresponding to the i-th first type of array is a defective repair node when the cosine similarity between the first feature vector and the second feature vector is less than a similarity threshold; determine the defective code lines in each of the defective repair nodes to obtain a set of defective code lines, where the defective code lines are the code lines where the defective codes in the first code file are located.
[0015] Optionally, the defective code tracing module is further configured to obtain all the code files corresponding to the first code file before the version of the first code file to obtain a set of code files; determine the commit versions of each defective code line in the set of defective code lines based on the set of code files.
[0016] In a third aspect, a computer device is further provided, including: a memory and a processor, where at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to execute the defective code tracing method based on a program dependence graph in the above embodiments.
[0017] In a fourth aspect, a computer-readable storage medium is further provided, where at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to execute the defective code tracing method based on a program dependence graph in the above embodiments.
[0018] In a fifth aspect, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method described in the first aspect is implemented.
[0019] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure at least include: In the defect code tracing method based on the program dependence graph in the related art, usually only the similarity between the code strings of different versions is simply compared, and then the defective code is determined based on the similarity between the strings. However, the similarity between strings can only reflect whether the literals are close, and it is easy to misidentify the formal modification as the defect repair code.
[0020] In the embodiments of the present disclosure, by obtaining the set of modified nodes, and obtaining the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array in the set of modified nodes, then calculating the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array, and finally tracing the defective code based on the similarity; wherein the modified nodes indicated by the first type of array exist in the first code file and the second code file. Since the program dependence graph can reflect the logical semantic information, the similarity between the program dependence graphs can reflect the similarity of the code in different versions in terms of logical semantics. Tracing the defective code based on the similarity in logical semantics can effectively reduce the situation where the formal modification is identified as the defective code and improve the accuracy of defect code tracing. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 Shows a flowchart of a defect code tracing method based on a program dependence graph provided by an exemplary embodiment of the present disclosure; Figure 2 Shows a flowchart of a defect code tracing method based on a program dependence graph provided by another exemplary embodiment of the present disclosure; Figure 3 Shows a schematic structural diagram of a defect code tracing device based on a program dependence graph provided by an exemplary embodiment of the present disclosure; Figure 4 Is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Unless otherwise defined, technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second", "third" and similar terms used in the specification and claims of this patent application of the disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a limitation of quantity, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or items appearing before "comprising" or "including" cover the elements or items listed after "comprising" or "including" and their equivalents, and do not exclude other elements or items.
[0024] To make the objectives, technical solutions and advantages of the disclosure more clear, the embodiments of the disclosure will be further described in detail below with reference to the accompanying drawings.
[0025] Figure 1 The flowchart of a defect code tracing method based on a program dependence graph provided by an exemplary embodiment of the disclosure is shown, and this method can be executed by a computer device. Refer to Figure 1 This method includes: In step 101, a first code file and a second code file are obtained.
[0026] The first code file and the second code file are code files of adjacent versions. There is at least one defect in the first code file, and at least one defect in the first code file is fixed in the second code file.
[0027] That is to say, the second code file is the code file of the next version of the first code file.
[0028] Exemplarily, a code version control tool (such as git) is used to extract the version information of the first code file and the version information of the second code file. According to the version information of the first code file and the version information of the second code file, the first code file and the second code file can be obtained by using the code version control tool.
[0029] In step 102, based on the first code file and the second code file, a set of changed nodes is obtained.
[0030] The set of changed nodes includes multiple arrays, each array is used to indicate a changed node, and the multiple arrays of the set of changed nodes include a first type of array, and the changed nodes indicated by the first type of array exist in both the first code file and the second code file.
[0031] Optionally, the set of changed nodes is obtained through an abstract syntax tree. This method includes the following four steps.
[0032] The first step is to construct a first abstract syntax tree based on the first code file.
[0033] The second step is to construct a second abstract syntax tree based on the second code file.
[0034] Exemplarily, after obtaining the first code file and the second code file, a syntax parsing tool can be used to implement the first step and the second step.
[0035] The first abstract syntax tree and the second abstract syntax tree include multiple nodes. The nodes of the abstract syntax tree represent different compilation units, and the edges of the abstract syntax tree represent the hierarchical relationships between different units. The node name of any node in the abstract syntax tree is determined based on the hierarchical information of the compilation unit corresponding to the node, and the node name is used to uniquely indicate the node. The node value of any node in the abstract syntax tree is the complete code text content of the compilation unit corresponding to the node. The abstract syntax tree can represent the complete hierarchical structure of the code text. For example, the first abstract syntax tree can represent the complete hierarchical structure of the first code file, and the second abstract syntax tree can represent the complete hierarchical structure of the second code file.
[0036] The third step is to traverse the first abstract syntax tree and the second abstract syntax tree using the depth-first search algorithm to construct a first hash map and a second hash map.
[0037] The first hash map BugASTHash is used to store each node in the first abstract syntax tree, and the second hash map FixASTHash is used to store each node in the second abstract syntax tree.
[0038] The depth-first search algorithm (DFS) is an algorithm used to traverse or search a tree or a graph. The core idea of DFS is to start from the starting node and visit the nodes as deeply as possible along a path until it is no longer possible to continue, and then backtrack and try other paths. There are many implementation methods of DFS in related technologies, which are omitted here for detailed description.
[0039] Through the third step, each node in the first abstract syntax tree can be stored in the first hash map, and each node in the second abstract syntax tree can be stored in the second hash map, which is convenient for storing the modified nodes in the modified node set later.
[0040] The fourth step is to traverse the first hash map and the second hash map, and store the first modified node and the second modified node in the first type of array of the modified node set.
[0041] The first modified node is a node in the first abstract syntax tree, and the second modified node is a node in the second abstract syntax tree. The positions of the first modified node and the second modified node in the first abstract syntax tree and the second abstract syntax tree are the same, and the value of the first modified node is different from the value of the second modified node.
[0042] For the modified nodes in the first abstract syntax tree and the second abstract syntax tree, there are actually the following three cases.
[0043] In the first case, a certain node exists in the first abstract syntax tree, and this node also exists in the second abstract syntax tree, and the value of this node in the first abstract syntax tree is different from the value in the second abstract syntax tree.
[0044] In the second case, a certain node exists in the first abstract syntax tree, but this node does not exist in the second abstract syntax tree. The first case indicates that this node has been deleted, and the deleted node can be directly defined as a defective node without further processing.
[0045] In the third case, a certain node does not exist in the first abstract syntax tree, but this node exists in the second abstract syntax tree. The third case indicates that this node is a newly added node. Such a newly added node cannot be traced for defects. Therefore, the nodes in the third case belong to the modified nodes, but it cannot be determined whether the nodes in the third case belong to the defective nodes.
[0046] The defect code tracing in the embodiments of the present disclosure mainly analyzes the first case. Taking the fourth step as an example, in the process of traversing the first hash map and the second hash map, if the first modified node exists in the first abstract syntax tree, the second modified node exists in the second abstract syntax tree, the values of the first modified node and the second modified node are different, and the positions of the first modified node and the second modified node in the first abstract syntax tree and the second abstract syntax tree are the same, it means that the first modified node and the second modified node meet the first case, that is, the first modified node and the second modified node are actually the states before and after the modification of the same node (the first modified node is the state before the modification, and the second modified node is the state after the modification). At this time, the first modified node and the second modified node can be stored in a certain first type of array in the modified node set. This first type of array is used to indicate the modified node and includes the states before and after the modification of the modified node (that is, it includes the first modified node and the second modified node).
[0047] In addition, the modified node set also includes a second type of array and a third type of array. The modified nodes in the second case can be stored in the second type of array, and the modified nodes in the third case can be stored in the third type of array.
[0048] When implemented, the array in the modified node set can be represented in the following form:
[0049] Among them, represents the j-th array in the modified node set, and the j-th array includes the names of the modified nodes in the first code file and values , the names of the modified nodes in the second code file and values as well as the status of the j-th array . The of the j-th array is used to distinguish the first type of array, the second type of array, and the third type of array. Among them, j is a positive integer and j > 0.
[0050] For example, when the j-th array is the first type of array, the of the j-th array is change. At this time, the j-th array is used to indicate the status before and after the modification of the same modified node, and the value is used to store the first modified node of the modified node (that is, the status of the modified node before the modification), and the value is used to store the second modified node of the modified node (that is, the status of the modified node after the modification).
[0051] When traversing the first hash map and the second hash map, each node in the first hash map is traversed first, and it is checked whether any node in the first hash map exists in the second hash map.
[0052] In this case, for the above three cases, the following three rules can be set to facilitate storing different types of modified nodes into the first type of array, the second type of array, and the third type of array respectively.
[0053] Rule 1: If a certain node exists in the first hash map and also exists in the second hash map, then further compare whether the value of this node in the first hash map is the same as the value of this node in the second hash map. If the value of this node in the first hash map is the same as the value of this node in the second hash map, then this node belongs to the unmodified node and does not need to be stored in the form of an array. If the value of this node in the first hash map is different from the value of this node in the second hash map, then this node is considered to be in the "modified" state, and it is necessary to further determine whether this node belongs to the defective node in the future.
[0054] The nodes that meet Rule 1 need to be stored in the first type of array. When storing the nodes that meet Rule 1 in the first type of array, first, the status of the array Assign it as "change", representing the first type of array; then the name of the first modified node corresponding to this modified node is required. And the value Also, the name of the second modified node corresponding to this modified node needs to be stored. And the value That is, both the first modified node and the second modified node are stored in the first type of array. Obviously, in the first type of array, And Are the same.
[0055] Rule 2: If a certain node exists in the first hash map and does not exist in the second hash map, then this node in the first abstract syntax tree is considered to be in the "deleted" state.
[0056] The nodes that meet Rule 2 need to be stored in the second type of array. When storing the nodes that meet Rule 2 in the second type of array, first the state of the array Is assigned as "delete", representing the second type of array; then, since this node only exists in the first code file, only the name of this node needs to be stored in And the value of this node is stored in Since the modified nodes in the second type of array do not exist in the second code file, the And In the second type of array can be stored as empty (Null).
[0057] Rule 3: If a certain node does not exist in the first hash map and exists in the second hash map, then this node in the second abstract syntax tree is considered to be in the "added" state.
[0058] The nodes that meet Rule 3 need to be stored in the third type of array. When storing the nodes that meet Rule 3 in the third type of array, first the state of the array Is assigned as "add", representing the third type of array; then, since this node only exists in the second code file, only the name of this node needs to be stored in And the value of this node is stored in Since the modified nodes in the third type of array do not exist in the first code file, the And In the third type of array can be stored as empty (Null).
[0059] In the fourth step, after traversing the nodes in the first hash map and the second hash map using Rules 1 to 3, multiple arrays can be obtained, and these multiple arrays are the set of modified nodes. Exemplarily, the set of modified nodes can be represented in the following form.
[0060]
[0061] Among them, represents the set of modified nodes, represents the j-th array in the set of modified nodes, n is the number of arrays in the set of modified nodes, and the value range of j is from 1 to n, where n is a positive integer.
[0062] In the embodiments of the present disclosure, the method further includes: removing duplicates from the arrays in the set of modified nodes. Since there is an inclusion relationship between the parent and child nodes in the abstract syntax tree, there may be duplicate arrays with an inclusion relationship in the hierarchy in the set of modified nodes; in addition, during the process of code version change, there may also be situations where only variable names (such as function names, class names, variable names, etc.) or modifiers (such as access modifiers and non-access modifiers, etc.) are modified, without changing the content of the executed code block. This situation will result in two nodes having the same executed code block but different definition lines. Since the first to fourth steps above first compare the names of the nodes (the identification of the name is achieved through the definition line) and then compare the values of the nodes, the first to fourth steps cannot identify the situation where the names of the nodes are different but the executed code blocks of the nodes are the same. Therefore, the arrays in the set of modified nodes need to be de-duplicated before further processing.
[0063] Exemplarily, the following two steps are used for de-duplication.
[0064] The first step is to remove duplicates from the arrays in the set of modified nodes that are not Null.
[0065] Optionally, the first step includes: for the arrays in the set of modified nodes that are not Null, check whether the corresponding to the array is a sub-string of the in other arrays. Here, can reflect the hierarchical information of the nodes, so the hierarchical relationship of the arrays can be screened by comparing .
[0066] If the corresponding to the array is a sub-string of the in other arrays, it means that the modified node is the parent node of other modified nodes. At this time, the array where the modified node is located can be removed.
[0067] If the name of the modified node corresponding to the array is not a sub-string of the name of other modified nodes, it means that the modified node is an independent modified node and there is no inclusion relationship with other modified nodes. At this time, the array can be retained.
[0068] In the second step, duplicate elimination is performed on the arrays in the modified node set that have at least one array with the name Null.
[0069] The situations targeted in the second step include being Null, the situation of being Null.
[0070] When performing the second step, first, the non-Null in all arrays of the two situations targeted in the second step are stored in the cache after removing the declaration and definition lines. Here, when the of a certain array is Null, the and of this array must be non-Null; similarly, when the of a certain array is Null, the and of this array must be non-Null. Therefore, multiple non-Null can be obtained, and then these non-Null are stored in the cache after removing the declaration and definition lines, obtaining multiple elements.
[0071] Then, the multiple elements in the cache can be merged to achieve duplicate elimination. When merging, if a certain element in the cache is equal to the of array A, and this element is equal to the of array B, then array B is merged with array A (the and of array B are merged into array A), and the status in array A is modified to change, and array B is deleted.
[0072] In this way, it can be ensured that there are no "added" or "deleted" error arrays generated in the modified node set in the case where the node names are different but the execution code blocks of the nodes are the same.
[0073] In step 103, based on the first code file and the second code file, the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array are obtained.
[0074] Here, the modified nodes corresponding to each first type of array all belong to the third type of modified nodes, and the modified node corresponding to the i-th first type of array is the i-th third type of modified node.
[0075] In this case, the first program dependence graph corresponding to the i-th first type of array is the program dependence graph obtained based on the i-th third type of modified node in the first code file, and the second program dependence graph corresponding to the i-th first type of array is the program dependence graph obtained based on the i-th third type of modified node in the second code file.
[0076] In step 104, calculate the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array.
[0077] In step 105, trace the defective code based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array.
[0078] Tracing the defective code includes determining the location of the defective code in the first code file and the committed version of each defective code line.
[0079] In the related art, the method for tracing defective code based on program dependence graphs usually only simply compares the similarity between code strings of different versions, and then determines the defective code based on the similarity between the strings. However, the similarity between strings can only reflect whether the literals are close, and it is easy to misidentify a formal modification as a defective code repair.
[0080] In the embodiments of the present disclosure, by obtaining a set of changed nodes, and obtaining the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array in the set of changed nodes, then calculating the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array, and finally tracing the defective code based on the similarity; where the changed nodes indicated by the first type of array exist in the first code file and the second code file. Since the program dependence graph can reflect logical semantic information, the similarity between program dependence graphs can reflect the similarity in logical semantics of different versions of code. Tracing the defective code based on the similarity in logical semantics can effectively reduce the situation where a formal modification is identified as a defective code, and improve the accuracy of tracing the defective code.
[0081] Figure 2 The flowchart of the method for tracing defective code based on program dependence graphs provided by another exemplary embodiment of the present disclosure is shown, and this method can be executed by a computer device. Refer to Figure 2 and this method includes: In step 201, obtain the first code file and the second code file.
[0082] In step 202, based on the first code file and the second code file, obtain a set of changed nodes.
[0083] The set of changed nodes includes multiple arrays, each array is used to indicate a changed node, and the multiple arrays in the set of changed nodes include first type of arrays, and the changed nodes indicated by the first type of arrays exist in the first code file and the second code file.
[0084] For the relevant content of steps 201 to 202, refer to the foregoing steps 101 to 102, and the detailed description is omitted here.
[0085] It should be noted that when performing the subsequent steps 203 to 205, the processed node set after deduplication using the method in step 102 is processed.
[0086] In step 203, based on the first code file and the second code file, the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array are obtained.
[0087] Optionally, step 203 includes: based on the control dependence and data dependence of the i-th third change node in the first code file, constructing the first program dependence graph corresponding to the i-th third change node, where the i-th third change node is the change node corresponding to the i-th first type of array. Based on the control dependence and data dependence of the i-th third change node in the second code file, constructing the second program dependence graph corresponding to the i-th third change node.
[0088] Before constructing the first program dependence graph corresponding to the i-th third change node, a directed graph structure can be initialized through a graph library first. In this directed graph structure, the nodes of the directed graph structure represent variables, methods, control conditions, etc. in the code, and the edges of the directed graph structure represent the dependence relationships between the nodes. This initialized directed graph structure can provide a basic graph structure for the subsequent analysis of control dependence and data dependence.
[0089] Then, the control dependence of the i-th third change node in the first code file can be obtained. For example, the control flow statements (such as if, while, for, switch, etc.) in the compilation unit of the i-th third change node in the first code file can be extracted through regular expressions, the control conditions in the control flow statements are added as nodes to the directed graph, the dependent variables in the control conditions are obtained, and the dependent variables are added to a node in the directed graph structure. Then, an edge is added between the two nodes of the dependent variable and the control condition to represent the control dependence.
[0090] Similarly, the data dependence of the i-th third change node in the first code file can also be obtained. For example, the variable assignments, method calls, and expressions (such as =, +, -, *, / , etc.) in the compilation unit of the i-th third change node in the first code file can be extracted through regular expressions, the variable assignments, method calls, and expressions are added as nodes to the directed graph structure, and the dependence relationships between the variables are analyzed, and an edge is added from the dependent variable to the target variable or method call to represent the data dependence.
[0091] After the control dependencies and data dependencies of the i-th third modified node in the first code file are both extracted into a directed graph structure, it is also necessary to merge the directed graph structure to finally obtain the first program dependence graph corresponding to the i-th third modified node. Exemplarily, the control dependence nodes and edges and the data dependence nodes and edges can be merged into the same directed graph with code statements or basic blocks as the link, and the merged directed graph is the program dependence graph.
[0092] The second program dependence graph corresponding to the i-th third modified node can also be obtained in the same way as the first program dependence graph corresponding to the i-th third modified node.
[0093]
[0094] Among them, is the set of program dependence graphs, represents the two program dependence graphs corresponding to the i-th third modified node, where represents the first program dependence graph corresponding to the i-th third modified node, represents the second program dependence graph corresponding to the third modified node. m is the total number of the first type of arrays, and the value range of i is from 1 to m.
[0095] In step 204, calculate the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array.
[0096] Optionally, step 204 includes the following steps a-c.
[0097] Step a, based on the graph embedding technology, convert the first program dependence graph corresponding to the i-th first type of array into a first feature vector.
[0098] Step b, based on the graph embedding technology, convert the second program dependence graph corresponding to the i-th first type of array into a second feature vector.
[0099] Through the graph embedding technology, each program dependence graph can be embedded into a global feature vector. Here, the first program dependence graph corresponding to the i-th first type of array is also the first program dependence graph corresponding to the i-th third modified node, denoted by ; the second program dependence graph corresponding to the i-th first type of array is also the second program dependence graph corresponding to the i-th third modified node, denoted by . In this case, the first feature vector can be expressed as: . Among them, represents the first feature vector, represents the d-th component of the first feature vector.
[0100] Similarly, the second feature vector can be expressed as: .in, represents the second eigenvector, represents the dth component of the second eigenvector.
[0101] Step c: Calculate the cosine similarity between the first eigenvector and the second eigenvector.
[0102] Optionally, the cosine similarity between the first eigenvector and the second eigenvector is calculated using the following formula:
[0103] in, Represents the first program dependency graph corresponding to the i-th first-class array The second program dependency graph corresponding to the i-th first-class array The similarity between represents the first eigenvector With the second eigenvector The cosine similarity between .
[0104] Through the above steps ac, the similarity between the first program dependency graph and the second program dependency graph corresponding to each first-type array can be calculated.
[0105] In step 205 , defect code tracing is performed based on the similarity between the first program dependency graph and the second program dependency graph corresponding to each first-category array.
[0106] Defective code tracing includes determining the location of the defective code in the first code file and the committed version of each defective code line.
[0107] Optionally, step 205 includes the following steps de.
[0108] Step d: When the cosine similarity between the first eigenvector and the second eigenvector is less than a similarity threshold, determining the modified node corresponding to the i-th first-category array as a defect repair node.
[0109] Corresponding to step d, when the cosine similarity between the first eigenvector and the second eigenvector is greater than or equal to the similarity threshold, it is determined that the modified node corresponding to the i-th first-category array is not a defect repair node.
[0110] Illustratively, the similarity threshold value ranges from 0.5 to 0.8, for example, it may be 0.5, 0.6, 0.7 or 0.8.
[0111] Step e: Determine the defective code lines in each defect repair node to obtain a defective code line set.
[0112] The defective code line is a code line where the defective code is located in the first code file.
[0113] For each of the third modified nodes, the above-mentioned first step can be executed, so that a plurality of defect repair nodes can be determined. These plurality of defect repair nodes form a defect repair node set, and this defect repair node set can be represented by the following set:
[0114] Among them, is the defect repair node set, represents the k-th defect repair node in the defect repair node set. Where k>0 and k is a positive integer.
[0115] Exemplarily, the defect code line corresponding to the k-th defect repair node in the defect repair node set is determined by the following three steps.
[0116] First step, obtain the and in the array where the k-th defect repair node is located.
[0117] Second step, generate a first code list based on and generate a second code list based on .
[0118] The first code list includes all the code lines in , and the second code list includes all the code lines in .
[0119] Exemplarily, the first code list is represented by the following set:
[0120] Among them, is the first code list, is the p-th line of code in the first code list, p is an integer, p>0. Optionally, the first code list may also include the in the array where the k-th defect repair node is located.
[0121] The second code list is represented by the following set:
[0122] Among them, is the second code list, is the q-th line of code in the first code list, q is an integer, q>0. Optionally, the second code list may also include the in the array where the k-th defect repair node is located.
[0123] In the third step, a line-by-line comparison algorithm is used to compare the first code list and the second code list to determine the defective code line corresponding to the k-th defect repair node.
[0124] Exemplarily, the line-by-line comparison algorithm is the longest common subsequence algorithm. There are many implementation methods of the longest common subsequence algorithm in the related art, which are omitted here for detailed description.
[0125] When using the line-by-line comparison algorithm to compare the first code list and the second code list, the following three situations also exist.
[0126] The first situation is that a certain line of code exists in the first code list but does not exist in the second code list.
[0127] The second situation is that a certain line of code does not exist in the first code list but exists in the second code list.
[0128] The third situation is that a certain line of code exists in the first code list and also exists in the second code list, and the relative positions of this line of code in the first code list and the second code list are different.
[0129] The first situation indicates that this line of code has been modified or deleted, that is, this line of code belongs to the defective code line; the third situation indicates that this line of code has been moved, that is, this line of code belongs to the defective code line. The second situation cannot determine whether it belongs to the defective code line, and only indicates that this line of code is newly added, so it is not processed.
[0130] Through the above first to third steps, the defective code lines in the k-th defect repair node can be determined. For defect repair nodes other than the k-th defect repair node, the above first to third steps can also be used to determine the defective code lines.
[0131] In addition, each line of code in
[0132] is a defective code line.
[0133] Finally, the defective code line set can be determined, and the defective code line set includes all defective code lines.
[0134] After determining the defective code line set, defect code tracing also includes: determining the committed version of each defective code line. This method includes the following two steps.
[0135] In the first step, obtain all code files corresponding to the first code file before the version of the first code file to obtain a code file set.
[0136] The first step can be implemented by a code version management tool.In the second step, based on the set of code files, determine the committed version of each defective code line in the set of defective code lines.
[0137] When implementing the second step, then match the first defective code line with the code files of each version in the set of code files one by one. The version corresponding to the code file of the earliest version where the first defective code line is located is the committed version of the first defective code line. Among them, the first defective code line is any defective code line in the set of defective code lines.
[0138] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference may be made to the above method embodiment.
[0139] Figure 3 The structural schematic diagram of a defective code traceability apparatus based on a program dependence graph provided by an exemplary embodiment of the present disclosure is shown. Refer to Figure 3 , the defective code traceability apparatus 300 based on the program dependence graph includes: a first acquisition module 301, a second acquisition module 302, a third acquisition module 303, a similarity calculation module 304, and a defective code traceability module 305.
[0140] The first acquisition module 301 is used to acquire a first code file and a second code file. The first code file and the second code file are code files of adjacent versions. There is at least one defect in the first code file, and at least one defect in the first code file is fixed in the second code file.
[0141] The second acquisition module 302 is used to acquire a set of changed nodes based on the first code file and the second code file. The set of changed nodes includes multiple arrays, and each array is used to indicate a changed node. The multiple arrays of the set of changed nodes include a first type of array, and the changed nodes indicated by the first type of array exist in both the first code file and the second code file.
[0142] The third acquisition module 303 is used to acquire a first program dependence graph and a second program dependence graph corresponding to the i-th first type of array based on the first code file and the second code file.
[0143] The similarity calculation module 304 is used to perform defective code traceability based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array. The defective code traceability includes determining the position of the defective code in the first code file and the committed version of each defective code line.
[0144] The defective code traceability module 305 is used to perform defective code traceability based on the cosine similarity.
[0145] Optionally, the second acquisition module 302 is further configured to construct a first abstract syntax tree based on the first code file; construct a second abstract syntax tree based on the second code file; traverse the first abstract syntax tree and the second abstract syntax tree by using a depth-first search algorithm to construct a first hash map and a second hash map, where the first hash map is used to store each node in the first abstract syntax tree, and the second hash map is used to store each node in the second abstract syntax tree; traverse the first hash map and the second hash map, and store the first changed node and the second changed node into a first type of array in the changed node set, where the first changed node is a node in the first abstract syntax tree, the second changed node is a node in the second abstract syntax tree, the node positions of the first changed node and the second changed node in the first abstract syntax tree and the second abstract syntax tree are the same, and the value of the first changed node is different from the value of the second changed node.
[0146] Optionally, the third acquisition module 303 is further configured to construct a first program dependence graph corresponding to the i-th third changed node based on the control dependence and data dependence of the i-th third changed node in the first code file, where the i-th third changed node is the changed node corresponding to the i-th first type of array; construct a second program dependence graph corresponding to the i-th third changed node based on the control dependence and data dependence of the i-th third changed node in the second code file.
[0147] Optionally, the similarity calculation module 304 is further configured to convert the first program dependence graph corresponding to the i-th first type of array into a first feature vector based on a graph embedding technique; convert the second program dependence graph corresponding to the i-th first type of array into a second feature vector based on the graph embedding technique; calculate the cosine similarity between the first feature vector and the second feature vector to obtain the similarity between the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array.
[0148] Optionally, the defective code tracing module 305 is further configured to, when the cosine similarity between the first feature vector and the second feature vector is less than a similarity threshold, determine the changed node corresponding to the i-th first type of array as a defective repair node; determine the defective code lines in each of the defective repair nodes to obtain a defective code line set, where the defective code line is the code line where the defective code in the first code file is located.
[0149] Optionally, the defective code tracing module 305 is further configured to obtain all the code files corresponding to the first code file before the version of the first code file to obtain a code file set; determine the commit version of each defective code line in the defective code line set based on the code file set.
[0150] It should be noted that when the defect code tracing device based on the program dependence graph provided in the above embodiments performs defect code tracing, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the defect code tracing device based on the program dependence graph provided in the above embodiments and the embodiments of the defect code tracing method based on the program dependence graph belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0151] The division of modules in the embodiments of the present disclosure is illustrative, merely a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present disclosure, the functional modules can be integrated in a processor, can exist separately physically, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0152] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a terminal device (which can be a personal computer, a mobile phone, or a communication device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0153] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. As Figure 4 shown, the computer device 400 includes: a processor 401 and a memory 402.
[0154] The processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0155] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the defect code tracing method based on the program dependency graph provided in the embodiments of the present disclosure.
[0156] Those skilled in the art can understand that Figure 4 the structure shown in
[0157] does not constitute a limitation on the computer device 400, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0158] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the computer device, the computer device can execute the defect code tracing method based on the program dependency graph provided in the embodiments of the present disclosure.
[0159] The above are only optional embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A defect code tracing method based on a program dependence graph, characterized in that The method includes: Obtain a first code file and a second code file, where the first code file and the second code file are code files of adjacent versions, there are at least one defect in the first code file, and at least one defect in the first code file is fixed in the second code file; Based on the first code file and the second code file, obtain a set of changed nodes, the set of changed nodes includes multiple arrays, each array is used to indicate a changed node, and the multiple arrays of the set of changed nodes include a first type of array, and the changed nodes indicated by the first type of array exist in the first code file and the second code file; Based on the first code file and the second code file, obtain a first program dependence graph and a second program dependence graph corresponding to the i-th first type of array; Calculate the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array; Perform defect code traceability based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array, and the defect code traceability includes determining the position of the defect code in the first code file and the committed version of each defect code line.
2. The method according to claim 1, wherein The obtaining of the set of changed nodes based on the first code file and the second code file includes: Construct a first abstract syntax tree based on the first code file; Construct a second abstract syntax tree based on the second code file; Traverse the first abstract syntax tree and the second abstract syntax tree using a depth-first search algorithm to construct a first hash map and a second hash map, where the first hash map is used to store each node in the first abstract syntax tree, and the second hash map is used to store each node in the second abstract syntax tree; Traverse the first hash map and the second hash map, and store a first changed node and a second changed node into the first type of array of the set of changed nodes, where the first changed node is a node in the first abstract syntax tree, the second changed node is a node in the second abstract syntax tree, the node positions of the first changed node and the second changed node in the first abstract syntax tree and the second abstract syntax tree are the same, and the value of the first changed node is different from the value of the second changed node.
3. The method according to claim 1, characterized in that, The obtaining of the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array based on the first code file and the second code file includes: Construct a first program dependence graph corresponding to the i-th third changed node based on the control dependence and data dependence of the i-th third changed node in the first code file, where the i-th third changed node is the changed node corresponding to the i-th first type of array; Construct a second program dependence graph corresponding to the i-th third changed node based on the control dependence and data dependence of the i-th third changed node in the second code file.
4. The method according to any one of claims 1 to 3, characterized in that The calculating of the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array includes: Based on graph embedding technology, convert the first program dependence graph corresponding to the i-th first type of array into a first feature vector; Based on graph embedding technology, convert the second program dependence graph corresponding to the i-th first type of array into a second feature vector; Calculate the cosine similarity between the first feature vector and the second feature vector to obtain the similarity between the first program dependence graph and the second program dependence graph corresponding to the i-th first type of array.
5. The method according to claim 4, wherein The defect code traceability based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array includes: When the cosine similarity between the first feature vector and the second feature vector is less than the similarity threshold, determine that the modified node corresponding to the i-th first type of array is a defect repair node; Determine the defective code lines in each of the defect repair nodes to obtain a set of defective code lines, where the defective code lines are the code lines where the defective codes in the first code file are located.
6. The method according to claim 5, characterized in that, The defect code traceability based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array further includes: Obtain all the code files corresponding to the first code file before the version of the first code file to obtain a set of code files; Based on the set of code files, determine the commit version of each defective code line in the set of defective code lines.
7. A defect code tracing device based on a program dependence graph, characterized in that, The device includes: A first acquisition module, configured to acquire a first code file and a second code file, where the first code file and the second code file are code files of adjacent versions, there is at least one defect in the first code file, and at least one defect in the first code file is repaired in the second code file; A second acquisition module, configured to acquire a set of modified nodes based on the first code file and the second code file, where the set of modified nodes includes multiple arrays, each array is used to indicate a modified node, and the multiple arrays in the set of modified nodes include first type of arrays, and the modified nodes indicated by the first type of arrays exist in the first code file and the second code file; A third acquisition module, configured to acquire a first program dependence graph and a second program dependence graph corresponding to the i-th first type of array based on the first code file and the second code file; A similarity calculation module, configured to calculate the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array; A defect code traceability module, configured to perform defect code traceability based on the similarity between the first program dependence graph and the second program dependence graph corresponding to each first type of array, where the defect code traceability includes determining the location of the defect code in the first code file and the commit version of each defective code line.
8. A computer device, characterized in that, The computer device includes: a memory and a processor, where at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.