A defect location method and system based on code knowledge graph
By building a code knowledge graph and using BiLSTM+CRF model for naming entity recognition, the problem of defect reporting and source code mismatch in traditional defect positioning methods is solved, and more efficient and accurate defect positioning is achieved.
Patent Information
- Application Number
- CN202211190016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-09-28
AI Technical Summary
The traditional defect positioning method based on information retrieval has the problem that the defect report text does not match the word text of the source code file, resulting in low positioning efficiency and poor accuracy.
By building a code knowledge graph, the source code structure information and relationship information are used for defect positioning, and the named entity recognition is combined with the BiLSTM+CRF model to filter redundant information and improve positioning accuracy.
It significantly improves the efficiency and accuracy of defect positioning and reduces the time and energy costs of code maintenance personnel.
Smart Images

Figure CN115629760B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software maintenance, and in particular to a defect location method and system based on a code knowledge graph. Background Art
[0002] As software systems become larger and more complex, it is difficult to identify every software defect before formal release due to limited software testing resources. Therefore, released software systems often contain defects.
[0003] In order to efficiently identify and fix defects in released software systems, defect reports are documents written in natural language that describe problems where the software fails to run as expected or does not meet the technical requirements of the system. They are a collection of descriptions of software defect phenomena and reproduction steps.
[0004] The existence of defects in software systems may cause serious losses. Therefore, more than one-third of the software development-related costs are used to locate and repair defects. Automatic defect location can significantly reduce the time spent by development testers and reduce the cost of software construction and maintenance. At present, some researchers use information retrieval technology (IR) to locate defects, treating defect reports as queries and source files as documents. By calculating the similarity between queries and documents, source files that may contain defects are sorted. Traditional defect location based on information retrieval has the problem that the defect report text does not match the word text of the source code file. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a defect location method and system based on code knowledge graph, so as to improve the effect of traditional defect location and enhance the efficiency of defect location.
[0006] Technical solution: The present invention provides a defect location method based on code knowledge graph, comprising the following steps:
[0007] Step 1: Get the source code of AspectJ, SWT, and Zxing from the Git version control system;
[0008] Step 2: Parse the source code through the code parser to generate the abstract syntax tree AST of the source code;
[0009] Step 3: Extract entities and relationships from the abstract syntax tree AST to build a code knowledge graph;
[0010] Step 4: Crawl defect reports from the open source Bugzilla defect tracking system and obtain the summary and description of the defect report;
[0011] Step 5: Use the NLTK toolkit to preprocess the defect report summary and description to obtain the defect report dataset;
[0012] Step 6: Perform named entity recognition on the defect report dataset to extract the defect report entity sequence;
[0013] Step 7: Vectorize the defect report entity sequence and the code knowledge graph using Word2Vec and knowledge graph embedding algorithms;
[0014] Step 8: Map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph, arrange the similarities from high to low, and generate a list of suspicious methods.
[0015] Furthermore, in step 2, the Spoon code parser is used to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables and comments are recorded.
[0016] Furthermore, in step 3, the code knowledge graph is composed of package, class, method, parameter, variable, and statement as entities, and hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, and Call as edge types, and is visualized through the Neo4j graph database.
[0017] Furthermore, in step 6, the BIO sequence labeling method is used to manually label the defect report dataset, and BiLSTM-CRF is used for named entity recognition to obtain the defect entity sequence.
[0018] The present invention provides a defect location system based on code knowledge graph, which includes a source code extraction module, a source code parsing module, a code knowledge graph construction module, a crawling defect report module, a data set construction module, a named entity recognition module, a vectorization module, and a similarity calculation module;
[0019] Extract source code module is used to obtain the source code of AspectJ, SWT, and Zxing from the Git version control system;
[0020] The source code parsing module is used to parse the source code through the code parser to generate an abstract syntax tree AST of the source code;
[0021] The code knowledge graph construction module is used to extract entities and relationships from the constructed AST to construct a code knowledge graph;
[0022] The bug report crawling module is used to crawl bug reports from the open source Bugzilla bug tracking system and obtain the summary and description of the bug reports;
[0023] The dataset construction module is used to preprocess the defect report summary and description through the NLTK toolkit to obtain the defect report dataset;
[0024] The named entity recognition module is used to perform named entity recognition on the defect report dataset and extract the defect report entity sequence;
[0025] The vectorization module is used to vectorize the defect report entity sequence and the code knowledge graph through Word2Vec and knowledge graph embedding algorithm;
[0026] The similarity calculation module is used to map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, and calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph, arrange the similarities from high to low, and generate a list of suspicious methods.
[0027] Furthermore, in the source code parsing module, the Spoon code parser is used to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables and comments are recorded.
[0028] Furthermore, in the code knowledge graph construction module, the code knowledge graph is composed of package, class, method, parameter, variable, statement as entities, hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, Call as edge types, and is visualized through the Neo4j graph database.
[0029] Furthermore, in the named entity recognition module, the BIO sequence labeling method is used to manually label the defect report dataset, and BiLSTM-CRF is used for named entity recognition to obtain the defect entity sequence.
[0030] Beneficial effect: Compared with the prior art, the present invention has the following notable features: by constructing a code knowledge graph, the structural information and relationship information of the source code can be fully utilized to locate defects, and redundant information in the source code can be filtered out; by observing the defect report, the entity elements related to the defect in the defect report can be identified, and the BiLSTM+CRF model is used for named entity recognition to retain defect-related information, remove useless words, and improve the accuracy of defect location. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the process of the present invention;
[0032] Figure 2 The code in the present invention is just a picture Schema diagram;
[0033] Figure 3 This is a partial screenshot of the bug report summary of Bugzilla in the present invention;
[0034] Figure 4 This is a screenshot of the Bugzilla defect report description in the present invention;
[0035] Figure 5 It is a schematic diagram of the BiLSTM+CRF model in the present invention. DETAILED DESCRIPTION
[0036] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] Example 1
[0038] The present invention provides a defect location method based on code knowledge graph, please refer to Figure 1 As shown, the following steps are included:
[0039] Step 1: Get the source code of AspectJ, SWT, Zxing from the Git version control system.
[0040] Step 2: Parse the source code through the code parser to generate the abstract syntax tree AST of the source code.
[0041] Use the Spoon code parser to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables, and comments are recorded.
[0042] Step 3: Extract entities and relationships from the abstract syntax tree AST to build a code knowledge graph.
[0043] Step 3-1: Please refer to Figure 2As shown, the package name is extracted from the abstract syntax tree AST, and the types contained in the package are obtained through control flow, such as classes and interfaces, and then the methods and variables declared by the class, as well as the parameters and variables in the methods, and a triple is generated.
[0044] Step 3-2: Store triples in RDF format and use the Neo4j graph database to visualize the code knowledge graph. The entity and relationship types of the knowledge graph are shown in Table 3-1 and Table 3-2.
[0045] Table 3-1 Code knowledge graph entity type table
[0046]
[0047]
[0048] Table 3-2 Code knowledge graph relationship type table
[0049]
[0050] The code knowledge graph is composed of package, class, method, parameter, variable, and statement as entities, and hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, and Call as edge types, and is visualized through the Neo4j graph database.
[0051] Step 4: Crawl defect reports from the open source Bugzilla defect tracking system and obtain the summary and description of the defect reports.
[0052] Step 5: Use the NLTK toolkit to preprocess the defect report summary and description to obtain the defect report dataset.
[0053] See also Figure 3 and Figure 4 As shown in the figure, there are identifier names in the defect report, namely the defect summary part and the defect description part, which are helpful for defect location. First, the defect report is preprocessed by removing stop words and stemming, and then the composite token of the camel case notation method is split, for example, "commonviewer" is split into "common" and "viewer".
[0054] Use the Porter Algorith stemming algorithm to extract the token stems and transform "common" and "viewer" into "common" and "view".
[0055] Step 6: Perform named entity recognition on the defect report dataset and extract the defect report entity sequence.
[0056] Step 6-1: Use the BIO sequence labeling method to manually label the defect report dataset and mark the entities related to the defect.
[0057] Step 6-2: Please refer to Figure 5 As shown, the BiLSTM-CRF model is trained to perform sequence labeling and extract defect report entity sequences.
[0058] Step 7: Vectorize the defect report entity sequence and the code knowledge graph together through Word2Vec and the knowledge graph embedding algorithm.
[0059] The entity sequence extracted from the defect report is vectorized and represented with the code knowledge graph. The structural information in the code knowledge graph refers to information such as Class and Method, and the relationship information refers to the inheritance relationship between classes and the calling relationship between methods. This information can be used to more deeply mine defect-related files, thereby improving the accuracy of defect location and reducing the time and energy costs of code maintenance personnel.
[0060] Step 8: Map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, and calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph. k , m i ) is calculated as follows, where is the vector representation of the defect report entity sequence, It is the vector representation of the method in the source file code knowledge graph.
[0061]
[0062] Arrange the cosine similarities between the vector representations from high to low to generate a list of suspicious methods. The code method element with the highest cosine similarity is the most suspicious defect method. The defect report and the corresponding list of suspicious defect methods are shown in Table 8-1.
[0063] Table 8-1 Defect report summary and description and corresponding list of suspicious methods
[0064]
[0065]
[0066] Example 2
[0067] Corresponding to the defect location method based on code knowledge graph provided in Example 1, this embodiment provides a defect location system based on code knowledge graph, please refer to Figure 1 As shown, it includes source code extraction module, source code parsing module, code knowledge graph construction module, crawling defect report module, data set construction module, named entity recognition module, vectorization module, and similarity calculation module.
[0068] The source code extraction module is used to obtain the source code of AspectJ, SWT, and Zxing from the Git version control system.
[0069] The source code parsing module is used to parse the source code through the code parser to generate an abstract syntax tree AST of the source code.
[0070] Use the Spoon code parser to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables, and comments are recorded.
[0071] The code knowledge graph construction module is used to extract entities and relationships from the constructed AST and construct a code knowledge graph, including triple units and visualization units.
[0072] The triple unit is used to extract the package name from the abstract syntax tree AST, and obtain the types contained in the package through control flow, such as classes and interfaces, and then to the methods and variables declared by the class, as well as the parameters and variables in the method, and generate triples, such as Figure 2 shown.
[0073] The visualization unit is used to store triples in RDF format and use the Neo4j graph database to visualize the code knowledge graph. The entity and relationship types of the knowledge graph are shown in Table 3-1 and Table 3-2.
[0074] Table 3-1 Code knowledge graph entity type table
[0075]
[0076]
[0077] Table 3-2 Code knowledge graph relationship type table
[0078]
[0079] The code knowledge graph is composed of package, class, method, parameter, variable, and statement as entities, and hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, and Call as edge types, and is visualized through the Neo4j graph database.
[0080] The bug report crawling module is used to crawl bug reports from the open source Bugzilla bug tracking system and obtain the summary and description of the bug reports.
[0081] The dataset construction module is used to preprocess the defect report summary and description through the NLTK toolkit to obtain the defect report dataset.
[0082] See also Figure 3 and Figure 4 As shown in the figure, there are identifier names in the defect report, namely the defect summary part and the defect description part, which are helpful for defect location. First, the defect report is preprocessed by removing stop words and stemming, and then the composite token of the camel case notation method is split, for example, "commonviewer" is split into "common" and "viewer".
[0083] Use the Porter Algorith stemming algorithm to extract the token stems and transform "common" and "viewer" into "common" and "view".
[0084] The named entity recognition module is used to perform named entity recognition on the defect report dataset and extract the defect report entity sequence, including the annotation unit and the extraction sequence unit.
[0085] The annotation unit is used to manually annotate the defect report dataset using the BIO sequence annotation method to annotate entities related to the defect.
[0086] Extract sequence units to train the BiLSTM-CRF model, perform sequence labeling, and extract defect report entity sequences, such as Figure 5 shown.
[0087] The vectorization module is used to vectorize the defect report entity sequence and the code knowledge graph through Word2Vec and the knowledge graph embedding algorithm.
[0088] The entity sequence extracted from the defect report is vectorized and represented with the code knowledge graph. The structural information in the code knowledge graph refers to information such as Class and Method, and the relationship information refers to the inheritance relationship between classes and the calling relationship between methods. This information can be used to more deeply mine defect-related files, thereby improving the accuracy of defect location and reducing the time and energy costs of code maintenance personnel.
[0089] The similarity calculation module is used to map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, and calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph. k , m i ) is calculated as follows, where is the vector representation of the defect report entity sequence, It is the vector representation of the method in the source file code knowledge graph.
[0090]
[0091] Arrange the cosine similarities between the vector representations from high to low to generate a list of suspicious methods. The code method element with the highest cosine similarity is the most suspicious defect method. The defect report and the corresponding list of suspicious defect methods are shown in Table 8-1.
[0092] Table 8-1 Defect report summary and description and corresponding list of suspicious methods
[0093]
Claims
1. A defect location method based on code knowledge graph, characterized in that: The following steps are involved: Step 1: Get the source code of AspectJ, SWT, and Zxing from the Git version control system; Step 2: Parse the source code through the code parser to generate the abstract syntax tree AST of the source code; Step 3: Extract entities and relationships from the abstract syntax tree AST to build a code knowledge graph; Step 4: Crawl defect reports from the open source Bugzilla defect tracking system and obtain the summary and description of the defect report; Step 5: Use the NLTK toolkit to preprocess the defect report summary and description to obtain the defect report dataset; Step 6: Perform named entity recognition on the defect report dataset to extract the defect report entity sequence; Step 7: Vectorize the defect report entity sequence and the code knowledge graph using Word2Vec and knowledge graph embedding algorithms; Step 8: Map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph, arrange the similarities from high to low, and generate a list of suspicious methods.
2. The defect location method based on code knowledge graph according to claim 1 is characterized in that: In step 2, the Spoon code parser is used to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables, and comments are recorded.
3. The defect location method based on code knowledge graph according to claim 1 is characterized in that: In step 3, the code knowledge graph is composed of package, class, method, parameter, variable, and statement as entities, and hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, and Call as edge types, and is visualized through the Neo4j graph database.
4. The defect location method based on code knowledge graph according to claim 1 is characterized in that: In step 6, the BIO sequence labeling method is used to manually label the defect report dataset, and BiLSTM-CRF is used for named entity recognition to obtain the defect entity sequence.
5. A defect location system based on code knowledge graph, characterized in that: It includes source code extraction module, source code parsing module, code knowledge graph construction module, crawling defect report module, data set construction module, named entity recognition module, vectorization module, and similarity calculation module; Extract source code module is used to obtain the source code of AspectJ, SWT, and Zxing from the Git version control system; The source code parsing module is used to parse the source code through the code parser to generate an abstract syntax tree AST of the source code; The code knowledge graph construction module is used to extract entities and relationships from the constructed AST to construct a code knowledge graph; The bug report crawling module is used to crawl bug reports from the open source Bugzilla bug tracking system and obtain the summary and description of the bug reports; The dataset construction module is used to preprocess the defect report summary and description through the NLTK toolkit to obtain the defect report dataset; The named entity recognition module is used to perform named entity recognition on the defect report dataset and extract the defect report entity sequence; The vectorization module is used to vectorize the defect report entity sequence and the code knowledge graph through Word2Vec and knowledge graph embedding algorithm; The similarity calculation module is used to map the code knowledge graph and the vector representation of the defect report entity sequence to the same vector space, and calculate the cosine similarity between the vector representation of the defect report entity sequence and the vector representation of the source file code knowledge graph, arrange the similarities from high to low, and generate a list of suspicious methods.
6. The defect location system based on code knowledge graph according to claim 5 is characterized in that: In the source code parsing module, the Spoon code parser is used to parse the source code into an abstract syntax tree AST. Starting from the source code package, the control flow moves to the types contained in the package, and then to the variables and methods declared in the class. Each method is analyzed and the parameters, variables, and comments are recorded.
7. The defect location system based on code knowledge graph according to claim 5 is characterized in that: In the code knowledge graph construction module, the code knowledge graph is composed of package, class, method, parameter, variable, and statement as entities, with hasPackage, hasVariable, hasMethod, hasParameter, Extend, hasStatement, and Call as edge types, and is visualized through the Neo4j graph database.
8. The defect location system based on code knowledge graph according to claim 5 is characterized in that: In the named entity recognition module, the BIO sequence labeling method is used to manually annotate the defect report dataset, and BiLSTM-CRF is used for named entity recognition to obtain the defect entity sequence.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to claim 1 to claim 4 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 1 to claim 4 are implemented.