A defect search method and system based on code knowledge graph

By building a defect search system based on code knowledge graph, the problem of inaccurate search results in the existing technology is solved, and multi-source information fusion and visualization of defect code are realized, and the efficiency and accuracy of defect repair are improved.

CN115562673BActive Publication Date: 2025-05-09YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211190008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-05-09
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

Existing defect search technologies based on keyword matching cannot accurately understand user intentions, resulting in inaccurate search results and low correlation, making it difficult to effectively support defect repair.

Method used

Build a defect search system based on code knowledge graph, build a defect code knowledge graph by crawling defect reports and post information from multiple open source platforms, and use LDA theme models and visualization tools for information fusion and visualization, combine NLTK and Nicad code cloning detection tools for similarity calculation, and return relevant defect codes and repair codes.

Benefits of technology

It realizes multi-source information fusion and visualization of defective code, expands the information coverage, improves the accuracy and efficiency of retrieval, and enables developers to quickly query repaired code information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115562673B_ABST
    Figure CN115562673B_ABST
Patent Text Reader

Abstract

The present invention discloses a defect search method and system based on code knowledge graph, which integrates the codes before and after defect repair of Mozilla@Bugzilla, Eclipse@Bugzilla, Github and Stack overflow websites from the perspective of text and code, crawls different topics, and builds a topic set, extracts defect code and correct code in posts at the same time, and establishes a code knowledge graph with defect code, correct code, title information in defect report, post title information and problem description information, and topic as entities, and visualizes the code knowledge graph with the help of visualization tools. Crawling from multiple platforms makes the coverage of defect code knowledge graph more and wider, and at the same time, by integrating code text and code information, popularizes code knowledge, so that developers can intuitively have a certain understanding of defect code, and developers can quickly query the repaired code information when searching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software security, and in particular to a defect search method and system based on a code knowledge graph. Background Art

[0002] Software defects are inevitable in the process of software development and maintenance. As the scale of modern software continues to grow, the number of software defects and the difficulty of repairing them have increased, causing huge economic losses to enterprises. Therefore, it is crucial to complete the work of repairing defects in a timely and effective manner. With the advent of the big data era, the booming development of knowledge graphs has made it possible for developers to repair program defects with high quality. Knowledge graphs can provide powerful background knowledge for understanding user intentions, support logical reasoning of knowledge and facts, and help achieve semantic-level search generalization; secondly, the repair code recommendation based on knowledge graphs has strong interpretability, which helps users understand the recommendation results and supports the optimization of recommendation rules and processes.

[0003] The defect code knowledge graph extracts entities, relationships, and attributes from natural language text and code knowledge for knowledge fusion, and then builds a system framework through ontology and stores them in the form of structured triples. Traditional information retrieval methods are based on keyword matching for information retrieval, but search engines do not understand user input and can only obtain keywords by segmenting the content input by users, and then perform similarity matching between keywords and data in the database, and then return the matching results to users according to a certain sorting algorithm. Finally, users browse and select the desired search results from the returned results. Because the real purpose of the user's search cannot be understood, the retrieval technology based on keyword matching has obvious defects. Not only is the information obtained inaccurate, but sometimes the results obtained are also of low relevance to the search content. Summary of the invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a defect search method and system based on a code knowledge graph, which integrates relevant multi-source information of the code before and after defect repair for identification and fusion, and visualizes it as a defect code knowledge graph to facilitate developers' query.

[0005] Technical solution: The present invention provides a defect search method based on code knowledge graph, which comprises the following steps:

[0006] Step 1: Crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla to obtain the Bug-ID and title description information in the bug reports;

[0007] Step 2: Crawl posts related to Java defects from the Stack Overflow website and extract post title information, question description information, answer description information, first defect code, and first repair code;

[0008] Step 3: Use the LDA topic model to extract topics from the title description information and post title information in the bug report to form a topic set;

[0009] Step 4: Crawl the commit information corresponding to the bug ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code;

[0010] Step 5: Establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the subject set as entities, and visualize the code knowledge graph with the help of a visualization tool;

[0011] Step 6: Enter the defect problem description and the corresponding defect code, and the system returns the relevant defect code and repair code information.

[0012] Furthermore, in step 2, posts related to Java defect repair are crawled from the Stack overflow forum with tags "java", "isaccepted:1", "hascode:1", and "fix", and post title information, question descriptive information, answer descriptive information, the first defect code, and the first repair code are extracted.

[0013] Furthermore, in step 5, a defect code knowledge graph is established with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, hasquestion, hasbuggycode, hasfixcode as edges, and related description information as attributes, and the defect code knowledge graph is visualized with the help of visualization tools.

[0014] Furthermore, in step 6, the defect problem description and the corresponding defect code are input, and the NLTK technology is used to extract keywords for the defect problem. The extracted keywords are matched with the topics in the defect code knowledge graph, and the similarity is calculated to find the optimal topic. Then, the Nicad code clone detection tool is used to perform a similarity detection on the input defect code and the defect code under the optimal topic to find the related code, and the first defect code, the first repair code, the second defect code, and the second repair code information are returned in the form of a list.

[0015] The present invention provides a defect search system based on code knowledge graph, including a crawling defect report module, a crawling post module, a theme set module, a crawling report module, a visualization module, and a query module;

[0016] The bug report crawling module is used to crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla, and obtain the Bug-ID and title description information in the bug report;

[0017] The post crawling module is used to crawl posts related to Java defects from the Stack Overflow website, and extract post title information, question description information, answer description information, first defect code, and first repair code;

[0018] The topic set module is used to extract topics from the title description information and post title information in the defect report using the LDA topic model to form a topic set together;

[0019] The crawling report module is used to crawl the Commit information corresponding to the Bug-ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code;

[0020] The visualization module is used to establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the subject set as entities, and visualize the code knowledge graph with the help of a visualization tool;

[0021] The query module is used for developers to input defect problem descriptions and corresponding defect codes, and the system returns relevant defect codes and repair codes and related information.

[0022] Furthermore, in the post crawling module, posts related to Java defect repair are crawled from the Stack overflow forum with tags "java", "isaccepted:1", "hascode:1", and "fix", and post title information, question descriptive information, answer descriptive information, the first defect code, and the first repair code are extracted.

[0023] Furthermore, in the visualization module, a defect code knowledge graph is established with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, hasquestion, hasbuggycode, hasfixcode as edges, and related description information as attributes, and the defect code knowledge graph is visualized with the help of visualization tools.

[0024] Furthermore, in the query module, the defect problem description and the corresponding defect code are input, and the NLTK technology is used to extract keywords for the defect problem. The extracted keywords are matched with the topics in the defect code knowledge graph, and the similarity is calculated to find the optimal topic. Then, the Nicad code clone detection tool is used to perform a similarity detection on the input defect code and the defect code under the optimal topic to find the related code, and the first defect code, the first repair code, the second defect code, and the second repair code information are returned in the form of a list.

[0025] Beneficial effect: Compared with the prior art, the present invention has the remarkable feature of crawling different topics from different open source websites, constructing a defect code knowledge graph of the code before and after repair, and crawling from multiple platforms, so that the defect code knowledge graph covers more content and a wider range. At the same time, by integrating code text and code information, code knowledge is popularized, so that developers can intuitively have a certain understanding of the defect code, and developers can quickly query the repaired code information when searching. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the process of the present invention;

[0027] Figure 2 The defect report diagrams of Mozilla@Bugzilla and Eclipse@Bugzilla in the present invention;

[0028] Figure 3 It is a related diagram of Stack Overflow forum posts in the present invention;

[0029] Figure 4 Schematic diagram of the code2vec model in the present invention;

[0030] Figure 5 This is a graph showing the result of using code2vec to predict a python method name in the present invention;

[0031] Figure 6 This is the Commit report diagram in GitHub in the present invention;

[0032] Figure 7 It is a partial schematic diagram of the defect code knowledge graph in the present invention;

[0033] Figure 8 This is the Nicad operation flow chart of the present invention. DETAILED DESCRIPTION

[0034] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Example 1

[0036] The present invention provides a defect search method based on code knowledge graph, please refer to Figure 1 As shown, the following steps are included:

[0037] Step 1: Crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla, and obtain the Bug-ID and title description information in the bug report. The screenshot of the bug report in Mozilla@Bugzilla is as follows: Figure 2 As shown in (a) in the figure, the screenshot of the defect report in Eclipse@Bugzilla is as follows Figure 2 As shown in (b) in .

[0038] Step 2: Crawl posts related to Java defects from the Stack overflow website and extract post title information, question descriptive information, answer descriptive information, first defect code, and first repair code.

[0039] Step 2-1: Crawl posts related to Java bug fixes from the Stack Overflow forum with tags "java", "isaccepted:1", "hascode:1", and "fix". Figure 3 As shown in (a) in the figure, it shows a screenshot of the crawled post problem part, which mainly includes the post title description information and the defect code part. Figure 3 As shown in (b) in the figure, it shows a screenshot of the crawled post answering the question.

[0040] Step 2-2: Filter the obtained posts and filter out posts with other code languages ​​in the tags, such as "C++", "python", etc.; filter out posts that do not contain code in questions or answers; in order to facilitate the extraction of the correspondence between correct code and defective code, only posts that only involve one code block in both questions and answers are retained, and posts containing multiple code blocks are filtered out; filter out short codes of three lines or less in questions because they do not have sufficient context information.

[0041] Step 2-3: Extract relevant information of the filtered posts, mainly including post tags, post title information, question descriptive information, answer descriptive information, first defect code, and first repair code.

[0042] Step 2-4: Please refer to Figure 3 As shown in (c) in the figure, the defective code is not a complete method. Code2vec is used to predict the method name of the defective code line, as shown in Figure 4 As shown, make it a complete method, such as Figure 5 Shown are the prediction results of code2vec.

[0043] Step 3: Use the LDA topic model to extract topics from the title description information and post title information in the bug report to form a topic set.

[0044] Step 4: Crawl the Commit information corresponding to the Bug-ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code.

[0045] Step 4-1: Search for the corresponding commit information on Github according to the Bug-ID in the defect report, such as Figure 6 The Commit report shown is matched according to the title information in the defect report and the title information on GitHub. If they are the same, the corresponding Commit information is obtained, otherwise the information of the defect report is discarded.

[0046] Step 4-2: Extract the diff from the Commit report, split the diff into the second defect code and the second repair code, and store them in the database one by one.

[0047] Step 5: Establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the topic set as entities, and visualize the code knowledge graph with the help of visualization tools.

[0048] Step 5-1: Convert the data in the database into triples, i.e. (entity 1, relationship, entity 2), and establish a code knowledge graph with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, has question (problem code), has buggycode (defect code), has fixedcode (fix code) as edges, and a defect code knowledge graph with related description information as attributes.

[0049] Step 5-2: Please refer to Figure 7 As shown in the figure, the code knowledge graph is visualized with the help of the open source NOSQL graph database Neo4j.

[0050] Step 6: The developer enters the defect problem description and the corresponding defect code, and the system returns the most relevant defect code and repair code and related information.

[0051] Step 6-1: The developer enters the defect problem description information and the corresponding defect code in turn, uses NLTK technology to extract keywords for the defect problem, and then matches the extracted keywords with the topics in the code knowledge graph to find the most similar topics.

[0052] Step 6-2: Use the Nicad code clone detection tool to perform similarity detection on the defect code entered by the developer and the defect code under the most similar topic set. Figure 8 As shown, this embodiment uses the version Nicad6 to find the most relevant code and returns the first defect code, the first repair code, the second defect code, the second repair code and related information to the developer in the form of a list.

[0053] Example 2

[0054] Corresponding to the defect search method based on code knowledge graph in Example 1, this Example 2 provides a defect search system based on code knowledge graph, please refer to Figure 1 As shown, it includes crawling defect report module, crawling post module, topic set module, crawling report module, visualization module and query module.

[0055] The bug report crawling module is used to crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla, and obtain the Bug-ID and title description information in the bug report. The screenshot of the bug report in Mozilla@Bugzilla is as follows: Figure 2 As shown in (a) in the figure, the screenshot of the defect report in Eclipse@Bugzilla is as follows Figure 2 As shown in (b) in .

[0056] The post crawling module is used to crawl posts related to Java defects from the Stack overflow website, and extract post title information, question descriptive information, answer descriptive information, first defect code, and first repair code.

[0057] The post crawling module includes a crawling unit, a filtering unit, an extraction unit, and a prediction unit;

[0058] The crawler unit is used to crawl posts related to Java defect fixes from the Stack Overflow forum with tags "java", "isaccepted:1", "hascode:1", and "fix". See Figure 3 As shown in (a) in the figure, it shows a screenshot of the crawled post problem part, which mainly includes the post title description information and the defect code part. Figure 3 As shown in (b) in the figure, it shows a screenshot of the crawled post answering the question.

[0059] The filtering unit is used to filter the acquired posts, filter out posts with other code languages ​​in the tags, such as "C++", "python", etc.; filter out posts that do not contain code in questions or answers; in order to facilitate the extraction of the correspondence between correct code and defective code, only posts that only involve one code block in both questions and answers are retained, and posts containing multiple code blocks are filtered out; short codes of three lines or less in questions are filtered out because they do not have sufficient context information.

[0060] The extraction unit is used to extract relevant information of the filtered posts, mainly including post tags, post title information, question descriptive information, answer descriptive information, first defect code, and first repair code.

[0061] The prediction unit is used to make the defective code complete through prediction, see Figure 3 As shown in (c) in the figure, the defective code is not a complete method. Code2vec is used to predict the method name of the defective code line, as shown in Figure 4 As shown, make it a complete method, such as Figure 5 Shown are the prediction results of code2vec.

[0062] The topic set module is used to extract topics from the title description information and post title information in the defect report using the LDA topic model to form a topic set.

[0063] The crawling report module is used to crawl the Commit information corresponding to the Bug-ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code.

[0064] The crawl report module includes search unit and split unit;

[0065] The search unit is used to search for the corresponding Commit information on Github according to the Bug-ID in the defect report, such as Figure 6 The Commit report shown is matched according to the title information in the defect report and the title information on GitHub. If they are the same, the corresponding Commit information is obtained, otherwise the information of the defect report is discarded.

[0066] The splitting unit is used to extract diff from the Commit report, split the diff into the second defect code and the second repair code, and store them in the database in a one-to-one correspondence.

[0067] The visualization module is used to establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the topic set as entities, and visualize the code knowledge graph with the help of visualization tools.

[0068] The visualization module includes a code knowledge graph construction unit and a visualization unit;

[0069] A code knowledge graph unit is constructed to convert the data in the database into the form of triples, namely (entity 1, relationship, entity 2), and a code knowledge graph is established with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, has question, has buggycode, hasfixedcode as edges, and a defect code knowledge graph with related description information as attributes.

[0070] The visualization unit is used to visualize the code knowledge graph with the help of the open source NOSQL graph database Neo4j, such as Figure 7 shown.

[0071] The query module is used for developers to input defect problem descriptions and corresponding defect codes, and the system returns the most relevant defect codes and repair codes and related information.

[0072] The query module includes a question input unit and a question return unit;

[0073] The input problem unit is used by developers to input defect problem description information and corresponding defect codes in sequence, extract keywords from the defect problems using NLTK technology, and then match the extracted keywords with the topics in the code knowledge graph to find the most similar topics.

[0074] The problem return unit is used to use the Nicad code clone detection tool to perform similarity detection on the defect code entered by the developer and the defect code under the most similar topic set. Figure 8 As shown, this embodiment uses the version Nicad6 to find the most relevant code and returns the first defect code, the first repair code, the second defect code, the second repair code and related information to the developer in the form of a list.

Claims

1. A defect search method based on code knowledge graph, characterized in that: The following steps are involved: Step 1: Crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla to obtain the Bug-ID and title description information in the bug reports; Step 2: Crawl posts related to Java defects from the Stack Overflow website and extract post title information, question description information, answer description information, first defect code, and first repair code; Step 3: Use the LDA topic model to extract topics from the title description information and post title information in the bug report to form a topic set; Step 4: Crawl the Commit information corresponding to the Bug-ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code; Step 5: Establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the subject set as entities, and visualize the code knowledge graph with the help of a visualization tool; Step 6: Enter the defect problem description and the corresponding defect code, and the system returns the relevant defect code and repair code information.

2. The defect search method based on code knowledge graph according to claim 1 is characterized in that: In step 2, posts related to Java defect repair are crawled from the Stack overflow forum with "java", "isaccepted:1", "hascode:1", and "fix" as tags, and post title information, question description information, answer description information, first defect code, and first repair code are extracted.

3. The defect search method based on code knowledge graph according to claim 1 is characterized in that: In step 5, a defect code knowledge graph is established with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, hasquestion, hasbuggycode, hasfixcode as edges, and related description information as attributes, and the defect code knowledge graph is visualized with the help of visualization tools.

4. The defect search method based on code knowledge graph according to claim 1 is characterized in that: In step 6, input the defect problem description and the corresponding defect code, use NLTK technology to extract keywords for the defect problem, match the extracted keywords with the topics in the defect code knowledge graph, calculate the similarity, find the optimal topic, and then use the Nicad code clone detection tool to perform a similarity test on the input defect code and the defect code under the optimal topic to find the related code, and return the first defect code, the first repair code, the second defect code, and the second repair code information in the form of a list.

5. A defect search system based on code knowledge graph, characterized in that: Contains crawling defect report module, crawling post module, topic set module, crawling report module, visualization module, and query module; The bug report crawling module is used to crawl bug reports from the open source bug tracking systems Mozilla@Bugzilla and Eclipse@Bugzilla, and obtain the Bug-ID and title description information in the bug report; The post crawling module is used to crawl posts related to Java defects from the Stack Overflow website, and extract post title information, question description information, answer description information, first defect code, and first repair code; The topic set module is used to extract topics from the title description information and post title information in the defect report using the LDA topic model to form a topic set together; The crawling report module is used to crawl the Commit information corresponding to the Bug-ID in the defect report from the open source Github website, extract the diff, and split the diff into the second defect code and the second repair code; The visualization module is used to establish a defect code knowledge graph with the first defect code, the first repair code, the second defect code, the second repair code, the title description information in the defect report, the post title information, and the subject set as entities, and visualize the code knowledge graph with the help of a visualization tool; The query module is used to input the defect problem description and the corresponding defect code, and the system returns the relevant defect code and repair code information.

6. The defect search system based on code knowledge graph according to claim 5 is characterized in that: In the post crawling module, posts related to Java defect repair are crawled from the Stack overflow forum with tags "java", "isaccepted:1", "hascode:1", and "fix", and post title information, question description information, answer description information, first defect code, and first repair code are extracted.

7. The defect search system based on code knowledge graph according to claim 5 is characterized in that: In the visualization module, a defect code knowledge graph is established with the first defect code, the first fix code, the second defect code, the second fix code, the title description information in the defect report, the post title information and the subject as entities, hasquestion, hasbuggycode, hasfixcode as edges, and related description information as attributes, and the defect code knowledge graph is visualized with the help of visualization tools.

8. The defect search system based on code knowledge graph according to claim 5 is characterized in that: In the query module, input the defect problem description and the corresponding defect code, use NLTK technology to extract keywords for the defect problem, match the extracted keywords with the topics in the defect code knowledge graph, calculate the similarity, find the optimal topic, and then use the Nicad code clone detection tool to perform similarity detection on the input defect code and the defect code under the optimal topic to find the related code and return the first defect code, the first repair code, the second defect code, and the second repair code in the form of a list.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to claim 1 to claim 4 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 1 to claim 4 are implemented.