Defect repair data set construction method for defect category and repair granularity

By extracting multi-repair semantic information from Defects4j, Git and SVN, combining word matching, static analysis, large language model deviation prediction and clustering technology, we construct defect repair data sets for defect categories and repair granularity, solving the repetition problem of existing automatic repair tools in the construction of data sets, and achieving efficient automatic repair data set construction.

CN120067676APending Publication Date: 2025-05-30NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510091023.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing automated repair tools remake wheels when building manual repair patch datasets, lacking effective dataset construction methods for defect categories and repair granularity.

Method used

By extracting multi-repair semantic information from Defects4j, Git and SVN, combining word matching, static analysis, large language model deviation prediction and clustering technology, we judge the repair category, granularity and structural similarity, and construct a defect repair data set for defect categories and repair granularity.

Benefits of technology

It realizes the effective mining of information from multi-repair semantics, builds high-quality repair data sets that can be learned by automatic repair programs, overcomes the lack of defect categories and repair granularity division in the existing technology, and improves the development and verification efficiency of data-driven automatic repair tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067676A_ABST
    Figure CN120067676A_ABST
Patent Text Reader

Abstract

The invention discloses a defect repair data set construction method for defect categories and repair granularity. The method comprises the following steps: step 1) inputting a configuration file for identifying a to-be-processed project; 2) aiming at the heterogeneity of Defects4j, SVN and Git management projects, completing validity analysis of repairing of the to-be-processed project and information extraction of multi-dimensional repairing semantics; 3) based on lexical element matching and a large language model, judging a repair category from multi-dimensional repair semantics; 4) based on a static analysis method and a large language model, judging the repair granularity and cutting the repair granularity from the multi-dimensional semantics; 5) identifying similar repair structures in the same repair based on clustering; and 6) carrying out persistence on the classification information, the operation configuration information, the repair structure information and the artificial identification information.According to the method, a repair data set which aims at various types and can be learned by an automatic repair program can be efficiently constructed from Defects4j, Git and SVN, and the development and verification efficiency of a data-driven automatic repair tool is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software quality assurance, and particularly to a method for constructing a defect repair dataset for defect categories and repair granularities. Background Art

[0002] Software defects often cause software programs to run in incorrect or other unexpected ways. There are various categories of software defects. Currently, mainstream software defects include programming language syntax errors, business logic errors, runtime errors, memory leaks, security vulnerabilities, concurrency errors, performance bottlenecks, etc. Defects not only affect the business functions of software but also damage the reliability and security of software. They can cause incorrect outputs, high latency in function feedback at the least, and system crashes, data corruption, user privacy leakage, etc. at the worst, making users extremely dissatisfied. Moreover, with the rapid development of computer science and technology, software has emerged in the economy and politics around the world and plays an important role. The crisis brought by software defects has thus become a new global challenge. If the defects of important software are attacked or exploited by hackers, the losses and costs are unpredictable.

[0003] The reasons for software defects are also diverse and may be introduced at any stage of development. The traditional software development process includes requirement collection, requirement analysis, software architecture design, software coding, software testing, software deployment, and software maintenance. The worst defects are introduced during the requirement analysis and software architecture design stages, which often result from misunderstandings of requirements, poor design decisions, or outdated design concepts. More narrowly, software defects refer to those introduced during the software coding stage. During the software development process of generations of programmers over several decades, some effective methodologies have gradually emerged to improve development efficiency and minimize the occurrence of defects as much as possible.

[0004] Despite many factors leading to the occurrence of defects, misunderstandings, negligence, and judgment errors of software developers are still the main reasons for the appearance of defects. Human developers will definitely make mistakes or not program according to best practices. Therefore, defect location and repair are still problems that software development must face. Our efforts should not only focus on proposing various tools and methodologies to reduce the introduction of defects but also develop various highly automated tools to improve the efficiency of software defect detection and defect repair. For this purpose, many automated defect location technologies and automatic software repair technologies have emerged like bamboo shoots after a spring rain.

[0005] The main purpose of defect localization technology is to identify specific locations or specific code segments in a software system that cause system failures. It plays a crucial role in the process of software developers debugging code. By narrowing down the search space when software developers locate defects, it helps to efficiently solve software defects. Common defect localization techniques include: statistical-based techniques, mutation-based techniques, slicing and machine learning-based techniques, etc.

[0006] The goal of automated software repair is to automate and simplify the parts of traditional software maintenance processes that rely on manual debugging and manual patching by developing techniques and tools for automatically identifying and fixing software defects. Automated software repair techniques or tools often need to use one or several of the above-mentioned defect localization techniques, and then go a step further on this basis to generate possible candidate repair solutions for software developers to refer to. Compared with defect localization techniques, it takes a further step in the automation and intelligence of defect repair. After 2016, automated program repair tools officially entered the data-driven era. They all need to collect a large number of defect repair code changes from the history of manual repairs, and then learn patch repair patterns from these changes, or use these changes as a validation set to test the effectiveness of automated repair tools. However, the number of manual code modifications required in data-driven automated repair tools often reaches thousands, and it is often not an easy task to collect enough manual modifications that meet the requirements. Based on the above insights, in order to fill the gap in the construction tool for the artificial defect repair dataset and solve the problem of reinventing the wheel in the construction of the artificial repair patch dataset for automated repair tools, the present invention provides a method for constructing a defect repair dataset for defect categories and repair granularities. Summary of the Invention

[0007] To solve the problems existing in the above-mentioned prior art, the present invention provides a method for constructing a defect repair dataset for defect categories and repair granularities, which can extract repair semantic information in multiple dimensions such as commit messages, repair file lists, specific repair content, and test case passing conditions based on Defects4j, Git, and SVN, and then try to obtain whether this repair is worth recording, the category of this repair, the granularity of this repair, and the structural similarity marker of this repair; extracting this information combines multiple techniques such as token matching, static analysis, large language model bias prediction, and clustering-based structural similarity measurement, using methods based on token matching and large language models to judge the repair category from multi-dimensional repair semantics, methods based on static analysis and large language models to judge the repair granularity from multi-dimensional semantics, and methods based on clustering to identify similar repair structures in the same type of repair; finally, use a relational database to persist the classification information, running configuration information, repair structure information, and manual identification information. The present invention is implemented through the following technical solutions.

[0008] A method for constructing a defect repair dataset for defect categories and repair granularities, characterized in that the method comprises the following steps: Step 1) Input a configuration file identifying the project to be processed; Step 2) For the heterogeneity of projects managed by Defects4j, SVN, and Git, perform effectiveness analysis of the repairs completed on the project to be processed and extract information on multi-repair semantics; Step 3) Judge the repair category from the multi-repair semantics based on token matching and large language models; Step 4) Judge the repair granularity and repair granularity cutting from the multi-dimensional semantics based on static analysis methods and large language models; Step 5) Identify similar repair structures in the same type of repair based on clustering; Step 6) Persist the classification information, running configuration information, repair structure information, and manual identification information.

[0009] Further, in the above method for constructing a defect repair dataset for defect categories and repair granularities, the configuration file in Step 1) needs to be a csv format file. The configuration file can specify projects in Github or Defects4j as the projects to be extracted; if the configuration file is to specify a project in Github, the column names of the csv need to be "url, branch"; where the first column url specifies the address of the project's github repository, and the second column branch specifies the specific branch of the project; if the configuration file is to specify a project in Defects4j, the column names of the csv need to be "repoId,bugId", where the first column repoId specifies the project number in the Defects4j dataset, and the second column bugId specifies the defect number in the project; the configuration file specifies which projects to extract the defect dataset from.

[0010] Further, in the above method for constructing a defect repair dataset for defect categories and repair granularities, Step 2) specifically comprises the following steps: Step 21) Parse the configuration file provided in Step 1) and verify the correctness of the configuration file. The correctness verification includes: whether the url can be accessed, whether the repoId exists, and whether the bugId exists; Step 22) Create a thread pool according to whether the current configuration file specifies a Git / SVN project or a Defects4j project, and start different parallel processing processes.

[0011] Further, the method for constructing a defect repair dataset for defect categories and repair granularities is characterized in that if the currently configured file specifies a Git / SVN project in step 22), the parallel processing flow includes the following steps: Step 201) For a project managed by Git, obtain the project source code and git files through the git clone command, and use git checkout to switch branches; for a project managed by SVN, directly obtain the project corresponding to the version number through svn checkout; Step 202) Compile the project through maven, run the current test set, and record which tests can pass normally; Step 203) For a project managed by Git, extract commit information from git log, obtain the file information modified in this repair through git diff, and if only a single file is modified, the specific modification information of the single file will also be extracted; for a project managed by SVN, obtain the commit information through svn log, obtain the file information modified in this repair through svn diff, and if only a single file is modified, the specific modification information of the single file will also be extracted; Step 204) Summarize the source code, test set passing situation, commit information, modified file information, and modified single file information to generate Git and SVN defect repair candidate datasets.

[0012] Further, the method for constructing a defect repair dataset for defect categories and repair granularities is characterized in that if the currently configured file specifies a Defects4j project in step 22), the parallel processing flow includes the following steps: Step 211) Extract the repaired source code from the Mysql database of Defects4j; Step 212) Use defects4j export to extract the test passing situation and repair category information; Step 213) Compile and run the tests, and check with the existing information; Step 214) Generate a Defects4j defect repair data candidate set; Step 215) Finally, merge the two to generate the final candidate dataset.

[0013] Further, the method for constructing a defect repair dataset for defect categories and repair granularities is characterized in that step 3) specifically includes the following steps: Step 31) If the configured file specifies a Defects4j project, directly obtain the defect classification; Step 32) If the configured file specifies a Git / SVN project, then execute step 33); Step 33): Input the commit message collected in Step 2) into the large language model to determine whether this modification is a defect fix; if the large language model determines it is not, end the process; if the large model determines it is, continue to execute Step 34); Step 34): Obtain the commit message of the previous fix after the repair, input it into the large language model, and determine whether it is the same as the previous fix for the same time; if it is, end the process and discard the code fix across commits; if not, continue to execute Step 35); Step 35): Perform lemmatization matching classification on the commit message; if the lemmatization match is successful, input the commit message and the lemmatization matching classification result into the large language model to obtain the confidence level; Step 36): Examine the confidence level of the large language model for this classification. If it is greater than 0.8, approve this classification and record the classification; otherwise, discard this candidate example; Step 37): If the lemmatization match fails, input the commit message and the source code modification information into the large language model to calculate the deviation of each category; the deviation calculation formula for a single category is as follows , where represents a specific example, y represents a category, is the set of all categories; and represents the confidence level of asking the large language model whether the example belongs to category y for example ; Step 38): Calculate the uniquely belonging category for the example through the formula , approve and record this classification.

[0014] Furthermore, the above method for constructing a defect repair dataset for defect categories and repair granularity is characterized in that the lemmatization matching in Step 35) includes the following lemmas: If the first type of keyword appears in the commit message, it will be directly determined as the corresponding category; if no keyword appears, but the complete word / phrase of the second type of keyword appears, it will also be assigned to the corresponding category.

[0015] Furthermore, the above method for constructing a defect repair dataset for defect categories and repair granularity is characterized in that Step 4) specifically includes: Step 41): Examine the list of modified files. If there are multiple source files modified, classify it as the file level and end the process of the entire Step 4); otherwise, execute Step 42); Step 42) Use static analysis method to parse a single file into SpoonAst, and examine whether the modifications in a file cross classes / functions; if they cross classes, classify them at the class level and end the entire process of Step 4); otherwise, execute Step 43); Step 43) At this time, this modification is classified at the multi-function body level, and execute Step 44) and Step 46) in parallel Step 44) Use static analysis method to parse a single function into SpoonAst, and examine whether the modification in a file is only a single operation at consecutive positions; the single operation means that the operation only includes one of addition, deletion, and modification; if so, classify it at the consecutive statement level and end the entire process of Step 4); otherwise, execute Step 45); Step 45) Input the modified content into a large language model to determine whether this modification is the movement of a single statement block; if the confidence level is not less than 0.8, classify it at the consecutive statement level; otherwise, classify it at the single function body level; end the entire process derived from Step 4); Step 46) Input the commit message of this cross-function modification into the LLM to examine the problems it modifies; if the result is the repair of multiple problems and the confidence level is not less than 0.8, execute the next step; otherwise, end the entire process derived from Step 4); Step 47) Split this cross-function modification into different single-function modifications; Step 48) Input the diff file of the single-function repair into the large model to determine whether it is a complete repair of a single problem; if the confidence level is less than 0.8, end the entire process derived from Step 4). If the confidence level is not less than 0.8, regard this split as a single-function repair and execute Step 44).

[0016] Furthermore, the above-mentioned method for constructing a defect repair dataset for defect categories and repair granularities is characterized in that Step 5) includes a data preprocessing process, a feature extraction process, and a clustering execution process, and the data preprocessing process includes: Step 51) Parse the function code information before and after repair into SpoonAst; Step 52) Remove local variable definitions and normalize local variables to corresponding positions; Step 53) Unify the exchange rate and place constants on the left; Step 54) Unify the statement blocks of conditional statements and loop statements to be located in Block; Step 55) Replace the variable names, function names, and string contents named manually in the code with normalized characters calculated according to the structure; Step 56) Generate a preprocessed diff file based on AST; The feature extraction process includes: Step 57) Calculate the word frequency and inverse word frequency in the code file; Step 58) Multiply the word frequency and inverse word frequency of each word in the document to form a feature vector; The clustering execution process includes: Step 59) Perform K-means clustering based on cosine similarity to obtain clustering labels.

[0017] Furthermore, the above method for constructing a defect repair dataset for defect categories and repair granularities is characterized in that the step 6) specifically includes: Step 61) Take out the code before and after the candidate commit and parse it into SpoonAst; Step 62) Remove local variable definitions and normalize local variables to corresponding positions; Step 63) Unify the exchange rate and place constants on the left; Step 64) Unify the statement blocks of conditional statements and loop statements to be located in Block; Step 65) Save the above rewrite result as a sample before and after modification and store it in the specified path; Step 66) Newly establish Mysql database entries to save the samples before and after the repair; Step 67) Save the paths of the samples before and after the modification, commit message, test pass situation, and running environment parameters to the Mysql database; Step 68) Save the repair classification label, repair granularity label, and clustering label to the Mysql database.

[0018] The present invention adopts the above technical solutions and has the following beneficial effects: The present invention uses the methods of static analysis and LLM-assisted verification, which can mine information from various code repair semantics and efficiently construct a repair dataset for multiple categories that can be learned by an automatic repair program from Defects4j, Git, and SVN. It overcomes the problem that the existing defect repair dataset lacks the division of defect categories and repair granularities, and improves the development and verification efficiency of data-driven automatic repair tools. Brief Description of the Drawings

[0019] Figure 1 It is a flowchart of the method for constructing a defect repair dataset for defect categories and repair granularities according to an embodiment of the present invention.

[0020] Figure 2 It is a schematic flowchart of the effectiveness analysis for completing the repair of a project to be processed and the information extraction of multi-repair semantics according to an embodiment of the present invention.

[0021] Figure 3This is a schematic flowchart of the method for judging the repair category from multi-dimensional repair semantics based on token matching and large language models in the embodiments of the present invention.

[0022] Figure 4 This is a schematic flowchart of the method for judging the repair granularity and repair granularity cutting from multi-dimensional semantics based on static analysis and large language models in the embodiments of the present invention.

[0023] Figure 5 This is a schematic flowchart of the method for identifying similar repair structures in the same kind of repair based on clustering in the embodiments of the present invention. Detailed implementation manners

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] The present invention provides a method for constructing a defect repair dataset for defect categories and repair granularities. This method can extract repair semantic information from multiple dimensions such as commit messages, repair file lists, specific repair contents, and test case passing situations from two mainstream code repositories, Defects4j and Github, and then try to obtain whether this repair is worth recording, the category of this repair, and the granularity of this repair; extracting this information combines multiple technologies such as token matching, static analysis, and large language model deviation prediction, judging the repair category from multi-dimensional repair semantics based on token matching and large language models, and judging the repair granularity from multi-dimensional semantics based on static analysis and large language models; finally, using a relational database to persist the classification information, running configuration information, and manual identification information. As Figure 1 shown, Figure 1 is a schematic flowchart of the method for constructing a defect repair dataset for defect categories and repair granularities in the embodiments of the present invention. The method of the present invention includes the following steps: Step 1) Input a configuration file identifying the project to be processed; the configuration file in step 1) needs to meet the following requirements: The configuration file needs to be in csv format. It is possible to specify the use of Git and SVN management or projects in Defects4j as the projects to be extracted. If a Git project is to be specified, the column names in the csv need to be "gitUrl, branch". The first column, gitUrl, specifies the address of the project's GitHub repository, and the second column, branch, specifies the specific branch of the project. If an SVN project is to be specified, the column names in the csv need to be "svnUrl, branch". The first column, svnUrl, specifies the pull address of the project's SVN repository, and the second column, branch, specifies the specific branch of the project. If a project in Defects4j is to be specified, the column names in the csv need to be "repoId, bugId". The first column, repoId, specifies the project number in the Defects4j dataset, and the second column, bugId, specifies the defect number in the project. These configuration files specify from which projects the defect datasets need to be extracted.

[0026] Step 2) For the heterogeneity of Defects4j, SVN, and Git managed projects, perform effectiveness analysis of the completed repairs for the projects to be processed and extract information on multi-repair semantics; as Figure 2 shown, Figure 2 This is a schematic flowchart of the process for the present invention to perform effectiveness analysis of the completed repairs for the projects to be processed and extract information on multi-repair semantics. The specific steps of step 2) include: Step 21) Parse the configuration file provided in step 1) and verify the correctness of the configuration file. The correctness verification includes: whether the url can be accessed, whether the repoId exists, and whether the bugId exists.

[0027] Step 22) According to whether the currently specified configuration file is a Git / SVN project or a Defects4j project, create a thread pool and start different parallel processing flows. If the configuration file specifies a Git / SVN project, the parallel processing flow specifically includes the following steps: Step 201) For a project managed by Git, obtain the project source code and git files through the git clone command and use git checkout to switch branches; for a project managed by SVN, directly obtain the project corresponding to the version number through svn checkout.

[0028] Step 202) Compile the project through maven, run the current test set, and record which tests can pass normally.

[0029] Step 203) For projects managed by Git, extract commit information from git log, and obtain the file information modified in this fix through git diff. If only a single file is modified, the specific modification information of the single file will also be extracted; for projects managed by SVN, obtain the commit information through svn log, and obtain the file information modified in this fix through svn diff. If only a single file is modified, the specific modification information of the single file will also be extracted.

[0030] Step 204) Aggregate the source code, test pass status, commit information, modified file information, and single file modification information to generate Git and SVN defect repair candidate datasets.

[0031] If the configuration file specifies a Defects4j project, the parallel processing flow specifically includes the following steps: Step 211) Extract the repaired source code from the Mysql database of Defects4j.

[0032] Step 212) Use defects4j export to extract the test pass status and repair category information.

[0033] Step 213) Compile and run the tests, and check against the existing information.

[0034] Step 214) Generate a Defects4j defect repair data candidate set.

[0035] Step 215) Finally, merge the two to generate the final candidate dataset.

[0036] Step 3) Determine the repair category from multi-repair semantics based on lemmatization matching and large language models; as Figure 3 shown, Figure 3 is a schematic flow diagram of the method for determining the repair category from multi-repair semantics based on lemmatization matching and large language models in an embodiment of the present invention. Step 3) specifically includes the following steps: Step 31) If the configuration file specifies a Defects4j project, directly obtain the defect classification.

[0037] Step 32) If the configuration file specifies a Git / SVN project, execute Step 33).

[0038] Step 33) Input the commit message collected in Step 2) into the large language model to determine whether this modification is a defect repair; if the large language model determines it is not, end the process; if the large model determines it is, continue to execute Step 34).

[0039] Step 34) Obtain the commit message after the first repair, input it into the large language model, and determine whether it is the same repair as before; if so, end the process and discard the code repair across commits; if not, continue to execute Step 35).

[0040] Step 35) Perform token matching classification on the commit message; if the token matching is successful, input the commit message and the token matching classification result into the large language model to obtain the confidence level; the token matching includes the following tokens: If the first type of keyword appears in the commit message, it will be directly determined as the corresponding category. If no keyword appears, but the complete word / phrase of the second type of keyword appears (i.e., both sides are word boundaries), it will also be assigned to the corresponding category.

[0041] Step 36) Examine the confidence level of the large language model for the token classification. If it is greater than 0.8, recognize the classification and record the classification; otherwise, discard the candidate example.

[0042] Step 37) If the token matching fails, input the commit message and the source code modification information into the large language model to calculate the deviation of each category; the deviation calculation formula for a single category is as follows where represents a specific example, y represents a category, is the set of all categories; and represents the confidence level of asking the large language model whether the example belongs to category y for the example .

[0043] Step 38) Through the formula , calculate the uniquely belonging category for the example, recognize and record the classification.

[0044] Step 4) Judge the repair granularity and repair granularity cutting from multi-dimensional semantics based on static analysis and the large language model; as Figure 4 shown, Figure 4 is a schematic flow diagram of the method for judging the repair granularity and repair granularity cutting from multi-dimensional semantics based on the static analysis method and the large language model. The specific steps of Step 4) include: Step 41) Examine the modified file list. If there are multiple source files modified, classify it as the file level and end the entire process of Step 4); otherwise, execute Step 42).

[0045] Step 42) Use static analysis method to parse a single file into SpoonAst, and examine whether the modifications in a file cross classes / functions; if they cross classes, classify them at the class level and end the entire process of Step 4); otherwise, execute Step 43).

[0046] Step 43) If they cross functions, classify them at the multi-function body level, and then execute Step 44) and Step 46).

[0047] Step 44) Use static analysis method to parse a single function into SpoonAst, and examine whether the modifications in a file are only a single operation at consecutive positions; the single operation refers to an operation that only includes one of addition, deletion, and modification; if so, classify it at the consecutive statement level and end the entire process of Step 4); otherwise, execute Step 45).

[0048] Step 45) Input the modified content into a large language model to determine whether this modification is the movement of a single statement block; if the confidence level is not less than 0.8, classify it at the consecutive statement level; otherwise, classify it at the single function body level; end the entire process derived from Step 4).

[0049] Step 46) Input the commit message of this cross-function modification into the LLM to examine the problems with the modification; if the result is the repair of multiple problems and the confidence level is not less than 0.8, execute the next step; otherwise, end the entire process derived from Step 4).

[0050] Step 47) Split this cross-function modification into different single-function modifications.

[0051] Step 48) Input the diff file of the single-function repair into the large model to determine whether it is a complete repair of a single problem; if the confidence level is less than 0.8, end the entire process derived from Step 4). If the confidence level is not less than 0.8, then regard this split as a single-function repair and execute Step 44).

[0052] Step 5) Identify similar repair structures in the same type of repairs based on clustering; as Figure 5 shown, Figure 5 is a schematic flow diagram of the method for identifying similar repair structures in the same type of repairs based on clustering. The specific process of Step 5) includes: Step 51) Parse the function code information before and after the repair into SpoonAst.

[0053] Step 52) Remove local variable definitions and normalize local variables to their corresponding positions.

[0054] Step 53) Unify the exchange rate and place constants on the left.

[0055] Step 54) The statement blocks of the unified conditional statements and loop statements are all located in Block.

[0056] Step 55) Replace the variable names, function names, and string contents named manually in the code with normalized characters calculated according to the structure.

[0057] Step 56) Generate a preprocessed diff file based on the AST.

[0058] The feature extraction process includes: Step 57) Calculate the word frequency and inverse word frequency in the code file.

[0059] Step 58) Multiply the word frequency and inverse word frequency of each word in the document to form a feature vector.

[0060] The clustering execution process includes: Step 59) Perform K-means clustering according to the cosine similarity to obtain the clustering labels.

[0061] Step 6) Persist the classification information, running configuration information, repair structure information, and manual identification information. The specific process of Step 6) includes: Step 61) Take out the code before and after the candidate commit and parse it into SpoonAst.

[0062] Step 62) Remove the local variable definitions and normalize the local variables to the corresponding positions.

[0063] Step 63) Unify the exchange rate and place the constants on the left.

[0064] Step 64) The statement blocks of the unified conditional statements and loop statements are all located in Block.

[0065] Step 65) Save the above rewritten results as before-and-after modification examples and store them in the specified path.

[0066] Step 66) Newly establish a Mysql database entry to save the before-and-after repair examples.

[0067] Step 67) Save the paths, commit messages, test passing situations, and running environment parameters of the before-and-after modification examples in the Mysql database.

[0068] Step 68) Save the repair classification labels, repair granularity labels, and clustering labels to the Mysql database.

[0069] The above are only the preferred embodiments of the present invention, but the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Any person skilled in the art, without departing from the principle and spirit of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a defect repair dataset for defect categories and repair granularity, characterized in that: The method comprises the following steps: Step 1) Enter the configuration file identifying the project to be processed; Step 2) In view of the heterogeneity of projects managed by Defects4j, SVN and Git, the effectiveness analysis of repairs and the information extraction of multi-dimensional repair semantics of the projects to be processed are completed; Step 3) Determine the repair category from the multi-dimensional repair semantics based on word unit matching and large language model; Step 4) Determine the repair granularity and repair granularity cutting from multi-dimensional semantics based on static analysis methods and large language models; Step 5) Identify similar repair structures in the same repair based on clustering; Step 6) Persist the classification information, operation configuration information, repair structure information and manual identification information.

2. The defect repair data set construction method according to claim 1, characterized in that: The configuration file in step 1 needs to be a file in csv format. The configuration file can specify a project in Github or Defects4j as the extracted project. If the configuration file is to specify a project in Github, the column names of the csv need to be "url, branch". The first column url specifies the address of the github repository of the project, and the second column branch specifies the specific branch of the project. If the configuration file is to specify a project in Defects4j, the column names of the csv need to be "repoId, bugId". The first column repoId specifies the project number in the Defects4j dataset, and the second column bugId specifies the defect number in the project. The configuration file specifies the projects from which the defect dataset needs to be extracted.

3. The defect repair data set construction method according to claim 1, characterized in that: The step 2) specifically includes the following steps: Step 21) Parse the configuration file provided in step 1) and verify the correctness of the configuration file; the correctness verification includes: whether the URL is accessible, whether the repoId exists, and whether the bugId exists; Step 22) Depending on whether the current configuration file specifies a Git / SVN project or a Defects4j project, a thread pool is created to start different parallel processing flows.

4. The defect repair data set construction method according to claim 3, characterized in that: In step 22), if the current configuration file specifies a Git / SVN project, the parallel processing flow includes the following steps: Step 201) For a project managed by Git, use the git clone command to obtain the project source code and git files, and use git checkout to switch branches; for a project managed by SVN, use svn checkout to directly obtain the project with the corresponding version number; Step 202) Compile the project through Maven, run the current test set, and record which tests can pass normally; Step 203) For projects managed by Git, the commit information is extracted from the git log, and the file information of this repair and modification is obtained through git diff. If only a single file is modified, the specific modification information of the single file is also extracted; for projects managed by SVN, the commit information is obtained through svn log, and the file information of this repair and modification is obtained through svn diff. If only a single file is modified, the specific modification information of the single file is also extracted; Step 204) Summarize the source code, test set passing status, commit information, modified file information, and modified single file information to generate Git and SVN defect repair candidate data sets.

5. The defect repair data set construction method according to claim 3, characterized in that: In step 22), if the current configuration file specifies a Defects4j project, the parallel processing flow includes the following steps: Step 211) Extract the repair source code from the Defects4j MySQL database; Step 212) Use defects4j export to extract test pass status and repair category information; Step 213) compile and run the test, and check with the existing information; Step 214) Generate a Defects4j defect repair data candidate set; Step 215) Finally, the two are merged to generate the final candidate data set.

6. The defect repair data set construction method according to claim 1, characterized in that: The step 3) specifically includes the following steps: Step 31) If the configuration file specifies a Defects4j project, the defect classification is obtained directly; Step 32) If the configuration file specifies a Git / SVN project, proceed to step 33); Step 33) Input the commit message collected in step 2) into the large language model to determine whether this modification is a bug fix; if the large language model determines that it is not, end the process; if the large model determines that it is, continue to step 34); Step 34) Get the commit message after the fix, input the large language model, and determine whether it is the same fix as the previous one; if so, end the process and abandon the cross-commit code fix; if not, continue to step 35); Step 35) Perform word-unit matching and classification in the commit message; if the word-unit matching is successful, input the commit message and the word-unit matching and classification results into the large language model to obtain the confidence level; Step 36) Check the confidence of the language model for the classification. If it is greater than 0.8, the classification is recognized and recorded; otherwise, the candidate sample is discarded; Step 37) If the word-unit matching fails, the commit message and source code modification information are input into the large language model to calculate the deviation of each category; the deviation calculation formula for a single category is as follows ,in represents a specific example, y represents a category, is the set of all categories; and For the sample Ask the large language model whether the sample is the confidence level of the input category y; Step 38) Use the formula , calculate the unique category to which the sample belongs, recognize and record the classification.

7. The method for constructing a defect repair data set according to claim 6, characterized in that: The word units in step 35) include the following word units when matching: If the first category of keywords appear in the commit message, it will be directly determined to be in the corresponding category; if no keywords appear, but the complete word / phrase of the second category of keywords appears, it will also be assigned to the corresponding category.

8. The defect repair data set construction method according to claim 1, characterized in that: The step 4) specifically includes: Step 41) Check the modified file list. If multiple source files are modified, classify them at the file level and end the entire process of step 4); otherwise, execute step 42); Step 42) Use static analysis methods to parse a single file as SpoonAst and examine whether the modification in a file is cross-class / cross-function; if it is cross-class, it is classified as class level and the entire process of step 4) is ended; otherwise, execute step 43); Step 43) At this time, this modification is classified as a multi-function body level, and steps 44) and 46) are executed in parallel); Step 44) Use static analysis method to parse a single function as SpoonAst, and check whether the modification in a file is a single operation of only continuous positions; the single operation means that the operation only includes one of addition, deletion, and modification; if so, classify it as a continuous statement level and end the entire process of step 4); otherwise, execute step 45); Step 45) Input the modified content into the large language model to determine whether this modification is the movement of a single sentence block; if the confidence is not less than 0.8, it is classified as a continuous sentence level; otherwise, it is classified as a single function body level; end the entire process derived from step 4); Step 46) Input the commit message of this cross-function modification into LLM to examine the modified issues; if the result is that multiple issues are fixed and the confidence level is not less than 0.8, execute step 47); otherwise, end the entire process derived from step 4); Step 47) Divide the cross-function modification into different single function modifications; Step 48) Input the diff file of the single function repair into the large model to determine whether it is a complete single problem repair; if the confidence is less than 0.8, end the entire process derived from step 4); if the confidence is not less than 0.8, regard this segmentation as a single function repair, and then execute the step 44).

9. The method for constructing a defect repair data set according to claim 1, characterized in that: The step 5) includes a data preprocessing process, a feature extraction process and a clustering execution process. The data preprocessing process includes: Step 51) Parse the function code information before and after the repair into SpoonAst; Step 52) Remove local variable definitions and normalize local variables to corresponding positions; Step 53) Normalize the exchange rate and put the constant on the left; Step 54) The statement blocks of the unified conditional statements and loop statements are all located in Block; Step 55) Replace the manually named variable names, function names, and string contents in the code with normalized characters calculated according to the structure; Step 56) Generate a preprocessed diff file based on AST; The feature extraction process includes: Step 57) Calculate the word frequency and inverse word frequency in the code file; Step 58) Multiply the word frequency and inverse word frequency of each word in the document to form a feature vector; The clustering execution process includes: Step 59) Perform K-means clustering based on cosine similarity to obtain cluster labels.

10. The defect repair data set construction method according to claim 1, characterized in that: The step 6) specifically includes: Step 61) Take out the code before and after the candidate commit is modified, and parse it into SpoonAst; Step 62) Remove local variable definitions and normalize local variables to corresponding positions; Step 63) Normalize the exchange rate and put the constant on the left; Step 64) The statement blocks of the unified conditional statements and loop statements are all located in Block; Step 65) Save the above rewriting results as samples before and after modification and store them in the specified path; Step 66) Create a new Mysql database entry to save the samples before and after the repair; Step 67) Save the path, commit message, test pass status, and operating environment parameters of the sample before and after the modification in a MySQL database; Step 68) Save the repaired classification labels, repaired granularity labels, and cluster labels to the MySQL database.