Submission method for automatically positioning introduced defects based on large language model
By constructing prompt words and expanding the context, the evaluation method of the large language model is optimized, which solves the problems of noise interference and context ignoring in the existing technology, realizes more efficient defect localization and analysis, and improves the recognition accuracy of the tool.
Patent Information
- Application Number
- CN202510944793.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
AI Technical Summary
Existing defect localization methods based on large language models are easily affected by noise and ignore contextual information when processing defect-fixing submissions, making it difficult to accurately locate the submissions that introduce defects. Furthermore, they perform poorly when handling multiple files and complex changes.
By constructing prompt words, shuffling the submission order, and using a large language model to analyze the root cause of defects and related documents, the context is expanded, the evaluation model's judgment ability is optimized, and context enhancement or ranking recognition methods are used to locate the submissions that introduce defects.
It improves the accuracy and efficiency of defect localization, reduces development costs, enhances the tool's identification precision, and provides reliable defect analysis recommendations.
Smart Images

Figure CN120848854A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of program analysis and defect localization technology, and in particular to a method for automatically locating the submission of introduced defects based on a large language model. Background Technology
[0002] Modern software development heavily relies on version control systems to manage source code and track changes. Version control systems use commits as the basic unit, with each commit representing a snapshot of the entire project. Among all types of commits, there are bug fix commits, which contain a wealth of information about the bug and are often used by developers to analyze the root cause of the bug and find the commit that introduced it. Unfortunately, many bug fix commits contain numerous code changes unrelated to the cause of the bug. This irrelevant code can hinder bug analysis and reduce the accuracy of tools for finding bug-inducing commits. While previous research has made some progress, several limitations remain. First, these methods focus only on the modified lines of code, ignoring the context of the entire patch. Previous research has shown that context (including unmodified lines near the modified lines) provides crucial information for models to understand the code. Sometimes, it may be an unmodified line of code that causes the bug, rather than the modified line. Second, these methods ignore the commit messages of bug fix commits. Typically, commit messages contain important information about why these changes were made, much of which describes how the bug arose and how the commit fixed it. This information is crucial for understanding the commit content and accurately locating the bugging line of code. Third, these methods assume that only deleted lines of code cause the defect, and therefore are not suitable for fixes that only contain added lines of code. Finally, methods for selecting the final commit that introduces the defect from a set of candidate commits often rely on heuristics, such as commit date or number of lines modified. These assumptions may not apply to all scenarios. Ideally, the final commit that introduces the defect should be determined based on the root cause of the defect and the content of the candidate commits.
[0003] The emergence of Large Language Models (LLMs) offers an opportunity to address the aforementioned limitations. Previous research has shown that LLMs can effectively understand code change and commit messages. A fundamental improvement is to leverage LLMs to analyze the root causes of defects and identify defective lines of code based on code change and commit messages. The improved approach traces these defective lines of code to obtain a set of candidate commits and asks the LLM to select the commit that introduced the defect. However, LLMs also have limitations. First, they perform poorly when handling complex fix commits involving numerous changes to multiple files and functions. These commits often contain noise unrelated to fixing the defect, thus impacting the performance of the LLM. Second, when asking the LLM to determine whether a commit contains a defect, the provided context must be carefully considered. Excessively long contexts can degrade performance, while excessively short contexts may miss crucial information necessary for the LLM to understand the code. Third, many types of defects are beyond the comprehension of LLMs, making it difficult for them to determine their presence in commits. Treating these defects the same as those that the LLM can understand negatively impacts overall performance. For example, if a large language model is determined to understand a defect and identify its presence in a submission, it can be used to select the final submission that introduces the defect from a set of candidate submissions; otherwise, it cannot. Finally, more information is needed to help the large language model determine the presence of a defect. The root cause of the defect and the content of the submission are often insufficient for this purpose.
[0004] Two examples are used here to illustrate the potential and challenges of large language models in this scenario. The first example involves fixing a defect commit eed6e41813d in Linux. The prompt "Based on the content of the defect fix commit, analyze the root cause of the defect and output the line of code that caused the defect" and the content of the defect fix commit are input into the large language model. The large language model successfully predicts the occurrence of the defect because the function list_for_each_entry_safe fails to protect the list when multiple threads add or delete nodes in parallel. It identifies lines 8 and 9 as the defective lines of code and excludes line 3. Tracing these two lines yields two candidate commits that introduced the defect: eb7fbc9fb11 and 25b4e70dcce. The large oracle model is then used to determine which candidate commit introduced the defect. The large oracle model finds that commit eb7fbc9fb11 introduced line 9, but only modified the second parameter of the dev_info function, which did not affect the existence of the defect. Therefore, submission eb7fbc9fb11 was excluded, and 25b4e70dcce was identified as the final submission to introduce the defect.
[0005] The second example involves commit c5153331c in Accumulo, which modifies four files, introducing 29 lines of insertion and deleting 8 lines. According to the commit message, the defect arises because the program fails to call the `getInstanceId` function to enforce a valid instance name, and the defect is only related to the `instanceName` variable. If the entire patch is directly input into the LLM and asked to identify the line of code causing the defect, it incorrectly points to the `@Test(expected=RuntimeException.class)` statement in another file called `ZooKeeperInstanceTest.java`. If other files are excluded and only the changes in the relevant file `ZooKeeperInstance.java` are input into the LLM, the LLM still fails to output the correct line of code. Specifically, if the commit message and the original patch content are provided to the LLM, it still incorrectly identifies line 14 as the defective line of code. This is due to insufficient context. According to the commit message, the defect is related to the `instanceName` variable. However, in the original patch content, lines 12 through 19, the only line of code related to the `instanceName` variable is line 15, which is used to fix the defect. The full contents of the ZooKeeperInstance constructor include the statement `this.instanceName = clientConf.get(ClientProperty.INSTANCE_NAME)`, which is related to the instanceName variable. However, this statement is not shown in the original patch. By providing extended context (lines 1 through 19), which includes the entire contents of the constructor, LLM correctly recognizes the code statement on line 9. Summary of the Invention
[0006] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for automatically locating and submitting defective content based on a large language model.
[0007] The objective of this invention is achieved through the following technical solution: a method for automatically locating defective submissions based on a large language model, comprising the following steps:
[0008] S1. Use a large language model to analyze the commits that fix the defects to obtain the root cause of the defects and the files related to fixing the defects, and remove all files that are not related to fixing the defects.
[0009] S2. Based on the root cause of the defect and the documents related to fixing the defect, evaluate whether the large language model can determine whether there is a defect in the submission for fixing the defect through context expansion and optimization. If the large language model can determine it, the context-enhanced identification method is used to locate the submission that introduced the defect; otherwise, the ranking-based identification method is used to locate the submission that introduced the defect.
[0010] Further, step S1 specifically includes:
[0011] First, prompt words for the large language model are constructed. Then, the order of multiple bug fix submissions is shuffled. The shuffled bug fix submissions and prompt words are then used as input to the large language model. The large language model analyzes patches to summarize the changes in the bug fix submissions and the relationships between these changes. Based on the changes and the bug fix submission information, it outputs the root cause of the bug and the names of the files related to the bug fix. The large language model is run three times for the same problem, shuffling the order of multiple bug fix submissions each time. Based on the outputs of the three runs, the root cause of the bug and the names of the files related to the bug fix are obtained, and all files irrelevant to the bug fix are removed.
[0012] Furthermore, step S2 specifically includes the following sub-steps:
[0013] S21. For changes submitted to fix defects, expand the context of each change based on the files related to the defect fix to obtain the expanded context;
[0014] S22. Construct a prompt word, and use the prompt word, the root cause of the defect, and the expanded context as input to the large language model to obtain the code statement that caused the defect and its corresponding reason, as well as the code statement that fixed the defect and its corresponding reason.
[0015] S23. Based on the files related to fixing the defect and the code statements that caused the defect, optimize the extended context to obtain the current defect-fixing commit and the optimized context corresponding to the previous commit.
[0016] S24. Construct prompt words based on the code statements that cause defects and their corresponding reasons, as well as the code statements that fix defects and their corresponding reasons. Based on the prompt words, the root cause of the defects, and the optimized context corresponding to the two defect-fixing submissions, evaluate whether the large language model can determine whether there are defects in the defect-fixing submissions. If the large language model can determine this, then use the context-enhanced recognition method to locate the submission that introduces the defect; otherwise, use the ranking-based recognition method to locate the submission that introduces the defect.
[0017] Further, step S21 specifically includes:
[0018] Changes in bug fix submissions include modifications to functions and changes made outside of functions;
[0019] For each modified function, its context is expanded by extracting the complete contents of its error version and the repair version to generate the differences between the two versions, thus obtaining the expanded context; wherein, the error version and the repair version are obtained based on the files related to the defect repair obtained in step S1;
[0020] For modified lines outside functions, the context is expanded by extracting the three unmodified lines above and below the modified line in the repaired version of the program to obtain the expanded context.
[0021] Furthermore, step S23 specifically includes:
[0022] First, let's assume the current commit for fixing the defect is C. gix Then the previous commit before the commit that fixes the defect is Then, based on the code statement that caused the defect obtained in step S22, the commit obtained in step S1... Extract the code statements that cause the defect from the corresponding files related to the defect fix, and then sort these code statements by their line numbers {l1,l2,…,l n Sort the rows in ascending order, where l1 is the smallest row number. n The largest line number; secondly, l min Defined as l1-N, l max Defined as l n +N, where N is a constant initially set to 3, and the value of N is gradually increased to ensure row l min and l max Able to map to commit The corresponding line in the file related to fixing the defect; then in the commit. Extract from line number l of the corresponding file related to defect repair. min to line number l max The content between them forms the submission. The corresponding optimized context; finally, line l min and l max Mapping to commit C fix The corresponding line number l in the file related to fixing the defect ′ min and l ′ max And in submitting C fix Extract from line number l of the corresponding file related to defect repair. ′min to line number l ′ max The content between them forms the submission C. fix The corresponding optimized context.
[0023] Further, step S24 specifically includes:
[0024] First, construct prompt words based on the code statements that cause the defects and their corresponding reasons, as well as the code statements that fix the defects and their corresponding reasons;
[0025] Then, submit the warning message, the root cause of the defect, and the current commit to fix the defect (C). fix The corresponding optimized context is used as input to the large language model to obtain the output of the large language model, which is then submitted as C. fix The result of the judgment on whether the corresponding context contains a defect; the prompt word, the root cause of the defect, and the previous submission to fix the defect. The corresponding optimized context is used as input to the large language model to obtain the output of the large language model. The result of the judgment on whether there is a defect in the corresponding context;
[0026] Secondly, a judgment is made based on two judgment results output by the large language model: if the large language model considers the submission... The corresponding context contains a defect, and submission C... fix If the corresponding context does not contain defects, it is assumed that the large language model can correctly identify the fixed version and the error version of the program, and the context-enhanced identification method is used to locate the submission that introduces the defect; otherwise, it is assumed that the large language model cannot correctly identify the fixed version and the error version of the program, and the ranking-based identification method is used to locate the submission that introduces the defect.
[0027] Furthermore, the method of using context-enhanced identification to locate submissions that introduce defects specifically includes:
[0028] First, extract all the code statements that cause the defect;
[0029] Then, trace the code statements that caused the defects, find all the commits that introduced these code statements to fix the defects, obtain a set of candidate commits, and sort them in descending order according to the commit date to form a candidate commit list {C1, C2, ..., C...}. j ,…,C i}, where C j Let j represent the j-th candidate submission, and i represent the total number of candidate submissions;
[0030] Then, an optimized context is generated for each candidate submission, resulting in each candidate submission C. j and its previous submission The corresponding optimized context;
[0031] Secondly, the large language model is used to examine each candidate submission in the candidate submission list from index 1 to i. If an index f can be found such that the large language model considers candidate submission C to be... f The corresponding context contains defects, and the candidate commits If the corresponding context does not contain defects, then candidate submission C is selected. f Submissions that introduce defects are identified; conversely, submissions that introduce defects are identified using a ranking-based identification method.
[0032] Furthermore, the method of using ranking-based identification to locate submissions that introduce defects specifically includes:
[0033] First, the constructed prompt words, the submission to fix the defect, the root cause of the defect, and the files related to fixing the defect are input into the large language model to obtain the code statements in the submission to fix the defect that caused the defect.
[0034] Then, using a listwise ranking algorithm based on a large language model, a prompt word is constructed, and the prompt word, the code statement that causes the defect, and the root cause of the defect are input into the large language model. The large language model sorts these code statements according to the correlation between the code statement that causes the defect and the root cause of the defect.
[0035] Secondly, for each file, retrieve the top N code statements, trace back to the commits that introduced them, and add these commits to the candidate commit list;
[0036] Finally, these candidate submissions are sorted according to their submission dates, with the most recent submission being the one that introduced the defect.
[0037] The beneficial effects of this invention are as follows: This invention overcomes the limitations of traditional methods by using a large language model to filter out files that are irrelevant to fixing defects; This invention can locate the root cause of defects based on more context, providing reliable suggestions for developers, reducing the time and cost spent by developers, and improving maintenance efficiency and software quality; This invention uses a large language model to locate the commits that introduce defects by removing noise in the commits, which helps to improve the accuracy of tools that identify commits that introduce defects. Attached Figure Description
[0038] Figure 1 A flowchart illustrating the architecture of the method for automatically locating defects based on a large language model, as described in this invention.
[0039] Figure 2 This is an example diagram illustrating the defects of large language models and context expansion in this invention;
[0040] Figure 3 This is an example diagram of the context-enhanced recognition of the present invention;
[0041] Figure 4 This is a flowchart illustrating the submission process for locating defects using the context-enhanced identification method of this invention. Detailed Implementation
[0042] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0043] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0044] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0045] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0046] This invention aims to locate the submission that causes the defect by fixing the defect submission, and consists of three stages, such as... Figure 1As shown. The first stage is the preparation stage, the second stage is the evaluation stage, and the third stage is the commit location stage. The main typical application scenario of this invention is to provide suggestions for developers to analyze and fix defective commits, and to improve the accuracy of tools that find commits that introduce defects. In the preparation stage, LLM is used to remove files unrelated to fixing defects from the commits. In the evaluation stage, more context is provided to the LLM to assess whether it can accurately understand the current defect. In the commit location stage, based on the evaluation results, different methods are used to locate the commit that caused the defect. This invention improves the efficiency of developers in analyzing code changes and helps to improve the accuracy of tools that find commits that introduce defects.
[0047] It's important to note that a bug fix commit is a code change submitted by a developer through a version control system (such as Git) to fix a discovered software defect (bug). Bug fix commits typically clearly state their purpose in the commit message, for example, using keywords such as `fix:`, `bugfix:`, or `Fixes#123` (associated with the issue ID in the issue tracking system). Example: "plaintext"
[0048] Fix: Fixed the issue of password verification failing during user login.
[0049] Root cause: The password hash comparison logic is flawed; null values were not handled.
[0050] Solution: Add null value checks and correct the hash comparison logic.
[0051] Fixes#456 (linked to Issue#456).
[0052] The present invention provides a method for automatically locating and submitting defective content based on a large language model, such as... Figure 1 As shown, the specific steps include:
[0053] S1. Use LLM to analyze the commits that fix the defects to find the root cause of the defects and the files related to fixing the defects, and remove all files that are not related to fixing the defects.
[0054] Specifically, first, construct the prompt words for the LLM, for example... Figure 1 The statement "Based on the summary and submission information, you need to:" indicates that...
[0055] 1. Analyze the root cause of this bug in detail.
[0056] 2. Output the filename most likely to cause the bug. Then, shuffle the order of multiple bug fix commits. Use the shuffled bug fix commits and the warning message as input to the LLM. Following the Chain of Trace (CoT) concept, the LLM analyzes patches to summarize the changes in the bug fix commits and the relationships between these changes. Based on these changes and the bug fix commit information, it outputs the root cause of the bug and the filenames related to the bug fix. Run the LLM three times for the same issue, shuffling the order of multiple bug fix commits each time. Based on the three outputs, obtain the root cause of the bug and the filenames related to the bug fix, and remove all files unrelated to the bug fix. The final set of filenames related to the bug fix is the union of the three outputs.
[0057] It should be noted that, to improve performance when processing large defect fix submissions containing multiple modified files, this invention employs two additional methods. First, by shuffling the order of multiple defect fix submissions, it ensures that each file has an equal chance of being identified as related to the root cause. Previous research has shown that LLM tends to ignore the middle content when processing long texts. Second, LLM is run three times for the same issue, each time shuffling the order of multiple defect fix submissions, i.e., shuffling the order of patches. This method is similar to a voting system; however, this invention does not only consider files with the most votes but also adopts a more conservative strategy: if a filename appears in any LLM output, it is considered related to the root cause of the defect. This strategy helps minimize the risk of missing important files. Therefore, LLM will have three outputs, which may be the same, different, or partially identical, but as long as a filename appears in any output, it will be retained and considered the root cause of the defect. Through this step, the root cause of the defect can be obtained, and all files irrelevant to fixing the defect can be filtered out.
[0058] S2. Based on the root cause of the defect obtained in step S1 and the files related to fixing the defect, evaluate whether the LLM can determine whether there is a defect in the submission to fix the defect through context extension and optimization. If the LLM can determine it, the context enhancement identification method is used to locate the submission that introduced the defect; otherwise, the ranking-based identification method is used to locate the submission that introduced the defect.
[0059] It should be noted that in the current step S2, the results obtained in step S1 need to be further processed in order to evaluate the LLM's understanding of the current defect.
[0060] S21. For the changes submitted in the defect repair submission, expand the context of each change based on the files related to the defect repair obtained in step S1 to obtain the expanded context.
[0061] Specifically, defect fix commits typically do not contain the complete content of the modified functions. However, partial content of functions in defect fix commits may hinder the LLM from understanding the functions and their modifications. Therefore, a context expansion process is needed to provide sufficient context for the LLM. Specifically, the changes in defect fix commits include function modifications and changes outside the functions. For each modified function, its context is expanded by extracting the complete content of its faulty and fixed versions of the program to generate differences between the two versions. The faulty version refers to the program version containing the defect before the defect fix commit, and the fixed version refers to the program version with the defect fixed after the defect fix commit. The faulty and fixed versions are obtained from the files related to the defect fix obtained in step S1. For changes outside the functions, their context is expanded by extracting the three lines of unmodified code above and below the modified line in the fixed version. Figure 2 As shown, an example of a context expansion process is provided, where gray code represents newly added parts, and the rest comes from the original patch; in this example, the complete function content is provided for the large language model, and the problematic code this.instanceName = clientConf.get(ClientProperty.INSTANCE_NAME) does not appear in the original patch, but in the expanded context, which shows the necessity of context expansion.
[0062] S22. Construct a prompt word, and use the prompt word, the root cause of the defect obtained in step S1, and the expanded context obtained in step S21 as input to the LLM to obtain the LLM output code statement that caused the defect and its corresponding reason, as well as the code statement that fixed the defect and its corresponding reason.
[0063] It should be noted that, because the LLM needs to identify the code statements that cause the defect and provide reasons, and also needs to identify the code statements that fix the defect and provide reasons, the prompt words need to be constructed to address both of these aspects. For example, such as... Figure 1 The prompt message displayed is: "You need:
[0064] 1. Explain how the bug occurred, including the code statements that might have caused it and their reasons.
[0065] 2. Explain how the bug was fixed, including the code statements that fixed the bug and their reasons. Then, input the constructed hints, the root cause of the defect obtained in step S1, and the expanded context obtained in step S21 into the LLM to obtain the LLM output. From the output, we can obtain the code statements that caused the defect and their corresponding reasons, as well as the code statements that fixed the defect and their corresponding reasons. Notably, when selecting the code statements that caused the defect and those that fixed the defect, the Big Oracle model is not limited to selecting statements from the deleted code lines; it can select any code statement from the expanded context.
[0066] S23. Based on the files related to fixing the defect obtained in step S1 and the code statements that caused the defect obtained in step S22, optimize the extended context obtained in step S21 to obtain the current defect-fixing submission and the optimized context corresponding to the previous submission.
[0067] Specifically, let's first assume that the current commit for fixing the defect is C. fix Then the previous commit before the commit that fixes the defect is Then, based on the code statement that caused the defect obtained in step S22, the commit obtained in step S1... Extract the code statements that cause the defect from the corresponding files related to the defect fix, and then sort these code statements by their line numbers {l1,l2,…,l n Sort the rows in ascending order, where l1 is the smallest row number. n The largest line number; secondly, l min Defined as l1-N, l max Defined as l n +N, where N is a constant initially set to 3, and the value of N is gradually increased to ensure row l min and l max Able to map to commit The corresponding line in the file related to fixing the defect, i.e., line l min or line l max Unable to map to commit In the corresponding lines of the files related to fixing the defects, the value of N needs to be gradually increased; then, in the commit... Extract from line number l of the corresponding file related to defect repair. min to line number l max The content between them forms the submission. The corresponding optimized context; finally, line l min and l max Mapping to commit C fix The corresponding line number l in the file related to fixing the defect ′min and l ′ max And in submitting C fix Extract from line number l of the corresponding file related to defect repair. ′ min to line number l ′ max The content between them forms the submission C. fix The corresponding optimized context.
[0068] It's important to note that before assessing the LLM's ability to understand the current defect—that is, before evaluating whether the LLM can determine if a defect exists in the commit that fixes it—the extended context needs to be optimized to obtain an optimized context. This step is necessary because the extended context may contain much content unrelated to the defect. For example, the extended context might include an entire function with hundreds of lines, of which only a few lines are relevant to the defect. Directly providing the extended context to the LLM might weaken its ability to determine if a defect exists in the commit that fixes it.
[0069] S24. Based on the code statement that caused the defect and its corresponding reason obtained in step S22, and the code statement that repaired the defect and its corresponding reason, construct a prompt word. Based on the prompt word, the root cause of the defect obtained in step S1, and the optimized context corresponding to the two defect repair submissions obtained in step S23, check whether the LLM can determine whether there is a defect in the defect repair submission. If the LLM can determine it, the context enhancement recognition method is used to locate the submission that introduced the defect; otherwise, the ranking-based recognition method is used to locate the submission that introduced the defect.
[0070] Specifically, after obtaining the current commit for fixing the defect and the optimized context of its previous commit, the process begins to check whether the LLM can determine whether a defect exists in the commit for fixing the defect. First, a hint word is constructed based on the code statement causing the defect and its corresponding reason obtained in step S22, as well as the code statement for fixing the defect and its corresponding reason. Then, this hint word, the root cause of the defect obtained in step S1, and the current commit C for fixing the defect obtained in step S23 are used to... fix The corresponding optimized context is used as input to the LLM to obtain the LLM output submission C. fix The judgment result of whether the corresponding context contains a defect; the submission of the prompt word, the root cause of the defect obtained in step S1, and the previous defect repair obtained in step S23. The corresponding optimized context is used as input to the LLM to obtain the LLM output submission. The judgment result is whether the corresponding context contains a defect. Secondly, based on the two judgment results output by the LLM, a judgment is made, and then, according to the LLM evaluation result, different methods are used to locate the submission that introduced the defect: if the LLM considers the submission... The corresponding context contains a defect, and LLM considers the submission of C to be flawed. fix If the corresponding context does not contain defects, it is assumed that LLM can correctly identify the fixed version and the faulty version of the program, and then the context-enhanced identification method is used to locate the commit that introduced the defect; otherwise, it is assumed that LLM cannot correctly identify the fixed version and the faulty version of the program, and then the ranking-based identification method is used to locate the commit that introduced the defect.
[0071] Furthermore, given a commit that fixes a defect, after verifying that the LLM can understand the defect through the extended and optimized context, a context-enhanced identification method is used to locate the commit that introduced the defect. The specific process is as follows: Figure 4 As shown, the specific steps include: First, through steps S1, S21, and S22, all code statements that cause a given defect are obtained corresponding to the commit that fixes the defect; that is, all code statements that cause defects are extracted from step S22. Then, these code statements that cause defects are traced to find all commits that introduce these code statements to fix the defect, resulting in a set of candidate commits. These code statements that cause defects can be traced using version control tools and sorted in descending order according to the commit date to form a candidate commit list {C1, C2, ..., C...}. j ,…,C i}, where C j Let represent the j-th candidate commit, and i represent the total number of candidate commits. Then, in step S23, an optimized context is generated for each candidate commit, resulting in C for each candidate commit. j and its previous submission The corresponding optimized context. Next, through step S24, the LLM is used to examine each candidate commit in the candidate commit list from index 1 to i. If an index f can be found such that the LLM considers candidate commit C... f The corresponding context contains defects, and the candidate commits If the corresponding context does not contain defects, then candidate submission C is selected. f This means that the submission that introduces the defect is considered defective; conversely, if no such candidate submission is found, the system conservatively reverts to the ranking-based identification method to locate the submission that introduces the defect.
[0072] Figure 3An example is provided to demonstrate the workflow of context-enhanced identification. First, all code statements that cause defects are extracted. Then, all defective code statements are traced back, identifying two candidate commits: candidate commit eb7fbc9fb11 (denoted as C1) and candidate commit 25b4e70dcce (denoted as C2). These are then sorted in descending order by commit date, resulting in a candidate commit list {C1, C2}. First, commits C1 and C2 are examined. The corresponding optimized context, following the steps above, requires LLM to determine whether the optimized contexts for these two versions contain defects. LLM identifies commit C1 and... All versions of the program contain defects; therefore, this candidate commit C1 is not the commit that introduces a defect. Next, we examine commits C2 and... Follow the same steps. Here, C2 is the initial commit for the imported file, therefore, The context is empty. LLM found that commit C2 contained a defect, while commit... The context contained no code statements and therefore no defects. Therefore, commit C2 was ultimately determined to be the commit that introduced the defect.
[0073] Furthermore, the ranking-based identification method addresses the issue that LLMs cannot fully understand defects and cannot determine the existence of defects in the program. The ranking-based identification method locates the commits that introduce defects, specifically including: First, the constructed hint words, the commit that fixes the defect, the root cause of the defect obtained in step S1, and the files related to fixing the defect are input into the LLM to obtain the code statements in the commit that cause the defect. At this stage, only the hint words, the commit that fixes the defect, the root cause of the defect obtained in step S1, and the files related to fixing the defect are input; no expanded context is provided because if the LLM cannot understand the defect, additional context would actually degrade its performance. Then, using an LLM-based listwise ranking algorithm, a hint word is constructed, and this hint word, the code statement that causes the defect, and the root cause of the defect obtained in step S1 are input into the LLM. The LLM ranks these code statements according to the correlation between the code statement that causes the defect and the root cause of the defect obtained in step S1. Second, for each file, the top N code statements are retrieved, the corresponding commits that introduced them are traced, and these commits are added to the candidate commit list. Finally, these candidate commits are sorted according to the commit date, and the most recent commit is the commit that introduced the defect.
[0074] In summary, the method described in this invention can automatically locate commits that introduce defects, provide reliable suggestions to developers when they analyze the commits, reduce the time and cost spent by developers, and improve maintenance efficiency and software quality.
[0075] The purpose of this invention is to locate commits that introduce defects. To this end, this invention selects three datasets for evaluation: the DS_LINUX dataset, the DS_GITHUB dataset, and the DS_APACHE dataset. These three datasets are based on publicly available datasets of commits that fix defects and datasets of commits that introduce defects.
[0076] The method described in this invention was tested against baseline methods such as B-SZZ, AG-SZZ, MA-SZZ, R-SZZ, L-SZZ, RA-SZZ, and Neural-SZZ on the collected dataset. Precision, recall, and F1-score were used to evaluate the method's effectiveness in locating defective submissions. Tables 1 and 2 illustrate the localization performance of the method described in this invention, with Table 1 based on a C language dataset and Table 2 based on a Java language dataset. The method described in this invention (referred to as LLM4SZZ) achieved the best results across all metrics. LLM4SZZ was more accurate than all baseline methods in identifying defective submissions, improving precision by 1.4% to 6.9%. Furthermore, LLM4SZZ also achieved a significant improvement in F1-score, increasing it by 8.9% to 16.4% compared to the best baseline method. LLM4SZZ also showed more consistent and stable performance across the three datasets. While improving accuracy and F1-score, LLM4SZZ did not significantly sacrifice recall.
[0077] Table 1: The effect of automated defect localization on submission (C language)
[0078]
[0079] Table 2: The effect of submitting defects introduced by automated defect location (Java language)
[0080]
[0081] To better evaluate why LLM4SZZ outperforms other methods, the evaluation results were manually examined. In this embodiment, the effectiveness of key components of LLM4SZZ was investigated. In LLM4SZZ-raw, the most basic setup was implemented. First, LLM analysis was used to identify the root cause of the defect, and then the statements leading to the defect were located based on the root cause. Finally, these defect-leading statements were traced back, the commits that introduced them were identified, and these commits were marked as defect-introducing commits. In LLM4SZZ-r, context-enhanced evaluation was removed, and a ranking-based identification method was applied to all test cases. The LLM4SZZ-re variant is built on top of LLM4SZZ-r, providing an extended LLM context in the ranking-based identification, instead of the original patch. In LLM4SZZ-c, context-enhanced evaluation was also removed, but context-enhanced identification was used in all scenarios. The LLM4SZZ-h variant, based on LLM4SZZ-c, omits the hints and only provides the context of the commit and the root cause of the defect when requesting LLM to determine whether a given commit contains a defect. This allows for the evaluation of the contribution of the hints to LLM's ability to identify the presence of defects. The performance of LLM4SZZ and its variants in identifying defect-introducing commits is shown in Table 3.
[0082] Table 3: Performance evaluation results of LLM4SZZ and its variants
[0083]
[0084] As shown in Table 3, the LLM4SZZ-raw variant performs poorly, indicating that directly applying LLM to the SZZ algorithm does not significantly improve performance. For example, LLM4SZZ-raw only improves the F1 score by 4.1% relative to the best baselines R-SZZ and B-SZZ on the DS_LINUX dataset, and its performance on the DS_APACHE dataset is far inferior to Neural-SZZ.
[0085] The LLM4SZZ-r variant employs a ranking-based recognition method across all scenarios, outperforming LLM4SZZ-raw on the F1 score, with improvements ranging from 7.2% to 21.6% across the three datasets. In contrast, LLM4SZZ-re underperformed LLM4SZZ-r on all metrics, indicating that providing additional context weakens performance if the LLM cannot understand bugs.
[0086] Table 3 also shows that LLM4SZZ-c, employing context-enhanced recognition in all test cases, improved accuracy compared to LLM4SZZ-r. However, in some test cases, LLM4SZZ-c may produce null results, leading to lower recall. Furthermore, omitting prompts in context-enhanced recognition has a significant negative impact on both accuracy and recall; LLM4SZZ-h performed worse than LLM4SZZ-c across all metrics.
[0087] Overall, LLM4SZZ outperforms all other baselines on the F1-score, highlighting the performance improvement achieved by combining ranking-based recognition with context-enhanced recognition. This also demonstrates the effectiveness of context-enhanced evaluation, which assesses the capabilities of LLM and identifies suitable recognition methods to effectively handle diverse test cases. This invention is the first to employ a large language model to locate defect-introducing submissions. Accordingly, its evaluation results show that LLM4SZZ can more effectively identify defect-introducing submissions.
[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for automatically locating defective submissions based on a large language model, characterized in that, The following steps are involved: S1. Use a large language model to analyze the commits that fix the defects to obtain the root cause of the defects and the files related to fixing the defects, and remove all files that are not related to fixing the defects. S2. Based on the root cause of the defect and the documents related to fixing the defect, evaluate whether the large language model can determine whether there is a defect in the submission to fix the defect through context expansion and optimization. If the large language model can determine it, then use the context enhancement recognition method to locate the submission that introduced the defect. Conversely, a ranking-based identification method is used to locate submissions that introduce defects.
2. The method for automatically locating and submitting defective data based on a large language model according to claim 1, characterized in that, Step S1 specifically includes: First, prompt words for the large language model are constructed. Then, the order of multiple bug fix submissions is shuffled. The shuffled bug fix submissions and prompt words are then used as input to the large language model. The large language model analyzes patches to summarize the changes in the bug fix submissions and the relationships between these changes. Based on the changes and the bug fix submission information, it outputs the root cause of the bug and the names of the files related to the bug fix. The large language model is run three times for the same problem, shuffling the order of multiple bug fix submissions each time. Based on the outputs of the three runs, the root cause of the bug and the names of the files related to the bug fix are obtained, and all files irrelevant to the bug fix are removed.
3. The method for automatically locating and submitting defective data based on a large language model according to claim 1, characterized in that, Step S2 specifically includes the following sub-steps: S21. For changes submitted to fix defects, expand the context of each change based on the files related to the defect fix to obtain the expanded context; S22. Construct a prompt word, and use the prompt word, the root cause of the defect, and the expanded context as input to the large language model to obtain the code statement that caused the defect and its corresponding reason, as well as the code statement that fixed the defect and its corresponding reason. S23. Based on the files related to fixing the defect and the code statements that caused the defect, optimize the extended context to obtain the current defect-fixing commit and the optimized context corresponding to the previous commit. S24. Construct prompt words based on the code statements that cause defects and their corresponding reasons, as well as the code statements that fix defects and their corresponding reasons. Based on the prompt words, the root cause of the defects, and the optimized context corresponding to the two submissions to fix defects, evaluate whether the large language model can determine whether there are defects in the submissions to fix defects. If the large language model can determine this, then use the context enhancement recognition method to locate the submission that introduces the defects. Conversely, a ranking-based identification method is used to locate submissions that introduce defects.
4. The method for automatically locating and submitting defective content based on a large language model according to claim 3, characterized in that, Step S21 specifically includes: Changes in bug fix submissions include modifications to functions and changes made outside of functions; For each modified function, its context is expanded by extracting the complete contents of its error version and the repair version to generate the differences between the two versions, thus obtaining the expanded context; wherein, the error version and the repair version are obtained based on the files related to the defect repair obtained in step S1; For modified lines outside functions, the context is expanded by extracting the three unmodified lines above and below the modified line in the repaired version of the program to obtain the expanded context.
5. The method for automatically locating and submitting defective content based on a large language model according to claim 3, characterized in that, Step S23 specifically includes: First, let's assume the current commit for fixing the defect is C. fix Then the previous commit before the commit that fixes the defect is Then, based on the code statement that caused the defect obtained in step S22, the commit obtained in step S1... Extract the code statements that cause the defect from the corresponding files related to the defect fix, and then sort these code statements by their line numbers {l1,l2,…,l n Sort the rows in ascending order, where l1 is the smallest row number. n The largest line number; secondly, l min Defined as l1-N, l max Defined as l n +N, where N is a constant initially set to 3, and the value of N is gradually increased to ensure row l min and l max Able to map to commit The corresponding line in the file related to fixing the defect; then in the commit. Extract from line number l of the corresponding file related to defect repair. min to line number l max The content between them forms the submission. The corresponding optimized context; finally, line l min and l max Mapping to commit C fix The corresponding line number l′ in the file related to fixing the defect min and l′ max And in submitting C fix Extract from line number l′ in the corresponding file related to defect repair. min to line number l′ max The content between them forms the submission C. fix The corresponding optimized context.
6. The method for automatically locating defective submissions based on a large language model according to claim 3, characterized in that, Step S24 specifically includes: First, construct prompt words based on the code statements that cause the defects and their corresponding reasons, as well as the code statements that fix the defects and their corresponding reasons; Then, submit the warning message, the root cause of the defect, and the current commit to fix the defect (C). fix The corresponding optimized context is used as input to the large language model to obtain the output of the large language model, which is then submitted as C. fix The result of the judgment on whether the corresponding context contains a defect; the prompt word, the root cause of the defect, and the previous submission to fix the defect. The corresponding optimized context is used as input to the large language model to obtain the output of the large language model. The result of the judgment on whether there is a defect in the corresponding context; Secondly, a judgment is made based on two judgment results output by the large language model: if the large language model considers the submission... The corresponding context contains a defect, and submission C... fix If the corresponding context does not contain defects, it is assumed that the large language model can correctly identify the fixed version and the error version of the program, and the context-enhanced identification method is used to locate the submission that introduces the defect; otherwise, it is assumed that the large language model cannot correctly identify the fixed version and the error version of the program, and the ranking-based identification method is used to locate the submission that introduces the defect.
7. The method for automatically locating defective submissions based on a large language model according to claim 6, characterized in that, The method of using context-enhanced identification to locate submissions that introduce defects specifically includes: First, extract all the code statements that cause the defect; Then, trace the code statements that caused the defects, find all the commits that introduced these code statements to fix the defects, obtain a set of candidate commits, and sort them in descending order according to the commit date to form a candidate commit list {C1, C2, ..., C...}. j ,…,C i }, where C j Let j represent the j-th candidate submission, and i represent the total number of candidate submissions; Then, an optimized context is generated for each candidate submission, resulting in each candidate submission C. j and its previous submission The corresponding optimized context; Secondly, the large language model is used to examine each candidate submission in the candidate submission list from index 1 to i. If an index f can be found such that the large language model considers candidate submission C to be... f The corresponding context contains defects, and the candidate commits If the corresponding context does not contain defects, then candidate submission C is selected. f Submissions that introduce defects are identified; conversely, submissions that introduce defects are identified using a ranking-based identification method.
8. The method for automatically locating defective submissions based on a large language model according to claim 1, 3, 6, or 7, characterized in that, The method of using ranking-based identification to locate submissions that introduce defects specifically includes: First, the constructed prompt words, the submission to fix the defect, the root cause of the defect, and the files related to fixing the defect are input into the large language model to obtain the code statements in the submission to fix the defect that caused the defect. Then, using a listwise ranking algorithm based on a large language model, a prompt word is constructed, and the prompt word, the code statement that causes the defect, and the root cause of the defect are input into the large language model. The large language model sorts these code statements according to the correlation between the code statement that causes the defect and the root cause of the defect. Secondly, for each file, retrieve the top N code statements, trace back to the commits that introduced them, and add these commits to the candidate commit list; Finally, these candidate submissions are sorted according to their submission dates, with the most recent submission being the one that introduced the defect.