Software fault positioning method based on bidirectional association vector retrieval, fault tracking system and application

By using the two-way correlation vector retrieval (BLER) method in software fault location, the two-way correlation information between software faults and code submission records is analyzed, and the problem of low positioning accuracy caused by the missing fault repair association is solved, and more efficient software fault location is achieved.

CN119988189APending Publication Date: 2025-05-13SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410540888.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In software fault location, the prior art has low positioning accuracy due to the failure repair association (BFLM) phenomenon, which increases labor and time costs.

Method used

The two-way correlation vector search (BLER) method is used to generate code file scores that cause and/or repair the fault by analyzing the bidirectional association information in the fault repair history, including repairing the association and raising the association.

Benefits of technology

It improves the accuracy of software fault location, and can rank fault code submissions in the top 5 with an 85% probability and 70% probability in the top 3 with a significantly improved positioning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004820987700000011
    Figure HDA0004820987700000011
  • Figure HDA0004820987700000012
    Figure HDA0004820987700000012
  • Figure HDA0004820987700000021
    Figure HDA0004820987700000021
Patent Text Reader

Abstract

The invention discloses a software fault positioning method based on bidirectional association vector retrieval, and the method comprises the steps: receiving a newly obtained software fault report, and extracting a fault description; according to the extracted fault description text, performing similarity comparison on the fault description text in the newly obtained software fault report and the fault description text in each historical fault report by using a word embedding model; according to the obtained similarity comparison result, generating a code file score for causing and / or repairing the fault by analyzing bidirectional association information between the software fault and the code submission record by utilizing a bidirectional association vector retrieval method; and matching submitted codes possibly causing software faults with the obtained code file scores, obtaining scores of the submitted codes, sorting the submitted codes, and obtaining a code fault positioning result. The invention further discloses a system and application for implementing the method, and the system and the application have wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer technology and artificial intelligence technology, and relates to a software fault locating method based on bidirectional correlation vector retrieval, a fault tracking system and an application. Background Art

[0002] Mature software development companies often require testers to provide relevant code submission records of the introduction of the fault while providing software fault information. If a simple method is used to repeatedly test each code submission, it is often time-consuming and laborious, increasing the high manpower and time costs.

[0003] In addition, in enterprises, due to the negligence or bad habits of engineers, when repairing the code, the engineer may forget to fill in the association information between the repaired code submission and the corresponding fault report (such as the defect report number, etc.), resulting in the problem of missing fault fix association (BFLM), which will affect the fault location technology based on information retrieval (IR). Since the fault location technology based on information retrieval (IR) uses the association relationship between the code submission for fixing the fault and the fault report, and the BFLM phenomenon will cause the loss of such association information, thus making this fault location method ineffective.

[0004] The IR technology is used to retrieve information matching the query condition from a large amount of text. In software fault location, the IR technology can be used to find code submissions related to a specific fault and analyze the association information between the fault report and the code submission record.

[0005] Sliwerski et al. (2005) [1] The SZZ method for defect prediction research is proposed. The SZZ method analyzes the code change history and related defect reports to determine which previous code submissions may have introduced defects that were later discovered. The method is based on the condition that development engineers need to fill in the corresponding defect report number and other information when submitting a fix for a software defect. However, in the actual development process, engineers often ignore this action when fixing the code, resulting in the loss of direct correlation information between the defect and the fix. This will ultimately seriously affect the effectiveness of the method. Summary of the invention

[0006] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a software fault location method, fault tracking system and application based on bidirectional association vector retrieval. The present invention uses the Bi-directional Links Embedding Retrieval (BLER) method to provide a time-saving and labor-saving solution for quickly locating software fault codes.

[0007] The present invention proposes a new fault code location method, which automatically and quickly locates the defective code that causes software faults through a self-created bidirectional association vector retrieval method. Since the existing software fault location technology does not take into account the fault repair association loss phenomenon existing in software development companies, the present invention has made a brand-new improvement on this: on the basis of the existing SZZ method, by adding an association dimension, a bidirectional association vector retrieval (Bi-directional Links Embedding Retrieval, referred to as BLER) is proposed to solve the problem of low positioning accuracy caused by the fault repair association missing (BFLM, Bug / Fixing Link Missing) phenomenon; the two directions in the bidirectional association vector retrieval contain information in two directions: the association information between historical faults and repair codes, and the association information between the code that caused the historical fault and the historical fault. The present invention also introduces the embedding model commonly used in the semantic big model to improve the retrieval accuracy.

[0008] In order to solve the failure impact caused by BFLM, we can first understand the nature of IR-based fault location technology. There are two correlations between software faults and program codes: one is the correlation between the code that fixes the fault and the fault, and the other is the correlation between the code that causes the fault and the fault.

[0009] like Figure 2 As shown in the figure, there are two kinds of associations between the historical fault report H and the code:

[0010] Association 1 can also be called "repair association", that is, the association between the code that fixes the fault and the fault, which is created by the development engineer when submitting the repair code. After submitting the repair code, the development engineer fills in the number of the fault report fixed by the code submission in the description of the code submission, thereby establishing an association relationship from the code submission CL_F to the historical fault report H (indicated by the dotted arrow from CL_F to H in the figure).

[0011] The second association can be called "causing association", that is, the association between the code that caused the fault and the fault, which is created by the test engineer when updating the fault report. After discovering the fault, the test engineer first submits the fault report, and then finds the code submission CL_O that caused the fault by testing and comparing the compiled versions of each recent code submission. Then fill CL_O into the fault report to help the development engineer fix the fault as soon as possible, thus establishing an association relationship from the historical fault report H to the code submission CL_O (indicated by the dotted arrow from H to CL_O in the figure).

[0012] The principle of the current fault location technology based on information retrieval is to use the repair association of historical faults to infer the triggering association of new faults. However, from the perspective of probabilistic causality, using the triggering association of historical faults to infer the triggering association of new faults is obviously more consistent with causality. Moreover, the existence of the BFLM phenomenon in enterprises will cause the repair association between the fault repair code and the fault to be missing, thus causing the current fault location technology based on information retrieval to fail.

[0013] The present invention takes into account that the association between code and fault is bidirectional (referring to repair association and triggering association). In addition to the repair association of historical faults, the triggering association of historical faults can also be used to infer the code file that causes new faults. The BLER method used in the present invention can not only track the repair code of the fault by analyzing the historical data of fault repair, but also determine the code submission that initially introduced the fault, thereby providing a more accurate and comprehensive fault location solution when the BFLM problem exists. The bidirectional correlation vector retrieval (BLER) proposed in the present invention is based on this idea.

[0014] The present invention provides a method for locating software faults by using bidirectional correlation vector retrieval, the method comprising the following steps:

[0015] Step 1: Receive a newly acquired software fault report and extract the fault description text;

[0016] Step 2: Based on the fault description text extracted in step 1, a word embedding model is used to compare the similarity between the fault description text in the newly obtained software fault report and the fault description text in each historical fault report;

[0017] Step 3: Based on the similarity comparison result obtained in step 2, a bidirectional correlation vector retrieval (BLER) method is used to generate a code file score that causes and / or repairs the fault by analyzing the bidirectional correlation information between the software fault and the related code submission record;

[0018] Step 4: Match the recent code submissions that may cause software failures with the code file scores obtained in step 3, obtain the code submission scores and generate code submission rankings to obtain code failure location results.

[0019] In step one, a fault report submitted by the test team and / or end user through the fault tracking system is received; the fault report includes a detailed description of the fault, such as the fault phenomenon, reproduction steps, impact range, etc.; key information in the fault report is extracted through a predetermined processing flow, including the fault description text, and the fault description text is subsequently used as an input for information retrieval to help locate the specific code location that caused the fault.

[0020] In the process of extracting the fault description text, the fault report also needs to be preprocessed including format conversion and filtering, converting the fault report from HTML format to ordinary text format, filtering out meaningless special symbols, extracting the fault title and fault description in the fault report, and merging them into a complete fault description text.

[0021] In step 2, a word embedding model such as Word2Vec or BERT is used to process the fault description text obtained in step 1; the word embedding model can convert the fault description text into a vector representation in a high-dimensional space, thereby capturing the deep semantic information in the text. By using methods such as cosine similarity calculation, the similarity between the fault description vector and the description vector in the code submission record is directly calculated and compared in the word embedding model, so that the code submission related to the fault can be located more accurately.

[0022] In step 3, the core of the bidirectional correlation vector retrieval (BLER) method is to analyze and utilize the bidirectional correlation information in the fault repair history: that is, tracing back from the repaired code to the fault report (repair correlation), and tracing back from the fault report to the code submission that introduced the fault (trigger correlation). The newly obtained software fault report is related to one or more historical fault reports, and the similarities are Sim1, Sim2, ..., ..., Sim n-1 , Sim n , a code file is associated with the newly obtained software fault report and / or one or more historical fault reports, and the score of the code file is the sum of the similarity scores of the one or more historical fault reports associated with the code file and the newly obtained software fault report.

[0023] In step 4, recent code submissions are matched with scored code files, the code files contained in the code submissions are mapped to the scored code files, and the scores of each code submission are calculated to generate a ranking of recent code submissions; based on the ranking order, code submissions with high scores are given priority for review, repair, and optimization.

[0024] The present invention also provides a fault tracking system for implementing the above-mentioned software fault locating method, the fault tracking system comprises: a fault report receiving and extracting module, an information retrieval module, a bidirectional correlation vector retrieval module, and a code submission tracking module;

[0025] The fault report receiving and extracting module is used to receive the software fault report obtained by transmission, and extract the fault description from the software fault report;

[0026] The information retrieval module is used to retrieve and extract historical fault reports related to the newly received fault report;

[0027] The bidirectional relevance vector retrieval module uses a word embedding model to compare the similarity between the text in the newly received fault report and the text in the historical fault report, and performs a file score on the code file;

[0028] The code submission tracking module is used to match the code submission that may cause the new fault with the code file with a score, calculate the score of the code submission, generate a code submission ranking, and obtain the code submission that is most relevant to the new fault.

[0029] The present invention also provides the software fault location method or the application of the fault location system as described above in automated and / or non-automated software construction and testing.

[0030] The present invention also provides a hardware system for implementing the above-mentioned software fault locating method, the hardware system comprising: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned software fault locating method is implemented.

[0031] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned software fault locating method is implemented.

[0032] The beneficial effects of the present invention include: through experiments on actual enterprise projects, the results show that there is an 85% probability that the fault code submission appears in the top 5 of the fault code sorting list, and there is about a 70% probability that it is ranked in the top 3 of the fault code sorting list. This accuracy is much higher than that of previous methods such as SZZ (the accuracy is only 60%).

[0033] In addition, the BLER method of the present invention solves the problem of the SZZ method encountering a fault repair association missing (BFLM) by applying the association between the code repairing the fault and the fault and the association between the code causing the fault and the fault, which will directly cause the SZZ method to fail. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1This is a block diagram of the fault location method of the present invention. First, the fault description is extracted from the newly received fault report and used as the query information for vector retrieval. The description text in all historical fault reports is compared with the similarity based on the embedding model. Then, based on the two types of associated information (the two dotted lines in the figure), the code files that cause the fault and the code files that fix the fault are scored. Finally, the recent code submissions are matched with the previously generated scored code files to calculate the scores of each code submission, thereby generating a ranking of recent code submissions.

[0036] Figure 2 It is a bidirectional association diagram between faults and codes. It includes the association between the code that causes the fault and the fault, and the association between the code that fixes the fault and the fault.

[0037] Figure 3 The method of calculating the code file score using BLER is described. In the figure, code file F3 has caused or repaired historical faults H1 and H4. Assuming that the similarity between the description of fault H1 and the new fault description is Sim1, and the similarity between the description of fault H4 and the new fault description is Sim4, then the score of file F3 is Sim1+Sim4.

[0038] Figure 4 It is the historical fault report screening process.

[0039] Figure 5 Demonstrates how to calculate scores for recent code commits. DETAILED DESCRIPTION

[0040] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0041] The present invention provides a method for locating software faults by using bidirectional correlation vector retrieval. Figure 1 This is a block diagram of the fault location method of the present invention. First, the fault description text is extracted from the newly received fault report and used as the query information for vector retrieval. The fault description text extracted from the newly received fault report is compared with the description text in all historical fault reports based on the similarity of the embedding model. Then, based on the two types of associated information (the two dotted lines in the figure), the code files that cause the fault and the code files that fix the fault are scored. Finally, the recent code submissions are matched with the previously generated scored code files to calculate the scores of each code submission, thereby generating a ranking of recent code submissions, and selecting the top ranked code submissions as the code submissions for priority review, repair, and optimization.

[0042] First, the title text and description text of the newly received fault report in the fault tracking system are extracted, and after text preprocessing including format conversion and filtering, the text is input as a query to the information retrieval module. Then, based on the extracted title text and description text of the newly received fault report, the information retrieval module extracts historical fault reports related to the project from the fault tracking system, and the titles and descriptions of these fault reports will be used as the text database for information retrieval. The bidirectional relevance vector retrieval module (BLER) performs text preprocessing and indexing on the fault description text in the historical fault report, and uses the Embedding Retrieval method to compare the similarity between the fault description text in the new fault report and the fault description text in each historical fault report. According to the text similarity, the code file that caused the historical fault and the code file that repaired the historical fault are scored. The code file may be associated with both the historical fault description and the new fault description at the same time. The more similar the historical fault description associated with the code file is to the new fault description, the higher the file score of the code file. Finally, the recent code submissions since the last regression test (which did not contain the new fault at that time) are obtained from the version control system, and these code submissions contain code changes that caused the new fault. By matching these code submissions with the scored code files, the score of each code submission is calculated, thereby generating a code submission ranking and obtaining the code submission most relevant to the new fault.

[0043] In order to perform bidirectional correlation information retrieval on the historical fault reports of enterprise projects, firstly, all the historical fault reports of a certain software project are extracted from the fault tracking system, and then these fault reports are screened. Figure 4 As shown, first filter out those software fault reports that do not involve code modification. There are many reasons for such fault reports: the fault report itself may be invalid or wrong; or the project team may decide not to fix it for the time being after discussion.

[0044] Another situation that needs to be filtered out is that there are some special purpose fault numbers in the enterprise that are used to make some special code submissions. For example, adding comments to the code, code cleaning, etc. Although these fault numbers also have code changes, their purpose is not to fix the fault, so they also need to be filtered out. Common special purpose fault numbers in enterprises include but are not limited to those shown in Table 1 below:

[0045] Table 1 Common special purpose fault numbers in enterprises

[0046] type use 1 Code promotion between different code lines 2 Software version number change (version change) 3 Document management 4 Source code clean up 5 Some simple test tool development (test development)

[0047] Table 1 lists some common special-purpose fault codes in enterprises. In different software enterprises, engineers will define some fault codes to meet special needs according to their own needs. Some are used for software version management (such as types 1 and 2), some are used for non-logical modification of documents and codes (such as types 3 and 4), and some are used for test tool development (such as type 5).

[0048] After filtering, the remaining software faults are all real and effective fault reports that development engineers have fixed by modifying the code. From these fault reports, the fault title (called Summary in Jira) and fault description of the report are extracted, and then the text of the fault title and the fault description are connected to form a complete fault description text.

[0049] For newly received fault reports, the fault title and fault description text are also extracted and merged into the fault description text. The fault description text is converted from HTML format to ordinary text format, and meaningless special symbols are filtered out. The final text is used for information retrieval query, that is, compared with the historical fault description text stored in the csv file to calculate the text similarity.

[0050] The similarity between the new fault report and each historical fault report is calculated by using the Embedding Retrieval method to compare the fault descriptions. The code file that causes or fixes a historical fault report is scored based on the similarity. Figure 3 The score of a code file F in F It is the cumulative value of the similarity between all historical faults caused or repaired by it and the new fault.

[0051] For example, Figure 3 The code file F3 has caused or repaired historical faults H1 and H4. Assuming that the similarity between the description of fault H1 and the new fault description is Sim1, and the similarity between the description of fault H4 and the new fault description is Sim4, the score of file F3 is Sim1+Sim4.

[0052] After the above steps, we get the file scores of all code files that caused or repaired historical faults. Then we can calculate the scores of recent code submissions based on them. The main method is to map the code files contained in a recent code submission to the code files scored in the previous section, add up the scores of these files, and calculate the total value.

[0053] The score of each code submission is equal to the total score of each code file it contains. If the code file in the code submission has ever caused or fixed a historical fault report, then there will be a corresponding file score. For code files in the code submission that have never caused or fixed historical faults, their file scores are zero.

[0054] like Figure 5 As shown in the figure, suppose there is a code submission CL_N recently, which contains two code files that have caused or repaired historical faults: file F3 and file F5. In addition, the submission also contains a code file that has not caused or repaired historical faults: file F9. Then according to the above algorithm, the score of F9 is zero, and the score of the code submission is S3+S5+0, that is, S3+S5.

[0055] Calculate the score of each recent code submission according to the above steps, sort the code submissions according to the score, and select the top-ranked code submissions for priority review, repair, and optimization.

[0056] References

[0057] [1] J.Sliwerski, T.Zimmermann, and A.Zeller, "When Do Changes InduceFixes?" Proc.Int'lConf.Mining Software Repositories, 2005, pp.1-5.

[0058] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A software fault location method based on bidirectional correlation vector retrieval, characterized in that: The software fault locating method comprises: Step 1: Accept the newly acquired software fault report and extract the fault description; Step 2: Based on the fault description text extracted in step 1, a word embedding model is used to compare the similarity between the fault description text in the newly obtained software fault report and the fault description text in each historical fault report; Step 3: Based on the similarity comparison result obtained in step 2, a bidirectional correlation vector retrieval method is used to generate a code file score that causes and / or repairs the fault by analyzing the bidirectional correlation information between the software fault and the code submission record; Step 4: Match the code submissions that may cause software failures with the code file scores obtained in step 3, obtain the scores of each code submission, sort the code submissions, and obtain the code failure location results.

2. The software fault location method according to claim 1, characterized in that: In step 1, the software fault report is submitted by the testing team and / or the end user through the fault tracking system, and includes a detailed description of the fault; the fault description text is extracted from the fault report as subsequent input; During the fault description text extraction process, the fault report is formatted and filtered preprocessed to convert the fault report from HTML format to ordinary text format, and meaningless special symbols are filtered out to extract the fault title and fault description in the fault report and merge them into a complete fault description text.

3. The software fault location method according to claim 1, characterized in that: In step 2, the word embedding model includes Word2Vec and BERT; the word embedding model is used to convert the fault description text into a vector representation in a high-dimensional space to capture the deep semantic information in the text; and the similarity between the fault description vector and the description vector in the code submission record is directly calculated and compared through the word embedding model to locate the code submission related to the fault.

4. The software fault location method according to claim 1, characterized in that: In step 3, the newly obtained software fault report is related to one or more historical fault reports, and the similarity comparison results are Sim1, Sim2, ..., ..., Sim n-1 , Sim n A code file is associated with the newly obtained software fault report and / or one or more historical fault reports, and the score of the code file is the sum of the similarity scores of the one or more historical fault reports associated with the code file and the newly obtained software fault report.

5. The software fault location method according to claim 1, characterized in that: In step 4, the code files contained in the code submission are mapped to the code files with scores obtained in step 3, the scores of each code submission are calculated, and the code submissions are sorted according to the scores. According to the sorting order, the code submissions with high scores are preferentially selected for review, repair and optimization.

6. A fault tracking system for implementing the software fault location method according to any one of claims 1 to 5, characterized in that: The fault tracking system includes: a fault report receiving and extracting module, an information retrieval module, a bidirectional correlation vector retrieval module, and a code submission tracking module; The fault report receiving and extracting module is used to receive the software fault report obtained by transmission, and extract the fault description from the software fault report; The information retrieval module is used to retrieve and extract historical fault reports related to the newly received fault report; The bidirectional relevance vector retrieval module uses a word embedding model to compare the similarity between the text in the newly received fault report and the text in the historical fault report, and performs a file score on the code file; The code submission tracking module is used to match the code submission that may cause the new fault with the code file with a score, calculate the score of the code submission, generate a code submission ranking, and obtain the code submission that is most relevant to the new fault.

7. Application of the software fault location method according to any one of claims 1 to 5, or the fault tracking system according to claim 6 in automated and / or non-automated software construction and testing.

8. A hardware system for implementing the software fault location method according to any one of claims 1 to 5, characterized in that: The hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the software fault locating method according to any one of claims 1 to 5 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the software fault locating method according to any one of claims 1 to 5 is implemented.