Interpretable defect prediction method, and related device

WO2026199896A1PCT designated stage Publication Date: 2026-10-01HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/128077
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-10-16
Publication Date
2026-10-01

Smart Images

  • Figure CN2025128077_01102026_PF_FP_ABST
    Figure CN2025128077_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An interpretable defect prediction method, and a related device, which relate to the technical fields of software analysis and defect prediction, and are used for interpretably describing a decision-making process during software defect prediction. The method in the present application comprises: constructing a defect data set on the basis of historical data of a software project to be analyzed, and obtaining, from current data of said software project, code modules to be analyzed in said software project, and code module features; on the basis of the defect data set, determining an actual defect rate and an expected defect rate of any code module feature in any code module to be analyzed; calculating the difference between the actual defect rate and the expected defect rate, so as to obtain a feature defect score; calculating the feature defect scores of all the code module features in any code module to be analyzed, so as to obtain a module defect score; and after the module defect scores of all the code modules to be analyzed are obtained, sorting all the module defect scores, and sequentially processing all the code modules to be analyzed in said software project.
Need to check novelty before this filing date? Find Prior Art

Description

An interpretable defect prediction method and related equipment

[0001] This application claims priority to Chinese Patent Application No. 202510361564.6, filed on March 26, 2025, entitled "An Explainable Defect Prediction Method and Related Equipment", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of software analysis and defect prediction technology, and in particular to an interpretable defect prediction method and related equipment. Background Technology

[0003] Interpretability is a crucial aspect of practical defect prediction. Existing workload-aware software defect prediction (predicting whether a code module has defects) methods train a defect prediction model using a defect dataset (features of the code module and the number of defects in that module). This model is typically a complex machine learning model, such as a deep learning model. Finally, the code modules are ranked based on the prediction results. However, the black-box decision-making process of these models leads to poor interpretability of the prediction results.

[0004] To improve the interpretability of defect prediction models, researchers have recently begun employing model-agnostic interpretation methods (such as Local Interpretable Model-agnostic Explanations (LIME) or machine learning model explanation tools (SHAP)) to interpret individual predictions. For a new code module, the defect prediction model predicts its defect probability. Then, a model-agnostic interpretation method (such as LIME) is used to interpret this prediction. These interpretation methods generate explanations based on complex perturbation mechanisms.

[0005] Specifically, to generate feature importance for a specific prediction, these methods typically modify the instance to be explained and observe the changes in the model output. The explanation method outputs the contribution of each feature to the prediction result, i.e., the feature importance. However, these explanation methods have several problems: First, they treat the prediction model as a black box, reducing developers' trust in the explanation; second, these methods rely on complex perturbation mechanisms, which may generate inaccurate or even misleading explanations; finally, the generated explanations may be unstable due to randomness. Summary of the Invention

[0006] This application provides an interpretable defect prediction method and related equipment for interpretably describing the decision-making process in software defect prediction.

[0007] The first aspect of this application provides an interpretable defect prediction method, including:

[0008] A defect dataset is constructed based on historical data of the software project to be analyzed. Each code module to be analyzed in the software project and the code module features corresponding to each code module to be analyzed are obtained from the current data of the software project to be analyzed. The code module features are used to characterize the numerical features of the code module to be analyzed.

[0009] If the module feature value of any code module feature in any code module to be analyzed is the target feature value, the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed are determined according to the defect dataset; wherein, the expected defect rate is the ratio of the total number of lines of code modules in all defective code modules in the defect dataset whose module feature values ​​are the target feature value to the total number of lines of code modules in all defective code modules in the defect dataset, and the actual defect rate is the ratio of the total number of defects in all defective code modules in the defect dataset to the total number of defects in all defective code modules in the defect dataset; the defect dataset includes the total number of lines of code modules and the total number of defects in code modules corresponding to different module feature values, as well as the total number of lines of code modules and the total number of defects in code modules corresponding to all defective code modules; the defective code module is the code module in the defect dataset;

[0010] Calculate the difference between the actual defect rate and the expected defect rate to obtain the feature defect score of any code module feature in any code module;

[0011] Calculate the feature defect scores of all code module features in any code module to be analyzed to obtain the module defect score of any code module to be analyzed; wherein, the module defect score is used to characterize the defect tendency of the code module to be analyzed;

[0012] After obtaining the module defect scores of all code modules to be analyzed, the module defect scores are sorted so that all code modules in the software project to be analyzed are processed in descending order of the module defect scores.

[0013] A second aspect of this application provides an interpretable defect prediction apparatus, comprising:

[0014] Central processing unit, memory, input / output interfaces, wired or wireless network interfaces, and power supply;

[0015] The memory is either a short-term storage memory or a persistent storage memory;

[0016] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the interpretable defect prediction method described in the first aspect.

[0017] A third aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the interpretable defect prediction method described in the first aspect.

[0018] A fourth aspect of this application provides a computer program product including instructions that, when executed on a computer, cause the computer to perform the interpretable defect prediction method described in the first aspect.

[0019] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: The interpretable defect prediction method provided by these embodiments evaluates the defect tendency and defect score of each code module in a software project by calculating the offset between the actual defect rate and the expected defect rate, and sorts the defect scores to indicate the importance of module features to the defect score of the code module. The evaluation process is described interpretably, improving the accuracy and stability of the interpretation. Simultaneously, developers can adjust the code modules sequentially according to the sorting order, improving the efficiency of software defect repair. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0021] Figure 1 is a schematic diagram of a ScoringDP system architecture provided in an embodiment of this application;

[0022] Figure 2 is a flowchart illustrating an interpretable defect prediction method provided in an embodiment of this application;

[0023] Figure 3 is a flowchart illustrating another interpretable defect prediction method provided in an embodiment of this application;

[0024] Figure 4 is a flowchart illustrating another interpretable defect prediction method provided in an embodiment of this application;

[0025] Figure 5 is a flowchart illustrating another interpretable defect prediction method provided in an embodiment of this application;

[0026] Figure 6 is a flowchart illustrating another interpretable defect prediction method provided in an embodiment of this application;

[0027] Figure 7 is a schematic diagram of the algorithm logic for scoring a code module provided in an embodiment of this application;

[0028] Figure 8 is a schematic diagram of the algorithm logic for module feature scoring provided in an embodiment of this application;

[0029] Figure 9 is a schematic diagram of a window-based actual defect rate estimation provided in an embodiment of this application;

[0030] Figure 10 is an example diagram illustrating the feature importance of ScoringDP according to an embodiment of this application;

[0031] Figure 11 is a Chinese-English comparison chart of the feature names of a module feature;

[0032] Figure 12 is a schematic diagram of the structure of an interpretable defect prediction system provided in an embodiment of this application;

[0033] Figure 13 is a schematic diagram of the structure of an interpretable defect prediction device provided in an embodiment of this application. Detailed Implementation

[0034] Software defect prediction is an important research area in software engineering, aiming to predict which modules may have defects by analyzing the characteristics of code modules (such as code complexity and number of lines of code). Traditional defect prediction methods typically rely on complex machine learning models (such as deep learning models) to improve prediction performance. However, these models often lack interpretability, leading to reduced developer confidence in the prediction results and thus affecting their practical application.

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0036] Please refer to Figure 1, which is a schematic diagram of the system architecture of ScoringDP provided in an embodiment of this application. As shown in Figure 1, this embodiment mainly provides a method for defect prediction by scoring each code module. Specifically, in order to rank the code modules in a software project, the Scoring Defect Prediction (ScoringDP) system first calculates a score for each code module, and then ranks the modules based on these scores. These scores are called defect scores. Figure 1 shows the process of ScoringDP calculating defect scores, where the defect score of a code module is an aggregation of the defect scores of its features. The defect scores of module features can indicate the importance of module features to the code module's defect score. The defect dataset can be constructed based on the historical change data of the code modules (obtained from analysis of the code repository); if historical data is insufficient, data from similar projects can be used to construct the dataset, and no specific limitation is made here. It should also be noted that module features mainly describe various attributes of the code, such as cohesion, coupling, complexity, or scale. Examples include Average Method Complexity (AMC), Average Cyclomatic Complexity (AVG CC), and Lines of Code (LOC), which will not be elaborated upon here. Please refer to Figure 11 for details; Figure 11 shows a Chinese-English translation of the feature names for a module feature. For ease of description, this will not be elaborated upon further.

[0037] To address the aforementioned technical problems, please refer to Figure 2, which is a flowchart illustrating an interpretable defect prediction method provided in an embodiment of this application. It includes steps 201-205.

[0038] 201. Construct a defect dataset based on the historical data of the software project to be analyzed, and obtain each code module to be analyzed and the corresponding code module characteristics from the current data of the software project to be analyzed.

[0039] In this embodiment, a defect dataset is constructed using historical data from the software project to be analyzed. Then, by analyzing the current data of the software project, all code modules to be analyzed and the code module characteristics of each code module are obtained. It should be noted that the code module characteristics are used to characterize the numerical features of the code modules to be analyzed. The software project to be analyzed is the one to be analyzed in this embodiment; specifically, the defects of this software project need to be analyzed and predicted.

[0040] In one specific embodiment, referring to Figure 1, during the analysis of a software project, it is necessary to construct a dataset by combining historical data of the software project or data from similar projects. This dataset can be understood as a defect dataset, which is a tabular dataset where each row represents a code module. Each code module is described by one or more numerical features, such as the number of lines of code (LOC) and cyclomatic complexity, etc., without specific limitations. Furthermore, if each row in the defect dataset D represents a code module cm, and code module cm is described by m numerical features, then the defect dataset D includes m features and a corresponding label, representing the defect data of that module. Therefore, it can be understood that the defect dataset includes the total number of lines of code in the code modules, the total number of defects in the code modules, and the total number of lines of code in the total number of defects in the code modules of the software project.

[0041] 202. When the module feature value of any code module feature in any code module to be analyzed is the target feature value, determine the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed based on the defect dataset.

[0042] Furthermore, when the target feature value is a module feature that satisfies the characteristics of any code module within any code module in the software project to be analyzed, then each code module feature within each code module to be analyzed can be analyzed based on the total number of lines of code in each code module, the total number of lines of code in all code modules, the total number of defects in each code module, and the total number of defects in all code modules in the defect dataset. This allows for the calculation of the actual defect rate and the expected defect rate for each code module feature. It should be noted that the expected defect rate is the ratio of the total number of lines of code in the defective code module to the total number of lines of code in all defective code modules in the defect dataset, while the actual defect rate is the ratio of the total number of defects in the defective code module to the total number of defects in all defective code modules in the defect dataset. The defective code module refers to the code module in the defect dataset, and the code module to be analyzed refers to the code module in the software project to be analyzed.

[0043] In one specific implementation, ScoringDP calculates the defect score of a code module feature by quantifying the deviation between two defect rates: the actual defect rate and the expected defect rate. For a module feature fv of a code module feature f, the expected defect rate is defined as the ratio of the total number of lines of code in the code module with module feature fv to the total number of lines of code in all code modules within the defect dataset. The calculation of the expected defect rate is based on the intuitive assumption that defects in a software project are uniformly distributed across the lines of code in all code modules within the project. Similarly, the actual defect rate is defined as the ratio of the total number of defects in those code modules with module feature fv to the total number of defects in all code modules. For ease of understanding, the definitions of expected and actual defect rates will not be described further below.

[0044] Furthermore, the calculation of the expected defect rate is based on an intuitive assumption that defects in a software project are uniformly distributed across lines of code in all code modules within the project. However, the scoring of the aforementioned code module characteristics contains an implicit assumption: the defect dataset used contains an infinite number of code modules. However, in real-world scenarios, the size of the defect dataset is finite, so at least two scenarios may occur. For ease of understanding, please refer to the embodiments shown in Figure 3 or Figures 4 to 6, respectively.

[0045] 203. Calculate the difference between the actual defect rate and the expected defect rate to obtain the feature defect score of any code module feature in any code module.

[0046] After obtaining the actual defect rate and the expected defect rate of the code module features, the difference between the actual defect rate and the expected defect rate can be calculated, thereby obtaining the feature defect score of any code module feature in any code module.

[0047] In one specific embodiment, the defect score of a code module feature is defined as the difference between the actual defect rate and the expected defect rate. Furthermore, the feature defect score can be obtained by subtracting the expected defect rate from the actual defect rate.

[0048] 204. Calculate the feature defect scores of all code module features in any code module to be analyzed, and obtain the module defect score of any code module to be analyzed.

[0049] Therefore, based on steps 202-203, the feature defect scores of all code module features in any code module to be analyzed can be calculated, thus obtaining the module defect score of any code module to be analyzed. It should be noted that the module defect score is used to characterize the defect tendency of the code module to be analyzed.

[0050] In one specific embodiment, when analyzing all the code module features in a certain code module to be analyzed, the feature defect scores of all the code module features in the code module to be analyzed can be accumulated to obtain the module defect score of the code module to be analyzed.

[0051] 205. After obtaining the module defect scores of all code modules to be analyzed, sort all module defect scores so that all code modules to be analyzed in the software project to be analyzed can be processed in descending order of module defect scores.

[0052] Therefore, after obtaining the module defect scores of all code modules in the software project to be analyzed, the defect scores of all modules can be sorted. Then, according to the order of the defect scores from largest to smallest, and according to the priority, all code modules in the software project to be analyzed can be processed in turn. Furthermore, more effective actions can be taken to reduce defects based on the defect scores of the code module characteristics, such as modifying the code to prioritize changing the characteristics with higher defect scores.

[0053] In one specific implementation, defect scores for all code module features are summed, a simple summation method that enhances the interpretability of ScoringDP.

[0054] This embodiment provides an interpretable defect prediction method that assesses the defect tendency and defect score of each code module in a software project by calculating the offset between the actual defect rate and the expected defect rate. The defect scores are then ranked to indicate the importance of module features to the defect score. The evaluation process is described interpretably, improving the accuracy and stability of the interpretation. Furthermore, developers can adjust code modules sequentially according to the ranking order, improving the efficiency of software defect repair.

[0055] Figure 3 illustrates one implementation of steps 201-202 in Figure 2, specifically obtaining the actual defect rate and expected defect rate of the code module features. For ease of understanding, please refer to Figure 3, which is a flowchart illustrating another interpretable defect prediction method provided in this application embodiment. It includes steps 301-304.

[0056] 301. Perform line-by-line analysis on the software project to be analyzed, set each line of the software project to be analyzed as any code module to be analyzed, and determine the characteristics of the code modules contained in any code module to be analyzed.

[0057] This embodiment primarily describes one possible implementation of steps 201-202. Specifically, by analyzing the software project to be analyzed line by line, each line in the software project can be set as a corresponding code module to be analyzed (referred to as any code module to be analyzed). Thus, the code module characteristics contained in any code module to be analyzed can be determined.

[0058] In one specific embodiment, by analyzing the software project to be analyzed line by line, the code module to be analyzed represented by each line in the software project to be analyzed and the corresponding code module characteristics are obtained.

[0059] 302. Based on the defect dataset, obtain the total number of lines of code in the code module, the total number of lines of code in the total code module, the total number of defects in the code module, and the total number of defects in the total code module.

[0060] It should be noted that steps 302-304 in this embodiment are a calculation of module feature defect scores based on an ideal state. Specifically, in step 302, the calculation of the expected defect rate is based on an intuitive assumption that defects in the defect dataset are uniformly distributed across all lines of code in all code modules. Therefore, when the defect data in the defect dataset is uniformly distributed across all defective code modules, the total number of lines of code in the parsed defect dataset, the total number of lines of code in all code modules, the total number of defects in the code modules, and the total number of defects in the total code modules are obtained.

[0061] Furthermore, during the analysis of the defect dataset, since code module features and related defects are labeled on the defective code modules, one or more code module features and a corresponding module label can be obtained for each defective code module. This label indicates the number of defects in the code module. Therefore, the total number of lines of code in each module, the total number of lines of code in all code modules, the total number of defects in each code module, and the total number of defects in all code modules can be obtained.

[0062] 303. Input the total number of lines of code in the code module and the total number of lines of code in the expected defect rate calculation formula to obtain the expected defect rate.

[0063] Then, input the total number of lines of code in each module and the total number of lines of code in all modules into the expected defect rate calculation formula to obtain the expected defect rate. It should be noted that the expected defect rate calculation formula is:

[0064] Where expected_defect_ratio is the expected defect rate, cm is the defect code module in the defect dataset, D is the defect dataset, f is any code module feature, fv is the target feature value, LOC is the number of lines of code in the defect code module in the defect dataset, Sum({cm.LOC|cm.f=fv,cm∈D}) is used to represent the total number of lines of code modules, and Sum({cm.LOC|cm∈D}) is used to represent the total number of lines of code modules accumulated when the module feature value of the defect code module in the defect dataset is the target feature value.

[0065] 304. Input the total number of defects in the code modules and the total number of defects in the total code modules into the actual defect rate calculation formula to obtain the actual defect rate.

[0066] Corresponding to step 303, the total number of defects in the code modules and the total number of defects in the total number of code modules can be input into the actual defect rate calculation formula to obtain the actual defect rate. It should be noted that the actual defect rate calculation formula is:

[0067] Where actual_defect_ratio is the actual defect rate, defects is the number of defects in the defect code modules in the defect dataset, Sum({cm.defects|cm.f=fv,cm∈D}) is used to represent the total number of code module defects accumulated when the module feature value of the defect code module in the defect dataset is the target feature value, and Sum({cm.LOC|cm∈D}) is used to represent the total number of code module defects.

[0068] In this embodiment, one input to the scoring algorithm is a defect dataset. Each row of the dataset represents a code module, including its features and the number of defects. This dataset is typically constructed based on historical data from the project. For example, features of code module cm are statistically analyzed a year ago, and then the number of defects found in it during that year is observed. However, it should be noted that the embodiment shown in Figure 3 primarily assumes that the defect dataset is large enough that, for any given feature value, there are enough code modules with that feature value. In this case, the expected defect rate and the actual defect rate will not be incorrectly calculated as zero; they will change smoothly with the feature values, thus helping to reasonably estimate these two defect rates.

[0069] The interpretable defect prediction method provided in this embodiment calculates the defect scores of module features, thereby effectively generating reasonable expected defect rates and actual defect rates. This allows for a quantitative and intuitive understanding of the defect status of the code module features of the code module to be analyzed.

[0070] It should be noted in advance that Figures 4, 5, and 6 are alternative embodiments for obtaining the actual defect rate and expected defect rate of code module features compared to Figure 3. Figure 4 describes an embodiment for obtaining the initial interval window; please refer to Figure 4 for details. Figure 4 is a flowchart illustrating another interpretable defect prediction method provided by the embodiments of this application, including steps 401-406.

[0071] 401. Obtain the module feature values ​​of all code module features in any defect code module from the defect dataset, sort all module feature values ​​in ascending order, and obtain a feature value sorting list corresponding to all module feature values.

[0072] As mentioned above, this embodiment mainly describes determining the size of the initial interval window based on the defect dataset. Specifically, since the module feature values ​​of all code module features have been obtained through the defect dataset in the embodiment shown in Figure 3 above, all module feature values ​​can be sorted in ascending order to obtain a feature value sorting list corresponding to all module feature values.

[0073] In one specific embodiment, after obtaining the module feature values ​​of all code module features, all module feature values ​​are sorted in ascending order to obtain a sorted list corresponding to all module feature values, i.e., a feature value sorted list.

[0074] 402. In the feature value sorting list, check whether there are interval window feature values ​​among all module feature values.

[0075] Then, in the feature value sorting list, all module feature values ​​are analyzed, and the presence of interval window feature values ​​is determined. It should be noted that the interval window feature value is the ratio of the number of code modules within a certain interval window to the total number of code modules in the dataset, which must satisfy a preset interval ratio. For example, 20% or 30%, etc., are not restricted here.

[0076] Furthermore, based on different inspection results, any one of steps 403, 404, 405, or 406 can be executed respectively.

[0077] 403. If there are interval window feature values ​​in the feature value sorting list, use the interval window feature values ​​as the insertion points in the feature value sorting list to determine the index of the initial interval window.

[0078] Based on step 402, when a range window feature value exists in the feature value sorting list, the range window feature value is used as the insertion point in the feature value sorting list, and the index of the initial range window is determined. It should be noted that the index includes a left index and a right index. The left and right indices of the initial range window are the interval positions where the module feature values ​​in the feature value sorting list are located, and the insertion point is the center of the initial range window.

[0079] In one specific embodiment, if the interval window feature value fv exists in the feature value sorting list, then the left and right indices of the initial interval window can be set to the indices of the interval window feature value fv in the feature value sorting list. It is necessary to satisfy, as far as possible, that the center of the left and right indices of the initial interval window is the interval window feature value fv.

[0080] 404. If the insertion point is located after the interval position of the feature value of the interval window and the feature value of the last module in the feature value sorting list, set the left and right indices of the initial interval window to the feature value of the last module in the feature value sorting list.

[0081] Based on step 402, when the insertion point is located after the interval position of the feature value of the interval window and the feature value of the last module in the feature value sorting list, the left index and right index of the initial interval window are set to the feature value of the last module in the feature value sorting list.

[0082] In one specific embodiment, if the interval window feature value fv, i.e. the insertion point, is after the last element (module feature value) in the feature value sorting list, then the left and right indices of the initial interval window at this time can be set to the last index in the feature value sorting list, that is, the left and right indices of the initial interval window at this time are set to the last module feature value in the feature value sorting list.

[0083] 405. If the insertion point is located in the interval between the feature value of the interval window and the feature value of the first module in the feature value sorting list, set the left and right indices of the initial interval window to the feature value of the first module in the feature value sorting list.

[0084] Based on step 402, when the insertion point is located in the interval position between the feature value of the interval window and the feature value of the first module in the feature value sorting list, the left index and right index of the initial interval window are set to the feature value of the first module in the feature value sorting list.

[0085] In one specific embodiment, if the insertion point is at the first position in the feature value sorting list (i.e., the module feature value that is first in the sorting list), then the left and right indices of the initial interval window at this time can be set to the index of the first element in the feature value sorting list, that is, the left and right indices of the initial interval window at this time are set to the first module feature value in the feature value sorting list.

[0086] 406. If the insertion point is located before the feature value of the first module in the feature value sorting list, set the left index of the initial interval window to the index of the previous sorting position of the insertion point, and the right index to the index of the insertion point.

[0087] Based on step 402, when the insertion point is located before the feature value of the first module in the feature value sorting list, the left index of the initial interval window is set to the index of the previous sorting position of the insertion point, and the right index is the index of the insertion point.

[0088] In one specific embodiment, if the interval window feature value fv, i.e. the insertion point, is before the first element (module feature value) of the feature value sorting list, then the left index of the initial interval window at this time can be set to the index of the interval position before the insertion point, and the corresponding right index can be set to the index of the insertion point.

[0089] The interpretable defect prediction method provided in this embodiment can effectively guide the initial interval window and the corresponding left and right indices, providing conditions for finding the final interval window in the future, and reducing the workload.

[0090] Figure 5 illustrates an embodiment of obtaining the target interval window based on Figure 4. For ease of understanding, please refer to Figure 5, which is a flowchart illustrating another interpretable defect prediction method provided by an embodiment of this application. It includes steps 501-507.

[0091] 501. Set the preset interval window value.

[0092] To obtain the target interval window from the initial interval window, it is necessary to ensure that there are enough code modules within the initial interval window for estimation. Therefore, a reference interval window value, i.e., a preset interval window value, needs to be set. Furthermore, this preset interval window value can be understood as the interval size θ, which will be described in detail later for ease of understanding. It should be noted that the preset interval window value is used to represent the ratio of the preset total number of lines of code modules to the total number of lines of all code modules.

[0093] It should also be noted that the execution order of step 501 is not restricted in this embodiment. When executing step 201 of the embodiment shown in FIG2, step 501 can also be executed simultaneously.

[0094] 502. Calculate the ratio of the total number of lines in the code module to be analyzed within the initial interval window to the total number of lines in the total code module, and obtain the ratio of the number of lines to be analyzed.

[0095] Since the total number of lines of code in the code module, the total number of lines of code in the total code module, the total number of defects in the code module, and the total number of defects in the total code module have already been obtained in the aforementioned steps, they will not be repeated here. Specifically, in this embodiment, the ratio of the total number of lines of code in the defective code module located within the initial interval window to the total number of lines of code in the total code module can be calculated to obtain the ratio of the number of lines to be analyzed.

[0096] In one specific embodiment, the ratio of the total number of lines in the defective code module to the total number of lines in the total code module can be calculated within the interval window at this time, thereby obtaining the ratio of the number of lines to be analyzed.

[0097] Then, based on the calculation results of different ratios of the number of rows to be analyzed, steps 502-504 or steps 505-507 are executed respectively.

[0098] 503. When the ratio of the number of rows to be analyzed is less than half of the preset interval window value, and any index of the initial interval window is not located at the list boundary of the feature value sorting list, the initial interval window is iterated to move the left index of the initial interval window one unit to the left or move the right index of the initial interval window one unit to the right, so as to satisfy that the feature value of the interval window is located at the center of the iterated initial interval window.

[0099] Furthermore, when the ratio of rows to be analyzed is less than half of the preset interval window value, and no index of the initial interval window is located at the boundary of the feature value sorting list, the initial interval window can be iterated, thereby shifting the left index of the initial interval window one unit to the left. Alternatively, the right index of the initial interval window can be shifted one unit to the right, ultimately satisfying the condition that the feature value of the interval window is located at the center of the iterated initial interval window.

[0100] In one specific embodiment, when the ratio of rows to be analyzed is less than half of the preset interval window value, and an index (left or right index) of the interval window is not located at a boundary (left or right boundary) of the feature value sorting list, the index can be adjusted. Specifically, the interval window can be expanded. For example, if the interval window is the initial interval window (the interval window that has not yet been iterated), the right index of the initial interval window can be moved one unit to the right in the feature value sorting list. Alternatively, the left index of the initial interval window can be moved one unit to the left in the feature value sorting list. It should be noted that moving to the right or left is acceptable, but the feature value of the interval window must be located at the center of the initial interval window after iteration.

[0101] Furthermore, if the interval window at this time is the initial interval window after iteration, it can be expanded further, and step 504 can be executed.

[0102] 504. Calculate the ratio of the number of rows to be analyzed in the initial interval window after iteration, until the ratio of the number of rows to be analyzed is not less than half of the preset interval window value, and any index of the initial interval window is located at the list boundary of the feature value sorting list. Stop the iteration and determine the initial interval window after iteration as the target interval window.

[0103] Based on step 503, after completing one iteration, the ratio of the number of rows to be analyzed in the initial interval window after the iteration can be calculated until the ratio of the number of rows to be analyzed is not less than half of the preset interval window value, and any index of the initial interval window is located at the list boundary of the feature value sorting list. Then the iteration stops and the initial interval window after the iteration is determined as the target interval window.

[0104] In one specific embodiment, after completing one iteration, the ratio of the number of rows to be analyzed in the initial interval window after the iteration needs to be recalculated. The specific calculation method can be found in the steps described above. If the ratio of the number of rows to be analyzed in the initial interval window after the iteration is still less than half of the preset interval window value, and no index of the initial interval window is located at the boundary of the feature value sorting list, then step 503 continues. However, when the ratio of the number of rows to be analyzed in the initial interval window after the iteration is not less than half of the preset interval window value (i.e., not less than θ*0.5), and an index of the initial interval window after the iteration reaches the corresponding boundary of the list, the iteration stops. The initial interval window after the iteration at this point can be determined as the target interval window.

[0105] Furthermore, in other feasible technical solutions, the following situations also exist: If the right index of the iterated initial interval window reaches the right boundary of the feature value sorting list, then the left index of the iterated initial interval window is shifted to the right by one unit interval position. Or, if the left index of the iterated initial interval window reaches the left boundary of the feature value sorting list, then the right index of the iterated initial interval window is shifted to the left by one unit interval position.

[0106] 505. When the ratio of rows to be analyzed is not less than the preset interval window value, the interval window feature value is used as the window center point, and the interval distance between the left or right index of the initial interval window and the window center point is calculated.

[0107] Based on step 502, when the ratio of rows to be analyzed is not less than the preset interval window value, the feature value of the interval window can be used as the center point of the window, and the interval distance between the left index or right index of the initial interval window and the center point of the window can be calculated.

[0108] In one specific embodiment, when the ratio of the number of rows to be analyzed (the specific calculation process has been described in the preceding steps and will not be repeated here) is not less than the preset interval window value, that is, when the code volume ratio of the initial interval window is not less than θ, the module feature value corresponding to the initial interval window at this time (i.e., the interval window feature value described above) can be calculated. Then, the interval window feature value is used as the window center point to calculate the interval distance from the left index of the initial interval window to the window center point, and the interval distance from the right index of the initial interval window to the window center point.

[0109] Then, based on the judgment result at this time, step 506 or step 507 are executed respectively.

[0110] 506. When the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the feature value sorting list, the left index of the initial interval window is moved to the left by one unit interval position, so that the moved initial interval window is determined as the target interval window.

[0111] Based on step 505, when the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the feature value sorting list, the left index of the initial interval window can be moved to the left by one unit interval position, thereby determining the moved initial interval window as the target interval window.

[0112] In one specific embodiment, when the left endpoint corresponding to the left index is closer to the feature value fv of the interval window and the left index has not reached the corresponding boundary, the left index of the initial interval window is moved to the left by one unit interval position, and the initial interval window after moving one unit to the left is taken as the target interval window.

[0113] 507. When the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the feature value sorting list, the right index of the initial interval window is moved to the right by one unit interval position, so that the moved initial interval window is determined as the target interval window.

[0114] Based on step 505, when the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the feature value sorting list, the right index of the initial interval window can be moved to the right by one unit interval position, thereby determining the moved initial interval window as the target interval window.

[0115] In one specific embodiment, when the right endpoint corresponding to the right index is closer to the feature value fv of the interval window and the right index has not reached the corresponding boundary, the right index of the initial interval window is moved to the right by one unit interval position, and the initial interval window after moving one unit to the right is taken as the target interval window.

[0116] Based on the embodiments shown in Figures 4 and 5, this can be understood as a module feature scoring algorithm. Specifically, please refer to Figure 8, which is a schematic diagram of the algorithm logic for module feature scoring provided by an embodiment of this application. Algorithm 2 details the algorithm for calculating the defect score of code module feature f (module feature name fn, module feature value fv) based on the defect dataset D. The last input of this algorithm is θ, used to determine the interval size in our proposed window-based estimation method. The algorithm is divided into three parts: the first part calculates some basic data (lines 1 to 17), the second part determines the interval for window-based estimation (lines 18 to 29), and the third part calculates the final defect score (lines 30 to 36). It should be noted that Figure 8 includes a feature value list (fvlist), a sorted feature value list (sFvList), a total number of defects (totalDefects), an existence status (isExist), and an insertion index (insIdx). Figure 8 is mainly divided into two parts.

[0117] The first part retrieves unique values ​​of the input features from the input defect dataset D (line 1) and sorts them in ascending order (line 2). The resulting sorted list is denoted as sFvList. The algorithm then determines the total number of defects and the total number of lines of code for all code modules in dataset D (lines 3 and 4). Lines 5 through 17 are used to calculate the initial interval, which is a minimal interval almost centered on fv, located on the list sFvList, or, if fv exceeds the list's boundaries, the interval closest to fv. The left and right indices of the interval are denoted as lIdx and rIdx, respectively. It's important to note that the endpoints of the interval are inclusive. Specifically, it first checks if the input feature value fv exists in sFvList (line 5) and calculates the insertion point if fv is inserted into sFvList (line 6). Then, determine the indices of the initial interval: if fv exists in sFvList, then lIdx and rIdx are both set to the indices of fv in sFvList (line 8); if the insertion point is after the last element of sFvList, then lIdx and rIdx are both set to the last index of sFvList (line 11); if the insertion point is the first position of sFvList, then lIdx and rIdx are both set to the index of the first element (line 13); otherwise, lIdx is set to the index of the position before the insertion point, and rIdx is set to the index of the insertion point (line 15).

[0118] The second part expands the initial interval to ensure it contains enough code modules for efficient window-based estimation. The algorithm iteratively expands the interval by shifting the left index of the initial interval to the left or the right index to the right. The iteration stops when either condition (1) or condition (2) is met (where condition 1 is that the proportion of code in the current interval is not less than θ, and condition 2 is that the proportion of code in the current interval window is not less than θ*0.5, and an index of the current interval window reaches the corresponding boundary of the list). It should be noted that condition (2) is used to stop the iteration early to avoid the situation where the interval cannot represent the input feature value. Furthermore, in each iteration, Algorithm 2 first calculates the proportion of code in the current interval (lines 19 to 20), and then decides whether to stop the iteration (lines 21 to 23). If it does not stop, the index continues to be adjusted. When, in condition (1), the left endpoint is closer to fv than the right endpoint and the left index has not reached the corresponding boundary, or, in condition (2), the right index reaches the corresponding boundary, the left index is shifted to the right. Otherwise, Algorithm 2 shifts the right index to the left. The purpose of this index shift is to ensure that the interval is centered around fv as much as possible. Therefore, the interval window must be expanded after each iteration, and the iteration stopping condition must be met after a certain number of iterations.

[0119] The final part first calculates the expected defect rate and the actual defect rate using formulas (4) and (5) (lines 30 to 34). Then, the defect score is calculated using formula (3), representing the output of Algorithm 2 (line 35), where formulas (3), (4), and (5) are described in the embodiment shown in Figure 6.

[0120] It should be noted that the interval increases iteratively, and the current interval is the interval in the iterative process, eventually leading to the target interval window.

[0121] The interpretable defect prediction method provided in this embodiment, through a window-based defect rate estimation method, can effectively obtain the defect score of code feature modules, making the estimated defect rate more reasonable and accurate.

[0122] Figure 6 illustrates an embodiment of obtaining feature defect scores based on Figure 5. For ease of understanding, please refer to Figure 6, which is a flowchart illustrating another interpretable defect prediction method provided by an embodiment of this application. It includes steps 601-605.

[0123] 601. Analyze the defect dataset and filter out defective code modules whose feature values ​​are range window feature values.

[0124] In this embodiment, by analyzing the defect dataset, defective code modules can be selected where the module feature value of each code module is the interval window feature value. It should be noted that the interval window feature value is used to characterize the ratio of the number of defective code modules located within the target interval window centered on the interval window feature value to the total number of all defective code modules in the defect dataset, satisfying a preset ratio.

[0125] In one specific embodiment, for ease of understanding, please refer to Figure 9, which is a schematic diagram of a window-based actual defect rate estimation provided by an embodiment of this application. It should be noted that this embodiment mainly illustrates the principle of the window-based actual defect rate estimation method (the principle of the expected defect rate is similar). For a point (x, y) on the line in Figure 9, x represents the value of code module feature f, and y represents the total number of defects in all code modules in the defect dataset where the value of code module feature f is x. It is worth noting that for some code module feature f values, due to the limited size of the defect dataset used, there may not be a corresponding point in the figure. When using the window-based estimation method, calculating the defect score of code module feature f as fv involves the following steps:

[0126] First, ScoringDP defines an interval centered at fv (an interval is a special case of a window, i.e., the target interval window), denoted as intvl, to represent the module feature value fv. In this step, it is necessary to ensure that the dataset contains a sufficient number of code modules whose feature f values ​​fall within this interval.

[0127] It should be added that the target interval window is determined based on the dataset analysis. For example, for a defective code module, its code module feature f has a value of fv. Then, an interval centered on fv is determined, so that the number of code modules in the dataset whose code module feature f is within this interval accounts for 20% of the total number of code modules in the dataset.

[0128] Then, steps 602-605 can be executed.

[0129] 602. Based on the number of defective code modules and the total number of all defective code modules, determine the total number of lines of code modules, the total number of lines of code modules, the total number of defects in code modules, and the total number of defects in code modules.

[0130] In this embodiment, step 602 is similar to step 302 in Figure 3 above. However, it should be noted that in this embodiment, the total number of lines of code modules, the total number of lines of code modules, the total number of code module defects, and the total number of code module defects can also be determined based on the number of defective code modules and the total number of all defective code modules.

[0131] 603. Input the total number of lines of code in the code module and the total number of lines of code in the estimated expected defect rate calculation formula to obtain the expected defect rate.

[0132] Based on step 602, ScoringDP can input the total number of lines of code in a module and the total number of lines of code in all modules into the formula for estimating the expected defect rate, thereby obtaining the expected defect rate. It should be noted that the formula for estimating the expected defect rate is:

[0133] Where expected_defect_ratio is the expected defect rate, cm is the defect code module in the defect dataset, D is the defect dataset, f is the code module feature, intvl is the interval window feature value, LOC is the number of lines of code in the defect code module in the defect dataset, Sum({cm.LOC|cm.f=intvl,cm∈D}) is used to represent the total number of lines of code modules accumulated when the module feature value of the defect code module in the defect dataset is the interval window feature value, and Sum({cm.LOC|cm∈D}) is used to represent the total number of lines of code modules.

[0134] 604. Input the total number of defects in the code modules and the total number of defects in the total code modules into the formula for calculating the estimated actual defect rate to obtain the actual defect rate.

[0135] Based on step 602, ScoringDP can input the total number of defects in the code modules and the total number of defects in the total number of code modules into the formula for estimating the actual defect rate, thereby obtaining the actual defect rate. It should be noted that the formula for estimating the actual defect rate is:

[0136] Where actual_defect_ratio is the actual defect rate, defects is the number of defects in the defect code modules in the defect dataset, Sum({cm.defects|cm.f=intvl,cm∈D}) is used to represent the total number of code module defects accumulated when the module feature value of the defect code module in the defect dataset is the feature value of the interval window, and Sum({cm.defects|cm∈D}) is used to represent the total number of code module defects.

[0137] 605. Subtract the expected defect rate from the actual defect rate to obtain the characteristic defect score.

[0138] Then, the actual defect rate can be subtracted from the expected defect rate to obtain the characteristic defect score. Specifically,

[0139] defect_score=actual_defect_ratio—expected_defect_ratio, (Formula 3);

[0140] Where defect_score is the feature defect score. Finally, ScoringDP uses these two defect rates to calculate the defect score of feature f of the code module using formula (3). Therefore, based on the above window-based estimation method, the defect score can be reasonably calculated for features with any value (including those feature values ​​that do not appear in the defect dataset used).

[0141] Referring to the embodiment shown in Figure 6 and Algorithm 1, see Figure 7, which is a schematic diagram of the algorithm logic for code module scoring provided in this application embodiment. As shown in Figure 7, the working process of Algorithm 1 is as follows: It first calculates the defect score of each feature of the input code module cm sequentially (rows 1 to 7). The parameters required for calculating the feature defect score of each code module to be analyzed, and the algorithm for obtaining these parameters, are detailed in Algorithm 2. Subsequently, it derives the defect score of cm by summing the feature defect scores of these code module features (row 8). Finally, it sums the feature defect scores. This summing method enhances the interpretability of ScoringDP.

[0142] In Figure 7, it should be noted that fn is the feature name, fv is the feature value, and fns are multiple feature names. Further details will not be elaborated here.

[0143] Furthermore, for code modules, a higher defect score indicates a higher defect density. Therefore, quality assurance resources should be prioritized for allocation to code modules with higher defect scores to improve resource utilization. To interpret the prediction results for a specific code module, ScoringDP uses the model feature defect score calculated during the prediction process as feature importance. Refer to Figure 10, which is an example diagram illustrating the interpretation of feature importance in ScoringDP according to an embodiment of this application.

[0144] Figure 10 illustrates an example of ScoringDP interpretation, where the horizontal axis represents the defect score of a feature (i.e., the calculated defect score), and the vertical axis represents the code module features sorted in descending order of defect score (along with their values, i.e., the true values). For the various features described above, please refer to Figure 11. Figure 11 is a bilingual (English and Chinese) representation of the feature names for a module feature.

[0145] Specifically, the defect score of a single code module feature can be interpreted in two ways: sign and magnitude. A positive (or negative) defect score indicates that the actual defect rate is higher than the expected defect rate, meaning that the code module has more (or fewer) defects than expected when the value of code module feature X is x. A larger absolute value of the defect score indicates the degree to which the total number of defects in the code module is more (or fewer) than expected. Based on the ranked features, developers can make evidence-based decisions: to reduce the defect predisposition of a code module, developers should modify the code module to prioritize changing the features with larger defect scores, i.e., the features ranked higher in Figure 10. The specific module feature names can be found in Figure 11.

[0146] The interpretable defect prediction method provided in this embodiment has the following three advantages: (1) The interpretation is relatively intuitive: ScoringDP's interpretation comes directly from its internal decision-making process and does not rely on external interpretation methods (such as LIME and SHAP), which improves the accuracy and stability of the interpretation. (2) It is highly operable: Developers can take more effective actions to reduce defects based on the defect scores of module features, such as modifying code to prioritize changing features with higher defect scores. (3) It has excellent predictive performance: According to experiments, ScoringDP outperforms or is equivalent to existing state-of-the-art methods on most performance metrics.

[0147] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0148] Please refer to Figure 12, which is a schematic diagram of the structure of an interpretable defect prediction system provided in an embodiment of this application.

[0149] The acquisition unit 1201 is used to acquire each code module to be analyzed in the software project to be analyzed, and to determine the defect dataset and code module features corresponding to the code module to be analyzed; wherein, the defect dataset is constructed from the historical data of the code module to be analyzed in the software project to be analyzed, and the code module features are used to characterize the numerical features of the code module to be analyzed;

[0150] The determining unit 1202 is configured to, when the module feature value of any code module feature in any code module to be analyzed is a target feature value, determine the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed based on the defect dataset; wherein, the expected defect rate is the ratio of the total number of lines of code modules in all defective code modules in the defect dataset whose module feature values ​​are the target feature value to the total number of lines of code modules in all defective code modules in the defect dataset, and the actual defect rate is the ratio of the total number of defects in all defective code modules in the defect dataset whose module feature values ​​are the target feature value to the total number of defects in all defective code modules in the defect dataset; the defect dataset includes the total number of lines of code modules and the total number of defects in code modules corresponding to different module feature values, and the total number of lines of code modules and the total number of defects in code modules corresponding to all defective code modules; the defective code module is the code module in the defect dataset;

[0151] The calculation unit 1203 is used to calculate the difference between the actual defect rate and the expected defect rate to obtain the feature defect score corresponding to the feature of any code module in any code module.

[0152] The calculation unit 1203 is also used to calculate the feature defect scores of all code module features in any code module to be analyzed, so as to obtain the module defect score of any code module to be analyzed; wherein, the module defect score is used to characterize the defect tendency of the code module to be analyzed.

[0153] The sorting unit 1204 is used to sort all the module defect scores after obtaining the module defect scores of all code modules to be analyzed, so as to process all code modules to be analyzed in descending order of all module defect scores.

[0154] Please refer to Figure 13 below. Figure 13 is a schematic diagram of the structure of an interpretable defect prediction device provided in an embodiment of this application. The interpretable defect prediction device 1300 may include:

[0155] Central processing unit 1301, memory 1305, input / output interface 1304, wired or wireless network interface 1303, and power supply 1302;

[0156] Memory 1305 is either a short-term storage memory or a persistent storage memory;

[0157] The central processing unit 1301 is configured to communicate with the memory 1305 and execute instructions in the memory 1305 to perform the interpretable defect prediction method in any of the embodiments shown in Figures 2 to 6.

[0158] This application also provides a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a line. The at least one processor is used to run computer programs or instructions to execute the interpretable defect prediction method in any of the embodiments shown in Figures 2 to 6.

[0159] This application also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the interpretable defect prediction method in any of the embodiments shown in Figures 2 to 6.

[0160] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the interpretable defect prediction method in any of the embodiments shown in Figures 2 to 6.

[0161] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0162] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0163] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0164] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0165] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An interpretable defect prediction method, characterized in that, The method includes: A defect dataset is constructed based on historical data of the software project to be analyzed. Each code module to be analyzed in the software project and the code module features corresponding to each code module to be analyzed are obtained from the current data of the software project to be analyzed. The code module features are used to characterize the numerical features of the code module to be analyzed. If the module feature value of any code module feature in any code module to be analyzed is the target feature value, the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed are determined according to the defect dataset; wherein, the expected defect rate is the ratio of the total number of lines of code modules in all defective code modules in the defect dataset whose module feature values ​​are the target feature value to the total number of lines of code modules in all defective code modules in the defect dataset, and the actual defect rate is the ratio of the total number of defects in all defective code modules in the defect dataset to the total number of defects in all defective code modules in the defect dataset; the defect dataset includes the total number of lines of code modules and the total number of defects in code modules corresponding to different module feature values, as well as the total number of lines of code modules and the total number of defects in code modules corresponding to all defective code modules; the defective code module is the code module in the defect dataset; Calculate the difference between the actual defect rate and the expected defect rate to obtain the feature defect score of any code module feature in any code module; Calculate the feature defect scores of all code module features in any code module to be analyzed to obtain the module defect score of any code module to be analyzed; wherein, the module defect score is used to characterize the defect tendency of the code module to be analyzed; After obtaining the module defect scores of all code modules to be analyzed, the module defect scores are sorted so that all code modules in the software project to be analyzed are processed in descending order of the module defect scores.

2. The interpretable defect prediction method according to claim 1, characterized in that, The step of obtaining each code module to be analyzed and the corresponding code module features from the current data of the software project to be analyzed includes: The software project to be analyzed is analyzed line by line. Each line of the software project to be analyzed is set as any code module to be analyzed, and the characteristics of the code module contained in any code module to be analyzed are determined.

3. The interpretable defect prediction method according to claim 1, characterized in that, The step of determining the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed based on the defect dataset includes: Based on the defect dataset, obtain the total number of lines of code in the code module, the total number of lines of code in the total code module, the total number of defects in the code module, and the total number of defects in the total code module; Input the total number of lines of code in the code module and the total number of lines of code in the expected defect rate calculation formula to obtain the expected defect rate; wherein, the expected defect rate calculation formula is: The expected_defect_ratio is the expected defect rate, cm is the defect code module in the defect dataset, D is the defect dataset, f is the feature of any code module, fv is the target feature value, LOC is the number of lines of code in the defect code module in the defect dataset, Sum({cm.LOC|cm.f=fv,cm∈D}) is used to represent the total number of lines of code in the defect code module when the module feature value of the defect code module in the defect dataset is the target feature value, and Sum({cm.LOC|cm∈D}) is used to represent the total number of lines of code in the total code module. The actual defect rate is obtained by inputting the total number of defects in the code module and the total number of defects in the total code modules into the actual defect rate calculation formula; wherein, the actual defect rate calculation formula is: The actual_defect_ratio is the actual defect rate, the defects is the number of defects in the defect code modules in the defect dataset, and Sum({cm.defects|cm.f=fv,cm∈D}) is used to represent the total number of code module defects accumulated when the module feature value of the defect code module in the defect dataset is the target feature value. Sum({cm.defects|cm∈D}) is used to represent the total number of code module defects.

4. The interpretable defect prediction method according to claim 1, characterized in that, The step of determining the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed based on the defect dataset includes: Analyze the defect dataset and filter out defective code modules whose module feature value is an interval window feature value; wherein, the interval window feature value is used to characterize that, in the defect dataset, the ratio of the number of defective code modules located in the target interval window centered on the interval window feature value to the total number of all defective code modules in the defect dataset satisfies a preset number ratio. Based on the number of defective code modules and the total number of all defective code modules, determine the total number of lines of code modules, the total number of lines of code modules, the total number of defects in code modules, and the total number of defects in code modules; Input the total number of lines of code in the code module and the total number of lines of code in the total code module into the formula for estimating the expected defect rate to obtain the expected defect rate; wherein, the formula for estimating the expected defect rate is: The expected_defect_ratio is the expected defect rate, cm is the defect code module in the defect dataset, D is the defect dataset, f is the code module feature, intvl is the interval window feature value, LOC is the number of lines of code in the defect code module in the defect dataset, Sum({cm.LOC|cm.f=intvl,cm∈D}) is used to represent the total number of lines of code modules accumulated when the module feature value of the defect code module in the defect dataset is the interval window feature value, and Sum({cm.LOC|cm∈D}) is used to represent the total number of lines of code modules; The total number of defects in the code modules and the total number of defects in the total code modules are input into the formula for estimating the actual defect rate to obtain the actual defect rate; wherein, the formula for estimating the actual defect rate is: The actual_defect_ratio is the actual defect rate, the defects is the number of defects in the defect code modules in the defect dataset, and Sum({cm.defects|cm.f=intvl,cm∈D}) is used to represent the total number of code module defects accumulated when the module feature value of the defect code module in the defect dataset is the interval window feature value. Sum({cm.defects|cm∈D}) is used to represent the total number of code module defects.

5. The interpretable defect prediction method according to claim 4, characterized in that, Before filtering out defective code modules whose module feature values ​​are interval window feature values, the method further includes: Obtain the module feature values ​​of all code module features in any defect code module from the defect dataset, sort all module feature values ​​in ascending order, and obtain a feature value sorting list corresponding to all module feature values; In the feature value sorting list, check whether the interval window feature value exists among all the module feature values; If the interval window feature value exists in the feature value sorting list, the interval window feature value is used as the insertion point in the feature value sorting list to determine the index of the initial interval window; wherein, the left and right indices of the initial interval window are the interval positions where the module feature value in the feature value sorting list is located, and the insertion point is the center of the initial interval window; or, If the insertion point is located after the interval position of the feature value of the interval window and the last module feature value of the feature value sorting list, then the left and right indices of the initial interval window are set to the last module feature value of the feature value sorting list; or, If the insertion point is located within the interval between the feature value of the interval window and the feature value of the first module in the feature value sorting list, then the left and right indices of the initial interval window are set to the feature value of the first module in the feature value sorting list; or, If the insertion point is located before the interval position of the feature value of the interval window that is located before the first module feature value in the feature value sorting list, the left index of the initial interval window index is set to the index of the previous sorting position of the insertion point, and the right index is the index of the insertion point.

6. The interpretable defect prediction method according to claim 5, characterized in that, The methods for obtaining the target interval window include: Set a preset interval window value; wherein, the preset interval window value is used to represent the ratio of the preset total number of lines of code module to the preset total number of lines of code module; Calculate the ratio of the total number of lines in any defective code module located within the initial interval window to the total number of lines in all code modules, and obtain the ratio of the number of lines to be analyzed; When the ratio of the number of rows to be analyzed is less than half of the preset interval window value, and any index of the initial interval window is not located at the list boundary of the feature value sorting list, the initial interval window is iterated to move the left index of the initial interval window to the left by one unit interval position, or to move the right index of the initial interval window to the right by one unit interval position, so as to satisfy that the feature value of the interval window is located at the center of the iterated initial interval window; The iteration continues until the ratio of the number of rows to be analyzed in the initial interval window after iteration is not less than half of the preset interval window value, and any index of the initial interval window is located at the list boundary of the feature value sorting list. Then the iteration stops and the initial interval window after iteration is determined as the target interval window.

7. The interpretable defect prediction method according to claim 6, characterized in that, Before terminating the iteration and determining the initial interval window after iteration as the target interval window, the method further includes: If the right index of the iterated initial interval window reaches the right boundary of the feature value sorting list, then the left index of the iterated initial interval window is shifted one unit to the right; or, If the left index of the iterated initial interval window reaches the left list boundary of the feature value sorting list, the right index of the iterated initial interval window is shifted to the left by one unit interval position.

8. The interpretable defect prediction method according to claim 6, characterized in that, The method further includes: When the ratio of the number of rows to be analyzed is not less than the preset interval window value, the interval window feature value is used as the window center point, and the interval distance between the left index or right index of the initial interval window and the window center point is calculated. When the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the feature value sorting list, the left index of the initial interval window is shifted to the left by one unit interval position, so that the shifted initial interval window is determined as the target interval window; or, When the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the feature value sorting list, the right index of the initial interval window is moved to the right by one unit interval position, so that the moved initial interval window is determined as the target interval window.

9. An interpretable defect prediction device, characterized in that, The device includes: Central processing unit, memory, input / output interfaces, wired or wireless network interfaces, and power supply; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the interpretable defect prediction method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on a computer, cause the computer to perform the interpretable defect prediction method as described in any one of claims 1 to 8.