An interpretable defect prediction method and related devices

By calculating the difference between the actual and expected defect rates of each code module in the software project, evaluating the defect tendency of the code module and sorting it, the problem of poor interpretability of defect prediction results in the prior art is solved, and the defect repair efficiency and accuracy of interpretation are improved.

CN119883869BActive Publication Date: 2025-05-30HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361564.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-30
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing software defect prediction methods rely on complex machine learning models, resulting in poor interpretability of defect prediction results, reducing developers' trust in explanations and may generate inaccurate or misleading explanations.

Method used

By constructing the defect dataset, the difference between the actual defect rate and the expected defect rate of each code module is calculated to obtain a feature defect score and used to evaluate the defect tendency and defect score of the code module, and then sort it to indicate the importance of the module characteristics to the defect score of the code module.

Benefits of technology

It improves the interpretability and stability of the defect prediction model, enhances developers' trust in explanation, and prioritizes the code modules with higher defect scores, improving the software's defect repair efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883869B_ABST
    Figure CN119883869B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an interpretable defect prediction method and related devices, which are used to interpret the decision-making process in software defect prediction. The method of the embodiments of the present application includes: constructing a defect data set based on the historical data of the software project to be analyzed, and obtaining the code module to be analyzed and the code module features in the software project to be analyzed from the current data of the software project to be analyzed; determining the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed according to the defect data set; calculating the difference between the actual defect rate and the expected defect rate to obtain the feature defect score; calculating the feature defect scores of all code module features in any code module to be analyzed to obtain the module defect score; after obtaining the module defect scores of all code modules to be analyzed, sorting all the module defect scores, and processing all the code modules to be analyzed in the software project to be analyzed in turn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of software analysis and defect prediction, and in particular, to an interpretable defect prediction method and related devices. Background Art

[0002] Interpretability is a key aspect of practical defect prediction. Existing workload-aware software defect prediction (predicting whether a code module has defects) methods use a defect dataset (features of the code module and the number of defects in the module) to train a defect prediction model. This model is usually a complex machine learning model, such as a deep learning model. Finally, the code modules are sorted based on the prediction results. However, the black-box decision-making process of these models leads to poor interpretability of the prediction results.

[0003] To improve the interpretability of the defect prediction model, in recent years, researchers have begun to adopt model-agnostic explanation methods (such as local interpretable model-agnostic explanations (LIME) or SHapley Additive exPlanations (SHAP)) to explain individual prediction results. For a new code module, the defect prediction model will predict its defect probability. Then, a model-agnostic explanation method (such as LIME) is used to explain the prediction result. These explanation methods generate explanations based on complex perturbation mechanisms.

[0004] Specifically, to generate the feature importance for a specific prediction, these methods usually modify the instance to be explained and observe the change in the model output. The explanation method will output the contribution degree of each feature to the prediction result, that is, the feature importance. However, these explanation methods have some problems: First, they regard the prediction model as a black box, reducing the trust of developers in the explanation; second, these methods rely on complex perturbation mechanisms, which may generate inaccurate or even misleading explanations; finally, the generated explanations may be unstable due to randomness. Summary of the Invention

[0005] The embodiments of the present application provide an interpretable defect prediction method and related devices for interpretably describing the decision-making process in software defect prediction.

[0006] The first aspect of the embodiments of the present application provides an interpretable defect prediction method, including:

[0007] Construct a defect dataset based on the historical data of the software project to be analyzed, and obtain each code module to be analyzed in the software project to be analyzed and the code module features corresponding to each code module to be analyzed from the current data of the software project to be analyzed; wherein, the code module features are used to characterize the numerical features of the code module to be analyzed.

[0008] If the module feature value of any code module feature in any code module to be analyzed is the target feature value, determine the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed according to the defect dataset; wherein, the expected defect rate is the ratio of the total number of code module lines of the defective code modules with the module feature value being the target feature value in the defect dataset to the total number of code module lines of all defective code modules in the defect dataset, and the actual defect rate is the ratio of the total number of code module defects of the defective code modules with the module feature value being the target feature value in the defect dataset to the total number of code module defects of all defective code modules in the defect dataset; the defect dataset includes the total number of code module lines and the total number of code module defects corresponding to different module feature values, as well as the total number of code module lines and the total number of code module defects corresponding to all defective code modules; the defective code modules are the code modules in the defect dataset.

[0009] Calculate the difference between the actual defect rate and the expected defect rate to obtain the feature defect score of any code module feature in any code module.

[0010] Calculate the feature defect scores of all code module features in any code module to be analyzed to obtain the module defect score of any code module to be analyzed; wherein, the module defect score is used to characterize the defect tendency of the code module to be analyzed.

[0011] After obtaining the module defect scores of all code modules to be analyzed, sort all the module defect scores, and process all the code modules to be analyzed in the software project to be analyzed in descending order of all the module defect scores in turn.

[0012] A second aspect of the embodiments of the present application provides an interpretable defect prediction device, including:

[0013] A central processing unit, a memory, an input / output interface, a wired or wireless network interface, and a power supply;

[0014] The memory is a transient storage memory or a persistent storage memory;

[0015] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the interpretable defect prediction method described in the first aspect.

[0016] In the third aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the interpretable defect prediction method described in the first aspect.

[0017] In the fourth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes instructions that, when run on a computer, cause the computer to execute the interpretable defect prediction method described in the first aspect.

[0018] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: Through an interpretable defect prediction method disclosed in the embodiments of the present application, the defect tendency and defect score of each code module in a software project are evaluated by calculating the deviation between the actual defect rate and the expected defect rate of the code module, and the defect scores are sorted, thereby indicating the importance of module features for the defect scores of the code module and interpretably describing the evaluation process, improving the accuracy and stability of the interpretation. At the same time, developers can adjust the code modules in sequence according to the sorting order, improving the defect repair efficiency of the software. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic diagram of the system architecture of ScoringDP disclosed in the embodiments of the present application;

[0021] Figure 2 It is a schematic flowchart of an interpretable defect prediction method disclosed in the embodiments of the present application;

[0022] Figure 3 It is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application;

[0023] Figure 4 It is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application;

[0024] Figure 5 It is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application;

[0025] Figure 6 It is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application;

[0026] Figure 7 Schematic diagram of the algorithm logic for scoring code modules disclosed in the embodiments of the present application;

[0027] Figure 8 Schematic diagram of the algorithm logic for scoring module features disclosed in the embodiments of the present application;

[0028] Figure 9 Schematic diagram of the estimation of the actual defect rate based on a window disclosed in the embodiments of the present application;

[0029] Figure 10 Exemplary diagram for explaining the feature importance of ScoringDP disclosed in the embodiments of the present application;

[0030] Figure 11 Diagram for comparing the Chinese and English names of the feature names of a module feature;

[0031] Figure 12 Schematic diagram of the structure of an interpretable defect prediction system disclosed in the embodiments of the present application;

[0032] Figure 13 Schematic diagram of the structure of an interpretable defect prediction device disclosed in the embodiments of the present application. Detailed implementation manners

[0033] Software defect prediction is an important research area in software engineering, aiming to predict which modules may have defects by analyzing the features of code modules (such as code complexity, number of lines of code, etc.). Traditional defect prediction methods usually rely on complex machine learning models (such as deep learning models) to improve prediction performance. However, these models often lack interpretability, resulting in a reduced level of trust in the prediction results by developers, and thus affecting their practical applications.

[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0035] Please refer to Figure 1 , Figure 1 Schematic diagram of the system architecture of ScoringDP disclosed in the embodiments of the present application. Consisting of Figure 1It can be seen that the embodiment of the present application is mainly a method for defect prediction by scoring each code module. Specifically, in order to sort the code modules in a software project, the Scoring Defect Prediction (ScoringDP) system first calculates a score for each code module, and then sorts the modules according to these scores. These scores are called defect scores. Figure 1 The process of calculating defect scores for ScoringDP, where the defect score of a code module is the aggregation of the defect scores of its features. The defect score of a module feature can indicate the importance of the module feature for the defect score of the code module. The defect dataset can be constructed based on the historical change data of the code module (analyzed from the code repository); if the historical data is insufficient, data from similar projects can be used to construct the dataset, and specific details are not limited here. It should also be noted that module features mainly describe various attributes of the code, such as cohesion, coupling, complexity, or size, etc. For example, Average Method Complexity (AMC), Average McCabe's Cyclomatic Complexity (AVG CC), or Lines of Code (LOC), etc., and specific details are not elaborated here. For details, please refer to Figure 11 , Figure 11 a Chinese-English comparison chart of the feature names of a module feature. For the convenience of description, it will not be elaborated further hereinafter.

[0036] To solve the above-mentioned technical problems, please refer to Figure 2 , Figure 2 is a schematic flowchart of an interpretable defect prediction method disclosed in the embodiment of the present application. It includes steps 201 - step 205.

[0037] 201. Construct a defect dataset based on the historical data of the software project to be analyzed, and obtain each code module to be analyzed in the software project to be analyzed and the code module features corresponding to each code module to be analyzed from the current data of the software project to be analyzed.

[0038] In this embodiment, a defect dataset is constructed based on the historical data of the software project to be analyzed, and then, by analyzing the current data of the software project to be analyzed, all the code modules to be analyzed in the software project to be analyzed and the code module features of each code module to be analyzed are obtained. It should be noted that the code module features are used to characterize the numerical features of the code module to be analyzed. The software project to be analyzed is the software project that needs to be analyzed in this embodiment, and specifically, the defects of this software project need to be analyzed and predicted.

[0039] In one specific embodiment, in combination with Figure 1 As shown, in the process of analyzing a software project, it is necessary to construct a data set in combination with the historical data of the software project or the data of similar projects. Among them, the data set can be understood as a kind of defect data set, which is a tabular data set, and each row represents a code module. At the same time, the code module has one or more numerical feature descriptions, such as the number of lines of code LOC and the cyclomatic complexity, etc., and specific limitations are not made here. Furthermore, it can be understood that if each row in the defect data set D represents a code module cm, and the code module cm is described by m numerical features, then the defect data set D includes m features and a corresponding label, representing the defect data of the module. Therefore, it can be understood that the defect data set includes the total number of lines of code modules, the total number of defects in code modules. And the total number of lines of all code modules and the total number of defects in all code modules in the software project.

[0040] 202. When the module feature value of any code module feature in any code module to be analyzed is the target feature value, according to the defect data set, determine the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed.

[0041] Furthermore, when it satisfies that the module feature value of any code module feature in any code module to be analyzed in the software project to be analyzed is the target feature value, then according to the total number of lines of code modules, the total number of lines of all code modules, the total number of defects in code modules, and the total number of defects in all code modules in the defect data set, analyze each code module feature in each code module to be analyzed, and then obtain the actual defect rate and the expected defect rate of each code module feature. It should be noted that the expected defect rate is the ratio of the total number of lines of code modules of defective code modules to the total number of lines of all defective code modules in the defect data set, and the actual defect rate is the ratio of the total number of defects in defective code modules in the defect data set to the total number of defects in all defective code modules in the defect data set. The defective code module is the code module in the defect data set, and the code module to be analyzed is the code module in the software project to be analyzed.

[0042] In one specific embodiment, ScoringDP can calculate the defect score of a code module feature by quantifying the deviation between two defect rates: the actual defect rate and the expected defect rate. For the module feature value fv of a code module feature f, the expected defect rate is defined as the ratio of the total number of code lines of the code module when the module feature value of the code module feature f is fv in the defect dataset to the total number of code lines of all code modules. The calculation of the expected defect rate is based on an intuitive assumption that defects in a software project are evenly distributed among the code lines of all code modules within the project. Similarly, the actual defect rate is defined as the ratio of the total number of defects in those code modules when the module feature value of the code module feature f is fv to the total number of defects in all code modules. For ease of understanding, the definitions of the expected defect rate and the actual defect rate will not be described further hereinafter.

[0043] Furthermore, since the calculation of the expected defect rate is based on an intuitive assumption that defects in a software project are evenly distributed among the code lines of all code modules within the project. However, there is an implicit assumption in the scoring of the above code module features: the defect dataset used contains an infinite number of code modules. However, in an actual scenario, the size of the defect dataset is finite, so there may be at least two cases. For ease of understanding, reference may be made respectively to Figure 3 the embodiment shown in Figures 4 to 6 the embodiment shown.

[0044] 203. Calculate the difference between the actual defect rate and the expected defect rate to obtain the characteristic defect score of any code module feature in any code module.

[0045] After obtaining the actual defect rate and the expected defect rate of the code module feature, the difference between the actual defect rate and the expected defect rate can be calculated, thereby obtaining the characteristic defect score of any code module feature in any code module.

[0046] In one specific embodiment, the defect score of the code module feature is defined as the difference between the actual defect rate and the expected defect rate. Furthermore, by subtracting the expected defect rate from the actual defect rate, the characteristic defect score can be obtained.

[0047] 204. Calculate the characteristic defect scores of all code module features in any code module to be analyzed to obtain the module defect score of any code module to be analyzed.

[0048] Thus, based on step 202 - step 203, the characteristic defect scores of all code module features in any code module to be analyzed can be calculated, and then the module defect score of any code module to be analyzed can be obtained. It should be noted that the module defect score is used to characterize the defect tendency of the code module to be analyzed.

[0049] In one specific embodiment, when analyzing all the code module features in a code module to be analyzed, the feature defect scores of all the code module features within the code module to be analyzed can be accumulated to obtain the module defect score of the code module to be analyzed.

[0050] 205. After obtaining the module defect scores of all the code modules to be analyzed, sort all the module defect scores, and in the order from largest to smallest of all the module defect scores, process all the code modules to be analyzed in the software project to be analyzed one by one.

[0051] Thus, after obtaining the module defect scores of all the code modules to be analyzed in the software project to be analyzed, all the module defect scores can be sorted, and in the order from largest to smallest of all the module defect scores, in the order of priority, process all the code modules to be analyzed in the software project to be analyzed one by one. Furthermore, more effective actions can be taken according to the defect scores of the code module features to reduce defects, such as modifying the code to preferentially change the features with higher defect scores.

[0052] In one specific embodiment, the defect scores of all the code module features are summarized by summation, and this simple summarization method can enhance the interpretability of ScoringDP.

[0053] Through an interpretable defect prediction method disclosed in this embodiment, by calculating the deviation between the actual defect rate and the expected defect rate of each code module in the software project, the defect tendency and defect score of the code module are evaluated, and the defect scores are sorted, thereby indicating the importance of the module features for the defect scores of the code module and interpretably describing the evaluation process, improving the accuracy and stability of the interpretation. At the same time, developers can adjust the code modules one by one in the sorting order, improving the defect repair efficiency of the software.

[0054] Figure 3 For Figure 2 one implementation manner of step 201 - step 202, specifically obtaining the actual defect rate and the expected defect rate of the code module features. For easy understanding, please refer to Figure 3 , Figure 3 is a flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application. It includes steps 301 - step 304.

[0055] 301. Analyze the software project to be analyzed line by line, set each line of the software project to be analyzed as any code module to be analyzed, and determine the code module features included in any code module to be analyzed.

[0056] In this embodiment, it is mainly one of the realizable ways from step 201 to step 202. Specifically, by analyzing the software project to be analyzed line by line, each line in the software project to be analyzed can be set as a corresponding code module to be analyzed (expressed as any code module to be analyzed). Thus, the code module features included in any module to be analyzed can be determined.

[0057] In one specific embodiment, by analyzing the software project to be analyzed line by line, the code module to be analyzed represented by each line in the software project to be analyzed and the corresponding code module features are obtained.

[0058] 302. Based on the defect dataset, obtain the total number of code lines of the code module, the total number of code lines of all code modules, the total number of defects in the code module, and the total number of defects in all code modules.

[0059] It should be noted that steps 302 to 304 in this embodiment are a calculation based on the defect score of the module feature under an ideal state. Among them, in step 302, the calculation of the expected defect rate is based on an intuitive assumption, that is, the defects in the defect dataset are evenly distributed among the code lines of all code modules. Therefore, when the defect data in the defect dataset is evenly distributed among all defective code modules, obtain the total number of code lines of the code module, the total number of code lines of all code modules, the total number of defects in the code module, and the total number of defects in all code modules after parsing the defect dataset.

[0060] Furthermore, in the process of analyzing the defect dataset, since the code module features and related defects are marked on the defective code module, furthermore, one or more code module features of the defective code module and a corresponding module label can be obtained, and the defect quantity of the code module is noted on this label. Thus, the total number of code lines of the code module, the total number of code lines of all code modules, the total number of defects in the code module, and the total number of defects in all code modules can be obtained.

[0061] 303. Input the total number of code lines of the code module and the total number of code lines of all code modules into the expected defect rate calculation formula to obtain the expected defect rate.

[0062] Then, input the total number of code lines of the code module and the total number of code lines of all code modules into the expected defect rate calculation formula to obtain the expected defect rate. It should be noted that the expected defect rate calculation formula is:

[0063] , formula (1); where, expected_defect_ratio is the expected defect rate, cm is the defective code module in the defect dataset, D is the defect dataset, f is any code module feature, fv is the target feature value, LOC is the number of code lines of the defective code module in the defect dataset, used to represent the total number of code lines of the code module, The total number of lines of code modules accumulated when the module feature value of the defective code module in the defect dataset is the target feature value.

[0064] 304. Input the total number of code module defects and the total number of defects in all code modules into the actual defect rate calculation formula to obtain the actual defect rate. Corresponding to step 303, the total number of code module defects and the total number of defects in all code modules can be input into the actual defect rate calculation formula to obtain the actual defect rate. It should be noted that the actual defect rate calculation formula is: , formula (2); where actual_defect_ratio is the actual defect rate, and defect is the number of defects of the defective code module in the defect dataset. The total number of defective code modules accumulated when the module feature value of the defective code module in the defect dataset is the target feature value. Used to represent the total number of defects in all code modules.

[0065] In this embodiment, one input of the scoring algorithm is the defect dataset. Each row of the dataset represents a code module, including various features and the number of defects of the code module. This dataset is usually constructed based on the historical data of the project. For example, the features of the code module cm were counted one year ago, and then the number of defects found in it within this year was observed. However, it should also be noted that Figure 3 The illustrated embodiment mainly assumes that under the condition that the defect dataset is large enough, for any given feature value, there are enough code modules with that feature value. In this case, the expected defect rate and the actual defect rate will not be wrongly calculated as zero, and they will change smoothly with the change of the feature value, thus helping to reasonably estimate these two defect rates.

[0066] Through an interpretable defect prediction method disclosed in this embodiment, by calculating the module feature defect score, reasonable expected defect rate and actual defect rate can be effectively generated, and the defect situation of the code module features of the code module to be analyzed can be quantified and visually seen.

[0067] It should be noted in advance that Figure 4 、 Figure 5 and Figure 6 is another embodiment for obtaining the actual defect rate and the expected defect rate of the code module features compared with Figure 3 . Among them, Figure 4 is an embodiment for describing obtaining the initial interval window, and for details, reference can be made to Figure 4 , Figure 4 is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application. It includes steps 401 - step 406.

[0068] 401. Obtain the module feature values of all code module features in any defect code module from the defect data set, sort all the module feature values in ascending order, and obtain a feature value sorted list corresponding to all the module feature values.

[0069] As described above, this embodiment mainly describes determining the interval size of the initial interval window based on the defect data set. Specifically, since Figure 3 in the foregoing embodiment, the module feature values of all code module features have been obtained through the defect data set, therefore, all the module feature values can be sorted in ascending order, so as to obtain a feature value sorted list corresponding to all the module feature values.

[0070] In one specific embodiment, after obtaining the module feature values of all code module features, sort all the module feature values in ascending order, so as to obtain a sorted list corresponding to all the module feature values, that is, a feature value sorted list.

[0071] 402. In the feature value sorted list, check whether there is an interval window feature value among all the module feature values.

[0072] Then, in this feature value sorted list, analyze all the module feature values and screen out whether there is an interval window feature value. It should be noted that the interval window feature value is that the ratio of the number of code modules in a certain interval window to the total number of code modules in the data set should meet a preset interval ratio. For example, 20% or 30%, etc., which is not limited here.

[0073] Furthermore, according to different inspection results, any one of steps 403, 404, 405 or 406 is executed respectively.

[0074] 403. If there is an interval window feature value in the feature value sorted list, use the interval window feature value as the insertion point of the feature value sorted list to determine the index of the initial interval window.

[0075] Based on step 402, when there is an interval window feature value in the feature value sorted list, use the interval window feature value as the insertion point of the feature value sorted list, and determine the index of the initial interval window. It should be noted that the index includes a left index and a right index. The left index and the right index of the initial interval window are the interval positions where the module feature values in the feature value sorted list are located, and the insertion point is the center of the initial interval window.

[0076] In one specific embodiment, if the interval window eigenvalue fv exists in the eigenvalue sorting list, then the left index and the right index of the initial interval window at this time can be set to the index of the interval window eigenvalue fv in the eigenvalue sorting list, wherein it is necessary to satisfy as much as possible that the center of the left index and the right index of the initial interval window is the interval window eigenvalue fv.

[0077] 404. If the insertion point is located after the interval position of the interval window eigenvalue and the last module eigenvalue in the eigenvalue sorting list, the left index and the right index in the index of the initial interval window are set to the last module eigenvalue in the eigenvalue sorting list.

[0078] Based on step 402, when the insertion point is located after the interval position of the interval window feature value and the last module feature value in the feature value sorting list, the left index and the right index in the index of the initial interval window are set to the last module feature value in the feature value sorting list.

[0079] In one specific embodiment, if the interval window eigenvalue fv, that is, the insertion point is after the last element (module eigenvalue) in the eigenvalue sorted list, then the left index and the right index of the initial interval window at this time can be set to the last index in the eigenvalue sorted list, that is, the left index and the right index of the initial interval window at this time are set to the last module eigenvalue in the eigenvalue sorted list.

[0080] 405. If the insertion point is located at the interval position between the interval window eigenvalue and the first module eigenvalue in the eigenvalue sorting list, the left index and the right index in the index of the initial interval window are set to the first module eigenvalue in the eigenvalue sorting list.

[0081] Based on step 402, when the insertion point is located at the interval position between the interval window feature value and the first module feature value in the feature value sorting list, the left index and the right index in the index of the initial interval window are set to the first module feature value in the feature value sorting list.

[0082] In one specific embodiment, if the insertion point is at the first position in the eigenvalue sorting list (i.e., the module eigenvalue corresponding to the first position in the sorting), then the left index and the right index of the initial interval window at this time can be set to the index of the first element in the eigenvalue sorting list, that is, the left index and the right index of the initial interval window at this time are set to the first module eigenvalue in the eigenvalue sorting list.

[0083] 406. If the insertion point is located before the interval position of the first module eigenvalue in the eigenvalue sorting list in the interval window eigenvalue, the left index in the index of the initial interval window is set to the index of the previous sorting position of the insertion point, and the right index is the index of the insertion point.

[0084] Based on step 402, when the insertion point is before the interval position of the eigenvalue of the interval window feature value that is the first module eigenvalue in the sorted list of eigenvalue, the left index in the index of the initial interval window is set to the index of the previous sorted position of the insertion point, and the right index is the index of the insertion point.

[0085] In one specific embodiment, if the interval window feature value fv, that is, the insertion point is before the first element (module eigenvalue) in the sorted list of eigenvalues, then the left index of the initial interval window at this time can be set to the index of the previous interval position of the insertion point, and the corresponding right index is set to the index of the insertion point.

[0086] Through an interpretable defect prediction method disclosed in this embodiment, it can effectively guide the initial interval window and the corresponding left and right indices, providing conditions for finding the final interval window and reducing the workload.

[0087] For Figure 5 , Figure 5 For the embodiment of obtaining the target interval window based on Figure 4 . For easy understanding, please refer to Figure 5 , Figure 5 which is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application. It includes steps 501 - step 507.

[0088] 501. Set a preset interval window value.

[0089] To obtain the target interval window from the initial interval window, it is necessary to ensure that there are enough code modules in the initial interval window for estimation. Therefore, it is necessary to set a reference interval window value, that is, the preset interval window value. Further, the preset interval window value can be understood as the interval size θ, which will be described in detail later for easy understanding. It should be noted that the preset interval window value is used to represent the ratio of the total number of lines of the preset code modules to the total number of lines of all code modules.

[0090] It should also be noted that the execution order of step 501 is not restricted in this embodiment. When executing step 201 of the embodiment shown in Figure 2 , step 501 can also be executed synchronously.

[0091] 502. Calculate the ratio of the total number of lines of the code modules to be analyzed in the initial interval window to the total number of lines of all code modules to obtain the ratio of the lines to be analyzed.

[0092] Since the total number of lines of code modules, the total number of lines of all code modules, the total number of defects in code modules, and the total number of defects in all code modules have been obtained in the foregoing steps, they will not be elaborated here. Specifically, in this embodiment, the ratio of the total number of lines of code modules with defective code within the initial interval window to the total number of lines of all code modules can be calculated to obtain the ratio of lines to be analyzed.

[0093] In one specific embodiment, the ratio of the total number of lines of code modules with defects in the interval window at this time to the total number of lines of all code modules can be calculated to obtain the ratio of lines to be analyzed.

[0094] Furthermore, based on the calculation results of different ratios of lines to be analyzed, steps 502 - 504, or steps 505 - 507 are respectively executed.

[0095] 503. When the ratio of lines to be analyzed is less than half of the preset interval window value and any index of the initial interval window is not at the list boundary of the sorted list of eigenvalue, iterate the initial interval window to move the left index of the initial interval window one unit of interval position to the left, or move the right index of the initial interval window one unit of interval position to the right to make the eigenvalue of the interval window located at the center of the iterated initial interval window.

[0096] Furthermore, when the ratio of lines to be analyzed is less than half of the preset interval window value and any index of the initial interval window is not at the list boundary of the sorted list of eigenvalue, the initial interval window can be iterated, so as to move the left index of the initial interval window one unit of interval position to the left. Or, move the right index of the initial interval window one unit of interval position to the right and finally make the eigenvalue of the interval window located at the center of the iterated initial interval window.

[0097] In one specific embodiment, when the ratio of lines to be analyzed is less than half of the preset interval window value and one index (left index or right index) of the interval window at this time is not at a certain list boundary (left boundary or right boundary) of the sorted list of eigenvalue, the index can be adjusted. Specifically, expand the interval window at this time. For example, if the interval window at this time is the above-mentioned initial interval window (the interval window that has not been iterated), the right index of the initial interval window can be moved one unit to the right in the sorted list of eigenvalue. Or, move the left index of the initial interval window one unit to the left in the sorted list of eigenvalue. It should be noted that either moving to the right or to the left is acceptable, but it is necessary to make the eigenvalue of the interval window located at the center of the iterated initial interval window.

[0098] Furthermore, if the interval window at this time is the iterated initial interval window, it can also be continuously expanded and step 504 is executed.

[0099] 504. Calculate the ratio of the number of lines to be analyzed in the initial interval window after iteration until the ratio of the number of lines to be analyzed is not less than half of the preset interval window value, and when any index of the initial interval window is at the list boundary of the sorted eigenvalue list, abort the iteration and determine the initial interval window after iteration as the target interval window.

[0100] Based on step 503, after one iteration is completed, the ratio of the number of lines to be analyzed in the initial interval window after iteration can be calculated until it meets the condition that the ratio of the number of lines to be analyzed is not less than half of the preset interval window value, and when any index of the initial interval window is at the list boundary of the sorted eigenvalue list, abort the iteration and determine the initial interval window after iteration as the target interval window.

[0101] In one specific embodiment, after one iteration is completed, it is necessary to recalculate the ratio of the number of lines to be analyzed in the initial interval window after iteration. The specific calculation method can refer to the above steps. If the ratio of the number of lines to be analyzed in the initial interval window after iteration at this time is still less than half of the preset interval window value, and any index of the initial interval window is not at the list boundary of the sorted eigenvalue list, then continue to execute step 503. However, when the ratio of the number of lines to be analyzed in the initial interval window after iteration is not less than half of the preset interval window value (i.e., not less than ) and one index of the initial interval window after iteration reaches the corresponding boundary of the list, the iteration stops. The initial interval window after iteration at this time can be determined as the target interval window.

[0102] Furthermore, in other realizable technical solutions, there are also the following situations. If the right index of the initial interval window after iteration reaches the right list boundary of the sorted eigenvalue list, then move the left index of the initial interval window after iteration one unit of interval position to the right. Or, if the left index of the initial interval window after iteration reaches the left list boundary of the sorted eigenvalue list, then move the right index of the initial interval window after iteration one unit of interval position to the left.

[0103] 505. When the ratio of the number of lines to be analyzed is not less than the preset interval window value, taking the eigenvalue of the interval window as the window center point, calculate the interval distance between the left index or the right index of the initial interval window and the window center point.

[0104] Based on step 502, when the ratio of the number of lines to be analyzed is not less than the preset interval window value, the eigenvalue of the interval window can be used as the window center point, and the interval distance between the left index or the right index of the initial interval window and the window center point can be calculated.

[0105] In one specific embodiment, when the ratio of the number of lines to be analyzed (the specific calculation process has been described in the foregoing steps and will not be elaborated here) is not less than the preset interval window value, that is, when the proportion of the code volume in the initial interval window is not less than θ, the module eigenvalue corresponding to the initial interval window at this time (i.e., the interval window eigenvalue described above) can be calculated. Then, taking this interval window eigenvalue as the window center point, the interval distance from the left index of the initial interval window to the window center point and the interval distance from the right index of the initial interval window to the window center point are calculated at this time.

[0106] Furthermore, according to the judgment result at this time, step 506 or step 507 is executed respectively.

[0107] 506. When the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the eigenvalue sorted list, move the left index of the initial interval window one unit of interval position to the left to determine the moved initial interval window as the target interval window.

[0108] Based on step 505, when the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the eigenvalue sorted list, the left index of the initial interval window can be moved one unit of interval position to the left, so as to determine the moved initial interval window as the target interval window.

[0109] In one specific embodiment, when the left endpoint distance corresponding to the left index is closer to the interval window eigenvalue fv and the left index does not reach the corresponding boundary, move the left index of the initial interval window one unit of interval position to the left, and take the initial interval window moved one unit to the left as the target interval window.

[0110] 507. When the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the eigenvalue sorted list, move the right index of the initial interval window one unit of interval position to the right to determine the moved initial interval window as the target interval window.

[0111] Based on step 505, when the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the eigenvalue sorted list, the right index of the initial interval window can be moved one unit of interval position to the right, so as to determine the moved initial interval window as the target interval window.

[0112] In one specific embodiment, when the right endpoint distance interval window eigenvalue corresponding to the right index is closer, and the right index does not reach the corresponding boundary, the right index of the initial interval window is moved one unit of interval position to the right, and the initial interval window after moving one unit to the right is used as the target interval window.

[0113] Based on Figure 4 and Figure 5 the embodiments shown, it can be understood as a scoring algorithm for module features. Specifically, reference can be made to Figure 8 , Figure 8 which is a schematic diagram of the algorithm logic for scoring module features disclosed in the embodiments of the present application. Among them, Algorithm 2 details the algorithm for calculating the defect score of the code module feature f (the module feature name is fn, and the module feature value is fv) based on the defect data set D. The last input of this algorithm is θ, which is used to determine the interval size in the window-based estimation method we proposed. The algorithm is divided into three parts: the first part calculates some basic data (lines 1 to 17), the second part determines the interval for window-based estimation (lines 18 to 29), and the third part calculates the final defect score (lines 30 to 36). It should be added that in Figure 8 there are the feature value list (fvlist, feature value list), the sorted feature value list (sFvList, sorted feature value list), the total number of defects (totalDefects, total number of defects), whether it exists (isExist, is Exist), and the subscript of the insertion position (insIdx, insert index). Among them, Figure 8 is mainly divided into two parts.

[0114] The first part first retrieves the unique values of the input features from the input defect data set D (line 1) and sorts them in ascending order (line 2). The obtained sorted list is denoted as sFvList. Subsequently, the algorithm determines the total number of defects and the total number of lines of code in all code modules in the data set D (lines 3 and 4). The code from lines 5 to 17 is used to calculate the initial interval, which is a minimum interval centered almost on and is located on the list sFvList, or if exceeds the boundary of the list, then take the interval closest to . The left index and right index of the interval are denoted as and respectively. It should be noted that the endpoints of the interval are included. Specifically, first check the input feature value Whether it exists in the sFvList (line 5), and calculate if is inserted into the sFvList, the insertion point (line 6). Then, determine the indices of the initial interval: If exists in the sFvList, then and are both set to its index in the sFvList (line 8); if the insertion point is after the last element of the sFvList, then and are both set to the last index of the sFvList (line 11); if the insertion point is the first position of the sFvList, then and are both set to the index of the first element (line 13); otherwise, is set to the index of the position before the insertion point, is set to the index of the insertion point (line 15).

[0115] The second part expands the initial interval to ensure that it contains enough code modules for effective window-based estimation. The algorithm iteratively expands the interval by moving the left index of the initial interval to the left or the right index to the right. When any one of the conditions (1) or (2) is met (where condition 1 is that the proportion of the code volume in the current interval is not less than , condition 2 is that the proportion of the code volume in the current interval window is not less than , and one of the indices of the current interval window reaches the corresponding boundary of the list), the iteration stops (line 21). It should be noted that condition (2) is used to stop the iteration in advance to avoid the situation where the interval cannot represent the input eigenvalue. Further, in each iteration, the algorithm 2 first calculates the proportion of the code volume in the current interval (lines 19 to 20), and then decides whether to stop the iteration (lines 21 to 23). If not, continue to adjust the indices. When in condition (1), the left endpoint is closer to fv than the right endpoint and the left index has not reached the corresponding boundary, or, in condition (2), the right index reaches the corresponding boundary, then move the left index to the right. Otherwise, the algorithm 2 moves the right index to the left. The purpose of this index movement is to ensure that the interval is centered around fv as much as possible. Therefore, the interval window must be expanded after each iteration, and the iteration stop condition must be met after a certain number of iterations.

[0116] The last part first calculates the expected defect rate and the actual defect rate through Figure 6 Formulas (4) and (5) in the embodiments shown (lines 30 to 34). Subsequently, calculate the defect score through Formula (3), which represents the output of algorithm 2 (line 35), where Formulas (3), (4), and (5) are inFigure 6 as described in the illustrated embodiments.

[0117] It should be noted that the interval becomes larger iteratively. The current interval is the interval during the iterative process, and finally the target interval window is obtained.

[0118] Through an interpretable defect prediction method disclosed in this embodiment, by a window-based defect rate estimation method, the defect score of the code feature module can be effectively obtained, making the estimated defect rate more reasonable and accurate.

[0119] For Figure 6 , Figure 6 is an embodiment for obtaining the feature defect score based on Figure 5 . For easy understanding, please refer to Figure 6 , Figure 6 is a schematic flowchart of another interpretable defect prediction method disclosed in the embodiments of the present application. It includes step 601 - step 605.

[0120] 601. Analyze the defect data set, and screen out the defective code modules whose module feature values of the code module features are the interval window feature values.

[0121] In this embodiment, by analyzing the defect data set, it is possible to screen out the defective code modules under the condition that the module feature value of each code module feature is the interval window feature value. It should be noted that the interval window feature value is used to characterize that in the defect data set, in the target interval window centered on the interval window feature value, the ratio of the number of defective code modules within the target interval window to the total number of all defective code modules in the defect data set satisfies a preset quantity ratio.

[0122] In one specific embodiment, for easy understanding, please refer to Figure 9 , Figure 9 is a schematic diagram of a window-based actual defect rate estimation disclosed in the embodiments of the present application. It should be noted that this embodiment mainly shows the principle of the window-based actual defect rate estimation method (the principle of the expected defect rate is similar). For Figure 9 a point (x, y) on the line in

[0123] First, ScoringDP determines an interval centered on fv (the interval is a special case of the window, i.e., the target interval window), denoted as intvl, to represent the module feature value fv. In this step, it is necessary to ensure that the dataset used contains a sufficient number of code modules whose feature f values fall within this interval.

[0124] It should be added that the target interval window is analyzed based on the dataset. For example, for a defective code module, the value of its code module feature f is fv, and then an interval centered on fv is determined such that the number of code modules in the dataset whose feature f values are within this interval accounts for 20% of the total number of code modules in the dataset.

[0125] Then, steps 602 - 605 can be executed.

[0126] 602. Determine the total number of lines of code modules, the total number of lines of all code modules, the total number of code module defects, and the total number of code module defects of all defective code modules according to the number of defective code modules and the total number of all defective code modules.

[0127] In this embodiment, step 602 is similar to step 302 in the foregoing Figure 3 . However, it should be noted that in this embodiment, the total number of lines of code modules, the total number of lines of all code modules, the total number of code module defects, and the total number of code module defects of all defective code modules can also be determined according to the number of defective code modules and the total number of all defective code modules.

[0128] 603. Input the total number of lines of code modules and the total number of lines of all code modules into the estimated expected defect rate calculation formula to obtain the expected defect rate. Based on step 602, ScoringDP can input the total number of lines of code modules and the total number of lines of all code modules into the estimated expected defect rate calculation formula, thereby obtaining the expected defect rate. It should be noted that the estimated expected defect rate calculation formula is:

[0129] , formula (4); where expected_defect_ratio is the expected defect rate, cm is the defective code module in the defect dataset, D is the defect dataset, f is the code module feature, intvl is the interval window feature value, LOC is the number of lines of code of the defective code module in the defect dataset, used to represent the total number of lines of code modules accumulated when the module feature value of the defective code module in the defect dataset is the interval window feature value, used to represent the total number of lines of all code modules.

[0130] 604. Input the total number of defects in the code module and the total number of defects in all code modules into the estimated actual defect rate calculation formula to obtain the actual defect rate. Based on step 602, ScoringDP can input the total number of defects in the code module and the total number of defects in all code modules into the estimated actual defect rate calculation formula, thereby obtaining the actual defect rate. It should be noted that the estimated actual defect rate calculation formula is:

[0131] , formula (5); where actual_defect_ratio is the actual defect rate, defect is the number of defects in the defective code module in the defect dataset, used to represent the total number of defects in the code module accumulated when the module feature value of the defective code module in the defect dataset is the interval window feature value, used to represent the total number of defects in all code modules.

[0132] 605. Subtract the actual defect rate from the expected defect rate to obtain the characteristic defect score.

[0133] Then, the actual defect rate can be subtracted from the expected defect rate to obtain the characteristic defect score. Specifically, defect_score = actual_defect_ratio - expected_defect_ratio, formula (3); where defect_score is the characteristic defect score. Finally, ScoringDP uses these two defect rates to calculate the defect score of the code module feature f through formula (3). Therefore, based on the above window-based estimation method, for features with any value (including those feature values that do not appear in the used defect dataset), their defect scores can be reasonably calculated.

[0134] Combined with Figure 6 the embodiments shown and Algorithm 1, reference can be made to Figure 7 , Figure 7 which is a schematic diagram of the algorithm logic for scoring code modules disclosed in the embodiments of the present application. As can be seen from Figure 7 , the working process of Algorithm 1 is as follows. It first calculates the defect scores of each feature of the input code module cm in sequence (lines 1 to 7). During the process of calculating the characteristic defect scores of the features of each code module to be analyzed, the required parameters and the algorithms for obtaining the parameters are detailed in Algorithm 2. Subsequently, it derives the defect score of cm by summarizing the characteristic defect scores of the code module features of these code modules to be analyzed (line 8). Finally, the defect scores of the features are summarized by summation. This summarization method can enhance the interpretability of ScoringDP.

[0135] In Figure 7It should be noted that fn is the feature name, fv is the feature value, and fns is multiple feature names. The details are not repeated here.

[0136] Furthermore, for code modules, higher defect scores indicate higher defect density. Therefore, quality assurance resources should be allocated to code modules with higher defect scores to improve resource utilization. To interpret the prediction results for a specific code module, ScoringDP uses the model feature defect scores calculated during the prediction process as feature importance. Figure 10 , Figure 10 This is an example diagram for explaining the feature importance of ScoringDP disclosed in the embodiment of the present application. Figure 10 An example of ScoringDP explanation is shown, where the horizontal axis represents the defect score of the feature (i.e. the calculated defect score) and the vertical axis represents the code module features (and their values, i.e. the true value) sorted in descending order according to the defect score. For each feature described above, see Figure 11 . Figure 11 A Chinese-English comparison chart of the feature name of a module feature.

[0137] Specifically, the defect score of a single code module feature can be interpreted from two aspects: sign and magnitude. A positive (or negative) defect score indicates that the actual defect rate is higher than the expected defect rate, which means that the code module has more (or fewer) defects than expected when the value of code module feature X is x. A larger absolute value of the defect score indicates the extent to which the total number of defects in the code module is more (or less) than the expected number of defects. Based on the ranked features, developers can make evidence-based decisions: in order to reduce the defect proneness of a code module, developers should modify the code module to prioritize changing features with larger defect scores, i.e. Figure 10 The specific module feature names can be found in Figure 11 .

[0138] The explainable defect prediction method disclosed in this embodiment has the following three advantages: (1) The explanation is relatively intuitive: The explanation of ScoringDP comes directly from its internal decision-making process, without relying on external explanation methods (such as LIME and SHAP), which improves the accuracy and stability of the explanation. (2) Strong operability: Developers can take more effective actions to reduce defects based on the defect scores of module features, such as modifying the code to prioritize changing features with higher defect scores. (3) Excellent prediction performance: According to experiments, ScoringDP is better than or equal to the existing state-of-the-art methods in most performance indicators.

[0139] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0140] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of an interpretable defect prediction system disclosed in an embodiment of the present application.

[0141] An acquisition unit 1201 is configured to acquire each code module to be analyzed in a software project to be analyzed, and determine a defect data set corresponding to the code module to be analyzed and code module features; wherein, the defect data set is constructed from historical data of the code module to be analyzed in the software project to be analyzed, and the code module features are used to characterize the numerical features of the code module to be analyzed;

[0142] A determination unit 1202 is configured to, when the module feature value of any code module feature in any code module to be analyzed is a target feature value, determine the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed according to the defect data set; wherein, the expected defect rate is the ratio of the total number of lines of code modules of defective code modules with the module feature value being the target feature value in the defect data set to the total number of lines of all code modules of defective code modules in the defect data set, and the actual defect rate is the ratio of the total number of defective code modules of defective code modules with the module feature value being the target feature value in the defect data set to the total number of defective code modules of all defective code modules in the defect data set; the defect data set includes the total number of lines of code modules and the total number of defective code modules corresponding to different module feature values, as well as the total number of lines of all code modules and the total number of defective code modules corresponding to all defective code modules; the defective code module is the code module in the defect data set;

[0143] A calculation unit 1203 is configured to calculate the difference between the actual defect rate and the expected defect rate to obtain a feature defect score corresponding to any code module feature in any code module.

[0144] The calculation unit 1203 is further configured to calculate the feature defect score of all code module features in any code module to be analyzed, so as to obtain the module defect score of any code module to be analyzed; wherein, the module defect score is used to characterize the defect tendency of the code module to be analyzed.

[0145] The sorting unit 1204 is configured to, after obtaining the module defect scores of all code modules to be analyzed, sort all the module defect scores, so as to process all the code modules to be analyzed in sequence according to the order from large to small of all the module defect scores.

[0146] Please refer to Figure 13 , the structural schematic diagram of an interpretable defect prediction device disclosed in an embodiment of the present application includes:

[0147] A central processing unit 1301, a memory 1305, an input / output interface 1304, a wired or wireless network interface 1303, and a power supply 1302;

[0148] The memory 1305 is a transient storage memory or a persistent storage memory;

[0149] The central processing unit 1301 is configured to communicate with the memory 1305 and execute the instruction operations in the memory 1305 to execute the interpretable defect prediction method in any of the foregoing Figures 2 to 6 shown embodiments.

[0150] An embodiment of the present application further provides a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instruction to execute the interpretable defect prediction method in any of the foregoing Figures 2 to 6 shown embodiments.

[0151] An embodiment of the present application further provides a computer-readable storage medium, which includes instructions. When the instructions run on a computer, the computer is caused to execute the interpretable defect prediction method in any of the foregoing Figures 2 to 6 shown embodiments.

[0152] An embodiment of the present application further provides a computer program product containing instructions. When the computer program product runs on a computer, the computer is caused to execute the interpretable defect prediction method in any of the foregoing Figures 2 to 6 shown embodiments.

[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0154] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0155] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0157] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs and other various media that can store program codes.

Claims

1. An explainable defect prediction method, characterized in that: The method comprises: Constructing a defect data set based on historical data of the software project to be analyzed, and obtaining each code module to be analyzed in the software project to be analyzed and a code module feature corresponding to each code module to be analyzed from current data of the software project to be analyzed; wherein the code module feature is used to characterize a numerical feature of the code module to be analyzed; If the module characteristic value of any code module characteristic in any code module to be analyzed is the target characteristic value, the defect data set is analyzed according to the defect data set, and defect code modules whose module characteristic value of the code module characteristic is the interval window characteristic value are screened out; wherein the interval window characteristic value is used to characterize that in the target interval window centered on the defect data set, the ratio of the number of defect code modules located in the target interval window to the total number of all defect code modules in the defect data set meets a preset quantity ratio; Determine the actual defect rate and expected defect rate of any code module feature in any code module to be analyzed; wherein the expected defect rate is the ratio of the total number of code module lines of the defective code modules whose module feature values ​​in all the defective code data sets are the target feature values ​​to the total number of code module lines of all the defective code modules in the defective code data sets, and the actual defect rate is the ratio of the total number of code module defects of the defective code modules whose module feature values ​​in all the defective code data sets are the target feature values ​​to the total number of code module defects of all the defective code modules in the defective code data sets; the defective code data sets include the total number of code module lines and the total number of code module defects corresponding to different module feature values, and the total number of code module lines and the total number of code module defects corresponding to all the defective code modules; the defective code module is the code module in the defective data sets; Calculating a difference between the actual defect rate and the expected defect rate to obtain a feature defect score of any code module feature in any code module; Calculating feature defect scores of all code module features in any code module to be analyzed to obtain a module defect score of any code module to be analyzed; wherein the module defect score is used to characterize the defect tendency of the code module to be analyzed; After obtaining the module defect scores of all the code modules to be analyzed, sorting all the module defect scores, so as to process all the code modules to be analyzed in the software project to be analyzed in descending order of the module defect scores; Before filtering out defective code modules whose module characteristic values ​​of code module characteristics are interval window characteristic values, the method further includes: Obtaining module feature values ​​of all code module features in any defective code module in the defect data set, sorting all module feature values ​​in ascending order, and obtaining a feature value sorting list corresponding to all module feature values; In the eigenvalue sorting list, checking whether the interval window eigenvalue exists in the eigenvalues ​​of all modules; If the interval window eigenvalue exists in the eigenvalue sorting list, the interval window eigenvalue is used as an insertion point of the eigenvalue sorting list to determine the index of the initial interval window.

2. The explainable defect prediction method according to claim 1, characterized in that: The step of obtaining each code module to be analyzed in the software project to be analyzed and a code module feature corresponding to each code module to be analyzed from the current data of the software project to be analyzed includes: The software project to be analyzed is analyzed line by line, each line of the software project to be analyzed is set as any code module to be analyzed, and the code module features contained in any code module to be analyzed are determined.

3. The explainable defect prediction method according to claim 1, characterized in that: Determining the actual defect rate and the expected defect rate of any code module feature according to the defect data set includes: Based on the defect data set, obtaining the total number of lines of the code module, the total number of lines of the total code module, the total number of defects of the code module and the total number of defects of the total code module; The total number of lines of the code module and the total number of lines of the total code module are input into the expected defect rate calculation formula to obtain the expected defect rate; wherein the expected defect rate calculation formula is: , the expected_defect_ratio is the expected defect rate, the cm is the defect code module in the defect data set, the D is the defect data set, the f is the feature of any code module, the fv is the target feature value, the LOC is the number of lines of code of the defect code module in the defect data set, The total number of code module rows accumulated when the module characteristic value of the defective code module in the defective data set is the target characteristic value, Used to represent the total number of lines of the total code module; The total number of code module defects and the total number of total code module defects are input into the actual defect rate calculation formula to obtain the actual defect rate; wherein the actual defect rate calculation formula is: , the actual_defect_ratio is the actual defect rate, the defect is the number of defects in the defect code module in the defect data set, and the The total number of code module defects accumulated when the module characteristic value of the defective code module in the defect data set is the target characteristic value, Used to represent the total number of defects in the total code module.

4. The explainable defect prediction method according to claim 1, characterized in that: Determining the actual defect rate and the expected defect rate of any code module feature in any code module to be analyzed according to the defect data set includes: Determine, according to the number of defective code modules and the total number of all defective code modules, the total number of lines of the code modules, the total number of defects of the code modules, and the total number of defects of the total code modules; The total number of lines of the code module and the total number of lines of the total code module are input into the calculation formula for estimating the expected defect rate to obtain the expected defect rate; wherein the calculation formula for estimating the expected defect rate is: , the expected_defect_ratio is the expected defect rate, the cm is the defect code module in the defect data set, the D is the defect data set, the f is the code module feature, the intvl is the interval window feature value, the LOC is the number of code lines of the defect code module in the defect data set, The total number of code module rows accumulated when the module characteristic value of the defective code module in the defective data set is the interval window characteristic value, Used to represent the total number of lines of the total code module; The total number of code module defects and the total number of total code module defects are input into a calculation formula for estimating the actual defect rate to obtain the actual defect rate; wherein the calculation formula for estimating the actual defect rate is: , the actual_defect_ratio is the actual defect rate, the defect is the number of defects in the defect code module in the defect data set, and the The total number of code module defects accumulated when the module characteristic value of the defective code module in the defect data set is the interval window characteristic value, Used to represent the total number of defects in the total code module.

5. The explainable defect prediction method according to claim 4, characterized in that: The left index and the right index of the initial interval window are the interval positions of the module feature values ​​in the feature value sorting list, and the insertion point is the center of the initial interval window. The module feature value of the code module feature is filtered out before the defective code module whose feature value is the interval window feature value. The method further includes: If the insertion point is located after the interval position of the interval window feature value and the last module feature value in the feature value sorting list, the left index and the right index in the index of the initial interval window are set to the last module feature value in the feature value sorting list; or, If the insertion point is located at the interval position between the interval window feature value and the first module feature value in the feature value sorting list, the left index and the right index in the index of the initial interval window are set to the first module feature value in the feature value sorting list; or, If the insertion point is located before the interval position of the first module feature value in the feature value sorting list where the interval window feature value is located, the left index in the index of the initial interval window is set to the index of the previous sorting position of the insertion point, and the right index is the index of the insertion point.

6. The explainable defect prediction method according to claim 5, characterized in that: The method for obtaining the target interval window includes: Setting a preset interval window value; wherein the preset interval window value is used to represent the ratio of the total number of lines of a preset code module to the total number of lines of a preset total code module; Calculating the ratio of the total number of lines of the code module of any defective code module located in the initial interval window to the total number of lines of the total code modules to obtain a ratio of the number of lines to be analyzed; When the ratio of the number of rows to be analyzed is less than half of the preset interval window value, and any index of the initial interval window is not located at the list boundary of the feature value sorting list, the initial interval window is iterated to move the left index of the initial interval window to the left by one unit of interval position, or to move the right index of the initial interval window to the right by one unit of interval position, so as to satisfy that the interval window feature value is located at the center of the iterated initial interval window; Calculate the ratio of the number of rows to be analyzed of the initial interval window after iteration, until the ratio of the number of rows to be analyzed is not less than half of the preset interval window value, and any index of the initial interval window is located at the list boundary of the feature value sorted list, terminate the iteration, and determine the initial interval window after iteration as the target interval window.

7. The explainable defect prediction method according to claim 6, characterized in that: Before the iteration is terminated and the initial interval window after the iteration is determined as the target interval window, the method further includes: If the right index of the iterated initial interval window reaches the right list boundary of the eigenvalue sorting list, the left index of the iterated initial interval window is moved rightward by one unit interval position; or, If the left index of the initial interval window after the iteration reaches the left list boundary of the eigenvalue sorting list, the right index of the initial interval window after the iteration is moved to the left by a unit interval position.

8. The explainable defect prediction method according to claim 6, characterized in that: The method further comprises: When the ratio of the number of rows to be analyzed is not less than the preset interval window value, taking the interval window characteristic value as the window center point, calculating the interval distance between the left index or the right index of the initial interval window and the window center point; When the left interval distance of the left index of the initial interval window is less than the right interval distance of the right index of the initial interval window, and the left index of the initial interval window is not located at the list boundary of the feature value sorted list, the left index of the initial interval window is moved to the left by a unit interval position, so as to determine the moved initial interval window as the target interval window; or, When the right interval distance of the right index of the initial interval window is less than the left interval distance of the left index of the initial interval window, and the right index of the initial interval window is not located at the list boundary of the feature value sorted list, the right index of the initial interval window is moved to the right by one unit of interval position to determine the moved initial interval window as the target interval window.

9. An explainable defect prediction device, characterized in that: The device comprises: CPU, memory, input and output interfaces, wired or wireless network interfaces, and power supply; The memory is a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the explainable defect prediction method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises instructions, and when the instructions are executed on a computer, the computer is caused to perform the explainable defect prediction method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Software defect prediction method and device, electronic equipment and computer storage medium

    CN113656325A

  • Code detection method and device, equipment and storage medium

    CN115344491A