A workload-aware just-in-time software defect prediction method based on linear programming
Through the workload-aware real-time software defect prediction method (LPDP) based on linear programming, the problem of time-consuming and labor-consuming parameter tuning in the existing technology is solved, and efficient real-time software defect prediction is achieved, which improves the defect recognition effect and reduces development costs.
Patent Information
- Application Number
- CN202210788768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-06
AI Technical Summary
The existing classifier-based real-time software defect prediction method has the problem of time-consuming and labor-consuming parameter tuning, resulting in unsatisfactory defect recognition effect.
The workload-aware real-time software defect prediction method (LPDP) based on linear programming is adopted. By extracting software metric elements and labeling software entities, class imbalance data processing and data normalization are performed, and the LPDP model is constructed to predict whether newly submitted software code changes introduce defects.
It realizes a simple and effective real-time software defect prediction without parameters, avoids the parameter tuning process, improves the defect identification effect, reduces development costs, and ensures the quality and reliability of software products.
Smart Images

Figure CN115033493B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect prediction, and in particular to a workload-aware real-time software defect prediction method based on linear programming. Background Art
[0002] With the advent of the information age, the computer software industry has developed vigorously, people's demand for software is increasing, and software technology is affecting everyone around us. However, in the process of software development, some unavoidable problems will arise, such as system compatibility issues of new software technologies, communication problems between project members, etc., which will inevitably introduce software defects. Discovering these defects as early as possible and repairing them is extremely important for the later development and maintenance of the software. The process of detecting software defects is called software defect prediction.
[0003] Traditional software defect prediction technology mainly focuses on coarse-grained measurement of software code or software development process, such as defect prediction at method level, class level, file level or package level. However, after predicting that a module has defect tendency, this coarse-grained measurement needs to locate the location of the defect in the method, class, file or package and repair it. Developers need to recall the details of the development. At this time, the process of locating and repairing defects takes a lot of time and effort. The cost of traditional software defect prediction in practical applications is relatively high. Therefore, researchers proposed real-time software defect prediction, which is to predict whether the code changes submitted by developers within a period of time will introduce defects during the software development process. Developers are still very clear about the code they have just submitted, and it is relatively easy to locate the specific location of the defect and repair the defect in time. The measurement of real-time software defect prediction is at the code change level. Compared with traditional software defect prediction, it is fine-grained, real-time and easy to trace. Therefore, it can predict more defects with relatively less development cost.
[0004] Most of the existing classifier-based real-time software defect prediction methods are parameterized models, which usually use default parameter settings, and their defect prediction effects are often not ideal, such as Naive Bayes (NB), decision tree (DT), random forest (RF), etc. If these classifier algorithms are tuned, the optimization process is usually time-consuming and labor-intensive, which will hinder the application of the above methods in practice and result in poor defect identification effects. Summary of the invention
[0005] In view of the above problems existing in the existing methods, the present invention conducts in-depth mining of defect data from the perspective of linear programming to identify whether the code submission entity will introduce defects, and proposes a workload-aware real-time software defect prediction method (LPDP) based on linear programming, which is a parameter-free, simple and effective method.
[0006] In order to achieve the above object, the present invention provides the following technical solution: a workload-aware real-time software defect prediction method based on linear programming, comprising the following steps:
[0007] Extract software metrics and annotate software entities to construct an instant software defect dataset;
[0008] Perform class imbalance data processing on real-time software defect datasets;
[0009] The metric elements of the software entities in the processed real-time software defect dataset are normalized to eliminate the impact of the dimensions of different metric elements of the software entities;
[0010] Build the LPDP model on the real-time software defect prediction dataset obtained by data normalization;
[0011] The LPDP model is used to predict the newly submitted software code changes and obtain their predicted labels.
[0012] Preferably, in the process of class imbalance data processing, a data sampling algorithm is used to perform data class imbalance processing, and the distribution of original data is changed by randomly increasing the number of defective software entities or reducing the number of non-defective software entities to achieve a quantitative balance between the two types of data.
[0013] Preferably, the data sampling algorithm adopts a random upsampling algorithm.
[0014] Preferably, the data normalization processing method is:
[0015] The normalization processing method is the z-score method, and the calculation method is as follows:
[0016]
[0017] where x i is the raw value of the ith metric of software entity x, is x i The normalized value, μ x is the average value of software entity x, σ x is the standard deviation of software entity x.
[0018] Preferably, the LPDP model is constructed by defining a probability labeling matrix, constructing a discriminant distance matrix and constructing an objective function.
[0019] Preferably, the definition of the probability labeling matrix definition includes:
[0020] Use C = {c 1 , c 2} represents the class of the training set, that is, the defect class and the non-defect class; μ={μ 1 , μ 2} represents the class center of the training set; Ω S ={x 1 , x 2 , ..., x m} represents the software entity in the training set; p = {0, 1} is used to represent the actual label of the software entity in the training set, and Ω T ={x 1 , x 2 , ..., x n} represents the software entity in the test set, and represents the predicted label of the software entity in the test set, m and n represent the number of software entities in the training set and the test set respectively, and d represents the number of metrics of the software entity;
[0021] Let M represent the probability labeling matrix, then M ij A software entity x representing a test set n The probability of belonging to class C;
[0022] The class center μ i The calculation formula is as follows:
[0023]
[0024] Where p(x j ) represents the software entity x in the training set j The actual label.
[0025] Preferably, D is used to represent the discriminant distance matrix, D ij Represents the software entity x in the test set j To the cluster center μ i The Euclidean distance, D ij The calculation method is as follows:
[0026] D ij =||μ i -x j || 2 (3)
[0027] Preferably, the objective function construction method of the LPDP is as follows:
[0028] The probability labeling matrix M is learned, and the product of the discriminant distance matrix D and the probability standard matrix M is minimized as the objective function of the LPDP model. The formula is:
[0029]
[0030] Among them, the constraints of the objective function are:
[0031] M ij Represents the software entity x j Belongs to class c i The probability value is expressed as:
[0032] 0≤M ij ≤1 (5)
[0033] Software entity x in the test set j Belongs to each category c i The probability sum is 1, expressed as:
[0034]
[0035] Because the software entities in the training set and the software entities in the test set have the same label space, each class c i contains at least one software entity x in the test set j , the constraint is expressed as:
[0036]
[0037] Then the objective function of the LPDP model is:
[0038]
[0039]
[0040] Preferably, the method for obtaining the predicted probability or label through the LPDP model is as follows:
[0041] Probabilistically label the software entities t in the test set in M j Belongs to a class c i When the probability of is the largest, the software entity is determined to belong to this category, which is expressed as:
[0042]
[0043] in, Represents the software entity x in the test set j The predicted label of .
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. From the perspective of linear programming, the present invention designs a parameter-free workload-aware real-time software defect prediction method, which avoids the setting of relevant parameters and the time-consuming and labor-intensive parameter tuning process. At the same time, the present invention performs class-imbalanced data processing and normalization on the real-time software defect data set, eliminates the impact of dimensions between different measurement elements of software entities, and expands the number of defective software entities in the data sampling algorithm, thereby realizing the use of fewer test resources to discover as many defects as possible, which is beneficial to improving the defect identification effect and helps to better ensure the quality and reliability of software products. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flow chart of the workload-aware real-time software defect prediction method based on linear programming of the present invention;
[0047] Figure 2 This is a comparison chart of the experimental effects of the workload-aware real-time software defect prediction method based on linear programming of the present invention and the existing method on the workload-aware indicator EA_Fmeasure. DETAILED DESCRIPTION
[0048] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0049] The present invention provides a workload-aware real-time software defect prediction method based on linear programming, such as Figure 1 As shown, a workload-aware real-time software defect prediction method based on linear programming includes the following steps:
[0050] S1: Constructing an instant software defect dataset, which mainly includes extracting software metrics and labeling software entities; by analyzing the software development process, extracting metrics related to software defects, extracting software entities from the software history warehouse, and mining relevant defect information to label whether each software entity introduces defects, constructing an instant software defect dataset.
[0051] Specifically, the software metrics in S1 refer to the intrinsic properties related to the external manifestations of defects, also known as features or attributes; software entities refer to the code changes submitted by developers during the software development process.
[0052] S2: Class imbalance data processing: Data sampling algorithm is used to process data class imbalance. The distribution of original data can be changed by randomly increasing the number of software entities in the minority class (i.e. defect class) or reducing the number of software entities in the majority class (i.e. defect-free class), so as to balance the two types of data as much as possible.
[0053] Specifically, the data sampling algorithms in S2 include random upsampling, random downsampling, etc., wherein random upsampling mainly involves random duplication from software entities in the minority class to increase the number of software entities in the minority class, and random downsampling involves randomly selecting some from software entities in the majority class and removing them to balance the distribution of the two types of data. Considering that random downsampling removes some software entities, resulting in the loss of information about these entities, the effect of the learned model may not be optimal. Therefore, the present invention uses a random upsampling algorithm to deal with the problem of imbalanced data classes.
[0054] S3: Data preprocessing: normalize the metrics of the software entities in the defect data set to eliminate the impact of the dimensions of different metrics of the software entities.
[0055] Specifically, the normalization method in S3 is:
[0056] The normalization processing method is the z-score method, and the calculation method is as follows:
[0057]
[0058] where x i is the raw value of the ith metric of software entity x, is x i The normalized value, μ x is the average value of software entity x, σ x is the standard deviation of software entity x.
[0059] S4: Constructing LPDP model on real-time software defect prediction data; mainly including probability annotation matrix definition; identification (discrimination) distance matrix construction; objective function construction, etc.
[0060] Specifically, the probability labeling matrix in S4 is defined as follows:
[0061] Use C = {c 1 ,c 2} represents the class of the training set, that is, the defect class and the non-defect class; μ={μ 1 ,μ 2} represents the class center of the training set; Ω S ={x 1 ,x 2 ,...,x m} represents the software entity in the training set; p = {0,1} is used to represent the actual label of the software entity in the training set, and Ω T ={x 1 ,x 2 ,...,x n} represents the software entity in the test set, and represents the predicted label of the software entity in the test set, m and n represent the number of software entities in the training set and the test set respectively, and d represents the number of metrics of the software entity.
[0062] Let M represent the probability labeling matrix, then M ij A software entity x representing a test set n The probability of belonging to class c.
[0063] In order to obtain the predicted probability of the software entity, it is necessary to learn the probability labeling matrix M, and use the minimum loss function as the objective function of the LPDP model, which is expressed as:
[0064]
[0065] D represents the discriminant distance matrix, D ij Represents the software entity x in the test set j To the cluster center μ i The Euclidean distance is calculated as:
[0066] D ij =||μ i -x j || 2 (3)
[0067] where μ of the class center i The calculation formula is as follows:
[0068]
[0069] Where p(x j ) represents the software entity x in the training set j The actual label.
[0070] There are three constraints on the objective function:
[0071] First, Mij represents the probability value, so M ij The value of is between 0 and 1 and is represented by:
[0072] 0≤M ij ≤1 (5)
[0073] Secondly, the software entity x in the test set j To each class c i The probability of the sum is equal to 1, expressed as:
[0074]
[0075] Finally, because the software entities in the training set and the test set have the same label space, each class center contains at least one software entity in the test set. The experiment is to predict whether the software entity introduces defects, that is, Mij The result is a binary 0 or 1. A value of 0 indicates that the software entity does not introduce defects, that is, the code change does not introduce defects. A value of 0 indicates that the code change introduces defects. The constraint is expressed as:
[0076]
[0077] In summary, the LPDP model is formed as follows:
[0078]
[0079]
[0080] S5: Solve the linear programming problem to predict the newly submitted software code changes and obtain their predicted labels.
[0081] The probability that the software entity in the test set belongs to a certain class in the probability labeling matrix M is the largest, then the software entity is judged to belong to this class, which can be expressed as:
[0082]
[0083] in, Represents the software entity x in the test set j The predicted label of .
[0084] In order to fully demonstrate the feasibility of the workload-aware real-time software defect prediction method based on linear programming proposed in the present invention, a verification experiment was carried out.
[0085] Experimental environment configuration: windows11 system, matlab2020b
[0086] Experimental data: This experiment uses six public datasets Bugzilla (BUG), Columba (COL), EclipseJDT (JDT), EclipsePlatform (PLA), Mozilla (MOZ) and PostgreSQL (POS) as experimental datasets. The 14 measurement information of software entities in the datasets are shown in Table 1 below:
[0087]
[0088]
[0089] Commonly used evaluation indicators such as precision, recall, and the comprehensive indicator F-measure require reviewing all software entities in the test set, which will consume more test resources. In real life, test resources are limited, and fewer test resources are needed to discover as many software defects as possible. The present invention uses workload-aware evaluation indicators to evaluate the performance of the proposed method and existing methods. The specific implementation process is briefly introduced below.
[0090] Input: training set
[0091] Output: Whether each software entity in the test set introduces defects, 1 indicates that the software entity introduces defects, and 0 indicates that the software entity does not introduce defects.
[0092] In the first step, we use the time series cross-validation experimental setting to divide the data set. Each data set is divided into subsets according to the submission time. Each subset is used as the training set in turn. The interval is set to 4, that is, the test set is the 4th subset after the training set. The model is built and traversed in turn until the last subset is used as the test set.
[0093] The second step is to process the data class imbalance and normalize the data.
[0094] The third step is to calculate the class center of the training set. There are two classes in the experiment, defect class and non-defect class. The Euclidean distance from the software entity in the test set to each class center is calculated.
[0095] Step 4: Build the LPDP model on the software defect dataset to obtain the probability labeling matrix M, and then get the predicted probability or label.
[0096] Step 5: Arrange the software entities in the test set in ascending order according to the workload cost required to review them.
[0097] Step 6. Check these software entities from top to bottom according to the sorting results. According to the 80 / 20 rule in software testing, 80% of the defects exist in 20% of the software entities. Take out all the software entities that we can detect with 20% of the workload from the list, and count the number of defects introduced in these software entities.
[0098] Introduce the EA_Fmeasure indicator, and the calculation formula is as follows:
[0099]
[0100] Where P 20% represents the precision rate (Precision) with 20% workload, R 20% Represents the recall rate under 20% workload.
[0101] If the number of software entities in the test set is n, the number of software entities that actually introduce defects in the test set is k, and the number of all software entities that can be detected with 20% of the workload is n 20% , the number of defects actually introduced into the software entity that can be detected by 20% of the workload is k 20% , P 20% and R 20% The calculation formula is as follows:
[0102]
[0103]
[0104] Under the same experimental environment, the method is compared with EALR, naive Bayes (NB), logistic regression (LR), random forest (RF), decision tree (J48), sequential minimal optimization algorithm (SMO), unsupervised learning methods (NS, ND, NF, Entropy, LT, FIX, NDEV, AGE, NUC, EXP, REXP, SEXP), Oneway and CCUM, and the experimental results are compared on the EA_Fmeasure indicator.
[0105] The present invention uses a linear programming method to predict whether a software entity in a real-time software defect dataset introduces defects. The method of the present invention is compared with EALR, NB, LR, RF, J48, SMO, unsupervised learning methods (NS, ND, NF, Entropy, LT, FIX, NDEV, AGE, NUC, EXP, REXP, SEXP), OneWay and CCUM methods on six public datasets using a time series cross-validation experimental setting. The experimental results are as follows: Figure 2 As shown, the feasibility of this method is verified. Figure 2 The lower limit of the box plot is 25%, the upper limit is 75%, and the horizontal line is the median of the LPDP method. A value higher than the median of the LPDP method indicates that the method is better than LPDP, and a value parallel to the median of the LPDP method indicates that there is no significant difference between the method and LPDP. A value higher than the median of the LPDP method indicates that the method is worse than LPDP. Experiments show that the LPDP method is better than most methods that require parameter tuning in terms of the EA_Fmeasure indicator.
[0106] The embodiments described above are only preferred specific implementation modes of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of the technical solutions that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belong to the protection scope of the present invention.
Claims
1. A workload-aware real-time software defect prediction method based on linear programming, characterized in that: The steps include: Extract software metrics and annotate software entities to construct an instant software defect dataset; Perform class imbalance data processing on real-time software defect datasets; The metric elements of the software entities in the processed real-time software defect dataset are normalized to eliminate the impact of the dimensions of different metric elements of the software entities; The LPDP model is constructed on the real-time software defect prediction dataset obtained by data normalization; specifically: The LPDP model is constructed by defining a probability labeling matrix, constructing a discriminant distance matrix and constructing an objective function; The definition of the probability labeling matrix includes: C = {c1, c2} represents the class of the training set, i.e., defect class and non-defect class; μ = {μ1, μ2} represents the class center of the training set; Ω S ={x1, x2, ..., x m } represents the software entity in the training set; p = {0, 1} is used to represent the actual label of the software entity in the training set, and Ω T ={x1, x2, ..., x n } represents the software entity in the test set, and represents the predicted label of the software entity in the test set, m and n represent the number of software entities in the training set and the test set respectively, and d represents the number of metrics of the software entity; Let M represent the probability labeling matrix, then M ij A software entity x representing a test set j Belongs to class c i probability; The class center μ i The calculation formula is as follows: Where p(x j ) represents the software entity x in the training set j The actual label of Let D represent the discriminant distance matrix, D ij Represents the software entity x in the test set j To the cluster center μ i The Euclidean distance, D ij The calculation method is as follows: D ij =||μ i -x j ||2 (3) The objective function construction method of the LPDP is as follows: The probability labeling matrix M is learned, and the product of the discriminant distance matrix D and the probability standard matrix M is minimized as the objective function of the LPDP model. The formula is: Among them, the constraints of the objective function are: M ij Represents the software entity x j Belongs to class c i The probability value is expressed as: 0≤M ij ≤1 (5) Software entity x in the test set j Belongs to each category c i The probability sum is 1, expressed as: Because the software entities in the training set and the software entities in the test set have the same label space, each class c i contains at least one software entity x in the test set j , the constraint is expressed as: Then the objective function of the LPDP model is: The LPDP model is used to predict the newly submitted software code changes and obtain their predicted labels.
2. The workload-aware real-time software defect prediction method based on linear programming according to claim 1, characterized in that: In the process of class imbalance data processing, a data sampling algorithm is used to process data class imbalance, and the distribution of original data is changed by randomly increasing the number of defective software entities or reducing the number of non-defective software entities, so as to achieve a quantitative balance between the two types of data.
3. The workload-aware real-time software defect prediction method based on linear programming as claimed in claim 2, characterized in that: The data sampling algorithm adopts a random upsampling algorithm.
4. The workload-aware real-time software defect prediction method based on linear programming according to claim 1, characterized in that: The data normalization processing method is the z-score method, and the calculation method is as follows: where x i is the raw value of the ith metric of software entity x, is x i The normalized value, μ x is the average value of software entity x, σ x is the standard deviation of software entity x.
5. The workload-aware real-time software defect prediction method based on linear programming according to claim 1, characterized in that: The method of obtaining the predicted probability or label through the LPDP model is as follows: Probabilistically label the software entities t in the test set in M j Belongs to a class c i When the probability of is the largest, the software entity is determined to belong to this category, which is expressed as: in, Represents the software entity x in the test set j The predicted label of .
Citation Information
Patent Citations
Software defect prediction method, device, storage medium and electronic equipment
CN108647138A
Semi-supervised change-level software defect prediction method based on three-branch decisions
CN109543707A