Cross-project defect prediction sample filtering method and prediction method based on isolation forest
Through the sample filtering method based on isolated forest, the problem of noise and data imbalance in software isomorphic cross-project defect prediction is solved, the model performance and efficiency are improved, and efficient prediction without relying on the target dataset is achieved.
Patent Information
- Application Number
- CN202210393192.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-14
AI Technical Summary
There is noisy data and data imbalance in software isomorphic cross-project defect prediction, resulting in low performance of prediction models, low efficiency of existing methods and relying on target datasets.
A cross-project defect prediction sample filtering method based on isolated forest is adopted, and data quality is improved by balancing data, building isolated forests, weighting and sample filtering is carried out.
The performance of the software defect prediction model and the efficiency of building the model are improved, the noise and data imbalance problems are solved, and efficient sample filtering that does not depend on the target dataset is realized.
Smart Images

Figure CN114756461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect prediction, and more specifically to a cross-project defect prediction sample filtering method and a prediction method based on isolation forest. Background Art
[0002] In recent years, software defect prediction has become an active field in software engineering. When the software under test has no historical version or the amount of software historical data is too small, the historical knowledge of other software ("source data") with the same name and number of metric elements as the software under test is used to predict the defects of the software under test ("target data"), which is called "isomorphic cross-project defect prediction". Knowing the quality status of the software in advance can provide certain guidance for software engineering personnel, so that they can reasonably allocate resources, save costs, and improve software testing efficiency.
[0003] The data samples for cross-project defect prediction mainly come from other projects. Therefore, for the software under test, there is a part of noise data, and the cross-project data samples have serious data imbalance, that is, the defect data is far less than the non-defect data. The two characteristics of noise and data imbalance in the data samples will reduce the performance of the constructed software cross-project defect prediction model, resulting in inaccurate prediction of defects in the software under test. Therefore, how to filter and screen the data is the key to improving the performance of the cross-project defect prediction model.
[0004] At present, many scholars have proposed different methods for processing homogeneous cross-project data, which can be mainly divided into three categories: methods based on sample selection, dimensionality conversion, and improved algorithms.
[0005] There are seven methods based on sample selection: Burak Filter (BF) based on k-nearest neighbors, Transfer Naive Bayes Bayes, TNB), Hybrid Instance Selection Using Nearest-Neighbor (HISNN), data filter by agglomerative clustering (DFAC), the hierarchical selection-based filter (HSBF), effort-aware supervised cross-project defect prediction (EASC), and collaborative filtering based source projects selection (CFPS). There are two methods based on dimensionality conversion: transfer component analysis extension (TCA+) and joint feature representation with double marginalized denoising autoencoders (DMDA-JFR). There are three main methods based on improved algorithms: the naive leader method ( Bellwether, referred to as BNaive), the bellwether method with transfer principal component analysis extension (Bellwether TCA+, referred to as BTCA+), and the weighted naive Bayes bellwether method (Bellwether TNB, referred to as BTNB).
[0006] However, most of the above methods consider the similarity between samples of source project data and target data through sample selection (noise removal) or reduce the distribution difference between source project data and target data through feature space transformation. But these methods are strongly dependent on the target data set. In addition, when the amount of source project data is large, the time cost of sample selection or feature transformation is very high. For example, the computational complexity of BF and HISNN is exponential, which leads to low efficiency in building defect prediction models. In addition, except for DMDA-JFR and the leader method, most methods do not use defect label data, which is very important when the boundary of defect prediction is fuzzy and can improve the performance of the model. Model performance, cost and efficiency are what software engineers care about most.
[0007] Therefore, how to provide a simple, easy-to-use, target-project-independent and efficient cross-project defect prediction sample filtering method and prediction method is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0008] In view of this, the present invention provides a cross-project defect prediction sample filtering method and a prediction method based on isolation forests, which aims to solve the problems of data noise, over-dependence on target data sets and low model building efficiency in software isomorphic cross-project defect prediction, and can improve the performance of software defect prediction models and the efficiency of model building while deleting noise samples and utilizing defect label data.
[0009] In order to achieve the above object, the present invention adopts the following technical solution:
[0010] A cross-project defect prediction sample filtering method based on isolation forest, comprising:
[0011] S1. Extract the dataset of isomorphic cross-project software as the source project dataset;
[0012] S2. Balancing the source project data set to obtain balanced data;
[0013] S3. Divide the balanced data into positive sample data and negative sample data, wherein the positive sample data is defective data and the negative sample data is non-defective data;
[0014] S4. constructing an isolation forest for the positive sample data and the negative sample data respectively;
[0015] S5. Perform weighted processing on the isolation forest and perform sample filtering: calculate the weighted path length of each sample data on the isolation tree, calculate the average weighted path length of each sample data in the weighted isolation forest based on the weighted path length, calculate the outlier value of each sample data based on the average weighted path length, remove the abnormal samples of the weighted isolation forest according to the preset abnormal ratio and synthesize the remaining positive sample data and negative sample data to obtain the filtered source data set.
[0016] Preferably, before balancing the source project data in S2, the step further includes preprocessing the source project data set, and the specific content of the data preprocessing includes:
[0017] S21. Binarize the defect label information of the numerical type of the source project data set: when the defect label is greater than or equal to 1, it is marked as 1, indicating a defect; when the defect label is 0, it remains unchanged, indicating no defect;
[0018] S22. Select the source project data set as the source data sample;
[0019] S23. Eliminate duplicate samples: when there are exactly the same samples in the source data samples, only one sample is retained;
[0020] S24. Data standardization: Use Z-Score standardization technology to convert the deduplicated source data samples and target metrics into data with a mean of 0 and a variance of 1.
[0021] Preferably, the dividing of the balanced data into positive sample data and negative sample data in S3 is performed according to the defect label information.
[0022] Preferably, the specific content of S4 includes:
[0023] S41. Randomly select m samples from the positive sample data and the negative sample data as training data sets Each sample consists of feature F, and the isolation forest consists of t isolation trees, iForest = {iTree1, iTree2, ..., iTree t}, select a subsample X' from X by random sampling without replacement, the size of X' is φ, and the limit height of the isolation tree is h lim ,h lim =ceiling(log2φ);
[0024] S42. Initialize iTree: construct a root node, which contains all subsamples X';
[0025] S43. Randomly select a feature f, f∈F, from all features, and randomly select a partition point p, the partition point is between the maximum values of the selected feature, f min ≤p≤f max ;
[0026] S44. Input sample Compare the value of the feature f of the current root node with the selected partition point, and divide the root node into two child nodes, that is, if f < p, the sample is placed on the left child node, and if f ≥ p, the sample is placed on the right child node;
[0027] S45. Repeat S42-S44 until all samples are isolated, and the isolation conditions are: (a) X' contains only one sample, or all samples have the same eigenvalue; (b) iTree reaches the limit height.
[0028] Preferably, the specific content of S5 includes:
[0029] S51. Calculate the weighted path length h of the positive sample and the negative sample on the isolation tree iTree w (x):
[0030] Among them, the weighted edge calculation of each layer of nodes is:
[0031]
[0032] Among them, φ1 is the number of real samples of the current node, and φ2 is the number of artificially synthesized samples of the current node;
[0033] h w (x) is the number of weighted edges traversed when the traversal from the root node to an external node on the iTree terminates for each sample x, that is, the total number of weighted edges traversed in the WiTree:
[0034]
[0035] Among them, when h≤h lim When h is the total height of WiTree, when h>h lim When h=h lim , c(φ) is the harmonic parameter;
[0036]
[0037] Where H(i) is the harmonic number, which can be estimated as ln(i)+0.5772156649 (Euler constant);
[0038] S52. Calculate the average weighted path length E(h w (x)), which is the average number of weighted edges of the sample on all WiTrees:
[0039]
[0040] Where, t is the number of WiTrees contained in WiForest;
[0041] S53. Calculate the normalized abnormal scores s(x) of the positive sample and the negative sample, and determine whether the sample is abnormal;
[0042]
[0043] Among them, E(h w (x)) has a value range of [0,1];
[0044] S54. Set the anomaly ratio in the data set to α, then the anomaly score in the sample set is The samples of are abnormal samples, among which It represents rounding. The samples with abnormal scores in the first α proportion are removed from the forest composed of positive samples and the forest composed of negative samples respectively. The remaining positive samples and negative samples are synthesized to form the filtered source data set.
[0045] A cross-project defect prediction method based on isolation forests, based on the cross-project defect prediction sample filtering method based on isolation forests, comprising:
[0046] Step 1: randomly select a preset proportion of source project data sets of isomorphic cross-project software as source data sets for sample filtering to obtain a filtered source data set;
[0047] Step 2: Using a machine learning algorithm as a classifier, inputting the filtered source data set into the classifier to train the classifier and obtain a defect prediction model;
[0048] Step 3: Input the target data set of the tested software into the defect prediction model to obtain the prediction result of the target data set;
[0049] Step 4: Based on the prediction results, the performance evaluation index of the classification task is used to evaluate the performance of the tested software.
[0050] Preferably, before inputting the target data of the tested software into the defect prediction model in step three, the target data set is standardized using Z-score.
[0051] Preferably, steps 1 to 4 are repeated n times, where n is greater than 1.
[0052] It can be seen from the above technical solution that compared with the prior art, the present invention discloses a cross-project defect prediction sample filtering method and a prediction method based on isolation forest. For the classification task in software defect prediction, it is based on the simple and easy-to-use isolation forest method (iForest). By improving the isolation forest, the source data set is filtered, and the quality of the source project data in software defect prediction is improved. The problems of strong dependence on the target project, low efficiency, and poor prediction model performance in the current sample filtering method are solved, and data selection guidance for the software prediction model is realized, thereby shortening the software development cycle and saving costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0054] Figure 1 The accompanying drawing is a flow chart of the cross-project defect prediction method provided by the present invention;
[0055] Figure 2The accompanying drawings are statistical data of 36 open source Java projects provided by an embodiment of the present invention;
[0056] Figure 3 The accompanying drawings are SkewedF-Measure comparison results of the WIFLF method and the baseline method using RF as a classifier provided by an embodiment of the present invention;
[0057] Figure 4 The accompanying drawings show the G-Measure results of the WIFLF method and the baseline method using RF as a classifier according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] The embodiment of the present invention discloses a cross-project defect prediction sample filtering method based on isolation forest, comprising:
[0060] S1. Extract the dataset of isomorphic cross-project software as the source project dataset;
[0061] S2. Balancing the source project data set to obtain balanced data;
[0062] S3. Divide the balanced data into positive sample data and negative sample data, where the positive sample data is defective data and the negative sample data is non-defective data;
[0063] S4. Construct isolation forests for positive sample data and negative sample data respectively;
[0064] S5. Perform weighted processing on the isolation forest and perform sample filtering: calculate the weighted path length of each sample data on the isolation tree, calculate the average weighted path length of each sample data in the weighted isolation forest based on the weighted path length, calculate the outlier value of each sample data based on the average weighted path length, remove the abnormal samples of the weighted isolation forest according to the preset abnormal ratio and synthesize the remaining positive sample data and negative sample data to obtain the filtered source data set.
[0065] In order to further implement the above technical solution, before balancing the source project data in S2, the source project data set is also preprocessed. The specific contents of the data preprocessing include:
[0066] S21. Binarize the defect label information of the numerical type of the source project data set: when the defect label is greater than or equal to 1, it is marked as 1, indicating a defect; when the defect label is 0, it remains unchanged, indicating no defect;
[0067] S22. Selecting a source project data set as a source data sample;
[0068] S23. Eliminate duplicate samples: When there are exactly the same samples in the source data samples, only one sample is retained;
[0069] S24. Data standardization: Use Z-Score standardization technology to convert the deduplicated source data samples and target metrics into data with a mean of 0 and a variance of 1.
[0070] In this implementation, the SMOTE algorithm is used to balance the source project data set to obtain balanced data.
[0071] In this embodiment, the SMOTE algorithm is used to increase the unbalanced data in the source data to balanced data. For example, if there are 20 positive samples and 15 negative samples in the source data, the negative samples are increased to 20 through the SMOTE method.
[0072] In order to further implement the above technical solution, the balanced data is divided into positive sample data and negative sample data in S3 according to the defect label information.
[0073] In order to further implement the above technical solution, the specific contents of S4 include:
[0074] S41. Randomly select m samples from the positive sample data and negative sample data as training data sets Each sample consists of feature F, and the isolation forest consists of t isolation trees, iForest = {iTree1, iTree2, ..., iTree t}, select a subsample X' from X by random sampling without replacement, the size of X' is φ, and the limit height of the isolation tree is h lim ,h lim =ceiling(log2φ);
[0075] S42. Initialize iTree: construct a root node, which contains all subsamples X';
[0076] S43. Randomly select a feature f, f∈F, from all features, and randomly select a partition point p, which is between the maximum values of the selected feature, f min ≤p≤f max ;
[0077] S44. Input sample Compare the value of the feature f of the current root node with the selected partition point, and divide the root node into two child nodes, that is, if f < p, the sample is placed on the left child node, and if f ≥ p, the sample is placed on the right child node;
[0078] S45. Repeat S42-S44 until all samples are isolated. The isolation conditions are: (a) X' contains only one sample, or all samples have the same eigenvalue; (b) iTree reaches the limit height.
[0079] In order to further implement the above technical solution, the specific contents of S5 include:
[0080] S51. Calculate the weighted path length h of positive samples and negative samples on the isolation tree iTree w (x):
[0081] Among them, the weighted edge calculation of each layer of nodes is:
[0082]
[0083] Among them, φ1 is the number of real samples of the current node, and φ2 is the number of artificially synthesized samples of the current node; real samples refer to the data originally existing in the source data, and artificially synthesized samples refer to the data added after the SMOTE method.
[0084] h w (x) is the number of weighted edges traversed when the traversal from the root node to an external node on the iTree terminates for each sample x, that is, the total number of weighted edges traversed in the WiTree:
[0085]
[0086] Among them, when h≤h lim When h is the total height of WiTree, when h>h lim When h=h lim , c(φ) is the harmonic parameter;
[0087]
[0088] Where H(i) is the harmonic number, which can be estimated as ln(i)+0.5772156649 (Euler constant);
[0089] S52. Calculate the average weighted path length E(h) of positive samples and negative samples w (x)), which is the average number of weighted edges of the sample on all WiTrees:
[0090]
[0091] Where, t is the number of WiTrees contained in WiForest;
[0092] S53. Calculate the normalized abnormal scores s(x) of the positive and negative samples, and determine whether the samples are abnormal;
[0093]
[0094] Among them, E(h w (x)) has a value range of [0,1];
[0095] S54. Set the anomaly ratio in the data set to α, then the anomaly score in the sample set is The samples of are abnormal samples, among which It represents rounding. The samples with abnormal scores in the first α proportion are removed from the forest composed of positive samples and the forest composed of negative samples respectively. The remaining positive samples and negative samples are synthesized to form the filtered source data set.
[0096] In this embodiment, the abnormal ratio α≤0.5.
[0097] The traditional isolation forest constructs a forest for all data. In this embodiment, before constructing the weighted isolation forest, the source data is first divided into two groups according to the defect labels, and then the isolation forest is constructed. After the isolation forest is constructed, the normality and abnormality of the evaluation samples are evaluated by weighting the path length according to the ratio of real samples and artificially synthesized samples, thereby reducing the impact of artificially synthesized samples on the evaluation results and avoiding the evaluation of all samples being treated equally.
[0098] A cross-project defect prediction method based on isolation forests, based on a cross-project defect prediction sample filtering method based on isolation forests, comprising:
[0099] Step 1: randomly select a preset proportion of source project data sets of isomorphic cross-project software as source data sets for sample filtering to obtain a filtered source data set;
[0100] Step 2: Use a machine learning algorithm as a classifier, input the filtered source data set into the classifier to train the classifier, and obtain a defect prediction model;
[0101] Step 3: Input the target data set of the tested software into the defect prediction model to obtain the prediction result of the target data set;
[0102] Step 4: Based on the prediction results, use the performance evaluation indicators of the classification task to evaluate the performance of the software under test.
[0103] In order to further implement the above technical solution, before inputting the target data of the tested software into the defect prediction model in step three, the Z-score is used to standardize the target data set.
[0104] In order to further implement the above technical solution, steps 1 to 4 are repeated n times, where n is greater than 1.
[0105] In practical applications, n is greater than or equal to 30.
[0106] In this embodiment, 90% of the source project data sets are selected as source data samples, and steps 1 to 4 are repeated 30 times.
[0107] In this embodiment, if Figure 2In this paper, three open source JAVA defect databases are used for sample filtering and defect prediction, namely PROMISE (JURECZKO M, MADEYSKI L. Towards identifying software project clusters with regard to defect prediction;proceedings of the Proceedings of the 6th International Conference on Predictive Models in Software Engineering, F, 2010[C].), AEEEM (D'AMBROS M, LANZA M, ROBBES R. An extensive comparison of bug prediction approaches;proceedings of the Proceedings of MSR 2010(7th IEEE Working Conference on Mining Software Repositories), Cape Town, South africa, F, 2010[C]. IEEE Computer Society.) and ReLink (WU R, ZHANG H, KIM S, et al. ReLink: recovering links between bugs and changes;proceedings of the Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering,Szeged,Hungary,F,2011[C].ACM:2025120.).Defect datasets of 36 software projects are selected as source project data and divided into 4 groups. Each sample represents a software class module. Therefore, the granularity of all datasets is class-level. The metrics in the samples are independent variables and the class labels are dependent variables. The specific statistical information of these defect datasets includes project name and version, number of metrics, number of defects and defect rate. The numerical class labels are converted into binary types, that is, the number of defects in the last column of defect data labels in the dataset is set to 0 (no defects), and the number of defects greater than or equal to 1 is set to 1 (defective). The defect rate of a defect dataset refers to the ratio of the number of defective samples in the dataset to the total number of samples.
[0108] In this embodiment, in order to compare with the method proposed in the present invention (abbreviated as WIFLF), 12 comparison methods are selected, namely: BF, HISNN, TNB, DFAC, HSBF, CFPS, EASC, Bellwether (BNaive), Bellwether+TNB (BTNB), TCA+, Bellwether+TCA+ (BTCA+) and DMDA_JFR used the common machine learning algorithm Random Forest (RF). The experiments were conducted with the data processing software Python. All algorithms used default parameters. In order to avoid the impact of the randomness of the algorithm, each data set was executed 30 times, and the experimental results used the mean.
[0109] The commonly used classification model performance evaluation indicators skewed F-Measure and G-Measure are used for evaluation. In the classification model, the prediction result of the true positive class as the positive class is called a true positive example (TP), the prediction result of the true positive class as the negative class is called a false positive example (FP), the prediction result of the true negative class as the negative class is called a true negative example (TN), and the prediction result of the true negative class as the positive class is called a false negative example (FN).
[0110] Accuracy It indicates the proportion of positive examples among the predicted positive examples. The accuracy range is [0,1]. The larger the value, the higher the accuracy of the model prediction.
[0111] Recall It indicates the proportion of predicted positive examples among the true positive examples. The recall rate ranges from [0,1]. The larger the value, the higher the recall rate predicted by the model.
[0112] False alarm rate The value range is [0,1]. The smaller the PF value, the better the model.
[0113] The calculation method of skewed F-Measure is as follows, using β = 2, for model performance evaluation built on unbalanced data:
[0114]
[0115] G-Measure is the harmonic mean of PD and (1-PF), and its value range is [0,1]. The larger the value, the better the model:
[0116]
[0117] In order to evaluate whether there is a significant statistical difference between the method of the present invention and the comparative method, this paper adopts Wilcoxon Signed-Rank test (WILCOXON F. Individual Comparisons by Ranking Methods [J]. Biometrics, 1945, 1 (6).) to evaluate the performance of the WIFLF method and the baseline method respectively; if the p-value of the test result of the two groups of samples is less than 0.05, it means that the method of the present invention and the baseline method are significantly different at a confidence level of 95%; at the same time, on a data set, if WIFLF is significantly better than a certain method, it is recorded as "Win", if it is significantly worse than a certain method, it is recorded as "Loss", otherwise, it is recorded as "Tie", and "W\T\L" indicates the significant advantages and disadvantages of the method of the present invention and the comparative method.
[0118] Some representative results are as follows Figure 3 and Figure 4 As shown, except for "W\T\L", the values in the figure are omitted.
[0119] from Figure 2-Figure 4 From the analysis, it can be seen that the sample filtering method for isomorphic cross-project defect prediction based on isolation forest proposed in the present invention has achieved good results on most of the random versions of these 36 projects:
[0120] from Figure 3 It can be seen that the average skewed F-Measure obtained by WIFLF is 50.54%, which is 14.64% higher than other methods. In addition, it can be seen from the results of "W\T\L" that WIFLF is significantly better than the baseline method at 95% confidence level, because WIFLF wins 25 out of 36 data sets compared with BNaive, the best performing method in the baseline.
[0121] from Figure 4It can be seen that the average G-Measure obtained by WIFLF is 63.32%, which is 4.90% higher than other methods. In addition, it can be seen from the results of "W\T\L" that WIFLF is significantly better than the baseline method at a confidence level of 95%, because WIFLF wins 17 out of 36 data sets compared with BNaive, the best performing method in the baseline.
[0122] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0123] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cross-project defect prediction sample filtering method based on isolation forest, characterized in that: include: S1. Extract the dataset of isomorphic cross-project software as the source project dataset; S2. Balancing the source project data set to obtain balanced data; S3. Divide the balanced data into positive sample data and negative sample data, wherein the positive sample data is defective data and the negative sample data is non-defective data; S4. constructing an isolation forest for the positive sample data and the negative sample data respectively; S5. Perform weighted processing on the isolation forest and perform sample filtering: calculate the weighted path length of each sample data on the isolation tree, calculate the average weighted path length of each sample data in the weighted isolation forest based on the weighted path length, calculate the outlier value of each sample data based on the average weighted path length, remove the outlier samples of the weighted isolation forest according to a preset outlier ratio and synthesize the remaining positive sample data and negative sample data to obtain a filtered source data set; The specific contents of S5 include: S51. Calculate the weighted path length h of the positive sample and the negative sample on the isolation tree iTree w (x): The weighted edge of each layer node is calculated as: Among them, φ1 is the number of real samples of the current node, and φ2 is the number of artificially synthesized samples of the current node; h w (x) is the number of weighted edges traversed when the traversal from the root node to an external node on the iTree terminates for each sample x, that is, the total number of weighted edges traversed in the WiTree: Among them, when h≤h lim When h is the total height of WiTree, when h>h lim When h=h lim , c(φ) is the harmonic parameter; Where H(i) is the harmonic number, which can be estimated as ln(i)+0.5772156649 (Euler constant); S52. Calculate the average weighted path length E(h w (x)), which is the average number of weighted edges of the sample on all WiTrees: Where t is the number of WiTrees contained in WiForest, and i is the i-th weighted isolation tree; S53. Calculate the normalized abnormal scores s(x) of the positive sample and the negative sample, and determine whether the sample is abnormal; Among them, E(h w (x)) The value range is [0,1]; S54. Set the anomaly ratio in the data set to α, then the anomaly score in the sample set is The samples of are abnormal samples, among which It represents rounding. The samples with abnormal scores in the first α proportion are removed from the forest composed of positive samples and the forest composed of negative samples respectively. The remaining positive samples and negative samples are synthesized to form the filtered source data set.
2. According to the method for filtering cross-project defect prediction samples based on isolation forests in claim 1, it is characterized in that: Before balancing the source project data in S2, the source project data set is also preprocessed. The specific content of the data preprocessing includes: S21. Binarize the defect label information of the numerical type of the source project data set: when the defect label is greater than or equal to 1, it is marked as 1, indicating a defect; when the defect label is 0, it remains unchanged, indicating no defect; S22. Select the source project data set as the source data sample; S23. Eliminate duplicate samples: when there are exactly the same samples in the source data samples, only one sample is retained; S24. Data standardization: Use Z-Score standardization technology to convert the deduplicated source data samples and target metrics into data with a mean of 0 and a variance of 1.
3. The method for filtering cross-project defect prediction samples based on isolation forest according to claim 2, characterized in that: The dividing of the balanced data into positive sample data and negative sample data in S3 is performed according to the defect label information.
4. The method for filtering cross-project defect prediction samples based on isolation forest according to claim 1, characterized in that: The specific contents of S4 include: S41. Randomly select m samples from the positive sample data and the negative sample data as training data sets Each sample consists of feature F, and the isolation forest consists of t isolation trees, iForest = {iTree1, iTree2, ..., iTree t }, select a subsample X' from X by random sampling without replacement, the size of X' is φ, and the limit height of the isolation tree is h lim ,h lim =ceiling(log2φ); S42. Initialize iTree: construct a root node, which contains all subsamples X'; S43. Randomly select a feature f, f∈F, from all features, and randomly select a partition point p, the partition point is between the maximum values of the selected feature, f min ≤p≤f max ; S44. Input sample Compare the value of the feature f of the current root node with the selected partition point, and divide the root node into two child nodes, that is: if f < p, the sample is placed on the left child node, and if f ≥ p, the sample is placed on the right child node; S45. Repeat S42-S44 until all samples are isolated, and the isolation conditions are: (a) X' contains only one sample, or all samples have the same eigenvalue; (b) iTree reaches the limit height.
5. A cross-project defect prediction method based on isolation forest, based on the cross-project defect prediction sample filtering method based on isolation forest according to claims 1-4, characterized in that: include: Step 1: randomly select a preset proportion of source project data sets of isomorphic cross-project software as source data sets for sample filtering to obtain a filtered source data set; Step 2: Using a machine learning algorithm as a classifier, inputting the filtered source data set into the classifier to train the classifier and obtain a defect prediction model; Step 3: input the target data set of the tested software into the defect prediction model to obtain the prediction result of the target data set; Step 4: Based on the prediction results, the performance evaluation index of the classification task is used to evaluate the performance of the tested software.
6. The cross-project defect prediction method based on isolation forest according to claim 5 is characterized in that: In step 3, before inputting the target data of the tested software into the defect prediction model, the target data set is normalized using Z-score.
7. The cross-project defect prediction method based on isolation forest according to claim 6 is characterized in that: Repeat steps 1 to 4 for n times, where n is greater than 1.
Citation Information
Patent Citations
Cross-project defect prediction method based on data screening and data oversampling
CN107391369A
Prediction method for unbalanced data set based on isolated forest learning
CN112070125A