An Instant Software Defect Prediction Method Based on Imbalanced Distribution of Data Classes

By performing data distribution analysis and region division on the instant software defect prediction data set, a more targeted undersampling strategy is adopted to process noise and boundary samples, the class imbalance problem is solved and the recognition rate and accuracy of the model are improved.

CN114138632BActive Publication Date: 2025-06-03HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111341822.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-06-03
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

There is a class imbalance problem in real-time software defect prediction, which leads to a low recognition rate of defect changes in the model, and traditional undersampling technology fails to effectively consider the impact of data distribution and noise data.

Method used

By analyzing the data distribution of the samples, the samples are divided into noise areas, boundary areas and safety areas, and a more targeted undersampling strategy is adopted to process noise samples and boundary samples, and finally trained through a random forest model.

Benefits of technology

This improves the recognition rate of defect changes by the model, reduces model training errors, and enhances the accuracy and efficiency of real-time software defect prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114138632B_ABST
    Figure CN114138632B_ABST
Patent Text Reader

Abstract

The present invention relates to an instant software defect prediction method based on unbalanced data class distribution. First, the present invention calculates the sample nearest neighbor set, analyzes the data distribution of the samples, identifies the samples in the noise area, deletes the defective samples in the noise area, and replaces the labels of the non-defective samples in the noise area with the labels of the defective samples; then recalculates the nearest neighbor set of each sample, analyzes the data distribution of the current relatively clean data set, divides the boundary area and the safe area, and deletes the non-defective samples in the boundary area; finally, performs random undersampling on the training data set to balance the number of samples. By analyzing the data distribution of the samples, the present invention removes the noise samples in the sample set and the samples that have a negative impact on the identification of defective samples with a more accurate strategy, thereby improving the identification performance of the model for defective samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is a method for dealing with imbalance problems based on data distribution, aiming to use this technology to find more defective samples under limited resource conditions and improve the potential defects detected by each code inspection work unit. Specifically, it relates to an instant software defect prediction method based on data class imbalance distribution. Background Art

[0002] Instant software defect prediction is a new idea in the field of software defect prediction, which takes the code changes submitted by developers each time as the object and predicts whether defects will occur. Instant software defect prediction has the following advantages: (1) Fine-grained. (2) Immediacy. (3) Easy traceability. As shown in Table 1, there is a relatively obvious class imbalance problem in the instant software defect prediction dataset, that is, the number of defective changes is much less than the number of non-defective changes. Due to the insufficient number of defective samples, the model pays less attention to defective changes. And the classifier considers the global performance index when training the model, and the model tends to predict changes as non-defective changes, which makes the recognition rate of defective changes in the instant software defect prediction dataset low. However, the correct recognition of defective changes is the main purpose of instant software defect prediction.

[0003] The current mainstream strategies for solving class imbalance can be divided into the following three categories: 1) Sampling techniques. Delete majority class samples or synthesize minority class samples to solve the problem at the data level and balance the number of majority class samples and minority class samples. 2) Cost-sensitive learning. Set a higher misclassification cost for the minority class, so that the classifier pays more attention to minority class samples. 3) Ensemble learning. Use the better generalization performance of ensemble learning to improve classification accuracy.

[0004] However, these technologies still have certain limitations. For example: Sampling techniques generally solve the imbalance problem at the data level. The most commonly used random undersampling method balances the number of majority class and minority class. However, simply random undersampling only considers the imbalance in the number of the two types of samples, and fails to consider the complex data distribution of the two types of samples and the negative impact of the noise data in the dataset on imbalance processing and model learning. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention proposes an instant software defect prediction method based on data class imbalance distribution.

[0006] Starting from the perspective of data distribution, the present invention divides samples into three regions - a noise region, a boundary region, and a safe region by analyzing the data distribution of samples in the current dataset. At different stages, two types of samples in each region are processed separately, replacing the original random undersampling with a more targeted undersampling strategy. First, by calculating the nearest neighbor set of samples, the data distribution of samples is analyzed, and noise samples / outlier samples existing in the dataset are found and processed differently; then, by analyzing the data distribution of the current relatively clean dataset, the boundary region and the safe region are divided for processing; finally, the quantity of the training dataset is balanced.

[0007] An instant software defect prediction method based on data class imbalance distribution of the present invention specifically includes the following steps:

[0008] Step 1) Data preprocessing.

[0009] Calculate the correlation between metrics, process highly correlated metrics, and perform logarithmic transformation on the metrics. Specifically: Calculate the correlation between each metric. Since NF and ND, REXP and EXP are highly correlated, ND and REXP are excluded. Since LA and LD are highly correlated with LT, LA and LD are normalized by dividing by LT. Since LT and NUC are highly correlated with NF, LT and NUC are normalized by dividing by NF; perform logarithmic transformation on each metric, except FIX. The specific description of the metrics is shown in Table 2.

[0010] Step 2) Obtain the nearest neighbor set of each sample.

[0011] 2-1. Calculate the distance between samples.

[0012] Assume that for any sample x i , x il represents the respective attribute values of sample x i , Y i represents the label of sample x i , n represents the number of attributes of sample x i , represents the nearest neighbor set of sample x i . Use formula (1) to calculate the distance d(x i , x j ) between sample x i and sample x j .

[0013]

[0014] 2-2. Determine the nearest neighbor set of the sample according to the sample spacing d(x i , x j ).

[0015] Sample x iBy i ,x j ) Arrange in ascending order, select the first k samples, and get the distance sample x i The most recent K samples make up sample x i The neighbor set of Where k is a parameter set by yourself.

[0016] Step 3) Identify and process the noise samples in the training sample set.

[0017] This method determines the data distribution around the sample by analyzing the sample's nearest neighbor set. The first stage is mainly to deal with the noise problem in the training sample set, so the samples in the training sample set are classified into noise area samples and non-noise area samples. Different treatments are performed on the two types of samples in the noise area.

[0018] 3-1. Calculate the defect density of the training sample set.

[0019] Assume that the original training set is T, in which the defective samples are the minority class, denoted as T + ; The samples without defects are the majority class, denoted as T - . Use formula (2) to calculate the defect density D of the training sample set T , where nums(T + ) represents the defective sample set T + nums(T) represents the number of samples in the original training set T.

[0020] D T =nums(T + ) / nums(T) (2)

[0021] 3-2. Calculate the defect density of the sample.

[0022] The nearest neighbor set of each sample in step 3-1 is calculated by formula (3) The defect density D i , where Y nsj Represents the neighbor set of step 2-2 The label of the jth sample in Y nsj 0 means step 2-2 neighbor set The jth sample of Y is a non-defective sample. nsj 1 means step 2-2 neighbor set The j-th sample of is a defective sample.

[0023]

[0024] 3-3. Divide the samples into regions.

[0025] As shown in formula (4), if the sample x i is a defect-free sample (majority class sample), by comparing the defect density D i of the neighborhood set of the current sample x i with the defect density D T of the training sample set, the region of the sample x i is divided. When the defect density D i is greater than the defect density D T of the training sample set, the sample x i is divided into the noise region; when the defect density D i is less than the defect density D T of the training sample set, the sample x i is judged as a non-noise region. In the instant software defect prediction dataset, defective samples are minority class samples, with a small number of samples and being the classes that need to be identified as the key in this field. Therefore, the noise identification for defective samples is more stringent. As shown in formula (5), if the sample x i is a defective sample (minority class sample), when the defect density D i is equal to 0, that is, all samples in the neighborhood set i of the current defective sample x are defect-free samples, then the sample x i is judged to be in the noise region; when the defect density D i is greater than 0, the sample x i is divided into the non-noise region.

[0026] if Y i is 0 (defect-free sample):

[0027]

[0028] if Y i is 1 (defective sample):

[0029]

[0030] 3-4. Process samples in the noise region.

[0031] As shown in formula (6), discard the defective (minority class) samples in the noise region; convert the labels of the defect-free (majority class) samples in the noise region into the labels of the defective (minority class) samples, so as to increase the number of defective (minority class) samples.

[0032]

[0033] Step 4) Identify and process boundary samples.

[0034] After obtaining a relatively clean training sample set T′ through step 3), the second stage of data processing is entered to identify and process the defect-free (majority class) samples in the boundary area.

[0035] 4-1. Return to step 2) and recalculate the neighbor set T′ of each sample xins .

[0036] 4-2. Determine sample x i Is it a non-defective (majority class) sample? i It is a non-defective (majority class) sample, and goes to step 3-2 to calculate the current sample x i The defect density D′ i When the sample x i The defect density D′ i If it is greater than 0, that is, if there are defective (minority class) samples among the k neighbors around the current defect-free (majority class) sample, the sample is considered to be in the boundary area, and the defect-free (majority class) samples in the boundary area are discarded to form a new training sample set T″.

[0037] Step 5) Balance the data set. Determine whether the new training sample set T″ is a balanced data set. If T″ is not a balanced sample set, perform random undersampling to achieve a balance in quantity and obtain a new training sample set T′″.

[0038] Step 6) Use the random forest model to train the model. Train the sample set obtained in step 5) to obtain a prediction model.

[0039] Beneficial effects of the present invention:

[0040] Starting from the perspective of data distribution, this technology discusses the possible data distribution between the two types of samples and performs targeted processing on samples with different data distributions. It aims to restore the original distribution of the training data set more realistically when balancing the training data set, and to clarify the distribution boundary of the two types of samples, so that the model has a better training effect.

[0041] This technology takes into account the noise samples present in the data set and improves the traditional undersampling technology, reducing the model training error caused by noise samples, enabling the model to more accurately identify defect changes and reduce the cost of real-time software defect prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Overall flow chart of the instant software defect prediction method based on data class imbalance distribution processing method. DETAILED DESCRIPTION

[0043] The present invention is described in detail below with reference to the accompanying drawings and in combination with the real-time software defect prediction data set.Figure 1 As shown below, the specific steps are as follows:

[0044] Step 1: Data preprocessing. Calculate the correlation between each metric. Since NF and ND, REXP and EXP are highly correlated, ND and REXP are excluded. Since LA and LD are highly correlated with LT, LA and LD are normalized by dividing by LT. Since LT and NUC are highly correlated with NF, LT and NUC are normalized by dividing by NF. Perform logarithmic transformation on each metric (except FIX); the specific descriptions of the metrics are shown in Table 2.

[0045] Step 2: Use the ten-fold cross-validation method to divide the data set in Step 1 into a training set and a test set. The training set is used to train the model, and the test set is used to test the model.

[0046] Step 3: Calculate the distance between each sample in the training set obtained in Step 2. Select the k samples closest to the sample according to the distance between samples to form the neighbor set of the sample.

[0047] Step 4: Calculate the defect density of the sample according to the neighbor set of each sample in Step 3. Compare the defect density of the training sample set with the defect density of each sample, and divide the samples into the noise area and the non-noise area.

[0048] Step 5: Discard the minority (defective) samples in the noise area, and convert the labels of the majority (non-defective) samples in the noise area into minority (defective) labels to obtain a new training set.

[0049] Step 6: Repeat Steps 3-4 for the new sample set to obtain a new sample neighbor set and the sample defect density. Compare the defect density of the training sample set with the defect density of each sample, and divide the samples into the safe area and the boundary area.

[0050] Step 7: Delete the non-defective (majority) samples in the boundary area, and perform random undersampling on the remaining samples to form the final training set samples.

[0051] Step 8: Use the random forest model to train the sample set obtained in Step 7 to obtain a prediction model.

[0052] Project Time Number of Changes Defective Changes Non-Defective Changes Imbalance Rate bugzilla 1998 / 08-2006 / 12 4620 1696 2924 1.724 columba 2002 / 11-2006 / 07 4455 1361 3094 2.273 jdt 2001 / 05-2007 / 12 35386 5089 30297 5.953 platform 2001 / 02-2007 / 12 64250 9452 54798 5.798 mozilla 2000 / 01-2006 / 12 98275 5149 93126 18.086 postgres 1996 / 07-2010 / 05 20431 5119 15312 2.991

[0053] Table 1

[0054]

[0055] Table 2.

Claims

1. An instant software defect prediction method based on the imbalanced distribution of data classes, characterized in that, it includes the following steps: Step 1) Data preprocessing; Calculate the correlation between metrics, process highly correlated metrics, and perform logarithmic transformation on the metrics; Step 2) Obtain the neighbor set of each sample; Step 2-1: Calculate the distance between samples; Assume that for any sample x i , x il represents the respective attribute values of sample x i , Y i represents the label of sample x i , n represents the number of attributes of sample x i ; represents the neighbor set of sample x i ; calculate the distance d(x i , x j ) between sample x i and sample x j using formula (6); Step 2-2: Determine the nearest neighbor set of the samples according to the sample spacing d(x i , x j ). Sample x i By sorting d(x i , x j ) in ascending order and selecting the first k samples, the K samples closest to sample x i are obtained to form the nearest neighbor set of sample x i ; where k is a parameter set by oneself; Step 3) Identify and process the noise samples in the training sample set; Step 3-1: Calculate the defect density of the training sample set; Assume that the original training set is T, and the defective samples are the minority class, denoted as T + ; the non-defective samples are the majority class, denoted as T - ; Use formula (1) to calculate the defect density D of the training sample set T , where nums(T + ) represents the number of samples in the defective sample set T + ; nums(T) represents the number of samples in the original training set T; D T = nums(T + ) / nums(T) (1) Step 3-2: Calculate the defect density of the sample; Calculate the nearest neighbor set of each sample through formula (2). The defect density D i , where Y nsj represents the label of the j-th sample in the nearest neighbor set , Y nsj being 0 indicates that the j-th sample in the nearest neighbor set is a defect-free sample, and Y nsj being 1 indicates that the j-th sample in the nearest neighbor set is a defective sample; Step 3-3: Divide the samples into regions; As shown in formula (3), if the sample x i is a defect-free sample, by comparing the defect density D i of the neighborhood set of the current sample x i with the defect density D T of the training sample set, the region of the sample x i is divided; when the defect density D i is greater than the defect density D T of the training sample set, the sample x i is divided into the noise region; when the defect density D i is less than the defect density D T of the training sample set, the sample x i is judged as a non-noise region; as shown in formula (4), if the sample x i is a defective sample, when the defect density D i is equal to 0, that is, all samples in the neighborhood set i of the current defective sample x are defect-free samples, then the sample x i is judged to be in the noise region; When the defect density D i is greater than 0, the sample x i is classified into the non-noise region; Defect-free samples: Defective samples: Step 3-4: Process the samples in the noise region; As shown in formula (5), discard the defective samples in the noise region; convert the labels of the defect-free samples in the noise region into the labels of the defective samples to increase the number of defective samples; Step 4) Identify and process the boundary samples; Step 4-1: After obtaining the relatively clean training sample set T′ through Step 3), enter the second stage of data processing to identify and process the defect-free samples in the boundary region; Step 4-2: Return to Step 3) to recalculate the nearest neighbor sets of each sample Step 4-3: Determine whether the sample x i is a defect-free sample; If the sample x i is a defect-free sample, proceed to step 3-2 to calculate the defect density D i ' of the current sample x i ; when the defect density D i ' of the sample x i ' is greater than 0, that is, there are defective samples among the k neighbors around the current defect-free sample, then it is considered that the sample is in the boundary region, and the defect-free samples in the boundary region are discarded to form a new training sample set T″; Step 5) Balance the data set; Step 6) Use a random forest to train the model.

2. The instant software defect prediction method based on the imbalanced distribution of data classes according to claim 1, characterized in that The specific implementation of the data preprocessing described in Step 1 is as follows: Calculate the correlation between each metric. Since NF and ND, REXP and EXP are highly correlated, ND and REXP are excluded; since LA and LD are highly correlated with LT, LA and LD are normalized by dividing by LT; since LT and NUC are highly correlated with NF, LT and NUC are normalized by dividing by NF; perform logarithmic transformation on each metric, except for FIX.

3. The instant software defect prediction method based on the imbalanced distribution of data classes according to claim 1, characterized in that: The specific implementation of the balanced data set described in Step 5 is as follows: Judge whether the new training sample set T″ is a balanced data set. If T″ is not a balanced sample set, perform random undersampling to achieve balance in quantity and obtain a new training sample set T”'.

Citation Information

Patent Citations

  • Multi-feature software defect comprehensive prediction method based on unbalanced noise set

    CN111782512A

  • Software defect prediction method based on class imbalance learning algorithm

    CN112465040A