A Software Defect Prediction Method, Device and Storage Medium Based on Multi-Model Selection
By incrementally training the random forest model, detecting concept drift, balancing data blocks and selecting the best-performing model for prediction, the problems of data dynamics and category imbalance of software modules are solved, and the timely identification and prediction of software defects is realized, and the reliability and stability of the software are improved.
Patent Information
- Application Number
- CN202210137455.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-02-15
AI Technical Summary
During the software development process, the dynamicity and category imbalance of software module data lead to increased difficulty in predicting software defects, and it is difficult for the existing technology to effectively and timely identify defective software modules.
The software defect prediction method based on multi-model selection is adopted, and the random forest model is trained incrementally, concept drift is detected, data blocks are balanced using the SMOTE algorithm, multiple stream data integration classification models are established, and the best performance model is selected for prediction.
It realizes the rapid and effective identification of defective software modules, improves the reliability and stability of the software, and adapts to the dynamic changes of software module data.
Smart Images

Figure CN114546847B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a software defect prediction method, device and storage medium based on multi-model selection. Background Art
[0002] With the rapid development of technologies such as big data, cloud computing, and parallel computing, corresponding application scenarios have become increasingly rich, such as transportation, commerce, and medical and health. At the same time, the development of high-tech has also accelerated the emergence and development of various software. During the software development process, it is necessary to strictly follow user requirements, otherwise the software development process is prone to errors. Such problems that affect the normal operation of software or programs are called software defects. Software defects will seriously affect software development. If not detected and corrected in time, software defects will further accumulate or spread, thus affecting the reliability and stability of the software. Therefore, predicting software defects is a very important task with great research and practical value.
[0003] The software defect prediction task is to timely and effectively identify software modules that may have defects for defect correction to ensure the correctness of software development. Software module data is generated in real time during the software development process, and its data distribution will continuously change with factors such as software development conditions. Therefore, software module data can be regarded as stream data, and the dynamic nature of its data distribution is called concept drift. Software module data is also called software module stream data, so the method of stream data classification can be used to predict software defects.
[0004] Compared with the manual method for software defect detection, the method based on stream data classification can more effectively and real-time ensure the reliability and stability of the software. Software module data is divided into two categories: defective and non-defective, and the number of defective software models is usually less than that of non-defective software modules. If software defect prediction is regarded as a binary classification problem of stream data, then this classification problem faces a data stream environment with class imbalance. Among them, defective software module data belongs to small samples, while non-defective software module data belongs to large samples.
[0005] Software module flow data is generated in real time, and its data distribution changes continuously over time. This phenomenon is called concept drift. According to the occurrence speed of concept drift, concept drift can be divided into three types: abrupt, gradual, and incremental. The processing mechanisms of concept drift are divided into two types: proactive and passive. The proactive type uses concept drift detection to identify the dynamics of the software module data distribution based on the stability of statistics. After detecting concept drift, the current flow data classification model is adjusted or reconstructed in a timely manner to adapt to the new environment. The passive type method does not require an additional concept drift detection mechanism. By adaptively adjusting the weights of the base classifiers in the flow data integrated classification model, it can passively adapt to the changing software module flow data environment. Compared with the passive type method, the proactive drift detection mechanism can adjust the model more timely with software module flow data that conforms to the new data distribution. Summary of the Invention
[0006] The present invention aims to provide a software defect prediction method, device, and storage medium based on multi-model selection; the present invention can quickly and effectively improve the model's ability to identify defective software modules, thereby ensuring the reliability and stability of the software.
[0007] One aspect of the present invention provides a software defect prediction method based on multi-model selection, including the following steps:
[0008] Step 1) Use the first data collection mechanism to collect newly arrived software module flow data and incrementally train the random forest model M0. At the same time, update the statistics in the confusion matrix and the statistics of the sample mean using the new data.
[0009] Step 2) Use the sample mean updated at the current moment in the concept drift detection mechanism to obtain the small sample balanced data blocks D1 and D2.
[0010] Step 3) Based on the SMOTE algorithm, perform oversampling on the obtained data blocks D1 and D2 to obtain the class distribution balanced data blocks D1' and D2' respectively.
[0011] Step 4) On the obtained data blocks D1, D2, D1', and D2', establish random forest classification models M1, M2, M3, and M4 respectively.
[0012] Step 5) Calculate the G-mean performance values of the trained flow data classification models M0, M1, M2, M3, and M4 for the latest software module flow data, and obtain the software defect prediction model M based on multi-model selection.
[0013] Step 6) Use the software defect prediction model M to predict the class of software defect data.
[0014] Preferably, in step 2), the concept drift detection mechanism ADWIN is used to detect whether there is concept drift in the current data.
[0015] Preferably, the concept drift detection mechanism ADWIN includes: a warning level and a drift level, and data blocks D1 and D2 are formed based on these two levels.
[0016] Preferably, ADWIN identifies the stability of software module flow data by detecting the change in the current sample mean. If the warning level is reached, the first data collection mechanism no longer collects software module flow data, and data block D1 is formed; otherwise, a second data collection mechanism is created to collect software module flow data after the warning level until the current data distribution environment reaches the drift level, thereby forming data block D2.
[0017] Preferably, in step 3), the SMOTE algorithm obtains data blocks D1' and D2' with balanced class distributions by generating new small-sample balanced data distributions of data blocks D1 and D2.
[0018] Preferably, in step 4), M3 and M4 are streaming data integration classification models established on balanced data blocks, and the class distributions of the training data in M1 and M2 are usually unbalanced.
[0019] Preferably, in step 5), the G-mean performance value is calculated based on the values in the confusion matrix. Based on the G-mean performance value, the model M with the best performance among M0, M1, M2, M3, and M4 is selected to replace the streaming data integration classification model M0 that is currently being incrementally trained.
[0020] Preferably, in step 6), if the prediction result is +1, it is a defective class sample; if the prediction result is -1, it is determined as a non-defective class sample.
[0021] On the other hand, the present invention provides a software defect prediction device based on multi-model selection, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the software defect prediction method based on multi-model selection.
[0022] On yet another aspect, the present invention provides a computer-readable storage medium storing a computer program for executing the above-mentioned software defect prediction method based on multi-model selection.
[0023] Compared with the prior art, the beneficial effects produced by the present invention are:
[0024] By effectively identifying defective software modules in real time, the present invention facilitates timely adjustment by software developers, thereby ensuring the reliability and stability of the software. First, a random forest model M0 is trained incrementally one by one using the incremental learning method. At the same time, a first data collection mechanism is used to retain the software module flow data arriving one by one. Then, the ADWIN concept drift detection mechanism is used to detect the dynamics of the sample mean, where two thresholds are set: the warning level and the drift level. All software module flow data before the warning level are retained in the first data collection mechanism, thereby obtaining the data block D1. And all software module flow data between the warning level and the drift level are retained in the second data collection mechanism, thereby obtaining the data block D2. Then, the SMOTE algorithm is used to balance the class distributions in D1 and D2, obtaining the data blocks D1' and D2' respectively. For the obtained data blocks D1, D2, D1', and D2', four classification models M1, M2, M3, and M4 are established based on the random forest model respectively. Then, based on the characteristics of the data distribution, the model with the best performance among the five software defect flow data classification models M0, M1, M2, M3, and M4 is selected as the final software defect prediction model M. Finally, the class of software defect data is predicted based on the M classification model of the random forest, thereby realizing software defect prediction. Description of the Drawings
[0025] Figure 1 It is a schematic diagram of a software defect prediction method based on multi-model selection proposed by the present invention.
[0026] Figure 2 It is a structural diagram of the device of the present invention. Detailed Embodiment
[0027] The present invention mainly includes the following steps:
[0028] Step 1) Use the first data collection mechanism to collect software module flow data, update the confusion matrix and the sample mean statistic value at the same time, and incrementally train the random forest model M0 using each newly arrived software module flow data. Compared with the single classifier model, the random forest model has better generalization ability by retaining multiple decision trees. Among them, the software module flow data has the characteristics of concept drift and class imbalance.
[0029] Step 2) Use the updated sample mean statistic in the concept drift detection of software module flow data to obtain data blocks D1 and D2. The ADWIN method is used for concept drift detection. First, if the mean of the software module flow data reaches the warning level, the first data collection mechanism stops collecting software module flow data. Meanwhile, a second data collection mechanism is created to collect software module flow data from the warning level to the drift level. Among them, the software module flow data in the first data collection mechanism forms data block D1, and the software module flow data in the second data collection mechanism forms data block D2.
[0030] Step 3) For the obtained data blocks D1 and D2, use the SMOTE algorithm to obtain data blocks D1’ and D2’ with balanced class distributions respectively. The SMOTE algorithm can balance the class distribution of the data block by generating minority samples. The number of new software module flow data to be generated in each data block depends on the number of non-defective and defective software module flow data items in that data block.
[0031] Step 4) For the obtained data blocks D1, D2, D1’, and D2’, train stream data integrated classification models M1, M2, M3, and M4 respectively. Among them, the used stream data classification model adopts a random forest model. By retaining multiple decision trees, it can ensure that the classification model has good generalization performance. Data blocks D1 and D2 are usually imbalanced in class distribution, while data blocks D1’ and D2’ are balanced in class distribution.
[0032] Step 5) For the trained models M0, M1, M2, M3, and M4, evaluate their classification performance on the latest arriving software module flow data and obtain the final software defect prediction model M. Based on the confusion matrix, the G-mean value of the integrated classifier can be obtained. This performance value has good robustness to the class imbalance problem. Select the model with the highest G-mean value among the five stream data integrated models to replace the current incrementally trained stream data integrated classification model M0, that is, select the software defect prediction model based on the characteristics of the data distribution.
[0033] Step 6) Use the stream data integrated classification model M as the software defect prediction model to predict the class of the software module data. If the prediction result is +1, it is judged as a defective class sample; if the prediction result is -1, it is judged as a non-defective class sample.
[0034] Preferably, in step 1), the training of the M0 model is obtained by incremental training, and the random forest model is adopted. Since the class distribution of software model flow data is usually unbalanced, the class distribution of the training data used to train M0 is often unbalanced, and the recognition rate of defective software model flow data is usually not high. The first data collection mechanism retains newly arrived software module flow data items, which can be used for the model training of M0, and updates the sample mean and confusion matrix with the data therein.
[0035] Preferably, in step 2), the stability of the sample mean can be used for the dynamic detection of software module flow data. The ADWIN mechanism is used, and two levels are set therein: the warning level and the drift level. If the change in the sample mean reaches the warning level, the first data collection mechanism stops collecting data, and the data in the first data collection mechanism constitutes data block D1. Then, a second data collection mechanism is created to collect software module data after the warning level until the current data distribution environment reaches the drift level, thereby constructing data block D2.
[0036] Preferably, in step 3), the class distributions in data block D1 and data block D2 are usually unbalanced. SMOTE generates new defective software module data to obtain data blocks D1' and D2' with balanced class distributions respectively. Among them, the number of newly generated defective class samples is equal to the number of large samples in the data block minus the number of small samples.
[0037] Preferably, in step 4), the model training of M1, M2, M3, and M4 is respectively based on the obtained data blocks D1, D2, D1', and D2'.
[0038] Preferably, in step 5), from the statistics in the confusion matrix, the G-mean value of the model can be obtained, which has good robustness to the class imbalance problem. First, calculate the G-mean performance values of M1, M2, M3, M4, and M0 on the latest software module flow data. Based on the data distribution characteristics of the new software module flow data, select the model with the best performance as the software defect prediction model.
[0039] Preferably, in step 6), use the trained software defect prediction model to predict newly arrived software module flow data.
[0040] Example, as Figure 1 shown, the model mainly includes a software module flow data integration classification mechanism based on incremental learning, a software module flow data block division mechanism based on concept drift detection, a software module flow data block resampling mechanism based on SMOTE, a software module flow data integration classification mechanism based on data block division, and a software defect prediction mechanism based on model selection.
[0041] The software module flow data has the problems of concept drift and class imbalance. Data items arrive one by one, and the random forest model M0 is obtained through incremental training. At the same time, a data collection mechanism is used to save each newly arrived software module data, and the statistics in the confusion matrix and the statistics of the sample mean are updated. Then, based on the ADWIN concept drift detection mechanism, the stability of the sample mean is monitored in real time, and data blocks D1 and D2 are obtained respectively. The software module flow data distributions in D1 and D2 are different. Then, D1 and D2 are oversampled using SMOTE to obtain data blocks D1' and D2' with balanced class distributions respectively. Among them, the number of small samples to be generated in the data block depends on the number of large samples and small samples in the data block respectively.
[0042] Based on the software module flow data blocks D1, D2, D1' and D2', forest classification models M1, M2, M3 and M4 are obtained respectively. Among them, M3 and M4 are streaming data integrated classification models established on balanced data blocks. The class distributions of the training data in M1 and M2 are usually unbalanced. Then, based on the confusion matrix, the G-mean classification performance of M0, M1, M2, M3 and M4 on the latest software module flow data is obtained, and the model with the best performance is selected as the software defect prediction model M. For software module flow data distributed in different feature spaces, models suitable for their data distribution characteristics are selected for classification respectively. Finally, the software defect prediction model M is used to predict the class of newly arrived software module data. Among them, the class distribution of the software module flow data is usually unbalanced. If the prediction result is +1, it is judged as a defective class sample; if the prediction result is -1, it is judged as a non-defective class sample.
[0043] Embodiments of the present invention can be applied to network devices. The embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. The computer program is used to execute the method determined in the above steps 1)-6). From the hardware level, as Figure 2 shown, it is the hardware structure diagram of the software defect prediction device based on multi-model selection of the present invention. In addition to Figure 2 the shown processor, network interface, memory and non-volatile memory, the device usually may also include other hardware for expansion at the hardware level. On the other hand, the present application also provides a computer-readable storage medium storing a computer program, and the computer program is used to execute the method determined in the above steps 1)-6).
[0044] For the embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative, and can be understood and implemented by those of ordinary skill in the art without creative efforts.
[0045] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.
[0046] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device.
[0047] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A software defect prediction method based on multi-model selection, characterized in that: It includes the following steps: Step 1) Use the first data collection mechanism to collect the newly arrived software module flow data and incrementally train the random forest classification model M0; at the same time, use the new data to update the statistics in the confusion matrix and the statistics of the sample mean; Step 2) Use the sample mean updated at the current moment in the concept drift detection mechanism to obtain the small sample balanced data blocks D1 and D2; Step 3) Based on the SMOTE algorithm, perform oversampling on the obtained data blocks D1 and D2 to obtain the data blocks D1' and D2' with balanced class distributions respectively; the SMOTE algorithm generates the data distributions of the new small sample balanced data blocks D1 and D2, thereby obtaining the data blocks D1' and D2' with balanced class distributions; Step 4) On the obtained data blocks D1, D2, D1' and D2', establish the random forest classification models M1, M2, M3 and M4 respectively; M3 and M4 are the random forest classification models established on the data blocks D1' and D2'; Step 5) Calculate the G-mean performance values of the trained random forest classification models M0, M1, M2, M3 and M4 for the latest software module flow data, and obtain the software defect prediction model M based on multi-model selection; Step 6) Use the software defect prediction model M to predict the category of software defect data.
2. The software defect prediction method based on multi-model selection according to claim 1, characterized in that: In step 2), the concept drift detection mechanism ADWIN is used to detect whether there is a concept drift in the current data.
3. The software defect prediction method based on multi-model selection according to claim 2, characterized in that: The concept drift detection mechanism ADWIN includes: a warning level and a drift level, and forms the data block D1 and the data block D2 based on the warning level and the drift level.
4. The software defect prediction method based on multi-model selection according to claim 3, characterized in that: The concept drift detection mechanism ADWIN identifies the stability of the software module flow data by detecting the change of the current sample mean; If the warning level is reached, the first data collection mechanism no longer collects the software module flow data to form the data block D1; and then create a second data collection mechanism to collect the software module flow data after the warning level until the current data distribution environment reaches the drift level, thereby forming the data block D2.
5. The software defect prediction method based on multi-model selection according to claim 1, characterized in that: In step 5), the G-mean performance value is calculated based on the values in the confusion matrix, and based on the G-mean performance value, the model with the best performance among M0, M1, M2, M3 and M4 is selected as the software defect prediction model M to replace the currently incrementally trained random forest classification model M0.
6. The software defect prediction method based on multi-model selection according to claim 1, characterized in that: In step 6), if the prediction result is +1, it is a defective class sample, and if the prediction result is -1, it is judged as a non-defective class sample.
7. A software defect prediction device based on multi-model selection, characterized in that, it includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the computer program, it implements the software defect prediction method based on multi-model selection according to any one of claims 1-6 above.
8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and the computer program is used to execute the software defect prediction method based on multi-model selection according to any one of claims 1-6 above.
Citation Information
Cited By
Method and system for classification and / or prediction on unbalanced datasets
US12699931B2