Unsupervised workload perception defect prediction method based on stack generalization hierarchical clustering
Through the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering, the problem of ignoring resource allocation and defect repair costs in traditional defect prediction is solved, and more efficient defect detection and repair is achieved, and software quality and development efficiency are improved.
Patent Information
- Application Number
- CN202510132189.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
Smart Images

Figure CN120066962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect prediction, and particularly to an unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering. Background Art
[0002] With the rapid progress of information technology, various application software has become indispensable in daily life. However, the diversification of software functions is accompanied by the complexity of program development, resulting in a significant increase in the occurrence frequency of software defects. Software defect prediction is one of the most popular research topics in software engineering. By using historical data and machine learning techniques, potential defects in software systems can be identified in advance. Its goal is to predict possible problems in the code, optimize resource allocation, improve development efficiency, and guide developers to focus on high-risk areas for testing and repair. Software defect prediction includes key steps such as data collection and preprocessing, feature extraction and selection, model training, and defect prediction and analysis. Software defect prediction can not only improve the reliability of software products, but also reduce maintenance costs and accelerate the software release cycle. With the progress of artificial intelligence technology, new methods and application scenarios are constantly introduced in software defect prediction research. It not only provides important support in theoretical research, but also shows significant value in practical applications.
[0003] Classification-based defect prediction is a method that uses classification algorithms to classify software modules to predict which modules may contain defects. It usually involves selecting appropriate classification algorithms and using software features for classification. The aim is to improve software quality by identifying potential defect areas in advance. However, traditional classification-based defect prediction research often ignores the issues of resource allocation and defect repair costs. Due to the significant difference in the number of lines of code of different software modules, the workload required for reviewing and repairing defects is also different, and it may not be able to effectively guide the development team to make the best decisions.
[0004] Workload-aware defect prediction is a method that introduces resource and cost considerations into defect prediction. Different from traditional classification-based defect prediction methods, workload-aware defect prediction not only focuses on the presence or absence of defects but also takes into account the workload required to fix these defects (such as time, human resources, etc.). Workload-aware defect prediction analyzes code features, historical defect data, and workload data, and uses machine learning algorithms for model training and optimization to achieve more efficient defect detection and repair. Workload-aware defect prediction preferentially checks defective modules with fewer lines of code, that is, modules with a high defect density will be checked first, and more software defects can be found when the code inspection workload is the same. Workload-aware defect prediction combines software defect prediction with the evaluation of development workload, providing a more efficient resource allocation strategy and a more accurate defect management plan for software development teams. In this invention, a software defect prediction is mainly carried out based on an unsupervised workload-aware method of Stacking hierarchical clustering. The following problems still need to be solved in current research:
[0005] Traditional classification-based defect prediction research often ignores the issues of resource allocation and defect repair costs. Due to the significant differences in the number of lines of code in different software modules, the workload required for reviewing and fixing defects also varies, which may not effectively guide the development team to make the best decisions. In the unsupervised workload-aware defect prediction method based on Stacking generalization hierarchical clustering, workload-aware features are introduced, enabling the model to flexibly adapt to different workloads and environmental changes. This method ensures better performance of the model in terms of diversity and complexity, improving the overall performance and adaptability. However, this method uses different hierarchical clustering methods to process the dataset to determine which methods are more effective in specific situations.
[0006] To solve the above problems, this invention proposes an unsupervised workload-aware defect prediction method based on Stacking generalization hierarchical clustering. Summary of the Invention
[0007] The purpose of this invention is to propose an unsupervised workload-aware defect prediction method based on Stacking generalization hierarchical clustering to solve the problems raised in the background technology:
[0008] Traditional classification-based defect prediction research often ignores the issues of resource allocation and defect repair costs. Due to significant differences in the number of lines of code in different software modules, the amount of work required for reviewing and fixing defects also varies, and it may not be able to effectively guide the development team to make optimal decisions. In the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering, the workload-aware feature is introduced, enabling the model to flexibly adapt to different workloads and environmental changes. This method ensures better performance of the model in terms of diversity and complexity, improving the overall performance and adaptability. However, this method uses different hierarchical clustering methods to process the dataset to determine which methods are more effective in specific situations.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] An unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering, comprising the following steps:
[0011] S1: Extract relevant software project data from the dataset, and isolate these basic components from the project for further analysis and processing;
[0012] S2: Divide the dataset. When constructing the supervised EADP model, a cross-version verification strategy is adopted, using the previous version as the training set and the current version as the test set;
[0013] S3: Use different clustering algorithms at multiple levels to extract and fuse the clustering information of the data layer by layer. The clustering results of each layer are used as features and passed to the next layer. Finally, the original information is comprehensively integrated to generate the final clustering result, and marking is performed based on the clustering result;
[0014] S4: Prediction and sorting. Ascending sorting is performed on the defect clusters and clean clusters respectively according to the feature values of each module. Modules with smaller feature values usually have a higher defect density.
[0015] Preferably, in the S3, in the first layer, initial clustering is performed by using KMedoids, KMeans++, and CURE. In the second layer, the EMA and MeanShift algorithms are used to comprehensively optimize the clustering results of the first layer and the original input features. The randomized majority voting method is used for the integrated results of the second layer.
[0016] Preferably, the steps of the KMedoids clustering method are specifically as follows:
[0017] S3.1.1: Select the initial center points (medoids). First, randomly select K samples as the initial center points (medoids). Denote these initial center points as {m 1 , m 2 , …, mk}。
[0018] S3.1.2: Assign the sample to the nearest medoid. For each sample x in the dataset i , calculate its distance to each center point m k , and assign x i to the cluster represented by the nearest medoid. The distance formula is as follows:
[0019] dist(x i , m k ) = ||x i - m k ||
[0020] The sample x i is assigned to the center point m k where the distance is minimized. The formula is as follows:
[0021]
[0022] where C j represents the j-th cluster; J represents the cluster number to which the sample point x i is assigned, that is, the index of the cluster C j to which the sample point belongs;
[0023] S3.1.3: Update the medoids. For each cluster C j , calculate the total distance from each sample x i within the cluster to each sample, and select the sample with the minimum total distance as the new center point m j . The formula for selecting the new center point is as follows:
[0024]
[0025] S3.1.4: Repeat the iteration. Repeatedly execute the assignment and update steps until the medoids no longer change or reach the predetermined number of iterations.
[0026] S3.1.5: Convergence. When the algorithm converges, each cluster will be represented by a sample that best represents the cluster.
[0027] Preferably, the KMeans++ steps are as follows:
[0028] S3.2.1: Randomly select a data point as the first centroid.
[0029] S3.2.2: Calculate the minimum distance from each data point to the selected centroid.
[0030] S3.2.3: Select the next centroid probabilistically according to the square of the distance. The farther a point is, the greater the probability of being selected. For each data point x, calculate the distance D(x) to the nearest centroid. The probability of selecting the next centroid is proportional to D(x)^2.
[0031] S3.2.4: Repeat S3.2.2 and S3.2.3 until k centroids are selected.
[0032] S3.2.5: Use the selected centroids as the initial points and run the standard KMeans algorithm.
[0033] Preferably, the CURE steps are as follows:
[0034] S3.3.1: Initial sampling. Randomly draw a sample subset from the dataset to reduce the computational complexity while maintaining the representativeness of the dataset.
[0035] S3.3.2: Initial clustering. Perform initial clustering on the drawn sample subset, usually using classical hierarchical clustering methods such as single-link or complete-link clustering, in order to generate initial clusters.
[0036] S3.3.3: Select representative points. Select several representative points from each initial cluster. These representative points are evenly distributed within the cluster to more accurately describe the shape and distribution of the cluster.
[0037] S3.3.4: Shrink the representative points. Move each representative point a certain distance towards the cluster center, usually a fixed proportion of the cluster center. This step helps to reduce the distance between representative points and enhance the tightness of the cluster.
[0038] S3.3.5: Merge clusters. Based on the distance between representative points, gradually merge the clusters until a predetermined stopping condition is met. A common merging strategy is to merge the clusters with the smallest distance between representative points.
[0039] S3.3.6: Extend to the entire dataset. Apply the initial clustering results and representative points to the entire dataset, and assign the unsampled data points to the nearest cluster through the nearest neighbor method.
[0040] Preferably, the MeanShift steps are as follows:
[0041] S3.4.1: Select a bandwidth parameter h, which is the radius of the kernel function. For each data point x, initialize the current mean point m = x.
[0042] S3.4.2: Iterative process. Calculate the mean of all points within the neighborhood of each data point x and update m. Use the kernel function to calculate the weighted average of the points within the neighborhood. Iterate this process until the mean point m converges.
[0043] S3.4.3: Convergence and clustering. When the mean points of all points converge, the similar mean points are grouped into the same cluster center. Finally, each data point is assigned to the corresponding cluster according to the converged mean point.
[0044] Preferably, the EMA step is specifically as follows:
[0045] The objective function is optimized by repeatedly executing two main steps, namely the Expectation Step (E-step) and the Maximization Step (M-step). The goal is to find the maximum likelihood estimate of the model parameters.
[0046] Preferably, S3 is specifically as follows:
[0047] In workload-aware defect prediction, unsupervised clustering techniques are applied to divide software modules into different clusters, and the sum of the feature values of each module is calculated. Subsequently, the risk level of each cluster is evaluated by calculating the average SFM of each cluster.
[0048] Preferably, S4 is specifically as follows:
[0049] S4.1: Sort the defect clusters and clean clusters in ascending order according to the feature values of each module, because the modules with smaller feature values usually have higher defect densities.
[0050] S4.2: According to the principle of preferentially checking defective modules, arrange the defective modules in the front and the clean modules in the back to form the final module priority ranking list.
[0051] S4.3: Evaluate the top 20% of the modules in terms of LOC according to the order of the software modules.
[0052] Compared with the prior art, the present invention provides an unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering, having the following beneficial effects:
[0053] The present invention proposes an unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering. In the first layer, KMedoids, KMeans++, and CURE are used for preliminary clustering to fully capture the diversity and complex structure of the data and more comprehensively reveal different features in the data. In the second layer, EMA and MeanShift are used to further optimize the clustering results of the first layer and the original input features. By combining the original features and the clustering features, more information is retained and the internal structure of the data is captured more accurately. The randomized majority voting method is used for prediction on the clustering results. The prediction category with the highest frequency is selected as the final result, and if there is a tie, a category is randomly selected. The present invention selects data from 29 different versions of 11 open-source projects in the PROMISE dataset for experiments. The research results show that the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering is superior to most common unsupervised and some supervised methods in multiple key metrics. Notably, it performs excellently in reducing the IFA value and the number of false alarms. It significantly improves the reliability of the model and the work efficiency of the test team. Description of the Drawings
[0054] Figure 1 It is the overall flowchart of the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering mentioned in Embodiment 1 of the present invention;
[0055] Figure 2 It is the process framework diagram of the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering mentioned in Embodiment 1 of the present invention. Detailed Embodiments
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0057] The present invention proposes an unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering. In the first layer, KMedoids, KMeans++, and CURE are used for preliminary clustering to fully capture the diversity and complex structure of the data, and more comprehensively reveal different features in the data. In the second layer, EMA and MeanShift are used to further optimize the clustering results of the first layer and the original input features. By combining the original features and the clustering features, more information is retained, and the internal structure of the data is captured more accurately. The randomized majority voting method is used for prediction on the clustering results. The prediction category with the highest frequency is selected as the final result, and if there is a tie, a category is randomly selected. This provides a useful guidance for future work and emphasizes the potential advantages of the hierarchical clustering method in unsupervised workload-aware defect prediction. The following will illustrate the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering proposed by the present invention with specific examples, which specifically includes the following content.
[0058] Example 1:
[0059] Please refer to Figure 1-2 , the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering of the present invention includes:
[0060] S1: Extract relevant software project data from the dataset, and isolate these basic components from the project for further analysis and processing; specifically as follows:
[0061] The dataset includes the historical records of multiple software projects, covering various information such as source code, defect reports, and metrics. Specifically, in each project, the source code is usually organized into multiple modules (such as classes or files). These modules represent the basic components of the code. The process of extracting modules involves isolating these basic components from the project for further analysis and processing. This process helps to identify and extract the key features of each module, thus supporting subsequent defect prediction and software quality assessment.
[0062] S2: Divide the dataset, and adopt a cross-version validation strategy when constructing the supervised EADP model. The previous version is used as the training set, while the current version is used as the test set; specifically as follows:
[0063] To compare with the supervised EADP method, a cross-version validation strategy is adopted in constructing the supervised EADP model in this paper. Cross-version validation aims to evaluate the generalization ability of the model in software systems of different versions. Specifically, by validating in multiple versions, the performance of the model in different versions is evaluated. If the model performs excellently in each version, it indicates that it has strong robustness and reliability, can adapt to version changes, and has high practical value in actual applications. Cross-version validation not only evaluates the accuracy of the model, but also ensures its continuous effectiveness in evolving software systems, thus becoming a trustworthy defect prediction tool. During the cross-version validation process, the previous version is usually used as the training set, while the current version is used as the test set.
[0064] S3: Use different clustering algorithms at multiple levels to extract and fuse the clustering information of data layer by layer. Take the clustering results of each layer as features and pass them to the next layer. Finally, integrate the original information to generate the final clustering result and perform labeling based on the clustering result; specifically, the following two processing steps:
[0065] The first step is clustering. Specifically, in the first layer, initial clustering is performed using KMedoids, KMeans++, and CURE to effectively capture the diversity and complex structure of the data. Using these algorithms, the clustering in the first layer can more comprehensively reveal different features in the data. In the second layer, the EMA and MeanShift algorithms are used to comprehensively optimize the clustering results of the first layer and the original input features. This process combines the original features and the clustering features, thus retaining more valuable information. The randomized majority voting method is adopted for the integrated results of the second layer. Specifically, each model makes predictions for each sample, and then statistical voting is performed on these prediction results to select the prediction category with the highest frequency as the final result. If there is a voting tie, the final prediction is determined by randomly selecting one of the categories. This method effectively improves the robustness and reliability of the prediction based on the fusion of the prediction results of multiple models. Even in the case of a tie, it can maintain high prediction performance and stability.
[0066] The second step is labeling. Specifically, in workload-aware defect prediction, unsupervised clustering technology is applied to divide software modules into different clusters, and the sum of the feature values of each module is calculated. Subsequently, the risk level of each cluster is evaluated by calculating the ASFM of each cluster. For binary clustering, the cluster with a higher ASFM is labeled as a defective module, and the cluster with a lower ASFM is labeled as a clean module; for multi-clustering, the cluster with an ASFM higher than the average ASFM of all clusters is labeled as a defective module, and the remaining clusters are labeled as clean modules.
[0067] S4: Prediction and Sorting: Ascendingly sort the defect clusters and clean clusters respectively according to the eigenvalue of each module. Modules with smaller eigenvalues usually have higher defect densities. Specifically as follows:
[0068] Ascendingly sort the defect clusters and clean clusters respectively according to the eigenvalue of each module. Since modules with smaller eigenvalues usually have higher defect densities. The eigenvalue reflects indicators such as the complexity and modification frequency of the code. A smaller eigenvalue means that these modules may have undergone fewer modifications, lower complexity, or fewer lines of code. In fact, these modules often concentrate more defects. Specifically, following the principle of preferentially checking defective modules, arrange the defective modules in the front and the clean modules in the back to form the final module priority ranking list. Finally, evaluate the top 20% of the modules in terms of LOC according to the order of the software modules. This method not only enhances the accuracy of defect prediction but also has significant advantages in optimizing resource allocation and defect repair priority management, thus providing an efficient and practical solution for software quality assurance.
[0069] Comparative Example:
[0070] This comparative example uses experiments to verify that the proposed method has better performance in unsupervised workload-aware defect prediction. Specifically, use common clustering methods for experiments and then compare the experimental results of the above models. Here, different clustering methods (KMedoids, KMeans++, CURE, EMA, MeanShift, ManualUp, CBS+, SERS) and the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering proposed by the present invention are mainly selected for comparison.
[0071] The present invention uses the PROMISE dataset. This dataset is publicly available and has been widely used in defect prediction research.
[0072] The main goal of EADP is to predict the defect density of software modules, that is, the ratio between the number of defects and the number of lines of code, and sort the modules according to this density. When evaluating the performance of the model, the PofB@20% metric is very crucial and accurate defect number information is required to calculate the coverage rate of actual defects in the top 20% of high-risk modules. To meet the requirement of the EADP task for the number of defects, the present invention selects the PROMISE dataset containing detailed defect records and code metrics as the basic data source for the research. This selection is based on the following reasons.
[0073] Detailed defect records. The PROMISE dataset contains rich and detailed defect records, which can provide the accurate defect quantity information required by the model, ensuring the accuracy of defect density calculation. This repository collects data from various actual projects, including source code metrics, defect information, project management data, etc.
[0074] Comprehensive code metrics. This dataset not only records the lines of code but also includes other important code metric indicators such as code complexity and modification frequency, which are crucial for constructing and evaluating the EADP model.
[0075] Widely used standard dataset. The PROMISE dataset is a widely used standard dataset in the field of software defect prediction, and its reliability and representativeness have been recognized by the industry and academia, which helps to ensure the credibility and comparability of research results. These datasets are provided in a standardized format and are accompanied by detailed metadata descriptions to assist researchers in processing and analysis.
[0076] Support for model development and evaluation. The detailed information and comprehensive metrics provided by the PROMISE dataset support the development and evaluation of the EADP method, ensuring that the model can accurately predict and rank the defect risks of software modules, thereby improving the effectiveness and practical value of the model in actual applications. Table 1 shows the specific information about the PROMISE dataset.
[0077] Table 1: Specific information of the PROMISE dataset
[0078]
[0079] Among them, "#Module" refers to the number of modules. "AvgDef" represents the average number of defects per module. "%Def" is the percentage of defective modules. "#Def" is the number of defects. The detailed defect records and comprehensive code metrics of the PROMISE dataset provide a solid data foundation for the EADP model, ensuring its high accuracy and effectiveness in predicting the defect risks of software modules.
[0080] The present invention uses indicators such as Precision@20%, Recall@20%, PofB@20%, PMI@20%, and IFA to evaluate the effectiveness of software defect prediction methods.
[0081] Precision@20% is used to evaluate the accuracy of the model when identifying the top 20% high-risk code modules. It is defined as the ratio of the number of actual defective modules to the number of predicted defective modules among the top 20% of the lines of code. A lower Precision@20% indicates that the model has poor accuracy in identifying high-risk modules, which may lead to a decrease in the confidence of the test team in the model's prediction results and thus waste valuable testing resources. Precision@20% focuses on the top 20% high-risk modules and provides an intuitive way to evaluate the accuracy of the model in identifying high-risk modules.
[0082] Recall@20% is used to measure the proportion of actual defective modules identified by the model in the top 20% of the LOCs to the total defective modules in the dataset. It is defined as the ratio of the number of defective modules actually identified in the top 20% of the lines of code to the total number of defective modules among the total defective modules. A lower Recall@20% indicates that the model has a weak ability to identify defects in high-risk areas, which may result in many actual defects not being discovered in a timely manner, thus affecting software quality and maintenance efficiency.
[0083] PofB@20% is a user-related effort perception metric for a user who operates based on the entity defect probability ranking provided by the classifier. This metric is used to evaluate the model's ability to cover defects in identifying the top 20% high-risk code modules. Specifically, it is defined as the proportion of the number of actual defective modules in the top 20% of the lines of code to the total number of defective modules. When each defective module contains only one defect, this metric is equal to Recall@20%. A higher PofB@20% indicates that the model has a stronger ability to detect more defects. On the contrary, a lower PofB@20% indicates that the model has a poor ability to identify defects in high-risk areas, resulting in many actual defects not being discovered, thus affecting software quality and maintenance efficiency.
[0084] PMI@20% is defined as the ratio of the number of modules predicted as defective modules in the top 20% of the lines of code to the total number of modules. A lower PMI@20% indicates that the model identifies fewer modules in high-risk areas. On the contrary, a higher PMI@20% means that the software testing team needs to check more modules in the same number of LOCs, thus increasing the actual workload and time cost and resulting in waste of resources when switching between different modules. By introducing clustering techniques with superior performance, the number of modules that the test team needs to check can be significantly reduced, thus greatly reducing the resource consumption when switching between modules.
[0085] The Initial False Alarm (IFA) is used to measure the false positive situation of the model. It is specifically defined as the number of false alarms encountered by the test team before finding the first truly defective module. A higher IFA value indicates that the model generates more false positives before actually detecting defects, which may cause the test team to waste valuable time and resources and weaken their trust in the model's prediction results. The level of the IFA value directly reflects the accuracy and effectiveness of the model. To improve the actual application effect of the model, it is crucial to reduce the IFA value.
[0086] Table 2, Table 3, Table 4, Table 5 and Table 6 show the experimental results of the present invention.
[0087] Table 2: Experimental results of this method on all metrics
[0088] Cross-version Precision@20% Recall@20% PofB@20% PMI@20% IFA Ant1.4-1.5 0.065 0.226 0.226 0.212 3 Ant1.5-1.6 0.258 0.211 0.211 0.162 0 Ant1.6-1.7 0.201 0.218 0.218 0.192 8 Camel1.2-1.4 0.120 0.248 0.248 0.345 1 Camel1.4-1.6 0.202 0.261 0.261 0.316 2 Ivy1.4-2.0 0.089 0.300 0.300 0.384 0 Jedit4.0-4.1 0.181 0.367 0.367 0.356 9 Jedit4.1-4.2 0.077 0.229 0.229 0.275 14 Log4j1.1-1.2 0.944 0.279 0.279 0.263 1 Lucene2.2-2.4 0.525 0.307 0.307 0.347 0 Poi2.0-2.5 0.549 0.290 0.290 0.200 1 Poi2.5-3.0 0.603 0.256 0.270 0.362 0 Synapse1.1-1.2 0.419 0.268 0.268 0.263 2 Velocity1.5-1.6 0.352 0.410 0.372 0.336 1 Xalan2.5-2.6 0.469 0.227 0.227 0.219 2 Xalan2.6-2.7 0.995 0.234 0.222 0.220 0 Xerces1.2-1.3 0.110 0.288 0.288 0.240 13 Xerces1.3-1.4 0.770 0.278 0.278 0.214 0 Average 0.385 0.272 0.270 0.273 3.167
[0089] Table 3: Precision@20% values of 9 methods in each cross-version experiment when checking the first 20% LOC
[0090] Cross-version UEADPSHC KMedoids KMeans++ CURE EMA MeanShift ManualUp CBS+ SERS Ant1.4-1.5 0.065 0.077 0.095 0.091 0.065 0.099 0.051 0.102 0.058 Ant1.5-1.6 0.258 0.271 0.224 0.224 0.227 0.203 0.104 0.423 0.305 Ant1.6-1.7 0.201 0.172 0.173 0.184 0.184 0.167 0.095 0.261 0.229 Camel1.2-1.4 0.120 0.167 0.120 0.120 0.120 0.108 0.119 0.120 0.213 Camel1.4-1.6 0.202 0.175 0.125 0.125 0.125 0.223 0.160 0.182 0.193 Ivy1.4-2.0 0.089 0.148 0.089 0.089 0.072 0.082 0.039 0.023 0.080 Jedit4.0-4.1 0.181 0.131 0.072 0.072 0.072 0.164 0.131 0.213 0.248 Jedit4.1-4.2 0.077 0.145 0.039 0.046 0.039 0.059 0.051 0.164 0.149 Log4j1.1-1.2 0.944 0.879 0.929 0.908 0.901 0.900 0.935 0.906 0.860 Lucene2.2-2.4 0.525 0.449 0.504 0.473 0.468 0.464 0.525 0.572 0.486 Poi2.0-2.5 0.549 0.591 0.522 0.450 0.461 0.600 0.573 0.750 0.677 Poi2.5-3.0 0.603 0.551 0.493 0.463 0.450 0.574 0.493 0.615 0.627 Synapse1.1-1.2 0.419 0.255 0.260 0.260 0.254 0.270 0.196 0.353 0.196 Velocity1.5-1.6 0.352 0.326 0.325 0.319 0.291 0.211 0.271 0.353 0.381 Xalan2.5-2.6 0.469 0.447 0.477 0.388 0.444 0.387 0.310 0.459 0.383 Xalan2.6-2.7 0.995 0.989 0.977 0.978 0.977 0.964 0.983 0.981 0.988 Xerces1.2-1.3 0.110 0.115 0.110 0.092 0.077 0.045 0.102 0.111 0.089 Xerces1.3-1.4 0.770 0.702 0.703 0.770 0.771 0.772 0.702 0.843 0.747 Average 0.385 0.366 0.347 0.336 0.333 0.350 0.324 0.413 0.384
[0091] Table 4: PofB@20% values of 9 methods in each cross-version experiment when checking the first 20% LOC
[0092]
[0093]
[0094] Table 5: PMI@20% values of 9 methods in each cross-version experiment when checking the first 20% LOC
[0095]
[0096]
[0097] Table 6: Average Recall@20% and IFA values of all methods
[0098] Cross-version UEADPSHC KMedoids KMeans++ CURE EMA MeanShift ManualUp CBS+ SERS Recall@20% 0.272 0.304 0.244 0.258 0.225 0.212 0.481 0.267 0.280 IFA 3.167 5.833 4.167 5.445 5.056 5.111 18.556 5.111 7.833
[0099] The evaluation on the Precision@20% metric shows that the unsupervised workload-aware defect prediction method based on stacked generalization hierarchical clustering outperforms its five basic methods, with an improvement range between 5.19% and 15.62%. Compared with the ManualUp method, the improvement of the UEADPSHC method reaches 18.83%. Although it is not dominant in the comparison with the CBS+ method, the UEADPSHC still shows a certain improvement compared with the SERS method. These results highlight the significant advantage of UEADPSHC in improving Precision@20%, demonstrating its ability to effectively perceive and predict workload-related defects without supervision, and further proving the efficiency and reliability of UEADPSHC in practical applications.
[0100] UEADPSHC improves by 5.43% to 28.30% compared with the basic methods in terms of Recall@20%. It decreases by 11.76% compared with the KMedoids method. In the comparison with the two supervised methods, UEADPSHC improves by 1.87% compared with the CBS+ method and is slightly inferior to the SERS method. This indicates that although UEADPSHC can effectively identify and detect a relatively high proportion of defects in most cases, it may not fully capture all potential defects in some cases. Generally speaking, UEADPSHC shows significant advantages in unsupervised workload-aware defect prediction, further verifying its effectiveness and superiority in practical applications.
[0101] In the evaluation of the PofB@20% metric, UEADPSHC performs better than most of the comparison methods. It improves by 4.65% to 25% compared with the basic methods. It decreases by 9.63% compared with the KMedoids method. However, in the comparison with the two supervised methods, UEADPSHC improves by 2.66% compared with the CBS+ method. UEADPSHC still shows its significant advantage in improving PofB@20%. This shows that UEADPSHC can not only provide a relatively high detection coverage rate in most cases, but also show good performance in unsupervised workload-aware defect prediction, thus optimizing the defect detection process and improving the overall efficiency and resource utilization rate.
[0102] A lower PMI@20% indicates that the model has better performance in identifying fewer modules in high-risk areas. The PMI@20% of UEADPSHC decreases by up to 21.33% compared with several basic methods, and significantly decreases by 60.89% compared with the ManualUp method. In the comparison with the two supervised methods, UEADPSHC decreases by 5.54% compared with the CBS+ method. These results highlight the excellent performance of UEADPSHC in reducing the number of identified modules in high-risk areas.
[0103] Almost all respondents said that checking more than 10 actual cleaning modules was beyond their acceptable range. UEADPSHC reduced IFA@20% by 24% to 46.56% compared to the basic method, and significantly reduced it by 82.93% compared to the ManualUp method. Compared with other clustering techniques, UEADPSHC performed well in significantly reducing IFA values, highlighting its excellent ability in optimizing feature sequences. It further verified its significant advantages in improving overall performance and efficiency, reflecting its high efficiency and reliability in practical applications.
[0104] According to the analysis of the above indicators, UEADPSHC achieves the best overall effort-aware performance by sorting the modules in ascending order based on the LOC feature. Its advantage lies in the introduction of hierarchical clustering technology in unsupervised effort-aware defect prediction. This method effectively retains more information, thereby improving the clustering quality and the adaptability of the model. By more accurately capturing the intrinsic structure of the data, the model significantly improves the performance of various indicators. Specifically, the improvement of UEADPSHC in various indicators fully demonstrates its efficiency and reliability in defect prediction. The model can more accurately assist the test team in defect detection and repair, thereby optimizing the software quality management and maintenance process. UEADPSHC significantly improves the utilization efficiency of test resources, ensuring that high-risk modules are identified and processed in the shortest time, thereby improving the efficiency and quality of overall software maintenance.
[0105] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. An unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering, characterized in that: The steps include: S1: Extract software project data from the data set, organize the source code of each software project into multiple modules, use the modules to represent the basic components of the code, and separate the basic components from the software project data for further analysis and processing; S2: Divide the dataset and build a supervised EADP model using a cross-version validation strategy, using the previous version as the training set and the current version as the test set; S3: Use different clustering algorithms at multiple levels to extract and fuse the clustering information of data layer by layer; The clustering results of each layer are passed as features to the next layer, and the final clustering results are generated by integrating the original information and marked based on the clustering results; S4: Prediction and sorting, the defective clusters and clean clusters are sorted in ascending order according to the eigenvalues of each module. Modules with small eigenvalues have high defect density.
2. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 1 is characterized in that: The S3 specifically includes the following contents: In workload-aware defect prediction, unsupervised clustering technology is applied to divide software modules into different clusters, and the sum of the eigenvalues of each module is calculated; then, the risk level of each cluster is evaluated by calculating the average SFM of each cluster.
3. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 2 is characterized in that: In S3, the first layer uses KMedoids, KMeans++ and CURE for preliminary clustering, the second layer uses MeanShift and EMA algorithms, the clustering results of the first layer and the original input features are comprehensively optimized, and the integrated results of the second layer are comprehensively optimized using randomized majority voting method.
4. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 3 is characterized in that: The preliminary clustering using KMedoids specifically includes the following: S3.1.1: Select initial center point; randomly select K samples as initial center points, and record the center points as {m1,m2,…,m k }; S3.1.2: Assign samples to the nearest center point; for each sample x in the dataset i , calculate its point m with each center k distance, and x i Assigned to the cluster represented by the nearest center point; each sample x i With each center point m k The calculation formula of the distance is as follows: dist(x i ,m k )=||x i -m k || The sample x i Assign to the center point m that minimizes the distance k The specific formula is as follows: Among them, C j represents the jth cluster; J represents the sample point x i The cluster number assigned, that is, the cluster C to which the sample point belongs j The index of S3.1.3: Update the center point; for each cluster C j , calculate each sample x in the cluster i The total distance to each sample, select the sample with the smallest total distance as the new center point m j ; The formula for selecting the new center point is as follows: S3.1.4: Repeat the iterations; repeatedly perform the allocation and update steps until the center point no longer changes or the predetermined number of iterations is reached; S3.1.5: Convergence: When the algorithm converges, each cluster is represented by a sample that best represents the cluster.
5. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 4 is characterized in that: The preliminary clustering using KMeans++ specifically includes the following: S3.2.1: Randomly select a data point as the first centroid; S3.2.2: Calculate the minimum distance from each data point to the selected centroid; S3.2.3: Select the next centroid in a probabilistic manner based on the square of the distance; the farther the point is, the greater the probability of being selected; for each data point x, calculate its distance D(x) to the nearest centroid, and the probability of selecting the next centroid is proportional to D(x)2; S3.2.4: Repeat S3.2.2 and S3.2.3 until k centroids are selected; S3.2.5: Using the selected centroids as the initial points, run the standard KMeans algorithm.
6. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 5, characterized in that: The preliminary clustering using CURE specifically includes the following: S3.3.1: Initial sampling: Randomly extract a subset of samples from the dataset to reduce computational complexity while maintaining the representativeness of the dataset; S3.3.2: Initial clustering: Perform initial clustering on the extracted sample subset to generate initial clusters; S3.3.3: Select representative points: Select a number of representative points from each initial cluster, wherein the representative points are evenly distributed within the cluster to describe the shape and distribution of the cluster; S3.3.4: Shrink representative points; move each representative point toward the center of the cluster, reduce the distance between representative points, and enhance the compactness of the cluster; S3.3.5: Merge clusters; based on the distance between representative points, gradually merge clusters until a predetermined stopping condition is met; S3.3.6: Expand to the entire dataset; apply the initial clustering results and representative points to the entire dataset, and assign unsampled data points to the nearest cluster using the nearest neighbor method.
7. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 6 is characterized in that: The MeanShift algorithm specifically includes the following contents: S3.4.1: Select a bandwidth parameter h, i.e. the radius of the kernel function; for each data point x, initialize it to the current mean point m = x; S3.4.2: Iterative process; for each data point x, calculate the mean of all points in its domain and update m; use the kernel function to calculate the weighted average of the points in the domain; iterate this process until the mean point m converges; S3.4.3: Convergence and clustering; when the mean points of all points have converged, the similar mean points are grouped into the same cluster center; finally, each data point is assigned to the corresponding cluster according to its converged mean point.
8. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 7, characterized in that: The EMA algorithm specifically includes the following contents: The objective function is optimized by repeatedly performing the expectation step and the maximization step, and the optimization goal is to find the maximum likelihood estimate of the model parameters.
9. The unsupervised workload-aware defect prediction method based on stacked generalized hierarchical clustering according to claim 1, characterized in that: The S4 specifically includes the following contents: S4.1: Sort the defective clusters and clean clusters in ascending order according to the eigenvalues of each module; S4.2: According to the principle of checking defective modules first, the defective modules are arranged in front and the clean modules are arranged in the back, so as to form the final module priority sorting list; S4.3: Evaluate the top 20% of the LOC modules according to the order of the software modules.