Cross-project software defect prediction method based on manifold combination features and joint distribution
By combining global and local transferable features in manifold feature space and combining iteratively learning pseudo-label strategies, the problem of data distribution differences and feature distortion in cross-project software defect prediction is solved, significantly improving prediction performance.
Patent Information
- Application Number
- CN202111564014.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-20
AI Technical Summary
Existing cross-project software defect prediction methods have problems with feature distortion and data distribution differences when processing high-dimensional nonlinear data, resulting in poor prediction performance, especially low pseudo-label accuracy.
The joint distribution matching method based on manifold combination features is adopted, and the data distribution difference is reduced through the linear combination of global transferable features and local transferable features, and a strategy of iteratively learning pseudo-labels is introduced to improve its accuracy by updating pseudo-labels multiple times.
It effectively reduces data distribution differences, solves feature distortion problems, and significantly improves pseudo-label accuracy, thereby improving the performance of cross-project software defect prediction.
Smart Images

Figure CN114253849B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software security and machine learning, and specifically relates to a cross-project software defect prediction method based on manifold combination features and joint distribution matching. Background Art
[0002] In today's society, software has become a very important part of production and life. Potential and difficult-to-find defects in software will affect the quality of the software and even cause huge losses to production and life. Therefore, it is necessary to predict defects before the software is released and minimize the losses. At the same time, it can also give software developers more opinions to improve software quality.
[0003] In order to predict defects in software, some researchers have proposed to train a classifier in the historical data of the same software project that has been released and apply the classifier to new software projects for defect prediction. This is intra-project defect prediction. In actual production life, intra-project software defect prediction still faces many challenges and difficulties. For example, some projects are newly launched and lack sufficient historical data. In addition, the cost of collecting data is also very expensive. Therefore, it is difficult to achieve good prediction results for intra-project defect prediction.
[0004] In order to solve the above difficulties, cross-project software defect prediction came into being. Cross-project defect prediction can greatly reduce the cost of data collection and solve the problem of insufficient data. Cross-project defect prediction uses the labeled data in the data of external projects (source projects) to build a prediction model and apply it to the new project (target project) for prediction. Although cross-project defect prediction solves the problem of missing data, it still faces challenges. Since the source project and the target project come from different fields, there are large differences in their data distribution, which will have a negative impact on the final prediction performance. In order to reduce the difference in data distribution, some studies have proposed to use the joint distribution matching method to reduce the data distribution difference, that is, to jointly adjust the marginal distribution difference and the conditional distribution difference in the feature space, and use the kernel mean matching (KMM) method to construct a source project instance weight vector, thereby reducing the data distribution difference. However, this method has two defects. First, because the original data is high-dimensional and nonlinear, it will cause feature distortion, which is not conducive to reducing the data distribution difference; second, this method fails to improve the accuracy of pseudo-labels, which has a negative impact on the final prediction performance.
[0005] In response to data distribution differences and feature distortions, the present invention considers using joint distribution matching based on manifold combination features to solve feature distortions and reduce data distribution differences; in response to the low accuracy of pseudo-labels, the present invention considers improving the accuracy of pseudo-labels by iteratively learning pseudo-labels, thereby improving the performance of cross-project software defect prediction. Summary of the invention
[0006] In view of the deficiencies of the prior art, the present invention proposes a cross-project software defect prediction method based on manifold combination features and joint distribution matching.
[0007] The present invention proposes a cross-project software defect prediction method based on manifold combination features and joint distribution matching, which overcomes the insufficiency of reducing data distribution differences in the original regenerated Hilbert space (RHKS) and the problem of feature distortion. The method chooses to consider using a global marginalized denoising autoencoder (DA-GMDA) with domain adaptation to extract global transferable features and a local subset marginalized denoising autoencoder (DA-LMDA) with domain adaptation in the manifold feature space to extract local transferable features, and then linearly combines the extracted global transferable features and local transferable features into new combined features and applies them to achieve joint distribution matching; secondly, in order to solve the problem of low pseudo-label accuracy, the method introduces an iterative learning pseudo-label strategy to improve the pseudo-label accuracy by updating the pseudo-label multiple times in a loop. The strategy achieves joint distribution matching by using new combined features, and then obtains an instance weight training model through joint distribution matching and updates the label again, and makes the updated label undergo a new round of combined feature extraction, pseudo-label update and joint distribution matching until the final prediction result converges.
[0008] The method of the present invention specifically comprises the following steps:
[0009] Step 1: Obtain defective data sets and non-defective data sets, and select some data for experimental verification;
[0010] Step 2: extracting global transferable features and local transferable features of the data in the manifold feature space;
[0011] Step 3: Combine the combined features in a linear manner and combine the combined features with the joint distribution matching. The specific operations are as follows:
[0012] That is, using DA-GMDA and DA-LMDA in the manifold feature space to generate globally transferable features from the source and target items and and the local transferable features of the source project and the local transferable features of the target project and Then and as well as and Linear combination into new combination features and After that, the pseudo-label is updated for the first time, and the Gram matrix is calculated using the combined features Joint distribution matching via Gram matrix;
[0013] Step 4: Obtain instance weights through joint distribution matching; train a prediction model based on the weights, and finally perform a secondary update of pseudo labels based on the prediction model;
[0014] Step 5: During the repeated training of the prediction model, the updated labels are used again to extract the combined features and update the pseudo labels and joint distribution matching until the final prediction converges stably, and the training is terminated.
[0015] Preferably, the global transferable features and local transferable features of the data are extracted in the manifold feature space; the specific operation is as follows:
[0016] Step 2-1: Extract transferable features in the manifold feature space; first, extract the transferable features in the Grassmann manifold G(d k ) to learn a mapping function g(.), where d k is the dimension of the data subspace of different items; the geodesic flow kernel GFK is used to learn the computational efficiency of g(.), and G(·) is regarded as a d k The set of subspaces of dimension;
[0017] For any two data features x i 、x j Constructing geodesic flow is equivalent to converting the original features into infinite-dimensional feature space; the new feature data is recorded as z = g(x), and the inner product of the new features produces a semi-positive GFK:
[0018]
[0019] Where G is a semi-positive definite matrix; the original data x is transformed into Grassmann manifold features, Project instance set X = {X S ,X T}, where X S is a collection of source project instances, X T is the set of target project instances. G is just an expression and cannot be directly calculated, so the square root is calculated using the Denman-Beavers algorithm; it is used in the following sections As a manifold feature representation;
[0020] Step 2-2: DA-GMDA Exploits X S and X T To extract global features with more transferable capacity; assuming the marginal distribution M s (x S )≠M t (x t ) and the conditional distribution C s (y S |x S)≠C t (y t |x t ), and the marginal distribution and conditional distribution are matched simultaneously; DA-GMDA aims to learn the mapping matrix W to reconstruct the transferable global feature space. The input data is corrupted by noise probability p. The objective function is defined as:
[0021]
[0022] in and Match the marginal distribution and match the conditional distribution respectively, Is the source project data X S A corrupted version of Is the target project data X T The corrupted version of ; λ is the regularization parameter. The hyperbolic tangent function is used to generate the global transferable features of the source project and target project global transferable features
[0023] Step 2-3: Use DA-LMDA to extract features of local subsets;
[0024] The source item instances and target item instances are divided into different local subsets according to the labels and pseudo labels. Secondly, DA-LMDA is performed to match the local subset distribution and reconstruct the local subspace to obtain richer transferable local features of different subsets; the objective function of DA-LMDA on the local subset with label c is defined as:
[0025]
[0026] in, represents an instance with label c from the source project, represents an instance with label c from the target project. yes A corrupted version of yes is a corrupted version of , and λ is the regularization parameter.
[0027] Get the local transferable features of the source item on label c and the target project's local transferable features
[0028] Beneficial results of the present invention:
[0029] 1. Different from the existing CPDP methods, which only consider reducing the distribution difference in the original feature space and do not consider the feature distortion problem, the present invention considers the combined features and the joint distribution matching based on the combined features in the manifold embedding feature space, which more effectively reduces the data distribution difference and solves the feature distortion problem caused by nonlinearity and high-dimensional data in the original data.
[0030] 2. Unlike the existing pseudo-label-based CPDP method which initializes the pseudo-label once before training, the method proposed in the present invention utilizes an iterative learning pseudo-label strategy to achieve joint distribution matching through new combined features, and updates the label multiple times in each round of learning to improve the accuracy of the pseudo-label. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 The overall framework diagram of the cross-project software defect prediction method based on popular combination features and joint distribution matching.
[0032] Figure 2 Specific features of the dataset. DETAILED DESCRIPTION
[0033] This paper proposes a novel cross-project software defect prediction method based on manifold combination features and joint distribution matching. First, global transferable features and local transferable features are combined into new combination features, and the new combination features are applied to joint distribution matching to reduce data distribution differences and solve feature distortion. Secondly, in order to solve the problem of low pseudo-label accuracy, this method introduces a strategy of iterative learning pseudo-labels.
[0034] The present invention will be described in detail below in conjunction with the accompanying drawings and PROMISE dataset. Figure 1 As shown, the specific steps are as follows:
[0035] (1) Pseudo-label initialization: Pseudo-label the unlabeled data in the dataset, specifically using a trained prediction model to label them.
[0036] (2) Iteratively learn pseudo labels. The accuracy of the initialized pseudo labels is low, which will have a negative impact on the experimental results, so this step will improve the accuracy of the pseudo labels.
[0037] (2-1) Combination features
[0038] First, on the Grassmann manifold G(d k ) to learn a mapping function g(.), where d k is the dimension of the subspace of different project data. Here we use the geodesic flow kernel (GFK) to learn the computational efficiency of g(.). G is regarded as a d kdimensional subspace set, using As manifold features, DA-GMDA is used to extract global transferable features with more transferable capacity, and DA-LMDA is used to extract local transferable features. Then, the global features and local features are used to form combined features in a linear manner.
[0039] (2-2) Joint Distribution Matching
[0040] Joint distribution matching is achieved based on combined features. Joint distribution means that when reducing data distribution differences, edge distribution differences and conditional distribution differences are considered simultaneously. In addition, the present invention takes into account the different degrees of influence of the two on reducing distribution differences, so different weights c1 and c2 are given to the two respectively. The constrained optimization problem is shown in formula (4):
[0041]
[0042] (2-3) Normalize the data instance weight α
[0043] (2-4) The prediction model is iteratively trained using instance weights, source item data labels, and pseudo-label data, and the model is used to predict target item data.
[0044] (3) When the model prediction results tend to be stable, the prediction results are output.
[0045] As shown in Table 1, Figure 2 As shown in the figure, the F value and G-mean value of the cross-project defect prediction based on the present invention on the PROMISE dataset are 0.539 and 0.670. Compared with other classic methods, the present invention improves the F value by 23.3% and the G-mean value by 6.2%. The effects are better than these classic methods.
[0046] Table 1 Cross-project defect prediction results
[0047]
Claims
1. A cross-project software defect prediction method based on manifold combination features and joint distribution, It is characterized in that The method specifically comprises the following steps: Step 1: Obtain defective data sets and non-defective data sets, and select some data for experimental verification; Step 2: Extract the global transferable features and local transferable features of the data in the manifold feature space; the specific operations are as follows: Step 2-1: Extract transferable features in the manifold feature space; first, extract the transferable features in the Grassmann manifold G(d k ) to learn a mapping function g(.), where d k is the dimension of the data subspace of different items; the geodesic flow kernel GFK is used to learn the computational efficiency of g(.), and G(·) is regarded as a d k The set of subspaces of dimension; For any two data features x i 、x j Constructing geodesic flow is equivalent to converting the original features into infinite-dimensional feature space; the new feature data is recorded as z = g(x), and the inner product of the new features produces a semi-positive GFK: Where G is a semi-positive definite matrix; the original data x is transformed into Grassmann manifold features, Project instance set X = {X S ,X T }, where X S is a collection of source project instances, X T is the set of target project instances; G is just an expression and cannot be calculated directly, so the square root is calculated using the Denman-Beavers algorithm; it is used in the following sections As a manifold feature representation; Step 2-2: DA-GMDA Exploits X S and X T To extract global features with more transferable capacity; assuming the marginal distribution M s (x S )≠M t (x t ) and the conditional distribution C s (y S |x S )≠C t (y t |x t ), and the marginal distribution and conditional distribution are matched simultaneously; DA-GMDA aims to learn the mapping matrix W to reconstruct the transferable global feature space; the input data is corrupted by the noise probability p; the objective function is defined as: in and Match the marginal distribution and match the conditional distribution respectively, Is the source project data X S A corrupted version of Is the target project data X T is a corrupted version of ; λ is the regularization parameter; the hyperbolic tangent function is used to generate the global transferable features of the source project and target project global transferable features Step 2-3: Use DA-LMDA to extract features of local subsets; The source item instances and the target item instances are divided into different local subsets according to the labels and pseudo-labels. Secondly, DA-LMDA is performed to match the local subset distribution and reconstruct the local subspace to obtain richer transferable local features of different subsets. The objective function of DA-LMDA on the local subset with label c is defined as: in, represents an instance with label c from the source project, represents an instance with label c from the target project; yes A corrupted version of yes is a corrupted version of , where λ is the regularization parameter; local transferable features of the source item on label c are obtained and the target project's local transferable features Step 3: Combine the combined features in a linear manner and combine the combined features with the joint distribution matching. The specific operations are as follows: That is, using DA-GMDA and DA-LMDA in the manifold feature space to generate globally transferable features from the source and target items and and the local transferable features of the source project and the local transferable features of the target project and Then and as well as and Linear combination into new combination features and After that, the pseudo-label is updated for the first time, and the Gram matrix is calculated using the combined features Joint distribution matching via Gram matrix; Step 4: Obtain instance weights through joint distribution matching; train a prediction model based on the weights, and finally perform a secondary update of pseudo labels based on the prediction model; Step 5: During the repeated training of the prediction model, the updated labels are used again to extract the combined features and update the pseudo labels and joint distribution matching until the final prediction converges stably, and the training is terminated.
Citation Information
Patent Citations
Cross-project software defect prediction method based on supervised expression learning
CN110751186A
Cross-project software defect prediction method based on shared hidden layer auto-encoder
CN111198820A