Cross-project defect prediction method and system based on domain self-adaption
Through a domain adaptive method, the features of the source project and the target project are mapped to the same feature space, and through training of primary and auxiliary classification networks, the data distribution differences and feature applicability problems in cross-project defect prediction are solved, improving the accuracy and adaptability of the prediction.
Patent Information
- Application Number
- CN202510489029.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
There are problems with data distribution differences and feature applicability in the prediction of defects in cross-projects in the prior art, resulting in inaccurate predictions.
The domain adaptive method is adopted to map the features of the source project and the target project to the same feature space through the feature map network, and train it using the primary classification network and the auxiliary classification network to reduce the distribution difference and align the features.
The accuracy of cross-project defect prediction is improved, allowing the model to better adapt to different software projects and data distribution, and provides software development team with accurate defect prediction.
Smart Images

Figure CN120011202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect prediction, and in particular to a cross-project defect prediction method and system based on domain adaptation. Background Art
[0002] Software defect prediction technology is a key research direction in the field of software engineering. In the early stages of software defect analysis, the number of predicted defects is mainly based on rich experimental data and accumulated experience, and is often positively correlated with the volume of software code, which helps to evaluate the possible defect rate of the software and reasonably configure test resources accordingly. The source code measurement criteria involve static evaluation of code data, focusing on two key dimensions: code volume and complexity.
[0003] In real-world application scenarios, newly launched projects often face the problem of lack of sufficient historical data, or the existing historical data is insufficient to support effective predictions, which poses a major challenge to defect prediction within the project. Many new or small-scale projects often face the problem of insufficient data availability, which makes it difficult to train effective internal defect prediction models. In addition, developing high-quality defect prediction models requires a lot of resources and time, increasing costs and development cycles. Each software project has its own unique development environment, team composition, and coding practices. These characteristics affect the pattern of defect occurrence and increase the challenges of model adaptability and robustness across different projects. In order to address these issues, the CPDP (Cross Project Defect Prediction) method has begun to attract attention.
[0004] The challenges faced by cross-project defect prediction are indeed diverse, but one of the most critical challenges is the difference in data distribution between the source project and the target project. This difference may come from many aspects, including but not limited to the application field, development process, project size, etc. For example, a large project developed in the field of e-commerce may have a completely different code structure and maintenance mode from an embedded system project. These differences will cause the features and models that perform well on the source project to not be directly applied to the target project, thus affecting the accuracy of the prediction. In addition to the difference in data distribution, the problem of feature applicability is also an important factor. The effective defect prediction features found in the source project may not be directly migrated to the target project because there may be significant differences between different projects in code structure, development practices, and even the use of programming languages. These differences may cause the originally effective features to lose their predictive power in the new environment, or their predictive power may be greatly reduced. Summary of the invention
[0005] The present invention provides a cross-project defect prediction method and system based on domain adaptation, which are used to solve the defects in the prior art that the data distribution difference and feature applicability problems between the source project and the target project lead to inaccurate cross-project defect prediction, and improve the accuracy of cross-project defect prediction.
[0006] The present invention provides a cross-project defect prediction method based on domain adaptation, comprising:
[0007] Extracting features of a source item and a target item, and mapping the features of the source item and the target item to the same feature space using a feature mapping network;
[0008] Inputting the mapped features of the source item and the target item into the main classification network respectively to obtain the first defect prediction results of the source item and the target item;
[0009] Inputting the mapped features of the source item and the target item into the auxiliary classification network respectively to obtain second defect prediction results of the source item and the target item;
[0010] Determining a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0011] Training the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function so that the distribution difference between the source item and the target item is reduced and the features of the source item and the target item are aligned;
[0012] The features of the item to be predicted are extracted, and the features of the item to be predicted are sequentially input into the trained feature mapping network and the trained main classification network to obtain the defect prediction result of the item to be predicted output by the main classification network.
[0013] The present invention also provides a cross-project defect prediction system based on domain adaptation, comprising:
[0014] A feature extraction module, used to extract features of source items and target items, and map the features of the source items and the target items to the same feature space using a feature mapping network;
[0015] A first prediction module, used to input the mapped features of the source item and the target item into a main classification network respectively, to obtain first defect prediction results of the source item and the target item;
[0016] A second prediction module, used to input the mapped features of the source item and the target item into an auxiliary classification network respectively, to obtain second defect prediction results of the source item and the target item;
[0017] A loss calculation module, used to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0018] A model training module, used for training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so that the distribution difference between the source project and the target project is reduced and the features of the source project and the target project are aligned;
[0019] The defect prediction module is used to extract the features of the item to be predicted, input the features of the item to be predicted into the trained feature mapping network and the trained main classification network in sequence, and obtain the defect prediction result of the item to be predicted output by the main classification network.
[0020] The cross-project defect prediction method and system based on domain adaptation provided by the present invention can more accurately predict cross-project defects by carefully tuning the feature mapping network and classifier, making the method a powerful tool that not only has excellent prediction performance within a project, but can also adapt to different software projects and data distributions, providing accurate defect prediction for software development teams. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 It is a flowchart of a cross-project defect prediction method based on domain adaptation provided by the present invention;
[0023] Figure 2 It is a schematic diagram of reducing distribution differences and improving project prediction performance in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0024] Figure 3 It is a schematic diagram of the model structure of the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0025] Figure 4 It is a flowchart of the ATC2Defect method in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0026] Figure 5 It is a flowchart of syntax tree feature extraction in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0027] Figure 6 It is a schematic diagram of the overall network structure of the ATC2Defect defect prediction method in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0028] Figure 7 It is a schematic diagram of class imbalance processing in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0029] Figure 8 It is a schematic diagram for comparing improved cross-project predictions in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0030] Fig. 9 It is a schematic diagram comparing the GAT and GCN methods in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0031] Fig.10 It is a schematic diagram of cross-project prediction ablation analysis in the cross-project defect prediction method based on domain adaptation provided by the present invention. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] Combine the following Figure 1 A cross-project defect prediction method based on domain adaptation of the present invention is described, comprising:
[0034] Step 101, extracting features of a source item and a target item, and mapping the features of the source item and the target item to the same feature space using a feature mapping network;
[0035] Step 102, inputting the mapped features of the source item and the target item into the main classification network respectively, to obtain the first defect prediction results of the source item and the target item;
[0036] Step 103, inputting the mapped features of the source item and the target item into the auxiliary classification network respectively, to obtain the second defect prediction results of the source item and the target item;
[0037] Step 104, determining a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0038] Step 105, training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so that the distribution difference between the source item and the target item is reduced, and the features of the source item and the target item are aligned;
[0039] Step 106, extracting the features of the item to be predicted, inputting the features of the item to be predicted into the trained feature mapping network and the trained main classification network in sequence, and obtaining the defect prediction result of the item to be predicted output by the main classification network.
[0040] The challenges faced by cross-project defect prediction are indeed diverse, but one of the most critical challenges is the difference in data distribution between the source project and the target project. This difference may come from many aspects, including but not limited to application fields, development processes, project scale, etc.
[0041] Domain adaptation technology provides a possible solution to the problem of data distribution differences. Domain adaptation aims to adapt to new data distributions by adjusting the model, such as Figure 2 As shown, the generalization ability of the model on the target project is improved, and the distribution difference problem between the source project and the target project is solved. We will explore how to use the data of the source project to build a defect prediction model that can effectively adapt to the characteristics of the target project. In this way, the performance of CPDP can be improved, so that it can accurately predict potential software defects in different project environments.
[0042] First, the features of the source project and the target project are extracted. This embodiment does not limit the method for extracting the features of the source project and the target project.
[0043] Assume that the feature representation set of the source project is ,
[0044] Represents a tag in the source project. represents the feature of the i-th sample in the source project feature set, d represents the dimension of the feature, Represents the corresponding label. The feature representation set of the target project is ,in is the number of instance samples in the target project.
[0045] like Figure 3 As shown, the first part of the model contains two key network components, namely the feature mapping network and the main classification network.
[0046] The feature mapping network aims to learn effective feature representations from the data of the source and target items. It provides high-quality feature representations for subsequent classification tasks. The feature mapping network maps the data of the source and target items into a shared feature space through a deep neural network structure, so that it can better capture the commonalities and differences between the two.
[0047] The main classification network is responsible for classifying features based on the feature mapping network and generating pseudo labels for the target items. Pseudo labels refer to labels predicted by the model based on the knowledge it has learned on data without clear labels. In this way, pseudo labels can be used to represent potential defect instances in the target items. These pseudo labels, combined with the characteristics of the target items, can treat the target items as new, labeled data instances. This step greatly enriches the training data of the target items and improves the learning effect of the model.
[0048] The second part of the model introduces an auxiliary classification network, which is similar to a discriminator. The role of the auxiliary classification network is to identify inaccuracies in the pseudo-labels generated by the main classification network. It analyzes the pseudo-labels to find possible errors and feeds this information back to the main classification network. The auxiliary classification network can challenge the performance of the main classification network by maximizing the domain adaptation loss and the pseudo-label learning loss. The domain adaptation loss is used to measure the distribution difference between the source and target items, and the pseudo-label learning loss is used to evaluate the accuracy of the pseudo-labels.
[0049] Through this adversarial training method, the auxiliary classification network and the main classification network improve each other in continuous feedback and adjustment. The auxiliary classification network challenges the main classification network, forcing it to continuously improve the prediction accuracy of the target project. After receiving feedback, the main classification network optimizes its classification decisions to generate more accurate pseudo labels. In this way, the distribution difference between the source project and the target project gradually decreases, and the generalization ability of the model continues to increase. Ultimately, this dual-network collaborative mechanism significantly improves the performance of cross-project defect prediction, enabling it to accurately predict potential software defects in different project environments.
[0050] The auxiliary classification network and the main classification network have the same structure in the prediction part, for example, both contain a fully connected layer and a softmax layer. However, there is one more gradient reversal layer (GRL) than the main classification network.
[0051] The main classification network focuses on the classification task of the source and target items. The auxiliary classification network not only participates in classification, but also achieves domain alignment through GRL. The auxiliary classification network and the classification network are trained independently, and the parameters are not shared with each other.
[0052] Before training, the features must be preprocessed. In addition, the data must be normalized. For any feature in the source or target project, normalize it according to the following formula:
[0053]
[0054] In order to effectively align the feature distributions of the source and target items, a two-layer feature mapping network is used to map the feature spaces of the two items into a shared, low-dimensional common subspace through two consecutive mapping steps. This design aims to minimize the distribution differences between the source and target items in the new feature space and improve the generalization ability and accuracy of cross-item predictions.
[0055] The mapping function of the feature mapping network can be expressed as ,in It is the feature set of the source project and the target project. The feature representation of the source item and the target item after mapping can minimize the distribution difference of the source item and the target item in the common subspace by optimizing the parameters of the mapping function.
[0056] This embodiment can more accurately predict cross-project defects by carefully tuning the feature mapping network and classifier, making this method a powerful tool that not only performs well in intra-project predictions, but can also adapt to different software projects and data distributions, providing accurate defect predictions for software development teams.
[0057] Based on the above embodiments, the loss function in this embodiment includes a domain adaptive loss function, which is determined based on the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project.
[0058] The target item instance generates pseudo labels using the classification network to ensure its accuracy and to bring the source and target items closer in the new feature space. The distribution difference between the source and target items is measured using the output of the two classification networks to obtain the domain adaptation loss function.
[0059] Based on the above embodiment, the formula of the domain adaptation loss function in this embodiment is:
[0060]
[0061] in, is the domain adaptation loss function, is a hyperparameter, n s is the number of samples in the source project, n tis the number of samples in the target project, are the parameters of the main classification network, are the parameters of the auxiliary classification network, is the feature after mapping of the i-th sample in the source project, is the feature after mapping of the i-th sample in the target project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, and is the defect prediction probability; is the first defect prediction result of the i-th sample in the target project output by the main classification network, and is the defect prediction probability; is the second defect prediction result of the i-th sample in the source project output by the auxiliary classification network, which is the defect prediction probability; is the second defect prediction result of the i-th sample in the target project output by the auxiliary classification network, which is the defect prediction probability.
[0062] is an adjustable hyperparameter used to adjust the difference between the two, that is, the proportion of source items. For the auxiliary classification network, the domain adaptation loss function should be maximized during training. The reason for this design is to align the feature representations of the source domain and the target domain as much as possible during the training process, thereby improving the performance and generalization ability of the model in the target domain.
[0063] On the basis of the above embodiment, the loss function in this embodiment further includes one or more of a supervised contrastive learning loss function of the source project, a classification loss function of the main classification network, and a pseudo label loss function generated by the main classification network for the target project;
[0064] The supervised contrastive learning loss function of the source project is determined according to the mapped features of the source project and the target project, and the defect label of the source project;
[0065] The classification loss function of the main classification network is determined according to the defect label of the source item and the first defect prediction result;
[0066] The pseudo-label loss function is determined according to the first defect prediction result and the second defect prediction result of the target item.
[0067] In order to make full use of the label information of labeled instances to mine discriminative information, supervised contrastive learning loss can be used , whose main goal is to improve the model's discriminative ability by maximizing the similarity between similar samples and minimizing the similarity between non-similar samples. Compared with the traditional cross-entropy loss, supervised contrastive learning loss can more effectively cluster similar samples in the embedding space while separating samples of different classes. This method introduces more information for supervision, especially in multi-class classification tasks, and can significantly improve the generalization ability and robustness of the model.
[0068] For the main classification network, in order to guide its learning, a classification loss function is used.
[0069] The pseudo labels of the target item instances are generated by the classification network, and the pseudo label loss function is used to improve the accuracy.
[0070] Based on the above embodiment, the calculation formula of the supervised contrastive learning loss function in this embodiment is:
[0071]
[0072] in, is the supervised contrastive learning loss function, I is the set of labeled samples of the source project, i is the index of the anchor sample, P(i) is the set of positive samples of the same type as the anchor sample i in the source project, A(i) is the other samples in the source project and the target project except the anchor sample i, , , They represent the vectors of anchor sample i, positive sample p and other samples a after feature mapping, and sim is the similarity algorithm usually calculated using dot product or cosine similarity s; is the temperature parameter, which is used to control the scaling of the similarity score and affects the smoothness of the probability distribution;
[0073] The calculation formula of the classification loss function is:
[0074]
[0075] in, is the classification loss function, n s is the number of samples in the source project, are the parameters of the main classification network, is the feature after mapping of the i-th sample in the source project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the defect label of the i-th sample in the source project;
[0076] The calculation formula of the pseudo-label loss function is:
[0077]
[0078] in, is the pseudo label loss function, n t is the number of samples in the target project, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the target item output by the auxiliary classification network. is the logarithm of the inverse probability predicted by the auxiliary classification network. By designing this way, it is to minimize the probability that the target item sample is misclassified. By minimizing this loss function, the main classification network is forced to The output of the auxiliary classification network The output of is opposite, forcing the classification network to find a better decision boundary for the features in the common space, thereby enhancing the accuracy of the pseudo-label. This definition also introduces the idea of adversarial learning. The two networks compete with each other and constrain each other by maximizing and minimizing different goals, thereby improving the generalization ability of the model.
[0079] On the basis of the above embodiment, in this embodiment, the feature mapping network, the main classification network and the auxiliary classification network are trained according to the loss function, including:
[0080] Determining a first loss value according to the domain adaptation loss function, the supervised contrastive learning loss function, the classification loss function and the pseudo label loss function;
[0081] Adjusting parameters of the feature mapping network and the main classification network so that the first loss value is minimized;
[0082] Determining a second loss value according to the domain adaptation loss function and the pseudo label loss function;
[0083] The parameters of the auxiliary classification network are adjusted so that the second loss value is maximized.
[0084] When constructing feature mapping networks and classification networks, the goal is to minimize their losses on standard classification tasks. However, when introducing an auxiliary classification network for adversarial training, it is hoped that this auxiliary network will perform as well as possible in distinguishing the data after feature mapping, even if these data are generated from different distributions. Therefore, for the auxiliary classification network, the goal is to maximize its loss in distinguishing these feature mapping data. So the adversarial loss is defined as:
[0085] ;
[0086] ;
[0087] in, is the classification loss function, is the supervised contrastive learning loss function, is the pseudo label loss function, is the domain adaptation loss function, are the parameters of the feature mapping network, are the parameters of the main classification network, is the parameter of the auxiliary classification network. If there is no triangle above the parameter, it means that the parameter is currently being optimized. If there is a triangle above the parameter, it means that the parameter is currently fixed and will not be updated. and is a coefficient used to balance several losses.
[0088] The pseudo code of the process of the method provided in this embodiment is shown in Algorithm 1, and the specific steps include:
[0089] Step 1: The feature mapping network is responsible for mapping the data of the source and target items into a shared feature space, learning the common differences between the two and supervising the contrastive loss.
[0090] Step 2: The main classification network is responsible for classifying the data based on the shared features generated by the feature mapping network and generating pseudo labels for the target items.
[0091] Step 3: The auxiliary classification network further optimizes the pseudo-labels, plays the role of a discriminator, evaluates the quality of the pseudo-labels, and challenges the main classification network.
[0092] Step 4: Continuously optimize pseudo labels and classification decisions through adversarial training between the main classification network and the auxiliary classification network.
[0093] Algorithm 1: Cross-project defect prediction method based on domain adaptation
[0094] Require: Source project feature vector and tags , the target item feature vector
[0095] Ensure: Target item prediction label
[0096] 1: Initialize network parameters , , and hyperparameters , , ,
[0097] 2: while the maximum number of iterations or convergence condition is not reached do
[0098] 3: Map the data of the source and target items to a common subspace:
[0099] 4: Calculate supervised contrast loss
[0100] 5: Calculate classification loss
[0101] 6: Computational Domain Adaptive Loss
[0102] 7: Calculate pseudo-label loss
[0103] 8: Update feature maps and classification network parameters by minimizing the loss of the formula
[0104] 9: Update the auxiliary classification sub-network parameters by maximizing the formula loss
[0105] 10: end while
[0106] 11: return target item prediction
[0107] Initially, the parameters of the feature mapping subnetwork, classification subnetwork, and auxiliary classification subnetwork are all random. The pull-in segment model attempts to reduce the distribution difference between the source and target domains to improve generalization and pseudo-label generation. The gradient information of the formula loss is back-propagated to the feature mapping subnetwork and the classification subnetwork to adjust their parameters. The classification subnetwork uses the updated parameters to predict the unlabeled data and generate pseudo-labels.
[0108] The auxiliary classification subnetwork in the boost phase tries to maximize the domain adaptation loss and pseudo-label loss , which means it tries to find the decision boundary that best distinguishes the feature representations between the source and target domains. The parameters of the auxiliary classification network are updated by maximizing the gradient of the formula loss based on the distribution difference between the two domains measured by the output of the two classification networks.
[0109] The push phase and the pull phase are performed alternately until the predetermined number of training iterations or convergence conditions are reached. In each iteration, the parameters of the model are continuously adjusted to achieve a balance between maximizing the domain adaptation loss and minimizing the pseudo-label loss. This process forms a dynamic adversarial cycle. Through continuous iterations, the model gradually improves its performance on unlabeled data, while improving its ability to adapt to the distribution differences between the source and target domains, thereby improving the performance of cross-project defect prediction.
[0110] Based on the above embodiments, in this embodiment, the features of the source project and the target project are extracted, including:
[0111] Convert the source project or the target project into an abstract syntax tree, and process the abstract syntax tree based on a convolutional neural network to obtain features of the abstract syntax tree;
[0112] Building a class dependency network according to the dependency relationship in the source project or the target project, and acquiring features of the class dependency network based on a network embedding algorithm;
[0113] After fusing the code metrics of the source project or the target project with the features of the class dependency network, the code metrics are again fusing with the features of the abstract syntax tree;
[0114] The re-fused features are embedded into the class dependency network, a graph attention network is constructed, and graph node features are obtained.
[0115] The PROMISE library is a public data repository specifically designed to support empirical research in the field of software engineering. It provides researchers with a wealth of software defect prediction-related datasets. The source and target projects can select datasets of 7 classic software projects published in the PROMISE library, totaling 21 different versions of data. These datasets have been widely used in the field of software defect prediction research. These datasets contain rich historical data, providing a solid foundation for reliable data analysis and model building. Table 1 provides a detailed description of these datasets, including key information such as the source, version information, number of instances, and defect rate of each dataset.
[0116] Table 1 Dataset
[0117]
[0118] The data sets of these projects are very representative, covering different types of software projects, and can represent various application scenarios and code structures. These data sets are not only used in this embodiment, but also involve multiple fields such as data management, search, and text processing, providing empirical research cases for this embodiment.
[0119] This embodiment adopts multi-view feature extraction, which includes traditional code metrics, Abstract Syntax Tree (AST) feature extraction and Class Dependency Network (CDN) feature extraction.
[0120] (1) Traditional software features include two major categories: one is static metrics based on structure, mainly including LOC metrics, McCabe metrics, Halstead metrics, etc. The other is the CK metric based on object-oriented characteristics. The LOC metric is a method to measure the size and complexity of software. This metric also has certain limitations because it cannot reflect the quality and functionality of the code. The McCabe metric measures the complexity of software code based on the number of nodes and edges in the program graph, which can help developers identify potential logical errors and difficult-to-test code paths. The Halstead metric is used to measure the complexity of software. Based on the number of operators and operands and their different combinations, these metrics include the length, volume, difficulty and complexity of the program. The CK metric is mainly used to measure the complexity of object-oriented software. This metric analyzes the complexity, inheritance structure, coupling and cohesion of object-oriented software, thereby guiding software design and optimization.
[0121] (2) Abstract Syntax Tree (AST): In the life cycle of software development, the compilation process of source code is a key step in converting the high-level language code written by developers into machine code that computers can understand and execute. This process usually includes several consecutive stages, each of which analyzes and converts the source code at different levels.
[0122] (3) Class dependency network features. In the method of the invention, the class dependency network can be analyzed and various features related to software quality and potential defects can be extracted. In addition to feature extraction, the class dependency network can also be used for software measurement and defect location. By measuring various attributes and relationships in the CDN, the development team can obtain quantitative information about the health of the software system. When defects occur, the CDN can help quickly locate possible problem modules, thereby accelerating the defect repair process. In addition, CDN also plays an important role in software reconstruction. It can help identify modules that need to be reconstructed and possible reconstruction paths, thereby improving the maintainability and scalability of the software system.
[0123] When modeling software dependency networks, usually only a single network structure information is considered, while the rich feature information contained in each node in the class dependency network itself is ignored. This method is not enough to fully utilize code features and may lead to a decrease in prediction performance. Node feature information includes code complexity, change history, developer information, etc., which are of great significance for accurately predicting defects. In addition, the grammatical semantic information of the code module where each node is located is also crucial. Combining multi-view information and node features for defect prediction is expected to significantly improve the performance and application value of the model. By integrating multi-view features, more comprehensive code semantics and structural information can be captured, improving the accuracy of prediction. Making full use of the rich feature information of nodes can further improve the model's ability to identify defects, and ultimately promote the development and maintenance of high-quality software.
[0124] like Figure 4 As shown in the figure, the process of the ATC2Defect method includes: (1) combining the traditional features of the dataset and the class dependency network embedding features. (2) parsing the source code into an abstract syntax tree, extracting the token sequence through specific rules, and then embedding the token sequence into a matrix to extract features through CNN. (3) taking the class dependency network as the data of the last view, the features obtained in the first two steps are fused and embedded into the class dependency network to construct a GCN (Graph Convolution Network) model. (4) Defect prediction is performed through the output of the graph convolution network.
[0125] This embodiment uses an AST traversal algorithm based on primary node priority to obtain the features of the abstract syntax tree. The source code is constructed as an abstract syntax tree to extract features. By using the abstract syntax tree, the structure and syntax information of the code can be captured. And by extracting the nodes and subtrees in the syntax tree, detailed information about the code structure can be obtained, which is helpful to identify potential defects in the code. In order to extract a node sequence containing important semantic information from the abstract syntax tree, this paper proposes a traversal algorithm based on primary node priority (PNPT). The method has the following advantages and characteristics: the algorithm introduces the concept of primary node, and gives priority to processing the node with the largest subtree during the traversal process. In this way, it can ensure that more important semantic information is retained as much as possible in the generated token sequence, reducing the possibility of information loss. By calculating the subtree size of each node, the algorithm can dynamically adjust the traversal order. No matter how complex the code structure is, the algorithm can adapt flexibly to ensure that key nodes are processed first. The primary node priority selection strategy can retain the key information in the code to the maximum extent. This has a significant effect improvement for subsequent tasks based on code semantic information. For token sequences that exceed the maximum length, the algorithm can intelligently prioritize and retain key nodes, thus avoiding the loss of important information due to truncation.
[0126] After generating the token sequence, each source file is converted into a set of tokens. In order to input the data into the neural network, the token set needs to be mapped, and each token is mapped to an integer. In this way, each source code file is converted into an integer vector. Considering that the length of the tokens generated by each file is different, it is necessary to fill the shorter vectors with zeros to ensure that the length of the vectors generated by each file is consistent. Consider a project named P, which contains n Java source files, denoted as Through the token extraction method, a series of tokens are extracted from these files, which are recorded as a set . Let the input word index sequence be , where n represents the sequence length. Assume that the embedding matrix is , where V represents the vocabulary size and d represents the word embedding dimension. Then the output of the embedding layer is The calculation formula is as follows:
[0127]
[0128] in, Expressive words The corresponding word embedding vector.
[0129] Assume that the input embedding vector sequence is ,in represents the word embedding vector. Let the convolution kernel be ,in Indicates the size of the convolution kernel. Software defect prediction based on multi-view feature learning indicates the number of convolution kernels or the number of output channels. Output of the convolution layer The calculation formula is as follows:
[0130]
[0131] in, Represents the sequence from word i to word b represents a bias vector, and ReLU represents a linear rectifier unit activation function.
[0132] For each feature map , take the feature value after the maximum pooling, , the feature vector after pooling is , the pooled feature vector is transformed through the fully connected layer, and finally the syntax tree feature is obtained The whole process of syntax tree feature extraction is as follows Figure 5 shown.
[0133] .
[0134] The class dependency network shows the dependency relationship between classes in the form of a graph, including inheritance, interface implementation, aggregation or method call, etc. After constructing the class dependency network, it is necessary to obtain the embedded representation of each node through network embedding technology. The network embedding vector can capture the structural information and semantic information of the node in the network, making the initial attributes of each node more expressive. This representation can better reflect the role and relationship of the node in the entire graph. In the modeling of the class dependency network in this embodiment, the node2vec algorithm is first used to embed the class dependency network in the network to capture the structural and semantic information in the network. The node2vec algorithm explores the local and global structures in the network by adjusting the two parameters p and q in the random walk.
[0135] We can set p=2 and q=0.15. This parameter selection aims to balance the homogeneity and structural roles of nodes in order to better reveal the similarities and differences between nodes. In this way, an embedding vector is generated for each node in the CDN. :
[0136]
[0137] Next, we normalize traditional software engineering features, such as code complexity and coupling, to eliminate the impact of different dimensions. Normalization can be performed using the following formula:
[0138]
[0139] in, is the i-th feature in the original feature vector, is the i-th feature in the normalized feature vector. Normalization ensures that all feature values are in the range of 0 to 1. After that, the normalized traditional feature vector is concatenated with the network embedding vector to form an enhanced feature representation .
[0140] In order to further reduce the dimension of the features and retain the most important information, this embodiment uses principal component analysis dimensionality reduction technology. The original features are converted to a new coordinate system through linear transformation, so that the data variance on the first coordinate axis of this new coordinate system is the largest, and each subsequent coordinate axis is orthogonal to the previous coordinate axis and has the largest variance. In this way, a low-dimensional feature representation can be obtained. Finally, the above features are combined with the features obtained from the abstract syntax tree By combining them, we can get the feature X of the graph node. The formula is as follows:
[0141]
[0142] Taking into account the different contributions from the two features, another combination method with weight parameters can be used. The formula is as follows:
[0143]
[0144] in, It is a combination method. In this study, three methods are selected: element-wise addition (add), element-wise multiplication (multiply) and concatenation (concat).
[0145] Following the graph convolutional network architecture and parameter settings proposed by Kipf, a double-layer convolution structure is also adopted. Using the node attribute feature matrix X obtained in the previous step, the final output result can be calculated by the following formula.
[0146]
[0147] In this formula, is the weight between the hidden layer and the input layer, represents the number of hidden layer units, Represents feature dimension is the weight coefficient between the hidden layer and the output layer. is the preprocessed adjacency matrix, is the degree matrix of the node, is the normalized adjacency matrix, is the adjacency matrix of the graph, of size N×N, Is the identity matrix, size N × N. Output feature matrix The dimension is , ReLU function is used as the nonlinear activation function in the network.
[0148] The final prediction output of the network is:
[0149]
[0150] Among them, W (2) It is the weight matrix of the last layer (output layer) of the model, which is used to map the feature matrix F to the final classification space.
[0151] The loss function is:
[0152]
[0153] in, is the true label, Represents the predicted probability after the softmax function of the prediction layer.
[0154] The overall structure of the ATC2Defect defect prediction method is as follows Figure 6 As shown in FIG. 1 , this method combines the abstract syntax tree (AST), traditional code metrics (Traditional Features) and class dependency network (Class Dependency Network), so this method is called ATC2Defect.
[0155] There are large differences in graph structures and identifiers in syntax trees between different projects, which will have a negative impact on cross-project prediction. Therefore, when using the ATC2Defect method for cross-project defect prediction, some improvements need to be made in these aspects.
[0156] First, when embedding the abstract syntax tree, avoid using identifiers directly like in the prediction task within the project, and use the type of the current node instead. Identifiers are usually highly project-specific because they are arbitrarily named by developers based on project requirements and coding style. Identifiers that are valid in the source project may be completely different in the target project, which will cause the model to overfit the source project data, thereby reducing the generalization ability of the model in the target project. Compared with identifiers, node types are more universal. AST node types represent the grammatical structure and logical relationships of the source code, rather than specific implementation details. This universality enables the model to learn more abstract code structure features without relying on specific identifiers. In this way, overfitting can be reduced and more universal features can be learned.
[0157] Secondly, when processing graph structures, you should choose a Graph Attention Network (GAT) with better generalization performance. Compared with GCN, GAT has the following advantages in generalization ability: On the one hand, in GCN, the feature representation of a node is aggregated through the average features of neighboring nodes, which cannot distinguish the importance of different neighboring nodes. GAT can adaptively assign different weights to each neighboring node, so that important neighboring nodes contribute more to the feature representation of the central node, thereby better capturing the key information in the graph structure. On the other hand, GAT uses multiple independent attention heads, each of which can capture different patterns in the graph structure. By combining these patterns together, GAT can improve the expressiveness and generalization ability of the model. The formula for GAT is as follows:
[0158]
[0159] in, Representation Node and nodes The attention coefficient between is a trainable parameter vector, is a trainable weight matrix, is a vector concatenation operation, and The nodes are and nodes Next, the attention weights It can be obtained by normalization:
[0160]
[0161] Finally, the node The feature representation of is updated as:
[0162]
[0163] in, Representation Node The set of neighbor nodes of is a nonlinear activation function. By introducing the graph attention network, we can better cope with the differences in graph structures between different projects and improve the accuracy and robustness of cross-project defect prediction. In addition, when processing abstract syntax trees, using node types instead of specific identifiers can significantly reduce the overfitting of the model, thereby improving the universality of features and the generalization ability of the model.
[0164] In this embodiment, the ATC2Defect network is first used to extract and characterize the code data of the source project and the target project, and two feature sets are obtained: the feature set of the source project and the feature set of the target item These two feature sets capture the essential characteristics of the source project and target project code respectively, providing a basis for subsequent defect prediction. , replace the original classifier with Figure 3 The network structure shown.
[0165] This example will use ATC2Defect to predict cross-project defects, and make improvements on this basis to improve the performance of cross-project defect prediction. We will focus on how to further improve the performance of ATC2Defect in cross-project defect prediction, hoping that ATC2Defect can better adapt to diverse software projects and provide more accurate defect prediction for software development teams, thereby helping them identify and fix potential software defects in advance and improve software quality and development efficiency.
[0166] Based on the above embodiments, after extracting the features of the source project and the target project, this embodiment further includes:
[0167] For a sample set in which the number of samples of any defect category in the source project or the target project is less than a preset threshold, determining the density of each sample in the sample set;
[0168] Determine the number of neighbors of each sample according to the density of each sample based on a density function;
[0169] Selecting the nearest neighbor sample of each sample according to the number of neighbors of each sample, and calculating the weight of each nearest neighbor sample;
[0170] Randomly select neighbor samples from the nearest neighbor samples of each sample to generate a synthetic sample together with each sample;
[0171] The synthesized sample is adjusted using the weight of the randomly selected neighbor sample to obtain a new sample.
[0172] Since the data for defect prediction is class imbalanced, the data needs to be resampled. However, for the graph structure, data sampling will destroy the original graph structure. On the other hand, for cross-version prediction and cross-project prediction, the graph structures of the two data sets will be completely different due to the addition of new instances. Therefore, the model of the method of the present invention is divided into two parts: pre-training part and downstream task. The feature F of each project is obtained through pre-training, and the downstream task samples these features to deal with the class imbalance problem, such as Figure 7 As shown, the specific steps include:
[0173] Step 1: The syntax tree features (AST), traditional metric features, and dependency network features are integrated to obtain comprehensive multi-view features.
[0174] Step 2: ATC2 fuses the three features into a model, preprocesses the features, and divides the training data into test data.
[0175] Step 3: There may be a data imbalance problem in the training data, and data sampling and test data are needed as new data sets.
[0176] Step 4: After processing the data, input it into the classifier model for training.
[0177] Step 5: The classifier predicts the category of each sample based on the features in the test dataset, determining whether it has a defect (Bug) or is clean (Clean).
[0178] This phased model building approach can not only effectively deal with the class imbalance problem, but also adapt to the needs of cross-version and cross-project prediction, improving the generalization ability and prediction accuracy of the model in different software projects.
[0179] In order to solve the problem of noise and unreasonable samples introduced, this embodiment proposes a density-weighted SMOTE algorithm, which is named DW-SMOTE algorithm, see the pseudo code of Algorithm 2. Its main feature is the dynamic selection of the number of neighbors, which is related to the local density of each sample. The most suitable number of neighbors is determined by the density function to ensure that the true distribution of the data can be better reflected when generating synthetic samples. The algorithm first calculates the density of each minority class sample, and dynamically adjusts the number of neighbors according to the density, and then calculates the weight of each neighbor sample, and the weight is used to guide the subsequent sample generation process. When generating synthetic samples, DW-SMOTE not only randomly selects neighbors, but also adjusts the position of synthetic samples according to the weight of the neighbors, thereby generating more diverse and high-quality synthetic data. Through the noise filtering step, possible abnormal samples are further removed, and the reliability of the synthetic data is improved. In this way, DW-SMOTE can better adapt to the local distribution characteristics of the data and improve the performance and stability of the classification model.
[0180] Algorithm 2: Density weighted upsampling algorithm DW-SMOTE
[0181] Require: minority class sample set S, oversampling rate N
[0182] Ensure: Increase the number of minority class samples S new
[0183] 1: Initialize S new ←
[0184] 2: for each sample x i ∈ S do
[0185] 3: Calculate sample x i The density ρ i
[0186] 4: Dynamically select the number of neighbors k i ← f (ρ i ), f is the density function
[0187] 5: Select x i K i Nearest neighbor samples
[0188] 6: Initialize sample weights ← []
[0189] 7:for each neighbor sample x i,j do
[0190] 8: Calculate sample weight w i,j
[0191] 9: weights.append(w i,j )
[0192] 10: end for
[0193] 11: end for
[0194] 12: for each sample x i ∈ S do
[0195] 13:for n= 1 to N do
[0196] 14: Randomly select a neighbor sample x i,j
[0197] 15: Generate synthetic sample x new ← x i + θ · (x i,j -x i ), where θ~ U(0,1)
[0198] 16: According to the weight w i,j Update x new
[0199] 17:S new .append(x new )
[0200] 18: end for
[0201] 19: end for
[0202] 20: For S new Perform noise filtering to remove abnormal samples
[0203] 21: return S new
[0204] like Figure 3 As shown, the overall framework is divided into two adversarial training parts: the first part is the feature mapping network and the classification network. The feature mapping network is composed of two layers of fully connected neural networks with 64 and 32 neurons respectively, and the activation function is ReLU. The classification network is a classification network composed of a single-layer fully connected network, equipped with two output nodes, and uses the softmax function for output conversion. The second part contains an auxiliary classification network with the same structure as the classification network. In order to optimize the predictive ability of the model, the present invention adopts a grid search method to determine a set of hyperparameters that can maximize the predictive performance of the model. Grid search is a systematic hyperparameter tuning strategy that searches for the best hyperparameter settings by traversing all possible combinations in a predefined set of hyperparameters. In this study, The value of is finally determined to be [0.95, 1.0, 2.0, 0.07]. In order to ensure the smooth progress of model training and the need to process large-scale data sets, the experiment was performed on a computer equipped with 32GB of memory, an NVIDIA GeForce GTX 1080Ti graphics card was used, and the framework was PyTorch.
[0205] The present invention refers to the improved cross-project prediction method as Table 2 and Table 3 show the comparison results of the improved method in this paper and the comparison method in terms of F1 and AUC (Area Under Curve, area under the ROC curve) indicators. The values shown in the table are the average values of cross-project predictions with the current project as the target project and the other data sets as the source projects.
[0206] From Table 2, The method shows a significant improvement in F1 scores in some projects (such as ant and poi), which shows that for specific projects, the method can effectively improve the prediction accuracy. The average F1 score across all projects is 0.522, which is higher than other methods, indicating that the improved method has better overall performance in the cross-project prediction scenario.
[0207] Table 2 Improved cross-project prediction F1
[0208]
[0209] According to Table 3, in terms of AUC index, It also performs well, especially in the poi and velocity projects, which are significantly improved compared with other methods, which emphasizes its effectiveness in distinguishing defect and non-defect classes. The AUC value also shows the performance differences between different projects, which may be related to the specific characteristics of the project and the defect distribution. For example, the AUC value of the velocity project is relatively low, which may be because the data characteristics of this project make defect prediction more difficult. On average, The average AUC value of the method in all projects is 0.665, which exceeds all other compared methods. This shows that the method performs well in balancing the prediction ability of positive and negative samples and has better generalization ability.
[0210] Table 3 Improved cross-project prediction AUC
[0211]
[0212] like Figure 8 As shown in (a) and (b), the horizontal axis represents the project names in Tables 2 and 3, and the corresponding coordinate points on each project name are the F1 scores or ACU indicators of different methods. The improved method for cross-project defect prediction of the present invention has significant improvements compared with the original method. In terms of F1 score, the average performance has been improved by 12.5%, and in terms of AUC indicator, the average performance has been improved by 18.3%. At the same time, compared with the best performing method in the benchmark model, the average improvement in the two indicators is 9.4% and 12.7%, respectively. These results show that through carefully designed feature mapping and adversarial training strategies, the data distribution differences between the source project and the target project can be effectively narrowed, while making full use of limited label information to improve the prediction ability of the model.
[0213] In order to verify the performance difference of using GCN and GAT in the improved method, the data of seven projects were trained and tested. The evaluation indicators include F1 score and AUC value, both of which are used to measure the classification performance of the model. The experimental results are shown in Figure 2. Fig. 9 As shown in (a) and (b) in the figure, the horizontal axis is the project name in Table 2 and Table 3, and the coordinate point corresponding to each project name is the F1 score or ACU index of the GCN and GAT methods.
[0214] On all datasets, GAT's F1 score increased by 3% on average. This shows that GAT performs better in accurately predicting defects and reducing false positives. GAT also shows a clear advantage in AUC value, with an average increase of about 4%. In summary, the experimental results show that GAT performs significantly better than GCN in cross-project defect prediction, especially when dealing with projects with large differences in data distribution. GAT can better adapt to different project characteristics, capture complex relationships in graph structures, and integrate information from different modes to improve the model's expressiveness and generalization capabilities. When the node feature noise is large, GAT can ignore irrelevant features through the attention mechanism, thereby reducing overfitting and improving the accuracy and reliability of predictions.
[0215] In this section, ablation analysis is performed to evaluate the impact of different components in the improved method on the cross-project defect prediction performance. In order to explore the impact of three different loss functions on the method of the present invention and gain a deeper understanding of the contribution of each component to the overall performance, a series of comparative experiments are designed, in which a specific loss function is removed each time to observe its impact on the model performance. First, the complete method containing all three loss functions is referred to as ATC. In order to evaluate the contribution of domain adaptation loss, the domain adaptation loss is removed, and this variant is called ATC-da. In order to evaluate the impact of supervised contrast loss, the supervised contrast loss is removed, and this variant is called ATC-sc. Finally, in order to evaluate the role of pseudo-label learning loss, the pseudo-label loss is removed, and this variant is called ATC-pl. Table 4 and Fig.10 The results of the above ablation experiments are shown. By comparing the performance of the model under different loss function combinations, we can more clearly understand the importance of each component in the improvement method. Fig.10 The horizontal axis in is the project name in Table 2 and Table 3, and the coordinate point corresponding to each project name is the F1 score or ACU index of different methods.
[0216] Table 4 Cross-project prediction ablation analysis
[0217]
[0218] The corresponding F1 scores and AUC values are recorded to evaluate the specific impact of each component. The experimental results show that removing any loss function will lead to a decrease in model performance, proving the importance of each loss in the model training process. Specifically, when removing the domain adaptation loss (i.e. ATC-da model), the average F1 score of the model dropped from 0.522 to 0.467, and the average AUC value dropped from 0.665 to 0.617. This result shows that It plays a key role in promoting the transfer learning of models between different projects and enhancing their generalization ability. Similarly, supervised contrast loss The removal of (i.e., ATC-sc model) also leads to a decrease in performance, although its impact is less than that of the removal of This reflects Contribution in optimizing the model’s ability to distinguish between different classes of defects, especially when faced with an imbalanced dataset. The performance degradation caused by removing and This shows that although It plays a positive role in using unlabeled data for semi-supervised learning, but its effect on improving model performance is not as good as .
[0219] In summary, these ablation experiments not only demonstrate the effectiveness of the overall method, but also reveal the importance and mechanism of each component. In particular, the domain adaptation loss It plays a vital role in improving the generalization ability of the model, while the supervised contrast loss and pseudo-label loss They respectively optimized the model's category distinction ability and the use of unlabeled data, which jointly promoted the performance improvement of the model in cross-project defect prediction tasks.
[0220] The present invention conducts an in-depth discussion and analysis on the application of the improved model of ATC2Defect in the field of cross-project defect prediction. Through a series of experiments, the model's ability to transfer learning between different projects was evaluated, and a variety of feature selection and engineering strategies were tried to improve the generalization performance of the model. By improving the abstract syntax tree embedding strategy, introducing the graph attention network, and using domain adaptation technology to compensate for the data distribution differences between the source project and the target project, the performance of cross-project defect prediction is improved. Through experiments, the present invention demonstrates the prediction performance of the improved method on multiple projects and compares it with other methods. The results show that the improved α-ATC (CPDP) version achieved the best F1 score and AUC value in most projects, which is about 12.5% and 18.3% higher than the original method in F1 value and AUC value, respectively, and 9.4% and 12.7% higher than several benchmark methods on average.
[0221] The cross-project defect prediction system based on domain adaptation provided by the present invention is described below. The cross-project defect prediction system based on domain adaptation described below and the cross-project defect prediction method based on domain adaptation described above can refer to each other.
[0222] The system includes:
[0223] A feature extraction module, used to extract features of source items and target items, and map the features of the source items and the target items to the same feature space using a feature mapping network;
[0224] A first prediction module, used to input the mapped features of the source item and the target item into a main classification network respectively, to obtain first defect prediction results of the source item and the target item;
[0225] A second prediction module, used to input the mapped features of the source item and the target item into an auxiliary classification network respectively, to obtain second defect prediction results of the source item and the target item;
[0226] A loss calculation module, used to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0227] A model training module, used for training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so that the distribution difference between the source project and the target project is reduced and the features of the source project and the target project are aligned;
[0228] The defect prediction module is used to extract the features of the item to be predicted, input the features of the item to be predicted into the trained feature mapping network and the trained main classification network in sequence, and obtain the defect prediction result of the item to be predicted output by the main classification network.
[0229] This embodiment can more accurately predict cross-project defects by carefully tuning the feature mapping network and classifier, making this method a powerful tool that not only performs well in intra-project predictions, but can also adapt to different software projects and data distributions, providing accurate defect predictions for software development teams.
[0230] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-project defect prediction method based on domain adaptation, characterized in that: include: Extracting features of a source item and a target item, and mapping the features of the source item and the target item to the same feature space using a feature mapping network; Inputting the mapped features of the source item and the target item into the main classification network respectively to obtain the first defect prediction results of the source item and the target item; Inputting the mapped features of the source item and the target item into the auxiliary classification network respectively to obtain second defect prediction results of the source item and the target item; Determining a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project; Training the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function so that the distribution difference between the source item and the target item is reduced and the features of the source item and the target item are aligned; The features of the item to be predicted are extracted, and the features of the item to be predicted are sequentially input into the trained feature mapping network and the trained main classification network to obtain the defect prediction result of the item to be predicted output by the main classification network.
2. The cross-project defect prediction method based on domain adaptation according to claim 1 is characterized in that: The loss function includes a domain adaptive loss function, which is determined according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project.
3. The cross-project defect prediction method based on domain adaptation according to claim 2 is characterized in that: The formula of the domain adaptation loss function is: ; in, is the domain adaptation loss function, is a hyperparameter, n s is the number of samples in the source project, n t is the number of samples in the target project, are the parameters of the main classification network, are the parameters of the auxiliary classification network, is the feature after mapping of the i-th sample in the source project, is the feature after mapping of the i-th sample in the target project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the source project output by the auxiliary classification network, is the second defect prediction result of the i-th sample in the target item output by the auxiliary classification network.
4. The cross-project defect prediction method based on domain adaptation according to claim 2 is characterized in that: The loss function further includes one or more of a supervised contrastive learning loss function of the source project, a classification loss function of the main classification network, and a pseudo-label loss function generated by the main classification network for the target project; The supervised contrastive learning loss function of the source project is determined according to the mapped features of the source project and the target project, and the defect label of the source project; The classification loss function of the main classification network is determined according to the defect label of the source item and the first defect prediction result; The pseudo-label loss function is determined according to the first defect prediction result and the second defect prediction result of the target item.
5. The cross-project defect prediction method based on domain adaptation according to claim 4 is characterized in that: The calculation formula of the supervised contrastive learning loss function is: ; in, is the supervised contrastive learning loss function, I is the set of labeled samples of the source project, i is the index of the anchor sample, P(i) is the set of positive samples of the same type as the anchor sample i in the source project, A(i) is the other samples in the source project and the target project except the anchor sample i, , , They represent the vectors of anchor sample i, positive sample p and other samples a after feature mapping, sim is the similarity algorithm, is the temperature parameter; The calculation formula of the classification loss function is: ; in, is the classification loss function, n s is the number of samples in the source project, are the parameters of the main classification network, is the feature after mapping of the i-th sample in the source project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the defect label of the i-th sample in the source project; The calculation formula of the pseudo label loss function is: ; in, is the pseudo label loss function, n t is the number of samples in the target project, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the target item output by the auxiliary classification network.
6. The cross-project defect prediction method based on domain adaptation according to claim 4 is characterized in that: The feature mapping network, the main classification network and the auxiliary classification network are trained according to the loss function, including: Determining a first loss value according to the domain adaptation loss function, the supervised contrastive learning loss function, the classification loss function and the pseudo label loss function; Adjusting parameters of the feature mapping network and the main classification network so that the first loss value is minimized; Determining a second loss value according to the domain adaptation loss function and the pseudo label loss function; The parameters of the auxiliary classification network are adjusted so that the second loss value is maximized.
7. The cross-project defect prediction method based on domain adaptation according to claim 6 is characterized in that: The first loss value is determined according to the domain adaptation loss function, the supervised contrastive learning loss function, the classification loss function and the pseudo label loss function by the following formula: ; The second loss value is determined according to the domain adaptation loss function and the pseudo label loss function by the following formula: ; in, is the classification loss function, is the supervised contrastive learning loss function, is the pseudo label loss function, is the domain adaptation loss function, are the parameters of the feature mapping network, are the parameters of the main classification network, is the parameter of the auxiliary classification network. If there is no triangle above the parameter, it means that the parameter is currently being optimized. If there is a triangle above the parameter, it means that the parameter is currently fixed and will not be updated. and is the balance coefficient.
8. The cross-project defect prediction method based on domain adaptation according to any one of claims 1 to 7, characterized in that: Extract features of source and target items, including: Convert the source project or the target project into an abstract syntax tree, and process the abstract syntax tree based on a convolutional neural network to obtain features of the abstract syntax tree; Building a class dependency network according to the dependency relationship in the source project or the target project, and acquiring features of the class dependency network based on a network embedding algorithm; After fusing the code metrics of the source project or the target project with the features of the class dependency network, the code metrics are again fusing with the features of the abstract syntax tree; The re-fused features are embedded into the class dependency network, a graph attention network is constructed, and graph node features are obtained.
9. The cross-project defect prediction method based on domain adaptation according to any one of claims 1 to 7, characterized in that: After extracting the features of the source and target items, it also includes: For a sample set in which the number of samples of any defect category in the source project or the target project is less than a preset threshold, determining the density of each sample in the sample set; Determine the number of neighbors of each sample according to the density of each sample based on a density function; Selecting the nearest neighbor sample of each sample according to the number of neighbors of each sample, and calculating the weight of each nearest neighbor sample; Randomly select neighbor samples from the nearest neighbor samples of each sample to generate a synthetic sample together with each sample; The synthesized sample is adjusted using the weight of the randomly selected neighbor sample to obtain a new sample.
10. A cross-project defect prediction system based on domain adaptation, characterized in that: include: A feature extraction module, used to extract features of source items and target items, and map the features of the source items and the target items to the same feature space using a feature mapping network; A first prediction module, used to input the mapped features of the source item and the target item into a main classification network respectively, to obtain first defect prediction results of the source item and the target item; A second prediction module, used to input the mapped features of the source item and the target item into an auxiliary classification network respectively, to obtain second defect prediction results of the source item and the target item; A loss calculation module, used to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project; A model training module, used for training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so that the distribution difference between the source project and the target project is reduced and the features of the source project and the target project are aligned; The defect prediction module is used to extract the features of the item to be predicted, input the features of the item to be predicted into the trained feature mapping network and the trained main classification network in sequence, and obtain the defect prediction result of the item to be predicted output by the main classification network.
Citation Information
Patent Citations
Domain-adaptive object surface defect detection method and system, and storage medium
CN115564752A
Cross-project software defect prediction method based on domain self-adaption
CN115658504A
Cross-project software defect prediction method based on multi-peak feature alignment
CN118363643A
Defect detection method, defect detection system and storage medium
CN119313967A
Method and apparatus for unsupervised domain adaptation
US20230136609A1
Cited By
Deep learning-based multi-source heterogeneous software data defect prediction method
CN120429725A
Defect prediction method for multi-source heterogeneous software data based on deep learning
CN120429725B
Complex scene bolt data synthesis and identification method based on domain self-adaption
CN122157227A