Cross-Project Defect Prediction Method and System Based on Domain Adaptation
Through domain adaptive technology, the characteristics of the source project and the target project are mapped to the same space, and the main and auxiliary classification network training model is used to solve the problems of data distribution differences and feature applicability in cross-project defect prediction, achieving high-accuracy defect prediction in different projects.
Patent Information
- Application Number
- CN202510489029.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
There are problems with data distribution differences and feature applicability in cross-project defect prediction, which leads to inaccurate prediction, especially when sufficient historical data is lacking in new startup projects or small-scale projects, it is difficult to train an effective defect prediction model.
The cross-project defect prediction method based on domain adaptation is adopted, and the features of the source project and the target project are mapped to the same feature space through the feature mapping network, and the main classification network and auxiliary classification network are trained to reduce the distribution difference, and the model is optimized in combination with the loss function to improve prediction accuracy.
Improves the accuracy of cross-project defect prediction, enables the model to adapt effectively in different software projects and provides accurate defect prediction results, reducing resource and time costs.
Smart Images

Figure CN120011202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect prediction, and in particular, to a cross-project defect prediction method and system based on domain adaptation. Background Art
[0002] As a key research direction in the field of software engineering, software defect prediction technology predicts the number of defects mainly based on rich experimental data and experience accumulation in the early stage of software defect analysis, and is often positively correlated with the volume of software code, which helps to evaluate the possible defect ratio of software and reasonably allocate test resources accordingly. The measurement criteria for source code involve static evaluation of code data, mainly focusing on two key dimensions: the volume and complexity of the code.
[0003] In real-world application scenarios, newly launched projects often face the problem of lack of sufficient historical data, or the existing historical data is not enough to support effective prediction, which poses a major challenge to in-project defect prediction. Many new projects or small-scale projects often face the problem of insufficient data availability, which makes it difficult to train an effective internal defect prediction model. In addition, developing a high-quality defect prediction model requires a large amount of resources and time, increasing the cost and development cycle. Each software project has its unique development environment, team composition, and coding practices, which will affect the pattern of defect occurrence and increase the challenges of the adaptability and robustness of the model among different projects. To address these issues, the CPDP (Cross Project Defect Prediction) method has begun to attract attention.
[0004] The challenges faced by cross-project defect prediction are indeed diverse, but one of the most critical challenges is the data distribution difference between the source project and the target project. This difference may stem from multiple aspects, including but not limited to application domain, development process, project scale, etc. For example, a large project developed in the e-commerce field may have a completely different code structure and maintenance mode from an embedded system project. These differences will cause the features and models that perform well in the source project to be unable to be directly applied to the target project, thus affecting the prediction accuracy. In addition to the data distribution difference, the feature applicability problem is also an important factor. The effective defect prediction features found in the source project may not be directly transferable to the target project because there may be significant differences in code structure, development practices, and even the use of programming languages between different projects. These differences may cause the originally effective features to lose their prediction ability in the new environment, or their prediction ability to drop significantly. Summary of the Invention
[0005] The present invention provides a cross-project defect prediction method and system based on domain adaptation, which is used to solve the defect that the cross-project defect prediction is inaccurate due to the data distribution difference and feature applicability problem between the source project and the target project in the prior art, and to improve the accuracy of cross-project defect prediction.
[0006] The present invention provides a cross-project defect prediction method based on domain adaptation, including:
[0007] Extracting the features of the source project and the target project, and using a feature mapping network to map the features of the source project and the target project into the same feature space;
[0008] Respectively inputting the mapped features of the source project and the target project into a main classification network to obtain the first defect prediction results of the source project and the target project;
[0009] Respectively inputting the mapped features of the source project and the target project into an auxiliary classification network to obtain the second defect prediction results of the source project and the target project;
[0010] Determining a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0011] Training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project and align the features of the source project and the target project;
[0012] Extracting the features of the project to be predicted, and successively inputting the features of the project to be predicted into the trained feature mapping network and the trained main classification network to obtain the defect prediction result of the project to be predicted output by the main classification network.
[0013] The present invention also provides a cross-project defect prediction system based on domain adaptation, including:
[0014] A feature extraction module, which is used to extract the features of the source project and the target project, and use a feature mapping network to map the features of the source project and the target project into the same feature space;
[0015] A first prediction module, which is used to respectively input the mapped features of the source project and the target project into a main classification network to obtain the first defect prediction results of the source project and the target project;
[0016] A second prediction module, which is used to respectively input the mapped features of the source project and the target project into an auxiliary classification network to obtain the second defect prediction results of the source project and the target project;
[0017] A loss calculation module, configured to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0018] A model training module, configured to train the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project and align the features of the source project and the target project;
[0019] A defect prediction module, configured to extract features of a project to be predicted, and sequentially input the features of the project to be predicted into the trained feature mapping network and the trained main classification network, so as to obtain a defect prediction result of the project to be predicted output by the main classification network.
[0020] The cross-project defect prediction method and system based on domain adaptation provided by the present invention can more accurately predict cross-project defects by finely tuning the feature mapping network and the classifier, making this method a powerful tool that can not only perform excellently in intra-project prediction, but also adapt to different software projects and data distributions, and provide accurate defect prediction for software development teams. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic flowchart of the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0023] Figure 2 It is a schematic diagram of reducing the distribution difference and improving the project prediction performance in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0024] Figure 3 It is a schematic diagram of the model structure of the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0025] Figure 4 It is a schematic flowchart of the ATC2Defect method in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0026] Figure 5 It is a schematic flowchart of grammar tree feature extraction in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0027] Figure 6 It is a schematic diagram of the overall network structure of the ATC2Defect defect prediction method in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0028] Figure 7 It is a schematic diagram of class imbalance processing in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0029] Figure 8 It is a schematic diagram of the comparison of the improved cross-project prediction in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0030] Figure 9 It is a schematic diagram of the comparison between the GAT and GCN methods in the cross-project defect prediction method based on domain adaptation provided by the present invention;
[0031] Figure 10 It is a schematic diagram of the ablation analysis of cross-project prediction in the cross-project defect prediction method based on domain adaptation provided by the present invention. Detailed implementation manners
[0032] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0033] The following combines Figure 1 Describe a cross-project defect prediction method based on domain adaptation of the present invention, including:
[0034] Step 101, extract the features of the source project and the target project, and map the features of the source project and the target project to the same feature space by using a feature mapping network;
[0035] Step 102, respectively input the mapped features of the source project and the target project into the main classification network to obtain the first defect prediction results of the source project and the target project;
[0036] Step 103, respectively input the mapped features of the source project and the target project into the auxiliary classification network to obtain the second defect prediction results of the source project and the target project;
[0037] Step 104, determine the loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0038] Step 105: Train the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project and align the features of the source project and the target project.
[0039] Step 106: Extract the features of the project to be predicted, and input the features of the project to be predicted into the trained feature mapping network and the trained main classification network in sequence, so as to obtain the defect prediction result of the project to be predicted output by the main classification network.
[0040] The challenges faced by cross-project defect prediction are indeed diverse. However, one of the most critical challenges is the data distribution difference between the source project and the target project. This difference may stem from multiple aspects, including but not limited to application domain, development process, project scale, etc.
[0041] Regarding the problem of data distribution difference, domain adaptation technology provides a possible solution. Domain adaptation aims to adjust the model to adapt to the new data distribution, as Figure 2 shown, so as to improve the generalization ability of the model on the target project and solve the distribution difference problem between the source project and the target project. It will explore how to use the data of the source project to build a defect prediction model that can effectively adapt to the features of the target project. Through this method, the performance of CPDP can be improved, enabling it to accurately predict potential software defects in different project environments.
[0042] First, extract the features of the source project and the target project. In this embodiment, the method for extracting the features of the source project and the target project is not limited.
[0043] Suppose the feature representation set of the source project is ,
[0044] where represents the label in the source project. Among them, represents the feature of the i-th sample in the source project feature set, d represents the dimension of the feature, represents the corresponding label. The feature representation set of the target project is , where is the number of instance samples in the target project.
[0045] As Figure 3 shown, the first part of the model includes two key network components, namely the feature mapping network and the main classification network.
[0046] The feature mapping network aims to learn effective feature representations from the data of source projects and target projects. It provides high-quality feature representations for subsequent classification tasks. Through a deep neural network structure, the feature mapping network maps the data of source projects and target projects into a shared feature space, enabling it to better capture the commonalities and differences between the two.
[0047] Based on the feature mapping network, the main classification network is responsible for classifying the features and generating pseudo-labels for the target project. Pseudo-labels refer to the labels predicted by the model based on the knowledge it has learned for data without explicit labels. In this way, pseudo-labels can be used to represent potential defective instances in the target project. Combining with the features of the target project, these pseudo-labels can regard the target project as new, labeled data instances. This step greatly enriches the training data of the target project and improves the learning effect of the model.
[0048] The second part of the model introduces an auxiliary classification network, which is similar to a discriminator. The role of the auxiliary classification network is to identify inaccuracies in the pseudo-labels generated by the main classification network. By analyzing the pseudo-labels, it finds possible errors and feeds this information back to the main classification network. The auxiliary classification network can challenge the performance of the main classification network by maximizing the domain adaptation loss and the pseudo-label learning loss. The domain adaptation loss is used to measure the distribution difference between the source project and the target project, and the pseudo-label learning loss is used to evaluate the accuracy of the pseudo-labels.
[0049] Through this adversarial training method, the auxiliary classification network and the main classification network improve each other in continuous feedback and adjustment. The auxiliary classification network challenges the main classification network, forcing it to continuously improve the prediction accuracy for the target project. After receiving the feedback, the main classification network optimizes its classification decision, thus generating more accurate pseudo-labels. In this way, the distribution difference between the source project and the target project gradually decreases, and the generalization ability of the model is continuously enhanced. Finally, this mechanism of the two networks working together significantly improves the performance of cross-project defect prediction, enabling it to accurately predict potential software defects in different project environments.
[0050] The structures of the auxiliary classification network and the main classification network in the prediction part are the same. For example, both contain a fully connected layer and a softmax layer. But the auxiliary classification network has one more gradient reversal layer (GRL) than the main classification network.
[0051] The main classification network focuses on the classification tasks of source projects and target projects. The auxiliary classification network not only participates in classification but also achieves domain alignment through GRL. The auxiliary classification network and the classification network are independently trained, and their parameters are not shared with each other.
[0052] Before training, the features must be preprocessed. In addition, the data must be normalized. For any feature in the source or target project, normalize it according to the following formula:
[0053]
[0054] In order to effectively align the feature distributions of the source and target items, a two-layer feature mapping network is used to map the feature spaces of the two items into a shared, low-dimensional common subspace through two consecutive mapping steps. This design aims to minimize the distribution differences between the source and target items in the new feature space and improve the generalization ability and accuracy of cross-item predictions.
[0055] The mapping function of the feature mapping network can be expressed as ,in It is the feature set of the source project and the target project. The feature representation of the source item and the target item after mapping can minimize the distribution difference of the source item and the target item in the common subspace by optimizing the parameters of the mapping function.
[0056] This embodiment can more accurately predict cross-project defects by carefully tuning the feature mapping network and classifier, making this method a powerful tool that not only performs well in intra-project predictions, but can also adapt to different software projects and data distributions, providing accurate defect predictions for software development teams.
[0057] Based on the above embodiments, the loss function in this embodiment includes a domain adaptive loss function, which is determined based on the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project.
[0058] The target item instance generates pseudo labels using the classification network to ensure its accuracy and to bring the source and target items closer in the new feature space. The distribution difference between the source and target items is measured using the output of the two classification networks to obtain the domain adaptation loss function.
[0059] Based on the above embodiment, the formula of the domain adaptation loss function in this embodiment is:
[0060]
[0061] in, is the domain adaptation loss function, is a hyperparameter, n s is the number of samples in the source project, n tis the number of samples in the target project, are the parameters of the main classification network, are the parameters of the auxiliary classification network, is the feature after mapping the i-th sample in the source project, is the feature after mapping the i-th sample in the target project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, which is the defect prediction probability; is the first defect prediction result of the i-th sample in the target project output by the main classification network, which is the defect prediction probability; is the second defect prediction result of the i-th sample in the source project output by the auxiliary classification network, which is the defect prediction probability; is the second defect prediction result of the i-th sample in the target project output by the auxiliary classification network, which is the defect prediction probability.
[0062] is an adjustable hyperparameter used to adjust the difference between the two, that is, the proportion of the source project. For the auxiliary classification network, the domain adaptation loss function should be maximized during training. The reason for this design is to make the feature representations of the source domain and the target domain as aligned as possible during the training process, thereby improving the performance and generalization ability of the model in the target domain.
[0063] Based on the above embodiments, in this embodiment, the loss function further includes one or more of the supervised contrastive learning loss function of the source project, the classification loss function of the main classification network, and the pseudo-label loss function generated by the main classification network for the target project;
[0064] The supervised contrastive learning loss function of the source project is determined according to the features after mapping the source project and the target project, and the defect labels of the source project;
[0065] The classification loss function of the main classification network is determined according to the defect labels of the source project and the first defect prediction result;
[0066] The pseudo-label loss function is determined according to the first defect prediction result and the second defect prediction result of the target project.
[0067] To make full use of the label information of labeled instances to mine discriminative information, the supervised contrastive learning loss can be used , whose main goal is to enhance the discriminative ability of the model by maximizing the similarity between similar samples and minimizing the similarity between dissimilar samples. Compared with the traditional cross-entropy loss, the supervised contrastive learning loss can more effectively cluster samples of the same class in the embedding space while separating samples of different classes. This method introduces more information for supervision, especially performs well in multi-class classification tasks, and can significantly improve the generalization ability and robustness of the model.
[0068] For the main classification network, in order to guide its learning, a classification loss function is used.
[0069] The pseudo-labels of the target project instances are generated by the classification network. In order to improve the accuracy rate, a pseudo-label loss function is used.
[0070] Based on the above embodiments, the calculation formula of the supervised contrastive learning loss function in this embodiment is:
[0071]
[0072] Where, is the supervised contrastive learning loss function, I is the set of labeled samples of the source project, i is the index of the anchor sample, P(i) is the set of positive samples of the same class as the anchor sample i in the source project, and A(i) is the other samples in the source project and the target project except the anchor sample i. 、 、 respectively represent the vectors of the anchor sample i, the positive sample p, and the other sample a after feature mapping. Sim is the similarity algorithm, usually calculated using dot product or cosine similarity; is the temperature parameter, used to control the scaling of the similarity score and affect the smoothness of the probability distribution;
[0073] The calculation formula of the classification loss function is:
[0074]
[0075] Where, is the classification loss function, n s is the number of samples in the source project, is the parameter of the main classification network, is the feature after mapping of the i-th sample in the source project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the defect label of the i-th sample in the source project;
[0076] The calculation formula of the pseudo-label loss function is:
[0077]
[0078] Among them, is the pseudo-label loss function, and n t is the number of samples in the target project, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the target project output by the auxiliary classification network. is the logarithm of the reverse probability predicted by the auxiliary classification network. By designing like this, it is to minimize the probability that the samples in the target project are misclassified. By minimizing this loss function, the main classification network is forced to have an output opposite to that of the auxiliary classification network to force the classification network to find a better decision boundary for the features in the common space, thereby enhancing the accuracy of the pseudo-labels. This definition method also introduces the idea of adversarial learning. The two networks compete with each other, restricting each other by maximizing and minimizing different objectives, and improving the generalization ability of the model.
[0079] Based on the above embodiments, in this embodiment, training the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function includes:
[0080] Determine a first loss value according to the domain adaptation loss function, the supervised contrastive learning loss function, the classification loss function, and the pseudo-label loss function;
[0081] Adjust the parameters of the feature mapping network and the main classification network to minimize the first loss value;
[0082] Determine a second loss value according to the domain adaptation loss function and the pseudo-label loss function;
[0083] Adjust the parameters of the auxiliary classification network to maximize the second loss value.
[0084] When constructing the feature mapping network and the classification network, the goal is to minimize their losses in the standard classification task. However, when introducing the auxiliary classification network for adversarial training, it is hoped that this auxiliary network can perform as well as possible in distinguishing the data after feature mapping, even if these data are generated from different distributions. Therefore, for the auxiliary classification network, the goal is to maximize its loss in distinguishing these feature mapping data. So the adversarial loss is defined as:
[0085] ;
[0086] ;
[0087] Among them, is the classification loss function, is the supervised contrastive learning loss function, is the pseudo-label loss function, is the domain adaptation loss function, are the parameters of the feature mapping network, are the parameters of the main classification network, are the parameters of the auxiliary classification network. Without a triangle above the parameter, it means that the parameter is currently being optimized. With a triangle above the parameter, it means that the parameter is currently fixed and not updated. and are the coefficients used to balance several losses.
[0088] The process pseudo-code of the method provided in this embodiment can be seen in Algorithm 1. The specific steps include:
[0089] Step 1: The feature mapping network is responsible for mapping the data of the source project and the target project into a shared feature space, learning the common differences between the two, and supervising the contrastive loss.
[0090] Step 2: Based on the shared features generated by the feature mapping network, the main classification network is responsible for classifying the data and generating pseudo-labels for the target project.
[0091] Step 3: The auxiliary classification network further optimizes the pseudo-labels, acts as a discriminator, evaluates the quality of the pseudo-labels, and challenges the main classification network.
[0092] Step 4: Through the adversarial training between the main classification network and the auxiliary classification network, continuously optimize the pseudo-labels and classification decisions.
[0093] Algorithm 1: Cross-project defect prediction method based on domain adaptation
[0094] Require: Source project feature vector and label , target project feature vector
[0095] Ensure: Target project prediction label
[0096] 1: Initialize network parameters , , and hyperparameters , , ,
[0097] 2: while not reaching the maximum number of iterations or convergence conditions do
[0098] 3: Map the data of the source project and the target project to a common subspace:
[0099] 4: Calculate the supervised contrastive loss
[0100] 5: Calculate the classification loss
[0101] 6: Calculate the domain adaptation loss
[0102] 7: Calculate the pseudo-label loss
[0103] 8: Update the feature mapping and classification network parameters by minimizing the formula loss
[0104] 9: Update the parameters of the auxiliary classification sub-network by maximizing the formula loss
[0105] 10: end while
[0106] 11: return the target project prediction
[0107] In the initial state, the parameters of the feature mapping sub-network, the classification sub-network, and the auxiliary classification sub-network are all random. The pull-in stage model attempts to reduce the distribution difference between the source and target domains to improve the generalization ability and improve the generation of pseudo-labels. Through the gradient information of minimizing the formula loss, it is backpropagated to the feature mapping sub-network and the classification sub-network to adjust their parameters. The classification sub-network uses the updated parameters to predict the unlabeled data and generate pseudo-labels.
[0108] During the push stage, the auxiliary classification sub-network attempts to maximize the domain adaptation loss and the pseudo-label loss , which means it attempts to find the decision boundary that can best distinguish the feature representations between the source and target domains. Measure the distribution difference between the two domains according to the outputs of the two classification networks, and update the parameters of the auxiliary classification network through the gradient of maximizing the formula loss.
[0109] The promotion phase and the pulling - closer phase alternate until a predetermined number of training iterations or convergence conditions are reached. In each iteration, the parameters of the model are continuously adjusted to achieve a balance between maximizing the domain - adaptation loss and minimizing the pseudo - label loss. This process forms a dynamic adversarial cycle. Through continuous iteration, the model gradually improves its performance on unlabeled data while enhancing its adaptability to the distribution differences between the source and target domains, thereby improving the performance of cross - project defect prediction.
[0110] Based on the above - mentioned embodiments, in this embodiment, extracting the features of the source project and the target project includes:
[0111] Converting the source project or the target project into an abstract syntax tree, and processing the abstract syntax tree based on a convolutional neural network to obtain the features of the abstract syntax tree;
[0112] Constructing a class - dependency network according to the dependency relationships in the source project or the target project, and obtaining the features of the class - dependency network based on a network - embedding algorithm;
[0113] Fusing the code metrics of the source project or the target project with the features of the class - dependency network, and then fusing the result with the features of the abstract syntax tree again;
[0114] Embedding the features after the second fusion into the class - dependency network, constructing a graph - attention network, and obtaining graph - node features.
[0115] The PROMISE library is a public data repository specifically designed to support empirical research in the field of software engineering. It provides researchers with rich datasets related to software defect prediction. The source project and the target project can select datasets of 7 classic software projects publicly available in the PROMISE library, totaling 21 different versions of data. These datasets have been widely used in the research field of software defect prediction. These datasets contain rich historical data, providing a solid foundation for reliable data analysis and model construction. Table 1 provides a detailed description of these datasets, including key information such as the source, version information, number of instances, and defect rate of each dataset.
[0116] Table 1 Datasets
[0117]
[0118] The datasets of these projects are highly representative, covering different types of software projects and being able to represent various application scenarios and code structures. These datasets are not only applied in this embodiment but also involved in multiple fields such as data management, search, and text processing, providing empirical research cases for this embodiment.
[0119] This embodiment adopts multi-view feature extraction, which includes traditional code metrics, Abstract Syntax Tree (AST) feature extraction, and Class Dependency Network (CDN) feature extraction.
[0120] (1) Traditional software features include two major categories: one is static metrics based on structure, mainly including LOC metrics, McCabe metrics, Halstead metrics, etc. The other is CK metrics based on object-oriented features. LOC metric is a method to measure software size and complexity, but this metric also has certain limitations because it cannot reflect code quality and functionality. McCabe metric measures the complexity of software code based on the number of nodes and edges in the program graph, which can help developers identify potential logical errors and code paths that are difficult to test. Halstead metric is used to measure software complexity, based on the number of operators and operands and their different combinations. These metrics include the length, volume, difficulty, and complexity of the program, etc. CK metric is mainly used to measure the complexity of object-oriented software. This metric analyzes aspects such as the complexity, inheritance structure, coupling degree, and cohesion of object-oriented software, so as to guide software design and optimization.
[0121] (2) Abstract Syntax Tree (AST). In the software development life cycle, the compilation process of source code is a key step in converting the high-level language code written by developers into machine code that can be understood and executed by a computer. This process usually includes several consecutive stages, and each stage analyzes and transforms the source code at different levels.
[0122] (3) Class Dependency Network features. In the inventive method, various features related to software quality and potential defects can be extracted by analyzing the Class Dependency Network. In addition to feature extraction, the Class Dependency Network can also be used for software metrics and defect localization. By measuring various attributes and relationships in the CDN, the development team can obtain quantitative information about the health status of the software system. When a defect occurs, the CDN can help quickly locate the possible problem modules, thus accelerating the defect repair process. In addition, the CDN also plays an important role in software refactoring. It can help identify the modules that need to be refactored and possible refactoring paths, thereby improving the maintainability and scalability of the software system.
[0123] When modeling a software dependency network, usually only the single network structure information is considered, while the rich feature information contained in each node of the class dependency network is ignored. This method is not sufficient to fully utilize code features and may lead to a decline in prediction performance. Node feature information includes code complexity, change history, developer information, etc. These information are of great significance for accurately predicting defects. In addition, the syntactic and semantic information of the code module where each node is located is also crucial. Defect prediction by combining multi-view information and node features is expected to significantly improve the performance and application value of the model. By integrating multi-view features, more comprehensive code semantic and structure information can be captured, improving the accuracy of prediction. Making full use of the rich feature information of nodes can further improve the model's ability to identify defects, ultimately promoting the development and maintenance of high-quality software.
[0124] As Figure 4 shown, the process of the ATC2Defect method includes: (1) Combining the traditional features of the dataset and the class dependency network embedding features. (2) Parsing the source code into an abstract syntax tree, extracting a token sequence through specific rules, then embedding the token sequence into a matrix, and extracting features through a CNN. (3) Using the class dependency network as the data of the last view, fusing the features obtained in the previous two steps and then embedding them into the class dependency network to construct a GCN (Graph Convolution Network) model. (4) Performing defect prediction through the output of the graph convolution network.
[0125] In this embodiment, a feature of the abstract syntax tree is obtained by using an AST traversal algorithm based on primary node priority. The source code is constructed into an abstract syntax tree to extract features. By utilizing the abstract syntax tree, the structure and syntax information of the code can be captured. And by extracting the nodes and subtrees in the syntax tree, detailed information about the code structure can be obtained, which helps to identify potential defects in the code. In order to extract a node sequence containing important semantic information from the abstract syntax tree, this paper proposes a traversal algorithm based on primary node priority (Primary Node Priority Traversal Algorithm, PNPT). This method has the following advantages and characteristics: The algorithm introduces the concept of a primary node and preferentially processes the node with the largest subtree during traversal. This can ensure that as much important semantic information as possible is retained in the generated token sequence, reducing the possibility of information loss. By calculating the subtree size of each node, the algorithm can dynamically adjust the traversal order. Regardless of how complex the code structure is, the algorithm can flexibly adapt to ensure that key nodes are processed first. The primary node priority selection strategy can maximize the retention of key information in the code. This has a significant effect on subsequent tasks based on code semantic information. For a token sequence exceeding the maximum length, the algorithm can intelligently prioritize retaining key nodes, thus avoiding the loss of important information due to truncation.
[0126] After generating the token sequence, each source file is converted into a set of a series of tokens. To input the data into the neural network, the token set needs to be mapped, and each token is mapped to an integer, so that each source code file is converted into an integer vector. Considering that the token lengths produced by each file are different, it is necessary to pad the shorter vectors with zeros to ensure that the final vector lengths generated by each file are the same. Consider a project named P that contains n Java source files, denoted as ... Through the token extraction method, a series of token tags are extracted from these files, denoted as the set ... Let the input word index sequence be ... where n represents the sequence length. Assume the embedding matrix is ... where V represents the vocabulary size and d represents the word embedding dimension. Then the output of the embedding layer is calculated as follows:
[0127]
[0128] where, represents the word embedding vector corresponding to the word ...
[0129] Assume the input embedding vector sequence is ... where Denote the word embedding vector. Let the convolution kernel be , where represents the convolution kernel size, and based on multi-view feature learning, the software defect prediction represents the number of convolution kernels or output channels. The output of the convolution layer is calculated as follows:
[0130]
[0131] where, represents the segment from the i-th word to the -th word in the input sequence, b represents the bias vector, and ReLU represents the rectified linear unit activation function.
[0132] For each feature map , take the maximum pooled eigenvalue, , and the pooled feature vector is . Transform the pooled feature vector through a fully connected layer, and finally obtain the syntax tree feature . The entire process of syntax tree feature extraction is as Figure 5 shown.
[0133] .
[0134] The class dependency network shows the dependency relationships between classes in the form of a graph, including inheritance, interface implementation, aggregation, or method calls, etc. After constructing the class dependency network, it is necessary to obtain the embedding representation of each node through network embedding technology. The network embedding vector can capture the structural information and semantic information of the node in the network, making the initial attributes of each node more expressive. This representation can better reflect the role and relationship of the node in the entire graph. In this embodiment, in the modeling of the class dependency network, the node2vec algorithm is first used to perform network embedding on the class dependency network to capture the structural and semantic information in the network. The node2vec algorithm explores the local and global structures in the network by adjusting two parameters p and q in the random walk.
[0135] It can be set that p = 2 and q = 0.15. Such a parameter selection aims to balance the node homogeneity and structural role in order to better reveal the similarities and differences between nodes. In this way, an embedding vector is generated for each node in the CDN:
[0136]
[0137] Next, traditional software engineering features, such as code complexity, coupling degree, etc., are normalized to eliminate the influence brought by different dimensions. Normalization can be carried out through the following formula:
[0138]
[0139] Among them, is the i-th feature in the original feature vector, is the i-th feature in the normalized feature vector. Normalization ensures that all feature values are within the range of 0 to 1. After that, the normalized traditional feature vector is concatenated with the network embedding vector to form an enhanced feature representation .
[0140] To further reduce the dimension of the features and retain the most important information, this embodiment adopts the principal component analysis dimensionality reduction technique. The original features are transformed into a new coordinate system through a linear transformation, such that the data variance on the first coordinate axis in this new coordinate system is the largest, and each subsequent coordinate axis is orthogonal to the previous one and has the largest variance. In this way, a low-dimensional feature representation can be obtained . Finally, the above features are combined with the features obtained from the abstract syntax tree to obtain the feature X of the graph node. The formula is as follows:
[0141]
[0142] Considering that the contributions from the two types of features are different, another combination method with weight parameters can also be adopted. The formula is as follows:
[0143]
[0144] Among them, is the combination method. In this study, three methods of element-wise addition (add), element-wise multiplication (multiply), and concatenation (concat) are selected.
[0145] Following the graph convolutional network architecture and its parameter settings proposed in Kipf's research, a two-layer convolutional structure is also adopted. Using the node attribute feature matrix X obtained in the previous step, the final output result can be calculated through the following formula.
[0146]
[0147] In this formula, is the weight between the hidden layer and the input layer, represents the number of hidden layer units, represents the feature dimension is the weight coefficient between the hidden layer and the output layer. is the preprocessed adjacency matrix, is the degree matrix of the nodes, is the normalized adjacency matrix, is the adjacency matrix of the graph, with size N×N, is the identity matrix, with size N×N. The output feature matrix has a dimension of , and the ReLU function is used as the non-linear activation function in the network.
[0148] The final predicted output of the network is:
[0149]
[0150] where W (2) is the weight matrix of the last layer (output layer) of the model, which is used to map the feature matrix F to the final classification space.
[0151] The loss function is:
[0152]
[0153] where is the true label, represents the predicted probability after the softmax function of the prediction layer.
[0154] The overall structure of the ATC2Defect defect prediction method is as shown in Figure 6 . This method combines the Abstract Syntax Tree (AST), Traditional Code Metrics, and Class Dependency Network, so this method is called ATC2Defect.
[0155] The graph structures and identifiers in the syntax trees between different projects vary greatly, which will have a negative impact on cross-project prediction. Therefore, when using the ATC2Defect method for cross-project defect prediction, some improvements need to be made in these aspects.
[0156] First, when embedding the abstract syntax tree, instead of directly using identifiers as in the in-project prediction task, the type of the current node should be used. Identifiers usually have a high degree of project specificity because they are arbitrarily named by developers according to project requirements and coding styles. Identifiers that are valid in the source project may be completely different in the target project, which will lead to overfitting of the model to the source project data and thus reduce the generalization ability of the model in the target project. Compared with identifiers, node types are more general. The AST node type represents the syntax structure and logical relationship of the source code, rather than specific implementation details. This generality enables the model to learn more abstract code structure features without relying on specific identifiers. In this way, overfitting can be reduced and more general features can be learned.
[0157] Secondly, when dealing with graph structures, the Graph Attention Network (GAT) with better generalization performance should be selected. Compared with GCN, GAT has the following advantages in generalization ability: on the one hand, in GCN, the feature representation of nodes is aggregated through the average features of neighbor nodes, and this method cannot distinguish the importance of different neighbor nodes. While GAT can adaptively assign different weights to each neighbor node, making important neighbor nodes contribute more to the feature representation of the central node, thus being able to better capture the key information in the graph structure. On the other hand, GAT uses multiple independent attention heads, and each head can capture different patterns in the graph structure. By combining these patterns together, GAT can improve the expression ability and generalization ability of the model. The formula of GAT is as follows:
[0158]
[0159] Among them, represents the attention coefficient between node and node , is a trainable parameter vector, is the trainable weight matrix, is the vector concatenation operation, and are the feature vectors of node and node respectively. Next, the attention weight can be obtained through normalization:
[0160]
[0161] Finally, the feature representation of node is updated as:
[0162]
[0163] Among them, represents the set of neighbor nodes of node , is the non-linear activation function. By introducing the graph attention network, it is possible to better handle the graph structure differences between different projects and improve the accuracy and robustness of cross-project defect prediction. In addition, when dealing with abstract syntax trees, using node types instead of specific identifiers can significantly reduce the overfitting phenomenon of the model, thereby improving the generality of features and the generalization ability of the model.
[0164] This embodiment first extracts and characterizes the code data of the source project and the target project through the ATC2Defect network to obtain two feature sets: the feature set of the source project and the feature sets of the target project . These two feature sets respectively capture the essential features of the source project and the target project code, providing a basis for subsequent defect prediction. Further combining the label information of the source project , replace the original classifier with Figure 3 the network structure shown
[0165] In this embodiment, cross-project defect prediction will be performed using ATC2Defect, and improvements will be made on this basis to improve the performance of cross-project defect prediction. It will focus on how to further improve the performance of ATC2Defect in cross-project defect prediction, expecting ATC2Defect to better adapt to diverse software projects, provide more accurate defect prediction for software development teams, thereby helping them identify and fix potential software defects in advance, and improve software quality and development efficiency.
[0166] Based on the above embodiments, after extracting the features of the source project and the target project in this embodiment, it further includes:
[0167] For a sample set in which the number of samples of any defect category in the source project or the target project is less than a preset threshold, determine the density of each sample in the sample set;
[0168] Based on the density function, determine the number of neighbors of each sample according to the density of each sample;
[0169] Select the nearest neighbor samples of each sample according to the number of neighbors of each sample, and calculate the weights of each nearest neighbor sample;
[0170] Randomly select neighbor samples from the nearest neighbor samples of each sample to generate synthetic samples together with each sample;
[0171] Adjust the synthetic samples using the weights of the randomly selected neighbor samples to obtain new samples.
[0172] Since the data for defect prediction has class imbalance, it is necessary to resample the data. However, for graph structures, data sampling will destroy the original graph structure. On the other hand, for cross-version prediction and cross-project prediction, the graph structures of the two data sets will be completely different due to the addition of new instances. Therefore, the model of the method of the present invention will be divided into two parts: a pre-training part and a downstream task. Through pre-training, the feature F of each project is obtained, and the downstream task then samples these features to handle the class imbalance problem, as Figure 7 shown, and the specific steps include:
[0173] Step 1: Integrate the syntax tree features (AST), traditional metric features, and dependency network features to obtain comprehensive multi-view features.
[0174] Step 2: After ATC2 fuses the three features, it preprocesses the features and divides the training data and test data.
[0175] Step 3: There may be a problem of data imbalance in the training data, and data sampling is required and then combined with the test data as a new data set.
[0176] Step 4: After processing the data, input it into the classifier model for training.
[0177] Step 5: The classifier predicts each sample category based on the features in the test data set to determine whether there are defects (Bugs) or no defects (Clean).
[0178] This phased model construction method can not only effectively handle the class imbalance problem, but also meet the needs of cross-version and cross-project prediction, improving the generalization ability and prediction accuracy of the model in different software projects.
[0179] To solve the problems introduced by noise and unreasonable samples, this embodiment proposes a density-weighted SMOTE algorithm, named DW-SMOTE algorithm. See the pseudocode of Algorithm 2. Its main feature is to dynamically select the number of neighbors, which is related to the local density of each sample. The most suitable number of neighbors is determined through a density function to ensure that the true distribution of the data can be better reflected when generating synthetic samples. The algorithm first calculates the density of each minority-class sample and dynamically adjusts the number of neighbors according to the density, and then calculates the weight of each neighbor sample. The weight is used to guide the subsequent sample generation process. When generating synthetic samples, DW-SMOTE not only randomly selects neighbors, but also adjusts the position of the synthetic samples according to the weights of the neighbors, thus generating more diverse and high-quality synthetic data. Through the noise filtering step, possible abnormal samples are further removed, improving the reliability of the synthetic data. Through this method, DW-SMOTE can better adapt to the local distribution characteristics of the data, improving the performance and stability of the classification model.
[0180] Algorithm 2: Upsampling Algorithm DW-SMOTE Based on Density Weight
[0181] Require: Minority-class sample set S, oversampling rate N
[0182] Ensure: Increased minority-class sample set S new
[0183] 1: Initialize S new ←
[0184] 2: for each sample x i ∈ S do
[0185] 3: Calculate the density ρ of the sample x i of i
[0186] 4: Dynamically select the number of neighbors k i ← f(ρ i ), where f is the density function
[0187] 5: Select the k nearest neighbor samples of x i of i
[0188] 6: Initialize the sample weights ← []
[0189] 7: For each neighbor sample x i,j do
[0190] 8: Calculate the sample weight w i,j
[0191] 9: weights.append(w i,j )
[0192] 10: end for
[0193] 11: end for
[0194] 12: For each sample x i ∈ S do
[0195] 13: For n = 1 to N do
[0196] 14: Randomly select a neighbor sample x i,j
[0197] 15: Generate a synthetic sample x new ← x i + θ · (x i,j - x i ), where θ ~ U(0,1)
[0198] 16: Update x according to the weight w i,j new
[0199] 17:S new .append(x new )
[0200] 18: end for
[0201] 19: end for
[0202] 20: Filter the noise of S new and remove abnormal samples
[0203] 21: return S new
[0204] As Figure 3 shown, the overall framework is divided into two parts for adversarial training: the first part is the feature mapping network and the classification network. The feature mapping network consists of two fully connected neural networks with 64 and 32 neurons respectively, and the activation function is ReLU. The classification network is a classification network composed of a single fully connected layer, equipped with two output nodes, and uses the softmax function for output conversion. The second part includes an auxiliary classification network with the same structure as the classification network. To optimize the prediction ability of the model, the present invention adopts a grid search method to determine a set of hyperparameters that can maximize the model prediction performance. Grid search is a systematic hyperparameter tuning strategy that searches for the best hyperparameter settings by traversing all possible combinations in a predefined set of hyperparameters. In this study, the values of [] are finally determined to be [0.95, 1.0, 2.0, 0.07]. To ensure the smooth progress of model training and the need to process large-scale datasets, the experiment is executed on a computer with 32GB of body memory. The experiment uses an NVIDIA GeForce GTX 1080Ti graphics card, and the framework is PyTorch.
[0205] The present invention names the improved cross-project prediction method as , and Table 2 and Table 3 respectively show the comparison results of the improved method in this paper and the comparison method under the F1 and AUC (Area Under Curve) metrics. The values shown in the table are the averages of cross-project predictions with the current project as the target project and the rest of the datasets as the source projects.
[0206] From Table 2, the method shows a significantly improved F1 score in some projects (such as ant and poi), which indicates that for specific projects, this method can effectively improve the prediction accuracy. On average, the average F1 score in all projects is 0.522, which is higher than other methods, indicating that in the cross-project prediction scenario, the improved method has better overall performance.
[0207] Table 2 Improved Cross-Project Prediction F1
[0208]
[0209] According to Table 3, in terms of the AUC metric, It also performs excellently, especially in the poi and velocity projects, showing a significant improvement compared to other methods, which emphasizes its effectiveness in distinguishing defect classes from non-defect classes. The AUC values also show the performance differences between different projects, which may be related to the specific characteristics of the projects and the defect distribution. For example, the relatively low AUC value of the velocity project may be because the data characteristics of this project make defect prediction more difficult. On average, 's method has an average AUC value of 0.665 in all projects, exceeding all other comparison methods. This indicates that this method performs excellently in balancing the prediction capabilities of positive and negative samples and has better generalization ability.
[0210] Table 3 Improved cross-project prediction AUC
[0211]
[0212] As Figure 8 shown in Figure (a) and Figure (b) therein, the abscissa represents the project names in Table 2 and Table 3, and the coordinate points corresponding to each project name are the F1 scores or ACU metrics of different methods. The improved method for cross-project defect prediction in the present invention has a significant improvement compared to the original method. In terms of the F1 score, the average performance has increased by 12.5%, and in terms of the AUC metric, the average performance has increased by 18.3%. At the same time, compared with the method with the best performance in the benchmark model, the average increases in the two metrics are 9.4% and 12.7% respectively. These results indicate that through carefully designed feature mapping and adversarial training strategies, the data distribution differences between the source project and the target project can be effectively reduced, and at the same time, the limited label information can be fully utilized to improve the prediction ability of the model.
[0213] To verify the performance differences of using GCN and GAT respectively in the improved method, the data of seven projects were trained and tested. The evaluation metrics include the F1 score and the AUC value, both of which are used to measure the classification performance of the model. The experimental results are as Figure 9 shown in Figure (a) and Figure (b) therein. Among them, the abscissa is the project names in Table 2 and Table 3, and the coordinate points corresponding to each project name are the F1 scores or ACU metrics of the GCN and GAT methods.
[0214] On all datasets, the F1 score of GAT has increased by an average of 3%. This indicates that GAT performs better in accurately predicting defects and reducing false positives. In terms of the AUC value, GAT also shows a significant advantage, with an average increase of approximately 4%. In summary, the experimental results show that the performance of GAT in cross-project defect prediction is significantly better than that of GCN. Especially when dealing with projects with large differences in data distribution, GAT can better adapt to different project characteristics, capture complex relationships in the graph structure, and integrate information of different patterns, improving the model's expressive ability and generalization ability. In the case of large node feature noise, GAT can ignore irrelevant features through the attention mechanism, thereby reducing the overfitting phenomenon and improving the accuracy and reliability of prediction.
[0215] An ablation analysis was conducted in this section to evaluate the impact of different components in the improved method on the performance of cross-project defect prediction. To explore the influence of three different loss functions on the method of the present invention and gain an in-depth understanding of the contribution of each component to the overall performance, a series of comparative experiments were designed. In each experiment, a specific loss function was removed each time, and its impact on the model performance was observed. First, the complete method containing all three loss functions was simply referred to as ATC. To evaluate the contribution of the domain adaptation loss, the domain adaptation loss was removed, and this variant was called ATC-da. To evaluate the influence of the supervised contrast loss, the supervised contrast loss was removed, and this variant was called ATC-sc. Finally, to evaluate the role of the pseudo-label learning loss, the pseudo-label loss was removed, and this variant was called ATC-pl. Table 4 and Figure 10 shows the results of the above ablation experiments. By comparing the performance of the model under different combinations of loss functions, it is possible to more clearly understand the importance of each component in the improved method. Figure 10 The abscissa in is the project names in Table 2 and Table 3, and the coordinate points corresponding to each project name are the F1 score or ACU metrics of different methods.
[0216] Table 4 Cross-project prediction ablation analysis
[0217]
[0218] The corresponding F1 scores and AUC values were recorded to evaluate the specific impact of each component. The experimental results show that removing any loss function will lead to a decline in the model performance, demonstrating the importance of each loss in the model training process. Specifically, when removing the domain adaptation loss (i.e., the ATC-da model), the average F1 score of the model decreased from 0.522 to 0.467, and the average AUC value decreased from 0.665 to 0.617. This result indicates that plays a key role in promoting the transfer learning of the model between different projects and enhancing its generalization ability. Similarly, the supervised contrast loss Removal (i.e., the ATC-sc model) also leads to a performance decline, although its impact is slightly smaller than that of removing . This reflects the 's contribution to optimizing the model's ability to distinguish different types of defects, especially when facing an imbalanced dataset. Removing the pseudo-label loss results in a performance decline that lies between removing and . This indicates that although plays a positive role in semi-supervised learning using unlabeled data, its improvement effect on the model performance is less than that of .
[0219] In summary, these ablation experiments not only demonstrate the effectiveness of the overall method but also reveal the importance and mechanism of each component. In particular, the domain adaptation loss plays a crucial role in improving the generalization ability of the model, while the supervised contrastive loss and the pseudo-label loss optimize the model's class discrimination ability and the utilization of unlabeled data respectively, jointly promoting the performance improvement of the model in the cross-project defect prediction task.
[0220] This invention deeply explores and analyzes the application of the improved model of ATC2Defect in the field of cross-project defect prediction. Through a series of experiments, the model's ability to perform transfer learning between different projects is evaluated, and various feature selection and engineering strategies are attempted to improve the model's generalization performance. By improving the abstract syntax tree embedding strategy, introducing the graph attention network, and using domain adaptation technology to bridge the data distribution differences between the source project and the target project, the performance of cross-project defect prediction is improved. Through experiments, this invention demonstrates the prediction performance of the improved method on multiple projects and compares it with other methods. The results show that the improved α-ATC(CPDP) version achieves the best F1 score and AUC value in most projects, with an approximate 12.5% and 18.3% increase in F1 value and AUC value respectively compared with the original method, and an average improvement of 9.4% and 12.7% compared with several baseline methods.
[0221] The following describes the cross-project defect prediction system based on domain adaptation provided by this invention. The cross-project defect prediction system based on domain adaptation described below can be correspondingly referred to the cross-project defect prediction method based on domain adaptation described above.
[0222] The system includes:
[0223] A feature extraction module for extracting the features of the source project and the target project and mapping the features of the source project and the target project to the same feature space using a feature mapping network;
[0224] A first prediction module, configured to respectively input the features after mapping the source project and the target project into a main classification network to obtain first defect prediction results of the source project and the target project;
[0225] A second prediction module, configured to respectively input the features after mapping the source project and the target project into an auxiliary classification network to obtain second defect prediction results of the source project and the target project;
[0226] A loss calculation module, configured to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project;
[0227] A model training module, configured to train the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project, and align the features of the source project and the target project;
[0228] A defect prediction module, configured to extract the features of the project to be predicted, and sequentially input the features of the project to be predicted into the trained feature mapping network and the trained main classification network to obtain the defect prediction result of the project to be predicted output by the main classification network.
[0229] In this embodiment, by finely tuning the feature mapping network and the classifier, cross-project defects can be predicted more accurately, making this method a powerful tool that not only performs excellently in intra-project prediction, but also can adapt to different software projects and data distributions, providing accurate defect prediction for software development teams.
[0230] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-project defect prediction method based on domain adaptation, characterized in that, including: extracting the features of the source project and the target project, and mapping the features of the source project and the target project into the same feature space by using a feature mapping network; inputting the mapped features of the source project and the target project into a main classification network respectively to obtain the first defect prediction results of the source project and the target project. Based on the feature mapping network, the main classification network is responsible for classifying the features and generating pseudo-labels for the target project; inputting the mapped features of the source project and the target project into an auxiliary classification network respectively to obtain the second defect prediction results of the source project and the target project. The structures of the auxiliary classification network and the main classification network in the prediction part are the same, but the auxiliary classification network has one more gradient reversal layer than the main classification network; determining a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project; the auxiliary classification network challenges the performance of the main classification network by maximizing the domain adaptation loss and the pseudo-label learning loss. The domain adaptation loss is used to measure the distribution difference between the source project and the target project, and the pseudo-label learning loss is used to evaluate the accuracy of the pseudo-labels; the adaptive loss function is determined according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project; the pseudo-label loss function is determined according to the first defect prediction result and the second defect prediction result of the target project; training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project and align the features of the source project and the target project; extracting the features of the project to be predicted, and sequentially inputting the features of the project to be predicted into the trained feature mapping network and the trained main classification network to obtain the defect prediction result of the project to be predicted output by the main classification network; the formula of the domain adaptation loss function is: ; Among them, is the domain adaptation loss function, is a hyperparameter, n s is the number of samples in the source project, n t is the number of samples in the target project, are the parameters of the main classification network, are the parameters of the auxiliary classification network, is the feature after mapping the i-th sample in the source project, is the feature after mapping the i-th sample in the target project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the source project output by the auxiliary classification network, is the second defect prediction result of the i-th sample in the target project output by the auxiliary classification network.
2. The cross-project defect prediction method based on domain adaptation according to claim 1, wherein the loss function further includes one or more of the supervised contrast learning loss function of the source project and the classification loss function of the main classification network; the supervised contrast learning loss function of the source project is determined according to the mapped features of the source project and the target project and the defect label of the source project; the classification loss function of the main classification network is determined according to the defect label and the first defect prediction result of the source project.
3. The cross-project defect prediction method based on domain adaptation according to claim 2, characterized in that the calculation formula of the supervised contrast learning loss function is: ; Among them, is the supervised contrastive learning loss function, I is the set of labeled samples of the source project, i is the index of the anchor sample, P(i) is the set of positive samples of the same class as the anchor sample i in the source project, and A(i) is the set of other samples in the source project and the target project except the anchor sample i. , , respectively represent the vectors of the anchor sample i, the positive sample p, and the other sample a after feature mapping, sim is the similarity algorithm, is the temperature parameter; the calculation formula of the classification loss function is: ; Among them, is the classification loss function, and n s is the number of samples in the source project, are the parameters of the main classification network, is the feature after mapping of the i-th sample in the source project, is the first defect prediction result of the i-th sample in the source project output by the main classification network, is the defect label of the i-th sample in the source project; the calculation formula of the pseudo-label loss function is: ; Among them, is the pseudo-label loss function, and n t is the number of samples in the target project, is the first defect prediction result of the i-th sample in the target project output by the main classification network, is the second defect prediction result of the i-th sample in the target project output by the auxiliary classification network.
4. The cross-project defect prediction method based on domain adaptation according to claim 2, wherein training the feature mapping network, the main classification network and the auxiliary classification network according to the loss function includes: determining a first loss value according to the domain adaptation loss function, the supervised contrast learning loss function, the classification loss function and the pseudo-label loss function; adjusting the parameters of the feature mapping network and the main classification network to minimize the first loss value; determining a second loss value according to the domain adaptation loss function and the pseudo-label loss function; Adjust the parameters of the auxiliary classification network to maximize the second loss value.
5. The cross-project defect prediction method based on domain adaptation according to claim 4, wherein, Determine the first loss value according to the domain adaptation loss function, the supervised contrastive learning loss function, the classification loss function, and the pseudo-label loss function through the following formula: ; Determine the second loss value according to the domain adaptation loss function and the pseudo-label loss function through the following formula: ; Among them, is the classification loss function, is the supervised contrastive learning loss function, is the pseudo-label loss function, is the domain adaptation loss function, are the parameters of the feature mapping network, are the parameters of the main classification network, are the parameters of the auxiliary classification network. Without a triangle above the parameter indicates that the parameter is currently being optimized, and with a triangle above the parameter indicates that the parameter is currently fixed and not updated, and are the balance coefficients.
6. The cross-project defect prediction method based on domain adaptation according to any one of claims 1-5, characterized in that Extract the features of the source project and the target project, including: Convert the source project or the target project into an abstract syntax tree, and process the abstract syntax tree based on a convolutional neural network to obtain the features of the abstract syntax tree; Construct a class dependency network according to the dependencies in the source project or the target project, and obtain the features of the class dependency network based on a network embedding algorithm; Fuse the code metrics of the source project or the target project with the features of the class dependency network, and then fuse them with the features of the abstract syntax tree again; Embed the features after the second fusion into the class dependency network, and construct a graph attention network to obtain graph node features.
7. The cross-project defect prediction method based on domain adaptation according to any one of claims 1-5, characterized in that After extracting the features of the source project and the target project, it further includes: For a sample set in which the number of samples of any defect category in the source project or the target project is less than a preset threshold, determine the density of each sample in the sample set; Determine the number of neighbors of each sample according to the density of each sample based on a density function; Select the nearest neighbor samples of each sample according to the number of neighbors of each sample, and calculate the weights of each nearest neighbor sample; Randomly select neighbor samples from the nearest neighbor samples of each sample to generate synthetic samples together with each sample; Adjust the synthetic samples using the weights of the randomly selected neighbor samples to obtain new samples.
8. A cross-project defect prediction system based on domain adaptation, characterized in that, Applied to the cross-project defect prediction method based on domain adaptation according to any one of claims 1-7, it includes: A feature extraction module, configured to extract the features of the source project and the target project, and map the features of the source project and the target project to the same feature space using a feature mapping network; A first prediction module, configured to input the mapped features of the source project and the target project into a main classification network respectively to obtain the first defect prediction results of the source project and the target project; A second prediction module, configured to input the mapped features of the source project and the target project into an auxiliary classification network respectively to obtain the second defect prediction results of the source project and the target project; A loss calculation module, configured to determine a loss function according to the first defect prediction result and the second defect prediction result of the source project, and the first defect prediction result and the second defect prediction result of the target project; A model training module, configured to train the feature mapping network, the main classification network, and the auxiliary classification network according to the loss function, so as to reduce the distribution difference between the source project and the target project and align the features of the source project and the target project; A defect prediction module, configured to extract the features of the project to be predicted, and input the features of the project to be predicted into the trained feature mapping network and the trained main classification network in sequence to obtain the defect prediction result of the project to be predicted output by the main classification network.