Cross-project software defect prediction method based on migration graph neural network

By using a migration graph neural network to process abstract syntax trees in cross-project software defect prediction, combining semantic features and metric meta features, the problems of information loss and incomplete features in the existing methods are solved, and higher prediction accuracy and generalization performance are achieved.

CN120179557APending Publication Date: 2025-06-20ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261553.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing cross-project software defect prediction methods have information loss when processing code structure information, and lack comprehensive consideration of structural characteristics and measurement characteristics, resulting in insufficient prediction accuracy and generalization performance.

Method used

Using a transfer graph neural network method, the abstract syntax tree is processed through the graph neural network, structural information is preserved, complex relationship features are extracted, and the semantic features and traditional metric meta features are combined, and the transfer learning strategy is used to share feature representation across projects.

Benefits of technology

It significantly improves the richness and accuracy of semantic feature expression, enhances the model's prediction ability on different projects, and improves the accuracy and generalization of cross-project defect prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179557A_ABST
    Figure CN120179557A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-project software defect prediction method based on a migration graph neural network, and relates to the technical field of software defect prediction. The method comprises the following steps: S1, source code analysis and vector mapping: analyzing a Java project source code through Javalang to obtain an abstract syntax tree, then constructing structure information of AST, and converting a node of the AST into a digital vector; s2, neural network model construction and feature extraction: inputting the digital vector and an AST graph structure into a graph neural network model, and introducing a transfer learning algorithm to extract transferable semantic features among different items; s3, constructing joint features: combining migratable traditional metric element features extracted through weighted feature selection (WBAD) with semantic features extracted by the network model to form joint features for cross-project defect prediction (CPDP); and S4, constructing a defect prediction model: training an LR classifier by using the joint features corresponding to the Java file and the corresponding labels. And then, using the trained model to predict whether the Java file of the target item has defects or not. According to the method, cross-project defect prediction is carried out in combination with a transfer learning method and a deep learning method, semantic features are extracted by using a neural network model, then features with high similarity among projects are matched by using the transfer learning method, and the features are combined with traditional metric element features, so that the performance of the model in the aspect of cross-project defect prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software defect prediction, and particularly relates to a cross-project software defect prediction method based on transfer learning. Background Art

[0002] With the rapid development of information technology, modern life increasingly relies on various complex software systems. However, due to the diversification of application scenarios and the complexity of development processes, there are inevitably a large number of potential defects in software. These defects not only affect the stability and security of software, but may also result in serious economic losses. Effectively predicting and fixing software defects is an important means to improve software quality. To achieve this goal, it is crucial for development and testing teams to establish a software defect prediction model, as this model can assist them in identifying and fixing defects in the early stage of development, reducing the maintenance cost after software release, and improving the overall efficiency and quality of the project.

[0003] Current software defect prediction models are mainly divided into within-project defect prediction, cross-project defect prediction, and cross-company defect prediction. The cross-project defect prediction method is particularly important. It relies on training a model in a project with rich data, and then applying the model to a target project lacking training samples through transfer learning. In particular, in the case of large differences between projects and inconsistent data distributions, transfer learning can effectively help the model adapt to the new project environment, thereby improving the prediction accuracy.

[0004] Although transfer learning techniques and deep learning models have shown remarkable results in many fields, in software defect prediction, especially in cross-project defect prediction, traditional methods still face the following deficiencies:

[0005] 1. Existing cross-project defect prediction methods based on abstract syntax trees mostly flatten the abstract syntax tree into a linear sequence and then input it into a neural network model, which will lead to the loss of structural information in the code and make the extracted semantic features not rich enough.

[0006] 2. In current research, there is a lack of comprehensive consideration of structural features and metric features related to defects in cross-project defect prediction. Due to the complexity and diversity of software defects, it is difficult to meet the prediction requirements by only considering a single type of feature (such as code metric features or basic semantic features), resulting in the difficulty of the prediction model to accurately describe the characteristics of defect distribution, thus affecting the generalization performance of the model.

[0007] To solve the above problems, the present invention proposes a cross-project software defect prediction method based on transfer graph neural network. Summary of the Invention

[0008] The object of the present invention is to propose a cross-project software defect prediction method based on a transfer graph neural network to solve the problems raised in the background art. The present invention introduces a transfer graph neural network, uses a graph neural network to process the abstract syntax tree, and extracts complex relationship features in the code by retaining the structural information of the AST, significantly improving the richness and accuracy of semantic feature expression. By comprehensively considering semantic features and traditional metric features, the model can fully learn multi-dimensional features related to defects, thereby improving the accuracy and generalization of prediction. In addition, in order to improve the feature matching ability across projects, the present invention uses a transfer learning strategy to share and transfer feature representations between different projects across projects, thereby enhancing the prediction ability of the model on different projects.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] The present invention proposes a cross-project software defect prediction method based on a transfer graph neural network, including the following steps:

[0011] S1. Source code parsing and vector mapping: By parsing the project source code in the PROMISE repository, the source project is transformed into the form of an abstract syntax tree (AST), then the nodes of the abstract syntax tree are transformed into token vectors, and then the token vectors are mapped into integer (numeric) vectors by constructing a dictionary. Then, the DGL library is used to create a DGL graph object through the extracted node and edge information, and a complete graph structure containing node features and edge information is constructed for the input of the next neural network.

[0012] S2. Neural network model construction and feature extraction: Construct a neural network model, then input the obtained integer (numeric) vectors and the graph structure of the AST into the constructed neural network model, and introduce a transfer learning algorithm to extract transferable semantic features between different projects.

[0013] S3. Construct combined features: Combine the transferable traditional metric features extracted by weighted feature selection (WBAD) with the semantic features extracted by the deep learning model to form combined features for cross-project defect prediction (CPDP).

[0014] S4. Construct a defect prediction model: Use the combined features corresponding to the Java files and the corresponding labels to train an LR classifier. Then, use the trained model to predict whether there are defects in the Java files of the target project.

[0015] Preferably, in step S1, when parsing the project source code of the PROMISE repository and converting the source project into an abstract syntax tree form, the specific steps are as follows: Parse the source code into an AST through the Javalang method and construct the corresponding token vector. Considering that the names of methods and variables are usually project-specific, use node types to label nodes instead of specific names. Then, a unified mapping dictionary is created between node types and integers to convert the token vector into an integer vector. Finally, use the DGL library to create a DGL graph object through the extracted node and edge information, and then construct a complete graph structure containing node features and edge information for the input of the next neural network.

[0016] Preferably, in step S2, when inputting the obtained digital vector and the constructed graph structure into the constructed neural network model to extract the transferable semantic features between different projects, the specific steps are as follows: Input the obtained integer vector and the graph structure into the transferable graph neural network model (TGNN). First, through the embedding layer, map the integer vector to the same dimension. Then, input the embedded node features and the graph structure into the graph convolutional layer to capture the structural information between nodes. Finally, through the matching layer, use the transfer learning algorithm to match the transferable features between projects and transfer the features with high similarity between different projects.

[0017] Preferably, the TGNN model adopts two layers of graph convolutional layers and one layer of graph attention network. In the above three layers of networks, each layer will input the graph structure to capture the structural information between nodes.

[0018] Preferably, when using the transfer learning algorithm to match the transferable features between projects and transfer the features with high similarity between different projects, the specific steps are as follows: Use the maximum mean discrepancy in transfer learning to measure the similarity between the transferable features between projects, and transfer the features with high similarity between different projects for matching.

[0019] Preferably, when combining the transferable metric meta-features and semantic features to construct transferable combined features, the specific steps are as follows: Use WBAD to transfer the metric meta-features, and use the concatenate function to combine the transferable semantic features and metric meta-features to construct transferable combined features.

[0020] Preferably, in step S4, when using the target project to train the constructed cross-project defect prediction model, the specific steps are as follows: Use the combined features and corresponding labels of the Java files in the source project to train the LR classifier. Then, use the trained model to predict whether the Java files in the target project have defects. Finally, use the commonly used metrics for evaluating the performance of cross-project defect prediction to evaluate the performance of the model.

[0021] Compared with the prior art, the present invention provides a cross-project software defect prediction method based on a transfer graph neural network, having the following beneficial effects:

[0022] The present invention innovatively combines transfer learning and graph neural network methods for cross-project software defect prediction. First, a graph neural network model is used to extract structural features and semantic features in the code, which can more comprehensively capture deep information related to defects. Then, transfer learning technology is used to match similar features between different projects, realizing cross-project transfer and sharing of features, thereby significantly improving the performance of the model in cross-project defect prediction. Specifically, the present invention combines the transfer learning method with the graph neural network, introducing a transferable graph neural network model in cross-project software defect prediction to capture the structural information between nodes in the code and generate richer semantic features. Further, these features are jointly processed with code metric features to optimize the defect prediction ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0024] Figure 1 is a schematic flowchart of a cross-project software defect prediction method based on a transfer graph neural network;

[0025] Figure 2 is a flowchart framework of a cross-project software defect prediction method based on a transfer graph neural network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details.

[0027] Embodiment 1:

[0028] Referring to Figure 1 and Figure 2 , a cross-project software defect prediction method based on a transfer graph neural network includes the following steps:

[0029] S1. Source code parsing and vector mapping: By parsing the project source code in the PROMISE repository, the source project is transformed into the form of an abstract syntax tree, and then the nodes of the abstract syntax tree are converted into token vectors and mapped into integer vectors for use by the neural network model;

[0030] Source code parsing refers to the process of analyzing and understanding the source code of a program. Source code parsing can be achieved through methods such as lexical analysis, syntax analysis, and semantic analysis. Vector mapping refers to the process of mapping one vector space to another. Vector mapping is often used to map high-dimensional feature vectors to low-dimensional representation spaces or to map different vector representations to the same representation space.

[0031] Specifically, parse the project source code collected from the PROMISE repository and convert the source project into the form of an abstract syntax tree. The specific steps are as follows: By parsing the project source code in the PROMISE repository, convert the source project into the form of an abstract syntax tree (AST), then convert the nodes of the abstract syntax tree into token vectors, and then map the token vectors into integer (numeric) vectors by constructing a dictionary. Then use the DGL library to create a DGL graph object through the extracted node and edge information, and construct a complete graph structure containing node features and edge information for the input of the next neural network.

[0032] S2. Neural network model construction and feature extraction: Construct a neural network model, then input the obtained integer (numeric) vectors and graph structure into the constructed neural network model, and introduce a transfer learning algorithm to extract transferable semantic features between different projects, and finally generate transferable semantic features;

[0033] Specifically, construct a transferable graph neural network (TGNN) model, and then input the obtained integer (numeric) vectors and graph structure into the constructed graph neural network model. However, before inputting into the neural network model, first use the oversampling method to preprocess the class imbalance problem. We only perform random oversampling on the source project and do not perform any processing on the target project.

[0034] In the present invention, two layers of GCN and one layer of GAT network models are selected here to extract the (key) semantic features of the source code. By adding a transfer learning method to the network model and constructing a TGNN model, the model can extract transferable semantic features between different projects.

[0035] Specifically, the obtained integer (numeric) vector and the constructed graph structure are input into the transferable TGNN neural network model. First, through the embedding layer, the integer (numeric) vector is mapped to the same dimension in this embedding layer; the mapped vector and the graph structure are input into the first layer of GCN, and then the output and the graph structure are input into the next layer of GAT. The output of GAT is input into the second layer of GCN together with the graph structure again. For each convolution, we will re-input the graph structure to prevent the loss of structural information. Finally, through the matching layer, in this matching layer, the transfer learning algorithm is used to match the transferable features between projects. The maximum mean discrepancy (MMD) in transfer learning is used to measure the similarity between the transferable features between projects, and the features with high similarity between different projects are transferred for matching. Specifically, the features of different projects are mapped into the reproducing kernel Hilbert space, and the maximum mean discrepancy (MMD) between different feature distribution curves of the source project and the target project is measured. The network is iteratively trained to continuously reduce the maximum mean discrepancy (MMD) distance, achieving the effect of transferring features with high similarity between different projects.

[0036] In the present invention, the most commonly used maximum mean discrepancy (MMD) in transfer learning is selected to measure the distance between different features of the source project and the target project. It is mainly used to measure the distance between two different but related distributions. Specifically, first, different features are mapped into the reproducing kernel Hilbert space, and then MMD is used in this space to calculate the distance between the feature distribution curves. Through the continuous iterative training of the network model, the distance between different features is minimized, and the features with similar distances are transferred.

[0037] The calculation formula of MMD is as follows:

[0038]

[0039] Among them, σ(·) represents the mapping, which is used to map the original variable into the reproducing kernel Hilbert space (RKHS); X = {x1, x2, …, x i} is the entity dataset of the source project software; Y = {y1, y2, …, y j} is the entity dataset of the target project software; H represents the kernel space, in which the inner product can be converted into a kernel function, and at this time MMD can be directly calculated through the kernel function.

[0040] In the present invention, sigmoid is selected as the activation function, and this function transforms the input into an output on (0, 1). The specific formula is expressed as follows:

[0041]

[0042] When training the model, set the batch size to 1, the number of epochs to 15, and the learning rate to 1e -5 In addition, the present invention selects stochastic gradient descent and the Adam optimizer to optimize the model.

[0043] S3. Construct joint features: Combine the transferable metric features and semantic features to construct transferable joint features;

[0044] Specifically, combining the transferable metric features and semantic features to construct transferable joint features, the specific steps are as follows: Use the concatenate function to combine the transferable traditional metric features extracted by weighted feature selection (WBAD) with the semantic features to form joint features for cross-project defect prediction.

[0045] S4. Construct a defect prediction model: Use the target project to train the constructed cross-project defect prediction model, and compare the performance of different models using common evaluation metrics for cross-project defect prediction.

[0046] Specifically, use the joint features corresponding to the Java files in the source project and the corresponding labels to train an LR classifier. Then, use the trained model to predict whether there are defects in the Java files of the target project. Use common metrics such as F-measure, ACC, and MCC to evaluate the model.

[0047] In this embodiment, the neural network models used are all written based on the Pytorch framework, and the hardware environment of the experiment is as follows: The processor is an Intel Xeon Platinum 8255C CPU, the memory is 43GB, the operating system is Ubuntu 18.04, the GPU is an NVIDIA RTX 2080Ti, and the development environment is Pytorch. The dataset used in this embodiment is a software defect dataset composed of the source code of Java projects publicly available in the PROMISE repository and metric features. In addition, the present invention uses the source project as the training set, and then selects the target project for testing to verify the effectiveness of the model.

[0048] During the implementation process, use the target project to evaluate the trained model, and select the common evaluation metrics F-measure, ACC, and MCC in cross-project defect prediction to evaluate the proposed model and judge its performance in cross-project defect prediction. The calculation formulas of F-measure, ACC, and MCC are as follows:

[0049]

[0050] Among them, TP (True Positive): the number of defects correctly predicted by the model; TN (True Negative): the number of non-defects correctly predicted by the model; FP (False Positive): the number of non-defects mispredicted as defects by the model; FN (False Negative): the number of defects mispredicted as non-defects by the model; and the larger the F-measure, ACC, and MCC, the more accurate the result of the model in cross-project defect prediction.

[0051] Comparative example:

[0052] The present invention uses a comparative experiment to verify that the proposed method has better performance in cross-project defect prediction. Specifically, given a target project, the remaining projects are used as source projects to train the model and make predictions respectively, and the average performance metrics are calculated. Experiments are carried out using traditional machine learning methods and common neural network models respectively, and then the experimental results of the above models are compared. Here, different models (Logistic Regression Algorithm (LR), Nearest Neighbor Filter Algorithm (NNFilter), Extended TCA (TCA+), and Deep Belief Network (DBN), Convolutional Neural Network (CNN), Transferable Convolutional Neural Network (TCNN), MANN, ASTNN, CodeBERT, SDPBB) in the neural network model and the proposed TGNN of the present invention are mainly selected for comparison.

[0053] Experiments are carried out on multiple Java software projects in the PROMISE library (respectively: Ant-1.7, Camel-1.6, Forrest-0.8, Ivy-2.0, jEdit-4.2, Log4j-1.2, Lucene-2.4, Poi-3.0, Synapse-1.2, Velocity-1.6.1, Xalan-2.7, and Xerces-1.4.4). The detailed information of the software projects such as the number of modules, defect rate (the proportion of defective modules in the total number of modules), and the maximum number of defects in the modules is shown in Table 1 below.

[0054] Table 1. Defect information of open-source projects in the PROMISE repository

[0055]

[0056]

[0057] Preprocess the above data sets, conduct experiments according to the foregoing process, and the obtained results are shown in Tables 2, 3, and 4.

[0058] Table 2. F-measure values of different models on the PROMISE data set

[0059]

[0060] Table 3 ACC values of different models on the PROMISE dataset

[0061]

[0062]

[0063] Table 4 MCC values of different models on the PROMISE dataset

[0064]

[0065] As can be seen from Tables 2 - 4, the average values of F - measure, ACC, and MCC of the TGNN method are 0.505, 0.581, and 0.188 respectively. Compared with the baseline method, our average values are the best in the corresponding metrics. When examining the W / T / L perspective, compared with each method, it is better than at least 8 target projects in the F - measure metric and better than at least 7 target projects in the ACC and MCC metrics.

[0066] It should be noted here that the term "metric meta - feature" refers to the attribute values calculated by static code analysis tools. In this paper, 20 metric attributes are selected, including the number of weighted methods (wmc), depth of inheritance tree (dit), number of children nodes (noc), coupling between object classes (cbo), response for a class (rfc), lack of cohesion in methods (lcom), coupling among attributes (ca), coupling out (ce), number of public methods (npm), improved lack of cohesion in methods (lcom3), lines of code (loc), data access metric (dam), metric of aggregation (moa), metric of functional abstraction (mfa), cohesion among methods (cam), inheritance coupling (ic), coupling between methods (cbm), average method complexity (amc), maximum cyclomatic complexity (max_cc), average cyclomatic complexity (avg_cc), etc.

[0067] Compared with existing machine learning algorithms and neural network models, the TGNN method proposed by the present invention can produce better performance in defect prediction. This improvement will more effectively help the software quality assurance team find the modules or files prone to defects and point out the specific number of defects, so that the team can allocate resources more effectively and resolve the potential hazards in the software in a timely manner.

[0068] The present invention is not limited to the above - mentioned specific embodiments. Those of ordinary skill in the art, starting from the above - mentioned concept and without creative labor, can make various transformations, all of which fall within the protection scope of the present invention.

Claims

1. A cross-project defect prediction method based on a migration graph neural network, characterized in that: The following steps are involved: S1. Source code parsing and vector mapping: Parse the source code of the project in the PROMISE repository, convert the source project into an abstract syntax tree, then retain the structural information of the AST and convert the nodes of the AST into digital vectors; S2. Neural network model construction and feature extraction: Construct a neural network model, input the obtained integer vector and graph structure into the constructed neural network model, and introduce the transfer learning algorithm to extract semantic features that can be transferred between different projects; S3. Construct joint features: Combine the transferable traditional metric meta-features extracted by weighted feature selection with the semantic features extracted by the deep learning model to form joint features for cross-project defect prediction; S4. Construct a defect prediction model: Use the joint features obtained in S3 and the corresponding labels to build and train a defect prediction model, and use the trained model to predict whether there are defects in the Java files of the target project.

2. The cross-project software defect prediction method based on migration graph neural network according to claim 1 is characterized in that: The S1 specifically includes the following contents: S1.

1. Parse the source code into AST through Javalang method and construct the corresponding token vector; S1.

2. Use node types to label nodes, then create a unified mapping dictionary between node types and integers to convert token vectors into integer vectors; S1.

3. Finally, the DGL library is used to create a DGL graph object through the extracted node and edge information, and then a complete graph structure containing node features and edge information is constructed.

3. The cross-project software defect prediction method based on migration graph neural network according to claim 1 is characterized in that: The S2 specifically includes the following contents: S2.

1. Input the obtained integer vector and graph structure into the transferable graph neural network model, and map the integer vector to the same dimension through the embedding layer; S2.2, input the embedded node features and graph structure into the graph convolution layer to capture the structural information between nodes; S2.

3. After the matching layer, the transfer learning algorithm is used to match the transferable features across projects and transfer the features with high similarity between different projects.

4. The cross-project software defect prediction method based on migration graph neural network according to claim 4 is characterized in that: The neural network model includes two graph convolution layers and one graph attention network. The graph structure is input into each layer of the two graph convolution layers and the one graph attention network to capture the structural information between nodes.

5. The cross-project software defect prediction method based on migration graph neural network according to claim 4 is characterized in that: S2.3 specifically includes the following contents: The maximum mean difference in transfer learning is used to measure the similarity between features that can be transferred across projects, and features with high similarity between different projects are transferred for matching.

6. The cross-project software defect prediction method based on migration graph neural network according to claim 1 is characterized in that: The S3 specifically includes the following contents: The metric meta-features are migrated through weighted feature selection, and the migratable semantic features and metric meta-features are combined using the concatenate function to construct a migratable joint feature.

7. The cross-project software defect prediction method based on migration graph neural network according to claim 1 is characterized in that: The S4 specifically includes the following contents: S4.

1. Use the joint features and corresponding labels corresponding to the Java files in the source project to train the LR classifier; S4.

2. Use the trained model to predict whether the Java file of the target project has defects; S4.

3. Use commonly used metrics for evaluating cross-project defect prediction performance to evaluate the performance of the model.