A CodeBert and Spatial Structure Based Code Defect Prediction Method

By combining CodeBert and spatial structure information, extracting code semantics and spatial structure features, and using bi-LSTM and LSTM models for defect prediction, the problem of failing to effectively utilize code space structure and semantic vectors in the prior art is solved, and more efficient and accurate code defect prediction is achieved.

CN115185730BActive Publication Date: 2025-06-13NANTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210849413.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-13
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize the spatial structure information of the code in code defect prediction, and the model training efficiency is low directly using code semantic vectors.

Method used

The code defect prediction method based on CodeBert and spatial structure is adopted, and the code semantics are extracted through the pre-trained model CodeBert, and the shortest path length matrix is ​​generated in combination with the abstract syntax tree to obtain the code space information. Finally, the code feature representation is further obtained through the bi-LSTM and LSTM models to more accurately predict whether the code has defects.

Benefits of technology

This method not only takes into account the spatial structure information of the code, but also better characterizes the code semantics through CodeBert and BERT-whitening technologies, improving the accuracy and reliability of defect prediction. It performs excellently on multiple performance indicators compared with the baseline method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115185730B_ABST
    Figure CN115185730B_ABST
Patent Text Reader

Abstract

The present invention provides a code defect prediction method based on CodeBert and spatial structure, belonging to the field of computer technology. It solves the problem that the code feature extraction part in the defect prediction model lacks the code spatial structure, resulting in the model obtaining more code feature information. The technical solution is as follows: It includes the following steps: S1: Collect a data set from issues and perform preprocessing operations; S2: Perform key feature extraction and dimensionality reduction; S3: Represent the code spatial structure information by the shortest path length; S4: Construct a bi-LSTM / LSTM neural network model; S5: Construct an Aast input neural network model; S6: Obtain the prediction result. The beneficial effect of the present invention is that the present invention extracts richer code semantics and structural features from the source code, thereby improving the quality and reliability of defect prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a code defect prediction method based on CodeBert and spatial structure. Background Art

[0002] In software development and maintenance, technicians spend a large amount of time on determining whether there are defects in the code. More accurate defect prediction technologies can identify the modules in the program that are prone to defects, thereby optimizing the software development and testing processes and improving the quality and reliability of the software. Therefore, they play an important role in software development and maintenance. However, in the actual process, due to limited project development and maintenance budgets, technicians lack programming experience, or some defects are hidden too deeply and are difficult to detect manually. Therefore, it is necessary to design an effective automatic defect prediction method for the program.

[0003] In recent years, machine learning has been widely applied to the field of defect prediction. Researchers only extract code semantics and directly apply them to the model, ignoring the spatial structure information of the code itself, and rarely pre-train the initially extracted code semantic vectors to better adapt to the neural network model. Recent research has modeled defect prediction as a classifier problem and proposed deep learning-based methods that have achieved good results. However, these methods usually require accurate extraction of code semantics to train the model. Therefore, it is necessary to explore a method that can better represent code semantics to solve the defect prediction problem.

[0004] How to solve the above technical problems has become the subject faced by the present invention. Summary of the Invention

[0005] The purpose of the present invention is to provide a code defect prediction method based on CodeBert and spatial structure, which can predict whether there are defects in program code.

[0006] The idea of the present invention is as follows: The present invention proposes a code defect prediction method based on CodeBert and spatial structure, that is, extracting code semantics through the pre-trained model CodeBert and generating the shortest path length matrix using the abstract syntax tree, so as to obtain the code spatial information, and then further obtaining the code feature representation through bi-LSTM and LSTM models, etc., so as to more accurately predict whether there are defects in the code.

[0007] In order to achieve the above invention purpose, the technical solution adopted by the present invention is specifically as follows: A code defect prediction method based on CodeBert and spatial structure, wherein, the method includes the following steps:

[0008] (1) Mine the top-ranked Java projects on GitHub by the number of stars, and then collect the submitted defective code segments, defect descriptions, and the code segments after repair from the issues through a crawler program. Translate the defect descriptions in non-English languages into English to obtain the dataset D. Set the format of the dataset as <project name, defective code, repaired code, defect description>. The specific preprocessing operations include the following steps:

[0009] (1-1) First, delete the data without defect descriptions;

[0010] (1-2) Translate the defect descriptions in non-English languages into English;

[0011] (2) Extract the code semantic information through CodeBert and BERT-whitening operations, which specifically include the following steps:

[0012] (2-1) For a given code segment, first split it according to the case naming rule to obtain the input sequence;

[0013] (2-2) Input the sequence into CodeBert, extract the hidden states of the first layer and the last layer in the output, and take their average to obtain the semantic feature vector of the code segment;

[0014] (2-3) Use BERT-whitening to process the semantic vector, and perform key feature extraction and dimensionality reduction through linear transformation;

[0015] (3) Consider the abstract syntax tree information of the dataset code, and represent the code space structure information by the shortest path length, which specifically includes the following steps:

[0016] (3-1) First, generate the corresponding abstract syntax trees (ASTs) for the code according to solidity-parser-antlr;

[0017] (3-2) Use the breadth-first search algorithm to generate the shortest path length matrix of the AST, denoted as A;

[0018] (3-3) Take the reciprocal of the non-zero elements Aij in the matrix A to obtain the spatial structure information matrix, denoted as Aast.

[0019] (4) Randomly divide the constructed dataset into a training set, a validation set, and a test set, and at the same time construct a bi-LSTM and an LSTM neural network, which specifically include the following steps:

[0020] (4-1) Divide the dataset obtained in step 1, and perform random division according to the ratio of 80%:10%:10% (training:testing:evaluation);

[0021] (4-2) The neural network uses a stacked bi-LSTM and LSTM, with the bi-LSTM in the front and the LSTM in the back. At the same time, the hidden state vector hi of the bi-LSTM is used as the input of the LSTM;

[0022] (5) The semantic feature vector Xi obtained in step (2) is used as the input of the bi-LSTM, and the hidden state vector of its last layer is hi;

[0023] (6) The spatial structure information matrix Aast obtained in step (3) and the hidden state vector hi obtained in step (5) are used as the input of the LSTM, and the hidden state vector of its last layer is hj;

[0024] (7) The semantic feature vector Xi obtained in step (2) is used as the input of the bi-LSTM, and the hidden state vector of its last layer is hi;

[0025] (8) Use the softmax regression model to normalize the hidden state information obtained in step (7) to determine whether the code has defects.

[0026] As a further optimization scheme of the code defect prediction method based on CodeBert and spatial structure provided by the present invention, in step (2), BERT-whitening in machine learning is used to further process the semantic vector, that is, key feature extraction and dimensionality reduction are performed on the semantic vector extracted by CodeBert through linear transformation, thereby reducing the storage space and greatly improving the training speed of the model.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: The code defect prediction method based on CodeBert and spatial structure proposed by the present invention not only considers the spatial structure information of the code, but also abandons the traditional method of directly entering the semantic vector into the model training. Instead, it uses the pre-trained model CodeBert that can better represent the semantics of the code, and also uses the BERT-whitening technology to perform key feature extraction and dimensionality reduction on the extracted semantic vector; compared with the previous model that simply uses the bi-LSTM neural network, adding an additional LSTM layer can better represent the structural information of the code; and, by introducing the Attention mechanism in step (7), the attention of the model can be further concentrated on the more important hidden state information, making the final prediction result more accurate and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0029] Figure 1 The system framework diagram of a code defect prediction method based on CodeBert and spatial structure provided by the present invention. Detailed implementation manners

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0031] Embodiment 1

[0032] See Figure 1 As shown, this embodiment provides a code defect prediction method based on CodeBert and spatial structure, including the following content:

[0033] (1) Mine the Java projects with the top-ranked Stars on GitHub, and then collect the submitted defective code segments, defect descriptions, and the code segments after repair from the issues through a crawler program. Translate the defect descriptions in non-English languages into English to obtain a data set D. Set the format of the data set as <project name, defective code, code after repair, defect description>. The first column in the data set example shows the project name, the second column shows the defective code, the third column shows the code after repair, and the fourth column shows the defect description; the specific preprocessing operations include the following steps:

[0034] (1-1) First, delete the data without defect descriptions;

[0035] (1-2) Translate the defect descriptions in non-English languages into English;

[0036] (2) Extract the code semantic information through CodeBert and BERT-whitening operations, specifically including the following steps:

[0037] (2-1) Segment the code segments according to the camel case naming rule to obtain the input sequence where M is the length of the code sequence. Then input it into CodeBert, extract the hidden states of the first layer and the last layer in the output, and take the average value to obtain the semantic feature vector of the code X∈Rd, where d is the hidden size of CodeBert. Finally, obtain the set of smart contract code line vectors in the training set where N is the size of the training set;

[0038] (2-2) Further use BERT-whitening in machine learning to process the semantic vectors, and perform key feature extraction and dimensionality reduction through linear transformation.

[0039] (3) Consider the abstract syntax tree information of the dataset code, and represent the code space structure information by the shortest path length. The specific steps are as follows:

[0040] (3-1) First, generate the corresponding abstract syntax trees (ASTs) from the code according to solidity-parser-antlr;

[0041] (3-2) Use the breadth-first search algorithm to generate the shortest path length matrix of the AST, denoted as A;

[0042] (3-3) Take the reciprocal of the non-zero elements Aij in the matrix A to obtain the space structure information matrix, denoted as Aast.

[0043] (4) Randomly divide the constructed dataset into a training set, a validation set, and a test set. At the same time, construct a bi-LSTM and an LSTM neural network. The specific steps are as follows:

[0044] (4-1) Divide the dataset obtained in step 1 according to a ratio of 80%:10%:10% (training:testing:evaluation) for random division;

[0045] (4-2) The neural network uses a stack of bi-LSTM and LSTM, with bi-LSTM in the front and LSTM in the back. At the same time, use the hidden state vector hi of bi-LSTM as the input of LSTM;

[0046] (5) Use the semantic feature vector Xi obtained in step 2 as the input of bi-LSTM, and its hidden state vector of the last layer is hi;

[0047] (6) Use the space structure information matrix Aast obtained in step 3 and the hidden state vector hi obtained in step 5 as the input of LSTM, and its hidden state vector of the last layer is hj;

[0048] (7) Use the semantic feature vector Xi obtained in step 2 as the input of bi-LSTM, and its hidden state vector of the last layer is hi;

[0049] (8) Use the softmax regression model to normalize the hidden state information obtained in step 7 to determine whether the code has defects.

[0050] (9) Evaluate the method of this embodiment and the existing defect methods on the same dataset, and use five performance metrics (i.e., accuracy, recall, precision, F1, and AUC) from the field of defect prediction research to automatically evaluate the quality of the model:

[0051] Table 4 Comparison table of the results of the method of this embodiment and other methods

[0052]

[0053]

[0054] Experiments show that the code defect prediction method based on CodeBert and spatial structure proposed in this embodiment can perform more reliable defect prediction compared with the baseline method. Specifically, the method in this embodiment integrates the semantic and code spatial structure information of the code, and can outperform these baseline methods. Among them, for Accuracy, the method in this embodiment can at least improve the performance by 1.27% respectively; for Precision, the method in this embodiment has at least improved the performance by 18.58%; for Recall, the method in this embodiment has not been significantly improved compared with other methods; for F1, the method in this embodiment can at least improve the performance by 3.06%; for AUC, the method in this embodiment can at least improve the performance by 1.67%. These results show that the method proposed in this embodiment has strong competitiveness.

[0055] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A code defect prediction method based on CodeBert and spatial structure, characterized in that, it includes the following steps: S1. Mine Java projects with top Stars numbers on GitHub, and then collect the submitted defective code segments, defect descriptions, and repaired code segments from issues through a crawler program. Translate the defect descriptions in non-English languages into English to obtain a dataset D, and set the format of the dataset as <project name, defective code, repaired code, defect description>; S2. Use CodeBert to extract code semantic features, and perform key feature extraction and dimensionality reduction through BERT-whitening; S3. Consider the abstract syntax tree information of the dataset code, and represent the code spatial structure information through the shortest path length; S4. Randomly divide the constructed dataset into a training set, a validation set, and a test set, and simultaneously construct a bi-LSTM and an LSTM neural network; S5. Input the semantic feature vectors extracted in step S2 into the bi-LSTM neural network, and its last layer outputs hidden state information; S6. Use the code spatial structure information obtained in step S3 and the hidden state obtained in step S5 as inputs to re-enter the LSTM neural network; S7. Introduce an Attention mechanism for the hidden states obtained in steps S5 and S6 to further obtain important hidden state information; S8. Use a softmax regression model to normalize the hidden state information obtained in step S7 to determine whether the code has defects; The S2 includes the following steps: S21: For a given code segment, split it according to the case rules to obtain an input sequence; S22: Input the sequence into CodeBert, extract the hidden states of the first layer and the last layer in the output, and take their average to obtain the semantic feature vector of the code segment; S23: Use BERT-whitening to process the semantic vector, and perform key feature extraction and dimensionality reduction through linear transformation; The S3 includes the following steps: S31: First, generate the corresponding abstract syntax tree according to solidity-parser-antlr, denoted as ASTs; S32: Use the breadth-first search algorithm to generate the shortest path length matrix of the AST, denoted as A; S33: Take the reciprocal of the non-zero elements Aij in the matrix A to obtain the spatial structure information matrix, denoted as Aast.

2. The code defect prediction method based on CodeBert and spatial structure according to claim 1, characterized in that, the S4 includes the following steps: S41: Divide the dataset obtained in S1, and randomly divide the training set, test set, and validation set according to the ratio of 80%:10%:10%; S42: The neural network uses bi-LSTM and LSTM stacked, with bi-LSTM in the front and LSTM in the back, and at the same time use the hidden state vector hi of bi-LSTM as the input of LSTM.

3. The code defect prediction method based on CodeBert and spatial structure according to claim 1, characterized in that, in step S5, the semantic feature vector Xi obtained in S2 is used as the input of the bi-LSTM, and the hidden state vector of its last layer is hi.

4. The code defect prediction method based on CodeBert and spatial structure according to claim 1, characterized in that, in step S6, the spatial structure information matrix Aast obtained in step S3 and the hidden state vector hi obtained in step S5 are used as the input of the LSTM, and the hidden state vector of its last layer is hj.

5. The code defect prediction method based on CodeBert and spatial structure according to claim 1, characterized in that, in step S7, the Attention mechanism is introduced to weight the hidden states of all steps, and the attention is focused on the more important hidden state information.

6. The code defect prediction method based on CodeBert and spatial structure according to claim 1, characterized in that, in step S8, the softmax regression model is used to normalize the hidden state information obtained in step S7, so as to judge whether the code has defects.

Citation Information

Patent Citations

  • Intelligent contract code annotation generation method based on information retrieval

    CN113743062A

  • Development pipeline integrated ongoing learning for assisted code remediation

    WO2022093250A1