GraphCodeBERT-based code quality analysis method
Through the combination of the GraphCodeBERT model and the Integrated Gradients algorithm, the interpretability of code and annotation consistency detection is achieved, the problems of insufficient universality and visualization in the existing technology are solved, and the accuracy of code quality evaluation and software development efficiency are improved.
Patent Information
- Application Number
- CN202510476927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
The existing technology lacks universality in code quality evaluation, insufficient analysis of model prediction reasons and visualization methods, resulting in insufficient interpretability of code and annotation consistency detection, affecting software development and maintenance efficiency.
The GraphCodeBERT model is used to detect the consistency between code and annotation through the encoder-decoder framework, and the contribution degree is calculated in combination with the Integrated Gradients algorithm, and visually presented it to provide the key dimensions and feature contribution degree of code quality analysis.
Improve the accuracy and interpretability of code quality evaluation, enable intuitive understanding of model decision-making processes, optimize code specifications and document management, identify high-risk areas, and improve software development and maintenance efficiency.
Smart Images

Figure CN120371313A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software engineering, and particularly relates to a code quality analysis method based on GraphCodeBERT. Background Art
[0002] In software code development, comments are natural language descriptions of source code. Although they do not participate in compilation, they record rich code-related information, such as function implementation principles, API usage methods, relationships and evolution processes of code segments, etc., providing important guidelines for developers to understand the code and communicate with each other. Most of the time (about 82%) of developers is spent on program understanding and navigation. The clear expression of code comments is crucial for assisting in understanding the source code and improving development efficiency. Moreover, the relationships between comments can reflect code quality problems, which are of great significance in software development and maintenance. Under the software quality assessment framework, the consistency detection of code comments is an important part of code quality analysis, which directly affects the understandability, maintainability and reusability of the code. High-quality software requires that the code comments have correct functions and are consistent with the implementation, which helps developers understand the code logic and avoid misunderstandings or decision-making mistakes caused by incorrect comments. However, in actual development, developers often do not update comments in a timely manner for various reasons, resulting in inconsistencies between comments and code, misleading developers to understand the code and affecting program understanding. Automated code and comment consistency detection technology can improve the accuracy of code quality assessment, timely discover potential problems and reduce maintenance costs. In the field of machine learning, especially deep learning, model interpretability has become the focus of attention in various fields. The automatic detection methods in the field of code quality analysis mostly rely on deep learning models, but these model architectures are complex and are often regarded as "black boxes", making it difficult to intuitively understand their decision-making processes. For example, in the GraphCodeBERT model, the model makes predictions by learning code structure and data flow information, but its internal decision-making mechanism is difficult to be directly analyzed. Therefore, improving the interpretability of the code and comment consistency detection model can not only enhance the credibility of the detection results, but also provide developers with detailed code quality analysis reports to help optimize code quality. However, there is less research on the interpretability of code-related models at present, especially in the aspect of code quality assessment tasks.
[0003] Some scholars have already conducted research on code and comment consistency detection. The Panthaplackel team proposed a deep instant consistency detection based on source code and comments. This method aims to improve the model performance and enhance the model's ability to detect the consistency between code and comments. At the same time, it pays attention to the evolution characteristics of code comments to help developers identify code quality problems. According to these characteristics, a model combining a recurrent neural network and a gated graph neural network was constructed, and certain progress has been made in code and comment consistency detection. Existing research has the following drawbacks: (1) Poor generality: In the existing research on the code comment consistency problem for code quality assessment, it is difficult to obtain the change characteristics of code comments. When different models are applied to this task, the preprocessing method of source code representation needs to be redesigned, lacking a general and effective quantitative solution. (2) Insufficient analysis of the reasons for model prediction: Related research mainly focuses on constructing models to improve the task accuracy rate, but does not deeply analyze the reasons why the model has high performance advantages. The lack of this interpretability research is not conducive to increasing users' trust in the model, nor can it help users clarify which features in the source code have a greater impact on consistency. (3) Lack of visualization presentation method: In the research on the interpretability of deep learning models, the visualization method is crucial. It can show the operation mode of the model and provide ideas for researchers. However, the existing research on code and comment consistency detection lacks a visualization method, resulting in developers' difficulty in intuitively understanding the consistency detection results and limiting the application value in software engineering practice.
[0004] To solve the above problems, the present invention proposes a code quality analysis method based on GraphCodeBERT. Summary of the Invention
[0005] The purpose of the present invention is to provide a code quality analysis method based on GraphCodeBERT, aiming to solve the problems raised in the above background technology.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A code quality analysis method based on GraphCodeBERT includes the following steps:
[0008] Step 1: Construction of source code representation;
[0009] Preprocess the code and comments in the source code respectively to obtain annotation tokens, code tokens, and data flow tokens;
[0010] Step 2: Model structure construction and fine-tuning;
[0011] The model adopts an encoder-decoder framework. The encoder uses the GraphCodeBERT-base pre-trained code representation model and fine-tunes it by inputting the source code representation generated in Step 1. The decoder is a fully connected layer that outputs a binary classification result to judge the consistency between the code and the comment.
[0012] Step 3: Contribution analysis;
[0013] Extract the code features and comment features in different ways, use the Integrated Gradients algorithm to calculate the contribution values of each dimension in the source code representation, and draw a distribution map of the contribution values to analyze the feature importance.
[0014] Step 4: Visualization presentation;
[0015] Normalize and standardize the contribution values, visualize the contribution degrees of different dimension features through colors, and summarize the multi-case analysis results to guide the optimization of code comments.
[0016] Furthermore, the specific process of Step 1 is as follows:
[0017] The comments are tokenized by BPE in the natural language processing manner to obtain comment tokens. The code is tokenized by BPE based on the CodeSearchNet corpus during the pre-training of GraphCodeBERT, retaining the structural information of the code and performing sub-word decomposition to obtain code tokens, and generating an attention mask when the tokens are converted into the actual input of the model. The code is parsed into an abstract syntax tree using a code parser according to the code structure, the data flow relationship is extracted based on the abstract syntax tree to construct the data flow graph of the code, and the data flow tokens are determined according to the variable relationship order in the data flow graph. The comment tokens, code tokens, and data flow tokens are input in the format of [CLS]comment tokens[SEP]code tokens[SEP]data flow tokens[SEP].
[0018] Furthermore, in Step 2, the GraphCodeBERT-base pre-trained code representation model uses the Transformer + GraphAttention mechanism to model the semantic relationship between the code and the variables, and the default parameter settings are used during fine-tuning. When training the model, the cross-entropy is used as the loss function, and the model is trained using the gradient descent method.
[0019] Furthermore, the specific process of Step 3 is as follows:
[0020] Feature extraction: The code features select line-level granularity features and abstract syntax tree-level features of the code; the comment features annotate the part-of-speech of the comment text, and analyze the same part-of-speech change relationship and parse the relationship between different parts-of-speech through the annotation information.
[0021] Contribution value calculation: Select a trained model and a baseline, and calculate the integrated gradient of the i-th dimension of the input x according to Equation 1 to obtain the contribution value of each feature x to the prediction result;
[0022] Equation 1:
[0023] where IntegratedGrads i (x):: is the integrated gradient (contribution value) of the i-th dimension feature; x i is the i-th dimension feature value of the input; x i ' is the i-th dimension feature value of the baseline; F is the prediction function of the model; x' is the baseline input (select a zero vector with the same dimension as the input x), and α is the scale coefficient;
[0024] Feature importance analysis: Conduct importance statistics on the features within the dimension, draw a distribution map of the contribution values, and observe the differences between the feature contribution values based on the distribution map to determine the basis for the model to discriminate the consistency between code and comments.
[0025] Compared with the prior art, the beneficial effects of the present invention are:
[0026] 1. Generalizability and scalability: The data processing process is general, and the processed data can be directly applied to other pre-trained code models. This makes it possible to conduct similar model interpretability analysis for different code models, study the reasons for the model prediction results, and apply them to code quality analysis, laying a foundation for further research and application of more pre-trained code models.
[0027] 2. Deep analysis ability: Adopt a contribution degree analysis method based on the model prediction results, fully consider the contribution values of each different dimension of the source code, and be able to obtain complete and accurate contribution degree data. This method provides sufficient dimensions and perspectives for the relevant analysis of the model, helps to deeply analyze the reasons for the model prediction results, and provides strong support for subsequent various studies.
[0028] 3. Visualization and auxiliary decision-making: Provide an effective visualization method for the analysis based on the GraphCodeBERT model, enabling researchers to comprehensively and intuitively understand the reasons for the prediction results. Visualize and analyze the interpretability of the model from multiple aspects, assist developers in understanding the importance of different features in the code quality assessment process, thereby optimizing code specifications and document management strategies, identifying high-risk code areas, and specifically optimizing the code comment management strategy, thereby improving the efficiency of software development and maintenance.
[0029] 4. Code quality assessment and review optimization: Implement code and comment consistency detection at the function level, and provide automated detection methods for code quality assessment. By comprehensively analyzing the impact of various dimensions of the source code on model prediction, quantitatively determine which code features are most critical to the consistency detection results, and deeply analyze the core factors of code quality assessment, provide strong data support for code review and optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The present invention is a flow chart of the method.
[0031] Figure 2 The build process represented for source code.
[0032] Figure 3 For the model structure and fine-tuning process.
[0033] Figure 4 A visual analysis diagram of the row-level granularity for a specific case. DETAILED DESCRIPTION
[0034] In order to have a clearer understanding of the technical features, purposes and beneficial effects of the present invention, the technical solution of the present invention is now described in detail below, but it should not be construed as limiting the applicable scope of the present invention.
[0035] The present invention provides a code quality analysis method based on GraphCodeBERT, which takes the consistency between code and comments as the key dimension of code quality analysis, can determine whether the semantics of code and comments are consistent at the function granularity, analyze the reasons for the consistency relationship between code and comments, quantitatively determine the contribution value of each division dimension of code and comments to the prediction result, and present the analysis result through a visual graph. The specific steps of the method are as follows (see Figure 1 ):
[0036] Step 1: Construction of source code representation (see Figure 2 );
[0037] In order to give full play to the advantages of GraphCodeBERT in obtaining code representation by using data streams and to mine deeper complex features hidden in the source code, the source code needs to be serialized to make it conform to the input method of GraphCodeBERT pre-training, that is, [CLS] annotation word [SEP] code word [SEP] data stream word [SEP]. This can comprehensively provide various features of code and annotations, and provide better support for interpretability analysis. The specific operations are as follows:
[0038] The source code consists of two parts: code and comments, which need to be preprocessed separately. The comments are tokenized using Byte-Pair Encoding (BPE) in the natural language processing manner to obtain comment tokens. The code is tokenized using BPE based on the CodeSearchNet corpus during the pre-training of GraphCodeBERT, preserving the structural information of the code and performing sub-word decomposition to obtain code tokens. An attention mask is generated when the tokens are transformed into the actual input of the model. The code is parsed into an abstract syntax tree using a code parser (such as tree-sitter) according to the code structure. The data flow relationship is extracted based on the abstract syntax tree to construct the data flow graph of the code, and the data flow tokens are determined according to the variable relationship order in the data flow graph. The comment tokens, code tokens, and data flow tokens are combined according to the input manner of GraphCodeBERT.
[0039] Step 2: Model Structure Construction and Fine-Tuning (see Figure 3 ).
[0040] The model adopts an encoder-decoder framework. The encoder uses the GraphCodeBERT-base pre-trained code representation model proposed by Microsoft, which adopts the Transformer + GraphAttention mechanism to model the semantic relationship between code and variables. The source code representation obtained previously (code tokens, comment tokens, and data flow tokens) is used as the model input for fine-tuning, following the default parameter settings of GraphCodeBERT-base. Finally, a fully connected layer is used as the decoder for consistency judgment. The final output is a binary classification result used to judge the relationship (consistent or inconsistent) between the code and the comments. The cross-entropy is used as the loss function during the entire training process, and the model is trained using the gradient descent method, which can provide sufficient source code features for GraphCodeBERT and provide sufficient support for analyzing code quality, model performance, and interpretability.
[0041] Step 3: Contribution Analysis;
[0042] To quantify the contribution of each dimension of the source code representation to the prediction result, the Integrated Gradients algorithm is adopted, which can calculate the contribution value of each dimension of the source code representation to the prediction result. The larger the contribution value, the greater the impact on the model to produce this prediction result. For clearer visualization later, the contribution of each dimension of the code tokens and comment tokens is used here. The specific process is as follows:
[0043] Feature extraction: To comprehensively represent various dimensions of comments and code in source code and obtain more in-depth interpretable data, different methods are used to extract code features and comment features. For code features, line-level granularity features and abstract syntax tree-level features that have been studied more are selected; for comment features, the part-of-speech of comment text is marked, and the relationship between the same part-of-speech changes and the relationship between different parts-of-speech are analyzed through the marked information to better understand the changes in sentence meaning and their similarities.
[0044] Contribution value calculation: Select a trained model and a baseline, and calculate the integrated gradient of the i-th dimension of the input x according to Equation 1 to obtain the contribution value of each feature x to the prediction result.
[0045] Equation 1:
[0046] where IntegratedGrads i (x):: is the integrated gradient (contribution value) of the i-th dimension feature; x i is the i-th dimension feature value of the input; x i ' is the i-th dimension feature value of the baseline; F is the prediction function of the model; x' is the baseline input (select a zero vector with the same dimension as the input x); α is the scale coefficient.
[0047] Feature importance analysis: Conduct importance statistics on the features within the dimension, draw a distribution map of contribution values, and observe the differences between feature contribution values based on the distribution map to determine the basis for the model to discriminate the consistency of code and comments.
[0048] Step 4: Visualization presentation;
[0049] Based on the basic data of contribution degree analysis, provide multi-dimensional and in-depth visualization analysis for the interpretability of model prediction results. The previous contribution degree analysis has calculated the importance (contribution degree) of each dimension feature of the model to the prediction result. Here, visualize the contribution values of features in different dimensions of specific cases as the basis for verifying and supplementing the prediction classification of different features by the model. Specifically as follows:
[0050] Normalize and standardize the contribution values of features in different dimensions of the visualization case to observe the color differences between different contribution values. Take the visualization analysis graph divided by line-level granularity as an example, such as Figure 4As shown, where green indicates a positive impact on the consistency of code and comments, and red indicates a negative impact on the consistency. The darker the color, the greater the contribution value of the corresponding part and the greater the impact on the predicted classification result. The importance (contribution degree) of this feature is represented by the color, and at the same time, the figure is used to understand the impact degree of different features on the predicted classification result in this case as a whole. Summarize the interpretability analysis data of multiple specific cases with different prediction results to obtain richer comparable information, which provides support for the analysis of the importance of different dimensional features of visual code and comments and writing habits. Based on the visual differences, give suggestions for code developers to maintain the consistency of code comments during software development and maintenance, thereby ensuring code quality and improving the efficiency of software development and maintenance.
[0051] The following describes the specific implementation of the present invention in detail in combination with specific embodiments.
[0052] Embodiment 1: To verify the experimental effect of the present method, a dataset provided by the Panthaplackel team is used, and the present method is compared with the model proposed by this team in a comparative experiment to evaluate the performance of the present method in the task of code and comment consistency detection on the same dataset.
[0053] Table 1 Comparison of evaluation metrics
[0054] Model Precision Recall F1 Score Accuracy SEQ 0.589 0.680 0.630 0.603 GRAPH 0.606 0.702 0.650 0.622 HYBRID 0.537 0.773 0.633 0.552 GraphCodeBert 0.812 0.718 0.762 0.776
[0055] Judging from the comparison of the four evaluation metrics (precision, recall, F1 score, and accuracy) in Table 1, the present method (GraphCodeBert) has significant performance advantages in the task of code and comment consistency detection. The specific data shows that the precision of GraphCodeBert is 0.812, the F1 score is 0.762, and the accuracy is 0.776, all of which are higher than those of other comparison models (SEQ, GRAPH, HYBRID). Combining this performance advantage with the model interpretability analysis data can not only improve the accuracy of code quality evaluation but also deeply explore the specific impact of different feature dimensions in the source code on code quality. This analysis method helps developers understand the code structure, identify potential risks, and optimize the code comment strategy, thus providing more valuable guidance during software development and maintenance.
[0056] The above is only the preferred embodiment of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A code quality analysis method based on GraphCodeBERT, characterized in that, It includes the following steps: Step 1: Construction of source code representation; Preprocess the code and comments in the source code respectively to obtain comment tokens, code tokens and data flow tokens; Step 2: Model structure construction and fine-tuning; The model adopts an encoder-decoder framework. The encoder uses the GraphCodeBERT-base pre-trained code representation model, and the source code representation generated in Step 1 is input for fine-tuning; The decoder is a fully connected layer, and the binary classification result is output to judge the consistency between the code and the comments; Step 3: Contribution analysis; Extract the code features and comment features in different ways, use the Integrated Gradients algorithm to calculate the contribution values of each dimension in the source code representation, and draw a contribution value distribution diagram to analyze the feature importance; Step 4: Visualization presentation; Perform standardization and normalization processing on the contribution values, visualize the contribution degrees of different dimension features through colors, and summarize the multi-case analysis results to guide the optimization of code comments.
2. The method for code quality analysis based on GraphCodeBERT according to claim 1, wherein The specific process of Step 1 is as follows: The comments are tokenized by BPE in the way of natural language processing to obtain comment tokens; the code is tokenized by BPE based on the CodeSearchNet corpus during the pre-training of GraphCodeBERT, retaining the structural information of the code and performing sub-word decomposition to obtain code tokens, and an attention mask is generated when the tokens are converted into the actual input of the model; Use a code parser to parse the code into an abstract syntax tree according to the code structure, extract the data flow relationship based on the abstract syntax tree to construct the data flow graph of the code, and determine the data flow tokens according to the variable relationship order in the data flow graph; Input the comment tokens, code tokens and data flow tokens in the format of [CLS] comment tokens [SEP] code tokens [SEP] data flow tokens [SEP].
3. The code quality analysis method based on GraphCodeBERT according to claim 1, wherein In Step 2, the GraphCodeBERT-base pre-trained code representation model uses the Transformer+GraphAttention mechanism to model the semantic relationship between the code and the variables, and the default parameter settings are used during fine-tuning; when the model is trained, the cross-entropy is used as the loss function, and the model is trained using the gradient descent method.
4. The method for code quality analysis based on GraphCodeBERT according to claim 1, wherein The specific process of Step 3 is as follows: Feature extraction: The code features select line-level granularity features and abstract syntax tree-level features of the code; The comment features label the part-of-speech of the comment text, and analyze the same part-of-speech change relationship and parse the relationship between different parts-of-speech through the annotation information; Contribution value calculation: Select the trained model and the baseline, calculate the integrated gradient of the i-th dimension of the input x according to Equation 1, and obtain the contribution value of each feature x to the prediction result; Formula 1: Among them, IntegratedGrads i (x):: is the integrated gradient of the i-th dimensional feature; x i is the i-th dimensional feature value of the input; x i ' is the i-th dimensional feature value of the baseline; F is the prediction function of the model; x' is the baseline input; α is the scale coefficient; Feature importance analysis: Conduct importance statistics on the features within the dimension, draw a contribution value distribution diagram, and observe the difference situation between the feature contribution values based on the distribution diagram to determine the basis for the model to judge the consistency between the code and the comments.