Method for enhancing code abstract based on LLM enhanced data set and AST hierarchical perception model
Through the two-stage LLM quality feedback and the dirGCN model with directed syntax graph, the problems of low-quality training data and insufficient utilization of hierarchical relationships in code summarization are solved, high-quality code summaries are generated, and the efficiency of code understanding and maintenance is improved.
Patent Information
- Application Number
- CN202510703555.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing code summary generation methods rely on low-quality training data and fail to fully utilize the hierarchical relationship of the abstract syntax tree (AST), resulting in poor quality of the generated code summaries.
A two-stage LLM quality feedback strategy is adopted to evaluate and optimize the code dataset. A directed syntax graph is constructed and an improved directed graph convolutional network (dirGCN) is used to extract hierarchical syntax features. A decoder and a similarity comparator are combined to generate high-quality summaries.
Through code quality assessment and optimization methods and hierarchical perception models, higher quality code summaries are generated, which solves the problems of low-quality training data and insufficient utilization of hierarchical relationships, and improves the efficiency and accuracy of code summaries.
Smart Images

Figure BDA0005425159000000021 
Figure BDA0005425159000000024 
Figure BDA0005425159000000031
Abstract
Description
Technical Field
[0001] The invention belongs to a method for automatically generating a code summary, and relates to the technical field of software engineering and natural language processing. Background Art
[0002] Code summarization is a key task in software engineering, which aims to generate natural language summaries of source code snippets. It plays a vital role in helping developers understand and maintain software. In the actual software development process, developers often need to quickly understand the code. However, most software projects lack high-quality summaries because summaries are very time-consuming during the development process and often become outdated during maintenance, which reduces the efficiency of software development. Therefore, it is necessary to design an effective automated code summarization method to generate high-quality code summaries.
[0003] Existing research primarily aims to improve code representation quality by mining source code structural features, with methods based on abstract syntax tree (AST) modeling being the most representative. For example, Hu et al. pioneered a linearized AST representation method to effectively extract syntactic structural features; Tang et al. integrated AST ancestor and sibling node relationships through a tree-based attention mechanism; and Guo et al. introduced ternary position encoding into AST representation and employed GraphSAGE for undirected graph encoding. However, existing AST-based deep learning methods still suffer from two major limitations. Limitation 1: Low-quality training data constrains model performance. Deep learning-based code summarization methods rely on high-quality training data that must meet requirements for syntactic validity and structural integrity. Currently widely used Java and Python benchmark datasets are derived from GitHub projects before 2018. This paper's sample evaluation found that 36% and 45% of low-quality code samples in these datasets, respectively, are present. Limitation 2: Insufficient utilization of hierarchical semantics. Serialized ASTs struggle to fully express hierarchical relationships, and existing graph neural network methods lack explicit mechanisms for processing directed graphs. This paper proposes a collaborative optimization method for code quality assessment and enhancement, and a hierarchy-aware code summarization model to enhance automatic source code summarization models. Summary of the Invention
[0004] The purpose of this invention is to solve the problems mentioned above among researchers in automatic source code summary generation, such as poor dataset quality and failure to fully utilize AST hierarchical information.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] S1. Use LLM to evaluate and optimize the code dataset, and let LLM assist in defect correction in the second stage based on the quality feedback generated in the first stage.
[0007] S2. Build an abstract syntax tree for the source code and convert it into a directed syntax graph based on the design.
[0008] S3. Send the directed grammatical graph to a directed graph neural network that can process directed graphs to extract and encode hierarchical grammatical features.
[0009] S4. Use the decoder to decode the feature representation of the encoder in S3 and predict the generated code summary.
[0010] First, we designed a two-stage LLM prompting strategy with quality feedback, including (1) instructions: guiding the LLM to perform code quality assessment and optimization. (2) constraints: enforcing the output of structured responses to standardize the result format. (3) prior knowledge: we conducted a preliminary analysis of the code on Java and Python datasets and identified five categories of low-quality issues. We provide the following five typical categories of low-quality code issues with descriptions and examples to guide GPT in evaluating and optimizing code.
[0011] Complex code: often contains deep nesting, excessive branching or looping, or redundant constructs.
[0012] Ambiguous identifiers: Variable names that do not accurately reflect their purpose or meaning.
[0013] Missing or improper code for exception handling: Lack of mechanisms to handle possible exceptions or improper handling.
[0014] Redundant code: including repeated logic, unused variables or uncalled functions, etc.
[0015] Incomplete code: lacks the necessary parts to implement certain functions.
[0016] (4) Task input: Provide the original code as input. The code dataset and prompt template are input into the LLM for data quality assessment and improvement.
[0017] The code snippet is parsed into an abstract syntax tree, and some non-leaf nodes are deleted while maintaining the integrity of the code structure. A directed syntax graph (DSG) is then constructed based on the pruned AST, replacing the original AST edges with directed parent-child edges, and adding directed code edges between leaf nodes to represent the code sequence.
[0018] In order to effectively capture the hierarchical and sequential features of the directed grammar graph, the improved directed graph convolutional network dirGCN is used as the encoder. This network is based on the improved graph convolution framework and processes directed edges through a bidirectional aggregation mechanism. Figure 1In the DirGCN part, two separate input and output aggregators AGG← and AGG→ are first used to aggregate neighbor messages for incoming and outgoing edges respectively, with independent parameter sets for the two aggregation stages. For the parent-child directed edges in the DSG, AGG← aggregates the information of the parent node, and AGG→ aggregates the information of the child node. Processing these two relationships separately helps the model identify the parent-child hierarchical relationship of the DSG; for directed leaf node edges, AGG← aggregates the context information of the code sequence, and AGG→ aggregates the context information of the code sequence. Combining the two aggregations, the context information of the code sequence can be obtained. At the kth layer, for DSG node i, its state is updated by aggregating the incoming and outgoing edges of its neighbors. In the first step, separate aggregations are performed on the inner and outer neighbors respectively:
[0019]
[0020] Where E represents the set of directed edges in DSG; represents the vector of the jth neighbor in the k-1th layer; Agg represents the vector of the jth outer neighbor in the k-1th layer; ← and Agg → Denote the aggregation function of inner neighbors and outer neighbors respectively. In the second step, the inner and outer neighbor groups are converted and aggregated to update the state of the node. The formula is as follows:
[0021]
[0022] in Represents the vector of node i output by the previous dirGCN layer; W, W ← , is a learnable weight matrix.
[0023] When all node states are updated, the state vectors are concatenated and input into the ReLU activation function for nonlinear transformation, which is formulated as follows:
[0024]
[0025] By stacking more dirGCN layers, nodes can collect information from farther distances, thereby capturing global hierarchical and sequential information. To alleviate the gradient vanishing and excessive vector offset caused by multi-layer computation, we use residual connections and layer normalization in each layer.
[0026]
[0027] Summary generation is done by the decoder and the similarity comparator (SC) module. Given summary tokens from the (k-1)th layer and the extracted DSG node representation The decoding process of the K-th Transformer layer is as follows:
[0028]
[0029] MaskAtt and Att represent masked multi-head attention and standard multi-head attention, respectively. LayerNorm represents the layer normalization operation. Finally, the decoded vector is fed into the FFN for nonlinear transformation.
[0030] After decoding, Hisum implements a multi-head attention-based copy mechanism on the decoder and encoder to generate subsequent summary tokens. The decoder loss function is defined as:
[0031]
[0032] Where N is the total number of training data, n is the sentence length of each target summary, is the jth word in the i-th sentence, represents the probability of generating the jth word.
[0033] Due to the use of the Beam Search strategy, the decoder only focuses on the local semantics of the previous words and ignores the overall semantics. Therefore, the SC module is developed to compare the difference between the generated summary p and the reference summary y, which helps the decoder module learn global semantic information. The SC module is mainly composed of Sentence Bert. After embedding, the reference summary and the predicted summary are converted into vectors respectively. and Then, calculate y e and p e In addition, a RELU activation is added after the cosine similarity to ensure that the result is greater than zero.
[0034] SC module: To solve the problem of local semantic deviation in Beam Search, the Sentence-BERT encoder is introduced to map the reference summary y and the generated summary p into vectors respectively. and Calculate cosine similarity and add ReLU activation:
[0035] S = ReLU(cos(y e ,p e ))
[0036] The loss function of the SC module is:
[0037]
[0038] Joint training: The total loss function is as follows:
[0039] L=L decoder +L SC
[0040] Finally, the training method of the hierarchical-aware code summary generation model is as follows:
[0041] The experiment was conducted in Python 3.9, using PyTorch 1.21 and CUDA 11.6, and supported by two NVIDIA 3080Ti GPUs. The embedding dimensions of the summary token and DSG node were initialized to 256. The summary generation model consisted of a 4-layer dirGCN encoder and an 8-layer decoder. The Adam optimizer was used for training, with an initial learning rate of 5e^(-4) that decayed by 5% after each round of training until it dropped to 5e^(-4). The dropout rate was set to 0.2, and the batch sizes for the Java and Python datasets were 160 and 200, respectively. Training was terminated early if the validation set performance did not improve for 10 consecutive rounds or if the total number of training rounds reached 100. Beam search (beam width 5) was used in the decoding phase.
[0042] Compared with the existing technology, the beneficial effects of the present invention are:
[0043] 1) An innovative method integrating code quality assessment and optimization is proposed. This method uses a prompt engineering mechanism based on code quality feedback to enable large language models to automatically correct code defects in the dataset during generation.
[0044] 2) We design a directed grammar graph with hierarchical semantics preservation capabilities and train a hierarchical-aware code summarization model to generate higher-quality code summaries. The effectiveness of the model is verified through experiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is the model structure diagram;
[0046] Figure 2 This is the result of ablation experiment;
[0047] Figure 3 Study the impact of encoder, decoder size and embedding size on model performance; DETAILED DESCRIPTION
[0048] The accompanying drawings are only for illustrative purposes and should not be construed as limiting this patent.
[0049] The present invention is further described below with reference to the accompanying drawings and embodiments.
[0050] Figure 1 The overall framework diagram designed for the present invention is as follows: Figure 1 As shown in Figure 2, the method can be divided into three parts: A) Data augmentation, B) Directed grammar graph construction, and C) Hierarchy-aware code summarization model.
[0051] Figure 2 This is an ablation experiment. During 100 iterations, the performance results of the model after removing various components are shown. Overall, the performance results corresponding to the complete model are more ideal and stable, which further verifies the effectiveness of the model structure.
[0052] Figure 3 To understand the impact of model size on performance, this experiment is used to find the most appropriate model size. The experiment shows that the selection of the number of encoding layers, decoding layers, and embedding size is crucial when optimizing model performance. Overall, through in-depth analysis of the above experimental results, we have come to the following important findings: (1) The scale setting of the encoding layer and decoder has a significant impact on the performance of Hisum. Too large or too small a number of layers is not conducive to improving model performance. The appropriate scale setting is particularly critical. (2) When the embedding size is set to 256, the feature extraction effect of the code reaches the best state. These findings provide us with valuable reference for optimizing model configuration and improving model performance in practical applications.
[0053] Finally, the details of the above embodiments of the present invention are merely examples for explaining the present invention. For those skilled in the art, any modifications, improvements and replacements of the above embodiments should be included in the scope of protection of the claims of the present invention.
Claims
1. A method for enhancing code summarization based on an LLM-enhanced dataset and an AST hierarchical awareness model, the method comprising the following steps: S1. Use LLM to evaluate and optimize the code dataset, and let LLM assist in defect correction in the second stage based on the quality feedback generated in the first stage. S2. Build an abstract syntax tree for the source code and convert it into a directed syntax graph based on the design. S3. Send the directed grammatical graph to a directed graph neural network encoder that can process directed graphs to extract and encode hierarchical grammatical features. S4. Use the decoder to decode the feature representation of the encoder in S3 and predict the generated code summary.
2. According to the LLM-enhanced dataset of claim 1, the specific process of S1 is: A two-stage LLM hinting strategy with quality feedback is designed, including (1) instructions: guiding the LLM to perform code quality assessment and optimization. (2) constraints: forcing the output of structured responses to standardize the result format. (3) prior knowledge. Five categories of low-quality issues are identified. Five typical low-quality code issues are provided with descriptions and examples to guide the LLM to optimize the code. (4) task input: providing the original code as input. The code dataset and hint template are input into the LLM for data quality assessment and improvement.
3. The construction of the directed syntax graph according to claim 1, wherein the specific process of S2 is as follows: The code snippet is parsed into an abstract syntax tree, and some non-leaf nodes are deleted while maintaining the integrity of the code structure. A directed syntax graph (DSG) is then constructed based on the pruned AST. The original AST edges are replaced with directed parent-child edges, and directed code edges are added between leaf nodes to represent the code sequence. The introduction of directed parent-child edges enables the graph structure to model hierarchical tree relationships. The directed code edges between leaf nodes map the code sequence to the syntax graph structure, representing the contextual dependencies of the code tokens.
4. The hierarchical perception encoder according to claim 1, wherein the specific process of S3 is: At the kth level, for DSG node i, its state is updated by aggregating its neighbors’ inbound and outbound edges. In the first step, separate aggregations are performed on the inbound and outbound neighbors respectively: Where E represents the set of directed edges in DSG; represents the vector of the jth neighbor in the k-1th layer; Agg represents the vector of the jth outer neighbor in the k-1th layer; ← and Agg → Denote the aggregation function of inner neighbors and outer neighbors respectively. In the second step, the inner and outer neighbor groups are converted and aggregated to update the state of the node. The formula is as follows: in Represents the vector of node i output by the previous dirGCN layer; W, W ← , is a learnable weight matrix. When all node states are updated, the state vectors are concatenated and input into the ReLU activation function for nonlinear transformation, which is formulated as follows: By stacking more dirGCN layers, nodes can collect information from farther distances, thereby capturing global hierarchical and sequential information. To alleviate the gradient vanishing and excessive vector offset caused by multi-layer computation, we use residual connections and layer normalization in each layer.
5. The decoder according to claim 1 generates a code summary, wherein the specific process of S4 is: Summary generation is done by the decoder and the similarity comparator (SC) module to generate the final code summary. Given the summary from the (k-1)th layer and the extracted DSG node representation The decoding process of the K-th Transformer layer is as follows: MaskAtt and Att represent masked multi-head attention and standard multi-head attention, respectively. LayerNorm represents the layer normalization operation. Finally, the decoded vector is fed into the FFN for nonlinear transformation. After decoding, Hisum implements a multi-head attention-based copy mechanism on the decoder and encoder to generate subsequent summary tokens. The decoder loss function is defined as: in, N is the total number of training data, n is the sentence length of each target summary, is the jth word in the i-th sentence, represents the probability of generating the jth word. Due to the use of the Beam Search strategy, the decoder only focuses on the local semantics of the previous words and ignores the overall semantics. Therefore, the SC module is developed to compare the difference between the generated summary p and the reference summary y, which helps the decoder module learn global semantic information. The SC module is mainly composed of Sentence Bert. After embedding, the reference summary and the predicted summary are converted into vectors respectively. and Then, calculate y e and p e In addition, a RELU activation is added after the cosine similarity to ensure that the result is greater than zero. SC module: To solve the problem of local semantic deviation in Beam Search, the Sentence-BERT encoder is introduced to map the reference summary y and the generated summary p into vectors respectively. and Calculate cosine similarity and add ReLU activation: S=ReLU(cos(y e ,p e )) The loss function of the SC module is: Joint training: The total loss function is as follows: L=L decoder +L SC 。