A code summary generation method based on multi-granularity feature fusion
Through the multi-grained feature fusion and cross-attention mechanism, the deviation and information fragmentation problems of existing code summary generation methods in semantic extraction are solved, and the high accuracy and readability of code summary generation is achieved, reducing the computational overhead.
Patent Information
- Application Number
- CN202510260828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing code summary generation methods have biases in semantic extraction, poor scalability, and information fragmentation, making it difficult to generate code summary with high accuracy and readability.
Using a method based on multi-grained feature fusion, three granular features, namely word sequence, abstract syntax tree and control flow diagram, are extracted, and a multi-grained feature encoder is built to achieve feature fusion through the cross attention mechanism, and finally input into the Transformer decoder to generate a code summary.
Improves the accuracy and readability of code summary, reduces computational overhead, and solves the information fragmentation problem caused by single-grained processing.
Smart Images

Figure CN119739857B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the intersection of software engineering and deep learning, and in particular relates to a code summary generation method based on multi-granularity feature fusion. Background Art
[0002] In recent years, technological innovation, open source ecology and changing user needs have driven the rapid development of software systems, which increasingly replace traditional interaction scenarios and solve a wider range of practical problems. However, as the scale and complexity of software systems have increased significantly, developers need to spend more energy and time to maintain existing functions rather than develop new functions. In this regard, generating code summaries is a feasible solution, which can create concise natural language descriptions for source code to help developers quickly understand software systems. However, code summaries in software systems are often unreadable or completely missing, and cannot achieve the intended purpose. Automatic code summarization technology effectively alleviates this problem by generating high-quality summaries without manually reading all the code. At first, early code summary generation methods adopted rule-based and template-based methods. The set template is filled according to a set of predetermined information extraction rules to generate summaries. Later, based on the assumption of the "naturalness" of the code, researchers began to use deep learning-based methods to generate more accurate code summaries. Most of the existing studies choose abstract syntax trees (ASTs) as feature representations of code. However, the above methods all have certain disadvantages. The rule-based and template-based methods mainly extract surface semantics, and inappropriate identifier naming will seriously affect the accuracy of keyword extraction and lack scalability. Methods based on deep learning usually only focus on AST, but due to the sparse information in AST nodes and the large number of nodes, the information becomes fragmented. This makes it impossible to provide the model with continuous grammatical information and a comprehensive view of control dependencies. In view of the problems existing in the current methods, the present invention proposes a code summary generation method based on multi-granularity feature fusion, which can improve the accuracy and readability of the generated code summary and reduce the computational overhead generated in the process. Summary of the invention
[0003] Purpose of the invention: The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art, overcome the semantic extraction bias, poor scalability, and information fragmentation of the existing models in the process of code summary generation, and build a new multi-granularity fusion model to generate code summaries for various types of texts.
[0004] The present invention specifically provides a code summary generation method based on multi-granularity feature fusion, comprising the following steps:
[0005] Step 1: Preprocess the samples in the dataset and process the source code into three granular features: word sequence, abstract syntax tree, and control flow graph to obtain the overall semantic information of the code;
[0006] Step 2: Build a source code multi-granularity feature encoder. Based on the data scale and data structure differences of the three granularity features of word sequence, abstract syntax tree, and control flow graph, set different encoding methods and obtain the context vectors of the three granularity features.
[0007] Step 3, the outputs of feature encoders with different granularities are subjected to cross-attention calculation to achieve granular feature fusion;
[0008] Step 4: Input the fused word sequence fusion features with abstract syntax tree AST features and control flow graph fusion features into the Transformer-based decoder, and output the predicted words at the current time step through the output probabilities of all words in the vocabulary at each time step to form a summary.
[0009] Furthermore, in step 1, the samples in the data set are annotated and filtered, and abnormal samples are cleaned to ensure the validity and accuracy of the data set to better support model training. The present invention processes the source code into three granular features: word sequence, abstract syntax tree, and control flow graph through preprocessing operations to obtain the overall semantic information of the code, which specifically includes the following steps:
[0010] Step 1.1, obtain the word-gram sequence features of the source code: mark the function names of the word-gram sequences in the source code, use the function names as the natural language NL part of the sequence, replace the marked positions with "func name", and then use the sequence after source code word segmentation as the program language PL part. The natural language NL part and the program language PL part are combined as the overall word-gram sequence;
[0011] Step 1.2, obtain the abstract syntax tree features of the source code: For Python language, the abstract syntax tree AST string representation of the source code is obtained through the ast.dump() method of the abstract syntax tree AST module. In order to better generate a tree structure representation of the AST character, a method for generating an abstract syntax tree AST graph structure is adopted to generate a tree structure representation of the abstract syntax tree AST string: the null value attributes in the abstract syntax tree AST string are screened out and deleted to reduce redundant information. Then, the abstract syntax tree AST string is divided into nodes (when the string element is followed by "=value", "=number" or "=word" and other formats, it is divided into separate nodes), and the brackets are retained to obtain the structural hierarchy information between nodes; finally, according to the source code structure represented by the brackets, each divided node is pointed to the node's unique parent node, thereby constructing a complete tree structure; this method can effectively reduce the interference of invalid information and generate a more streamlined AST graph structure, which is convenient for subsequent analysis and processing.
[0012] For the Java language, use the Javalang module to get the string representation of the source code abstract syntax tree AST;
[0013] Step 1.3, obtain the control flow graph features of the source code: introduce other identifiable node types based on the pycfg module in Python, such as try, with, raise, etc., and correct potential error paths, so that the generated control flow graph CFG (Control Flow Graph) graph structure consists of statement-based nodes, connecting all nodes to reconstruct the content of the code; for the Java language, use the anger tool to reprocess the abstract syntax tree AST features, and combine it with the Soot open source tool to obtain the control flow graph CFG information;
[0014] Step 1.4, in order to alleviate the OOV problem, the Byte Pair Encode (BPE) algorithm is used to segment all source code features, namely the three granular features of word sequence, abstract syntax tree and control flow graph, and map them to unique indexes in the vocabulary. The vocabulary index is finally embedded into a dense vector. For the three granular features of word sequence, abstract syntax tree and control flow graph, the maximum length of the feature sequence is set to 300, 500 and 40 respectively. Sequences shorter than the maximum length are <pad>Mark for padding and truncate sequences greater than the maximum length;
[0015] For a word sequence, add <cls>and <sep>The corresponding index in the vocabulary to meet the input format of the pre-trained model;
[0016] For the abstract syntax tree AST, the nodes whose word unit sequence number after word segmentation is greater than one are aggregated into a single vector representation using the maximum pooling operation. In the present invention, a node refers to a node of a tree structure formed by dividing the abstract syntax tree AST string into nodes;
[0017] For the control flow graph CFG, the CodeBERT model is used to pre-train each node as a word sequence. <cls>The embedding vector of a token represents the overall meaning of the node.
[0018] During the training process, the model can bypass the embedding layer that requires training a large number of parameters, obtain the representation of each node of each single vector, and improve the training speed by pre-embedding node features.
[0019] Furthermore, step 2 includes the following steps:
[0020] Step 2.1, build a word sequence token encoder;
[0021] Step 2.2, build the abstract syntax tree AST encoder;
[0022] Step 2.3, build the control flow graph CFG encoder.
[0023] Step 2.1 includes: obtaining basic grammatical information by encoding the word sequence, using the CodeBERT model as the word sequence encoder, and the encoding process is expressed as:
[0024] ,
[0025] ,
[0026] in, It is the context vector of the word sequence directly output by the CodeBERT model pre-training. is the context vector of the word sequence output after training with the fully connected layer. is the matrix composed of the word-unit sequence feature context vector after fine-tuning by the fully connected layer, l represents the word-unit sequence length, is the word embedding dimension, It is a fully connected layer, which is used to further fine-tune the general pre-trained model through a learnable fully connected layer to suit the current task after encoding.
[0027] Step 2.2 includes: serializing the nodes of the abstract syntax tree AST by pre-order traversal, and storing the index sequence of the leaf nodes in the abstract syntax tree AST in the pre-order traversal sequence through an array, defining Matrix and The matrix stores the parent-child relationship and sibling relationship between nodes respectively, and N is the total number of nodes. From top to bottom and from left to right are defined as positive directions. If the i-th node in the sequence is the grandfather node of the j-th node, then the shortest path distance between the i-th node and the j-th node is 2, recorded as = 2, and at the same time = -2;
[0028] In order to further constrain the aggregation range of nodes and reduce the spatial complexity of processing data, a maximum relative distance threshold k is set for two types of relationships, namely, parent-child relationship and sibling relationship. The value is generally 7. If the relative distance between two nodes exceeds k, it is determined that there is no relationship between the two nodes, and the corresponding position in the matrix is set to infinity. The specific rules for defining the ancestor-descendant relationship matrix are as follows:
[0029] ,
[0030] ,
[0031] in, Indicates the relative distance between node i and node j when node i is a descendant node or ancestor node of node j; Indicates the horizontal relative distance between node i and node j when node i and node j are sibling nodes; Represents the shortest path distance between the i-th node and the j-th node in the parent-child relationship. Represents the shortest path distance between the i-th node and the j-th node in the parent-sibling relationship. Next, the relative distance information between nodes stored in the matrix is converted into relative position embedding, using the intermediate parameter Unified Representation or , and according to The value of to obtain the unique relative distance index:
[0032] ,
[0033] in, is the relative distance index between the ith node and the jth node. If the relative distance index between a node and the ith node is greater than 0, the node is determined to be a node with a strong relationship with the ith node.
[0034] In the above process, since the relative distance between nodes has positive and negative directions, the embedding methods of relative distances with the same absolute value are different. In addition, the offset value P+1 is applied to ensure that the final index must be a positive value. The following formula is used to calculate the inner product between the Q vector of the i-th node (including two Q vectors based on content embedding and relative distance) and the K vector of the j-th node (including two K vectors based on content embedding and relative distance): ,
[0035] in represents the inner product between the Q vector of the i-th node and the K vector of the j-th node, and Represent the embedding of the i-th node and the j-th node respectively; Q is the content-based query function, K is the content-based key function, Based on the relative distance index The query function, Based on relative position The key function of is, T represents transpose. In the above way, the relative position index is converted into an independent relative position vector, and the influence of the relative position between nodes on the final correlation between nodes is evaluated through additional content-to-position and position-to-content calculations.
[0036] The present invention uses the decoupled attention mechanism to calculate the attention coefficient of a node and its strong relationship node. The formula is as follows:
[0037] ,
[0038] Among them, i is the overall attention coefficient of node i and node j. i,j It consists of three items, and the scaling factor is adjusted accordingly. The formula defines both the value function V based on content and the value function V based on relative distance. P . And during the update of the last node feature only pay attention to Nodes greater than 0.
[0039] Step 2.3 includes: the number of nodes in the CFG of the source code is relatively small, the in-and-out degrees of non-branch and non-node nodes are very limited, and the in-and-out degrees of branch nodes (such as loops and judgment nodes) are relatively large. The decoupled attention operation cannot significantly reduce the number of nodes processed. Therefore, the graph convolutional network GCN is first used to perform preliminary feature fusion on the control flow graph CFG embedding matrix. The formula is:
[0040] ,
[0041] ,
[0042] Among them, the above formula is the overall description of matrix update, The feature matrix representing the control flow graph CFG node of layer l, represents the updated CFG node feature matrix, Represents the sigmoid mathematical function, is the adjacency matrix containing the node connection information in the control flow graph CFG, for The degree matrix of represents the inverse square root of the degree matrix, Represents a learnable weight matrix, which is used to transform the node feature representation of each layer. The following formula describes the update of specific node features. represents the set of neighbor nodes of node i, j can refer to any neighbor node of node i, represents the degree of node i, represents the degree of node j, represents the feature vector of node j at the lth iteration, Indicates the +1st iteration of the feature vector of node i; It is used to achieve a specific normalization effect so that neighboring nodes with higher degrees contribute less to the feature integration of the node.
[0043] Due to the high density of semantic information in CFG nodes and the implicit control and data dependency paths that may exist between distant nodes, after the initial aggregation of node features through the GCN layer, the node features are updated through the multi-head self-attention module (the attention mechanism further captures the global dependencies between CFG features), the formula is:
[0044] ,
[0045] in, represents the characteristic matrix of the control flow graph CFG, V represents the value function, It is the final global representation of the control flow graph CFG node sequence.
[0046] Step 3 includes the following steps:
[0047] Step 3.1, for the abstract syntax tree feature matrix, after the decoupled attention encoding of the AST encoder, the leaf node feature matrix is extracted from the abstract syntax tree AST overall feature matrix through the pre-stored leaf node index information, and the leaf node feature matrix is used for cross-attention fusion with the context representation vector generated in the word unit sequence encoder;
[0048] Step 3.2: fuse the features of different granularities. Use the following formula to fuse the intermediate representations of the three different granularity features together:
[0049] ,
[0050] V ,
[0051] in, represents the context vector after being processed by the word sequence encoder, Represents the context vector after being processed by the control flow graph CFG encoder, Represents the context vector after being processed by the abstract syntax tree AST encoder, Represents a sequence of leaf nodes extracted from the abstract syntax tree AST. is the feature vector after the word sequence granularity feature and the abstract syntax tree granularity feature are integrated. It is the feature vector after the fusion of the control flow graph granularity features and the abstract syntax tree granularity features. Softmax is a normalized exponential function. By aggregating the features of three different granularities, namely, word sequence, abstract syntax tree, and control flow graph, the final global representation of the fused code features is obtained.
[0052] Step 4 includes the following steps:
[0053] Step 4.1, build a decoder, the decoder includes N stacked decoder layers, each decoder layer is divided into three parts, the first part includes masked multi-head self-attention, residual connection and normalization, wherein the masked multi-head self-attention ensures that the decoder part in the method of the present invention only depends on the information output before the current time step (a unit used to describe each moment in the sequence data, which is a word element index in the code summary output at each time step in this generation method); the second part includes cross attention, residual connection and normalization, wherein the cross attention is used to process the final global representation obtained in step 3; the third part includes a feedforward network, residual connection and normalization, which is used to capture deep features;
[0054] In step 4.2, a projection layer with a dimension from the embedding dimension to the vocabulary capacity and a softmax activation function are used to obtain the output probability of all word units in the vocabulary at each time step, and the index corresponding to the highest probability in each time step is selected as the final output in this time step. The final output is the vocabulary index, and the vocabulary index generates the final code summary based on the vocabulary.
[0055] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0056] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0057] The present invention has the following beneficial effects: 1. In the data preprocessing stage, relevant statistical principles are used to screen word sequence data, exclude abnormal data, ensure the reliability of the data, standardize the format of word sequence data and partially add search indexes, making the data easier to process and compare, greatly improving the quality of the data set.
[0058] 2. A code summary generation method based on multi-granularity feature fusion is proposed. Through multi-granularity fusion, the accuracy of keyword extraction, lack of scalability, information fragmentation and other problems that arise from single-granularity processing are avoided. While improving the accuracy and readability of the generated code summary, the computational overhead generated in the process is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flow chart of the method of the present invention.
[0060] Figure 2 It is a model diagram of the present invention. DETAILED DESCRIPTION
[0061] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0062] like Figure 1 As shown, the embodiment of the present invention provides a code summary generation method based on multi-granularity feature fusion, Figure 2 The specific structure of the model involved in the method of the present invention is shown, wherein: Figure 2 BLEU, METEOR and ROUGE in the evaluation part represent three machine translation evaluation indicators, BLEU is a measurement indicator based on precision, METEOE is a measurement indicator based on precision and recall, ROUGE is a measurement indicator based on recall, RoBERTa in the embedding part and Transformer in the model part are both existing model names, and the method of the present invention includes: obtaining the word-unit sequence characteristics of the source code through a preprocessing step. Mark the function names of the word-unit sequence in the source code, and use the function name as the NL part of the sequence. In the source code, the marked positions are uniformly replaced with "func name", and then the sequence after the source code is segmented is regarded as the PL part. The two parts are merged as the overall word-unit sequence.
[0063] The abstract syntax tree features of the source code are obtained through the preprocessing step. For Python language, the AST string representation of the source code is obtained through the ast.dump() method of the AST module; for Java language, it is obtained through the Javalang module. In order to better generate a tree structure representation of AST characters, a method for generating an AST graph structure is designed: first, the null value attributes are filtered out and deleted to reduce redundant information. Then, the AST string is divided into nodes (when the string element is followed by "=value", "=number" or "=word" and other formats, it is divided into separate nodes), and the brackets are retained to obtain the structural hierarchy information between nodes. Finally, according to the source code structure represented by the brackets, each divided node is pointed to its unique parent node to construct a complete tree structure. This method can effectively reduce the interference of invalid information and generate a more streamlined AST graph structure, which is convenient for subsequent analysis and processing.
[0064] The control flow graph features of the source code are obtained through the preprocessing step. Based on the pycfg module in Python, other identifiable node types are introduced, and potential error paths are corrected, so that the generated CFG graph structure consists of statement-based nodes, and all nodes are connected to reconstruct the content of the code. For the Java language, CFG information is obtained by further processing the AST and combining it with the Soot open source tool.
[0065] In order to alleviate the OOV problem, the present invention uses the BPE algorithm to segment all features of the source code and map them to unique indexes in the vocabulary, which are finally embedded into dense vectors through the vocabulary index. For each granular feature, the maximum length of the feature sequence is set, and sequences shorter than the maximum length are <pad>The token is filled, and if it is larger, it is truncated. For a word sequence, add it at the beginning and end of the index sequence. <cls>and <sep>The corresponding word table index meets the input format of the pre-trained model. For AST, the maximum pooling operation is used to aggregate nodes with more than one token after word segmentation into a single vector representation. For CFG, the CodeBERT model is used to pre-train each node as a word sequence, with the first output <cls>The labeled embedding vector represents the overall meaning of the node. During the training process, the model can bypass the embedding layer that requires training a large number of parameters, obtain the representation of each node in each single vector, and improve the training speed by pre-embedding node features.
[0066] Build a token encoder. The most basic grammatical information is obtained by encoding the word sequence. This paper uses the CodeBERT model as the token encoder. The encoding process is as follows:
[0067] ,
[0068] ,
[0069] in, is the context vector of the word sequence, It is a fully connected layer, which is used to further fine-tune the general pre-trained model through a learnable fully connected layer to suit the current task after encoding.
[0070] Build an AST encoder. Use pre-order traversal to serialize the nodes of the AST, and use an array to store the index sequence of the leaf nodes in the pre-order traversal sequence in the AST. Definition Matrix and The matrix stores the ancestor-descendant relationship and sibling relationship between nodes respectively, and N is the total number of nodes. Define top-to-bottom and left-to-right as positive directions. If the i-th node in the sequence is the grandfather node of the j-th node, the shortest path distance between the two nodes is 2, recorded as = 2, corresponding to = -2. In order to further constrain the aggregation range of nodes and reduce the spatial complexity of processing data, a maximum relative distance threshold K is set for the two types of relationships. If the relative distance between two nodes exceeds K, it is considered that there is no relationship between them, and the corresponding position in the matrix is set to infinity. The specific rules for defining the ancestor-descendant relationship matrix are as follows:
[0071] ,
[0072] ,
[0073] Next, the relative distance information between nodes stored in the matrix is converted into relative position embedding. Using the intermediate parameters Unified Representation or , and according to The value of to obtain the unique relative distance index:
[0074] ,
[0075] Nodes whose relative position index with respect to node i is greater than 0 are considered to have a strong relationship with node i. In the above process, since the relative distance between nodes has positive and negative directions, the embedding methods of relative distances with the same absolute value are different. In addition, the offset value P+1 is applied to ensure that the final index is positive. The process of calculating the correlation between node i and node j is as follows:
[0076] ,
[0077] in and Represent the embedding of the i-th node and the j-th node respectively. Q is the content-based query function, K is the content-based key function, Based on relative position The query function, Based on relative position The key function of is, T represents transpose. Through the above formula, the relative position index is converted into an independent relative position vector, and the influence of the relative position between nodes on the final correlation between nodes is evaluated through additional content-to-position and position-to-content calculations.
[0078] By calculating the attention coefficient of a node and its strong relationship nodes, the present invention uses the decoupled attention mechanism to update the original features, only considering the sibling relationship and parent-child relationship of each node, and connecting the output results. Since three items are calculated in the process of defining the matrix, the scaling factor is adjusted accordingly. The present invention defines both a value function V based on content and a value function V based on relative distance. P . And during the update of the last node feature only pay attention to Nodes greater than 0. The decoupled attention mechanism is used to calculate the attention coefficient of a node and its strong relationship nodes. The formula is as follows:
[0079] ,
[0080] Build a CFG encoder. The number of nodes in the CFG of the source code is relatively small, the in-degree of non-branch and non-node nodes is very limited, and the in-degree of branch nodes (such as loops and judgment nodes) is relatively large. The decoupled attention operation cannot significantly reduce the number of nodes to be processed. Based on this, we first use the graph convolutional network (GCN) to perform preliminary feature fusion on its embedding matrix. The GCN processing process is as follows. The two formulas describe the same propagation mechanism at different granularities:
[0081] ,
[0082] ,
[0083] The above formula is an aggregate description of the entire graph. (l) represents the feature matrix of the CFG node at level l, H (l+1) represents the updated node feature matrix, is the adjacency matrix containing the node connection information in CFG, for The degree matrix of Represents the inverse square root of the degree matrix. The following formula describes the update of a specific node. A specific normalization effect is achieved so that neighboring nodes with higher degrees contribute less to the feature integration of the node. Represents a learnable weight matrix used to transform the node feature representation of each layer.
[0084] Due to the high density of semantic information in CFG nodes and the possible implicit control and data dependency paths between distant nodes, after the initial aggregation of node features through the GCN layer, the multi-head self-attention module will further update the node features. The attention mechanism further captures the global dependency between CFG features. The process is as follows:
[0085] ,
[0086] in, represents the characteristic matrix of CFG, and V represents the value function; It is the final global representation of the control flow graph CFG node sequence.
[0087] For AST structural features, after decoupled attention encoding, the word-unit sequence feature matrix encoded based on AST features is extracted through the pre-stored leaf node index information. This matrix is used for cross-attention fusion with the context representation vector generated in the token encoder.
[0088] The construction of the cross-attention module follows the principle of integrating high-granularity information into low-granularity features, fusing the intermediate representations of three different granularity features together. The specific steps are as follows:
[0089] ,
[0090] V ,
[0091] in, Represents the context vector after being processed by the token encoder, represents the context vector after being processed by the CFG encoder, represents the context vector after being processed by the AST encoder, It represents the leaf node sequence extracted from the AST feature. By aggregating features of three different granularities, the final global representation of the fused code features is finally obtained.
[0092] Build a decoder. The decoder consists of N stacked decoder layers, each of which is divided into three parts. The first part includes masked multi-head self-attention, residual connection and normalization, where the mask mechanism ensures that the model only relies on the information that has been output before the current time step; the second part includes cross attention, residual connection and normalization, where the cross attention model is used to process the final fusion representation obtained in step three, and these functions are combined with the functions of the previous part; the third part includes a feedforward network, residual connection and normalization to further capture deep features;
[0093] Based on the previous step, using a The linear layer of dimension and the softmax activation function calculate the occurrence probability of each word in the vocabulary in the current time step, and take the word corresponding to the highest probability index as the output word of this time step.
[0094] In order to evaluate the performance of the method of the present invention, a series of comparative experiments were conducted on eight baseline models. According to the results of the comparative experiments, the model of the present invention has achieved the highest score in each key indicator. For example, compared with the Dual model model based on code sequence, the model of the present invention exceeds the Dual model model by 10.3%, 16.14% and 5.35% in the three indicator scores of BLUE-4, Meteor and Rouge-L respectively; in the model based on code structure features, the model of the present invention is more advantageous than the DeepCom model, and has achieved high scores of 17.71%, 29.79% and 7.23% in the three indicator scores of BLUE-4, Meteor and Rouge-L; compared with the CSA-Trans model based on tree structure modeling AST, on the Java dataset, the three indicator scores of BLUE-4, Meteor and Rouge-L increased by 1.67%, 1.83% and 0.67%, respectively, and on the Python dataset, they increased by 1.56%, 2.72% and 2.15%, respectively. In general, the model of the present invention has significantly improved the code summary generation quality index compared with other current code summary generation models.
[0095] The present invention provides a code summary generation method based on multi-granularity feature fusion. There are many methods and ways to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.< / cls> < / sep> < / cls> < / pad> < / cls> < / sep> < / cls> < / pad>
Claims
1. A code summary generation method based on multi-granularity feature fusion, characterized in that: The following steps are involved: Step 1: Preprocess the samples in the dataset and process the source code into three granular features: word sequence, abstract syntax tree, and control flow graph to obtain the overall semantic information of the code; Step 2: Build a source code multi-granularity feature encoder. Based on the data scale and data structure differences of the three granularity features of word sequence, abstract syntax tree, and control flow graph, set different encoding methods and obtain the context vector of each granularity feature. Step 3, the outputs of feature encoders with different granularities are subjected to cross-attention calculation to achieve granular feature fusion; Step 4: Input the fused word sequence fusion features with abstract syntax tree AST features and control flow graph fusion features into the Transformer-based decoder, output the predicted words at the current time step through the output probabilities of all words in the vocabulary at each time step, and finally form a summary; Step 1 includes the following steps: Step 1.1, obtain the word-gram sequence features of the source code: mark the function names of the word-gram sequences in the source code, use the function names as the natural language NL part of the sequence, replace the marked positions with func name, and then use the sequence after source code word segmentation as the program language PL part. The natural language NL part and the program language PL part are combined as the overall word-gram sequence; Step 1.2, obtain the abstract syntax tree features of the source code: For the Python language, obtain the abstract syntax tree AST string representation of the source code through the ast.dump() method of the abstract syntax tree AST module, and use a method for generating an abstract syntax tree AST graph structure to generate a tree structure representation of the abstract syntax tree AST string: filter out and delete the null value attributes in the abstract syntax tree AST string, divide the abstract syntax tree AST string into nodes, and retain brackets to obtain the structural hierarchy information between nodes; according to the source code structure represented by the brackets, point each divided node to the node's only parent node, thereby constructing a complete tree structure; For the Java language, use the Javalang module to get the string representation of the source code abstract syntax tree AST; Step 1.3, obtain the control flow graph features of the source code: introduce identifiable node types based on the pycfg module in Python, and correct potential error paths, so that the generated control flow graph CFG graph structure consists of statement-based nodes, connecting all nodes to reconstruct the content of the code; for the Java language, use the anger tool to reprocess the abstract syntax tree AST features, and combine with the Soot open source tool to obtain the control flow graph CFG information; Step 1.4: Use the byte pair encoding (BPE) algorithm to segment all source code features, i.e., word sequence, abstract syntax tree, and control flow graph, and map them to unique indexes in the vocabulary. Finally, embed them into dense vectors through the vocabulary index. For the three granular features of word sequence, abstract syntax tree, and control flow graph, set the maximum length of the feature sequence respectively. Sequences shorter than the maximum length are counted as <pad> Mark for padding and truncate sequences greater than the maximum length;< / pad> For a word sequence, add <cls>and <sep> The corresponding index in the vocabulary to meet the input format of the pre-trained model;< / sep> < / cls> For the abstract syntax tree AST, the maximum pooling operation is used to aggregate the nodes whose word sequence number after word segmentation is greater than one into a single vector representation; For the control flow graph CFG, the CodeBERT model is used to pre-train each node as a word sequence. <cls> The embedding vector of a token represents the overall meaning of the node.< / cls> 2. The method according to claim 1, characterized in that: Step 2 includes the following steps: Step 2.1, build a word sequence token encoder; Step 2.2, build the abstract syntax tree AST encoder; Step 2.3, build the control flow graph CFG encoder.
3. The method according to claim 2, characterized in that Step 2.1 includes: obtaining basic grammatical information by encoding the word sequence, using the CodeBERT model as the word sequence encoder, and the encoding process is expressed as: in, It is the context vector of the word sequence directly output by the CodeBERT model pre-training. is the context vector of the word sequence output after training with the fully connected layer. is the matrix composed of the word sequence feature context vector after fine-tuning by the fully connected layer, l represents the word sequence length, d model is the word embedding dimension, W t is a fully connected layer.
4. The method according to claim 3, characterized in that Step 2.2 includes: serializing the nodes of the abstract syntax tree AST by pre-order traversal, and storing the index sequence of the leaf nodes in the abstract syntax tree AST in the pre-order traversal sequence through an array, defining S∈R N×N Matrix and P∈R N×N The matrix stores the parent-child relationship and sibling relationship between nodes respectively, N is the total number of nodes; from top to bottom and from left to right are defined as positive directions. If the i-th node in the sequence is the grandfather node of the j-th node, then the shortest path distance between the i-th node and the j-th node is 2, denoted as p ji =2, and p ij = -2; For two types of relationships, namely, parent-child relationship and sibling relationship, a maximum relative distance threshold k is set. If the relative distance between two nodes exceeds k, it is determined that there is no relationship between the two nodes, and the corresponding position in the matrix is set to infinity. The specific rules for defining the ancestor-descendant relationship matrix are: Where PAR(i, j) represents the relative distance between node i and node j when node i is a descendant node or ancestor node of node j; SIB(i, j) represents the horizontal relative distance between node i and node j when node i and node j are sibling nodes; p ij Represents the shortest path distance between the i-th node and the j-th node in the parent-child relationship, s ij Represents the shortest path distance between the i-th node and the j-th node in the parent-sibling relationship. Next, the relative distance information between nodes stored in the matrix is converted into relative position embedding, using the intermediate parameter r ji Unified representation p ji or ji , and according to r ji The value of to obtain the unique relative distance index: Wherein, δ(i, j) is the relative distance index between the i-th node and the j-th node. If the relative distance index between a node and the i-th node is greater than 0, the node is determined to be a node having a strong relationship with the i-th node. The correlation between the i-th node and the j-th node is calculated using the following formula: where a i , j represents the correlation between the i-th node and the j-th node, x i With x j Represent the embedding of the i-th node and the j-th node respectively; Q is the content-based query function, K is the content-based key function, is a query function based on the relative distance index δ(i, j), is a key function based on the relative position δ(i, j), and T represents the transpose.
5. The method according to claim 4, characterized in that Step 2.3 includes: using the graph convolutional network GCN to perform preliminary feature fusion on the control flow graph CFG embedding matrix, the formula is: Among them, H (l) The feature matrix of the control flow graph CFG node representing layer l, H (l+1) represents the updated CFG node feature matrix, σ represents the sigmoid mathematical function, is the adjacency matrix containing the node connection information in the control flow graph CFG, for The degree matrix of represents the inverse square root of the degree matrix, W (l) represents a learnable weight matrix; N(i) represents the set of neighbor nodes of node i, j refers to any neighbor node of node i, d i represents the degree of node i, d j represents the degree of node j, represents the feature vector of node j at the lth iteration, represents the feature vector of node i at the l+1th iteration; The node features are updated through the multi-head self-attention module, and the formula is: Among them, x c represents the feature matrix of the control flow graph CFG, V represents the value function, H cfg It is the final global representation of the control flow graph CFG node sequence.
6. The method according to claim 5, characterized in that Step 3 includes the following steps: Step 3.1, for the abstract syntax tree feature matrix, after the decoupled attention encoding of the AST encoder, the leaf node feature matrix is extracted from the abstract syntax tree AST overall feature matrix through the pre-stored leaf node index information, and the leaf node feature matrix is used for cross-attention fusion with the context representation vector generated in the word unit sequence encoder; Step 3.2: fuse the features of different granularities. Use the following formula to fuse the intermediate representations of the three different granularity features together: Among them, H Tok represents the context vector after being processed by the word sequence encoder, H Cfg represents the context vector after being processed by the control flow graph CFG encoder, H Ast Represents the context vector after being processed by the abstract syntax tree AST encoder, H Ast_tok Represents the leaf node sequence extracted from the abstract syntax tree AST, H cross_Token H is the feature vector after the word sequence granularity feature and the abstract syntax tree granularity feature are integrated. cross_AST It is the feature vector after the fusion of the control flow graph granularity features and the abstract syntax tree granularity features. Softmax is a normalized exponential function. By aggregating the features of three different granularities, namely, word sequence, abstract syntax tree, and control flow graph, the final global representation of the fused code features is obtained.
7. The method according to claim 6, characterized in that Step 4 includes the following steps: Step 4.1, build a decoder. The decoder consists of N stacked decoder layers. Each decoder layer is divided into three parts. The first part includes masked multi-head self-attention, residual connection and normalization. The second part includes cross attention, residual connection and normalization, where cross attention is used to process the final global representation obtained in step 3; the third part includes feedforward network, residual connection and normalization to capture deep features; In step 4.2, a projection layer with a dimension from the embedding dimension to the vocabulary capacity and a softmax activation function are used to obtain the output probability of all word units in the vocabulary at each time step, and the index corresponding to the highest probability in each time step is selected as the final output in this time step. The final output is the vocabulary index, and the vocabulary index generates the final code summary based on the vocabulary.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Automatic annotation generation method based on multi-modal code representation
CN115756597A
Cited By
Code annotation generation method driven by mixed entropy
CN121435935A