An Automated Software Refactoring Method Based on Transfer Learning
By combining BERT, GraphSAGE, Code2Vec and Transformer models, the tag mapping matrix and feature migration are constructed, and the integration of code structure and semantic information is achieved, solving the shortcomings of the existing software reconstruction methods in terms of accuracy and automation, and achieving efficient and accurate software reconstruction and cross-task knowledge sharing.
Patent Information
- Application Number
- CN202510407372.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing software reconstruction methods are insufficient in terms of accuracy and automation, and it is difficult to effectively integrate code structure and semantic features, and lack cross-task knowledge sharing capabilities.
Using a transfer learning-based method, combined with the BERT model, Graph Neural Network (GraphSAGE), Code2Vec and Transformer models, the integration of code structure and semantic information is achieved by constructing label mapping matrix and feature migration, and end-to-end automated reconstruction is carried out.
It realizes efficient and accurate software reconstruction prediction, improves the generalization ability and scalability of the model, adapts to the needs of different tasks and fields, solves the problem of insufficient reconstruction data, and realizes systematic cross-task knowledge sharing.
Smart Images

Figure CN119917163B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software engineering, and particularly relates to an automated software refactoring method based on transfer learning. Background Art
[0002] Software refactoring is an important activity in software engineering. In recent years, researchers have begun to explore the application of deep learning to various aspects of software refactoring. Existing literature discusses different automatic refactoring methods, whose purpose is to help practitioners detect code smells. Some of them can even suggest refactoring activities that practitioners should perform to eliminate detected code smells and reduce technical debt. Most methods are either rule-based, use search-based algorithms or machine learning methods. At the same time, in order to meet the needs of model training, researchers use traditional refactoring mining tools (such as RefactoringMiner, RefDiff, etc.) combined with deep learning technology to automatically identify and extract refactoring instances from existing code bases. Create a well-labeled dataset to train a deep learning model. Here are some representative existing software refactoring methods: Heuristic-based smell detection: JMove proposed by Sale et al. They use dependency-based metrics to determine whether method m should be moved to C by calculating the similarity between the dependency set of method m and other methods in class C. They compared JMove with related tools in terms of precision, recall, and runtime performance and found that JMove performs better when processing large methods, and in some cases outperforms JDeodorant and inCode. JMove's recommendation algorithm relies on the similarity of static dependencies, and its method for judging feature envy is relatively limited, which leads to its poor performance in some cases. In some cases, the real refactoring opportunities cannot be identified. In addition, the similarity calculation based on method calls and field accesses is not comprehensive enough. The responsibilities and behaviors of methods depend not only on dependencies, but also on their logical structure and context. It is difficult to fully capture the actual semantics of the method with the similarity function alone, which may lead to incorrect refactoring suggestions; Code smell detection based on federated learning: ALAWAD et al. use FedCSD to simulate the collaboration of multiple software development companies to improve the quality of their software development projects without sharing their code settings. In each client of federated learning, the long short-term memory (LSTM) algorithm was used in the study to simulate 10 different companies, representing these companies by dividing the dataset into different blocks. , the flow of information is regulated by input gate, forget gate and output gate to ensure that unimportant information is forgotten, thereby improving the accuracy of detection. Experimental results show that FedCSD performs well in detecting code smells, and its model is superior to traditional centralized methods in terms of accuracy, stability and learning ability. Through federated learning, FedCSD can capture and learn new changes in coding culture and provide stronger model generalization capabilities. However, the code has hierarchical and structured characteristics (such as functions, classes and loop nesting), while LSTM is good at processing sequential data and has limited ability to represent tree or graph-like structured code. This limitation will cause it to be insufficient in capturing code context information, affecting the accuracy of code smell detection;LLM and IDE Combined to Handle Refactoring Issues: BAO et al. proposed EM-Assist, which is a large language model (LLM)-based approach focusing on generating refactoring suggestions for extraction methods. It analyzes the given long method code, identifies code segments suitable for extraction, and thus suggests method separation to developers. This process not only relies on the LLMs' learning of code patterns but also combines a series of rules to ensure the effectiveness of the generated suggestions. To prevent code failures, EM-Assist filters out extraction method suggestions that cannot be compiled, ensuring that details such as return values and exception handling are properly addressed. Meanwhile, with program slicing technology, EM-Assist avoids generating small methods with excessive parameters, optimizing the code structure and better conforming to developers' refactoring habits. Experimental results show that EM-Assist performs excellently in both public datasets and open-source projects, being more effective than existing static analysis tools and machine learning-based models, and can provide developers with more reliable and practical code refactoring solutions. However, despite multiple rounds of filtering and enhancement after EM-Assist generates suggestions, the final suggestions still rely on developers' feedback and decisions, which may lead to developers spending additional time and effort to evaluate and select appropriate suggestions in practical applications, conflicting with our envisioned automated refactoring; Refactoring of Variable Renaming Based on BERT Pretraining: LIU et al. designed a two-stage pre-training framework, RefBERT, for variable name refactoring. It first predicts the number of sub-tokens in the new name and then generates sub-tokens accordingly, incorporating several techniques, including constrained masked language modeling, contrastive learning, and bag-of-tokens loss, to customize it for automatic variable name refactoring. Extensive experiments on the constructed refactoring dataset show that the variable names generated by RefBERT are more accurate and meaningful than those generated by existing methods. However, since RefBERT uses a two-stage method to first predict the number of sub-words and then generate sub-words, this step-by-step design may lead to a decrease in accuracy. Incorrect prediction of the number of sub-words directly affects the generation result, causing error accumulation in the sub-word generation stage and thus affecting the quality of the final variable name. Meanwhile, the model generation process of RefBERT ignores some specific naming conventions (such as prefixes and abbreviations used in specific projects), and this difference may make the generated variable names not conform to the naming style of the actual project, affecting the acceptability of renaming;Feature Dependence Detection of Graph Neural Networks: YU et al. represent the call relationships between methods in code by constructing a directed graph. They embed the code metrics of methods into the nodes in the graph and use a graph neural network model to extract features and classify methods. After determining that a method has feature envy, they calculate its call intensity with each external class and select the class with the highest call intensity as the recommended target class. The authors' method performs well in the three evaluation metrics of precision, recall, and F1-score, and its average performance is significantly better than the baseline method. However, this paper only considers the most basic indirect call relationships, lacks research on the deep connections between codes, and only collects seven code metrics, which is not enough to comprehensively describe the features of methods and limits the understanding ability of the model. In addition, the refactoring recommendation algorithm only judges the target class based on the call intensity, and this calculation method lacks analysis of deeper relationships (such as code semantics and functional relevance); Introduction of the combination of genetic algorithm and deep neural network: The purpose of the model designed by KARAKATI et al. is to evaluate and optimize code refactoring schemes through a deep neural network. It generates and evaluates refactoring schemes through a genetic algorithm and uses a deep neural network to predict and improve the quality of these schemes. The core idea of the model is to automatically extract rules through machine learning algorithms, thereby reducing the complexity and challenges of manually defining refactoring rules. This method is superior to other existing models in terms of code change score, automatic refactoring score, defect association ratio, refactoring accuracy ratio, defect detection ratio, and execution time. Although this method automates the refactoring process to a certain extent, developers still need to manually evaluate refactoring schemes in the initial stage. This method relies on the numerical representation of code snippets and uses technologies such as Code2Vec to generate code embeddings. If the representation of code snippets is not accurate or comprehensive enough, it may affect the quality of refactoring schemes.; Summary of the Invention
[0003] The object of the present invention is to provide an automated software refactoring method based on transfer learning, which can integrate code structure and semantic features, achieve end-to-end automated refactoring, solve the problem of insufficient refactoring data, and achieve systematic cross-task knowledge sharing.
[0004] The technical solution provided by the present invention is as follows:
[0005] An automated software refactoring method based on transfer learning, comprising:
[0006] Step 1: Train a BERT-based label prediction model using a code dataset; construct a label mapping matrix using the auxiliary task soft labels and refactoring task soft labels calculated by the BERT-based label prediction model;
[0007] Step 2: Process the code dataset according to the node type and edge type to construct a code refactoring structure diagram; use a graph neural network to extract code structure information, use Code2Vec to extract code semantic information, and fuse the code structure information and the code semantic information to generate a hybrid vector; input the hybrid vector into a fully connected layer, and use the activation function to generate a refactoring task label and an auxiliary task label; input the auxiliary task label into the label mapping matrix to generate a refactoring mapping label;
[0008] Step 3: Input the refactoring task label, the refactoring mapping label, and the hybrid vector into a Transformer model based on feature transfer to obtain the refactored code.
[0009] Preferably, the structure information of the code is extracted by constructing an abstract syntax tree and a control flow graph of the code and provided to the inductive framework GraphSAGE, and the GraphSAGE learns graph topology information and node embedding representations according to changes in node neighbor relationships.
[0010] Preferably, the code structure information generates a node embedding vector through the GraphSAGE, uses Code2Vec to extract code semantic information to generate an overall semantic vector of the code snippet, and fuses the node embedding vector and the overall semantic vector of the code snippet into a hybrid vector in a splicing manner in the fusion layer.
[0011] Preferably, the node embedding vector is:
[0012] ;
[0013] where is the node embedding vector of node ; is the activation function ReLu; is the attention coefficient; is the input feature matrix of node ; is the node embedding vector of the neighbor node set of node ; is the input feature matrix of the neighbor node set of node ; is the neighbor node set of node ; is the neighbor node of node ; represents connecting the input feature matrix of node with the node embedding vector of the neighbor node set of node ; is the mean function;
[0014] Preferably, the overall semantic vector of the code snippet is:
[0015]
[0016]
[0017] where is the overall semantic vector of the code snippet; is the weight of each path; is the path vector; is the total number of paths; is a learnable vector for generating attention scores; is the hyperbolic tangent activation function; is a learnable matrix for mapping the path vector to a latent space; is the source node embedding; is the embedding of the path itself; is the target node embedding; represents concatenating the embeddings of the source node, path, and target node.
[0018] Preferably, the label mapping matrix is obtained using the least squares method;
[0019] The objective function of the least squares method is:
[0020]
[0021] where is the label mapping matrix; is the soft label of the auxiliary task; is the soft label of the reconstruction task; is the Frobenius norm for measuring — error;
[0022] Minimize the objective function to obtain the label mapping matrix:
[0023]
[0024] Preferably, the Transformer model based on feature transfer includes: a Transformer-based reconstruction task code generation module, a Transformer-based auxiliary task code generation module, and an attention mechanism-based feature fusion and self-enhancement module.
[0025] Preferably, the attention mechanism-based feature fusion and self-enhancement module shares the attention mechanism among different tasks through transfer learning, and weights and fuses the features in the Transformer-based reconstruction task code generation module and the Transformer-based auxiliary task code generation module. The fused feature is:
[0026]
[0027] Where is the fused feature; is the query for the key attention coefficient; collectively refers to the query matrix, key matrix, and value matrix; is the th feature; is the number of Transformer models;
[0028] The fused feature is then enhanced using the self-attention mechanism to obtain the fused and self-enhanced feature as:
[0029]
[0030] Where is the fused and self-enhanced feature; , , are the learnable weights of the self-attention layer; is the concatenated feature, = ; is the original feature; is the reconstruction label; is the feature concatenation operation; is the dimension of the feature vector.
[0031] The beneficial effects of the present invention are:
[0032] The automated software refactoring method based on transfer learning provided by the present invention integrates technologies such as the BERT model, the graph neural network (GraphSAGE), Code2Vec, and the Transformer model to extract multi-dimensional information from the structure, semantics, and change patterns of code, achieving efficient and accurate software refactoring prediction; by constructing a label mapping matrix to assist knowledge transfer, the sampling strategy of GraphSAGE improves the efficiency of large-scale data processing, and at the same time the model exhibits good generalization ability to adapt to the requirements of different tasks and fields, demonstrating efficiency, accuracy, and scalability; by applying the Transformer model based on feature transfer, end-to-end automated refactoring is achieved, solving the problem of insufficient refactoring data and achieving systematic cross-task knowledge sharing. Description of the Drawings
[0033] Figure 1 It is a flowchart of the training process of the BERT-based label prediction model described in the present invention and the construction process of the label mapping matrix.
[0034] Figure 2 It is a flowchart of the automated software refactoring method based on transfer learning described in the present invention. Detailed Description of the Invention
[0035] The following further describes the present invention in detail with reference to the drawings, so that those skilled in the art can implement it according to the description in the specification.
[0036] As Figure 1 shown, a BERT-based label prediction model is trained using a code data set; an auxiliary task soft label and a refactoring task soft label calculated using the BERT-based label prediction model are used to construct a label mapping matrix.
[0037] The training process of the BERT-based label prediction model is as follows:
[0038] Step 1: Use the code data set to construct a BERT model input token sequence, where the code data set includes code segments before and after refactoring; input the token sequence of the code segment into the BERT model encoder;
[0039] Step 2: Perform forward propagation. Forward propagation refers to the process of inputting the token sequence of the code snippet into the BERT model and passing it from the input layer to the output layer of the BERT model. The input reconstructed and post-reconstructed code snippets are transformed into discrete sub-word units, and each sub-word is mapped to a unique index using BERT's vocabulary. At the same time, an attention mask is generated to indicate which tokens are valid and which are padding placeholders. To complete the label prediction for the reconstruction task, a task-specific classification head needs to be added on top of the BERT model. Therefore, a task-specific fully connected layer is added after the output layer of the BERT model. The number of neurons in the task-specific fully connected layer is equal to the number of categories of the task type. The context feature vectors of each input token output by the output layer of the BERT model are input into the task-specific fully connected layer. The scores of each category output by the task-specific fully connected layer pass through the activation function to generate the category probability distribution of each task type, thereby obtaining the soft labels predicted by each task code change model;
[0040] Step 3: Calculate the loss function. The difference between the predicted value and the actual value of the model is called the loss. The cross-entropy is used as the loss function, which is a commonly used loss function in classification problems. The loss function is used to measure the gap between the model's prediction and the actual label. The optimization goal is to minimize the loss function. Specifically, assume that the model predicts a set of reconstruction type probabilities for each data, and use this probability and the true misusage label to calculate the cross-entropy loss. This loss value represents the "distance" between the model's predicted value and the actual value. The smaller the loss value, the better the model's prediction effect;
[0041] Step 4: Perform backpropagation to calculate the gradient. Backpropagation is an optimization algorithm used in neural network models in machine learning. It is usually used to calculate the error gradient of the model and update the model parameters through the gradient descent method. Backpropagation relies on the chain rule to gradually calculate the gradient of the composite function. The gradient information is used by the Adam optimizer to update the parameters of the BERT model, with the goal of minimizing the value of the loss function;
[0042] Step 5: The Adam optimizer updates the parameters of the BERT model. The Adam optimizer can calculate the first-order momentum (mean) and second-order momentum (variance) of the gradient simultaneously. Its core idea is to dynamically adjust the learning rate of each parameter using historical gradient information. Utilization of gradient mean: The Adam optimizer tracks the historical change trend of the gradient of each parameter, calculates its mean, and uses it to smooth out short-term gradient fluctuations. This process is similar to the gradient descent algorithm with momentum, which can incorporate long-term gradient information into parameter updates to help the model optimize more stably. The Adam optimizer also tracks the weighted average of the squares of the gradients of each parameter to estimate the degree of gradient fluctuation. For parameters with large fluctuations, the algorithm will automatically reduce their learning rates, while for parameters with small fluctuations, it will appropriately increase the learning rates. This mechanism avoids the problem of unstable training caused by too large parameter update steps. In each iteration, the Adam optimizer updates the parameters of the BERT model once according to the calculated gradient information. In this training, only the parameters of the fully connected layer specific to the task are trained through backpropagation, while the parameters of the BERT encoder are frozen.
[0043] After the above process, a trained BERT-based label prediction model is obtained.
[0044] The construction process of the label mapping matrix is as follows:
[0045] In this embodiment, the auxiliary task code dataset containing the code snippets before and after reconstruction is used as the input to construct the BERT model input token sequence. Use as the start marker of the sequence, and separate the code before and after reconstruction with respectively. The token sequence in the input code snippet is represented as:
[0046]
[0047] where is each token in the code snippet before reconstruction; is each token in the code snippet after reconstruction; is the number of tokens;
[0048] By training the change patterns of the code before and after reconstruction, the code change rules similar to the reconstruction task are found in the auxiliary task, so as to establish a mapping relationship between the auxiliary task labels and the reconstruction task labels.
[0049] Train a BERT-based label prediction model specific to each auxiliary task, that is, the BERT-based auxiliary task label prediction model, to calculate the auxiliary task soft labels of the auxiliary task code dataset; since the labels generated by different tasks are in different label spaces, the auxiliary task soft labels are used to represent the probability distribution of the auxiliary task in the label space, and the label prediction of the auxiliary task itself is achieved through the BERT-based auxiliary task label prediction model:
[0050] Taking the calculation of the auxiliary task soft labels by the BERT-based auxiliary task label prediction model as an example, the calculation formula of the activation function of the BERT-based label prediction model is:
[0051] ;
[0052] where, is the class probability distribution of the th auxiliary task, that is, the auxiliary task soft label, which can capture the specific label features of each task; is the activation function; is the feature output by the output layer of the BERT model; is the weight of the fully connected layer of the BERT-based auxiliary task label prediction model specific to the th auxiliary task; is the bias of the fully connected layer of the BERT-based auxiliary task label prediction model specific to the th auxiliary task.
[0053] In this embodiment, the number of categories of the reconstruction task is set to , that is, the label space size, and the number of categories of the auxiliary task is set to . For the data of each auxiliary task, the auxiliary task soft labels of the data are calculated through the BERT-based auxiliary task label prediction model, and the auxiliary task soft labels represent the probabilities of the data in the current auxiliary task label space; the auxiliary task soft labels are expressed as:
[0054] ;
[0055] where, is the probability of the th class of the th auxiliary task;
[0056] Merge the auxiliary task soft label vectors of each dataset to obtain the auxiliary task soft label of the dataset. The dimension of the auxiliary task soft label is , where, is the number of data samples (i.e., there are data items in total);
[0057] Input the auxiliary task code data set into the BERT-based label prediction model specific to the reconstruction task, that is, the BERT-based reconstruction task label prediction model, to calculate the reconstruction task soft labels of the auxiliary task code data set; the reconstruction task soft labels represent the probabilities of the auxiliary task code data set in the reconstruction task label space, and represent the reconstruction task soft labels as:
[0058] ;
[0059] where is the probability of the auxiliary task code data set in the reconstruction task label space, that is, the reconstruction task soft label;
[0060] Merge the reconstruction task label vectors of each data set to obtain the reconstruction task soft label of this data set , and the dimension of the reconstruction task soft label is ;
[0061] Generate a label mapping matrix through the mapping relationship between the auxiliary task soft labels and the reconstruction task soft labels, and use the least squares method to find the label mapping matrix such that = ,
[0062] where the objective function of the least squares method is:
[0063]
[0064] where is the label mapping matrix; is the Frobenius norm, used to measure the — error;
[0065] Minimizing this objective function can obtain the label mapping matrix. Through formula derivation, the closed-form solution formula of the least squares method is obtained as:
[0066]
[0067] where the label mapping matrix is a matrix;
[0068] Using the least squares method, optimize and generate the label mapping matrix using all data at the same time. The label mapping matrix can capture the relationship between the probability distributions of the auxiliary task and the reconstruction task; for any new auxiliary task soft label vector , it can be used to predict the soft label probability vector of the corresponding refactoring task.
[0069] When implementing the mapping, the label mapping matrix pays more attention to the change rules before and after code refactoring than the features of the original code, and is more suitable for the code generation task of the Transformer model based on feature migration; in this embodiment, the label mapping matrix is used to assist the unified mapping of task labels to refactoring task labels, unify the label spaces of different auxiliary tasks into the label space of the refactoring task, realize the unification of the label space and cross-task transfer learning, and each value of the label mapping matrix represents the mapping relationship from the auxiliary task label to the refactoring task label, which can efficiently handle the label space differences between tasks and transfer the knowledge between tasks.
[0070] As Figure 2 shown, the present invention provides an automated software refactoring method based on transfer learning, and the specific implementation process is as follows:
[0071] Step 1: Train a BERT-based label prediction model using a code dataset; construct a label mapping matrix using the auxiliary task soft labels and refactoring task soft labels calculated by the BERT-based label prediction model.
[0072] Step 2: Perform data preprocessing on the input code dataset and construct a code refactoring structure diagram according to the node type and edge type; extract code structure information using a graph neural network, extract code semantic information using Code2Vec, fuse the code structure information and the code semantic information to generate a mixed vector; input the mixed vector into a fully connected layer, and use an activation function to generate refactoring task labels and auxiliary task labels; input the auxiliary task labels into the label mapping matrix to generate refactored mapping labels.
[0073] The specific process of the second step is:
[0074] Step 1: Perform data preprocessing on the input code dataset. Obtain node types and edge types through the Abstract Syntax Tree (AST) and Control Flow Graph (CFG), and construct a code refactoring structure graph based on the node types and edge types. The node types include computation nodes, control nodes, and parameter nodes, and the edge types include control dependency edges, data dependency edges, call dependency edges, and inheritance dependency edges. Based on various elements in the code dataset, according to the types of nodes and edges, extract node representations through a feature extractor, abstract different code elements into corresponding nodes, and connect each node with an edge according to the actual dependency relationships or association logics existing between elements, aggregating the information of different code elements and their relationships together, thereby constructing a graph structure that can reflect code structure information, enabling the embedding of each node to not only reflect the characteristics of the node itself but also contain its relationships with surrounding elements, thus generating a code refactoring structure graph with global context, providing a basic framework and key information support for subsequent code analysis and refactoring work.
[0075] Step 2: Extract code structure information by constructing the Abstract Syntax Tree (AST) and Control Flow Graph (CFG) of the code, and provide it to the inductive framework GraphSAGE. The GraphSAGE learns graph topology information and node embedding representations according to changes in node neighbor relationships. The GraphSAGE is a type of graph neural network. The code structure information generates node embedding vectors through the GraphSAGE. Select an arbitrary node in the code refactoring structure graph , first obtain the node 's node embedding vectors of the neighbor node set. The node embedding vectors of the neighbor node set are:
[0076] ;
[0077] Among them, is the node embedding vector of the neighbor node set of node ; is the input feature matrix of the neighbor node set of node ; is the neighbor node set of node ; is the neighbor node of node ; is the mean function;
[0078] Then obtain the node embedding vector of node as:
[0079] ;
[0080] Among them, is the node embedding vector of node ; is the activation function ReLu; is the attention coefficient; is the node input feature matrix; represents concatenating the input feature matrix of node with the node embedding vectors of the neighbor node set of node .
[0081] Step 3, Use Code2Vec to extract semantic features from the code, generate context-aware semantic vectors through the model, and capture the functions and logics of the code; Code2Vec extracts the paths from each source node to the target node in the code dataset; In this embodiment, set the path to represent a specific path from the source node to the target node , and the obtained path set is:
[0082]
[0083] wherein, is all paths from the source node to the target node; is the th specific path from the source node to the target node;
[0084] For each path , Code2Vec represents the path as a combination of the embedding of the source node , the embedding of the path itself, and the embedding of the target node , so as to obtain the vector of the path as:
[0085]
[0086] wherein, is the path vector; is the embedding of the source node ; is the embedding of the path itself; is the embedding of the target node ; represents concatenating the embeddings of the source node, the path, and the target node;
[0087] To generate the semantic vector representation of the entire code snippet, use the attention mechanism to weight each path, and the weight of each path is obtained through an attention scoring function. The calculation of the weight of each path is as follows:
[0088]
[0089] Among them, is the weight of each path; is a learnable matrix for mapping the path vector to a latent space; is a learnable vector for generating attention scores; is the hyperbolic tangent activation function;
[0090] By the weight of each path perform a weighted sum of each path vector to obtain the overall semantic vector of the code snippet as:
[0091]
[0092] Among them, is the overall semantic vector of the code snippet; is the total number of paths.
[0093] Step 4, fuse the node embedding vector and the overall semantic vector of the code snippet in the fusion layer through the vector mixing technique. In this embodiment, the node embedding vector and the overall semantic vector of the code snippet are combined by splicing to obtain a mixed vector; the mixed vector contains both the syntax, structure, and semantic information of the code, and can comprehensively consider the syntax, semantic, and structural features of the code in multi-task learning, optimize the task execution of code refactoring, and ensure prediction accuracy; pass the mixed vector output by the fusion layer to the fully connected layer as the input of the refactoring type classifier. The number of output nodes of the fully connected layer is equal to the number of possible refactoring types, such as method extraction, variable renaming, loop optimization, etc. Use activation function to convert the output of the fully connected layer into a probability distribution and generate a refactoring task label to predict the most likely refactoring type;
[0094] Among them, the activation function is:
[0095]
[0096] Among them, is the probability that the sample belongs to the category ; is the unactivated output of the sample in the category ; is the total number of categories;
[0097] Use the cross-entropy loss function to evaluate the performance of the model in predicting the refactoring type category. For soft labels, the cross-entropy loss calculates the difference between the predicted category probability and the actual soft label probability;
[0098] Suppose there are samples, each sample has categories, and the label probabilities of each category are given in the soft labels. The cross-entropy loss function is:
[0099]
[0100]
[0101] where is the soft label probability of sample in category ; is the predicted probability that the model predicts sample belongs to category ; is the cross-entropy loss function; is the total loss function.
[0102] In this embodiment, a separate fully-connected layer adapted to the specific task is prepared for each auxiliary task to output the auxiliary task label. The number of output nodes of the fully-connected layer is equal to the number of types of the task; the auxiliary task label is input into the label mapping matrix to be converted into a reconstructed mapping label, which is responsible for guiding the generation task of the Transformer model based on feature migration.
[0103] The reconstructed task label, the auxiliary task label, and the reconstructed mapping label are all soft labels.
[0104] Step 3: Input the reconstructed task label, the reconstructed mapping label, and the mixed vector into the Transformer model based on feature migration to obtain the reconstructed code. The Transformer model based on feature migration is provided with a Transformer-based reconstructed task code generation module, a Transformer-based auxiliary task code generation module, and an attention mechanism-based feature fusion and self-enhancement module.
[0105] The attention mechanism-based feature fusion and self-enhancement module shares the attention mechanism among different tasks through transfer learning, weights and fuses the features in the Transformer-based reconstructed task code generation module and the Transformer-based auxiliary task code generation module, and then enhances the fused features using the self-attention mechanism. The process of obtaining the fused and self-enhanced features is as follows:
[0106] Step 1: Generate a cross-model attention representation by calculating the query vector of the input feature and the key and value vectors of the external feature . The calculation formula is as follows:
[0107]
[0108] Among them, is the input feature; is the query matrix; is the key matrix; is the value matrix;
[0109] The query matrix, key matrix, and value matrix are collectively referred to as the matrix , and the attention calculation between multiple tasks or multiple features will use the same matrix , that is, this matrix is the same between different tasks, rather than assigning different matrices to each task or each feature, so that the model can share the cross-task attention mechanism and thus achieve cross-task information transfer;
[0110] For the Transformer-based reconstruction task code generation module, the input feature is the feature directly input into the Transformer-based reconstruction task code generation module, and the external feature is the feature from the Transformer-based auxiliary task code generation module; for the Transformer-based auxiliary task code generation module, the input feature is the feature directly input into the Transformer-based auxiliary task code generation module, and the external feature is the feature from the Transformer-based reconstruction task code generation module;
[0111] Calculate the similarity between the query and the key. The similarity between the query matrix and the key matrix represents the relationship strength between the query vector and the key vector. For the th query and the th key , the similarity (score) can be expressed as:
[0112]
[0113] Among them, is the th query; is the th key;
[0114] Perform the operation on the calculated similarity to obtain the normalized attention weights. These weights indicate the contribution size of each key to the query. After , they will be scaled to a probability distribution with a sum of 1:
[0115]
[0116] Among them, is the attention coefficient for querying the key ; is the dimension of the key vector ; is the total number of keys; is the th key;
[0117] The attention weight is determined by the attention coefficient , and the value matrix is weighted and summed using the attention weight. The output corresponding to each query is the weighted sum of all values :
[0118]
[0119] Among them, is the attention mechanism; is the total number of ;
[0120] The query will extract information from all values, but the amount extracted is determined by its similarity to each key (i.e., the attention weight).
[0121] Step 2, Use the calculated attention coefficients to perform weighted summation on all features to obtain the fused features. The formula for calculating the fused features is:
[0122]
[0123] Among them, is the fused feature, which contains information learned from all other task or model features, further enhancing the model's ability to understand and integrate multi-model features; is the th feature; is the number of Transformer models;
[0124] The goal of performing multi-model code information feature fusion operation is to integrate features of different tasks into a comprehensive representation, so that the model can learn richer context information.
[0125] Step 3, Further enhance the fused features through the self-attention mechanism to strengthen the mutual dependence between features; Calculate the concatenated features, and the concatenated features contain the original features, the fused features, the difference between the two, and the reconstruction label; The formula for the concatenated features is:
[0126]
[0127] Among them, is the concatenated feature; is the original feature; is the reconstruction label; is the feature concatenation operation;
[0128] Apply the self-attention mechanism to calculate the self-enhanced feature, and the calculation formula is:
[0129]
[0130] Among them, is the fused self-enhanced feature; , , are the learnable weights of the self-attention layer, is the dimension of the feature vector.
[0131] Through the feature fusion and self-enhancement module based on the attention mechanism, a fused self-enhanced feature representation is generated in each layer of the encoder , and the fused self-enhanced feature not only contains the information of multiple tasks, but also contains the features related to the target task, especially the task guidance provided by the reconstruction label.
[0132] Step 4: The internal structures of the Transformer-based reconstruction task code generation module and the Transformer-based auxiliary task code generation module are the same. Taking the Transformer-based reconstruction task code generation module as an example, the encoder part consists of multiple layers stacked. A standard Transformer model usually has 6 layers, that is, the encoder part has 6 layers. The features after fusion and self-enhancement are input into the next layer of the encoder part of the Transformer model. The Transformer model will gradually extract important information from these features after fusion and self-enhancement according to the self-attention mechanism and the multi-head attention mechanism, and perform in-depth feature representation learning. The Transformer model adopts an encoder-decoder structure. After being processed by the encoder, the obtained features will be passed to the decoder. The goal of the decoder is to convert these features into the reconstructed code. The specific process is as follows: In the decoder part, the features after fusion and self-enhancement will be further processed through the multi-head self-attention mechanism and the cross-attention mechanism. Through cross-attention, the decoder not only focuses on the input features (features after fusion and self-enhancement) themselves, but also generates corresponding code according to the context of the target task. In this way, the decoder can generate code that meets the expectations according to the specific requirements of the target task (such as the reconstruction type). This process is similar to the "decoding" process in natural language generation tasks. After multiple layers of processing, the decoder will output a probability distribution, indicating the generation probability of each possible code segment or character. The model will select the most likely code segment as the output according to this probability distribution. These code segments are the reconstructed code generated by the model according to the input features. After passing through the output layer of the decoder, the Transformer-based reconstruction task code generation module will generate the reconstructed code, which is the final prediction result of the model.
[0133] The Transformer-based auxiliary task code generation module will also generate the reconstructed code. However, the main purpose of the present invention is to obtain the reconstructed code generated by the Transformer-based reconstruction task code generation module. Therefore, the Transformer-based auxiliary task code generation module plays an auxiliary role and is not used as the final output result.
[0134] In order for the software refactoring model in the present invention to be able to use datasets in other fields to solve the problem of lack of datasets, transfer learning is introduced in the Transformer model part. The refactoring task label, the refactoring mapping label, and the hybrid vector are input into the Transformer model based on feature transfer to obtain the refactored code. Due to differences in data features, task objectives, or label spaces, the model for the target task may not be able to well adapt to the knowledge in the source field. Especially when there are large differences in the feature distributions between the target task and the source task, the model is prone to performance degradation or "catastrophic forgetting"; in addition, the model for the target task needs to directly adapt to the data and tasks in the source field, which may not be entirely suitable for the target task. Therefore, by transfer learning, the attention mechanism is shared among different tasks, and multi-task features are integrated into a comprehensive representation to enhance the model's understanding of multi-task features; cross-model feature fusion usually integrates multi-source information, but the feature dimensions and distributions of different models or tasks may be inconsistent, and direct fusion may lead to information "confusion" or lack of focus on the target task. Therefore, an adaptive mechanism is introduced to strengthen the association between features through self-attention, extract and emphasize features related to the target task, and weaken irrelevant noise. In this embodiment, the refactoring label feature is added, which is equivalent to providing a specific constraint for the fused features to ensure that the generated features are as relevant as possible to this refactoring type and reduce the ambiguity that may be introduced in multi-task fusion; the feature correlation between different tasks or models is calculated through the shared attention mechanism to generate attention coefficients, and weighted feature fusion is performed based on this. The fused features are further enhanced through the self-attention mechanism, thereby effectively improving the mutual dependence and fusion effect of the features. This structure is particularly suitable for solving the problem of inconsistent feature distributions between the source task and the target task in transfer learning and effectively improving the performance and stability of the model. The attention mechanism dynamically determines the dependence relationship between features during the feature fusion process of multiple tasks and models, enabling the model to more accurately capture the complex relationships between tasks and models during feature fusion and task transfer, thereby improving the effect of transfer learning.
[0135] The transfer learning in the present invention is achieved through feature transfer, specifically referring to directly using the features extracted from the source field in the target field model. These features can be feature vectors extracted through a pre-trained model or hand-designed features. In the source field, the model learns some high-quality features, which can be directly used in the target task without having to extract them from scratch. By combining or aggregating the features of the source field and the target field, the information of both can be utilized simultaneously.
[0136] The automated software refactoring method based on transfer learning provided by the present invention combines the BERT model, the Graph Neural Network (GraphSAGE), Code2Vec, and the Transformer model, enabling the extraction of multi-dimensional information from the structure, semantics, and change patterns of code. Through the global construction of a label mapping matrix, effective knowledge transfer between auxiliary tasks and refactoring tasks is achieved. The prediction of refactoring types is realized by fusing the structural information extracted by GraphSAGE and the semantic information extracted by Code2Vec, achieving more accurate refactoring type recognition. This fusion mode not only improves the accuracy of refactoring prediction but also enhances the model's ability to capture cross-task knowledge, making this method highly effective in the field of software refactoring. The present invention exhibits high computational efficiency when dealing with large-scale codebases. GraphSAGE can effectively process large-scale graph data. Compared with traditional Graph Neural Networks (GNNs) that usually need to traverse and calculate the entire graph, dealing with large-scale graphs is very memory and computationally resource-intensive, while GraphSAGE solves this problem by introducing a sampling strategy. Secondly, the feature fusion based on the attention mechanism and the design of the self-enhancement module enable the interaction between different features to be calculated in a short time. Through efficient feature mapping matrix calculation and parallel processing, the system can maintain good time complexity and computational efficiency on large-scale datasets. The present invention has excellent generalization ability and does not depend on specific programming languages or refactoring types. The label mapping matrix of its BERT-based label prediction model can easily adapt to different programming languages and refactoring tasks. By constructing a general code representation and label mapping matrix, the present invention can handle various types of refactoring tasks and adapt to the differences between the source domain and the target domain. In addition, by learning the change patterns of code rather than relying on manual features, the present invention can automatically adapt to different codebases and refactoring types, demonstrating strong generalization. In summary, the present invention not only has efficient and accurate technical effects in automated software refactoring tasks but also enables the model to adapt to the needs of different tasks and domains through flexible transfer learning and feature fusion strategies, having good scalability and adaptability.
[0137] The automated software refactoring method based on transfer learning provided by the present invention adopts dual feature extraction, combines structural information (extracting structural information such as control flow and data dependence of code through a graph neural network (GNN) like GraphSAGE) and semantic information (extracting path information from original fragments through Code2Vec), fuses the structural vector and the semantic vector in a splicing manner to generate a hybrid vector, and integrates the two kinds of information into a unified feature representation, so that the model can not only capture the function and logic of the code, but also take into account structural information such as control dependence and data dependence of the code. The hybrid vector is then provided to a Transformer model based on feature transfer for generating the refactored code, solving the problem of fusing code structure and semantic features.
[0138] The present invention realizes a complete process from code analysis to refactoring suggestions and then to refactored code generation. A refactoring type prediction module (such as extracting methods, renaming variables, etc.) is added to the model, and the hybrid vector is used for classifying refactoring tasks, so as to provide accurate refactoring type information for the Transformer model based on feature transfer, thereby generating appropriate code modification suggestions and realizing end-to-end automated refactoring; through the method of transfer learning, a label mapping matrix is constructed to transfer data and knowledge in other task domains (such as API recommendation, program repair) to the refactoring task, map the soft labels of other tasks to the label space of the refactoring task, and guide the model to learn richer refactoring patterns; during the training process of the label prediction model based on BERT, by generating soft labels (i.e., probability distributions) for each auxiliary task instead of traditional hard labels, the generalization ability of the model to different refactoring types is improved; the labels of the auxiliary tasks can provide valuable training signals for the main task, make up for the influence brought by insufficient data sets, and solve the problem of insufficient refactoring data; the constructed label mapping matrix unifies the label spaces of different tasks, ensures that the labels of different tasks can be shared in the same space, enables the models of different tasks to utilize each other's data and features during the training process, and improves the learning efficiency of the model. By introducing cross-task feature fusion based on the attention mechanism in the model, the features are weighted and fused using the similarity between different tasks; by calculating the similarity between tasks (such as the code change pattern between refactoring and repair), the model can automatically learn the information transfer and sharing between different tasks. Aiming at the possible noise problem during cross-task feature fusion, by introducing the self-attention mechanism, the model can dynamically adjust the weights of different features, strengthen the features related to the target task, weaken the irrelevant noise, and improve the task adaptability of the model, achieving systematic cross-task cognition and knowledge sharing and effective transfer and utilization of cross-task knowledge.
[0139] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the examples shown and described herein.
Claims
1. An automated software refactoring method based on transfer learning, characterized in that, Including: Step 1: Train a BERT-based label prediction model using a code dataset; Construct a label mapping matrix using the auxiliary task soft labels and the reconstruction task soft labels calculated by the BERT-based label prediction model; Step 2: Process the code dataset according to the node type and edge type to construct a code refactoring structure diagram; use a graph neural network to extract code structure information, use Code2Vec to extract code semantic information, and fuse the code structure information and the code semantic information to generate a hybrid vector; input the hybrid vector into a fully connected layer and use the activation function to generate a refactoring task label and an auxiliary task label; Input the auxiliary task labels into the label mapping matrix to generate reconstructed mapping labels; Step 3: Input the reconstruction task labels, the reconstructed mapping labels, and the mixed vector into a Transformer model based on feature transfer to obtain the reconstructed code.
2. The automated software refactoring method based on transfer learning according to claim 1, characterized in that Extract the structural information of the code by constructing the abstract syntax tree and control flow graph of the code, and provide it to the inductive framework GraphSAGE. The GraphSAGE learns the graph topology information and node embedding representation according to the changes in node neighbor relationships.
3. The automated software refactoring method based on transfer learning according to claim 2, characterized in that The code structure information generates node embedding vectors through the GraphSAGE, and Code2Vec is used to extract the semantic information of the code to generate the overall semantic vector of the code snippet. The node embedding vector and the overall semantic vector of the code snippet are fused into a mixed vector in the fusion layer by concatenation.
4. The automated software refactoring method based on transfer learning according to claim 3, wherein The node embedding vector is: ; Among them, is the node embedding vector; is the activation function ReLu; is the attention coefficient; is the node input feature matrix; is the node node embedding vector of the neighbor node set; is the node input feature matrix of the neighbor node set; is the node neighbor node set; is the node neighbor node; represents concatenating the input feature matrix of the node with the node embedding vectors of the neighbor node set of the node ; is the mean function.
5. The automated software refactoring method based on transfer learning according to claim 4, characterized in that The overall semantic vector of the code snippet is: ; ; Among them, is the overall semantic vector of the code snippet; is the weight of each path; is the path vector; is the total number of paths; is a learnable vector used to generate attention scores; is the hyperbolic tangent activation function; is a learnable matrix used to map the path vector to a latent space; is the source node 's embedding; is the embedding of the path itself; is the target node 's embedding; represents concatenating the embeddings of the source node, path, and target node.
6. The automated software refactoring method based on transfer learning according to claim 1, characterized in that, Obtain the label mapping matrix using the least squares method; The objective function of the least squares method is: ; Among them, is the label mapping matrix; is the soft label of the auxiliary task; is the soft label of the reconstruction task; is the Frobenius norm, used to measure — the error of Minimize the objective function to obtain the label mapping matrix: 。 7. The automated software refactoring method based on transfer learning according to claim 1, wherein The Transformer model based on feature transfer includes: a Transformer-based reconstruction task code generation module, a Transformer-based auxiliary task code generation module, and an attention mechanism-based feature fusion and self-enhancement module.
8. The automated software refactoring method based on transfer learning according to claim 7, characterized in that The attention mechanism-based feature fusion and self-enhancement module shares the attention mechanism between different tasks through transfer learning, and weights and fuses the features in the Transformer-based reconstruction task code generation module and the Transformer-based auxiliary task code generation module. The fused features are: ; Among them, is the fused feature; is the query for the key attention coefficient; is a general term for the query matrix, key matrix, and value matrix; is the th feature; is the number of Transformer models; The fused features are further enhanced using the self-attention mechanism to obtain the fused and self-enhanced features as: ; Among them, is the feature after fusion and self-enhancement; , , are the learnable weights of the self-attention layer; is the feature after connection, = ; is the original feature; is the reconstruction label; is the feature connection operation; is the dimension of the feature vector.
Citation Information
Patent Citations
Two-stage pre-training framework for automatic renaming reconstruction
CN117473307A
Feature attachment detection method based on multi-level code representation
CN119025418A