Code translation method based on tree instruction large language model
Through the two-stage instruction fine-tuning technology and similarity model evaluation, the generality and readability problems in code translation are solved, efficient and accurate code translation and function retention are achieved, and the translation cost is reduced.
Patent Information
- Application Number
- CN202510078021.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology lacks universality in code translation, the generated code has dead code and readability problems, and the translation cost is high. Traditional rules have good performance based on methods but are not universal. The code generated by natural language processing methods cannot be run through compilation.
Two-stage instruction fine-tuning technology is adopted to combine syntax information in abstract syntax tree (AST) with code snippets, and use large language models to better understand the syntax and semantics of the code in fine-tuning, and evaluate the readability and functional consistency of the translated code through a similarity model.
It realizes efficient and accurate code translation, reduces translation costs, generates code readability and complete functions, and the method is highly scalable, and is suitable for different large language models.
Smart Images

Figure CN120010854A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of software engineering, and in particular relates to a code translation method based on a tree instruction large language model, which can be used to convert a source programming language into a target programming language. Background Art
[0002] Code translation, which is converting a piece of code from one programming language to another while retaining the original functionality, plays a vital role in software maintenance programs. With the development of the digital field, code translation has become an emerging and rapidly developing task in software engineering. For example, migrating old programming languages (such as COBOL) to more reliable and stable popular programming languages (such as Java) platforms; or making software projects run on multiple platforms, all require the help of code translation technology. However, manually migrating code to another programming language is tedious and time-consuming. Therefore, how to translate the code efficiently and accurately while reducing the translation cost is a key issue.
[0003] The traditional approach to solving this problem is the rule-based translation method, the core idea of which is to develop common templates for the source programming language, such as loop templates, selection templates, exception handling templates, etc., and then parse the grammatical rules of the source language into an abstract syntax tree (AST) or a control flow graph, extract key information from it and embed it into the template. This solution has good performance and provides a more complete supplementary module for the target code while retaining basic functions. For example, OpenCOBOL4J will create corresponding classes and objects for each variable, and will also generate a more complete exception handling statement for each statement. However, this solution can only handle specific situations and is not universal. For example, a large amount of dead code is generated, which affects readability. In addition, the time required to develop a template for a project is usually in years, which greatly affects the progress of the project.
[0004] Another approach to solving this problem is to translate the code as natural language, and complete the code translation task by inputting large amounts of parallel data sets into machine learning models. This solution is more flexible than traditional solutions, the generated code is more readable, and the translation cost is greatly reduced. However, unlike natural languages, programming languages follow strict format and grammar requirements, and ambiguity in the translation results is not allowed. Translation models with small parameters cannot learn the semantics and grammar of the source language from parallel data sets, and the translated results usually cannot be compiled and run. Although there are some invention patents for code translation, these existing invention patents still have problems such as the inability to align the source code with the target code and low translation success rate. For example, the training method, device, equipment and storage medium of the programming language translation model (patent number: CN202110021389.8) encodes each word in the code, extracts feature vectors through a two-layer encoder, and inputs the extracted feature vectors into the decoder to predict the code translation result corresponding to the first answer code; an intelligent source code conversion method based on supervised learning (patent number: CN202110021389.8) captures a large amount of parallel data from the code website, constructs a mapping between programming languages, converts the language code into a machine-recognizable code, and inputs the machine-recognized code into the model for training to obtain a code translation model. However, the above two schemes are limited by the number of model parameters and the particularity of the programming language, and cannot accurately and efficiently complete the code translation task.
[0005] Therefore, in order to complete the code translation task efficiently and accurately, the present invention utilizes two-stage instruction fine-tuning to combine the grammatical information in AST with the code snippets in a more fine-grained manner, so that the large language model can better understand the syntax (generate correct code that can be compiled and run) and semantics (retain functions) of the code during fine-tuning; at the same time, the present invention establishes a new indicator to evaluate the readability of the code translated by the large model, thereby better improving the user-friendliness of the code translation model. Summary of the invention
[0006] The technical problem to be solved by the present invention is to complete the code translation task based on the tree instruction large language model. In order to enable the large language model to extract information from AST more accurately, a scheme that can accurately express the key information in AST is first designed as the first step to align the tree structure and code information; secondly, in the first stage of fine-tuning, the AST is decomposed into tree tags, and a similarity model is trained to make the source tree tags and the code snippets of the target language correspond one to one, and at the same time, the instruction paradigm of the code and tree structure is introduced to construct a self-supervised instruction data set, which is delivered to the large language model for fine-tuning. The large language model further enhances the understanding of the tree structure by aligning the tree tags with the code snippets; then, in the second stage of instruction fine-tuning, the downstream task code translation is customized, and the large language model generates a program according to the grammar of the target language by understanding the tree structure, thereby completing the translation task; finally, a new evaluation indicator, functional alignment, is added, and the similarity score of the structure and identifier naming of the AST generated by the translated code is obtained through the similarity model obtained through training, so as to evaluate the degree to which the code translated by the large model retains the function and identifier naming of the source code.
[0007] The technical solution of the present invention:
[0008] A code translation method based on a tree instruction large model, the specific steps are as follows:
[0009] Step (1): Parallel code and token alignment dataset construction.
[0010] The parallel codes of the corresponding functions are obtained from CodeGeeX and AVATAR datasets, and a parallel dataset Dp containing three key information: task number (task_id), source language code (source_code) and target language code (target_code) is constructed; the token alignment dataset Dt containing three key information: token number (token_id), source language tag (source_token) and target language tag (target_token) is obtained from CodeGeeX;
[0011] Step (2): Abstract syntax tree preprocessing.
[0012] For each source_code and target_code data in the parallel dataset Dp and the token alignment dataset Dt, the parser obtains its abstract syntax tree (AST), traverses the tree structure to convert it into a linear sequence, and filters the key information of the AST tree that retains the conversion result; the source tree sequence AST of the linearized AST is obtained from Dp source and the target tree sequence AST target , get the linearized source tree token DtToken from Dtsource and the target tree token DtToken target ;
[0013] Step (3): Similarity model training.
[0014] Mark the source tree obtained in step (2) with DtToken source and the target tree token DtToken target The corresponding and Using the corresponding aligned tree encoding and Train a similarity model Γ;
[0015] Step (4): First stage instruction fine-tuning prompt template construction.
[0016] Design a prompt template to be filled in 1: "Given a tree tag: <tree tag>, here is a <code>, please get the relevant <target language> code snippet information based on the tree tag: <code snippet>";
[0017] Step (5): First stage instruction fine-tuning.
[0018] The source tree sequence AST obtained in step (2) source and the target tree sequence AST target Further decomposed into tree tokens composed of internal nodes, and the source tree tokens DpToken are obtained respectively source and the target tree token DpToken target ; Obtain the corresponding structured coding scheme in step (3) and The obtained model is used to obtain a similarity score X, and the source tree tags and the target tree tags with the largest similarity score are matched one by one; the target tree tags are parsed into code snippets to obtain a sequence of source tree tags aligned with the target code snippets; the obtained source tree tag sequence and the target code snippet are used to fill the prompt template 1 provided in step (4), and it is added as the first stage self-supervised instruction dataset to the large language model for fine-tuning;
[0019] Step (6): The second stage instruction fine-tuning prompt template is constructed.
[0020] Adjust the prompt template 1 in step (4) to prompt template 2: "Given a tree tag: <tree tag>, please generate relevant <target language> code snippet information based on the tree tag: <target code snippet>";
[0021] Step (7): First stage instruction fine-tuning.
[0022] The AST in step (2)Source The target code target_code field corresponding to the parallel data set Dp in step (1) is filled into the prompt template 2 in step (6), and it is added as the second-stage self-supervised instruction data set to the fine-tuning large model obtained in step (5) for fine-tuning;
[0023] Step (8): Model inference.
[0024] Convert the test set code of CodeGeeX into a linearized tree structure and add it to prompt template 2. The "<code snippet>" is no longer displayed and is input into the fine-tuned large model for reasoning.
[0025] Step (9) Evaluation of translation result functional consistency and identifier consistency.
[0026] The code generated by the large language model is converted into a linearized tree structure, and the similarity score between the linearized tree structure of the target code and the linearized tree structure of the generated code is obtained through the similarity model obtained in step (4) to evaluate the translation result.
[0027] Furthermore, step (2) specifically includes the following steps:
[0028] 2-1) Traverse the AST tree:
[0029] a) Get the AST tree through the AST tree parser;
[0030] b) Input the root node of the AST tree;
[0031] c) If the root node is a leaf node, its name is directly generated: <node attribute, left> node name <node attribute, right>, and only the key information is retained: node attribute, node name and value;
[0032] d) Recursively repeat steps b and c with other nodes as root nodes.
[0033] Furthermore, step (3) specifically includes the following steps:
[0034] 3-1) Mark the source tree in step (2) as DtToken source and the target tree token DtToken target Through the Transformer encoder f r Generate source tree encoding S and target tree encoding T. Here, the vanilla encoder vanillatransformer is selected, as shown in formula (1);
[0035] S=f r (DtToken source ),T=f r (DtTokentarget ) (1)
[0036] 3-2) The source tree encoding and the target tree encoding are obtained by norm (i.e., row-by-row L2 normalization) and As shown in formula (2);
[0037]
[0038] 3-3) Design a similarity model Γ, which is calculated as shown in formula (3), where τ∈R is a trainable temperature parameter used to scale the similarity value; and is the transformation function of different dimensions, expressed as three transformation formulas (4), (5) and (6), N is and The first dimension of the two multidimensional matrices is the number of nodes, i = (1, 2, ..., N), N i Represented as the i-th node, j∈N i ; The contrast loss function L of the similarity model Γ is formula (7), where CE is the cross entropy loss function, is the label vector used for comparative training.
[0039]
[0040] Furthermore, step (5) specifically comprises the following steps:
[0041] 5-1) Traverse the source tree sequence AST in step (2) source and the target tree sequence AST target If the node has child nodes, the node and its child nodes are put into a list as tree tokens, which are source tree tokens DpToken source and the target tree token DpToken target ; Obtained according to the structured coding scheme of step (3) and At the same time, the text of non-leaf nodes is used as the corresponding code snippet information;
[0042] 5-2) Calculate the similarity model obtained in step (3) and The similarity score X of Corresponding to the maximum similarity score
[0043] 5-3) Mark the target tree Object code and Corresponding The code snippet information is filled into the <tree tag>, <code> and <code snippet> in the prompt template 1 designed in step (4), and a self-supervised instruction fine-tuning dataset is constructed;
[0044] 5-4) Add the instruction fine-tuning dataset to the large language model and use the LoRA fine-tuning scheme to perform the first stage of self-supervised instruction fine-tuning.
[0045] Furthermore, step (9) specifically includes the following steps:
[0046] On the premise that the generated results can be compiled and run, the source code and the translated code are converted into a linearized AST structure, and the normalized tree encoding is obtained according to formulas (1) and (2): and The similarity score X is calculated using the model trained in step (3), and the result is used as the basis for determining the functional consistency and identifier naming consistency after translation.
[0047] Compared with the prior art, the present invention has the following advantages and effects:
[0048] The method of the present invention can automatically and efficiently translate source code into target language code. The present invention extracts key tree structure information from the linearized AST and trains a similarity model to align the tree structure with the code structure, thereby constructing a more accurate instruction data set and increasing the ability of the large language model to understand the AST structure; by combining with the large language model, human intervention is greatly reduced, and the cost of code translation is greatly reduced; by adding additional functional evaluation schemes, the friendliness of the translation results to users is improved; in addition, the method of the present invention is also highly scalable and can easily switch to use different large language models, which is conducive to improving user experience and reducing the threshold of professional skills required for use. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of the code translation method based on the tree instruction tuning and tuning large model of the present invention.
[0050] Figure 2 It is a linearization preprocessing flow subgraph in the abstract syntax tree preprocessing in the code translation method based on the tree instruction tuning and tuning large model of the present invention.
[0051] Figure 3 It is a structured coding process subgraph of the similarity model training phase in the code translation method based on the tree instruction tuning large model of the present invention.
[0052] Figure 4 It is a tree structure alignment process subgraph of the first stage instruction fine-tuning phase in the code translation method based on the tree instruction tuning large model of the present invention.
[0053] Figure 5 It is a flow subgraph of function and identifier consistency evaluation in the reasoning and verification phases of the code translation method based on the tree instruction tuning large model of the present invention. DETAILED DESCRIPTION
[0054] The method of the present invention is described in detail below in conjunction with the accompanying drawings, technical solutions and embodiments.
[0055] like Figure 1 As shown, the code translation method based on tree instruction tuning and tuning large model of the present invention is carried out according to the following process: first, the collected parallel data set is converted into an abstract syntax tree, and the abstract syntax tree is linearized and preprocessed to screen key information, so as to obtain an aligned data set of trees and tree tags; then, the tree tag data set is structured encoded, and a similarity model is trained using the encoding result to expand the tree tag alignment data set; next, the tree alignment data set is split into tree tag alignment data sets, and the source tree tags and target code fragments are aligned through the similarity model to fill the instruction prompt template 1, and the obtained tree instruction data set is input into the large language model to complete the first stage of self-supervised instruction fine-tuning; then the source tree-target code data set is aligned to fill the instruction prompt template 2, and the obtained tree instruction data set is input into the large language model to complete the second stage of self-supervised instruction fine-tuning; finally, the test set is converted into a linearized abstract syntax tree and added to the prompt template 2, and input into the fine-tuned large language model for reasoning. If the translated code can be run through compilation, it and the target code are converted into a linearized abstract syntax tree, and input into the similarity model for function and identifier consistency evaluation. The following takes the first Python-Java code alignment dataset and the corresponding token alignment dataset in CodeGeeX as an example to explain the implementation details of each process in detail.
[0056] The specific implementation is as follows:
[0057] (1) Obtain the parallel codes of the corresponding functions from CodeGeeX and AVATAR datasets, and construct a parallel dataset Dp containing three key information: task_id, source_code, and target_code; obtain the token alignment dataset Dt of the corresponding functions from CodeGeeX and containing three key information: token_id, source_token, and target_token;
[0058] (2) For each source_code and target_code data in the parallel dataset Dp and the token alignment dataset Dt, its abstract syntax tree (AST) is obtained through the parser, such as Figure 2As shown, the tree structure is traversed to convert it into a linear sequence, and the key information of the AST tree that retains the conversion result is filtered out; the source tree sequence AST of the linearized AST is obtained by Dp source and the target tree sequence AST target , get the linearized source tree token DtToken from Dt source and the target tree token DtToken target .
[0059] 2.1. Traversing the AST tree:
[0060] a) Get the AST tree through the AST tree parser;
[0061] b) Input the root node of the AST tree;
[0062] c) If the root node is a leaf node, its name is directly generated: <node attribute, left> node name <node attribute, right>, and only the key information is retained: node attribute, node name and value;
[0063] d) Recursively repeat steps b and c with other nodes as root nodes
[0064] Specifically:
[0065] In this implementation example, the flattened Python code in the parallel dataset is first organized into a compilable and executable Python code, and the Python and Java codes are converted into a linearized abstract syntax tree to retain key information.
[0066] Next, for step 2.1), use Python's tree-sitter tool to parse and get the AST root node object of the input code, and then Figure 2 The linearized abstract syntax tree preprocessing scheme mentioned above can obtain the source tree sequence AST of the linearized AST from Dp source and the target tree sequence AST target , get the linearized source tree token DtToken from Dt source and the target tree token DtToken target .
[0067] Linearized AST source tree sequence AST source and the target tree sequence AST target As shown in Table 1:
[0068] Table 1 Linearized abstract syntax tree representation
[0069]
[0070] The token alignment code is shown in Table 2:
[0071] Table 2 Token alignment code dataset example
[0072]
[0073] Convert to source tree token source and target tree token target As shown in Table 3:
[0074] Table 3 Source tree token DpToken source and target tree token target Conversion Examples
[0075]
[0076] (3) Mark the source tree obtained in step (2) DtToken sOurce and the target tree token DtToken target The corresponding and Using the corresponding aligned tree encoding and Train a similarity model Γ.
[0077] 3.1. Mark the source tree in step (2) as DtToken source and the target tree token DtToken target Through the Transformer encoder f r Generate source tree encoding S and target tree encoding T. Here, the vanilla encoder vanillatransformer is selected, as shown in formula (1);
[0078] S=f r (DtToken source ),T=f r (DtToken target ) (1)
[0079] 3.2. The source tree encoding and the target tree encoding are obtained by norm (i.e., row-by-row L2 normalization) and As shown in formula (2);
[0080]
[0081] 3.3. Design a similarity model Γ, which is calculated as shown in formula (3), where τ∈R is a trainable temperature parameter used to scale the similarity value. and is the transformation function of different dimensions, expressed as three transformation formulas (4), (5) and (6), N is and The first dimension, i.e. the number of nodes, i = (1, 2, ..., N), N i Represented as the i-th node, j∈N i ; The contrast loss function L of the similarity model Γ is formula (7), where CE is the cross entropy loss function, is the label vector used for comparative training.
[0082]
[0083] Specifically:
[0084] First, Figure 3 As shown, the source tree token obtained in the previous step is Token source and target tree token target The Transformer encoder is used for structured encoding, and then input into the similarity model constructed according to formula (3)-formula (7) for training. The structured encoding representation of the Token alignment dataset 1 is shown in Table 4:
[0085] Table 4 Structured encoding of Python and Java tokens
[0086] Python Token 1 [[0.1702467118063,0.00081718421667,…],…[…,0.15522731632]] Python Token 2 [[0.0276056586868,0.000132507161,…],…[…,0.309425932]] … … Java Token 1 [[2.71173204e-02,1.30163138e-04,…],…[…,3.03964676e-01]] Java Token 2 [[0.11361865683,0.00054536955,…],[…,0.10408252514598822]] … …
[0087] (4) Design a prompt template 1 to be filled: "Given a tree tag: <tree tag>, here is a piece of <code>, please obtain the relevant <target language> code snippet information based on the tree tag: <code snippet>".
[0088] (5) The source tree sequence AST obtained in step (2) source and the target tree sequence AST target Further decomposed into tree tokens composed of internal nodes, and the source tree tokens DpToken are obtained respectively source and the target tree token DpToken target ; Obtain the corresponding structured coding scheme in step (3) and The obtained model is used to obtain a similarity score X, and the source tree tags and the target tree tags with the largest similarity score are matched one by one; the target tree tags are parsed into code snippets to obtain a sequence of source tree tags aligned with the target code snippets; the obtained source tree tag sequence and the target code snippet are used to fill the prompt template 1 provided in step (4), and it is added to the large language model as the first stage self-supervised instruction dataset for fine-tuning.
[0089] 5.1. Traverse the source tree sequence AST in step (2) source and the target tree sequence AST target If the node has child nodes, the node and its child nodes are put into a list as tree tokens, which are source tree tokens DpToken source and the target tree token DpToken target ; Obtained according to the structured coding scheme of step (3) and At the same time, the text of non-leaf nodes is used as the corresponding code snippet information;
[0090] 5.2. Calculate the similarity model obtained in step (3) and The similarity score X of Corresponding to the maximum similarity score
[0091] 5.3. Mark the target tree Object code and Corresponding The code snippet information is filled into the <tree tag>, <code> and <code snippet> in the prompt template 1 designed in step (4), and a self-supervised instruction fine-tuning dataset is constructed;
[0092] 5.4. Add the instruction fine-tuning dataset to the large language model and use the LoRA fine-tuning scheme to perform the first stage of self-supervised instruction fine-tuning.
[0093] Specifically:
[0094] Python and Java are preprocessed through abstract syntax trees to obtain abstract syntax trees and Figure 4 Split it into tree tags; then perform structured encoding on Python and Java tree tags and input them into the similarity model for maximum similarity matching, thereby obtaining an aligned token matching sequence; finally, fill the Python tree tags, Java code, and Java code snippets corresponding to the matching Java tree tags into the prompt template to construct the first-stage self-supervised instruction dataset, and add it to the large language model for fine-tuning. The specific generated instruction prompts are as follows:
[0095] "Given a tree notation: ['AST#expression_statement#Left', 'AST#assignment_expression#Left', 'X', '=', ..., 'AST#expression_statement#Right'], here is a piece of code:
[0096] import java.util.*; class GFG{static int maxPresum(int[]a,int[]b){intX=Math.max(a[0],0); for(int i=1;i <a.length;i++){a[i]+=a[i-1];X=Math.max(X,a[i]);}int Y=Math.max(b[0],0);for(int i=1;i<b.length;i++){b[i]+=b[i-1];Y=Math.max(Y,b[i]);}return X+Y;}public static void main(String[]args){int[]A={2,-1,4,-5};int[]B={4,-3,12,4,-3};System.out.print(maxPresum(A,B));}}
[0097] Please follow the tree tags to get the relevant Java code snippet information:
[0098] int X=Math.max(a[0],0);”
[0099] (6) Adjust the prompt template 1 of step (4) to prompt template 2: "Given a tree tag: <tree tag>, please generate relevant <target language> code snippet information based on the tree tag: <target code snippet>".
[0100] (7) The AST in step (2) source The target code target_code field corresponding to the parallel data set Dp in step (1) is filled into the prompt template 2 in step (6), and it is added as the second-stage self-supervised instruction data set to the fine-tuning large model obtained in step (5) for fine-tuning.
[0101] The specific generated command prompts are as follows:
[0102] "Given a tree tag: ['AST#module#Left','AST#function_definition#Left','def','max','Pres','um','AST#parameters#Left','(','a',',','b',')','AST#parameters#Right',…,'AST#argument_list#Right','AST#call#Right',')','AST#argument_list#Right','AST#call#Right','AST#module#Right'], please generate relevant Java code snippet information based on the tree tag:
[0103] import java.util.*; class GFG{static int maxPresum(int[]a,int[]b){intX=Math.max(a[0],0); for(int i=1;i <a.length;i++){a[i]+=a[i-1];X=Math.max(X,a[i]);}int Y=Math.max(b[0],0);for(int i=1;i<b.length;i++){b[i]+=b[i-1];Y=Math.max(Y,b[i]);}return X+Y;}public static void main(String[]args){int[]A={2,-1,4,-5};int[]B={4,-3,12,4,-3};System.out.print(maxPresum(A,B));}}”
[0104] (8) Convert the CodeGeeX test set code into a linearized tree structure and add it to prompt template 2. The "<code snippet>" is no longer displayed and is input into the fine-tuned large model for reasoning.
[0105] The specific input is:
[0106] "Given a tree tag: ['AST#program#Left','AST#local_variable_declaration#Left',…,'False','AST#ERROR#Right','AST#program#Right'], please generate the relevant Java code snippet information based on the tree tag:".
[0107] (9) The code generated by the large language model is converted into a linearized tree structure, and the similarity score between the linearized tree structure of the target code and the linearized tree structure of the generated code is obtained through the similarity model obtained in step (4) to evaluate the translation result.
[0108] 9.1. On the premise that the generated results can be compiled and run, convert the source code and the translated code into a linearized AST structure, and obtain the normalized tree encoding according to formulas (1) and (2): and The similarity score X is calculated using the model trained in step (3), and the result is used as the basis for determining the functional consistency and identifier naming consistency after translation.
[0109] Specifically:
[0110] Get the output result of fine-tuning the large model:
[0111] import java.util.*; class GFG{static boolean has_close_elements(float[]numbers,float threshold){for(int idx=0;idx <numbers.length;idx++){for(intidx2=0;idx2<numbers.length;idx2++){if(idx!=idx2){float distance=Math.abs(numbers[idx]-numbers[idx2]);if(distance<threshold)return true;}}}returnfalse;}public static void main(String[]args){float[]numbers={1.0f,2.0f,3.0f,4.0f,5.0f};float threshold=2.0f;if(has_close_elements(numbers,threshold))System.out.println("Yes");else System.out.println("No");}}
[0112] The result can be compiled and run;
[0113] The code generated by the large language model is converted into a linearized AST tree. The similarity score between the AST tree of the target code and the AST tree of the generated code is obtained through the similarity model in step (3) to evaluate the translation result. The similarity score is 0.8714, indicating that the functions and identifier naming of the source code are retained to a great extent.
Claims
1. A code translation method based on a tree instruction large model, characterized in that: The specific steps are as follows: Step (1): Parallel code and token alignment dataset construction; Obtain the parallel codes of the corresponding functions from CodeGeeX and AVATAR datasets, and construct a parallel dataset Dp containing three key information: task number task_id, source language code source_code, and target language code target_code; obtain the token alignment dataset Dt of the corresponding functions from CodeGeeX and containing three key information: token number token_id, source language tag source_token, and target language tag target_token; Step (2): abstract syntax tree preprocessing; For each source_code and target_code data in the parallel dataset Dp and the token alignment dataset Dt, the parser obtains its abstract syntax tree AST, traverses the tree structure to convert it into a linear sequence, and filters the key information of the AST tree that retains the conversion result; the source tree sequence AST of the linearized AST is obtained from Dp source and the target tree sequence AST target , get the linearized source tree token DtToken from Dt source and the target tree token DtToken target ; Step (3): similarity model training; Mark the source tree obtained in step (2) with DtToken source and the target tree token DtToken target The corresponding and Using the corresponding aligned tree encoding and Train a similarity model Γ; Step (4): First stage instruction fine-tuning prompt template construction; Design a prompt template to be filled in 1: "Given a tree tag: <tree tag>, here is a <code>, please get the relevant <target language> code snippet information based on the tree tag: <code snippet>"; Step (5): fine-tuning of the first-stage instructions; The source tree sequence AST obtained in step (2) source and the target tree sequence AST target Further decomposed into tree tokens composed of internal nodes, and the source tree tokens DpToken are obtained respectively source and the target tree token DpToken target ; Obtain the corresponding structured coding scheme in step (3) and The obtained model is used to obtain a similarity score X, and the source tree tags are matched one by one with the target tree tags with the largest similarity score; the target tree tags are parsed into code fragments, so as to obtain a sequence of source tree tags aligned with the target code fragments; Use the obtained source tree tag sequence and the target code snippet to fill in the prompt template 1 provided in step (4), and add it as the first stage self-supervised instruction dataset to the large language model for fine-tuning; Step (6): Second stage instruction fine-tuning prompt template construction; Adjust the prompt template 1 in step (4) to prompt template 2: "Given a tree tag: <tree tag>, please generate relevant <target language> code snippet information based on the tree tag: <target code snippet>"; Step (7): fine-tuning of the first stage instructions; The AST in step (2) Source The target code target_code field corresponding to the parallel data set Dp in step (1) is filled into the prompt template 2 in step (6), and it is added as the second-stage self-supervised instruction data set to the fine-tuning large model obtained in step (5) for fine-tuning; Step (8): Model reasoning; Convert the CodeGeeX test set code into a linearized tree structure and add it to prompt template 2. "<code snippet>" is no longer displayed and is input into the fine-tuned large model for reasoning. Step (9) evaluating the functional consistency and identifier consistency of the translation results; The code generated by the large language model is converted into a linearized tree structure, and the similarity score between the linearized tree structure of the target code and the linearized tree structure of the generated code is obtained through the similarity model obtained in step (4) to evaluate the translation result.
2. A code translation method based on a tree instruction large model according to claim 1, characterized in that: Step (2) specifically includes the following steps: 2-1) Traverse the AST tree: a) Get the AST tree through the AST tree parser; b) Input the root node of the AST tree; c) If the root node is a leaf node, its name is directly generated: <node attribute, left> node name <node attribute, right>, and only the key information is retained: node attribute, node name and value; d) Recursively repeat steps b and c with other nodes as root nodes.
3. A code translation method based on a tree instruction large model according to claim 1, characterized in that: Step (3) specifically includes the following steps: 3-1) Mark the source tree in step (2) as DtToken source and the target tree token DtToken target Through the Transformer encoder f r Generate source tree encoding S and target tree encoding T. Here, the vanilla encoder vanillatransformer is selected, as shown in formula (1); S=f r (DtToken source ),T=f r (DtToken target ) (1) 3-2) The source tree encoding and the target tree encoding are obtained by norm (i.e., row-by-row L2 normalization) and As shown in formula (2); 3-3) Design a similarity model Γ, which is calculated as shown in formula (3), where τ∈R is a trainable temperature parameter used to scale the similarity value; and is the transformation function of different dimensions, expressed as formula (4), (5) and (6), N is and The first dimension, i.e. the number of nodes, i = (1, 2, ..., N), N i Represented as the i-th node, j∈N i ; The contrast loss function L of the similarity model Γ is formula (7), where CE is the cross entropy loss function, is the label vector used for comparative training; 4. A code translation method based on a tree instruction large model according to claim 1, characterized in that: Step (5) specifically includes the following steps: 5-1) Traverse the source tree sequence AST in step (2) source and the target tree sequence AST target If the node has child nodes, the node and its child nodes are put into a list as tree tokens, which are source tree tokens DpToken source and the target tree token DpToken target ; Obtained according to the structured coding scheme of step (3) and At the same time, the text of non-leaf nodes is used as the corresponding code snippet information; 5-2) Calculate the similarity model obtained in step (3) and The similarity score X of Corresponding to the maximum similarity score 5-3) Mark the target tree Object code and Corresponding The code snippet information is filled into the <tree tag>, <code> and <code snippet> in the prompt template 1 designed in step (4), and a self-supervised instruction fine-tuning dataset is constructed; 5-4) Add the instruction fine-tuning dataset to the large language model and use the LoRA fine-tuning scheme to perform the first stage of self-supervised instruction fine-tuning.
5. A code translation method based on a tree instruction large model according to claim 1, characterized in that: Step (9) specifically comprises the following steps: On the premise that the generated results can be compiled and run, the source code and the translated code are converted into a linearized AST structure, and the normalized tree encoding is obtained according to formulas (1) and (2): and The similarity score X is calculated using the model trained in step (3), and the result is used as the basis for determining the functional consistency and identifier naming consistency after translation.
Citation Information
Patent Citations
Training methods, devices, equipment, and storage media for programming language translation models
CN112346737B
Cited By
Declarative UI automatic cross-platform translation method
CN122331956A
A declarative ui automation cross-platform translation method
CN122331956B