Large model code translation method and device fusing code functions and styles

By building functional consistency and style-oriented datasets, the code translation method of the large model is optimized, and the problem of insufficient functionality and readability in code translation technology is solved, achieving efficient and flexible code translation effect.

CN120335814AActive Publication Date: 2025-07-18HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510343085.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing code translation technology based on big model has shortcomings in the functionality and readability of the target code after translation, which is difficult to meet the actual needs of software development, especially inconsistent problems in code structure and variable naming specifications.

Method used

By building functional consistency data sets and style-oriented data sets, similarity search, fine-grained scoring and difference testing are used, combined with instruction fine-tuning and style learning training, the code translation method of the large model is optimized to ensure the functional consistency and style consistency of the generated code.

Benefits of technology

It significantly improves the functional accuracy and readability of the code translation model, reduces the dependence on large-scale models, and enables small models to achieve comparable or even better performance than large models, reduces hardware costs and improves model deployment flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335814A_ABST
    Figure CN120335814A_ABST
Patent Text Reader

Abstract

The invention provides a large model code translation method and device fusing code functions and styles, and relates to the technical field of natural language processing. The method comprises the steps that code pairs composed of source codes and target codes are obtained from an online programming platform, the code pairs are processed according to similarity retrieval, fine granularity scoring and difference testing, and a function consistency data set is constructed; performing functional learning training on the large model according to the functional consistency data set and an instruction fine tuning method to obtain a large model subjected to functional learning training; obtaining a source code, generating positive sample translation and negative sample translation of the source code, and constructing a style-oriented data set; and according to the style-oriented data set, style learning training is carried out on the large model subjected to function learning training, and a trained code translation large model is obtained. According to the method, a low-cost and high-efficiency code translation model is developed around a large-scale language model, and the correctness and readability of translated codes are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a large model code translation method and device that integrate code functions and styles. Background Art

[0002] Code translation is the process of converting code in one programming language into code in another programming language, which is widely used in software development and maintenance, such as application porting or software migration. Traditional code translation methods mainly rely on rule-based strategies, which require skilled programmers to manually handle complex situations, with high costs and low efficiency. With the development of deep learning, learning-based methods have gradually replaced traditional methods, significantly improving the performance of code translation. For example, previous researchers proposed a learnable attention mechanism to convert the syntax tree of the source language into the syntax tree of the target language, thus reducing the need for manual intervention. In recent years, large language models have performed well in tasks such as code generation and code repair, promoting the further development of code translation technology.

[0003] Existing code translation methods based on large language models usually adopt strategies such as Direct Prompt, Chain of Thought (CoT), RAG (Retrieval-Augmented Generation), and Self-Debug Prompt Learning to improve translation performance. For example, the RAG strategy achieves higher-quality translation by referring to translation data similar to the code to be translated; the self-debug prompt learning strategy compiles the generated target code and detects errors, and feeds the error information back into subsequent prompts to guide corrections.

[0004] Currently, code translation methods based on large models generally adopt the following strategies:

[0005] Direct prompt learning strategy: Prompt large language models such as GPT-4 (Generative Pre-trained Transformer) with concise instructions containing source code to generate code in the target language.

[0006] Prompt learning strategy based on the chain of thought: Induce the model to think step by step through prompt words to achieve successful translation of difficult code.

[0007] Prompt learning strategy for retrieval-augmented generation: Refer to an external knowledge base to enhance the code translation ability of the model. The external knowledge base generally includes mapping examples at the sentence level, fragment level, and function level between source code and target code.

[0008] Self-Debugging Prompt Learning Strategy: By compiling the generated target code and detecting errors, feedback the error information into subsequent prompts to guide corrections.

[0009] Although existing large model-based code translation technologies have achieved more than rule-based and traditional deep learning-based work, there are still deficiencies in the functionality and readability of the target code after translation.

[0010] In terms of correctness, out-of-the-box large models still struggle to meet the actual software development requirements in terms of the correctness of the generated code under the optimized prompt learning strategy. For example, the average success rate of the StarCoder-3B model on the traditional code translation benchmark CodeNet is only 7%.

[0011] In terms of readability, even if the translated code functions correctly, it often fails to maintain the style of the source code, including code structure and variable naming conventions. This style inconsistency increases the reading burden on developers. For example, Figure 1 shows the Python code translated by Qwen-7B. In addition to the code correctness issues, there are many style problems in the translated Python, such as inconsistent variable naming with the source code, mismatched function encapsulation, and lack of I / O optimization.

[0012] Existing technologies mainly rely on rule-based methods, deep learning methods, and large model-based prompt strategies. However, these methods have limitations at different levels:

[0013] (1) Rule-based methods: Require professional programmers to manually define translation rules between languages, which is costly and inefficient.

[0014] (2) Deep learning-based methods: Although the translation performance has been improved through unsupervised pre-training and backtranslation strategies, its dependence on high-quality parallel data (i.e., code translation data containing source and target languages) limits its scope of application.

[0015] (3) Large model-based methods: Although the translation performance has been improved through strategies such as direct prompting, chain of thought reasoning, retrieval-augmented generation, and self-debugging prompts, these methods still struggle to ensure both functional correctness and code readability simultaneously. Summary of the Invention

[0016] To solve the technical problems existing in the prior art, that is, in terms of correctness, for out-of-the-box large models under the optimized prompt learning strategy, the correctness of the generated code still difficult to meet the actual software development requirements; in terms of readability, although the translated code functions correctly, it often fails to maintain the style of the source code, including code structure and variable naming conventions, and this style inconsistency increases the reading burden on developers, the embodiments of the present invention provide a large model code translation method and device that integrate code function and style. The technical solutions are as follows:

[0017] On the one hand, a large model code translation method that integrates code function and style is provided. This method is implemented by a large model code translation device, and the method includes:

[0018] S1. Obtain code pairs composed of source code and target code from an online programming platform, process the code pairs according to similarity retrieval, fine-grained scoring, and differential testing, and construct a functional consistency dataset.

[0019] S2. Perform functional learning training on the large model according to the functional consistency dataset and the instruction fine-tuning method to obtain a large model trained through functional learning.

[0020] S3. Obtain the source code, generate positive sample translations and negative sample translations of the source code, and construct a style-oriented dataset.

[0021] S4. Perform style learning training on the large model trained through functional learning according to the style-oriented dataset to obtain a trained code translation large model.

[0022] S5. Obtain the code to be translated, input it into the trained code translation large model, and obtain the code translation result.

[0023] Optionally, the processing of the code pairs according to similarity retrieval, fine-grained scoring, and differential testing in S1 to construct a functional consistency dataset includes:

[0024] S11. Perform similarity retrieval on the code pairs to obtain similar code pairs.

[0025] S12. Perform fine-grained scoring on the similar code pairs, and screen the similar code pairs according to the fine-grained scoring results to obtain the screened code pairs.

[0026] S13. Perform differential testing on the screened code pairs to obtain the tested code pairs.

[0027] S14. Construct a functional consistency dataset according to the tested code pairs.

[0028] Optionally, the fine-grained scoring of the similar code pairs in S1 includes:

[0029] The fine - grained scoring of similar code pairs is performed through the following formulas (1) and (2):

[0030] (1)

[0031] (2)

[0032] In the formulas, represents the finally determined fine - grained score, represents the source code and the target code 's similarity score, represents and with a score of 's softmax - normalized probability, represents the prompt template of the large model, represents the exponential function with base e, represents the and generated by the large model with a score of 's log - likelihood probability, represents the per - score calculation of all scores, represents the and generated by the large model with a score of 's log - likelihood probability.

[0033] Optionally, the optimization objective function for the functional learning training in S2 is as shown in the following formula (3):

[0034] (3)

[0035] In the formula, represents the optimization objective function, represents the target code 's th token, represents the prompt word input into the large model for translating the source code into the target code, represents the target code 's first tokens.

[0036] Optionally, the positive - sample translation and negative - sample translation of the generated source code in S3 include:

[0037] S31. Generate multiple translation candidates with consistent styles for the source code, screen out the functionally correct translations from the translation candidates through differential testing, and select the optimal translation as the positive - sample translation using the style - consensus selection mechanism, as shown in the following formula (4):

[0038] (4)

[0039] In the formula, represents the positive sample translation, Indicates the correct translation of all features filtered out by differential detection, represents the function used to measure the style similarity between two target language codes, Indicates the target code No. Tokens, Indicates the target code No. tokens.

[0040] S32, generating multiple negative translation candidates for the source code, and selecting translations whose style differences from the positive sample translations exceed a preset threshold from the negative translation candidates through a function as negative sample translations.

[0041] Optionally, the optimization objective function of the style learning training in S4 is as shown in the following equations (5) and (6):

[0042] (5)

[0043] (6)

[0044] In the formula, represents the optimization objective function of style learning training, Represents source code, represents the positive sample translation, represents the set of all negative translations, represents the exponential function with base e, Indicates that the large model is inputting source code When , a positive translation is successfully generated The probability of Represents the set of all target codes, Representation model for source code Generate object code The probability of Indicates that the large model is inputting source code And generate target code Before When the token is generated successfully The probability of a token, express No. Tokens, Represents the input to the large model for the source code The prompt for translating into target code denotes the first n token sequences

[0045] Optionally, the style learning training further includes instruction fine-tuning on the forward translation, as shown in Equation (7) below:

[0046] (7)

[0047] In the formula denotes the value of the loss function for style learning denotes the balance hyperparameter between the two losses denotes the optimization objective function for style learning training denotes and the instruction fine-tuning loss function between denotes the positive sample translation

[0048] On the other hand, a large model code translation device that fuses code function and style is provided. This device is applied to the large model code translation method that fuses code function and style. The device includes:

[0049] A functional consistency dataset construction module, which is used to obtain code pairs composed of source code and target code from an online programming platform, and process the code pairs according to similarity retrieval, fine-grained scoring, and differential testing to construct a functional consistency dataset

[0050] A functional learning training module, which is used to perform functional learning training on the large model according to the functional consistency dataset and the instruction fine-tuning method to obtain a large model trained by functional learning

[0051] A style-oriented dataset construction module, which is used to obtain source code, generate positive sample translations and negative sample translations of the source code, and construct a style-oriented dataset

[0052] A style learning training module, which is used to perform style learning training on the large model trained by functional learning according to the style-oriented dataset to obtain a trained large model for code translation

[0053] A code translation module, which is used to obtain the code to be translated and input it into the trained large model for code translation to obtain the code translation result

[0054] Optionally, the functional consistency dataset construction module is further used for:

[0055] S11. Perform similarity retrieval on the code pairs to obtain similar code pairs

[0056] S12. Perform fine-grained scoring on the similar code pairs, and filter the similar code pairs according to the fine-grained scoring results to obtain the filtered code pairs.

[0057] S13. Perform differential testing on the filtered code pairs to obtain the tested code pairs.

[0058] S14. Construct a functional consistency dataset based on the tested code pairs.

[0059] Optionally, the functional consistency dataset construction module is used for:

[0060] Perform fine-grained scoring on the similar code pairs through the following formulas (1) and (2):

[0061] (1)

[0062] (2)

[0063] In the formulas, represents the finally determined fine-grained score, represents the source code and the target code similarity score, represents and with a score of softmax normalization probability, represents the prompt template of the large model, represents the exponential function with base e, represents the and generated by the large model with a score of log-likelihood probability, represents the calculation of each score one by one for all scores, represents the and generated by the large model with a score of log-likelihood probability.

[0064] Optionally, the optimization objective function for functional learning training is as shown in the following formula (3):

[0065] (3)

[0066] In the formula, represents the optimization objective function, represents the target code the th token, represents the prompt word input into the large model for translating the source code into the target code. Represents the target code of the first tokens.

[0067] Optionally, the style - oriented dataset construction module is further used for:

[0068] S31. Generate multiple translation candidates with consistent styles for the source code, filter out the translations with correct functions from the translation candidates through differential testing, and select the optimal translation as the positive - sample translation from the translations with correct functions using the style consensus selection mechanism, as shown in the following formula (4):

[0069] (4)

[0070] In the formula, represents the positive - sample translation, represents all the translations with correct functions filtered out by differential detection, represents a function used to measure the style similarity between two target - language codes, represents the target code of the th token, represents the target code of the th token.

[0071] S32. Generate multiple negative - translation candidates for the source code, and filter out the translations with a style difference from the positive - sample translation exceeding a preset threshold as negative - sample translations through a function.

[0072] Optionally, the optimization objective function for style learning training is as shown in the following formulas (5) and (6):

[0073] (5)

[0074] (6)

[0075] In the formula, represents the optimization objective function for style learning training, represents the source code, represents the positive - sample translation, represents the set of all negative - sample translations, represents the exponential function with base e, represents the probability that the large - model successfully generates the positive - sample translation when the input is the source code of, represents the set of all target codes, represents that the model generates the target code for the source code The probability, indicating that when the large model processes the input source code and generates the target code for the first several tokens, the probability of successfully generating the th token, indicating the th token, indicating the prompt word input into the large model for translating the source code into the target code, indicating the first token sequence.

[0076] Optionally, the style learning training further includes instruction fine-tuning on the forward translation, as shown in the following formula (7):

[0077] (7)

[0078] In the formula, represents the value of the loss function for style learning, represents the balance hyperparameter between the two losses, represents the optimization objective function for style learning training, represents and the instruction fine-tuning loss function between them, represents the positive sample translation.

[0079] On the other hand, a large model code translation device is provided. The large model code translation device includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned large model code translation method that integrates code function and style is implemented.

[0080] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by the processor to implement any one of the methods in the above-mentioned large model code translation method that integrates code function and style.

[0081] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0082] In the present invention, the functional correctness of the code translation model is significantly improved: In the first stage of F2STrans, functional learning, by mining high-quality source-target code pairs and combining differential testing to ensure the functional consistency of the generated code. Specifically: First, use the code data on the online programming platform to screen out code pairs with the same input-output behavior. Then, differential testing further verifies the functional consistency of these code pairs, thus constructing a high-quality training dataset. Finally, function-oriented learning improves the model performance.

[0083] Experiments show that functional learning significantly improves the functional correctness of the generated code. For example, in the CodeNet benchmark test, the Qwen-0.5B model after functional learning outperformed the Qwen-32B model based on the RAG strategy.

[0084] Enhance the readability of code translation: In the second stage of F2STrans, style learning, through positive and negative style example contrast learning, enables the model to generate target code that is both functionally correct and style-consistent. Specifically, it includes: using a powerful LLM (such as Qwen32B) to generate multiple forward translation candidates with consistent styles, and screening out the optimal translation through differential testing. Suppress the generation results with inconsistent styles through negative translation examples. Introduce a style consensus selection mechanism and a contrast loss function (List-wise Loss Function) to further optimize the effect of style learning.

[0085] Style learning significantly improves the readability and style consistency of the generated code. For example, in the benchmark test of the present invention, the improved Qwen-0.5B model of F2STrans achieved a high score of 80.7 in the CCSim metric, exceeding GPT-4 and other baseline models.

[0086] Significantly reduce the model size dependence: F2STrans enables small models to achieve performance comparable to or even better than large models through an efficient training strategy. For example, Qwen-1.5B performs better than Qwen32B and GPT-4 under function-to-style guidance. This feature significantly reduces the dependence on large-scale models, enabling resource-constrained developers to use small models to complete high-quality code translation tasks. This not only reduces the hardware cost but also improves the deployment flexibility of the model. Brief Description of the Drawings

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0088] Figure 1 It is a schematic diagram of the problems existing in the correctness and readability of the code for large model translation provided by the embodiments of the present invention;

[0089] Figure 2 It is a flowchart of a large model code translation method that integrates code function and style provided by the embodiments of the present invention;

[0090] Figure 3 It is a flowchart of large model function learning provided by the embodiments of the present invention;

[0091] Figure 4 It is a flowchart of large model style learning provided by the embodiments of the present invention;

[0092] Figure 5 It is a block diagram of a large model code translation device that integrates code function and style provided by the embodiments of the present invention;

[0093] Figure 6 It is a schematic diagram of the structure of a large model code translation device provided by the embodiments of the present invention. Detailed implementation manners

[0094] Next, the technical solutions in the present invention will be described with reference to the accompanying drawings.

[0095] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" aims to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0096] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0097] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0098] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0099] An embodiment of the present invention provides a large model code translation method that integrates code functions and styles. This method can be implemented by a large model code translation device, which can be a terminal or a server. As Figure 2 shown in the flowchart of the large model code translation method that integrates code functions and styles, the processing flow of this method can include the following steps:

[0100] S1. Obtain a code pair consisting of source code and target code from an online programming platform, process the code pair according to similarity retrieval, fine-grained scoring, and differential testing, and construct a functional consistency dataset.

[0101] In a feasible implementation, to ensure the quality of training data, the present invention mines high-quality source-target code pairs from an online programming platform (such as Codeforces) and filters out code pairs with the same input-output behavior.

[0102] Specifically, function learning is the first stage of the function-to-style guidance paradigm F2STrans. Its core goal is to ensure that the generated target code is completely consistent with the source code in function. As Figure 3 shown, the specific implementation steps include the construction of functional consistency data in step S1 and the training of function learning in step S2.

[0103] Optionally, processing the code pair according to similarity retrieval, fine-grained scoring, and differential testing in S1 to construct a functional consistency dataset may include the following steps S11 - S14:

[0104] S11. Perform similarity retrieval on the code pair to obtain similar code pairs.

[0105] S12. Perform fine-grained scoring on the similar code pairs, and filter the similar code pairs according to the fine-grained scoring results to obtain the filtered code pairs.

[0106] S13. Perform differential testing on the filtered code pairs to obtain the tested code pairs.

[0107] S14. Construct a functional consistency dataset according to the tested code pairs.

[0108] In a feasible implementation, Relevance-driven CodePair Selection: Use the lightweight code embedding model Jina to retrieve similar code pairs, and perform fine-grained scoring on the candidate code pairs through a large model. The scoring formula is as follows:

[0109] (1)

[0110] (2)

[0111] Wherein, represents the finally determined fine-grained score, represents the source code and the target code similarity score, represents and the score is softmax normalization probability, represents the prompt template of the large model, represents the exponential function with base e, represents the and generated by the large model with a score of log-likelihood probability, represents the per-score calculation of all scores, represents the and generated by the large model with a score of log-likelihood probability.

[0112] Furthermore, differential testing: perform differential testing on the selected code pairs, that is, run the source code and the target code under the same input and compare their output results. Only the code pairs with exactly the same input-output behavior will be retained.

[0113] S2. Perform functional learning training on the large model according to the functional consistency dataset and the instruction fine-tuning method to obtain a large model trained through functional learning training.

[0114] In a feasible implementation, use the instruction fine-tuning (IFT) method to train the base model, and optimize the objective function as follows:

[0115] (3)

[0116] Wherein, represents the optimized objective function, represents the th token of the target code represents the prompt word input into the large model for translating the source code into the target code, represents the first tokens of the target code.

[0117] The present invention ensures the functional consistency of the generated code by mining high-quality source-target code pairs and combining differential testing (Differential Testing). A correlation-driven code pair selection mechanism and an LLM-based (LLM Judge) scoring system are used to screen out code pairs with the same input-output behavior, and a high-quality training dataset is constructed. Differential testing further verifies the functional consistency of these code pairs, thereby significantly improving the functional correctness of the generated code.

[0118] S3. Obtain the source code, generate positive and negative sample translations of the source code, and construct a style-oriented dataset.

[0119] In a feasible implementation, style learning is the second stage of F2S Trans, and its core goal is to improve the readability and style consistency of the generated code. As Figure 4 shown, the specific implementation steps include the style-oriented data construction in step S3 and the style learning training in step S4.

[0120] Optionally, generating positive and negative sample translations of the source code in S3 may include the following steps S31 - S32:

[0121] S31. Generate multiple translation candidates with consistent styles for the source code, screen out the functionally correct translations from the translation candidates through differential testing, and select the optimal translation as the positive sample translation using a style consensus selection mechanism.

[0122] In a feasible implementation, to help the model learn the style features of the source code, the present invention constructs positive and negative style examples. Among them, the positive translation uses a powerful large model (such as Qwen 32B) to generate multiple translation candidates with consistent styles, and screens out the functionally correct translations through differential testing. Subsequently, a style consensus selection mechanism (Style Consensus Selection, SCS) is used to select the optimal translation:

[0123] (4)

[0124] In the formula, represents the positive sample translation, represents all the functionally correct translations screened out by differential detection, represents a function used to measure the style similarity between two target language codes, represents the target code the th represents the target code the th

[0125] S32. Generate multiple negative translation candidates for the source code, and filter out the translations with a style difference from the positive sample translation exceeding a preset threshold from the negative translation candidates as negative sample translations through a function.

[0126] In a feasible implementation manner, the negative sample translation uses a base model to generate multiple negative translation candidates, and filters out the samples with a large style difference from the positive translation through a function.

[0127] S4. Perform style learning training on the large model trained through functional learning according to the style-oriented dataset to obtain a trained large model for code translation.

[0128] Optionally, the present invention designs a loss function with reference to the idea of contrastive learning. The involved list loss function encourages the model to generate target code with consistent styles while suppressing inconsistent translations:

[0129] (5)

[0130] (6)

[0131] In the formula, represents the optimization objective function of style learning training, represents the source code, represents the positive sample translation, represents the set of all negative sample translations, represents the exponential function with base e, represents that when the large model inputs the source code , it successfully generates the positive sample translation with a probability of, represents the set of all target codes, represents the probability that the model generates the target code for the source code of, represents that when the large model inputs the source code and generates the target code in the first tokens, the probability of successfully generating the th token is, represents the th token of, represents the prompt word input into the large model for translating the source code into the target code, represents the first token sequence of.

[0132] Optionally, the present invention also performs instruction fine-tuning on the forward translation to further emphasize the importance of forward translation:

[0133] (7)

[0134] In the formula, represents the value of the loss function for style learning, represents the balance hyperparameter between the two losses, represents the optimization objective function for style learning training, represents and the instruction fine-tuning loss function between represents the positive sample translation.

[0135] The present invention uses positive and negative style example contrast learning to enhance the model's understanding and reasoning ability of multimodal information; introduces a style consensus selection mechanism to select the optimal translation from multiple candidate translations to ensure the style consistency of the generated code; and proposes a list-wise loss function to further optimize the style learning effect by suppressing translation results with inconsistent styles.

[0136] The two-stage training framework from function to style proposed by the present invention: In the first stage, the translation correctness is optimized through function learning, and in the second stage, the translation readability and style consistency are improved through style learning. This framework significantly reduces the dependence on large-scale models, enabling small models to achieve performance comparable to or even better than large models.

[0137] S5. Obtain the code to be translated, input it into the trained large code translation model, and obtain the code translation result.

[0138] The present invention proposes a brand-new function-to-style guidance paradigm (F2STrans) for large models, aiming to significantly improve the performance of large language models in code translation tasks. This paradigm gradually optimizes the functional correctness and readability of the translated code through a two-stage training method:

[0139] (1) Functional Learning: Train using high-quality source-target code pairs to ensure that the generated target code is functionally consistent with the source code.

[0140] (2) Style Learning: Train based on positive and negative style examples to improve the readability and style consistency of the generated code.

[0141] Through the above method, the present invention not only solves the deficiencies of the prior art in terms of functional correctness, but also significantly improves the readability of the translated code, making it more in line with the coding habits of developers and team norms.

[0142] In the embodiments of the present invention, the functional correctness of the code translation model is significantly improved: in the first stage of F2STrans, function learning, by mining high-quality source-target code pairs and combining differential testing to ensure the functional consistency of the generated code. Specifically: First, use the code data on the online programming platform to screen out code pairs with the same input-output behavior. Then differential testing further verifies the functional consistency of these code pairs, thus constructing a high-quality training dataset. Finally, function-oriented learning improves the model performance.

[0143] Experiments show that function learning significantly improves the functional correctness of the generated code. For example, in the CodeNet benchmark test, the Qwen-0.5B model after function learning outperforms the Qwen-32B model based on the RAG strategy.

[0144] Enhance the readability of code translation: in the second stage of F2STrans, style learning, through positive and negative style example contrast learning, enables the model to generate target code that is both functionally correct and style-consistent. Specifically, it includes: using a powerful LLM (such as Qwen32B) to generate multiple forward translation candidates with consistent styles, and screening out the optimal translation through differential testing. Suppress the generation results with inconsistent styles through negative translation examples. Introduce a style consensus selection mechanism and a contrast loss function (List-wise Loss Function) to further optimize the effect of style learning.

[0145] Style learning significantly improves the readability and style consistency of the generated code. For example, in the benchmark test of the present invention, the improved Qwen-0.5B model of F2STrans achieved a high score of 80.7 in the CCSim metric, outperforming GPT-4 and other baseline models.

[0146] Significantly reduce the dependence on model scale: F2STrans enables small models to achieve performance comparable to or even better than large models through an efficient training strategy. For example, Qwen-1.5B performs better than Qwen32B and GPT-4 under function-to-style guidance. This feature significantly reduces the dependence on large-scale models, enabling resource-constrained developers to use small models to complete high-quality code translation tasks. This not only reduces the hardware cost but also improves the deployment flexibility of the model.

[0147] Figure 5It is a block diagram of a large model code translation device that integrates code function and style, which is used for the large model code translation method that integrates code function and style. Refer to Figure 5 This device includes a function consistency dataset construction module 310, a function learning and training module 320, a style-oriented dataset construction module 330, a style learning and training module 340, and a code translation module 350. Among them:

[0148] The function consistency dataset construction module 310 is used to obtain code pairs composed of source code and target code from an online programming platform, process the code pairs according to similarity retrieval, fine-grained scoring, and differential testing, and construct a function consistency dataset.

[0149] The function learning and training module 320 is used to perform function learning and training on the large model according to the function consistency dataset and the instruction fine-tuning method, and obtain a large model trained through function learning.

[0150] The style-oriented dataset construction module 330 is used to obtain source code, generate positive sample translations and negative sample translations of the source code, and construct a style-oriented dataset.

[0151] The style learning and training module 340 is used to perform style learning and training on the large model trained through function learning according to the style-oriented dataset, and obtain a trained code translation large model.

[0152] The code translation module 350 is used to obtain the code to be translated, input it into the trained code translation large model, and obtain the code translation result.

[0153] In the embodiment of the present invention, the functional correctness of the code translation model is significantly improved: in the first stage of F2STrans, function learning, by mining high-quality source-target code pairs and combining differential testing to ensure the functional consistency of the generated code. Specifically: First, use the code data on the online programming platform to screen out code pairs with the same input-output behavior. Then differential testing further verifies the functional consistency of these code pairs, thus constructing a high-quality training dataset. Finally, function-oriented learning improves the model performance.

[0154] Experiments show that function learning significantly improves the functional correctness of the generated code. For example, in the CodeNet benchmark test, the Qwen-0.5B model after function learning outperforms the Qwen-32B model based on the RAG strategy.

[0155] Enhancing the readability of code translation: In the second stage of F2STrans, style learning, through the comparison and learning of positive and negative style examples, enables the model to generate target code that is both functionally correct and stylistically consistent. Specifically, it includes: using a powerful LLM (such as Qwen32B) to generate multiple stylistically consistent positive translation candidates, and screening out the optimal translation through differential testing. Suppressing stylistically inconsistent generation results through negative translation examples. Introducing a style consensus selection mechanism and a contrastive loss function (List-wise Loss Function) to further optimize the effect of style learning.

[0156] Style learning significantly improves the readability and stylistic consistency of the generated code. For example, in the benchmark test of the present invention, the improved Qwen-0.5B model of F2STrans achieved a high score of 80.7 in the CCSim metric, surpassing GPT-4 and other baseline models.

[0157] Significantly reducing the dependence on model size: F2STrans enables small models to achieve performance comparable to or even better than large models through an efficient training strategy. For example, Qwen-1.5B performs better than Qwen32B and GPT-4 under function-to-style guidance. This feature significantly reduces the dependence on large-scale models, enabling resource-constrained developers to use small models to complete high-quality code translation tasks. This not only reduces hardware costs but also improves the deployment flexibility of the model.

[0158] Figure 6 is a schematic structural diagram of a large model code translation device provided by an embodiment of the present invention. As Figure 6 shown, the large model code translation device may include the above-mentioned Figure 5 large model code translation device that integrates code function and style. Optionally, the large model code translation device 410 may include a first processor 2001.

[0159] Optionally, the large model code translation device 410 may further include a memory 2002 and a transceiver 2003.

[0160] Among them, the first processor 2001, the memory 2002, and the transceiver 2003, such as, may be connected through a communication bus.

[0161] Next, in combination with Figure 6 specific introductions will be made to the various components of the large model code translation device 410:

[0162] Among them, the first processor 2001 is the control center of the large model code translation device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0163] Optionally, the first processor 2001 can execute various functions of the large model code translation device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0164] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 shown in

[0165] In a specific implementation, as an embodiment, the large model code translation device 410 may also include multiple processors, such as Figure 6 the first processor 2001 and the second processor 2004 shown in

[0166] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0167] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 6 not shown in) of the large model code translation device 410. The embodiments of the present invention do not make specific limitations in this regard.

[0168] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0169] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0170] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 6 not shown in) of the large model code translation device 410. The embodiments of the present invention do not make specific limitations in this regard.

[0171] It should be noted that Figure 6 the structure of the large model code translation device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0172] In addition, the technical effects of the large model code translation device 410 may refer to the technical effects of the large model code translation method that integrates code functions and styles described in the above method embodiments, and will not be elaborated here.

[0173] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0174] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0175] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0176] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0177] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0178] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0179] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0182] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0183] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0184] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0185] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A large model code translation method that integrates code functions and styles, characterized in that, The method includes: S1. Obtain a code pair consisting of source code and target code from an online programming platform, process the code pair according to similarity retrieval, fine-grained scoring, and differential testing, and construct a functional consistency dataset; S2. Perform functional learning training on a large model according to the functional consistency dataset and the instruction fine-tuning method to obtain a large model trained through functional learning; S3. Obtain the source code, generate positive sample translations and negative sample translations of the source code, and construct a style-oriented dataset; S4. Perform style learning training on the large model trained through functional learning according to the style-oriented dataset to obtain a trained code translation large model; S5. Obtain the code to be translated, input it into the trained code translation large model, and obtain the code translation result.

2. The method for translating large model code that integrates code functions and styles according to claim 1, wherein The processing of the code pair according to similarity retrieval, fine-grained scoring, and differential testing in S1 to construct a functional consistency dataset includes: S11. Perform similarity retrieval on the code pair to obtain similar code pairs; S12. Perform fine-grained scoring on the similar code pairs, screen the similar code pairs according to the fine-grained scoring results to obtain the screened code pairs; S13. Perform differential testing on the screened code pairs to obtain the tested code pairs; S14. Construct a functional consistency dataset according to the tested code pairs.

3. The method for translating large model code that integrates code functions and styles according to claim 1, wherein The fine-grained scoring of the similar code pairs in S1 includes: Perform fine-grained scoring on the similar code pairs through the following formulas (1) and (2): (1) (2) Wherein, represents the finally determined fine-grained score, represents the source code and the target code similarity score, represents and the score is softmax normalization probability, represents the prompt template of the large model, represents the exponential function with base e, represents the and generated by the large model with a score of log-likelihood probability, represents the per-score calculation of all scores, represents the and generated by the large model with a score of log-likelihood probability.

4. The method for translating large model code that integrates code functions and styles according to claim 1, wherein The optimization objective function of the functional learning training in S2 is as shown in the following formula (3): (3) In the formula, represents the optimization objective function, represents the target code of the th token, represents the prompt input into the large model for translating the source code into the target code, represents the first tokens of the target code.

5. The method for translating large model code that integrates code functions and styles according to claim 1, characterized in that, The generation of positive sample translations and negative sample translations of the source code in S3 includes: S31. Generate multiple translation candidates with consistent styles for the source code, screen out the translations with correct functions from the translation candidates through differential testing, and select the optimal translation as the positive sample translation from the translations with correct functions through a style consensus selection mechanism, as shown in the following formula (4): (4) In the formula, represents the positive sample translation, represents all the translations with correct functions screened by differential detection, represents a function used to measure the style similarity between two target language codes, represents the target code the -th token of represents the target code the -th token; S32. Generate multiple negative translation candidates for the source code, and screen out the translations with a style difference from the positive sample translation exceeding a preset threshold as the negative sample translations through a function.

6. The method for translating large model code that integrates code functions and styles according to claim 1, characterized in that, The optimization objective functions of the style learning training in S4 are as shown in the following formulas (5) and (6): (5) (6) In the formula, represents the optimization objective function for style learning training, represents the source code, represents the positive sample translation, represents the set of all negative sample translations, represents the exponential function with base e, represents that when the large model inputs the source code it successfully generates the positive sample translation with probability represents the set of all target codes, represents the probability that the model generates the target code for the source code with probability represents that when the large model inputs the source code and generates the target code in the first tokens, the probability of successfully generating the th token, represents the th token of represents the prompt input to the large model for translating the source code into the target code, represents the first token sequence of 7. The method for translating large model code that integrates code functions and styles according to claim 6, characterized in that, The style learning training further includes performing instruction fine-tuning on the forward translation, as shown in the following formula (7): (7) In the formula, represents the value of the loss function for style learning, represents the balance hyperparameter between the two losses, represents the optimization objective function for style learning training, represents and the instruction fine-tuning loss function between represents positive sample translation.

8. A large model code translation device that integrates code functions and styles, the large model code translation device that integrates code functions and styles is used to implement the large model code translation method that integrates code functions and styles as described in any one of claims 1-7, characterized in that, The device includes: A functional consistency dataset construction module, configured to obtain a code pair consisting of source code and target code from an online programming platform, process the code pair according to similarity retrieval, fine-grained scoring, and differential testing, and construct a functional consistency dataset; A functional learning training module, configured to perform functional learning training on a large model according to the functional consistency dataset and the instruction fine-tuning method to obtain a large model trained through functional learning; A style-oriented dataset construction module, configured to obtain the source code, generate positive sample translations and negative sample translations of the source code, and construct a style-oriented dataset; A style learning and training module, which is used to perform style learning and training on the large model trained by function learning according to the style-oriented dataset to obtain a trained large model for code translation; A code translation module, which is used to obtain the code to be translated and input it into the trained large model for code translation to obtain a code translation result.

9. A large model code translation device, characterized in that, The large model code translation device includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Program translation model obtaining method, program translation method and program translation device

    CN116795379A

  • Training / application method of representation learning model, equipment and medium

    CN118051774A

  • Simultaneous interpretation model training method and device, equipment and storage medium

    CN118395999A

  • Large language model code translation method based on prompt fine tuning

    CN118963756A

  • Code positioning method based on programming language migration

    CN119045880A

Cited By

  • Code generation method and system based on data enhancement and mixed similarity negative sampling

    CN122132012A