Code translation prompt augmented with compiled code candidates
By augmenting the prompt for a large language model with compiled code candidates from a database, the accuracy of source code translation across programming languages is enhanced, addressing the limitations of existing translation methods.
Patent Information
- Application Number
- PCT/CN2023/131162
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Large language models struggle to accurately translate source code from one programming language to another, especially when they have not been trained on the syntax or semantics of the target language.
The prompt for the large language model is augmented with compiled code candidates written in the target programming language, selected from a database based on embedding similarity and filtered for low compilation errors, to guide the model towards generating a more accurate translation.
This approach significantly improves the accuracy of source code translation by providing the model with context-specific code examples, reducing the need for fine-tuning and lowering computational costs.
Smart Images

Figure CN2023131162_22052025_PF_FP_ABST
Abstract
Description
CODE TRANSLATION PROMPT AUGMENTED WITH COMPILED CODE CANDIDATESBACKGROUND
[0001] A large language model (LLM) is a type of machine learning model trained on a massively-large training dataset of text and / or source code resulting in the model containing billions of parameters. The large language model is used to perform various tasks such as natural language processing, text generation, machine translation, and / or source code generation. The large language model is based on deep learning neural networks such as a neural transformer model with attention.
[0002] The large language model is typically given a user prompt that consists of text in the form of a question, an instruction, short paragraph and / or source code that instructs the model to perform a task and / or the format of the intended response. However, when the large language model has not been trained on the task, the large language model produces useless and / or vague responses. This occurs often when the task pertains to generating source code where the model has not seen the syntax of the programming language during training or the semantics of the source code written in the programming language.SUMMARY
[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] A prompt to a large language model for a translation of an input source code snippet written in a source programming language into a different target programming language is augmented with compiled code candidates written in the target programming language.
[0005] A database of code snippets in the target programming language is searched to find the top-k code candidates having an embedding that is closely similar to an embedding of the input source code snippet. A reward model generates a reward score for each of the top-k code candidates which is used to identify the top-n code candidates having less compilation errors. The top-n code candidates are included in the prompt for the large language model to guide the model towards generating a more accurate translation in the target programming language.
[0006] These and other features and advantages will be apparent from a reading of the following detailed description and a review of the associated drawings. It is to be understood that both the foregoing general description and the following detailed description are explanatory only and are not restrictive of aspects as claimed.BRIEF DESCRIPTION OF DRAWINGS
[0007] Fig. 1 illustrates an exemplary system for creating the reward model.
[0008] Fig. 2 illustrates an exemplary system for generating a prompt to a large language model for a code translation augmented with compiled code candidates.
[0009] Fig. 3 is a diagram of an exemplary configuration of the reward model.
[0010] Fig. 4 is a flow chart illustrating an exemplary method of the code translation system.
[0011] Fig. 5 is a flow chart illustrating an exemplary method for generating the reward model training dataset.
[0012] Fig. 6 is a flow chart illustrating an exemplary method for training the reward model with the reward model training dataset.
[0013] Fig. 7 is a flow chart illustrating an exemplary method for generating a code translation from a large language model.
[0014] Fig. 8 is a schematic diagram illustrating a first exemplary operating environment.
[0015] Fig. 9 is a block diagram illustrating a second exemplary operating environment.DETAILED DESCRIPTION
[0016] Overview
[0017] A prompt to a large language model for a translation of a first source code snippet written in a first programming language into a target programming language is augmented with examples or code candidates of compiled source code snippets written in the target programming language. The compiled code candidates are used to guide the large language model’s understanding of the syntax of the target programming language and the semantics of the first source code snippet.
[0018] Large language models have powerful reasoning capabilities having been trained on terabytes of natural language text and / or source code. It is often not possible to fine-tune a large language model to learn a desired task. Instead, the large language model learns the desired task from the examples provided in the prompt. The use of the compiled code candidates in a prompt enables the model to learn the correct syntax of the target programming language and the semantics of the input source code. This is beneficial when the large language model has not been trained on source code written in the first or source programming language or trained on source code in general. The examples guide the model toward making accurate inferences using the information the model has been trained to predict.
[0019] The examples are selected from a database of source code snippets. The source code snippets are extracted from methods and classes of publicly-available source code repositories. A snippet is of a pre-determined length and indexed by an embedding that represents the source code snippet. An embedding of an input source code written in a source programming language is used as a search index to find code candidates having closely-similar embeddings in the database of the target programming language. The source code snippet having the top-k closest embeddings in the database of the target programming language to the embedding of the input source code are extracted and considered the top-k code candidates, where k is a user-defined value.
[0020] Each of the top-k code candidates is then input into a reward model that generates a reward score for the candidate. In an aspect, the reward score is a predicted number of compiler errors for a code candidate. The candidates are ranked by reward score and the candidates having the top-n reward scores are selected and included in the prompt to the large language model, where n is a user-defined value.
[0021] The reward model is trained using reinforcement learning. A training dataset of training samples is generated from the top-n code candidates. Translations are generated for each of the top-k code candidates and compiled. The reward score is generated for each of the top-k code candidate and its translation. The training samples are ranked based on their respective reward score and put into pairs where each pair includes a positive sample and a negative sample. The paired samples are used to train the reward model to learn to predict the number of compilation errors for a prompt based on the input source code snippet and its code candidates.
[0022] Attention now turns to a more detailed description of the system, device, and methods of the code translation system.
[0023] System
[0024] Fig. 1 illustrates a block diagram of an exemplary system for generating the reward model 100. In an aspect, the system 100 includes several code databases 102A-102M ( “102” ) , a retrieval engine 104, an encoder 106, a prompt generator 108, a large language model 110, a compiler 112, a rank engine 114, a pairwise sampler 116, a reinforcement learning engine 118, and a reward model 120.
[0025] Each of the code databases 102 contains source code snippets or chunks written in a particular programming language. Each source code chunk is indexed by a respective embedding 122 produced by the encoder 106. The encoder 106 takes an input sequence of tokens from a source code chunk and produces a fixed-length vector representation referred to as an embedding. The embedding captures the semantic relationships between the tokens and places similar inputs close together in an embedding space. In an aspect, the encoder 106 is an encoder-only neural transformer model with attention. The encoder 106 generates an embedding for each code chunk which is then used as an embedding index.
[0026] The retrieval engine 104 searches for the top-k code candidates or chunks 124 in the database of the target programming language that are similar to an input source code snippet 128 by producing an embedding of the input source code snippet generated from the encoder 106.
[0027] The prompt generator 108 generates a prompt 128 for each of the top-k code candidates 126, one at a time, for the large language model 110 to generate a translation of each code candidate into the target programming language. The translated code 130 is then compiled by the compiler 112 and a count of the number of compiler errors 132 is output. The combination of the input source code snippet, code candidate 134 and number of compiler errors 132 forms a training sample 136.
[0028] A training sample is generated for each of the input source code snippets 128. The rank engine 114 ranks the training samples 136 by the count of compilation errors. The pairwise sampler 116 generates paired training samples 140 where each pair consists of a positive sample and a negative sample. A positive sample has less compilation errors than its paired negative sample. The paired samples 140 are then used to train a pre-trained model to learn to predict the number of compilation errors that the code contained in a prompt may produce.
[0029] In an aspect, the large language model is pre-trained on natural language text and source code. The training of a large language model requires a considerable amount of training data and computing resources which makes it impossible for some developers to create their own models. The large language model consists of billions of parameters (e.g., weights, embeddings) and trained on terabytes of data. Examples of the large language models include the conversational pre-trained generative neural transformer models with attention offered by OpenAI (i.e., ChatGPT and Codex models) , PaLM and Chinchilla by Google, and LLaMa by Meta.
[0030] In an aspect, the large language model may be implemented as a neural transformer model with attention. The neural transformer model with attention is one distinct type of machine learning model. Machine learning pertains to the use and development of computer systems that are able to learn and adapt without following explicit instructions by using algorithms and statistical models to analyze and draw inferences from patterns in data. Machine learning uses different types of statistical methods to learn from data and to predict future decisions. Traditional machine learning includes classification models, data mining, Bayesian networks, Markov models, clustering, and visual data mapping.
[0031] Deep machine learning differs from traditional machine learning since it uses multiple stages of data processing through many hidden layers of a neural network to learn and interpret the features and the relationships between the features. Deep machine learning embodies neural networks which differs from the traditional machine learning techniques that do not use neural networks. There are various types of deep machine learning models that generate source code, such as recurrent neural network (RNN) models, convolutional neural network (CNN) models, long short-term memory (LSTM) models, and neural transformers with attention.
[0032] Turning to Fig. 2, there is shown components of an exemplary code translation system 200. In an aspect, the code translation system 200 comprises several code databases 202A-202N ( “202” ) , a retrieval engine 204, an encoder 206, a reward model 208, a selector engine 210, a prompt generator 212 and a large language model 214.
[0033] An input source code snippet 220 is provided for translation into a target programming language. The retrieval engine 204 utilizes the encoder 206 to produce an embedding 216 of the input source code snippet 220 which is used to find the top-k closest embeddings in the code database of the target programming language 222 to the embedding 216 of the input source code snippet 220. The code candidates 218 associated with the k closest embeddings are retrieved 222.
[0034] The reward model 208 generates a reward score for each of the top-k code candidates 224. The selector engine 210 selects the top-n code candidates having the lowest number of predicted compiler errors, where n is a user-defined value. The prompt generator 212 generates a prompt 230 that includes the input source code snippet 220 and the top-n code candidates 226. The prompt is applied to the large language model 214 which predicts a source code translation in the target programming language 232.
[0035] Fig. 3 illustrates an exemplary configuration of the reward model. In an aspect, the reward model is configured as a decoder-only neural transformer model with attention pre-trained on source code written in one or more programming languages and optionally natural language text. The model is then trained via reinforcement learning (RL) on the paired training samples.
[0036] Pre-training is where the model is trained on unsupervised data to establish a broad and upper-level understanding of source code and / or natural language text. Pre-training is the process where the model’s parameters (e.g., embeddings, weights, biases) are learned from unsupervised data. The model learns the parameters through the optimization of the cost function used by the neural network layer of the model. The cost function determines the error loss from the previous epoch which is then backpropagated to the preceding layers of the model. The model’s parameters are updated through backpropagation based on the error loss determined by the cost function.
[0037] The decoder neural transformer model 300 includes an input layer 302, one or more decoder blocks 304A-304B ( “304” ) , and an output layer 306A or 306B. The decoder-only neural transformer model 300 is an auto-regressive model that produces an output one token at a time based on an input sequence initially. Thereafter, the input sequence contains embeddings of the outputs of the previous time steps. The input layer 302 to the first decoder block 304A includes an input embedding 308 containing embeddings of the input sequence and a positional embedding layer 310 which forms a context tensor 312. The positional embeddings 310 are used to retain the order of the tokens in the input sequence. The context tensor 312 contains the positional embeddings added to the input embedding 308.
[0038] During pre-training 334, the initial input to the first decoder block contains a sequence of embeddings that represent the tokens of a sample of the pre-training dataset. At each subsequent time step, the input is a shifted sequence of the output embeddings from the previous time step to which the positional embeddings are added.
[0039] During the reinforcement learning 334, the initial input to the first decoder block is a sequence of embeddings that represent the tokens of a paired input sequence. A first input sequence of the pair contains the tokens of the input source code, candidate A, and error count and the second input sequence of the pair contains the tokens of the input source code, candidate B, and an error count.
[0040] During inference 334, the initial input to the first decoder block is a sequence of embeddings that represent the tokens of a prompt that contains the input source code snippet and the top-n code candidates. At each subsequent time step, the input is a shifted sequence of the output embeddings from the previous time step to which the positional embeddings are added.
[0041] A decoder block 304 consists of two layers. The first layer includes a masked multi-head self-attention component 314 followed by a layer normalization component 316. The input to the masked multi-head self-attention component 314 has a residual connection to layer normalization 316. The output of layer normalization 316 is input into the feed-forward neural network 318 with a residual connection to layer normalization component 320. The output of the feed-forward neural network 318 is input into layer normalization component 320.
[0042] Each token / subtoken flows through all the decoder blocks along its own path. The masked multi-head self-attention component 314 allows the feed-forward neural network 318 to focus on certain features or inputs. The input embedding 308 to the decoder block 304A is added with the positional embeddings 310 forming context tensor 312. The decoder block 304 predicts each token / subtoken ti in the target language one-by-one at each time step conditioned on all previously-generated target tokens / subtokens t1, …ti-1.
[0043] The masked multi-head self-attention component 314 masks the input embeddings from future time steps. The feed-forward neural network 318 processes each input embedding separately. The layer normalization components 316, 320 are used between the layers in order to normalize the inputs across the features.
[0044] During pre-training, the output layer 306A consists of a linear layer 322 and a softmax layer 324. The linear layer 322 projects the vector produced by the stack of decoders 323 into a logits vector. The softmax layer 324 then turns the scores of the logits vector into probabilities for each token in the vocabulary which are positive and normalized.
[0045] During the reinforcement learning, the output layer 306B consists of a different linear layer 328 and softmax layer 330 that produces a scalar value 332 to represent the predicted number of compilation errors for a candidate. The input of the reinforcement learning phase includes an input source code snippet x with candidate code A and its error count and the input source code snippet x with candidate code B and its error count, where candidate code A has less compilation errors than candidate B. The model predicts a reward score for each of the input sequences of the paired training sample. The reward scores for the paired sample are subtracted 334 and the result is passed through a logarithmic function by the loss calculation engine 336 to obtain the final loss value. The final loss value 338 is then backpropagated to through the different layers of the reward model to update the weights at each layer of the model.
[0046] In other aspects, the reward model may be configured as an encoder-decoder neural transformer model with attention having a series of stacked encoder blocks coupled to a series of stacked decoder blocks or as an encoder-only neural transformer model with attention having stacked encoder blocks.
[0047] When configured as an encoder-decoder neural transformer model with attention, there is a different output layer used for pre-training and the reinforcement learning. The encoder-decoder neural transformer model with attention is pre-trained to learn to predict source code given a training dataset of source code snippets. During reinforcement learning, paired training samples are used where one training sample includes an input source code snippet, candidate A, and its error count and the second sample of the pair includes the input source code snippet, candidate B, and its error count, where candidate A has less compilation errors than candidate B. The first encoder block receives the input source code snippet which is encoded into a fixed-size vector. The first decoder block takes the candidates as input and their vectors interact in the cross-attention layer of the decoder. The key and value matrices are computed using the encoded information from the encoder while the query matrix is computed from the output of the previous decoder block. The output layer of the last decoder block outputs a scalar value representing the reward score.
[0048] In the embodiment where the reward model is configured as an encoder-only neural transformer model with attention, a pre-trained encoder is obtained having been trained on source code snippets written in different programming languages to produce embeddings for each token in the pre-training dataset. The pre-trained encoder is then altered with a different output layer for the reinforcement learning. The output layer includes a linear layer and a softmax layer that coverts the last hidden state output from the last encoder block into probabilities for each class. A class represents a number of compilation errors. The training samples include an input source code snippet, a candidate, and a label that represents the class associated with the number of compilation errors for the candidate. At inference, the encoder outputs a probability for each class and the class with the highest probability is the predicted reward score or number of compilation errors for the candidate.
[0049] Methods
[0050] Attention now turns to description of the various exemplary methods that utilize the system and device disclosed herein. Operations for the aspects may be further described with reference to various exemplary methods. It may be appreciated that the representative methods do not necessarily have to be executed in the order presented, or in any particular order, unless otherwise indicated. Moreover, various activities described with respect to the methods can be executed in serial or parallel fashion, or any combination of serial and parallel operations. In one or more aspects, the method illustrates operations for the systems and devices disclosed herein.
[0051] Fig. 4 is an exemplary method of the code translation system 400. Initially, an encoder is obtained that is trained to generate an embedding for each of the candidates in the databases and for an input source code snippet at inference (block 402) .
[0052] In an aspect, the encoder is configured as a bi-encoder that is jointly trained to learn an embedding of a sequence of tokens representing an input source code snippet and an embedding of its associated translation in a different programming language. The bi-encoder jointly embeds the input source code snippet and its corresponding translation so they are mapped to a continuous vector space in which the input source code snippet is close to its corresponding translation. The bi-encoder uses a first encoder to encode the sequence of tokens representing the input source code snippet and a second encoder to encode the corresponding translation. The first and second encoders learn an embedding for each token in isolation and then combines the token embeddings into a combined embedding. (Collectively, block 402) .
[0053] The first and second encoders update the embedding weights by performing a gradient descent on the cosine similarity across all the input source code snippets in the bi-encoder training dataset until convergence is achieved. The result is a joint embedding space with close embeddings for an input source code snippet and its related translation embeddings. (Collectively, block 402) .
[0054] Initially, a database of candidates for each programming language is constructed (block 404) . Source code snippets of various programming languages are extracted from different publicly-available source code repositories. The source code snippets may consist of methods / functions or classes from the different sources. The source code snippets are extracted into chunks where each chunk is composed of a pre-determine number of tokens. An embedding for each chunk is generated by the encoder and used as an index to access the corresponding source code snippet. (Collectively, block 404) .
[0055] Initially, a training dataset for the construction of the reward model is generated (block 406) . The training dataset is used to train a pre-trained model for the prediction of the reward score (block 408) . The reward model and the databases are then deployed in a code translation system (block 410) .
[0056] Turning to Fig. 5, there is shown an exemplary method 500 for generating the training dataset used to train the reward model. The training dataset uses source code snippets which are then translated and compiled. The source code snippets are retrieved from publicly-available repositories (i.e., projects, code bases, etc. ) (block 502) .
[0057] For each source code snippet, an embedding is generated by the encoder (block 504) . The retrieval engine uses the embedding of the input source code snippet to search the database of the target programming language for closely-similar embeddings which correspond to code candidates (block 506) .
[0058] For each of the top-k code candidates (block 508) , the prompt generator generates a prompt to apply to the large language model for each code candidate retrieved from the database (block 510) . The prompt includes instructions for the large language model to perform the translation, the input source code snippet and the code candidate (block 510) .
[0059] The prompt is given to the large language model which returns the translation (block 512) . In an aspect, the large language model may be hosted by a web service and accessed using an Application Programming Interface ( “API” ) that is transmitted to an endpoint of the web service via the Internet. The API contains the prompt. In another aspect, the large language model may reside in the same computing device as the code translation system.
[0060] The code translation is then compiled and the number of compiler errors is tracked (block 514) . The prompt or the input source code, code candidate and compilation error count is then used as a training sample that is part of the training dataset (block 516) .
[0061] Turning to Fig. 6, there is shown an exemplary method 600 for training the reward model through reinforcement learning.
[0062] Reinforcement learning differs from supervised learning and unsupervised learning. In supervised learning, a model learns from a training dataset of labeled examples. Each sample in the training dataset contains a correct action that the model should take. The model learns to generalize its actions in order to act in situations not present in the training dataset. In unsupervised learning, the model learns to find patterns or structure hidden in the training dataset of unlabeled data. By contrast, reinforcement learning maximizes a reward gradually observed on its outputs during its training instead of trying to find hidden patterns and structure in the unlabeled training dataset.
[0063] In reinforcement learning an actor interacts over time with its environment to achieve a goal and learns the actions that produce the most reward by trying them. The actor observes the current state of the environment (e.g., pairwise samples) to decide which action to take (e.g., prediction of next token in a translation) . The environment changes state and produces a reward for that action. The reward indicates whether the action was good or bad. A penalty is imposed when the action is bad. The cycle of observation, action, and reward is repeated until the learning is complete.
[0064] The actor uses a function or policy that maps the inputs into the actions or outputs. The environment uses the reward as feedback on the action. The goal of the training phase is for the reward model to learn the optimal policy. The feed forward neural network includes an activation function, a number of hidden layers, and a number of neurons in each layer. The learning algorithm generates the weights and biases for the nodes in the neural network that produce the optimal action.
[0065] In one aspect, the reward model is trained to predict a scalar value given two pairwise samples y ∈ {y1, y2} indicating human preferences. In an aspect, the loss calculation engine of the reward model maximizes a specific loss instead of optimizing a maximum-likelihood loss function: loss
[0066] where rθ (x, y) is the scalar output of the reward model for prompt x and completion y with parameters θ, yw is the preferred completion out of the pair yw, yl, D is the training dataset, and are the pairwise comparisons.
[0067] Turing now to Fig. 6, there is shown an exemplary method 600 for training the reward model. A training dataset is used to train a pre-trained model to learn to predict the number of compilation errors of a source code snippet (block 602) . The training samples are ranked according to the associated number of compiler errors where the samples having the lowest number of compiler errors is ranked higher than those with a large amount of compiler errors (block 604) .
[0068] The samples of the training data are configured into pairwise samples, where one sample of the pair has a low number of compiler errors, yw, and the other sample of the pair has a higher number of compiler errors, y1. (block 606) .
[0069] The pairwise samples are input into the reward model (block 608) . Each of the training samples of the training dataset is an input sequence that is transformed into a sequence of input embeddings. The input sequence is tokenized and each token in replaced with a respective embedding transforming the input sequence into a sequence of input embeddings. Each token embedding has a corresponding positional embedding. The neural transformer model does not read each token sequentially and as such, has no knowledge of the token’s position in a sequence without additional position information. The positional embedding is used to encode position information about a token’s position in a sequence into the neural transformer model. (Collectively, block 608) .
[0070] Neural transformer models are trained iteratively, making multiple passes over the training dataset before converging to a minimum. An epoch represents the entire training dataset passed forwards and backwards through the neural transformer blocks once. Since the training dataset is very large, it is partitioned into smaller batches. The training is iterative and the entire training dataset is passed through the reward model in multiple iterations. Each training iteration includes forward propagation, loss calculation, backpropagation steps followed by updating the weights. The training dataset is partitioned into batches with each batch of sequences running through the training process. (Collectively, block 608) .
[0071] For each input sequence of each batch in each epoch, the T-ordered sequences of tokens are then mapped into numeric vectors and then into respective token embeddings and positional embeddings. Initial values are generated for the token embedding and positional embeddings of each input sequence which are then used to form a context tensor. Thereafter, the neural transformer model learns the values for each embedding through backpropagation. Upon the completion of the training phase, the embeddings for each token and the positional embeddings are saved into respective matrices for later use. There is a token embedding matrix, We, that contains an embedding vector for each token ti, i=0…V of a particular programming language, and a positional embedding matrix, Wp, that contains an embedding vector Pj, j=0…T, for each position, where V is the size of the vocabulary for a particular programming language and T is the length of the token sequence. (Collectively, block 608) .
[0072] The feed forward neural networks in the decoder blocks are trained iteratively, making multiple passes over the training dataset before converging to a minimum. Each training iteration includes forward propagation, loss calculation, backpropagation steps followed by updating the weights by calculating the weight gradients. The loss function estimates the loss or error which is used to compare how good or bad the predicted results are. Once the loss is calculated, it is propagated backwards to the hidden layer that contributed directly to the output. In backpropagation, the partial derivatives of the loss function with respect to the trainable parameters are determined. The weight gradients are calculated as the difference between the old values and the new values of the weights. The weights are adjusted to make the loss as small as possible using a gradient descent technique. In one aspect, a Stochastic Gradient Descent (SGD) method is the optimization algorithm used to find the values of parameters of the function that minimizes the loss function. A backpropagation through time (BPTT) algorithm may be used to update the weights. (Collectively, block 608) .
[0073] At the completion of each batch, the parameters of the neural transformer model are updated at a preconfigured frequency. The parameters include the weights, biases, token embeddings and the positional embeddings which are stored in a respective embedding matrix. (Collectively, block 608) .
[0074] Fig. 7 illustrates an exemplary method of the code translation system for generating a code translation 700. The code translation system receives an input source code snippet in a first or source programming language that needs to be translated to a second programming language that differs from the first programming language (block 702) .
[0075] An embedding of the input source code snippet is generated and used to search for similar embeddings in the database of the target programming language. The code snippets in the database of the target programming language having an embedding similar to the embedding of the input source code snippet are extracted as the top-k code candidates (block 702) .
[0076] In an aspect, the retrieval engine may utilize an approximate k-nearest neighbor (ANN) method to perform the search. The ANN method returns the k-nearest embeddings rather searching for exact matches which is computationally expensive. The ANN method uses a distance measure to find the k-nearest embeddings such as Euclidean, Hamming, or Manhattan distance. (Collectively, block 704) .
[0077] Each of the top-k code candidates is input into the reward model to generate a reward score for each of the top-k code candidate (block 706) . The n code candidates having the highest reward scores or lowest errors are considered the top-n code candidates that are included in the prompt (block 708) .
[0078] The prompt generator forms a prompt to include a set of instructions, the input source code snippet and the top-n code candidates (block 710) . The prompt is transmitted to the large language model (block 712) which returns a response that includes the translated code (block 714) . The translated code is then output in the code translation system (block 716) .
[0079] Fig. 8 illustrates an exemplary inference system 800 that utilizes the code translation system. In an aspect, the inference system is a software development environment such as a source code editor or integrated development environment (IDE) 800. The IDE 800 contains a source code editor 802 having a user interface 804, a code translation system 806 and a large language model 808. The code translation system 806 may be a plug-in, add-in, or component that incorporates the code translation feature into the IDE 800. The code translation system 806 contains the code databases, the encoder, retrieval engine, reward model, selector engine and prompt generator shown in Fig. 2. In one aspect, the large language model 808 may be hosted on a web server that is external to IDE and in another aspect, the large language model 808 resides in the same computing device as the IDE.
[0080] A user using the IDE 800 may utilize the code translation system 806 to translate a source code snippet written in one programming language into a different programming language. The code translation system 806 receives the input source code snippet 810 from the user interface 804 and generates a prompt 814 to the large language model 808. The large language model 808 generates a response having the translation 812 which is then returned to the user interface 804.
[0081] It should be noted that the techniques described herein are not limited to a software development environment and may be used in any chatbot environment which requires that the conversational agent has the ability to translate a source code snippet. This includes not just a software development environment but any relevant task that requires a code translation.
[0082] Exemplary Operating Environment
[0083] Attention now turns to a discussion of an exemplary operating environment. Fig. 9 illustrates an exemplary operating environment 900 in which one or more computing devices 902 is communicatively coupled through a network 940 to one or more computing devices 942 hosting the large language model.
[0084] A computing device 902, 942 may be any type of electronic device, such as, without limitation, a mobile device, a personal digital assistant, a mobile computing device, a smart phone, a cellular telephone, a handheld computer, a server, a server array or server farm, a web server, a network server, a blade server, an Internet server, a work station, a mini-computer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a distributed computing system, multiprocessor systems, or combination thereof. The operating environment 900 may be configured in a network environment, a distributed environment, a multi-processor environment, or a stand-alone computing device having access to remote or local storage devices.
[0085] The computing device 902, 942 may include one or more processors 904, 944, one or more communication interfaces 906, 946, one or more storage devices 908, 948, one or more input / output devices 912, 952, and one or more memory devices 910, 950. A processor 904, 944 may be any commercially available or customized processor and may include dual microprocessors and multi-processor architectures. A communication interface 906, 946, facilitates wired or wireless communications between the computing device 902, 942 and other devices. A storage device 908, 948 may be a computer-readable medium that does not contain propagating signals, such as modulated data signals transmitted through a carrier wave. Examples of a storage device 908, 948 include without limitation RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) , or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, all of which do not contain propagating signals, such as modulated data signals transmitted through a carrier wave. There may be multiple storage devices 908, 948 in a computing device 902, 942. The input / output devices 912, 952 may include a keyboard, mouse, pen, voice input device, touch input device, display, speakers, printers, etc., and any combination thereof.
[0086] A memory device or memory 910, 950 may be any non-transitory computer-readable storage media that may store executable procedures, applications, and data. The computer-readable storage media does not pertain to propagated signals, such as modulated data signals transmitted through a carrier wave. It may be any type of non-transitory memory device (e.g., random access memory, read-only memory, etc. ) , magnetic storage, volatile storage, non-volatile storage, optical storage, DVD, CD, floppy disk drive, etc. that does not pertain to propagated signals, such as modulated data signals transmitted through a carrier wave. A memory device 910, 950 may also include one or more external storage devices or remotely located storage devices that do not pertain to propagated signals, such as modulated data signals transmitted through a carrier wave.
[0087] The memory device 910, 950 may contain instructions, components, and data. A component is a software program that performs a specific function and is otherwise known as a module, program, component, and / or application. Memory device 910 may include an operating system 914, code databases 916, retrieval engine 918, an encoder, a prompt generator 920, a compiler 922, a pairwise sampler 924, a reinforcement learning engine 926, a reward model 928, a selector engine 930, a software development environment 932, and other applications and data 934. Memory device 950 may include an operating system 954, a large language model 956, and other applications and data 958.
[0088] A computing device 902, 942 may be communicatively coupled via a network 940. The network 940 may be configured as an ad hoc network, an intranet, an extranet, a virtual private network (VPN) , a local area network (LAN) , a wireless LAN (WLAN) , a wide area network (WAN) , a wireless WAN (WWAN) , a metropolitan network (MAN) , the Internet, a portions of the Public Switched Telephone Network (PSTN) , plain old telephone service (POTS) network, a wireless network, a network, or any other type of network or combination of networks.
[0089] The network 940 may employ a variety of wired and / or wireless communication protocols and / or technologies. Various generations of different communication protocols and / or technologies that may be employed by a network may include, without limitation, Global System for Mobile Communication (GSM) , General Packet Radio Services (GPRS) , Enhanced Data GSM Environment (EDGE) , Code Division Multiple Access (CDMA) , Wideband Code Division Multiple Access (W-CDMA) , Code Division Multiple Access 2000, (CDMA-2000) , High Speed Downlink Packet Access (HSDPA) , Long Term Evolution (LTE) , Universal Mobile Telecommunications System (UMTS) , Evolution-Data Optimized (Ev-DO) , Worldwide Interoperability for Microwave Access (WiMax) , Time Division Multiple Access (TDMA) , Orthogonal Frequency Division Multiplexing (OFDM) , Ultra Wide Band (UWB) , Wireless Application Protocol (WAP) , User Datagram Protocol (UDP) , Transmission Control Protocol / Internet Protocol (TCP / IP) , any portion of the Open Systems Interconnection (OSI) model protocols, Session Initiated Protocol / Real-Time Transport Protocol (SIP / RTP) , Short Message Service (SMS) , Multimedia Messaging Service (MMS) , or any other communication protocols and / or technologies.
[0090] Technical Improvements
[0091] Aspects of the subject matter disclosed herein pertain to the technical problem of automatically translating a source code snippet in a first programming language into a functionally-equivalent translation in a different programming language. The technical features associated with addressing this problem includes incorporating code candidate translations of the first source code snippet in the target programming language into a prompt to the large language model in order to guide the model on the translation task. The candidate translations are selected from a database of source code snippets of the target programming language through an embedding search and then filtered based the compilability of the translated code. This selection is a two-stage ranking process where in the first stage, code candidates are identified using an embedding search and then ranked according to embedding similarity. The second stage ranks the top-k code candidates based on a reward score and the code candidates with the lowest compilation errors are selected.
[0092] The technical effect achieved is an increased accuracy of the translation without the computational burden on the computing device used to generate the translation. The augmentation of the compiled code candidates in the prompt reduces the frequency of invoking the large language model for a translation thereby leading to cost savings in the usage of the computing device hosting the large language model. Often, the web services that host a large language model charge for usage of the APIs that access a large language model.
[0093] In addition, there is a cost savings in the computing device generating the prompt since it reduces the number of prompts needed to obtain an accurate translation. Additionally, the prediction of the translation eliminates the need for building unit tests. In the absence of the technique disclosed herein, improving the performance of closed-source models, such as GPT-4, whose parameters cannot be adjusted, relies solely on the continuous trials of effective prompts. The techniques described herein are a significant improvement since the model is capable of providing the optimal candidate code for different inputs, thereby assisting models such as GPT-4 in generating high-quality code results. This approach significantly reduces the number of attempts needed to construct suitable prompts thereby enhancing the efficiency and effectiveness of pre-trained models like GPT-4.
[0094] The code completion system has to perform within tight timing requirements in order to be viable. In the scenario where the large language model resides on an external server that is accessed via a network, the operations used to generate the prompt need to be performed on a computing device. Hence, the operations performed are inherently digital. A human mind cannot interface directly with a CPU, or network interface card, or other processor, or with RAM or digital storage, to read and write the necessary data and perform the necessary operations and processing steps taught herein.
[0095] Embodiments are also presumed to be capable of operating “at scale” , that is capable of handling larger volumes, in production environments or in testing labs for production environments as opposed to being mere thought experiments.
[0096] The technique described herein is a technical improvement over prior solutions that did not utilized compiled translations in the prompt to the large language model or which fine-tuned the large language model for a translation task. Without the compiled translations included in the prompt, a large language model is likely to hallucinate and generate useless output. Fine-tuning a large language model on a translation task is not always possible due the considerable amount of resources needed to construct the fine-tuning data and the cost of fine-tuning a large language model. In some scenarios, it may not be possible to fine-tune a publicly-accessible large language model that has restrictions on its use. The augmentation of the prompt with compiled translations in the manner described herein avoids the costly fine-tuning step and improves the accuracy of the predictions generated by the large language model.
[0097] Conclusion
[0098] A system is disclosed comprising: a processor; and a memory that stores a program configured to be executed by the processor, the program comprising instructions that when executed by the processor perform acts that: obtain an input source code snippet associated with a first programming language; search for a plurality of translations of the input source code snippet in a second programming language, wherein the first programming language and the second programming language differ; select top-k translations from the plurality of translations based on a similarity to the input source code snippet; predict a number of compilation errors for each of the top-k translations; rank the top-k translations by respective compilation error count; select top-n translations from the top-k translations based on a low compilation error count; generate a prompt to a large language model to generate a translation of the input source code snippet into the second programming language, wherein the prompt comprises the input source code and the top-n translations; receive the translation generated by the large language model given the prompt; and output the translation generated by the large language model.
[0099] In an aspect, the program comprises instructions that when executed by the processor perform acts that: access a database comprising a plurality of source code snippets associated with the target programming language, wherein each of the plurality of source code snippets is associated with an embedding; and search the database associated with the target programming language for the plurality of translations based on a similar embedding to an embedding of the input source code snippet.
[0100] In an aspect, the program comprises instructions that when executed by the processor perform acts that: generate the embedding of the input source code snippet and embeddings of the plurality of translations from a bi-encoder trained to align multi-lingual source code samples.
[0101] In an aspect, the program comprises instructions that when executed by the processor perform acts that: predict the number of compilation errors of a top-k translation from a reward model given the input source code snippet and the top-k translation. In an aspect, the reward model is trained based on paired comparisons of a first sample and a second sample, wherein the first sample comprises a first translation having a first compilation error count, wherein the second sample comprises a second translation having a second compilation error count that is greater than the first compilation error count.
[0102] In an aspect, the reward model is a decoder-only neural transformer model with attention. In an aspect, the reward model is an encoder neural transformer model with attention. In an aspect, the large language model comprises a neural transformer model with attention.
[0103] A computer-implemented method is disclosed, comprising: obtaining an input source code snippet written in a first programming language to translate into a target programming language; obtaining a first plurality of code translations for the input source code snippet written in the target programming language; ranking the first plurality of code translations based on a comparison with the input source code snippet; selecting a second plurality of code translations extracted from the first plurality of code translation based on the ranking; generating a reward score for each of the code translations in the second plurality of code translations; ranking the second plurality of code translations based on the reward score; extracting a select number of the ranked second plurality of code translations based on a high reward score; generating a prompt to a large language model for a translation of the input source code snippet into the target programming language, wherein the prompt includes the input source code snippet and the select number of the ranked second plurality of code translations; and obtaining from the large language model the translation.
[0104] In an aspect, the computer-implemented method, further comprises: outputting the translation in a software development environment.
[0105] In an aspect, the computer-implemented method, further comprises: accessing a database of code snippets in the target programming language, wherein each code snippet in the database is indexed by a respective embedding; generating an embedding of the input source code snippet; searching for embeddings of code snippets matching the embedding of the input source code snippet; and extracting from the database in the target programming language, code snippets having closest matching embeddings to the embedding of the input source code snippet.
[0106] In an aspect, the computer-implemented method, further comprises: generating the embedding of the input source code snippet from a bi-encoder trained on joint embeddings of a source code snippet in a source programming language with a corresponding code snippet in the target programming language.
[0107] In an aspect, the computer-implemented method, further comprises: predicting a reward score for a code translation in the second plurality of code translations from a reward model given the code translation of the second plurality of code translations.
[0108] In an aspect, the reward model is a neural transformer model with attention trained via reinforcement learning on paired comparisons. In an aspect, the large language model is a conversational pre-trained generative neural transformer model with attention trained to generate source code.
[0109] A hardware storage device is disclosed having stored thereon computer executable instructions that are structured to be executed by a processor of a computing device to thereby cause the computing device to perform actions that: acquire an input source code snippet associated with a first programming language to translate into a target programming language; obtain a plurality of code translations for the input source code snippet in the target programming language; predict a reward score for each code translation, wherein the reward score is based on a compilation error count; select ones of the plurality of code translations having a high reward score; generate a prompt to a large language model for a translation of the input source code snippet into the target programming language, wherein the prompt comprises the input source code snippet and the select ones of the plurality of code translations; obtain the translation of the input source code snippet generated by the large language model; and output the translation in a user interface.
[0110] In an aspect, the hardware device includes further computer executable instructions that are structured to be executed by the processor of the computing device to thereby cause the computing device to perform actions that: extract the plurality of code translations from a database of source code snippets associated with the target programming language, wherein each of the plurality of code translations having an embedding that is similar to an embedding of the input source code snippet.
[0111] In an aspect, the hardware device includes further computer executable instructions that are structured to be executed by a processor of a computing device to thereby cause the computing device to perform actions that: select ones of the extracted plurality of code translations from the database of source code snippets associated with the target programming language having a closest embedding to the input source code snippet.
[0112] In an aspect, the hardware device includes further computer executable instructions that are structured to be executed by the processor of the computing device to thereby cause the computing device to perform actions that: access a reward model given a code translation to predict the reward score, wherein the reward model is a neural transformer model with attention.
[0113] In an aspect, the large language model is a conversational pre-trained generative neural transformer model with attention trained to generate source code on a plurality of programming languages.
[0114] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
[0115] It may be appreciated that the representative methods do not necessarily have to be executed in the order presented, or in any particular order, unless otherwise indicated. Moreover, various activities described with respect to the methods can be executed in serial or parallel fashion, or any combination of serial and parallel operations. In one or more aspects, the method illustrates operations for the systems and devices disclosed herein.
Claims
1.A system comprising:a processor; anda memory that stores a program configured to be executed by the processor, the program comprising instructions that when executed by the processor perform acts that:obtain an input source code snippet associated with a first programming language;search for a plurality of translations of the input source code snippet in a second programming language, wherein the first programming language and the second programming language differ;select top-k translations from the plurality of translations based on a similarity to the input source code snippet;predict a number of compilation errors for each of the top-k translations;rank the top-k translations by respective compilation error count;select top-n translations from the top-k translations based on a low compilation error count;generate a prompt to a large language model to generate a translation of the input source code snippet into the second programming language, wherein the prompt comprises the input source code and the top-n translations;receive the translation generated by the large language model given the prompt; andoutput the translation generated by the large language model.2.The system of claim 1, wherein the program comprises instructions that when executed by the processor perform acts that:access a database comprising a plurality of source code snippets associated with the target programming language, wherein each of the plurality of source code snippets is associated with an embedding; andsearch the database associated with the target programming language for the plurality of translations based on a similar embedding to an embedding of the input source code snippet.3.The system of claim 1, wherein the program comprises instructions that when executed by the processor perform acts that:generate the embedding of the input source code snippet and embeddings of the plurality of translations from a bi-encoder trained to align multi-lingual source code samples.4.The system of claim 1, wherein the program comprises instructions that when executed by the processor perform acts that:predict the number of compilation errors of a top-k translation from a reward model given the input source code snippet and the top-k translation.5.The system of claim 4, wherein the reward model is trained based on paired comparisons of a first sample and a second sample, wherein the first sample comprises a first translation having a first compilation error count, wherein the second sample comprises a second translation having a second compilation error count that is greater than the first compilation error count.6.The system of claim 4, wherein the reward model is a decoder-only neural transformer model with attention.7.The system of claim 4, wherein the reward model is an encoder neural transformer model with attention.8.The system of claim 1, wherein the large language model comprises a neural transformer model with attention.9.A computer-implemented method, comprising:obtaining an input source code snippet written in a first programming language to translate into a target programming language;obtaining a first plurality of code translations for the input source code snippet written in the target programming language;ranking the first plurality of code translations based on a comparison with the input source code snippet;selecting a second plurality of code translations extracted from the first plurality of code translation based on the ranking;generating a reward score for each of the code translations in the second plurality of code translations;ranking the second plurality of code translations based on the reward score;extracting a select number of the ranked second plurality of code translations based on a high reward score;generating a prompt to a large language model for a translation of the input source code snippet into the target programming language, wherein the prompt includes the input source code snippet and the select number of the ranked second plurality of code translations; andobtaining from the large language model the translation.10.The computer-implemented method of claim 9, further comprising:outputting the translation in a software development environment.11.The computer-implemented method of claim 9, wherein obtaining a first plurality of code translations for the input source code snippet written in the target programming language further comprises:accessing a database of code snippets in the target programming language, wherein each code snippet in the database is indexed by a respective embedding;generating an embedding of the input source code snippet;searching for embeddings of code snippets matching the embedding of the input source code snippet; andextracting from the database in the target programming language, code snippets having closest matching embeddings to the embedding of the input source code snippet.12.The computer-implemented method of claim 11, further comprising:generating the embedding of the input source code snippet from a bi-encoder trained on joint embeddings of a source code snippet in a source programming language with a corresponding code snippet in the target programming language.13.The computer-implemented method of claim 9, wherein generating a reward score for each of the code translations in the second plurality of code translations further comprises:predicting a reward score for a code translation in the second plurality of code translations from a reward model given the code translation of the second plurality of code translations.14.The computer-implemented method of claim 9, wherein the reward model is a neural transformer model with attention trained via reinforcement learning on paired comparisons.15.The computer-implemented method of claim 9, wherein the large language model is a conversational pre-trained generative neural transformer model with attention trained to generate source code.16.A hardware storage device having stored thereon computer executable instructions that are structured to be executed by a processor of a computing device to thereby cause the computing device to perform actions that:acquire an input source code snippet associated with a first programming language to translate into a target programming language;obtain a plurality of code translations for the input source code snippet in the target programming language;predict a reward score for each code translation, wherein the reward score is based on a compilation error count;select ones of the plurality of code translations having a high reward score;generate a prompt to a large language model for a translation of the input source code snippet into the target programming language, wherein the prompt comprises the input source code snippet and the select ones of the plurality of code translations;obtain the translation of the input source code snippet generated by the large language model; andoutput the translation in a user interface.17.The hardware device of claim 16, wherein obtain the plurality of code translations for the input source code snippet in the target programming language includes further computer executable instructions that are structured to be executed by the processor of the computing device to thereby cause the computing device to perform actions that:extract the plurality of code translations from a database of source code snippets associated with the target programming language, wherein each of the plurality of code translations having an embedding that is similar to an embedding of the input source code snippet.18.The hardware device of claim 17, having stored thereon computer executable instructions that are structured to be executed by a processor of a computing device to thereby cause the computing device to perform actions that:select ones of the extracted plurality of code translations from the database of source code snippets associated with the target programming language having a closest embedding to the input source code snippet.19.The hardware device of claim 16, wherein predict the reward score for each code translation includes further computer executable instructions that are structured to be executed by the processor of the computing device to thereby cause the computing device to perform actions that:access a reward model given a code translation to predict the reward score, wherein the reward model is a neural transformer model with attention.20.The hardware device of claim 19, wherein the large language model is a conversational pre-trained generative neural transformer model with attention trained to generate source code on a plurality of programming languages.
Citation Information
Patent Citations
Semi-supervised translation of source code programs using neural transformers
US20220308848A1
Code generation through reinforcement learning using code-quality rewards
US20230195428A1
Cited By
Multi-source heterogeneous code conversion method and device for cross-language development
CN121143797A