System and Method for Program Synthesis

The reinforcement learning-based framework addresses the limitations of existing program synthesis models by using an actor-critic approach to fine-tune pre-trained language models with unit tests, enhancing code generation accuracy and precision for complex tasks.

JP2025520071AActive Publication Date: 2025-07-01SALESFORCE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024569393
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-26
Filing Date
2023-05-19
Publication Date
2025-07-01
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

Existing language models for program synthesis fail to effectively utilize unit tests and other meaningful signals during the optimization process, leading to poor performance in generating accurate and functional code, especially for complex coding tasks.

Method used

A reinforcement learning-based framework that fine-tunes a pre-trained language model using an actor-critic approach, incorporating unit tests to evaluate the functional correctness of generated code and adjust the model's parameters to improve accuracy.

Benefits of technology

The framework significantly enhances the accuracy and precision of code generation by leveraging unit tests to refine and repair programs, resulting in improved performance on complex programming tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520071000030
    Figure 2025520071000030
  • Figure 2025520071000031
    Figure 2025520071000031
  • Figure 2025520071000032
    Figure 2025520071000032
Patent Text Reader

Abstract

The embodiments described in this specification provide a reinforcement learning-based framework using a pre-trained language model (LM) for program synthesis tasks. Specifically, the framework adopts a training strategy that optimizes the pre-trained LM for program synthesis tasks in the actor-critic approach.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - References] This international application claims priority based on U.S. Non - Provisional Application Nos. 17 / 896,942 and 17 / 896,946, which were filed on August 26, 2022, are co - pending, and are by the same applicant, and each of these applications is a non - provisional application of U.S. Provisional Application No. 63 / 344,900 filed on May 23, 2022, and claims the priority thereof based on 35 U.S.C. § 119.

[0002] All of the above applications are hereby expressly incorporated by reference in their entirety.

[0003] [Technical Field] Embodiments generally relate to machine learning systems, and more specifically, to systems and methods for program synthesis through pre - trained models and deep reinforcement learning.

Background Art

[0004] Program synthesis, also commonly referred to as code generation, is the task of generating a computer code program that meets a problem specification, such as sorting a list, merging two data tables, and / or the like. When program synthesis is treated as a sequence - to - sequence task, there are pre - trained language models that can be adapted to receive an input sequence as a natural - language problem specification and then generate a sequence of code as an output program. However, these existing language models may have limited code - generation performance because these models often train program - synthesis models only from natural - language problem descriptions and ground - truth programs according to standard supervised fine - tuning procedures. Such a paradigm largely ignores some important but potentially useful signals in problem specifications such as unit tests, resulting in poor performance when solving complex unseen coding tasks.

[0005] Therefore, an efficient and accurate program synthesis model is needed.

Brief Description of Drawings

[0006]

Figure 1

[0007]

Figure 2

[0008]

Figure 3

[0009]

Figure 4

[0010]

Figure 5

[0011]

Figure 6

[0012]

Figure 7

[0013]

Figure 8

[0014]

Figure 9

Figure 9B

[0015]

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

[0016]

Figure 20A

Figure 20B

Figure 20C

DETAILED DESCRIPTION OF THE INVENTION

[0017] In the figure, elements having the same name have the same or similar functions.

[0018] As used herein, the term "network" can include any hardware or software-based framework that includes any artificial intelligence network or system, neural network, or system, and / or any training or learning model implemented thereon or together therewith.

[0019] As used herein, the term "module" can include a hardware or software-based framework that performs one or more functions. In some embodiments, a module can be implemented on one or more neural networks.

[0020] Existing language models that can be used for program synthesis are often trained using the conventional next-token prediction (NTP) objective of maximizing the likelihood of the following ground-truth tokens. Training the model with only the next-token prediction objective in the "teacher forcing" manner often leads to cumulative errors during test time when tokens are generated by conditioning on previously sampled tokens rather than ground-truth tokens. This problem is more severe in the field of program synthesis where existing token matching scores such as BLEU may fail to measure the functional correctness of a complete program.

[0021] Furthermore, existing language models may fail to utilize potentially meaningful signals from unit tests that directly determine the model's performance by the functional correctness of the program. Current approaches ignore this important signal during model optimization as well as the generation procedure.

[0022] In view of the problems in existing program synthesis models, the embodiments described herein provide a reinforcement learning-based framework that uses a pre-trained language model (LM) for program synthesis tasks. Specifically, a pre-trained LM (e.g., pre-trained on publicly available code data, etc.) can be fine-tuned for the program synthesis task for pairs of natural language problem descriptions and corresponding solution programs. The fine-tuned LM may then function as an actor network, which synthetically samples sequences generated from this actor and forms sampled programs in response to inputs of the same problem description that include both correct and incorrect programs. These program samples are passed to a critic model trained as an error predictor to predict the test outcome of the input program given unit tests and to determine a return that evaluates the functional correctness of these program samples. The return generated from the critic model is then used to compute a policy gradient to minimize the expected return (negated). Next, the actor network is fine-tuned based on the policy gradient.

[0023] In this way, the pre-trained LM is fine-tuned for the program synthesis task in a reinforcement learning manner. For example, the pre-trained parameters of the LM may function as the probabilistic policy of the actor network, and accordingly, actions may be generated as predictions for each token of the output program. The pre-trained LM (actor network) receives a return measured by the functional correctness of the generated program, and the goal of reinforcement learning is to minimize the expected return.

[0024] In one embodiment, during inference, the fine-tuned LM via the RL framework can be used to generate one or more code programs in response to a natural language problem description. To improve the accuracy and precision of the resulting programs, programming improvement procedures and / or program repair procedures can be optionally used to refine and / or repair the generated programs based on the functional accuracy of the programs generated during test time. Specifically, exemplary unit tests and critic models are employed to filter and select "pass" programs (those that pass the unit tests) and "fail" programs (those that fail the unit tests) from the LM-generated programs, respectively.

[0025] The "pass" programs can then be used to improve program generation. That is, subsequences from the "pass" programs are used as "seeds" to initialize and condition the LM model to resample new tokens and obtain a new output program, for example, by generating subsequent tokens following the "seed" to form the output program. In this way, the regenerated programs already conditioned on the "pass" subsequences can result in a high likelihood of passing the unit tests.

[0026] The "fail" programs can be used to repair program generation. Among the "fail" programs, programs with a relatively high likelihood of passing the unit tests (compared to other "fail" programs) can be selected. These selected program candidates are concatenated with their respective error information (e.g., whether the "fail" program failed to compile, execute, or generate the correct test results, or whether a specific error such as a syntax error occurred). A program repair module can receive the concatenated input and generate an output code program. In this way, the regenerated (repaired) programs are generated based on the possible previous error information and thus can have a higher likelihood of being "repaired" and passing the unit tests.

[0027] FIG. 1 is a simplified block diagram showing an exemplary architecture 100 that employs an actor-critic framework 150 to optimize a pre-trained LM 110 (and a fine-tuned LM 120) for program synthesis according to embodiments described herein. In one embodiment, architecture 100 includes a pre-training / fine-tuning section 145 of one or more LMs 110 and an actor-critic framework 150. Specifically, the program synthesis task can be formulated as a reinforcement learning (RL) problem such that the actor-critic framework 150 can apply an actor-critic reinforcement learning (RL) approach to improve the performance of the fine-tuned pre-trained LM 120 for program synthesis from the pre-training / fine-tuning stage 145.

[0028] In one embodiment, at stage 145, the LM 110 can first be pre-trained on publicly available code data 102 (e.g., from Github). For example, the LM 110 can include a Transformer model as the backbone of the program synthesis system described herein. An example of such a pre-trained LM 110 can be a multilingual code-aware language model pre-trained on a large-scale source code corpus curated from Github, such as CodeT5 described in U.S. Non-Provisional Application No. 17 / 450,968, filed Aug. 27, 2021, which is co-pending and co-owned, the entire disclosure of which is hereby expressly incorporated by reference.

[0029] In one embodiment, the public code data 102 may include a Python pre-trained dataset such as the Github code dataset. The public code data 102 may have compiled public non-personal information from GitHub consisting of permissively licensed Python code (e.g., "mit", "apache-2", "bsd-3-clause", "bsd-2-126 clause", "cc0-1.0", "unlicense", "isc"). The resulting Python dataset (GCPY) has 10.5B tokens and is 10 times larger than the CodeSearchNet (CSN) corpus used in the original CodeT5 pre-training.

[0030] In one embodiment, the LM 110 may be pre-trained on a pre-training task similar to that used in CodeT5, such as masked span prediction (MSP). While the MSP task is useful for code understanding, it is significantly different from the purpose of program synthesis. To alleviate this gap, a pre-training task of next token prediction (NTP) may be used when pre-training the LM 110. Specifically, the pivot position is uniformly sampled for each code sample, and then the content preceding the pivot is passed to the encoder of the LM 110, and the rest is passed to the decoder of the LM 110. To control the lengths of the input and output sequences, the pivot may be restricted to within 10% - 90% of the original sequence.

[0031] After pre-training, the pre-trained LM 110 can be fine-tuned for a specific program synthesis task. Following the sequence-to-sequence approach, a program synthesis training pair of a natural language problem description 105 in the form of an input sequence D and a corresponding solution code program 106 may be used to fine-tune the pre-trained LM 110. In response to the input sequence D, the pre-trained LM 110 outputs a sequence of programs that can solve the problem

Number

[0032] Therefore, the model parameters θ of the pre-trained LM 110 are fine-tuned during the training time and can maximize the likelihood of the ground truth reference program. Specifically, if W = (w1,..., w T ) is the ground truth program, the objective is to minimize the cross-entropy loss 108, that is, [Number] in the formula, the conditional probability p θ is parameterized according to the above softmax function. During the inference time, the model can generate a sequence of programs by autoregressively sampling tokens [Number] from the conditional distribution [Number] .

[0033] In one embodiment, the fine-tuned LM 120 is evaluated against unit tests 112 corresponding to the problem description. Each test includes a pair of input and ground truth output. In some exemplary real-world program synthesis tasks, exemplary unit tests are often given as part of the problem specification.

[0034] In one embodiment, the fine-tuned LM 120 is then passed to the actor-critic framework 150 and serves as the actor network 130. Specifically, the learned parameters θ of the fine-tuned LM model 120 can be considered a probabilistic policy that determines actions as predictions of each token in the sampled program 133 in response to the input of the problem description 105. Following each action, the LM model 120 (or equivalently the actor network 130) updates its hidden state representation, which is used by the policy to determine the next action in the next decoding step. The generated tokens of the sampled program 133 can be sent to the critic network 140. At the end of the generation episode (i.e., <endoftext>Once the token is found), the actor network 130 receives the return r measured by the critic network 140 based on the functional accuracy of the generated program 133.

[0035] Specifically, for each token W t s sampled by the actor network 130 at the decoding time step t in each synthetic sample array W s =(W1 s ,...,W t s ), the critic network 140 can determine the return by checking its functional accuracy. On the one hand, the problem description 105 is associated with one or more unit tests 112 that include exemplary test inputs and corresponding outputs for solving the problem description 105. On the other hand, the generated program 133 is also passed to the compiler along with the corresponding unit test 112. The generated program 133 is then compiled, executed using the test inputs from the unit test 112, and generates an execution result. From the output of the execution, depending on whether the synthetic sample sequence W s can be fully compiled and executed, and if the execution is successful, depending on whether the execution result matches the test output of the unit test 112, the return r can be determined, that is,

Number

[0036] The determined reward r can then be used to calculate the reinforcement learning training objective of minimizing the expected return 135, that is,

Number

Number

[0037] In one embodiment, a "baseline" program may be employed in the RL training of the actor network 130. Specifically, a greedy decoding strategy may be used as the baseline, and any generated sample 133 that performs better than this baseline is given a positive return estimate, and a negative return estimate otherwise. This relative normalization technique allows the model to explore incomplete programs as long as their returns are better than those of the baseline. In other words, given the problem description 105, a baseline program sample sequence Wb can be generated using the baseline model. The return r(W b ) can be determined in the same way as r(W s ), and the expected gradient estimate can be calculated to reflect whether the sampled program sequence performs better than the baseline program sequence by including each reward.

Number

[0038] FIG. 2 is a simplified block diagram showing an exemplary program synthesis task according to one embodiment described herein. As shown in FIG. 2, an example of a program synthesis task includes a problem specification 105 in natural language that describes the problem of "not being a palindrome" and "printing the maximum length of a substring". The corresponding solution program 106 includes code segments that solve the problem described in the specification 105. The unit test 112 can include exemplary input-output test pairs corresponding to the problem specification 105. For example, when the input = "wuffuw", the output (e.g., the maximum length of the non-palindromic substring) = 5.

[0039] The pre-trained language model (LM 120) can be adapted to receive an input sequence as the problem specification 105 in natural language and generate a sequence of code as the output program. When the problem specification 105 is passed to a code generator (such as a pre-trained and fine-tuned LM 130), the expected output is a program that is checked for functional correctness against the unit test 112.

[0040] Figure 3 is a simplified block diagram showing an example of a reinforcement learning-based program synthesis framework 300 for an exemplary program synthesis task according to an embodiment described herein. The RL-based program synthesis framework 300 represents the dynamics within the actor-critic framework 145 of FIG. 1 in an RL manner. Specifically, in the RL network 300, the fine-tuned LM can function as an actor that, according to the learned parameters of the fine-tuned LM model θ, i.e., the policy, responds to the input of the problem specification 105 and determines an action 216 as the prediction of each token in the output program sequence. Then, the action 216 may be sent to a critic network 140 that functions as a value function for calculating a value 217 (e.g., a policy gradient) based on the current action 216 and updating the actor 130. The compiler 204 may serve as the environment for the actor 130 and the critic 140, receive the action 216, and generate a reward 213, for example, by compiling, executing a sampled program sequence consisting of predicted tokens (action 216) from the actor 130, and comparing the execution result with the unit test 112. The execution state 214 of the environment (compiler 204) may be shared with the actor 130 and the critic 140 such that each makes its respective prediction.

[0041] Figure 4 is a simplified block diagram showing an exemplary training procedure of the critic network 140 of FIG. 1 according to an embodiment described herein. The critic model 140 includes a sequence-to-sequence model 402, linear and softmax operators 404, a max pooling module 406, and a return estimation module 408.

[0042] In one embodiment, the critic model 140 receives a problem description D 105 and a sampled program W from the actor network 130 of FIG. 1 s =(W1 s ,...,W t s It is parameterized as a neural network with a parameter Φ that receives an input as []. The critic model may be trained as an error predictor that receives the problem specification 105 and the program 133 or 134 as input sequences and then predicts one of four possible test outcomes {CompileError; RuntimeError; FailedTest; PassedTest} as described in relation to the reward definition.

[0043] For example, the critic model 140 may include a Transformer model of a size smaller than the actor model 130 as the base architecture, i.e., the sequence-to-sequence model 402. The context hidden states {h1, …, h T} of the program tokens obtained from the critic model decoder are passed to the linear layer 404 and then max-pooled along the sequence length dimension via the max-pooling layer 206. h pool =Pooling(Linear (h1),…,Linear(h T ))(6) Next, the prediction of the critic for the unit test outcome is

Number

[0044] In this way, the training objective 409 of the parameters Φ of the critic model 130 can be calculated as the cross-entropy loss between the predicted unit test outcome from the max-pooling layer 406 of the critic model 130 and the ground truth unit test outcome 413. L critic (Φ)=-logpΦ(u│W s ,D) (8) Here, u is the sampled program sequence W s The ground truth unit test outcome 413 given by the compiler is shown after passing 133 to the unit test 112 corresponding to the problem. The calculated training objective L critic (Φ) is then used to update the critic model 140 (e.g., the max pooling layer 406, the linear and softmax operators 404, and the sequence-to-sequence model 402) via backpropagation.

[0045] After training the critic model 140, regarding the ground truth unit test output

Number

Number

Number

Number

Number

Number

[0046] In some embodiments, to improve and stabilize the training process, the baseline program 134 is considered, for example, by passing it to the unit test 112 to generate the baseline test result 414. In this way, by comparing the sample test result 413 and the baseline test result 414, a relative return is generated. Specifically, the return estimation module 408 can then calculate the policy gradient based on the intermediate return (in the baseline test result 414 generated by passing the baseline program sequence 134 to the unit test 112).

Number

[0047] In one embodiment, imitation learning can be employed to first warm-start the pre-trained LM model 110 with Lce for up to a maximum of 10 epochs. Then, the sampled program sequence is obtained from this actor network 130, and the critic model 140 is trained while keeping the parameters of the actor network 130 frozen. For example, if the actor network is the CodeT5 actor model, the CodeT5-small architecture is used for the critic model 140, and if the actor model is a GPT variant, the GPT2-small critic architecture can be used for the critic model 140.

[0048] In one embodiment, in addition to the synthesis program 133, the ground truth program 106 of the training samples can also be used to train the critic network 140. These samples are considered to be complete programs and always have the label of PassedTest. After training the critic, both Lce and Lrl are applied with equal weights to fine-tune the actor network 130. To optimize the LM actor network 130, in each training optimization step, the expected gradient is a single sample W s ~p θ can be approximated.

Number

[0049] FIG. 5 is a simplified block diagram showing a critic sampling (CS) framework 500 for program synthesis using the learned LM (actor network 130) from FIG. 1 during inference, according to an embodiment described herein. The LM 130 fine-tuned for the program synthesis task from FIG. 1 can be used to generate, improve, and repair programs based on the results for the exemplary unit tests of the corresponding problems. Specifically, a dual strategy (referred to as "critic sampling" (CS)) including a program repair procedure 560 and a programming improvement procedure 550 is implemented to generate and improve programs during inference from both successful cases (program improvements) and failed cases (program repairs). The program improvement procedure 550 and / or the program repair procedure 560 may be optionally implemented (as indicated by the dotted line in FIG. 5) after the LM 130 generates the output program 533, or may be implemented together.

[0050] In one embodiment, the test problem description 505 may be received by the fine-tuned LM 130 at the inference stage. The exemplary unit test input-output provided in the input problem description 505 may be used to improve the generation procedure during inference. For example, the exemplary input-output pair may be extracted from the problem description 505 and may form an exemplary unit test 112.

[0051] For each problem description 505, the fine-tuned LM 130 may generate N programs 533. Each of the generated programs 533 may then be passed to an exemplary unit test that is often embedded as part of the problem specification 505. Specifically, the generated programs 533 may be filtered by the exemplary unit test results in the filtering module 535, and the filtering module 535 selects the programs that passed the exemplary test as set

Number

Number

[0052] The generated programs 533 can go through a program improvement procedure 550 to generate the final program 555. Specifically, the qualified set

Number

[0053] In one implementation, subsequences from these program samples from the pass set

Number

Number

Number

[0054] Therefore, the subsequence 543 is <endoftext>It is used as a seed 545 to initialize and condition the (actor) LM 130 to resample new tokens up to the token. In this round, each seed sequence can be stacked N / |P| times for upsampling. This results in the same number of output programs N as in the initial generation. Finally, the N improved programs 555 generated can be evaluated against the hidden unit test 536.

[0055] In some situations, generating a program to solve a problem, especially a programming problem at the competition level, requires a huge search space of possible programs. In very many cases, a complete failure occurs where all programs fail the exemplary test, that is,

Number

[0056] In one embodiment, the same critic model (φ test ) used in the program improvement procedure 550 is used to sample top candidates from the failure set

Number

Number

[0057] In one embodiment, this program repair model 566 is designed as a Sequence-to-sequence generation model. The input sequence is the concatenation of the problem description D 505 and the buggy program W fail The additional signal received from the unit test results 112 includes the type of test outcome, for example, one of CompileError, RuntimeError, FailedTest, PassedTest, and error subtypes (e.g., syntax error, out-of-index error, and / or the like) may also be included in the input sequence. The error type is extracted from the error trace returned by the compiler.

[0058] To train the program repair model 566, the synthetic samples 133 that were originally used in RL training to train the actor-critic network 150 are used as the buggy program W fail =W s The ground truth program W 106 can be used as the expected correct program. The training objective of the program repair model is to minimize the cross-entropy loss, that is,

Number

[0059] In one embodiment, the program 533 is generated in mini - batches to improve efficiency during inference and can use kernel sampling with a batch size of N = 200. During program improvement, additional computational costs may occur to resample using the seed sequence 545, but it should be noted that only a partial program needs to be generated in the regeneration stage. In this way, the program improvement stage can be less expensive than conventional program synthesis.

[0060] [Computer Environment] FIG. 6 is a simplified diagram of a computing device 600 for implementing a reinforcement - learning - based program synthesis framework shown in FIGS. 1 - 5 according to some embodiments. As shown in FIG. 6, the computing device 600 includes a processor 610 coupled to a memory 620. The operation of the computing device 600 is controlled by the processor 610. Also, although the computing device 600 is illustrated as having only one processor 610, it is understood that the processor 610 can represent one or more central processing units, multi - core processors, microprocessors, microcontrollers, digital signal processors, field - programmable gate arrays (FPGAs), application - specific integrated circuits (ASICs), graphics processing units (GPUs), etc. within the computing device 600. The computing device 600 may be implemented as a stand - alone subsystem, as a board added to a computing device, and / or as a virtual machine.

[0061] Memory 620 may be used to store software executed by computing device 600 and / or one or more data structures used during the operation of computing device 600. Memory 620 may include one or more types of machine-readable media. Some common forms of machine-readable media include floppy (registered trademark) disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, any other optical media, punch cards, paper tape, any other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip, or cartridge, and / or any other media adapted to be read therefrom by a processor or computer.

[0062] Processor 610 and / or memory 620 may be arranged in any suitable physical arrangement. In some embodiments, processor 610 and / or memory 620 may be implemented on the same substrate, within the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), etc. In some embodiments, processor 610 and / or memory 620 may include distributed, virtualized, and / or containerized computing resources. Consistent with such embodiments, processor 610 and / or memory 620 may be arranged in one or more data centers and / or cloud computing facilities.

[0063] In some examples, memory 620 may include a non-transitory tangible machine-readable medium that includes executable code that, when executed by one or more processors (e.g., processor 610), may cause the one or more processors to execute methods described in more detail herein. For example, as illustrated, memory 620 may include instructions for program synthesis module 630 that may be used to implement and / or emulate systems and models and / or to implement any of the methods further described herein. Program synthesis module 630 may receive input 640 including a natural language problem specification via data interface 615 and generate a code program as output 650.

[0064] In some embodiments, program synthesis model 630 includes an actor network module 631 (similar to 130 in FIG. 1), a critic network module 632 (similar to 140 in FIG. 1), and a language model 633 (similar to 110 or 120 in FIG. 1). Details of program synthesis module 630 and its sub-modules 631-633, as well as their interactions, are described with respect to FIGS. 1-5.

[0065] In one embodiment, program synthesis module 630 and its sub-modules 631-633 may be implemented by hardware, software, and / or combinations thereof.

[0066] Some examples of computing devices, such as computing device 400, may include a non-transitory tangible machine-readable medium that includes executable code that, when executed by one or more processors (e.g., processor 410), may cause the one or more processors to execute a process of a method. Some common forms of machine-readable media that may include a process of a method may be, for example, a floppy (registered trademark) disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, any other optical medium, a punch card, a paper tape, any other physical medium having a pattern of holes, a RAM, a PROM, an EPROM, a FLASH-EPROM, any other memory chip or cartridge, and / or any other medium adapted to be read by a processor or a computer.

[0067] FIG. 7 is a simplified block diagram of a networked system suitable for implementing the program synthesis framework described in FIGS. 1-5 and other embodiments described herein. In one embodiment, block diagram 700 includes a user device 710 that can be operated by a user 740, database vendor servers 745, 770, and 780, a server 730, and other forms of devices, servers, and / or software components that operate to perform various methodologies according to the described embodiments. Exemplary devices and servers can include devices, stand-alone, and enterprise-class servers that can be similar to the computing device 100 described in FIG. 1, and operate an OS such as MICROSOFT® OS, UNIX® OS, LINUX® OS, or other suitable device and / or server-based OS. The devices and / or servers shown in FIG. 7 may be deployed and arranged in other ways, and the operations performed and / or services provided by such devices and / or servers may be combined or separated for a given embodiment, and it should be understood that more or fewer devices and / or servers may be involved. One or more devices and / or servers may be operated and / or maintained by the same or different entities.

[0068] The user device 710, database vendor servers 745, 770, and 780, and server 730 may communicate with each other via a network 760. The user device 710 may be utilized by a user 740 (e.g., a driver, system administrator, etc.) to access various features available on the user device 710, which may include processes and / or applications associated with the server 730 for receiving output data anomaly reports.

[0069] The user device 710, the database vendor server 745, and the server 730 may each include one or more processors, memories, and other appropriate components for executing instructions such as program code and / or data stored on one or more computer-readable media to implement the various applications, data, and steps described herein. For example, such instructions may be stored on one or more computer-readable media such as memories or data storage devices internal and / or external to the various components of the system 700 and / or accessible via the network 760.

[0070] The user device 710 may be implemented as a communication device that utilizes appropriate hardware and software configured for wired and / or wireless communication with the database vendor server 745 and / or the server 730. For example, in one embodiment, the user device 710 may be implemented as an autonomous vehicle, a personal computer (PC), a smartphone, a laptop / tablet computer, a wristwatch having appropriate computer hardware resources, glasses having appropriate computer hardware (e.g., GOOGLE GLASS (registered trademark)), other types of wearable computing devices, an embedded communication device, and / or other types of computing devices capable of transmitting and / or receiving data such as an IPAD (registered trademark) made by APPLE (registered trademark). Although only one communication device is shown, multiple communication devices may function similarly.

[0071] The user device 710 of FIG. 7 includes a user interface (UI) application 712 and / or other applications 716, which may correspond to executable processes, procedures, and / or applications having associated hardware. For example, the user device 710 may receive a message indicating a generated program from the server 730 and display the message via the UI application 712. In other embodiments, the user device 710 may include additional or different modules having dedicated hardware and / or software as desired in certain embodiments to provide functionality to the user device 710.

[0072] In various embodiments, the user device 710 includes other applications 716 as may be desired in certain embodiments to provide functionality to the user device 710. For example, the other applications 716 may include a security application for implementing client-side security functions, a programming client application for interfacing with an appropriate application programming interface (API) via the network 760, or other types of applications. The other applications 716 may also include communication applications such as email, text, voice, social networking, and IM applications that enable the user to send and receive emails, phone calls, texts, and other notifications via the network 760. For example, the other applications 716 may be an email or instant messaging application that receives a prediction result message from the server 730. The other applications 716 may include a device interface and other display modules that may receive input and / or output information. For example, the other applications 716 may include a software program for asset management executable by a processor, including a graphical user interface (GUI) configured to provide an interface for the user 740 to view the generated program.

[0073] The user device 710 may further include a database 718 stored in the temporary and / or non-temporary memory of the user device 710. The database 210 stores various applications and data and may be utilized during the execution of various modules of the user device 710. The database 718 may store a user profile regarding the user 740, predictions previously viewed or saved by the user 740, historical data received from the server 730, and / or the like. In some embodiments, the database 718 may be local to the user device 710. However, in other embodiments, the database 718 may be external to the user device 710 and accessible by the user device 710 including a cloud storage system and / or database accessible via the network 760.

[0074] The user device 710 includes at least one network interface component 719 adapted to communicate with a database vendor server 745 and / or the server 730. In various embodiments, the network interface component 719 may include a DSL (e.g., Digital Subscriber Line) modem, a PSTN (Public Switched Telephone Network) modem, an Ethernet® device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices including microwave, radio frequency, infrared, Bluetooth®, and near-field communication devices.

[0075] The database vendor server 745 may correspond to a server that hosts one or more of the databases 703a - n (or collectively referred to as 703) and provides a training data set including public code data to the server 730. The database 703 may be implemented by one or more relational databases, distributed databases, cloud databases, and / or the like.

[0076] The database vendor server 745 includes at least one network interface component 726 adapted to communicate with the user device 710 and / or the server 730. In various embodiments, the network interface component 726 may include DSL (e.g., Digital Subscriber Line) modems, PSTN (Public Switched Telephone Network) modems, Ethernet® devices, broadband devices, satellite devices, and / or various other types of wired and / or wireless network communication devices including microwave, radio frequency, infrared, Bluetooth®, and near-field communication devices. For example, in one embodiment, the database vendor server 745 may transmit asset information from the database 703 to the server 730 via the network interface 726.

[0077] The server 730 may be housed together with the program synthesis module 630 and its sub-modules described in FIG. 1. In some implementations, the module 630 may receive training code data from the database 719 in the database vendor server 745 via the network 760 to generate a program. The generated program may be transmitted to the user device 710 for review by the user 740 via the network 760.

[0078] The database 732 may be stored in the temporary and / or non-temporary memory of the server 730. In one embodiment, the database 732 may store data obtained from the database vendor server 745. In one embodiment, the database 732 can store the parameters of the program synthesis model 630. In one embodiment, the database 732 can store previously generated programs and problem descriptions, as well as the corresponding input feature vectors.

[0079] In some embodiments, database 732 may be local to server 730. However, in other embodiments, database 732 may be external to server 730 and accessible by server 730, including a cloud storage system and / or database accessible via network 760.

[0080] Server 730 includes at least one network interface component 733 adapted to communicate with user device 710 and / or database vendor servers 745, 770, or 780 via network 760. In various embodiments, network interface component 733 may comprise various other types of wired and / or wireless network communication devices including DSL (e.g., Digital Subscriber Line) modems, PSTN (Public Switched Telephone Network) modems, Ethernet® devices, broadband devices, satellite devices, and / or microwave, radio frequency (RF), and infrared (IR) communication devices.

[0081] Network 760 may be implemented as a single network or a combination of multiple networks. For example, in various embodiments, network 760 may include the Internet, one or more intranets, landline communication networks, wireless networks, and / or other suitable types of networks. Thus, network 760 may accommodate small-scale communication networks such as private or local area networks, or larger-scale networks such as wide area networks or the Internet, accessible by various components of system 700.

[0082] [Exemplary Workflow] FIG. 8 is an exemplary logical flow diagram showing a reinforcement learning-based training method for program synthesis based on the actor-critic framework shown in FIG. 1 according to some embodiments described herein. One or more of the processes of method 800 may be implemented in the form of executable code stored on a non-transitory tangible machine-readable medium that, when executed by one or more processors, may cause the one or more processors to perform one or more of the processes. In some embodiments, method 800 corresponds to the operation of program synthesis module 630 (e.g., FIGS. 6-7).

[0083] In step 802, a problem specification (e.g., 105 in FIG. 1) and a corresponding solution program (e.g., 106 in FIG. 1) may be received via an input interface (e.g., 615 in FIG. 6, 733 in FIG. 7).

[0084] In step 804, a pre-trained language model (e.g., 120 in FIG. 1) may be fine-tuned based on the problem specification and the corresponding solution program.

[0085] In step 806, the fine-tuned pre-trained language model (e.g., 130 in FIG. 1) may generate a sampled program (e.g., 133 in FIG. 1) in response to the problem specification (e.g., 105 in FIG. 1) at the decoding time step. For example, in one embodiment, the predicted token w of the sampled program t may be generated governed by the current parameters of the fine-tuned pre-trained language model (e.g., 130 in FIG. 1) at the decoding time step t. The hidden state representation of the fine-tuned pre-trained language model may be updated accordingly, and the next predicted token w of the sampled program t+1 may be generated using the updated hidden state representation of the fine-tuned pre-trained language model at the next decoding time step t + 1.

[0086] In step 808, a critic model (e.g., 140 in FIG. 1) can generate a return indicating the functional accuracy of the sampled program based on a comparison between the execution result of the sampled program and the test result of the problem specification. For example, in one embodiment, a return is generated when an end token is generated for the sampled program. The sampled program (e.g., 133 in FIG. 1) and the test (e.g., 112 in FIG. 1) may be passed to a compiler, and the first return value may be determined, for example, according to Equation (2), depending on whether the sampled program was successfully compiled and executed and whether the execution result matches the test result of the problem specification.

[0087] In step 810, a policy gradient of the expected value of the return when the current parameters of the fine-tuned pre-trained language model are given (e.g., the gradient of 135 in FIG. 1) can be calculated. For example, in one embodiment, the policy gradient is calculated as an estimated value based on the return and gradient of the conditional probabilities of the previous predicted token and the predicted token conditioned on the problem specification, for example, according to Equation (4).

[0088] In one implementation, the policy gradient can be calculated using a baseline comparison. For example, a baseline program generated by a base model in response to a problem specification can be input to a critic model (e.g., 140 in FIG. 1). The second return value can be determined depending on whether the baseline program was successfully compiled and executed and whether the execution result matches the test result of the problem specification. In this way, the policy gradient is calculated, for example, according to Equation (5), based on the difference between the first return value and the second return value, as well as the gradient of the conditional probabilities of the previous predicted token and the predicted token conditioned on the problem specification.

[0089] In other embodiments, the policy gradient is calculated, for example, according to Equation (9), based on the probability distribution of the predicted test outcomes generated by the critic model and the gradient of the conditional probability of the predicted tokens conditioned on the previous predicted tokens and the problem specification.

[0090] In step 812, the fine-tuned pre-trained language model (e.g., 130 in FIG. 1) may be updated according to the policy gradient.

[0091] In one implementation, a critic model (e.g., 140 in FIG. 1) is trained. The critic model receives a training sequence of the problem specification (e.g., 105 in FIG. 1) and the sampled program (e.g., 133 in FIG. 1), and the predicted test outcome corresponding to the sampled program can be calculated, for example, by the softmax operation of the maximum value pooled context hidden state of the decoder in the critic model according to Equations (6)-(7). For example, according to Equation (8), the cross-entropy loss by comparing the predicted test outcome with the execution result of the sampled program, and the critic model may be updated based on the cross-entropy loss.

[0092] FIG. 9 is an exemplary logical flow diagram showing a method of LM-based program synthesis shown in FIG. 5 according to some embodiments described herein. One or more of the processes of method 900, when executed at least in part by one or more processors, may cause the one or more processors to perform one or more of the processes, and may be implemented in the form of executable code stored on a non-transitory tangible machine-readable medium. In some embodiments, method 800 corresponds to the operation of program synthesis module 630 (e.g., FIGS. 6-7).

[0093] In step 902, the problem specification (e.g., 505 in FIG. 5) can be received in a pre-trained language model (e.g., 130 in FIG. 5) for program synthesis via an input interface (e.g., 615 in FIG. 6, 733 in FIG. 7).

[0094] In step 904, one or more unit test input-output pairs (e.g., 112 in FIG. 2) can be extracted from the problem specification (e.g., 105 in FIG. 2).

[0095] In step 906, the language model can generate a plurality of program samples (e.g., 533 in FIG. 5) from the problem specification.

[0096] In step 908, one or more unit tests (e.g., 112 in FIG. 5) can be applied to a plurality of program samples (e.g., 533 in FIG. 5) based on one or more unit test input-output pairs.

[0097] In step 910, from the plurality of program samples, a first set of program samples (e.g., 541 in FIG. 5) that passed one or more unit tests and a second set of program samples (e.g., 542 in FIG. 5) that failed can be determined. For example, the program samples in the second set include at least one of not passing at least one of a compile error, a runtime error, and a unit test.

[0098] In step 912, the critic model can determine, for example, according to formula (11), the value for a second program sample in the second set based on the predicted probability that the second program sample passes one or more unit tests.

[0099] In step 914, a subset of program samples having the highest value from the second set (e.g., 565 in FIG. 5) can be selected. An input sequence (e.g., 566) is formed by concatenating the problem specification, the selected program sample, and the error information corresponding to the selected program sample. For example, the error information includes either the unit test outcome corresponding to the selected program sample and the error subtype during compilation or runtime of the selected program sample.

[0100] In step 916, a program repair model can be used to generate a repaired program sample based on the input sequence. For example, the program repair model compares the program sample that failed the unit test with the ground truth program corresponding to the problem specification and is trained for training purposes conditioned on the unit test outcome and / or error subtype corresponding to the program sample.

[0101] In step 918, one or more subsequences (e.g., 543 in FIG. 5) can be selected from the first set of program samples via critical scoring. Each subsequence is a trimmed version of the program sample. For example, in one embodiment, the critical model can determine a value for each token of the first program sample in the first set based on the predicted probability that the subsequence up to that respective token passes one or more unit tests, e.g., according to Equation (10). A particular token of the first program sample having the highest value is identified, and the subsequence of the first program sample up to the particular token can be selected as the subsequence. If the selected subsequence includes a particular token where the corresponding subsequence is more likely to fail than pass one or more unit tests up to that particular token, the selected subsequence is further trimmed at the particular token.

[0102] In step 920, the language model can use, for example, a subsequence as a "seed" (e.g., 545 in FIG. 5) to generate the remaining tokens conditioned on one or more subsequences.

[0103] In step 922, the generated remaining tokens from step 920 can be combined with one or more subsequences to generate one or more improved program samples.

[0104] [Exemplary Data Experiments] In the exemplary data experiments of the proposed RL-based program synthesis framework shown in FIGS. 1-5 and the workflows of FIGS. 8-9, the CodeT5-large model (770M) is pre-trained from scratch according to the architecture of T5-large. In particular, the code-specific tokenizer is used from the CodeT5 research, and since the C / C# dataset is not publicly available, six programming languages (PLs) are used in CodeSearchNet (Husain et al., Codesearchnet challenge: Evaluating the state of semantic code search, Computing Research Repository (CoRR), abs / 1909.09436, 2019) (CSN) instead of the eight PLs in CodeT5. Only the pre-training task of masked span prediction (MSP) is applied, so the model does not need to parse the program into an abstract syntax tree (AST) to obtain identifier information.

[0105] The last preprocessing step was required in other original pre-training tasks such as masked identifier prediction in the original CodeT5 study. To further speed up training, data samples were concatenated to a batch size of 512 for pre-training using MSP, and the resulting number of tokens was 1.1B. To verify the advantages of using this new pre-trained CodeT5 as the base model (e.g., 130 in Figure 1), this model is evaluated on CodeXGLUE.

[0106] Exemplary data experiments were run on kubernetes using 16 A100-40G GPUs on Google Cloud Platform, and the total pre-training period was about 21 days. In the first pre-training stage using MSP, a corruption rate of 15%, a peak learning rate (LR) of 2e-4, and a batch size of 2048 were adopted. CSN was pre-trained over 150 epochs (10 days), and then 10 epochs (5 days) with GCPY. For the second stage of pre-training using NTP, a peak LR of 1e-4 and a batch size of 256, as well as 10 epochs (6 days) of pre-training were adopted. The maximum length is set to 768 and 600 for the source sequence and target sequence respectively for this purpose. For all experiments, the AdamW optimizer with a weight decay of 0:05 and a linear decay LR scheduler with 1000 warm-up steps were adopted.

[0107] The model is evaluated using the pass@k metric, which is the proportion of problems solved using k generated programs per problem, following (Hendrycks et al., Measuring coding challenge competence with apps, in proceedings of NeurIPS, 2021; Chen et al., Evaluating large language models trained on code, arXiv preprint, arXiv:2107.03374, 2021). The n@k metric is used, following (Li et al., Competition-level code generation with alphacode, arXiv preprint, arXiv:2203.07814, 2022), which considers only a subset of n candidates from the k generated programs per problem. The subset of n candidates is typically selected by a filtering method that passes the generated programs through exemplary tests given as part of the problem description.

[0108] Exemplary benchmarks for comparison include the APPS program synthesis benchmark (see Hendrycks et al.), because it has coding problems of various difficulties collected from multiple coding websites. APPS consists of 10,000 coding problems with a 50-50 training-test split. Each problem is accompanied by an average of 23.2 correct Python programs and 21.2 unit tests. The average length per problem is 293.2 words, and the average length per program is 18.0 lines. The dataset is classified into three difficulty levels, namely, Introductory (3639, training / test = 2639 / 1000), Interview (5000, training / test = 2000 / 3000), and Competition (1361, training / test = 361 / 1000). Each sample contains an average of 20 unit tests to verify the functional correctness of the program. The same preprocessing steps of Hendrycks et al. are used to formulate the input sequence from the problem description.

[0109] On APPS, the pre-trained CodeT5 is fine-tuned with the RL-based framework described in FIG. 1. To warm-start the CodeT5 model with Lce, a batch size of 64 and a warm-up LR from 0 to 2e-5 are used for the first 500 steps, decaying polynomially (power = 0.5) to 1e-5 by the end of 10 epochs, which takes about 30 hours on one A100 GPU. The maximum source and target sequence lengths are set to 600 and 512, respectively.

[0110] Additional benchmarks include the MBPP [Mostly Basic Programming Problems] benchmark, a smaller and simpler Python program synthesis dataset for evaluation (described in Austin et al., Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021). The dataset contains 974 instances with 374 / 90 / 500 instances for training / validation / test respectively, and 10 instances reserved for few-shot learning. The problems are typically short, usually each being one sentence in natural language description. Each problem is accompanied by one correct solution (on average 6.8 lines of code) and three unit tests in the form of assert statements to verify functional correctness. Unlike APPS, the unit tests in MBPP are not hidden and are explicitly incorporated into the source sequence for the program synthesis model. This may encourage the model to overfit to these assert statements, though very rarely, by hard-coding if expressions. However, for a fair comparison with baselines, the source sequence is constructed in the same way as in prior research. Specifically, using the same prompt format as Austin et al., the input sequence is prepared as follows: problem description + "Your code should satisfy these tests:" + three assert statements.

[0111] Experiments are conducted on MBPP in both zero-shot (Section 4.5) and fully fine-tuning setups. To fine-tune CodeT5, since the training set of MBPP is small, the model is fine-tuned for 60 epochs with a constant LR of 2e-5 and a batch size of 32, which takes less than 30 minutes on one A100. The maximum source and target lengths are set to 382 and 306 respectively.

[0112] Exemplary baselines include GPT2 (Radford et al., Language models are unsupervised multitask learners, OpenAI blog, 1(8):9, 2019), GPT-Neo (Black et al., GPT-NEO: Large scale autoregressive language modeling with mesh-tensorflow. URL https: / / doi.org / 10.5281 / zenodo, 5297715, 2021), and GPT3 (Brown et al., Language models are few-shot learners. Advances in neural information processing systems, 33:1877-1901, 2020) for comparison with the RL-based framework described herein (referred to as "CodeRL"). Results are also compared with Codex (see Chen et al.) and AlphaCode (see Li et al.). Note that by default, results for pre-trained LMs (except Codex and GPT3) are from models fine-tuned on APPS using only the standard loss Lce. Since CodeRL is model-independent, it can be integrated with GPT variants such as GPT-J and GPT-Neo.

[0113] Figure 10(a) shows that CodeRL using the CodeT5 model can achieve significant performance gains and outperforms many pre-trained LMs of much larger sizes. Specifically, CodeRL achieved new SOTA results of 2:69% pass@1, 6:81% pass@5, and 20:98% pass@1000. Figure 10(b) shows that CodeRL+CodeT5 can achieve SOTA results of 8:48% 1@k and 12:62% 5@k when evaluated on a subset of filtered code samples.

[0114] Second, the benefits of upsampling generation can be seen when increasing the number of generated samples k from 1 to 1000. Note that CodeRL incurs additional computational costs during inference in CS, but CodeRL requires only a much lower k to achieve performance comparable to other models.

[0115] Specifically, with just k = 1000, CodeRL performance is about as good as AlphaCode, which has a much larger generation budget of k = 50000. Finally, Figure 10(b) also shows that fine-tuning can significantly improve model performance for difficult programming tasks at the interview and competition levels. Specifically, Codex, which was not fine-tuned on APPS and was tested in the few-shot setting, was able to achieve good n@1000 results, but the model dramatically failed on synthetic tasks at the interview and competition levels. This observation indicates a significant gap between the pre-training stage and the downstream synthetic tasks.

[0116] Figure 11 shows the results of CodeT5-770M trained by different approaches to estimate the return of code samples. Overall, the CodeRL objective with relative token-level return estimates by the critic model (Model D) can achieve the best performance in pass@1 and pass@5. Second, note that using absolute returns without a baseline (Model B) can result in the largest performance degradation because this approach imposes a heavy penalty on all inaccurate samples (even if they are still better than naive baselines). Therefore, considering relative return estimation that can effectively utilize incomplete code can lead to a better synthesis system. Third, when there is no critic model, simply assigning the same reward to all tokens in the code sample (Model A) is disadvantageous because these return estimates are too restrictive to be used as feedback signals for RL training. For example, a program may be considered incorrect only by additional whitespace, which can lead to indentation errors in Python programs. Simply assigning the same reward to all tokens in this program would impose a large penalty on the correct parts of the program sequence. Finally, experiments were conducted using a distance-based critic that assumes that the token value

Number

[0117] Figure 12 shows the results for different combinations of Lce and Lrl. Since CodeRL is model-independent, experiments are conducted for both CodeT5 and GPT-Neo. Note that in these experiments, Lce and Lrl are applied to models that have already been warm-started / fine-tuned with Lce up to 10 epochs. First, when using only Lrl, a problem of gradient disappearance was observed during fine-tuning. Therefore, the final model actually deteriorates, leading to a performance decline. Second, by using only Lce for further fine-tuning, despite the improvement in loss during training, the model performance actually deteriorates during test time. Therefore, these models are expected to overfit to the training data, as is also seen in the analysis of the pre-trained models in Figure 16.

[0118] In addition, it is recognized that the naive approach of Lce using synthetic samples Ws (all of which are treated as correct codes with r(Ws)=1) still brings some performance improvement using GPT-Neo in terms of pass@5. However, in all other cases, this training strategy does not work as well as considering the critic model to estimate the return of Ws based on their test results. Finally, it is recognized that using both Lce and Lrl brings overall more consistent performance improvement in terms of pass@1 and pass@5 for both the GPT-Neo model and the CodeT5 model.

[0119] Figure 13 shows the ablation results of Critical Sampling (CS) during inference applied to the CodeT5 model. Different combinations of program improvement steps and repair steps are tested. Overall, across all metrics, the beneficial effects of CS that combines both program improvement and repair are seen, especially with more significant gains at pass@1000. Note that program improvement alone also helps improve performance, but its impact is reduced for the 1@1000 metric. Note that n@k measures the resolution rate between the subset P filtered from k samples. Since program improvement technically increases the size of this subset, the n@k metric considers an exponentially larger number of n sample options than before. This normalizes n@k by a larger pool of n candidate sets and reduces the impact of program improvement on model performance.

[0120] Second, (regarding the problem when P = ;), when program improvement is integrated with program repair, further performance is obtained in all metrics. Interestingly, when experimenting with different top-M selection methods, the best overall performance is found at M = 1, and the performance begins to decline from M = 2 to M = 4 (except for the pass@200 result). This observation shows the advantage of using a critic model rather than selecting multiple program candidates to focus on the best candidates for program repair. Furthermore, for larger M, each program candidate will have a smaller batch size (i.e., N = M). This reduces the likelihood that the program repair model will properly repair and generate the correct program.

[0121] In one embodiment, the data experiment examines a subset of the APPS test split that includes the highest difficulty level test samples (i.e., competitive programming tasks). Figure 14 shows the pass@k and n@k results for CodeRL+CodeT5 and CodeT5 only, for k in the range 1 to 200, and n = ⌈1.5⌉. Since CodeRL is model-independent, it is integrated with GPT-J and the results are reported. To focus on the impact of RL optimization, during test time, the model uses only nucleus sampling and is compared to the case without the CS procedure.

[0122] Figure 14 shows that the performance gains are fairly consistent for both GPT-J and CodeT5. In particular, as k increases, the performance gains of CodeRL become more prominent in both the GPT-J model and the CodeT5 model. These gains are thought to be due to the CodeRL learning objective Lrl that encourages the model to explore code solutions drawn from the model's sampling distribution. During test time when the k sampling budget increases, the model becomes able to generate diverse code solutions and the impact of Lrl becomes more prominent.

[0123] In one embodiment, the performance of the synthesis system is interrelated with the quality of the base model. FIG. 15 reports the results of CodeT5 using different configurations of model size, pre-training data, and pre-training objectives. For a fair comparison, all models are only fine-tuned / warm-started on the APPS with Lce up to 12 epochs. It can be seen that by scaling up the number of model parameters (from 60M to 770M), the model performance on downstream synthesis tasks can be significantly improved. When the pre-training data is improved by adding the GCPY dataset (10 times larger than the CSN dataset), good performance improvements can be seen, namely, from 1:3 to 1:56 pass@1, and from 1:72 to 2:06 pass@5. Finally, by combining the pre-training objectives from Masked Span Prediction (MSP) and Next Token Prediction (NTP), the model can better adapt to downstream synthesis.

[0124] FIG. 16 shows the performance of the CodeT5 model variants by fine-tuning the epochs and by the difficulty level of the programming tasks. In these experiments, it should be noted that the data experiments are only compared between the CodeT5 model variants by the pre-training strategy, and thus only involve Lce in the fine-tuning stage on the APPS. Consistent with our previous analysis, strengthening both the pre-training data (using larger data of GCPY) and the pre-training objective (using the NTP objective) generally improves the model performance over the training epochs. Furthermore, as shown by the analysis of the learning objective, using only Lce often results in overfitting performance, typically after epoch 10 in our case. Therefore, it is beneficial to utilize synthetic training samples to further fine-tune large-scale LMs and adopt the RL objective Lrl to avoid overfitting of the model.

[0125] Figure 17 reports the results of our CodeRL+CodeT5 on the MBPP benchmark compared to the fine-tuned GPT models up to 137B size. CodeRL+CodeT5(ZS) was trained on APPS and then evaluated on MBPP in zero-shot setting. CodeRL using the very small model size of CodeT5 is found to bring surprisingly good zero-shot performance, setting a new SOTA result of 63.0% pass@80 against 61.4% pass@80 of GPT-137B. This verifies the strong zero-shot transfer ability of CodeRL for unseen tasks.

[0126] A common concern regarding transfer learning is that the source (APPS) task and the target (MBPP) task may have overlap in their training data, and as a result, the source model may tend to memorize these substantially similar data when applied to the target task. To address this concern, following Austin et al., it is analyzed how many lines of code appear in both the training set of APPS and the programs of MBPP. For this analysis, code comments are discarded, whitespace is normalized for each line, and then lines that appear more than twice anywhere within MBPP are excluded as they are likely to be common Python keywords such as return and break.

[0127] Figure 18 shows the absolute number of duplicate lines (left) and the relative fraction of duplicate lines (right) in the MBPP programs. As can be seen from the figure, the overlap between APPS and MBPP seems to be minimal. Only 12.6% of the MBPP programs have more than half of their lines matching somewhere in the APPS training data. Furthermore, more than half of the programs (514 out of 974) have zero duplicates, and 90.9% have no more than 3 lines overlapping with the APPS training set. Additionally, if consecutive lines are required, there are no more than 2 consecutive duplicate lines.

[0128] Figure 19 shows the average percentage of generated programs per problem, grouped by their test outcomes. Specifically, CodeT5 or CodeRL+CodeT5 is used to generate programs, and 200 generated programs per test sample are randomly selected in the APPS test split. The programs are passed to either an exemplary unit test or a hidden unit test, and the output programs are grouped by their test outcomes. The outcomes are classified according to the definition of Equation (2) including CompileError, RuntimeError, FailedTest, and PassedTest.

[0129] First, it is observed that integrating CodeRL in both exemplary unit tests and hidden unit tests can increase the likelihood that a program passes the test and reduce the probability of failing one or more unit tests. The probability of passing the unit test is more significantly improved in entry-level programming problems.

[0130] Second, it is noted that the percentage of programs with compilation errors in programs created with CodeRL decreases and has a greater impact on interview and competition-level problems. Since the likelihood of compilation errors occurring in CodeRL programs is low, these programs still suffer from runtime errors. This leads to an increased probability that CodeRL programs contain runtime errors.

[0131] Note that there is a fairly significant performance gap between exemplary unit tests (Figure 19(a)) and hidden unit tests (Figure 19(b)) depending on the test outcome. This observation suggests that exemplary tests are not as comprehensive as hidden tests and thus limit the positive impact of the CodeRL generation procedure due to false positives.

[0132] Figure 20 shows an example of a programming problem from the APPS benchmark and the corresponding program generated by the CodeT5 variant. Specifically, a CodeT5 model fine-tuned only by Lce based on the same base pre-trained CodeT5 (pre-trained for GCPY data and NTP purposes) is compared with other models following the CodeRL framework. In CodeRL+CodeT5, the programs before and after applying the CS procedure are shown. By applying CodeRL, more appropriate programs can be generated, and it is seen that by using the CS procedure, their functional accuracy is further improved. For example, in Figure 20, the CodeT5 model misinterprets the problem and focuses on finding only the greatest common divisor between a and b. Instead, the CodeRL model avoids this error and addresses the problem of finding the greatest common divisor between the factorials of a and b.

[0133] Also, it has been found that CodeRL can improve the complexity of the generated programs, which is an important property in complex programming problems. For example, in the interview-level program of Figure 20, when CS is not applied, the generated program is functionally correct but fails during execution due to a timeout error. This program simply calculates the separate factorials of both a and b, which slows down execution in scenarios with extremely large a or b. Applying the CS procedure can condition the model on parts of the previous program and (re)generate new tokens to produce a more efficient program. In the example of Figure 20, to improve the efficiency of the program, the factorial is calculated for min(a,b). Therefore, the resulting final program can pass all the hidden unit tests (including tests with extremely large values) without a timeout error.

[0134] This description and the accompanying drawings, which illustrate aspects, embodiments, implementations, or applications of the invention, should not be construed as limiting. Various mechanical, compositional, structural, electrical, and operational changes can be made without departing from the spirit and scope of this specification and the claims. In some instances, well-known circuits, structures, or techniques are not shown or described in detail to avoid obscuring the embodiments of the present disclosure. Like numbers in two or more figures represent the same or similar elements.

[0135] In this description, specific details are set forth that describe several embodiments consistent with the present disclosure. Numerous specific details are set forth to provide a complete understanding of the embodiments. However, it will be apparent to those skilled in the art that some embodiments can be practiced without some or all of these specific details. The specific embodiments disclosed herein are meant to be illustrative but not limiting. Those skilled in the art can recognize other elements within the scope and spirit of the present disclosure that are not specifically described herein. Additionally, to avoid unnecessary repetition, one or more features shown and described in connection with one embodiment may be incorporated into other embodiments, unless specifically stated otherwise or unless one or more features render the embodiment non-functional.

[0136] Exemplary embodiments have been shown and described, but extensive modifications, changes, and substitutions are contemplated in the foregoing disclosure, and in some cases, some features of the embodiments may be employed without the corresponding use of other features. Those skilled in the art will recognize many variations, alternatives, and modifications. Accordingly, the scope of the present invention should be limited only by the following claims, and it is appropriate that the claims be construed broadly to be consistent with the scope of the embodiments disclosed herein.< / endoftext> < / endoftext>

Claims

1. A method of a reinforcement learning framework for program synthesis, comprising: receiving, via an input interface, a problem specification and a corresponding solution program; fine-tuning a pre-trained language model based on the problem specification and the corresponding solution program; generating, by the fine-tuned pre-trained language model, at a decoding time step, a sampled program in response to the problem specification; generating, by a critic model, a return indicating the functional accuracy of the sampled program based on comparing an execution result of the sampled program with a test result of the problem specification; calculating a policy gradient of an expected value of the return given current parameters of the fine-tuned pre-trained language model; updating the fine-tuned pre-trained language model according to the policy gradient; A method comprising the above steps.

2. The sampled program is generated by: generating, at the decoding time step, a predicted token of the sampled program governed by the current parameters of the fine-tuned pre-trained language model; updating a hidden state representation of the fine-tuned pre-trained language model; generating, at a next decoding time step, a next predicted token of the sampled program using the updated hidden state representation of the fine-tuned pre-trained language model; The method according to claim 1.

3. The method according to claim 2, wherein the return is generated when an end token is generated for the sampled program.

4. The return is generated by: passing the sampled program and a test to a compiler; determining a first return value according to whether the sampled program is successfully compiled and executed and whether the execution result matches the test result of the problem specification; The method according to claim 1.

5. The method according to claim 4, wherein the policy gradient is calculated as an estimated value based on the return and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification.

6. inputting a baseline program generated by a base model in response to the problem specification into the critic model; determining a second return value according to whether the baseline program is successfully compiled and executed and whether the execution result matches the test result of the problem specification; The method according to claim 4, further comprising:

7. The method according to claim 6, wherein the policy gradient is calculated based on the difference between the first return value and the second return value and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification.

8. The critic model: receiving, by the critic model, a training sequence of the problem specification and the sampled program; generating, by the critic model, a predicted test outcome corresponding to the sampled program; calculating a cross-entropy loss by comparing the predicted test outcome with the execution result of the sampled program; updating the critic model based on the cross-entropy loss; The method according to claim 1, wherein the method is trained by:

9. The method according to claim 8, wherein the predicted test outcome is calculated by a softmax operation on the maximum value pooled context hidden state of the decoder in the critic model.

10. The method according to claim 9, wherein the policy gradient is calculated based on the probability distribution of the predicted test outcome generated by the critic model and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification.

11. A system of a reinforcement learning framework for program synthesis, comprising: an input interface that receives a problem specification and a corresponding solution program; a memory that stores a plurality of processor-executable instructions; reading and executing the plurality of processor-executable instructions to Fine-tuning a pre-trained language model based on the problem specification and the corresponding solution program; Generating a sampled program in response to the problem specification at a decoding time step using the fine-tuned pre-trained language model; Generating a return indicating the functional accuracy of the sampled program based on comparing the execution result of the sampled program with the test result of the problem specification by a critic model; Calculating a policy gradient of the expected value of the return given the current parameters of the fine-tuned pre-trained language model; Updating the fine-tuned pre-trained language model according to the policy gradient; A processor that performs operations including; A system including.

12. The sampled program is Generating a predicted token of the sampled program governed by the current parameters of the fine-tuned pre-trained language model at the decoding time step; Updating the hidden state representation of the fine-tuned pre-trained language model; Generating the next predicted token of the sampled program using the updated hidden state representation of the fine-tuned pre-trained language model at the next decoding time step; Generated by The system according to claim 11.

13. The system according to claim 12, wherein the return is generated when an end token is generated for the sampled program.

14. The return is Passing the sampled program and test to a compiler; Determining a first return value according to whether the sampled program is successfully compiled and executed and whether the execution result matches the test result of the problem specification; Generated by The system according to claim 11.

15. The system according to claim 14, wherein the policy gradient is calculated as an estimate based on the return and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification.

16. The operations are Inputting the baseline program generated by the base model in response to the problem specification into the critic model; Determining a second return value according to whether the baseline program is successfully compiled and executed and whether the execution result matches the test result of the problem specification; The system according to claim 14, further comprising.

17. The system according to claim 16, wherein the policy gradient is calculated based on the difference between the first return value and the second return value and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification.

18. The critic model: Receiving, by the critic model, a training sequence of the problem specification and the sampled program; Generating, by the critic model, a predicted test outcome corresponding to the sampled program; Calculating a cross-entropy loss by comparing the predicted test outcome with the execution result of the sampled program; Updating the critic model based on the cross-entropy loss; Trained by: The system according to claim 11.

19. The predicted test outcome is calculated by a softmax operation on the maximum value pooled context hidden state of the decoder in the critic model. The policy gradient is calculated based on the probability distribution of the predicted test outcome generated by the critic model and the gradient of the conditional probability of the previous predicted token and the predicted token conditioned on the problem specification. The system according to claim 18.

20. A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for a reinforcement learning framework for program synthesis, the processor-executable instructions comprising: Receiving, via an input interface, a problem specification and a corresponding solution program; Fine-tuning a pre-trained language model based on the problem specification and the corresponding solution program; generating, by the fine-tuned pre-trained language model, a program sampled in response to the problem specification at a decoding time step; generating, by a critic model, a return indicating the functional accuracy of the sampled program based on comparing an execution result of the sampled program with a test result of the problem specification; calculating, given current parameters of the fine-tuned pre-trained language model, a policy gradient of an expected value of the return; updating the fine-tuned pre-trained language model according to the policy gradient; performed by one or more processors to perform operations including; a non-transitory processor-readable storage medium. [

21. ] A method for program synthesis, comprising: receiving, via an input interface, a problem specification in a pre-trained language model for program synthesis; extracting one or more unit test input-output pairs from the problem specification; generating, by the language model, a plurality of program samples from the problem specification; applying one or more unit tests to the plurality of program samples based on the one or more unit test input-output pairs; determining, from the plurality of program samples, a first set of program samples that pass the one or more unit tests; regenerating, by the language model, one or more improved program samples based on information from the first set of program samples; A method comprising. [

22. ] The regenerating, by the language model, one or more improved program samples comprises: selecting, via a critic scoring, one or more subsequences from the first set of program samples, each subsequence being a trimmed version of a program sample; regenerating, by the language model, subsequent tokens conditioned on the one or more subsequences to form the one or more improved program samples; The method according to claim 21, comprising. [

23. ] The one or more subsequences are Determining, by a critic model, a value for each token of a first program sample in the first set based on a predicted probability that a subsequence up to each respective token passes the one or more unit tests; Determining a specific token of the first program sample having the highest value; Selecting, as a subsequence, the subsequence of the first program sample up to the specific token; The method according to claim 22, as selected by.

24. Determining that the selected subsequence includes a specific token and that, up to the specific token, the corresponding subsequence has a higher probability of failing than of passing the one or more unit tests; Splitting the selected subsequence at the specific token; The method according to claim 23, further comprising.

25. Determining, from the plurality of program samples, a second set of unsuccessful program samples including at least one of a compilation error, a runtime error, and a failure to pass at least one of the unit tests; The method according to claim 21, further comprising.

26. Determining, by a critic model, a value for a second program sample in the second set based on a predicted probability that the second program sample passes the one or more unit tests; Selecting a subset of program samples having the highest value from the second set; Forming an input sequence by concatenating the problem specification, the selected program samples, and error information corresponding to the selected program samples; Generating, by a program repair model, a repaired program sample based on the input sequence; The method according to claim 25.

27. The error information is a unit test outcome corresponding to the selected program sample; and either an error subtype during compilation or runtime of the selected program sample. The method according to claim 26.

28. The method according to claim 26, wherein the program repair model is sequence-to-sequence.

29. The program repair model compares a program sample that failed the unit test with a ground truth program corresponding to the problem specification, and is trained for a training purpose conditioned on the unit test outcome and / or error subtype corresponding to the program sample, according to the method of claim 26.

30. Generating, by the program repair model, one or more repaired program samples using error information from a second set of the program samples; Improving, by the language model, the one or more repaired program samples; The method according to claim 25, further comprising.

31. A system for program synthesis, An input interface that receives a problem specification in a pre-trained language model for program synthesis; A memory storing a plurality of processor-executable instructions; Reading and executing the plurality of processor-executable instructions to Extract one or more unit test input-output pairs from the problem specification; Generating, by the language model, a plurality of program samples from the problem specification; Applying one or more unit tests to the plurality of program samples based on the one or more unit test input-output pairs; Determining, from the plurality of program samples, a first set of program samples that passed the one or more unit tests; Regenerating, by the language model, one or more improved program samples based on information from the first set of program samples; A processor that performs operations including; A system comprising.

32. The operation of regenerating one or more improved program samples by the language model is Selecting, via critical scoring, one or more subsequences from the first set of program samples, each subsequence being a trimmed version of the program sample; Regenerating subsequent tokens conditioned on the one or more subsequences by the language model to form the one or more improved program samples, The system according to claim 31.

33. The one or more subsequences are Determining, by a critic model, a value for each token of a first program sample in the first set based on a predicted probability that a subsequence up to each token passes the one or more unit tests; Determining a specific token of the first program sample having the highest value; Selecting, as a subsequence, the subsequence of the first program sample up to the specific token; The system according to claim 32, selected by.

34. The operation is Determining that the selected subsequence includes a specific token and that, up to the specific token, the corresponding subsequence has a higher probability of failing than of passing the one or more unit tests; Truncating the selected subsequence at the specific token; The system according to claim 33, further comprising.

35. The operation is Determining, from the plurality of program samples, a second set of unsuccessful program samples including at least one of a compile error, a runtime error, and a failure to pass at least one of the unit tests; The system according to claim 31, further comprising.

36. The operation is Determining, by a critic model, a value for a second program sample in the second set based on a predicted probability that the second program sample passes the one or more unit tests; Selecting a subset of program samples having the highest value from the second set; Forming an input sequence by concatenating the problem specification, the selected program sample, and error information corresponding to the selected program sample; Generating, by a program repair model, a repaired program sample based on the input sequence; The system according to claim 35, further comprising.

37. The error information is Either a unit test outcome corresponding to the selected program sample or an error subtype during compilation or runtime of the selected program sample, the system according to claim 36. The system according to claim 36, including either.

38. The program repair model compares the program sample that failed the unit test with the ground truth program corresponding to the problem specification, and is trained for training purposes, conditioned on the unit test outcome and / or error subtype corresponding to the program sample, of the system according to claim 36.

39. The operation is generating, by a program repair model, one or more repaired program samples using error information from a second set of the program samples; improving, by the language model, the one or more repaired program samples; The system according to claim 35, further comprising.

40. A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for program synthesis, the processor-executable instructions receiving, via an input interface, a problem specification in a pre-trained language model for program synthesis; extracting one or more unit test input-output pairs from the problem specification; generating, by the language model, a plurality of program samples from the problem specification; applying one or more unit tests to the plurality of program samples based on the one or more unit test input-output pairs; determining, from the plurality of program samples, a first set of program samples that pass the one or more unit tests; regenerating, by the language model, one or more improved program samples based on information from the first set of the program samples; to perform operations including, by one or more processors; A non-transitory processor-readable storage medium.

Citation Information

Patent Citations

  • Neural method completion based on natural language and source code

    US20210357187A1

  • Dual bayesian encoding-decoding technique for text to code transformations

    US20220035605A1

  • Search device, learning device, search method, learning method, and program

    WO2020194792A1

  • Program generation device, program generation method, and program

    WO2021144904A1