Methods for improving memory allocation of llm-generated code
By integrating memory sanitization, fuzzing, and metadata storage into the code generation process using LLM, the method addresses the challenges of code accuracy and memory efficiency in larger codebases, resulting in improved memory allocation and optimized code generation.
Patent Information
- Application Number
- JP2024182494
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-19
- Filing Date
- 2024-10-18
- Publication Date
- 2025-05-02
AI Technical Summary
Existing methods for automatic code generation using Large Language Models (LLM) face challenges in ensuring code accuracy and performance improvements, especially for larger codebases, and struggle to detect memory inefficiencies effectively.
The proposed method involves generating program code using LLM, compiling and instrumenting it with a memory sanitizer, performing fuzzing to monitor memory performance, and storing metadata in a database to improve memory allocation and optimize code generation.
This approach allows for iterative learning and improvement of LLM-generated code, effectively testing memory usage and generating more optimized code, thereby addressing the limitations of existing methods.
Smart Images

Figure 2025071047000001 
Figure 2025071047000002 
Figure 2025071047000003
Abstract
Description
[Background technology]
[0001] Language models, such as Large Language Models (LLMs), are increasingly being used to generate program code, for example as part of language translation, automatic program code translation, and code refactoring.
[0002] Using approaches for automatic code generation, such as those from LLM, always comes with challenges, as there are no guarantees about the correctness of the code, nor are there any guarantees of performance improvements: for example, while improvements can be generated for small code fragments, the approach is impractical for larger code bases.
[0003] Working with larger code bases reduces the usability of certain static methods, e.g., abstract interpretation, because correctness proofs either cannot be computed in a reasonable time or require approximations, resulting in overapproximation errors. Additionally, static methods are poorly suited to measuring software performance, e.g., memory consumption.
[0004] The sanitizer can detect all kinds of runtime errors, with a particular focus on memory errors and memory inefficiencies that occur at runtime. Partially initialized variables are particularly dangerous: a variable's value may be partially determined when it is initialized, for example by a write operation. Because the C++14 standard allows passing indeterminate values in some cases, the sanitizer usually reports many false positives. Partially initialized variables can be vulnerable when uninitialized values exceed the trusted range.
[0005] Uninitialized memory issues can often be resolved with memory sanitizers. However, LLM-based code generators may generate suboptimal but functionally correct code (e.g., increased heap memory allocations). Tests, e.g., unit tests to check optimizations of LLM-based code, may be poorly or unable to detect such inefficiencies. Summary of the Invention [Means for solving the problem]
[0006] A first general aspect of this disclosure relates to a method for improving memory allocation for code generated using a language model. This method is - providing the program code generated using the language model when a new version of the program code is available; - generating an executable file by compiling and instrumenting, where the memory sanitizer inserts instructions into the program code and / or the executable file; - performing fuzzing with a fuzzer, the fuzzer injecting inputs into the executable file; - monitoring memory performance, and optionally runtime information, behavior and / or output of executable files; - storing metadata generated from the allocated and freed memory in a memory metadata database, the metadata being stored during generation of the executable and / or during fuzzing based on the instructions; - outputting the program code if no memory performance degradation or other errors are found; Includes.
[0007] A second general aspect of the present disclosure relates to a method for training a language model configured for the automatic generation of program code. This method is - inputting the source code into a language model to generate program code; - improving program code using a method according to the first general aspect; - generating a reward for the language model, the reward being based on memory performance, optionally runtime information, monitoring of the behavior and / or output of the executable file, and / or metadata; - updating the weights of the language model with the reward values; Includes.
[0008] A third general aspect of the present disclosure relates to a computer system configured to carry out the method according to the first and / or second general aspects (or embodiments thereof). A fourth general aspect of the present disclosure relates to a computer program configured to carry out the method according to the first general aspect (or an embodiment thereof).
[0009] A fifth general aspect of the present disclosure relates to a computer-readable medium or signal storing and / or including a computer program according to the fourth general aspect (or an embodiment thereof).
[0010] The techniques of the first, second, third, fourth and fifth general aspects may, in some circumstances, have one or more of the following advantages. Effect of the Invention
[0011] The present disclosure uses a dynamic software testing methodology of fuzzing, specifically fuzzing augmented with a sanitizer. The memory sanitizer is used to visualize memory behavior (e.g., allocations, deallocations and their sizes). Fuzzing is used to cover as many different paths as possible through the software under test. The observed memory behavior can be used to evaluate whether the performance of the LLM-generated code has improved, allowing a feedback loop for better prompts for newer LLM-generated code.
[0012] This disclosure enables testing of software where there are no existing tests, such as unit tests that check whether patches in generated code are functionally correct. These tests typically do not measure memory consumption, nor can they measure improvements in memory performance. Existing tests fail when too much memory is used. When using LLMs for code generation, there is always a risk that the generated code will contain nonsensical or poorly functional code. This disclosure enables iterative learning to improve the code generated by the LLM over multiple generations.
[0013] A further challenge for LLM is refactoring and optimizing code for different purposes. A typical allocation of heap memory is to release it after all operations on the data are completed. In devices with limited memory space, such as IoT devices, refactoring can be beneficial to reuse heap memory between individual data operations so that not all variables need to be kept in memory. The present disclosure allows monitoring and improving memory allocation to support such refactoring. In other words, the present disclosure allows effectively testing the memory usage of code generated by a language model (e.g., LLM) so that the LLM can generate better, especially memory-optimized, code.
[0014] The present disclosure enables coverage as a measure of the quality of the generated code, which in turn allows for better selection of generated code. Such high-quality code is easier to test. This disclosure is relevant to any product based on automated testing, especially dynamic testing techniques, and to any product that has legacy code or performance issues.
[0015] In this disclosure, several terms are used as follows: A "language model" is a very general language model that includes, in particular, large-scale language models (LLMs), neural networks, recurrent neural networks (RNNs), transformer models, or code models as language models specialized for code, or even code. It also includes computer languages, program codes of computing devices such as computers, and the like. The language of the model includes not only natural languages, but also artificial languages, such as programming languages.
[0016] In software development, code "refactoring" refers to improving the structure of code while preserving the observable program behavior, i.e. functionality, e.g. to improve readability, understandability, maintainability and / or extensibility, with the goal of reducing the effort for error analysis and / or enhancements. A typical refactoring is, for example, changing variable names to more user-friendly names and / or extracting parts of code in a different way. Refactoring improves the quality of the code and therefore the quality of the software.
[0017] "Testing", "verification", or "comparison of source and target program code" includes formal verification of the same behavior of the source and target program code, e.g., by bounded model checking, testing in the source language, testing of contracts in the source language, and / or syntactic and stylistic testing, fuzzing, mutating inputs to a test harness, deriving from contracts in the source and / or target language, and / or deriving from language models, etc.
[0018] A "test harness" or test frame includes a collection of software and test data used for systematic automated testing of a program under various environmental conditions. A test harness typically includes a test execution engine responsible for processing the test logic and a test data repository or database that contains test scripts, test programs, and other test resources. Here, a test harness can be automatically generated, for example, by adding differentiation tests to the database. Tests can be initiated using specified or ready-made tests from the test database. Additionally, the system can automatically generate tests.
[0019] Here, the data can be the software code, including test cases and harnesses, plus additional (natural language) descriptions of areas of functionality and validity. In the case of language translation, we exemplarily describe C as the source language and Rust as the target language, but other combinations are possible. A translation from C to Rust is advantageous because while Rust offers features in the area of safety-critical systems, there is a lot of legacy code in other languages, especially C. In the case of refactoring, the source and target languages are the same.
[0020] A "memory sanitizer" is a tool that detects the use of uninitialized memory. Memory sanitizers can insert additional instructions and / or preserve metadata to operate on. Memory sanitizers can detect the use of uninitialized memory, memory leaks, out of memory space, hangs or infinite loops, and infinite recursion (stack overflow). Other sanitizers detect, for example, use-after-free, buffer overflows, data races, deadlocks, int / float overflows, or bitwise shifts by invalid amounts.
[0021] An "intermediate representation" (IR), or more broadly known as an intermediate language, is a data structure or code that is produced during the translation process by a compiler or virtual machine, at a level of abstraction between a higher-level source language and a target language that is usually closer to the machine. An intermediate representation can be used internally by a compiler or virtual machine to represent source code.
[0022] A "contract" is a component of contract-based programming or "design by contract", a software development concept aimed at optimizing the interaction of individual program modules by defining formal contracts for the use of their interfaces that go beyond their static definitions.
[0023] The term "codebase" refers to the set of source text files and associated configuration files that belong to a project, including various other files needed for the compilation process, such as so-called makefiles.
[0024] "Fuzzing" or "fuzz testing" is the automated process of sending randomly generated inputs from a fuzzer to a target or target program and observing the target's reaction.
[0025] A "fuzzer" or "fuzzing engine" is a program that generates inputs automatically. As such, they are not necessarily connected to the software under test and are not instrumented, but are capable of instrumenting code, generating test cases, and running the program under test. Known examples are afl and libfuzzer.
[0026] A "fuzz target" is a software program or function that is to be tested by fuzzing. The main characteristic of a fuzz target is that it is a binary file, a library, an application programming interface (API), or anything else that can process bytes as input.
[0027] The "glue code", "wrapper", "harness" or "fuzz driver" connects the fuzzer and the fuzz target. A "fuzz test" is a combined version of a fuzzer and a fuzz target. A fuzz target could be an instrumented piece of code with a fuzzer attached to its input. A fuzz test is executable. A fuzzer can also start, monitor and stop multiple running fuzz tests (typically hundreds or thousands per second) with slightly different inputs generated by the fuzzer.
[0028] A "test case" is an execution of a particular test with specific inputs from a test harness or fuzz test. Winning executions (discovery of new code paths or crashes) are saved to ensure reproducibility.
[0029] "Instrumentation" is used to make coverage metrics observable, for example at compile time. Instrumentation means inserting instructions into a program to get feedback on its execution. Instrumentation is usually achieved by a compiler and can for example describe which code blocks were reached during execution.
[0030] Coverage-guided fuzzing uses code coverage information as feedback during fuzzing to detect if an input causes the execution of a new code path / block.
[0031] "Mutation-based fuzzing" takes a set of known inputs (a corpus) and creates new inputs by applying random mutations to them. "Generation-based fuzzing" involves creating new inputs from scratch, for example using input models and input grammars.
[0032] A "mutator" is a function that takes a byte as input and outputs a small random mutation of the input. A "corpus" is a set of inputs. The initial inputs are the seeds. [Brief description of the drawings]
[0033] [Figure 1] 1 is a flow chart illustrating the techniques of this disclosure for improving memory allocation. [Diagram 2] FIG. 1 illustrates a schematic diagram of a system in which techniques of the present disclosure for improving memory allocation can be used. [Diagram 3] FIG. 1 illustrates a schematic diagram of a memory sanitizer that can use techniques of the present disclosure to improve memory allocation. [Figure 4] FIG. 1 illustrates a schematic diagram of a system in which techniques of the present disclosure for improving memory allocation can be used. [Diagram 5] 1 is a flow chart illustrating the techniques of this disclosure for training a language model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0034] Figure 1 is a flow chart illustrating a method 10 for verifying static warnings in generated code using a language model. According to Figure 1, the generated code is generated in an automatic program code translation from a source language to a target language. Alternatively, the program code may be generated by code refactoring.
[0035] The method 10 proposed in the present disclosure is directed to improving memory allocation of software code generated by a language model. The software may be configured to control, regulate and / or monitor at least one computing unit of a technical system, in particular a cyber-physical system, in particular a vehicle. In particular, the software may be embedded software configured to run on an embedded (i.e., e.g. task-specific) system.
[0036] In a first step, program code generated by a language model is provided 11. This providing includes both the mere input of the program code into the method or system, and the generation or incorporation of a language model into the method or system.
[0037] When a new version of the program code is available, compilation and instrumentation produces an executable file12, for example by a memory sanitizer, which inserts instructions into the program code at generation time or into the executable at run time. Sanitizers are best inserted into an intermediate representation to support many different languages. In memory sanitizers, a "shadow memory" is used to initialize and fill variables with "normal" instructions and information that it is shadow memory. And there are further instructions that do the trick. It is also possible to track when variables of what memory size are allocated and when they are freed. This allows you to track memory usage over time. But you can also look at high allocation counts.
[0038] Optionally, the corpus is provided with a corpus of inputs for the fuzzer, including initial test cases from a code repository of the program code and / or initial test cases from a provided test or test harness. The corpus may be initially populated and further populated during the method.
[0039] After this, fuzzing is performed, where the fuzzer injects inputs into the executable, with the goal of getting the best possible coverage of the file or code with the inputs injected by the fuzzer.
[0040] Subsequently, memory performance (e.g., memory usage, memory occupancy, etc.), optionally runtime information (e.g., actual runtime per test case), behavior and / or output of the executable are monitored.14 During monitoring, the executable to be fuzzed and the memory metadata database may be monitored. It may be contemplated that the memory performance, runtime information, behavior and / or output of the executable are fed back to the fuzzer.
[0041] Subsequently, metadata generated from the allocated and freed memory is stored in a memory metadata database15, where the metadata is stored during the generation of the executable and / or during the fuzzing based on the instructions. In this way, metadata can be generated from the allocated and freed memory of all different versions of the program code and stored in the memory metadata database, which allows for the aggregation and comparison of memory behavior between different versions of the program code.
[0042] Subsequently, if no memory performance degradation or other errors are found, output of the program code is performed.16 It may be contemplated that an abnormal termination of the method may be triggered during or after fuzzing if the memory performance is lower than an older entry from the memory metadata database.
[0043] Optionally, static testing can be performed in parallel with the execution of fuzzing or memory sanitizer. Static tests or checks can include, for example: - Contracts are either configured or extracted from the system environment.
[0044] - In Rust, static checking may be provided by the compiler, in other languages it may be provided by, for example, a linter. - Bounded model checking and / or abstract interpretation, achieved with commercial tools such as Astree, or with open source tools such as CBMC (Bounded Model Checker for C programs).
[0045] - Automatic construction of bounded model checking setups that statically compare target program code against a given contract. - Automatic construction of abstract interpretation setups that statically compare target program code against a given contract.
[0046] - Automatic construction of bounded model checking setups that check the functional equivalence of source and target program code. According to one embodiment, the method further comprises updating the program code or a portion of the program code using the memory performance, the runtime information, the behavior of the program code and / or the output of the program code, and optionally feeding back the updated program code as input to the language model.
[0047] As an alternative to the described translation, a refactoring of the program code may be envisaged. The refactoring of the program code may comprise a modification of the program code. The refactored program code may be software code, in particular software source code.
[0048] Fig. 2 shows a schematic representation of a computer system 20 that can use the techniques of the present disclosure for improving memory allocation of code generated by a language model. The computer system 20 is configured to execute the method 10 according to Fig. 1 and the method 50 according to Fig. 5. The computer system 20 can be realized in hardware and / or software. Thus, the system shown in Fig. 2 can be considered as a computer program configured to execute the method 10 according to Fig. 1 and the method 50 according to Fig. 5.
[0049] Source program code 21 in a source language, e.g., C, is provided to a language model 22, e.g., a large-scale language model (LLM), for translation into a target language, e.g., Rust. The language model 22 generates (target) program code 23 as a translation of the source program code 21. This domain of the computer system 20 may be referred to as the generation domain.
[0050] The language model is, for example, a Large Language Model (LLM), into which data, such as program code, and questions, such as translation or refactoring requests, are input via inputs (prompts).
[0051] Further inputs 24 to system 20 are tests in the source language and, optionally, a test harness. Alternatively or optionally, tests and / or test harnesses may be available in the target language. These inputs are provided to test harness 25. Test harness 25 records functions or tests in the target language.
[0052] Optionally, the static tests, quality assessments and / or contracts 26 may be provided to a static testing unit 27. These are managed for later checking of the program code 23.
[0053] The inputs of memory sanitizer with fuzzer 28 are connected to language model 22 for inputting program code 23 and to test harness 25 for inputting test routines. In memory sanitizer 28, program code 23 is tested against source program code 21 based on test harness 25. Memory sanitizer 28 and its functionality are described in more detail in conjunction with FIG.
[0054] The generation of program code 23 by the language model 22 may be performed iteratively under changed conditions, such as changing one or more hyperparameters of the language model, such as a temperature parameter, transforming the source program code, and / or changing the input to the language model, e.g., changing the task or prompt. Additionally, variables in the code may be renamed.
[0055] These processes generate variances that allow for inspection of the generated translations, improved quality assessment, and feedback to train the language model. As part of the improved quality assessment, it is possible to determine which of the generated program codes 23 are more suitable or less suitable. The fuzzer 28 then operates on these variants of the program code 23.
[0056] The inputs of the checking unit 29 are connected to the language model 22 for inputting the program code 23 and to a static testing unit 27 for inputting the static tests, quality assessments and / or contracts 26. In the checking unit 29, the program code 21 is checked by the static tests, quality assessments and / or contracts 26.
[0057] Upon successful completion of the tests in memory sanitizer 28 and test unit 29, a status message 30 is output indicating that program code 23 is OK. This area of computer system 20 is sometimes referred to as the test area.
[0058] The target program code 31 may have its quality evaluated using metrics 32. The metrics 32 may include code quality metrics, test quality metrics, and / or test counts. If the evaluation is successful, the quality and grade are output as program code output 33. This area of the computer system 20 may be referred to as the quality evaluation area.
[0059] A grade may be calculated based on the evaluation, and if multiple target program codes are generated, the solutions with the highest grade may be provided to the user. The program code may be evaluated using code quality metrics such as, for example, source code length, number of loops and / or branch depth, and test quality metrics such as, for example, coverage (branch coverage) and / or number of available or implemented tests.
[0060] FIG. 3 shows details of memory sanitizer 28 and test harness 25 and program code 23 as inputs to memory sanitizer 28. The memory sanitizer 28 includes the following parts for checking and monitoring memory performance: (Optional) A corpus 40 is populated with initial test cases from the code repository (translated or refactored) and / or from the provided tests and test harness. For this purpose, the corpus 40 is connected to the test harness 25.
[0061] A compilation and instrumentation unit 41 is provided for generating an executable file 43 from the program code 23. For this purpose, the unit 41 receives the program code 23 as input and is connected to the test harness 25 for compilation and instrumentation.
[0062] A fuzzer 42 is provided for generating and injecting inputs into executable file 43. For inputs, fuzzer 42 is connected to corpus 40 and / or test harness 25.
[0063] A memory metadata database 44 is provided for storing metadata. The memory metadata database 44 is generated or initially filled with information from the compilation and instrumentation unit 41. At runtime, the memory metadata database 44 is filled with runtime information from the allocated and freed memory of the executable file 43. For this purpose, the memory metadata database 44 is connected to the compilation and instrumentation unit 41 and the executable file 43. The memory metadata database 44 can store metadata from the allocated and freed memory of all different versions of the executable file 43, so that memory behavior can be seen across different versions.
[0064] The monitoring unit 45 measures or monitors memory performance and optionally runtime information, behavior and / or output of the executable. Code coverage at runtime can be monitored, typically in a grey-box setting. This can be stored in the monitoring unit 45. All this is fed back to the fuzzer 42 to generate better test cases. Additionally, the monitoring unit 45 is connected to a memory metadata database 44 to obtain information about memory metadata and include it in the monitoring. This can include, for example, comparison of memory data such as memory performance, memory occupancy, etc.
[0065] Fig. 4 shows a schematic representation of a computer system 20 that can use the techniques of the present disclosure for improving memory allocation. The computer system 20 can correspond to the computer system 20 of Fig. 2. The computer system 20 is configured to execute the method 10 according to Fig. 1 and the method 50 according to Fig. 5, and in particular the computer system 20 of Fig. 3 is configured to execute the training method 50 according to Fig. 5. The computer system 20 can be realized in hardware and / or software. The system shown in Fig. 3 can therefore be considered as a computer program configured to execute the method 10 according to Fig. 1 and the method 50 according to Fig. 5.
[0066] 4 shows the error handling mechanism of the computer system 20. If it is not possible to generate an error-free program code 31, an error is generated in the memory sanitizer 28. Thus, a message 34 is output that reduces the reliability of the translation. Furthermore, the error and the associated program code 31 are stored in an error module 35.
[0067] From the error module 35, the best program code so far, including the remaining errors, is fed back as information 36 to the language model 22 in order to generate a better, ideally error-free, target program code. This reduces the reliability and is taken into account in the quality decision.
[0068] Optionally, this may refer not only to errors in the inspection area, but also to errors in the quality assessment area. 5 is a flow chart illustrating a method 50 for training a language model. A language model is established for the automatic generation of program code.
[0069] In summary, the code generated by the language model is fed back to the language model as feedback to generate a new generation of source code. The idea is that already good (or better) code can be fine-tuned with updated input requirements, including new code, in terms of monitoring and behavioral output. For example, if already good code has runtime problems only in a specific place, that place can be fed back to the language model.
[0070] In a first step of the method, the source program code or source code is input 51 into a language model to generate a program code. The language model may already be pre-trained and possibly already fine-tuned. Alternatively, one may start with a new, untrained language model. The training in this case is based on reinforcement learning. The training is performed in a training environment, for example using PPO (Proximal Policy Optimisation).
[0071] Further, the program code of the predicted target program code is improved 52 using the method 10 described above with reference to FIG. A reward is then generated 53 for the language model, the reward being based on memory performance and optionally runtime information, monitoring 14 of the behavior and / or output of the executable 43, and / or metadata. Thus, a low reward may be given if memory performance is low and / or if a lot of memory is used or if not much memory is freed. A high reward may be given if memory performance is high and / or if a lot of memory is used or if a lot of memory is freed.
[0072] Finally, the weights of the language model are updated with the reward value 54. The result of the method is a language model that is better trained on new unlabeled data (here, for example, C code from the engine control), i.e., a language model that provides more reliable translations.
[0073] According to one embodiment, the method further includes approximating the reward by performing only a subset of the tests of the automated testing, which can accelerate training.
Claims
1. A method (10) for improving memory allocation of code generated using a language model (22), comprising the steps of: - providing (11) a program code (23) generated by using a language model (22) when a new version of the program code (23) is available; - generating (12) an executable file (43) by compiling and instrumenting, during which a memory sanitizer (28) inserts instructions into said executable file (43); - a step (13) of performing fuzzing by a fuzzer (42), said fuzzer (42) injecting inputs into said executable file (43); - monitoring (14) the memory performance, and optionally the runtime information, behavior and / or output of said executable file (43); - storing (15) metadata generated from allocated and freed memory in a memory metadata database (44), said metadata being stored during the generation (12) of said executable file (43) and / or during the execution of fuzzing (13) based on said instructions; - outputting (16) said program code (23) if no degradation of memory performance or other errors are found; A method comprising:
2. The method (10) of claim 1 , wherein instructions are inserted into an intermediate representation of the executable file (43).
3. 3. The method of claim 1 or 2, wherein a corpus of inputs for the fuzzer is provided, the corpus including initial test cases from a code repository of the program code and / or initial test cases from a provided test or test harness.
4. The method (10) of any one of claims 1 to 3, wherein an abnormal termination of the method is triggered during or after a fuzzing execution (13) if the memory performance is lower than an older entry from the memory metadata database (44).
5. The method (10) according to any one of claims 1 to 4, wherein for said monitoring (14), said executable file (43) to be fuzzed and said memory metadata database (44) are monitored.
6. The method (10) of any one of claims 1 to 5, wherein the memory performance, the runtime information, the behavior of the executable file (43) and / or the output of the executable file (43) are fed back to the fuzzer (42).
7. The method (10) according to any one of claims 1 to 6, wherein the memory performance, the runtime information, the behavior of the program code (23) and / or the output of the program code (23) are used to update the program code (23) or parts of the program code (23), and the updated program code is optionally fed back as input to the language model (22).
8. A method (50) for training a language model (22) configured for the automatic generation of program code (23, 31), comprising: - inputting (51) the source code into a language model (22) to generate a program code (23); - improving (52) said program code (23) using a method (10) according to any one of claims 1 to 7; - generating (53) a reward for said language model (22), said reward being based on memory performance, optionally runtime information, monitoring of the behavior and / or output of the executable file (43), and / or metadata; - updating (54) the weights of the language model with the reward values; A method comprising:
9. The method (50) of claim 8, wherein the reward is approximated by performing validation only.
10. A computer system (20) configured to carry out the method (10; 50) according to any one of claims 1 to 9.
11. A computer program adapted to carry out the method (10; 50) according to any one of claims 1 to 9.
12. A computer readable medium or signal storing and / or including a computer program according to claim 11.