Code generation optimization method based on adaptive planning framework
Through a multi-agent framework based on an adaptive planning framework, combined with the try-except code block and the HumanEval benchmarking module, the problem of the inability to compatible with open source models and code generation is solved, and unified access and efficient code generation of different types of large models are achieved.
Patent Information
- Application Number
- CN202510261501.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Existing code generation methods are not compatible with open source models, and it takes too long to generate code.
A multi-agent framework based on an adaptive planning framework is adopted, including four modules: coder, evaluator, debugger and planner. The evaluation system is built through the try-except code block and HumanEval benchmark test module, which is compatible with different types of large models, and significantly improves the testing efficiency through the two-stage mechanism of compilation testing and functional verification.
It realizes unified access to open source and closed source models, significantly reducing the time cost of code generation, and improving the accuracy and efficiency of code generation.
Smart Images

Figure CN120085846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer code optimization, and particularly to a code generation optimization method based on an adaptive planning framework. Background Art
[0002] In recent years, text generation technologies based on large language models (referred to as large models for short), such as models like DeepSeek, have been rapidly popularized. These models can understand the questions input by users and automatically generate corresponding answers, thus significantly improving the efficiency and accuracy of users in obtaining information and solving problems. Among numerous text generation tasks, code generation is a particularly important application. Code generation refers to the process in which a large model automatically generates a code solution for a corresponding problem according to the task requirements input by the user. However, different from the flexibility of natural language, code has strict structures and syntax rules and must be written according to precise rules to ensure that the computer can execute it correctly. This high degree of precision makes the code generation task more challenging and also poses higher requirements for the accuracy and efficiency of the generated results.
[0003] Some existing large models can generate thousands of lines of code within the time it takes developers to write a few lines of code, thus significantly reducing the coding workload and time cost of developers. By automatically generating code, this technology not only reduces the need for manual coding but also reduces the defects caused by human errors, while improving the consistency and maintainability of the code. However, in practical applications, the code generated by many existing large models still contains one or several errors, which not only prevents developers from directly using this code to improve development efficiency but also requires additional time to find and fix the hidden errors. This problem greatly limits the practical application value of code generation technology. To address this problem, researchers have proposed a class of methods called "multi-agent frameworks". Such frameworks decompose complex code generation tasks into multiple subtasks (such as logical verification, performance optimization, security review, etc.) and design multiple agents, with each agent responsible for handling a specific subtask; through division of labor and cooperation, the multi-agent framework can simplify complex tasks and thus improve the quality of code generation.
[0004] However, existing multi-agent frameworks have obvious limitations. First of all, they are usually only applicable to closed-source large models (such as ChatGPT) and are basically incompatible with open-source models; because closed-source models do not disclose their model itself and its implementation principles, while open-source models are the opposite. For open-source models, the improvement effect of these frameworks is extremely limited, and may even reduce the code generation quality, so the scope of application is greatly restricted. Secondly, applying these multi-agent frameworks will significantly increase the time required for code generation, and the time cost often increases by dozens of times, further restricting their practical application value. Therefore, there is an urgent need to develop a new type of multi-agent framework that can be applicable to all types of large models (including open-source and closed-source models), while having lower time consumption and better code generation effects. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the technical problem to be solved by the present invention is that the existing code generation methods cannot be compatible with open-source models, and the time-consuming for generating code is too long.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] An optimization method for code generation based on an adaptive planning framework, comprising the following steps:
[0008] S100: Construct a multi-agent framework model M, where M includes four modules: an encoder, an evaluator, a debugger, and a planner;
[0009] The encoder and the planner adopt existing large language models;
[0010] The evaluator is constructed by adopting a try-except code module and a benchmark test module, where several sample test cases are included in the benchmark test module;
[0011] The large language model, the try-except code module, the benchmark test module, and the Python script module are all prior arts. The large language models include CodeLlama-Python, DeepSeek-Coder, GPT-3.5-turbo, GPT-4, GPT-4-turbo, GPT-4o, etc.;
[0012] The debugger includes a debug database and a Python script module, and the debug database contains the names of all modules and the names of internal functions of the modules in the Python standard database;
[0013] The names of all modules in the Python standard database and the names of internal functions of the modules are divided into six major categories: system interaction and I / O management (including os, sys, shutil modules, covering more than 200 functions such as file reading and writing, path operations, etc.), data structures and algorithms (integrating 30 data structure operation methods of collections and itertools modules), numerical calculation and scientific analysis (including more than 450 mathematical functions of math and statistics modules and numpy compatibility interfaces), network and concurrent programming (integrating communication protocols and parallel computing tools of socket, asyncio, threading modules), text processing and regular expression engine (built-in 80 string operation and pattern matching methods of re and string modules), data serialization and persistence (supporting 12 data format conversion protocols such as json and pickle);
[0014] S200: The user edits a task description x and presets the maximum number of loops, then takes x as the input of the coder, and the output is the code solution C of x 1 ;
[0015] S300: Input C 1 to the evaluator for testing, and the output is the test result T 1 , where the T 1 includes the evaluation result R 1 of the evaluator on C 1 , the type E T that causes the error, and the error message E M ;
[0016] If R 1 is passed, the evaluator directly takes C 1 as the final code solution of x, and the process ends; passing the test indicates that C 1 can already be directly used by the user; otherwise, the evaluator will output T 1 , and execute the next step;
[0017] S400: Take C 1 and T 1 as the input of the debugger. The debugger analyzes E T and E M to fix the simple errors that occur in C 1 and obtains the repaired code solution C 2 ; The simple errors are code format, function reference, database reference, etc.;
[0018] S500: Input C 2 to the evaluator for testing, and the output is the test result T 2 , where the T2 including the evaluator's evaluation result R 2 for C 2 , the type E' of the error cause T , and the error message E' M ;
[0019] If R 2 is passed, the evaluator directly takes the output C 2 as the final code solution for x, and the process ends; otherwise, the evaluator outputs T 2 , and proceeds to the next step;
[0020] S600: Input both x and the E' 2 in T T and E' M into the planner, and then the planner outputs a step-by-step plan P;
[0021] S700: Input x and P again as the coder, and output a new code solution C' 1 , and detect whether the evaluation of C' 1 by the evaluator reaches the maximum number of loops. If it reaches the maximum number of loops, then C' 1 is the final code solution and is output, and the process ends; otherwise, let C' 1 = C 1 and return to S300.
[0022] Preferably, the benchmark test module in S100 includes the HumanEval method and the MBPP method. The HumanEval benchmark is widely adopted and more universal. Using this benchmark test can achieve an improvement in cross-model compatibility; traditional frameworks rely on model-specific output formats (such as the Markdown code blocks of ChatGPT), while the input-output assertion mechanism of HumanEval can adapt to the original output of any model.
[0023] Preferably, the specific steps of inputting C 1 into the evaluator for testing in S300 are as follows:
[0024] S310: The evaluator first embeds C 1 in a try-except code block for compilation. If no exception is caught during the compilation process, it indicates that the compilation is successful, and proceed to the next step; if an exception is caught during the compilation process, then the caught exception is output as T 1 ;
[0025] S320: Input the successfully compiled C 1 in S310 into the benchmark test module;
[0026] S330: Append any sample test case y in the benchmark test module to the end of the successfully compiled C 1 to obtain the recompiled C 1 ; Execute the test in another try-except code block. If no exception is caught during the execution of the other try-except code block, it indicates that the original C 1 passes the test, that is, the original C 1 can be used as the final code solution for x; 1
[0027] If an exception is caught during the execution of the other try-except code block, the exception will be temporarily stored as the test result T y corresponding to y;
[0028] S340: Repeat S330 to iterate through all sample test cases in the benchmark test module to obtain the test results T all corresponding to all sample test cases. After that, output T all as T 1 .
[0029] The try-except test mechanism decouples code compilation testing (try-except code block) and function verification (benchmark test) into independent stages. Through exception capture, the type of error can be accurately located; moreover, traditional frameworks need to execute the code completely to discover logical errors, while this method can intercept syntax errors (such as indentation errors, undefined variables) at the compilation stage, avoiding ineffective tests; it can significantly improve the test efficiency and reduce the test time cost.
[0030] Preferably, the specific steps to obtain the repaired code solution C 2 in S400 are as follows:
[0031] S410: Code filtering: Split C 1 by line and store it in a cache list, then check the indentation of each line and perform normalization processing to obtain C 11 ; Usually, the indentation should be four spaces. If the indentation of a certain line does not meet this standard, the debugger will normalize the indentation to solve the inconsistent indentation problem. The so-called normalization refers to the standard format requirements of computer coding.
[0032] S420: Code truncation: Compile the code of C 11 . If the compilation fails, split C 11 by line and store it in a list, then remove the last line and compile it again. Repeat this process until the code is successfully compiled or only one function remains, then stop compiling to obtain C 12 ;
[0033] S430: Missing module injection: Detect C 1 Corresponding E T Whether it is a NameError. If it is not a NameError, then use C 12 As the repaired code solution C 2 ;
[0034] If it is a NameError, extract the database name that causes the NameError through regular expressions, and match the database name with a pre-built debugging database. If the match is successful, execute the next step; otherwise, use C 12 As the repaired code solution C 2 ;
[0035] S440: Import the successfully matched database statement into C 12 After that, obtain C 13 At this time, C 13 Is the repaired code solution output by the debugger C 2 .
[0036] The debugger receives the test information provided by the evaluator. If the error type provided by the evaluator is NameError, it indicates that the code may have used unimported modules or functions. In this step, the name that causes the NameError (such as math, re, functools) is extracted through regular expressions and attempts to match it with a pre-built debugging database that contains all common module names and their internal function names; if the match is successful, the corresponding import statement (such as import math) is inserted; otherwise, it indicates that the name is only an undefined variable, rather than a module or library function.
[0037] The debugger adopts this method through a modular repair process based on scripts and rules, which can make the repair faster (with an average repair time of ten milliseconds per repair); moreover, by building a progressive repair workflow through a hierarchical processing strategy (code filtering → truncation → module injection), the repair will also be more accurate.
[0038] Compared with the prior art, the present invention has at least the following advantages:
[0039] 1. Through the collaborative design of the standardized evaluation module and the pre-built debugging database, this technical solution completely breaks through the interface dependence of the traditional framework on closed-source models. Specifically, the evaluation system constructed by using the try-except code block and the HumanEval benchmark test can directly parse the original output of open-source models (such as the unformatted code of DeepSeek-Coder), and at the same time be compatible with the API response format of closed-source models, realizing the unified access of all types of large models.
[0040] 2. This technical solution can greatly improve time efficiency. The "compilation test - functional test" two-stage mechanism designed by the present invention and the modular repair workflow form a synergistic effect, which can intercept low-level errors (such as Python indentation exceptions) at the compilation stage. Combined with the code truncation strategy, the single iteration time is compressed from dozens of seconds in the traditional solution to 10 milliseconds. Even if 5 complete loops are executed, the total time consumption is still less than 2 minutes, greatly enhancing the practical value of the framework.
[0041] 3. In terms of dealing with complex problems, the dynamic analysis mechanism constructed by the planner module shows strong advantages. This module identifies deep logical defects by parsing the error chain (E T →E′ T ) and generates a step-by-step plan P with priorities, which makes the passing rate of the test of this solution applied to open-source models increase by an average of 69.80%; among them, the Pass@1 index of DeepSeek-Coder-33B on the HumanEval dataset reaches 90.85%, significantly narrowing the performance gap with the closed-source model GPT-4 (96.95%). The synergistic effect of these technological breakthroughs enables this framework to comprehensively surpass the existing solutions on the HumanEval and MBPP datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the overall framework and flowchart of the present invention;
[0043] Figure 2 is the accuracy and time performance of the existing method on the HumanEval dataset;
[0044] Figure 3 is the accuracy and time performance of the existing method on the MBPP dataset;
[0045] Figure 4 is the accuracy and time performance of the method of the present invention on the HumanEval and MBPP datasets. DETAILED DESCRIPTION OF THE INVENTION
[0046] The present invention will be further described in detail below.
[0047] The present invention adopts an adaptive strategy: when initially generating code, a non-planning mechanism is used: the coder directly generates code based on the task description input by the user, and then repairs simple errors (such as inconsistent indentation, missing imports, etc.) through a rule-based method; this method can quickly handle common problems and avoid the additional overhead brought by the planning mechanism, and only enables the planning mechanism when necessary: when there are still errors in the code after the simple errors are repaired, the planner disassembles the task description, formulates a step-by-step plan, and provides it to the coder together with the task description again, and the coder regenerates the code according to this plan, thereby further improving the accuracy of the code; through this adaptive strategy, the present invention not only saves the time required for code generation, but also improves the quality of the generated code. In addition, due to the overall process being simple and clear, this framework is not only applicable to closed-source large models, but also can be widely applied to open-source models, significantly expanding its scope of application.
[0048] See Figures 1-4 , an optimized method for code generation based on an adaptive planning framework, comprising the following steps:
[0049] S100: Construct a multi-agent framework model M, where M includes four modules: a coder, an evaluator, a debugger, and a planner;
[0050] The coder and the planner adopt existing large language models;
[0051] The evaluator is constructed by adopting a try-except code module and a benchmark test module, wherein the benchmark test module contains several sample test cases;
[0052] The debugger includes a debug database and a Python script module, and the debug database contains the names of all modules and the names of internal functions of the modules in the Python standard database;
[0053] The benchmark test module in S100 includes the HumanEval method and the MBPP method.
[0054] S200: The user edits a task description x and presets a maximum number of loops, and then takes x as the input of the coder, and the output is the code solution C of x 1 ;
[0055] S300: Input C 1 into the evaluator for testing, and output the test result T 1 , where T 1 includes the evaluation result R of the evaluator on C 1 , the type E of the error caused 1 , and the error message E T ; M ;
[0056] If R 1 passes, the evaluator directly takes C 1 as the final code solution for x, and the process ends; otherwise, the evaluator outputs T 1 , and proceeds to the next step;
[0057] In S300, the specific steps for inputting C 1 into the evaluator for testing are as follows:
[0058] S310: The evaluator first embeds C 1 into a try-except code block for compilation. If no exception is caught during the compilation process, it indicates successful compilation, and the next step is executed; if an exception is caught during the compilation process, the caught exception is taken as T 1 for output;
[0059] S320: Input the successfully compiled C 1 from S310 into the benchmark testing module;
[0060] S330: Append any sample test case y in the benchmark testing module to the end of the successfully compiled C 1 to obtain the appended and compiled C 1 ; Execute the test for the appended and compiled C 1 in another try-except code block. If no exception is caught during the execution process of the other try-except code block, it indicates that the original C 1 passes the test, that is, the original C 1 can be used as the final code solution for x;
[0061] If an exception is caught during the execution process of the other try-except code block, the exception will be temporarily stored as the test result T y corresponding to y;
[0062] S340: Repeat S330 to traverse all sample test cases in the benchmark testing module to obtain the test results T all corresponding to all sample test cases. After that, take T all as T 1 for output.
[0063] S400: Take C 1 and T 1 as the input for the debugger. The debugger analyzes E T and E M to fix the simple errors occurring in C 1 to obtain the fixed code solution C 2 ;
[0064] The repaired code solution C obtained in S400 2 The specific steps are as follows:
[0065] S410: Code filtering: Split C 1 line by line and store it in a cache list, then check the indentation of each line and perform normalization to obtain C 11 ;
[0066] S420: Code truncation: Compile the code of C 11 . If the compilation fails, split C 11 line by line and store it in a list, then remove the last line and compile it again. Repeat this loop until the code is successfully compiled or only one function remains, then stop compiling to obtain C 12 ;
[0067] S430: Missing module injection: Detect whether the corresponding E 1 of C T is NameError. If it is not NameError, then use C 12 as the repaired code solution C 2 ;
[0068] If it is NameError, extract the database name that causes NameError through regular expressions, and match the database name with a pre-built debug database. If the match is successful, execute the next step; otherwise, use C 12 as the repaired code solution C 2 ;
[0069] S440: Import the successfully matched database statement into C 12 to obtain C 13 . At this time, C 13 is the repaired code solution C output by the debugger 2 .
[0070] S500: Input C 2 to the evaluator for testing, and output the test result T 2 . The T 2 includes the evaluation result R 2 of the evaluator on C 2 , the type of error E' T , and the error message E' M ;
[0071] If R 2 is passed, the evaluator directly uses the output C 2 as the final code solution of x, and the process ends; otherwise, the evaluator outputs T 2, and perform the next step;
[0072] S600: Input x and T 2 the E' in T and E' M into the planner, and then the planner outputs a step-by-step plan P;
[0073] S700: Input x and P into the coder again, and output a new code solution C' 1 , and detect whether the tester's test on C' 1 reaches the maximum number of loop iterations. If it reaches the maximum number of loop iterations, then C' 1 is the final code solution and output, and the process ends; otherwise, let C' 1 = C 1 and return to S300.
[0074] In the initial code generation of the present invention, a non-planning mechanism is preferentially adopted to ensure high efficiency and wide applicability; only when the code generated initially fails the test, will the planning mechanism be enabled to regenerate the code, so as to better cope with diverse task requirements.
[0075] Experimental content and results
[0076] The following proves this effect through a comparative experiment between the method of the present invention and the existing method.
[0077] Two datasets are selected for the comparative experiment: the HumanEval dataset created and published by OpenAI, and the MBPP dataset created and published by Google. The former contains 164 programming problem samples of various difficulty levels, and the latter contains about 500 programming problems, mainly for junior to intermediate programmers, covering basic Python programming tasks. These problems cover a wide range of topics and programming concepts, aiming to comprehensively evaluate the programming ability of the model. Each sample in the two datasets contains five parts: sample number, problem description, reference solution, test case, and entry function name. Among them, the problem description, test case, and entry function name of each sample are the parts we need to use in the test process.
[0078] The evaluation metric is accuracy Pass@k, and the accuracy is calculated based on the formula . Among them, Pass@k is the metric we want to calculate, representing the probability that the model solves at least one problem within k attempts. Here, we use the Pass@1 result when k = 1 as the evaluation metric, represents the average value of all problems, meaning that we are calculating the average performance on multiple problems. n is the total number of problems in the dataset, and c is the number of problems that the model can correctly solve. is the number of combinations of choosing \(k\) questions from \(n\) questions, representing the total number of combinations tried. is the number of combinations of choosing \(k\) questions from the questions that the model cannot solve (i.e., \(n - c\) questions), representing the number of failed combinations.
[0079] To ensure the validity of the experiment, we selected 10 large language models (LLMs) with different numbers of parameters that are currently popular and mainstream for the experiment. These models include:
[0080] 1) CodeLlama - Python: An open - source model released by Meta AI, providing three versions with 7B, 13B, and 34B parameters;
[0081] 2) DeepSeek - Coder: An open - source model released by DeepSeek, providing three versions with 1.3B, 6.7B, and 33B parameters;
[0082] 3) GPT series models: Closed - source models released by OpenAI, including GPT - 3.5 - turbo, GPT - 4, GPT - 4 - turbo, and GPT - 4o, all of which have more than 100B parameters.
[0083] In addition, we also selected four existing methods for comparison with the present invention, which are:
[0084] 1) AgentCoder: Proposed by Huang et al., containing three agents;
[0085] 2) MapCoder: Proposed by Islam et al., containing four agents;
[0086] 3) INTERVENOR: Proposed by Wang et al., containing two agents;
[0087] 4) Self - Collaboration: Proposed by Dong et al., containing three agents.
[0088] For the open - source models, we tested the performance of the existing methods and the method of the present invention on the HumanEval and MBPP datasets respectively, and counted the accuracy and time used for code generation. For the closed - source models, since they can only be called through the official OpenAI interface and cannot be locally deployed, it is difficult to record the specific time consumption. For these models, we selected to test using the method of the present invention on the HumanEval and MBPP datasets and counted the accuracy.
[0089] The experimental results are as Figure 2 、 Figure 3 and Figure 4 shown, and the specific analysis is as follows:
[0090] Figure 2 and Figure 3 It shows that the performance of existing methods is unstable and shows obvious inapplicability to open-source models: on the HumanEval dataset, even MapCoder, which shows the most obvious improvement, only improves the baseline performance by 38.86%. On the contrary, the other two existing methods, AgentCoder and Self-Collaboration, even lead to a performance decline, which reflects the poor applicability of these methods to open-source models; in addition, these methods all significantly increase the time required for code generation, and the time cost increases several times to more than ten times; on the MBPP dataset, the performance of existing methods is even worse; the highest performance improvement is only 6.77%, while the lowest leads to a performance decline of up to 18.46%. At the same time, the time cost also increases several times to more than ten times.
[0091] Figure 4 It demonstrates the superiority of the method of the present invention: the results show that whether on the HumanEval or MBPP dataset, whether for open-source models or closed-source models, the method of the present invention shows significant performance advantages. On the HumanEval dataset, the average performance improvement reaches 54.58%; on the MBPP dataset, the average performance improvement reaches 45.41%; more importantly, the method of the present invention does not significantly increase the time cost. On the HumanEval dataset, the time cost only increases by 0.68 times; on the MBPP dataset, the time cost only increases by 1.34 times, which indicates that the method of the present invention achieves a good balance between performance improvement and time cost control.
[0092] In short, the present invention proposes an optimization method for code generation based on an adaptive planning framework. On the one hand, it enhances the applicability of the current multi-agent framework to different large models, and on the other hand, it significantly reduces the time cost of applying the multi-agent framework, and achieves performance beyond the existing multi-agent framework; the performance of the present invention is superior to existing methods and can be applied to actual development scenarios to contribute to improving the development efficiency of developers.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A code generation optimization method based on an adaptive planning framework, characterized in that: The steps include: S100: construct a multi-agent framework model M, wherein M includes four modules: an encoder, an evaluator, a debugger, and a planner; Coders and planners use existing large language models; The evaluator is constructed using a try-except code module and a benchmark module, where each benchmark module contains several sample test cases; The debugger includes a debugging database and a Python script module, wherein the debugging database contains the names of all modules in the Python standard database and the names of the internal functions of the modules; S200: The user edits a task description x and presets the maximum number of loops, then uses x as the input of the encoder, and the output is the code solution C1 of x; S300: Input C1 to the evaluator for testing, and output the test result T1, wherein T1 includes the evaluator's evaluation result R1 of C1, the type of error E1, and the error-causing T , and error message E M ; If R1 is passed, the evaluator directly uses C1 as the final code solution for x and the process ends; otherwise, the evaluator outputs T1 and executes the next step; S400: C1 and T1 are used as inputs by the debugger. The debugger analyzes E T and E M Fix the simple error that occurred in C1 and get the fixed code solution C2; S500: Input C2 to the evaluator for testing, and output the test result T2, which includes the evaluator's evaluation result R2 on C2, the type of error E′ T , and error message E′ M ; If R2 is passed, the evaluator directly outputs C2 as the final code solution for x, and the process ends; otherwise, the evaluator outputs T2 and executes the next step; S600: x and E′ in T2 T and E′ M All are input into the planner, and then the planner outputs a step-by-step plan P; S700: x and P are used as encoder input again, and a new code solution C′1 is output. The evaluator's test on C′1 is checked to see if it reaches the maximum number of cycles. If it reaches the maximum number of cycles, C′1 is the final code solution and is output, and the process ends; otherwise, C′1=C1 is set and returns to S300.
2. A code generation optimization method based on an adaptive programming framework as claimed in claim 1, characterized in that: The benchmark test module in S100 includes the HumanEval method and the MBPP method.
3. A code generation optimization method based on an adaptive programming framework as claimed in claim 2, characterized in that: The specific steps of inputting C1 into the evaluator for testing in S300 are as follows: S310: The evaluator first embeds C1 into a try-except code block for compilation. If no exception is caught during the compilation process, the compilation is successful and the next step is executed. If an exception is caught during the compilation process, the caught exception is output as T1. S320: input C1 compiled successfully in S310 into the benchmark test module; S330: append any sample test case y in the benchmark test module to the end of the successfully compiled C1 to obtain the appended compiled C1; execute the test on the appended compiled C1 in another try-except code block. If the execution process of the other try-except code block does not catch the exception, it indicates that the original C1 passes the test, that is, the original C1 can be used as the final code solution for x; If another try-except code block catches an exception during execution, the exception will be used as the test result T corresponding to y y be temporarily stored; S340: Repeat S330 to traverse all sample test cases in the benchmark test module and obtain the test results T corresponding to all sample test cases. all After that, T all Output as T1.
4. A code generation optimization method based on an adaptive programming framework as claimed in claim 3, characterized in that: The specific steps of obtaining the repaired code solution C2 in S400 are as follows: S410: Code filtering: Split C1 by line and store it in a cache list, then check the indentation of each line and perform normalization to obtain C 11 ; S420: Code truncation: C 11 Compile the code. If the compilation fails, C 11 The code is split into lines and stored in a list. The last line is then removed and compiled again. This cycle repeats until the code is successfully compiled or only one function is left. The compilation is stopped to get C. 12 ; S430: Missing module injection: Detect E corresponding to C1 T Is it a NameError? If not, C 12 As the fixed code solution C2; If it is a NameError, the database name that caused the NameError is extracted through a regular expression, and the database name is matched with the pre-built debugging database. If the match is successful, the next step is executed; otherwise, C 12 As the fixed code solution C2; S440: Import the successfully matched database statements into C 12 Then get C 13 , at this time C 13 This is the fixed code solution C2 output by the debugger.
Citation Information
Patent Citations
Deep neural network automatic generation method based on grammar rules
CN114265581A
Software engineering development system and method based on Internet
CN115794038A
Code generation and defect repair method and device
CN116909532A
Code generation optimization method and device for large language model, equipment and medium
CN117724695A
Code generation optimization method based on code standardization
CN118151943A