Interactive code generation method and system and storage medium
Through user natural language feedback optimization of the code generation model, the problem of inconsistent code quality in the existing technology is solved, and efficient code generation and security improvement is achieved.
Patent Information
- Application Number
- CN202510340481.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing code generation model training methods rely on static offline data, resulting in uneven generated code quality, security risks, and lack of effective feedback mechanisms.
By introducing user natural language feedback, an interactive code generation method is constructed, the initial language model is used to sample and generate candidate codes, and the correct code is generated based on user feedback, and the distribution optimization model is improved through the target distribution and model to achieve the improvement of code quality.
It significantly improves the quality and reliability of code generation, reduces data preparation costs, and enhances the interactivity and feedback efficiency of model training.
Smart Images

Figure CN120276720A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a code generation method, in particular to an interactive code generation method, system and storage medium, belonging to the technical field of code generation. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have shown powerful capabilities in many natural language processing (NLP) tasks. Especially in code generation tasks, models such as BLOOM and GitHub Copilot have demonstrated the potential to assist programmers in writing code more efficiently. However, there are still some problems with the training methods of existing code generation models.
[0003] Most of the training data of existing pre-trained models comes from web-crawled code libraries, which contain a large number of low-quality and security-vulnerable code samples. Moreover, the training of the model only relies on these static offline data and lacks corrective feedback on the model output results. This leads to uneven code quality generated by the model, and security risks may occur when directly applied to actual projects. Improving the effective training method of the model and enhancing its ability to generate reliable and high-quality code has become a very urgent problem to be solved. Summary of the Invention
[0004] In view of this, the present application provides an interactive code generation method, system and storage medium to solve or alleviate the technical problems existing in the prior art, and at least provide a beneficial option.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] S100: Based on the initial language model π θ Perform a sampling to generate a candidate code x0.
[0007] S200: Obtain the natural language feedback result f of the user according to the candidate code x0, and generate a corrected code x1 based on the natural language feedback result f.
[0008] And, S300: Determine whether the corrected code x1 meets the preset code requirements, and when it meets, obtain the target code R based on the corrected code.
[0009] Further preferably: The S100 further includes: obtaining a task description dataset D = {(t, u)}; where the task description dataset D at least includes a task description t and a corresponding test set u.
[0010] Further preferably: The S200 includes:
[0011] Based on the test set u, determine whether the candidate code x0 passes the test, and obtain the data set C = {(x0, t, u)|x0 ∼ π θ (·|t), EVAL(x0, t) = 0, (t, u) ∈ D} according to the test judgment result.
[0012] Based on the data set C, obtain the natural language feedback result f, and obtain the sample set C annotated = {(x0, f, t)|(x0, t, u) ∈ C} according to the natural language feedback result f.
[0013] Construct the model improvement distribution q according to the natural language feedback result f.
[0014] And, based on the sample set C annotated and the model improvement distribution q, obtain the target distribution of the initial language model where (x1, t) is the quality scoring function of the corrected code x1 corresponding to the task description t, R(x1, t) = EVAL(x1, t), and β is the sign constant parameter.
[0015] Further preferably: The S200 further includes:
[0016] According to the target distribution the initial language model π θ and the model improvement distribution q, obtain the sampling estimation loss function
[0017]
[0018] Minimize the sampling estimation loss function L θ (t) to enhance the initial language model π N times θ ; where N is initialized to 1.
[0019] Further preferably: The S200 further includes:
[0020] Based on the initial language model π enhanced N times θ perform secondary sampling to construct the feedback distribution p based on the secondary sampling result F .
[0021] Generate the corrected code x1 through the correction model π ψ according to the feedback distribution p F .
[0022] Further preferably: The S300 further includes:
[0023] Train the initial language model π after N enhancements based on the target code θ to enhance the initial language model π N + 1 times θ .
[0024] Further preferably: It further includes:
[0025] When the corrected code x1 does not meet the preset code requirements, obtain the error code x0 based on the corrected code x1.
[0026] And, train the correction model π through the error code x0 ψ to enhance the correction model π ψ .
[0027] Further preferably: It further includes:
[0028] Repeat S100 to S300 to obtain an enhanced language model based on the initial language model π after N + 1 enhancements θ
[0029] Further preferably: This application also provides an interactive code generation system, and the system includes:
[0030] A first generation module for performing one sampling based on the initial language model π θ to generate a candidate code x0.
[0031] A second generation module for obtaining the natural language feedback result f of the user according to the candidate code x0, and generating a corrected code x1 based on the natural language feedback result f.
[0032] And, a third generation module for determining whether the corrected code x1 meets the preset code requirements, and when it meets, obtaining the target code R based on the corrected code.
[0033] Further preferably: This application also provides a storage medium, and a computer program is stored in the storage medium, wherein the computer program is set to execute the interactive code generation method when running.
[0034] Due to the adoption of the above technical solutions in the embodiments of this application, it has the following advantages:
[0035] 1. By introducing the natural language feedback of the user, this application can achieve significant improvements with very little feedback, greatly reducing the data preparation cost in the application.
[0036] 2. By introducing the feedback of the user during the use process, the model training becomes interactive.
[0037] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present application will become readily apparent by reference to the drawings and the following detailed description. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 Flowchart of the interactive code generation method described in the present application;
[0040] Figure 2 For use Figure 1 System framework diagram of the interactive code generation method described above. Detailed Embodiments
[0041] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature rather than restrictive.
[0042] The embodiments of the present application will be described in detail below with reference to the drawings.
[0043] As Figure 1 shown, the embodiments of the present application provide an interactive code generation method, including the following steps:
[0044] S100: Perform a first sampling based on the initial language model to generate candidate codes.
[0045] S200: Obtain the natural language feedback result of the user according to the candidate code, and generate a corrected code based on the natural language feedback result.
[0046] And, S300: Determine whether the corrected code meets the preset code requirements, and when it meets, obtain the target code based on the corrected code.
[0047] In this embodiment, specifically: the S100 further includes: obtaining a task description data set; wherein, the task description data set includes at least a task description and a corresponding test set.
[0048] In a feasible implementation manner, input the initial language model πθ , the coding task description dataset D = {(t, u)}, where:
[0049] π θ : The pre-trained language model with parameter θ, defined as π θ : V * → [0, 1], which maps the token sequence on the vocabulary V to the interval [0, 1].
[0050] t: The task description, such as "write a function to calculate prime numbers".
[0051] u: The test set UNITTESTS(t) corresponding to the task t.
[0052] Furthermore, perform a sampling operation on the initial language model π θ to generate candidate codes.
[0053] Specifically,
[0054] For the task description t ~ p(t), sample from π θ to generate candidate code x0: x0 ~ π θ (·|t).
[0055] In this embodiment, specifically: The S200 includes:
[0056] Based on the test set, determine whether the candidate code passes the test, and obtain a dataset according to the test judgment result.
[0057] In a feasible implementation manner, check whether the candidate code x0 passes the test set u: calculate EVAL(x0, t), where EVAL(x, t) is defined as:
[0058]
[0059] Furthermore, filter and retain those codes that do not pass the test to obtain a dataset
[0060] C = {(x0, t, u)|x0 ~ π θ (·|t), EVAL(x0, t) = 0, (t, u) ∈ D}
[0061] Among them, obtaining the low-quality code sample C of the initial language model prepares for the next user feedback.
[0062] Based on the dataset, obtain the natural language feedback result, and obtain a sample set according to the natural language feedback result.
[0063] In a feasible implementation, for each code sample (x0, t, u) in C, the user is requested to check the code x0 and give a natural language feedback result f, describing the code error and how to correct it. A sample set with feedback is obtained.
[0064] C annotated ={(x0, f, t)|(x0, t, u) ∈ C}
[0065] Furthermore, the natural language feedback result f is from the feedback distribution p F (f|t, x0, EVAL(x0, t) = 0), representing the feedback distribution for the incorrect code x0.
[0066] Construct a model improvement distribution based on the natural language feedback result.
[0067] In a feasible implementation, use the natural language feedback result f provided by the user to construct a model improvement distribution q for improving code quality:
[0068] where q is defined as the distribution over all possible corrected codes x1, and q uses the natural language feedback result f to guide how to correct the initial code x0.
[0069] The specific form of q is as follows:
[0070]
[0071] where δ0 and δ1 are Dirac δ distributions centered at 0 and 1, and π ψ is a correction model that generates the corrected code x1 according to the natural language feedback result f.
[0072] In addition, obtain the target distribution of the initial language model based on the sample set and the model improvement distribution.
[0073] In a feasible implementation, obtain a sample set C with user feedback annotated , and the model improvement distribution q of the improvement model for generating corrected codes defines the target distribution θ of the initial language model π in the form of:
[0074]
[0075] where R(x1, t) is the quality scoring function of the code x1 corresponding to the task t, here set as the unit test result:
[0076] R(x1, t) = EVAL(x1, t)
[0077] where β is a sign constant parameter.
[0078] In this embodiment, specifically: The S200 further includes:
[0079] Obtain a sampling estimation loss function according to the target distribution, the initial language model, and the model improvement distribution.
[0080] And, minimize the sampling estimation loss function to enhance the initial language model N times; where N is initialized to 1.
[0081] Where N is initialized to 1, which means that when the minimization process is performed for the first time, it is one enhancement. If the initial language model is subsequently enhanced, it is two enhancements, and so on to represent the number of enhancements.
[0082] In a feasible implementation, π θ Approximate by minimizing the following KL divergence
[0083]
[0084] Expanding gives the target cross-entropy loss function:
[0085] L(θ) = -E t~p(t) [L θ (t)]
[0086] Where,
[0087]
[0088] Furthermore, use importance sampling to estimate the loss function:
[0089]
[0090] Where q is the model improvement distribution using feedback constructed above.
[0091] Furthermore, by minimizing the loss function L(θ) estimated by sampling, make π θ Approximate the That fuses feedback knowledge
[0092] In this embodiment, specifically: The S200 further includes:
[0093] Perform secondary sampling based on the initial language model enhanced N times to construct a feedback distribution based on the secondary sampling results.
[0094] And, generate a corrected code by the correction model according to the feedback distribution.
[0095] In a feasible implementation, from π θSample the initial code \(x_0\), retain those for which \(EVAL(x_0, t)=0\), request the user to give the natural language feedback result \(f\) for \(x_0\), and construct the feedback distribution \(p\). F, Use the model \(\pi\). ψ Generate the corrected code \(x_1\) based on the natural language feedback result \(f\), approximately sampling from the distribution \(\pi\). ψ , only retain \(x_1\) that passes the test, i.e., \(EVAL(x_1, t)=1\).
[0096] It should be noted that \(\pi\) in this step θ is already the enhanced initial language model.
[0097] In this embodiment, specifically: the step S300 further includes:
[0098] Train the enhanced initial language model \(N\) times based on the target code to enhance the initial language model \(N + 1\) times.
[0099] In a feasible implementation, use \(\pi\). ψ Generate the corrected code set \(R\) (i.e., the target code) according to the natural language feedback result \(f\), and fine - tune the initial model \(\pi\) on \(R\). θ , to obtain the enhanced
[0100] Among them, the optimization algorithm used for fine - tuning:
[0101] Input: the initial model \(\pi\). θ , the corrected code set \(R=\{(t, x_1)\}\).
[0102] For each sample \((t, x_1)\), input the task description \(t\) and the target code \(x_1\).
[0103] Calculate the cross - entropy loss: \(loss = -\log\pi\). θ (x_1|t).
[0104] Minimize the loss through the gradient descent algorithm and update the parameters \(\theta\) of \(\pi\). θ of \(\pi\).
[0105] Obtain the fine - tuned
[0106] In this embodiment, specifically: it further includes:
[0107] When the corrected code does not meet the preset code requirements, obtain the error code based on the corrected code.
[0108] And, train the correction model through the error code to enhance the correction model.
[0109] In a feasible implementation, train the correction model π ψ , including:
[0110] Input: error code x0, natural language feedback result f.
[0111] Objective: corresponding manual correction code x 1。
[0112] Optimization objective: maximize the likelihood of π ψ to generate x1.
[0113] In this embodiment, specifically: it further includes:
[0114] Repeat S100 to S300 to obtain an enhanced language model based on the initial language model enhanced N + 1 times.
[0115] In a feasible implementation, repeat the above process and continuously improve to obtain an enhanced language model.
[0116] Among them, learned how to correct the common errors of π θ , and integrated the knowledge of user feedback.
[0117] Furthermore, it is also possible to evaluate the performance on unit tests:
[0118] For the task description t on the test set, sample to generate code; check whether the code passes the corresponding test cases; and calculate evaluation metrics such as pass@k.
[0119] It should be noted that the above embodiments use the feedback in the form of natural language for interactive training. The natural language feedback has strong expressiveness, large amount of information, and is more likely to point out the root cause of the problem; using this feedback to target the current output of the model makes the learning more targeted and the feedback is also more convenient to collect and accumulate. Other feedback forms such as scoring are difficult to express rich information and are not direct enough to guide the improvement of the model. The natural language feedback makes the interactive training method of this application more efficient.
[0120] Furthermore, the embodiments of this application first train a model that can make full use of feedback for code correction, and then enhance the original model through the generated code. This separated two-stage training method gives full play to the guiding role of natural language feedback and avoids the problem of poor stability during direct training. The two-stage training works together, enabling the feedback knowledge to be better accumulated into the model, which is the key to the significant effect of this application.
[0121] Further preferably: such as Figure 2As shown in the figure, the present application also provides an interactive code generation system, which includes:
[0122] A first generation module, configured to perform a first sampling based on the initial language model π θ to generate a candidate code x0.
[0123] A second generation module, configured to obtain the natural language feedback result f of the user according to the candidate code x0, and generate a corrected code x1 based on the natural language feedback result f.
[0124] And a third generation module, configured to determine whether the corrected code x1 meets the preset code requirements, and when it meets, obtain the target code R based on the corrected code.
[0125] Further preferably: The present application also provides a storage medium, in which a computer program is stored, and wherein the computer program is set to execute the interactive code generation method when running.
[0126] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An interactive code generation method, characterized in that Including the following steps: S100: Based on the initial language model π θ Perform a sampling to generate a candidate code x0; S2 00: Obtain the natural language feedback result f of the user according to the candidate code x0, and generate a corrected code x1 based on the natural language feedback result f; And, S300: Determine whether the corrected code x1 meets the preset code requirements, and when it meets, obtain the target code R based on the corrected code.
2. The interactive code production method according to claim 1, wherein The S100 further includes: obtaining a task description dataset D = {(t, u)}; where the task description dataset D at least includes a task description t and a corresponding test set u.
3. The interactive code generation method according to claim 2, wherein The S200 includes: Based on the test set u, determine whether the candidate code x0 passes the test, so as to obtain the data set C = {(x0, t, u)|x0 ~ π θ (·|t), EVAL(x0, t) = 0, (t, u) ∈ D} according to the test judgment result; Obtain the natural language feedback result f based on the dataset C, and obtain the sample set C according to the natural language feedback result f annotated ={(x0, f, t)|(x0, t, u) ∈ C}; Construct a model improvement distribution q according to the natural language feedback result f; and, based on the sample set C annotated and the model improvement distribution q to obtain the target distribution of the initial language model where (x1, t) is the quality scoring function of the corrected code x1 corresponding to the task description t, R(x1, t) = EVAL(x1, t), and β is the symbolic constant parameter.
4. The interactive code generation method according to claim 3, wherein The S200 further includes: According to the target distribution Initial language model π θ and the model improvement distribution q to obtain a sampling estimation loss function Minimize the sampling estimation loss function L θ (t) to enhance the initial language model π N times θ ; where N is initialized to 1.
5. The interactive code generation method according to claim 4, wherein The S200 further includes: Initial language model π enhanced N times θ Perform secondary sampling to construct a feedback distribution p based on the secondary sampling results F ; By correcting the model π ψ According to the feedback distribution p F Generate the corrected code x1.
6. The interactive code generation method according to claim 5, wherein The S300 further includes: Train the initial language model π enhanced N times based on the target code θ to enhance the initial language model π N + 1 times θ .
7. The interactive code generation method according to claim 5, characterized in that Further includes: When the corrected code x1 does not meet the preset code requirements, obtain an error code x0 based on the corrected code x1; And , training the correction model π with the error code x0 ψ , to enhance the correction model π ψ .
8. The interactive code generation method according to claim 6, wherein Further includes: Repeat S100 to S300 to obtain an enhanced language model based on the initial language model π enhanced N + 1 times θ Obtain the enhanced language model 9. A system using the interactive code generation method according to any one of claims 1-8, characterized in that, The system includes: A first generation module, for generating a candidate code x0 by performing a single sampling based on an initial language model π θ ; and generating a candidate code x0 by performing a single sampling A second generation module for obtaining the natural language feedback result f of the user according to the candidate code x0, and generating a corrected code x1 based on the natural language feedback result f; And, a third generation module for determining whether the corrected code x1 meets the preset code requirements, and when it meets, obtaining the target code R based on the corrected code.
10. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is set to execute the interactive code generation method according to any one of claims 1-8 when running.
Citation Information
Patent Citations
Computer and large model interaction system based on manual dictation command
CN119479651A
Natural language processing method and device based on language model
CN119598964A
Artificial intelligence regulatory mechanisms
WO2023014985A1