Retrieval enhancement code generation method and system based on retrieval query statement

By generating search query statements and optimizing query statement generators with reinforcement learning framework, the problem of low- and medium-to-low correlation and redundancy in code base search is solved, and high-quality code generation is achieved, suitable for a variety of programming languages and application scenarios.

CN120335811APending Publication Date: 2025-07-18PEKING UNIV

Patent Information

Application Number
CN202510173226.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-25
Filing Date
2025-02-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, code base retrieval based on natural language requirements has problems with low correlation and information redundancy, resulting in poor code generation results.

Method used

By generating search query statements, using reinforcement learning technology to train the search query statement generator, optimize the query statement to retrieve relevant contexts from the code base, and combine the reinforcement learning framework and reasonable reward mechanism to generate high-quality code snippets.

Benefits of technology

It improves the accuracy and quality of code generation, reduces redundant information interference, improves the efficiency and adaptability of code generation, and is suitable for a variety of programming languages and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335811A_ABST
    Figure CN120335811A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement code generation method and system based on retrieval query statements. The method comprises the following steps: acquiring a plurality of retrieval query statements; screening out context information matched with the natural language requirement and the retrieval query statement from the code library; generating a first code segment according to the natural language requirement and the context information; through a plurality of first code segments in one-to-one correspondence with the plurality of retrieval query statements and a corresponding label of each first code segment, training and optimizing a retrieval query statement generator to obtain a target retrieval query statement generator; based on a target natural language requirement, utilizing a trained and optimized target retrieval query statement generator to generate a query statement; and using the query statement to retrieve the related code context, and inputting the target natural language demand and the retrieved code context into the code generation model to complete the code generation of retrieval enhancement. According to the embodiment, the interference of redundancy and irrelevant content can be effectively reduced, so that the code generation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a retrieval-enhanced code generation method and system based on retrieval query statements. Background Art

[0002] With the continuous growth of software development requirements, developers usually develop software based on a code library, and program generation at the code library level becomes particularly important. Program generation at the code library level refers to generating code snippets that meet the given natural language requirements and are compatible with the current code library according to the given natural language requirements. In this task, the generated code snippets usually call functions and classes already defined in the current code library to ensure the compatibility of the generated code snippets with the current code library. Therefore, current program generation technologies at the code library level are usually based on a retrieval approach.

[0003] Specifically, this series of methods retrieves relevant code snippets from the current code library according to the input natural language requirements and uses these retrieval results to assist code generation. However, simply retrieving according to the given natural language requirements usually has the following problems: low relevance: due to the ambiguity and diversity of natural language requirements, simple matching retrieval may return many contents irrelevant to the actual requirements. Information redundancy: the retrieval results usually also contain a lot of redundant information, and these information will interfere with the generation process and instead affect the program generation effect. Summary of the Invention

[0004] In view of this, the present disclosure proposes a retrieval-enhanced code generation method and system based on retrieval query statements to solve the problems of low retrieval relevance and retrieval information redundancy existing in the related technologies.

[0005] A first aspect embodiment of the present disclosure proposes a retrieval-enhanced code generation method based on retrieval query statements, including:

[0006] Obtaining a plurality of retrieval query statements, where the plurality of retrieval query statements are generated by a retrieval query statement generator according to natural language requirements;

[0007] For any one of the plurality of retrieval query statements, screening out context information that matches the natural language requirements and the retrieval query statement from the code library;

[0008] Generating a first code segment according to the natural language requirements and the context information;

[0009] Calculating a reward value of the first code segment, if the reward value is greater than or equal to a preset threshold, setting a first label for the first code segment; if the reward value is less than the preset threshold, setting a second label for the first code segment;

[0010] The retrieval query statement generator is trained and optimized through a plurality of first code segments corresponding one by one to the plurality of retrieval query statements, and corresponding tags of each first code segment, to obtain a target retrieval query statement generator; the training and optimization direction is to train the retrieval query statement generator to generate a retrieval query statement corresponding to a first tag;

[0011] Based on a given target natural language requirement, a query statement is generated by using the trained and optimized target retrieval query statement generator;

[0012] The query statement is used to retrieve relevant code contexts, and the target natural language requirement and the retrieved code contexts are input into a code generation model to complete retrieval-enhanced code generation.

[0013] In the embodiments of the present disclosure, screening out context information matching the natural language requirement and the retrieval query statement from a code library includes:

[0014] The natural language requirement and the retrieval query statement are concatenated to obtain a concatenated sequence;

[0015] The concatenated sequence is encoded to obtain a first vector representation;

[0016] A plurality of predefined code segments in the code library are respectively encoded to obtain a plurality of second vector representations;

[0017] The similarity between each second vector representation and the first vector representation is calculated to obtain a plurality of similarity scores;

[0018] The plurality of similarity scores are sorted from high to low, and the code segment with the highest similarity score is selected as the context information.

[0019] In the embodiments of the present disclosure, calculating a reward value of the first code segment includes:

[0020] For any one of the plurality of first code segments, determining an n-gram exact match result, a syntax match result, a semantic match result, and a structure match result between the first code segment and a reference code segment;

[0021] According to the n-gram exact match result, the syntax match result, the semantic match result, and the structure match result, a first metric is determined; the first metric is used to judge the code quality of the first code segment;

[0022] The reward value of the first code segment is calculated according to the first metric and a second metric; the second metric is used to judge whether the first code segment calls existing functions and / or classes in the code block.

[0023] In an embodiment of the present disclosure, determining a first metric according to the n-gram exact matching result, the syntax matching result, the semantic matching result, and the structure matching result includes:

[0024] CodeBLEU = α·nM + β·SM1 + γ·SM2 + δ·SM3

[0025] Where nM represents the n-gram exact matching result, α represents a first weight coefficient corresponding to nM, SM1 represents the syntax matching result, β represents a second weight coefficient corresponding to SM1, SM2 represents the semantic matching result, γ represents a third weight coefficient corresponding to SM2, SM3 represents the structure matching result, and δ represents a fourth weight coefficient corresponding to SM3.

[0026] In an embodiment of the present disclosure, determining the syntax matching result between the first code segment and the reference code segment includes:

[0027] Comparing a first similarity of AST node pairs of the first code segment and the reference code segment;

[0028] Determining the syntax matching result according to the first similarity.

[0029] In an embodiment of the present disclosure, determining the semantic matching result between the first code segment and the reference code segment includes:

[0030] Respectively extracting semantic information of the first code segment and the reference code segment, and constructing a first semantic graph of the first code segment and a second semantic graph of the reference code segment according to their respective semantic information;

[0031] Calculating a second similarity between the first semantic graph and the second semantic graph;

[0032] Determining the semantic matching result according to the second similarity.

[0033] In an embodiment of the present disclosure, determining the structure matching result between the first code segment and the reference code segment includes:

[0034] Comparing a third similarity between the first code segment and the reference code segment in terms of code structure, where the code structure includes the arrangement relationship of identifiers and syntax symbols in the code;

[0035] Determining the structure matching result according to the third similarity.

[0036] An embodiment of the second aspect of the present disclosure provides a retrieval-enhanced code generation system based on a retrieval query statement. The system includes a retrieval query statement generator, a retriever, and a code generator;

[0037] The retrieval query statement generator is used to generate a target retrieval query statement corresponding to a given target natural language requirement; the retrieval query statement generator is obtained by training and optimizing a retrieval enhancement code generation method based on a retrieval query statement provided by the embodiment of the first aspect above;

[0038] The retriever is used to screen out target context information that matches the target natural language requirement and the target retrieval query statement from the code library;

[0039] The code generator is used to generate target code according to the target natural language requirement and the target context information.

[0040] An embodiment of the third aspect of the present disclosure provides an electronic device, which includes a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method described in the first aspect above.

[0041] An embodiment of the fourth aspect of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method described in the first aspect above.

[0042] The present disclosure has the following technical effects:

[0043] 1. Innovatively propose a retrieval query statement generation method. Different from traditional retrieval technologies based on templates or natural language requirements, this framework first generates query statements and retrieves relevant code contexts through the query statements, thereby realizing the interpretability and correctness of code generation.

[0044] 2. Improve the correctness of code generation. Through the reinforcement learning framework, optimize the retrieval query statement generator to generate more accurate and relevant retrieval queries, which can accurately extract code contexts closely related to natural language requirements from the code library, ensuring that the retrieved context information can effectively reduce the interference of redundant and irrelevant content, thereby improving the quality of code generation.

[0045] 3. Improve retrieval performance. Optimize the query generation statement through a reasonable reward mechanism, and continuously optimize the retrieval query statement generator according to the actual performance of code generation to form a closed-loop optimization. The framework can more effectively utilize the retrieved code fragments in the generation stage to improve the correctness of the generated program.

[0046] 4. Adaptive optimization. This method can perform adaptive optimization based on specific code library data. Whether it is for a general code library (such as GitHub open-source code) or a dedicated code library in a specific field, it can dynamically adjust the strategy to improve adaptability.

[0047] 5. Cross - language and multi - domain adaptation. This framework is applicable to multiple programming languages (such as Python, Java, C++, etc.) and various application scenarios, and has high generality and promotion value in the field of intelligent software engineering.

[0048] Generally speaking, the retrieval - enhanced code generation framework based on reinforcement - learning - based retrieval query generation provides an innovative solution for code generation and automated programming, while expanding the application space of code generation in the fields of actual engineering and research. Brief Description of the Drawings

[0049] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as a limitation of the present disclosure. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.

[0050] In the drawings:

[0051] Figure 1 Shows a schematic flow chart of a retrieval - enhanced code generation method provided by an embodiment of the present disclosure based on a retrieval query statement;

[0052] Figure 2 Shows a schematic diagram of the training process of a retrieval query statement generator provided by an embodiment of the present disclosure;

[0053] Figure 3 Shows a schematic diagram of the inference process of a retrieval query statement generator provided by an embodiment of the present disclosure;

[0054] Figure 4 Shows a schematic structural diagram of a code generation system provided by an embodiment of the present disclosure;

[0055] Figure 5 Shows a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure;

[0056] Figure 6 Shows a schematic diagram of a storage medium provided by an embodiment of the present disclosure. Detailed Embodiments

[0057] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0058] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this disclosure should have the ordinary meanings understood by those skilled in the art to which this disclosure pertains.

[0059] The following describes the technical scenarios related to the embodiments of this disclosure.

[0060] With the continuous growth of software development requirements, developers usually develop software based on a code library, and program generation at the code library level becomes particularly important. Program generation at the code library level refers to generating code snippets that meet the given natural language requirements and are compatible with the current code library according to the given natural language requirements. In this task, the generated code snippets usually call functions and classes that have been defined in the current code library to ensure the compatibility of the generated code snippets with the current code library. Therefore, current program generation techniques at the code library level are usually based on a retrieval approach.

[0061] Specifically, this series of methods retrieves relevant code snippets from the current code library according to the input natural language requirements and uses these retrieval results to assist in code generation. However, simply retrieving based on the given natural language requirements usually has the following problems:

[0062] Low relevance: Due to the ambiguity and diversity of natural language requirements, simple matching retrieval may return many contents that are not relevant to the actual requirements.

[0063] Information redundancy: The retrieval results usually also contain a lot of redundant information, which will interfere with the generation process and instead affect the program generation effect.

[0064] Therefore, how to generate query statements for retrieval to retrieve more accurate and relevant code snippets from the current code library, thereby improving the quality of program generation, has become an urgent problem to be solved. The present invention aims to solve the problems of low retrieval relevance and redundant retrieval information existing in the prior art, and proposes a retrieval-enhanced code generation framework for retrieval query generation. This framework uses reinforcement learning technology to train a retrieval query statement generator to generate high-quality query statements specifically for retrieval, thereby retrieving more relevant code contexts and assisting in generating high-quality program generation. Compared with the prior art, the present invention can generate high-quality retrieval query statements, effectively retrieve more relevant code contexts, avoid retrieving redundant code snippets, utilize high-relevance context information, thereby generating correct code snippets, meeting the actual software development requirements, and improving software development efficiency.

[0065] According to an embodiment of the present disclosure, an embodiment of a code generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0066] The present disclosure provides a retrieval-enhanced code generation method based on retrieval query statements, and this optimization method is implemented based on reinforcement learning. Among them, the external environment (Environment) is a code generator and a code retriever, the policy model (Policy Model) is a retrieval query statement generator, the action (Action) is to generate a series of query statements for retrieval, and the reward (Reward) is related to the correctness of the generated code. The specific implementation of the optimization method of its retrieval query statement generator can refer to the following embodiments:

[0067] Embodiment 1:

[0068] An embodiment of the present disclosure proposes a retrieval-enhanced code generation method based on retrieval query statements, as Figure 1 shown, the process includes the following steps:

[0069] Step S101, obtain multiple retrieval query statements.

[0070] Specifically, each retrieval query statement is generated by a retrieval query statement generator according to natural language requirements.

[0071] In some specific embodiments, multiple retrieval query statements can be generated in the following ways:

[0072] The retrieval query statement generator receives the natural language requirements of the user as input, understands the semantics of the natural language requirements to determine the core content and context of the user requirements. Based on the understanding of the natural language requirements, the retrieval query statement generator generates one or more retrieval query statements. Among them, each retrieval query statement is used to retrieve context information related to the natural language requirements from the code library. The multiple retrieval query statements generated by the retrieval query statement generator are different in terms of keywords, query structures, or retrieval strategies, and are used to capture context related to the natural language requirements from different perspectives.

[0073] Step S102, for any one of the multiple retrieval query statements, screen out context information from the code library that matches the natural language requirements and the retrieval query statement.

[0074] In some specific embodiments, the above step S102 is executed by a retriever, and its specific implementation steps include steps S1021 - S1025:

[0075] Step S1021: Concatenate the natural language requirement and the retrieval query statement to obtain a concatenated sequence.

[0076] Step S1022: Encode the concatenated sequence to obtain a first vector representation.

[0077] Among them, the encoding method can be selected according to the actual situation and will not be specifically limited here. For example, the encoding model CodeBERT encodes the above concatenated sequence.

[0078] Step S1023: Encode each of the multiple predefined code snippets in the code library to obtain multiple second vector representations.

[0079] Among them, the encoding method can be selected according to the actual situation and will not be specifically limited here. For example, the encoding model CodeBERT encodes the above multiple predefined code snippets.

[0080] Step S1024: Calculate the similarity between each second vector representation and the first vector representation to obtain multiple similarity scores.

[0081] Among them, the calculation method of similarity includes but is not limited to cosine similarity.

[0082] Step S1025: Sort the multiple similarity scores from high to low, and select the code snippet with the highest similarity score as the context information.

[0083] In some specific embodiments, the multiple similarity scores can be sorted in a descending order or in an ascending order. Whichever way is used, select the code snippet with the highest similarity score as the context information.

[0084] Step S103: Generate a first code segment according to the natural language requirement and the context information.

[0085] In some specific embodiments, the above Step S103 is executed by a code generator, and its specific implementation steps include:

[0086] Step S1031: Perform lexical analysis, syntactic analysis, and semantic analysis on the natural language requirement respectively;

[0087] Among them, lexical analysis: Decompose the natural language requirement into words and phrases; Syntactic analysis: Identify the structure of the sentence, such as subject, predicate, object and other components; Semantic analysis: Understand the meaning of the sentence, extract key intents and information.

[0088] Step S1032: Parse the context information, including: Code snippet analysis: Parse the retrieved code snippets to understand their functions and structures. Variable and method identification: Identify elements such as variables, methods, and classes in the code snippets. Dependency analysis: Understand the dependencies between code snippets, such as imported modules and used libraries.

[0089] Step S1033: Select a target template from the template library according to the natural language requirements and context information. The template library includes multiple code patterns and templates.

[0090] Step S1033: Combine the template and context information to generate a code segment; replace the placeholders in the template according to the specific parameters in the natural language requirements.

[0091] In some specific embodiments, it further includes: post-processing operations and test verification operations.

[0092] Among them, the generated code is optimized through post-processing operations to improve readability and efficiency. The generated code is tested and verified through test verification operations to ensure its correctness and functionality.

[0093] Step S104: Calculate the reward value of the first code segment. If the reward value is greater than or equal to the preset threshold, set the first label for the first code segment; if the reward value is less than the preset threshold, set the second label for the first code segment.

[0094] In the embodiments of the present disclosure, the preset threshold w can be selected according to the actual situation and is not specifically limited here; the reward value of the first code segment can be represented by R(q_i, s), where q_i represents the retrieval query statement corresponding to the first code segment.

[0095] In some specific embodiments, the above Step S104 can be represented by the following expression:

[0096] E(q_i, s) = 1 if R(q_i, s) ≥ w

[0097] E(q_i, s) = 0 if R(q_i, s) < w

[0098] In the embodiments of the present disclosure, if the reward value R(q_i, s) of the first code segment is greater than or equal to the preset threshold w, the first label will be set for the first code segment, that is, E(q_i, s) = 1; conversely, if the reward value R(q_i, s) of the first code segment is less than the preset threshold w, the second label will be set for the first code segment, that is, E(q_i, s) = 0.

[0099] In some specific embodiments, the above Step S104 includes Step S1041 - Step S1044:

[0100] Step S1041, for any one of the multiple first code segments, determine the n-gram exact match result, syntactic match result, semantic match result, and structural match result between the first code segment and the reference code segment.

[0101] In the embodiments of the present disclosure, the syntactic match result, semantic match result, and structural match result are used to measure the consistency of the syntax, semantics, and structure of the code. The n-gram Match is the n-gram exact match score of the traditional BLEU.

[0102] In some specific embodiments, determining the syntactic match result between the first code segment and the reference code segment includes steps a1 - a2:

[0103] Step a1, compare the first similarity of the AST node pairs of the first code segment and the reference code segment;

[0104] Step a2, determine the syntactic match result according to the first similarity.

[0105] In the embodiments of the present disclosure, the syntactic match considers the syntactic structure consistency of the code. CodeBLEU uses the abstract syntax tree (AST) of the code to represent the syntactic structure, compares the similarity of the AST node pairs of the generated code and the reference code, and uses the normalized tree edit distance (TED) to measure the difference between the two ASTs. The higher the syntactic match score, the closer the syntactic structure of the generated code is to the reference code.

[0106] In some specific embodiments, determining the semantic match result between the first code segment and the reference code segment includes steps b1 - b3:

[0107] Step b1, extract the semantic information of the first code segment and the reference code segment respectively, and construct the first semantic graph of the first code segment and the second semantic graph of the reference code segment according to their respective semantic information;

[0108] Step b2, calculate the second similarity between the first semantic graph and the second semantic graph;

[0109] Step b3, determine the semantic match result according to the second similarity.

[0110] In the embodiments of the present disclosure, whether the functions of the semantic information capture code of the semantic match (Semantic Match) code are consistent, CodeBLEU uses a static analysis tool to extract semantic information from the code, such as: function call relationships, variable dependency relationships, and constructs a semantic graph through this information to calculate the semantic graph similarity between the generated code and the reference code (for example, through a graph embedding method or an edit distance).

[0111] In some specific embodiments, determining the structural match result between the first code segment and the reference code segment includes steps c1 - step c2:

[0112] Step c1, comparing the third similarity between the first code segment and the reference code segment in terms of code structure, where the code structure includes the arrangement relationship of identifiers and syntax symbols in the code;

[0113] Step c2, determining the structural match result according to the third similarity.

[0114] In the embodiments of the present disclosure, structural match (Structural Match) in the code means that the arrangement of identifiers (such as variable names, function names) and syntax symbols in the code constitutes the structure of the code. CodeBLEU measures its structural similarity by comparing the DataFlow (data flow) of two code segments. Among them, the data flow represents the dependency relationship between variables and expressions in the code. The closer the data flow of the generated code is to the reference code, the higher the structural match score.

[0115] Step S1042, determining a first metric according to the n-gram exact match result, the syntax match result, the semantic match result, and the structural match result; the first metric is used to judge the code quality of the first code segment.

[0116] In some specific embodiments, the above step S1042 includes:

[0117] CodeBLEU = α·nM + β·SM1 + γ·SM2 + δ·SM3

[0118] Where nM represents the n-gram exact match result, α represents the first weight coefficient corresponding to nM, SM1 represents the syntax match result, β represents the second weight coefficient corresponding to SM1, SM2 represents the semantic match result, γ represents the third weight coefficient corresponding to SM2, SM3 represents the structural match result, and δ represents the fourth weight coefficient corresponding to SM3.

[0119] In the embodiments of the present disclosure, α, β, γ, and δ can be set according to actual situations and are not specifically limited here. For example, α, β, γ, and δ are set to 0.25.

[0120] The first metric CodeBLEU in the embodiments of the present disclosure is an evaluation metric for measuring the quality of generated code. It improves the traditional BLEU metric and provides a quality evaluation method more suitable for the code generation scenario by introducing syntax and semantic information unique to code. CodeBLEU not only measures the surface similarity between the generated code and the reference code, but also considers the structural and functional consistency of the code, thus more comprehensively evaluating the accuracy of code generation. Specifically, the design of CodeBLEU is based on the traditional BLEU metric, but adds three matching evaluations related to code characteristics, namely: Syntactic Match, Semantic Match, and Structural Match.

[0121] Step S1043, calculate the reward value of the first code segment according to the first metric and the second metric.

[0122] In the embodiments of the present disclosure, the second metric API_EM is used to determine whether the first code segment calls existing functions and / or classes in the code block.

[0123] In some specific embodiments, the above step S1042 includes:

[0124] R(q_i,s) = CodeBLEU + API_EM

[0125] Wherein, R(q_i,s) represents the reward value, CodeBLEU represents the first metric, and API_EM represents the second metric.

[0126] Step S105, train and optimize the retrieval query statement generator through a plurality of first code segments corresponding one by one to the plurality of retrieval query statements, and the corresponding labels of each first code segment, to obtain a target retrieval query statement generator.

[0127] In the embodiments of the present disclosure, the optimization training process of the retrieval query statement generator is, for example Figure 2 as shown.

[0128] The optimization objective of the retrieval query statement generator is as follows:

[0129]

[0130] Wherein, R(q_i,s) is the reward value of generating the retrieval query statement q_i, and log(p_θ(q_i|s)) is the probability of generating the retrieval query statement q_i. The optimization direction is to train the retrieval query statement generator to generate the retrieval query statement corresponding to the first label.

[0131] The gradient update strategy of the retrieval query statement generator is as follows:

[0132]

[0133] where α is the learning rate, is the updated gradient of the parameters of the retrieval query generator.

[0134] In some specific embodiments, the present invention crawls a large amount of code data on the open-source code platform Github for training the retrieval query statement generator, and optimizes the retrieval query generator to tend to generate retrieval query statements with high rewards.

[0135] Step S106, based on the given target natural language requirement, use the trained and optimized target retrieval query statement generator to generate a query statement;

[0136] Step S107, use the query statement to retrieve relevant code context, and input the target natural language requirement and the retrieved code context into the code generation model to complete retrieval-enhanced code generation.

[0137] Specifically, screen out target context information from the code library that matches the current natural language requirement and the target retrieval query statement; input the target natural language requirement and the target context information into a pre-trained code generation model to output enhanced target code.

[0138] In the embodiments of the present disclosure, through the above-mentioned Embodiment 1, a trained retrieval query statement generator is obtained, and then the trained retrieval query statement generator is used to generate a retrieval query statement corresponding to the natural language requirement. The retriever retrieves relevant context from the current code library according to the retrieval query statement, and finally inputs the natural language requirement and the retrieved context content into the code generator. The code generator completes program generation according to the input, and the process is as Figure 3 shown.

[0139] Embodiment 2:

[0140] The embodiments of the present disclosure also provide a retrieval-enhanced code generation system based on retrieval query statements, such as Figure 4 shown: The system includes a retrieval query statement generator, a retriever, and a code generator;

[0141] The retrieval query statement generator is used to generate a target retrieval query statement corresponding to the current natural language requirement; the retrieval query statement generator is trained and optimized based on the retrieval-enhanced code generation method based on retrieval query statements in the above-mentioned Embodiment 1;

[0142] The retriever is used to screen out target context information that matches the current natural language requirement and the target retrieval query statement from the code library;

[0143] The code generator is used to generate target code according to the target natural language requirement and the target context information.

[0144] In the embodiment of the present disclosure, the training and optimization process of the retrieval query statement generator can refer to Embodiment 1 above, and will not be elaborated here; the inference process of the retrieval query statement generator can refer to Embodiment 1 above, and will not be elaborated here.

[0145] Embodiment 3:

[0146] This framework uses reinforcement learning technology to train a retrieval query statement generator to generate high-quality query statements specifically for retrieval, so as to retrieve more relevant code contexts and assist in generating high-quality program generations.

[0147] In this framework, the external environment (Environment) is the code generator and the code retriever, the policy model (Policy Model) is the retrieval query statement generator, the action (Action) is to sample and generate a series of query statements for retrieval, and the reward (Reward) is related to the correctness of the generated code. The correctness of the generated code can be evaluated using metrics such as BLEU and CodeBLEU. Specifically, first, the natural language requirement is input into the retrieval query statement generator, and the retrieval query statement generator samples and generates M retrieval query statements for retrieving relevant contexts in the current code library; according to the natural language requirement and the generated retrieval query statements, the retriever is used to retrieve relevant code contexts in the current code library; the natural language requirement and the retrieved code contexts are input into the code generator to obtain M code fragments generated by the code generator (each query statement corresponds to a generated code fragment). For each query statement, this framework can use the retriever and the code generator to obtain a generated code fragment, and calculate the reward according to the generated code fragment. The reward value is calculated as follows:

[0148] R(q_i,s) = CodeBLEU + API_EM

[0149] R(q_i,s) = 1 if R(q_i,s) ≥ w

[0150] R(q_i,s) = 0 if R(q_i,s) < w

[0151] Among them, CodeBLEU is an evaluation metric for assessing the quality of generated code, API_EM is used to evaluate whether the functions and classes (referred to as APIs) existing in the current code library are correctly called in the generated code, R(q_i, s) is the reward for generating the query statement q_i, and w is the reward threshold. R(q_i, s) = 1 if R(q_i, s) ≥ w, which encourages the model to generate query statements with a reward value greater than w.

[0152] CodeBLEU is an evaluation metric specifically for code generation tasks. It improves the traditional BLEU metric by introducing code-specific syntax and semantic information, providing a quality assessment method more suitable for code generation scenarios. CodeBLEU not only measures the surface similarity between the generated code and the reference code but also considers the structural and functional consistency of the code, thus more comprehensively evaluating the accuracy of code generation. Specifically, the design of CodeBLEU is based on the traditional BLEU metric but adds three code-characteristic-related matching evaluations, namely: Syntactic Match, Semantic Match, and Structural Match. The total score of CodeBLEU is the result of weighted summation of these matching scores, and the formula is:

[0153] CodeBLEU = α·ngram Match + β·Syntactic Match + γ·Semantic Match + δ

[0154] ·Structural Match

[0155] Among them, α, β, γ, and δ are weight parameters (usually default to 0.25 each, indicating equal weights for each part), n-gramMatch is the n-gram exact match score of traditional BLEU, and the other three items measure the syntactic, semantic, and structural consistency of the code respectively. Syntactic Match considers the syntactic structure consistency of the code. CodeBLEU uses the Abstract Syntax Tree (AST) of the code to represent the syntactic structure, compares the similarity of AST node pairs between the generated code and the reference code, and uses the normalized Tree Edit Distance (TED) to measure the difference between two ASTs. The higher the syntactic match score, the closer the syntactic structure of the generated code is to the reference code. Semantic Match captures whether the semantic information of the code is consistent in terms of functionality. CodeBLEU uses static analysis tools to extract semantic information from the code, such as function call relationships and variable dependency relationships, constructs a semantic graph through this information, and calculates the similarity of the semantic graphs between the generated code and the reference code (for example, through graph embedding methods or edit distances). Structural Match of the code represents the arrangement of identifiers (such as variable names and function names) and syntactic symbols in the code to form the structure of the code. CodeBLEU measures its structural similarity by comparing the Data Flow of two pieces of code. Among them, the data flow represents the dependency relationship between variables and expressions in the code. The closer the data flow of the generated code is to the reference code, the higher the structural match score.

[0156] Then the optimization objective of the framework is as follows:

[0157]

[0158] Among them, R(q_i, s) is the reward for generating the query statement q_i, and log(p_θ(q_i|s)) is the probability of generating the query statement q_i. The optimization direction is to train the query generator to generate query statements with large reward values. Therefore, the above optimization objective is maximized in this framework.

[0159] The gradient update strategy is:

[0160]

[0161] Among them, α is the learning rate, is the updated gradient of the query generator parameters.

[0162] The training optimization process is as Figure 2 shown: The training process of the retrieval-enhanced code generation framework based on retrieval query generation

[0163] The present invention crawls a large amount of code data on the open-source code platform Github for the training of a retrieval-enhanced code generation framework based on retrieval query generation, and optimizes the query generator to tend to generate retrieval queries with high rewards. In the code generation inference stage, the framework first uses the trained retrieval query generator to generate queries for retrieval, the retriever retrieves relevant contexts from the current code library according to the queries, and finally inputs the natural language requirements and the retrieved content into the code generator, and the code generator completes the program generation according to the input. The inference process is as Figure 3 shown.

[0164] The embodiments of the present disclosure have the following technical effects:

[0165] 1. An innovative query statement generation method is proposed. Different from traditional retrieval techniques based on templates or natural language requirements, the framework first generates query statements and retrieves relevant code contexts through the query statements, thereby realizing the interpretability and correctness of code generation.

[0166] 2. Improve the correctness of code generation. Through the reinforcement learning framework, optimize the retrieval query statement generator to generate more accurate and relevant retrieval queries, which can accurately extract code contexts closely related to natural language requirements from the code library, ensuring that the retrieved context information can effectively reduce the interference of redundant and irrelevant content, thereby improving the quality of code generation.

[0167] 3. Improve retrieval performance. Optimize the query generation statement by designing a reasonable reward mechanism, and continuously optimize the retrieval query statement generator according to the actual performance of code generation to form a closed-loop optimization. The framework can more effectively utilize the retrieved code fragments in the generation stage and improve the correctness of the generated program.

[0168] 4. Adaptive optimization. This method can perform adaptive optimization based on specific code library data. Whether it is for a general code library (such as GitHub open-source code) or a dedicated code library in a specific field, it can dynamically adjust the strategy to improve adaptability.

[0169] 5. Cross-language and multi-domain adaptation. The framework is applicable to multiple programming languages (such as Python, Java, C++) and multiple application scenarios, and has high generality and promotion value in the field of intelligent software engineering.

[0170] Generally speaking, the retrieval-enhanced code generation framework based on retrieval query generation by reinforcement learning provides an innovative solution for code generation and automated programming, and at the same time expands the application space of code generation in the actual engineering and research fields.

[0171] The embodiments of the present disclosure also provide an electronic device to execute the above code generation method. Please refer to Figure 5, which shows a schematic diagram of an electronic device provided by some embodiments of the present disclosure. As Figure 5 shown, the electronic device 5 includes: a processor 500, a memory 501, a bus 502, and a communication interface 503. The processor 500, the communication interface 503, and the memory 501 are connected through the bus 502; a computer program that can run on the processor 500 is stored in the memory 501, and when the processor 500 runs the computer program, it executes the code generation method provided by any of the foregoing Figures 1 to 3 embodiments schematically illustrated.

[0172] Among them, the memory 501 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 503 (which can be wired or wireless), a communication connection is realized between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0173] The bus 502 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 501 is used to store a program, and after the processor 500 receives an execution instruction, it executes the program, and the Figures 1 to 3 code generation method disclosed in any of the foregoing embodiments schematically illustrated can be applied to the processor 500 or implemented by the processor 500.

[0174] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 500 or the instructions in the form of software. The above-mentioned processor 500 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501 and combines its hardware to complete the steps of the above method.

[0175] The electronic device provided by the embodiments of the present disclosure and the code generation method provided by the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0176] The embodiments of the present disclosure also provide a computer-readable storage medium corresponding to the code generation method provided by the foregoing embodiments. Please refer to Figure 6 which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the code generation method provided by any of the foregoing embodiments.

[0177] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0178] The computer-readable storage medium provided by the above embodiments of the present disclosure and the code generation method provided by the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0179] It should be noted that:

[0180] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0181] Similarly, it should be understood that in order to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present disclosure, the various features of the present disclosure are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed subject matter of the present disclosure requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present disclosure.

[0182] In addition, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features of different embodiments are meant to be within the scope of the present disclosure and form different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0183] The above is only a preferred specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A retrieval-enhanced code generation method based on retrieval query statements, characterized in that, The method includes: Obtaining a plurality of retrieval query statements, where the plurality of retrieval query statements are generated by a retrieval query statement generator according to natural language requirements; For any one of the plurality of retrieval query statements, screening out context information from the code library that matches the natural language requirements and the retrieval query statement; Generating a first code segment according to the natural language requirements and the context information; Calculating a reward value for the first code segment. If the reward value is greater than or equal to a preset threshold, setting a first label for the first code segment; if the reward value is less than the preset threshold, setting a second label for the first code segment; Training and optimizing the retrieval query statement generator through a plurality of first code segments corresponding one by one to the plurality of retrieval query statements, and the corresponding labels of each first code segment, to obtain a target retrieval query statement generator; the training and optimization direction is to train the retrieval query statement generator to generate retrieval query statements corresponding to the first label; Based on a given target natural language requirement, using the trained and optimized target retrieval query statement generator to generate a query statement; Using the query statement to retrieve relevant code context, and inputting the target natural language requirement and the retrieved code context into a code generation model to complete retrieval-enhanced code generation.

2. The method according to claim 1, characterized in that, Screening out context information from the code library that matches the natural language requirements and the retrieval query statement includes: Concatenating the natural language requirement and the retrieval query statement to obtain a concatenated sequence; Encoding the concatenated sequence to obtain a first vector representation; Encoding each of the defined multiple code segments in the code library to obtain multiple second vector representations; Calculating the similarity between each second vector representation and the first vector representation to obtain multiple similarity scores; Sorting the multiple similarity scores from high to low, and selecting the code segment with the highest similarity score as the context information.

3. The method according to claim 1 or 2, characterized in that, Calculating the reward value of the first code segment includes: For any one of the plurality of first code segments, determining the n-gram exact match result, syntax match result, semantic match result, and structural match result between the first code segment and a reference code segment; Determining a first metric according to the n-gram exact match result, the syntax match result, the semantic match result, and the structural match result; the first metric is used to judge the code quality of the first code segment; Calculating the reward value of the first code segment according to the first metric and a second metric; the second metric is used to judge whether the first code segment calls existing functions and / or classes in the code block.

4. The method according to claim 3, wherein Determining a first metric according to the n-gram exact match result, the syntax match result, the semantic match result, and the structural match result, including: CodeBLEU = α·nM + β·SM1 + γ·SM2 + δ·SM3 Wherein, nM represents the n-gram exact matching result, α represents the first weight coefficient corresponding to nM, SM1 represents the syntax matching result, β represents the second weight coefficient corresponding to SM1, SM2 represents the semantic matching result, γ represents the third weight coefficient corresponding to SM2, SM3 represents the structure matching result, and δ represents the fourth weight coefficient corresponding to SM3.

5. The method according to claim 3, characterized in that, Determining the syntax matching result between the first code segment and the reference code segment includes: Comparing the first similarity of the AST node pairs of the first code segment and the reference code segment; Determining the syntax matching result according to the first similarity.

6. The method according to claim 3, wherein Determining the semantic matching result between the first code segment and the reference code segment includes: Respectively extracting the semantic information of the first code segment and the reference code segment, and constructing the first semantic graph of the first code segment and the second semantic graph of the reference code segment according to their respective semantic information; Calculating the second similarity between the first semantic graph and the second semantic graph; Determining the semantic matching result according to the second similarity.

7. The method according to claim 3, characterized in that, Determining the structure matching result between the first code segment and the reference code segment includes: Comparing the third similarity in the code structure between the first code segment and the reference code segment, where the code structure includes the arrangement relationship of identifiers and syntax symbols in the code; Determining the structure matching result according to the third similarity.

8. A retrieval-enhanced code generation system based on retrieval query statements, characterized in that, The system includes a retrieval query statement generator, a retriever, and a code generator; The retrieval query statement generator is used to generate a target retrieval query statement corresponding to a given target natural language requirement; the retrieval query statement generator is trained and optimized based on the retrieval-enhanced code generation method according to claim 1; The retriever is used to screen out target context information that matches the target natural language requirement and the target retrieval query statement from the code library; The code generator is used to generate target code according to the target natural language requirement and the target context information.

9. A computer device, characterized in that, Including: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Generative adversarial network-based code generation and search method and system, and storage medium

    CN117519711A

  • Query result generation method and device

    CN117972077A

  • Retrieval enhanced code generation method and device, electronic equipment and storage medium

    CN118778941A

  • Code searching method and system and storage medium

    CN118860482A

Cited By

  • New code generation method and device based on historical code library, storage medium and electronic equipment

    CN121387255A

  • Code semantic understanding and mixed retrieval question and answer method and system oriented to AI programming

    CN122064775A