Large language model code generation system and method oriented to complex requirements

Through a large language model code generation system for complex needs, complex problems are broken down and functional consensus mechanism is adopted, the problem of low code generation accuracy under complex needs in the existing technology is solved, and higher code generation accuracy and reliability are achieved.

CN119938006APending Publication Date: 2025-05-06HARBIN INST OF TECH +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510306265.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-03-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing code generation methods are prone to causing chain errors when dealing with complex requirements, resulting in low accuracy in code generation.

Method used

A large language model code generation system for complex needs is adopted to decompose complex problems into simpler subfunctions through problem segmentation modules, and these subfunctions are gradually combined through problem solving modules, and the functional consensus mechanism is used to reduce the differences in code behavior.

Benefits of technology

By decomposing complex problems and adopting functional consensus mechanisms, the differences in code behavior are reduced, chain errors are alleviated, and the accuracy of code generation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938006A_ABST
    Figure CN119938006A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model code generation system and method oriented to complex requirements, and aims to introduce a new function to process specific sub-problems from a main problem. The new functions are decomposed recursively to finally form a function tree. These functions are then combined from bottom to top to achieve more and more complex goals. By decomposing the task into simpler sub-functions, the complexity can be gradually reduced. However, errors in sub-functions may propagate to the entire program, thereby impairing overall reliability. According to the function consensus mechanism provided by the invention, a plurality of functions can be sampled, the function showing the consensus in the candidate functions is selected, and the consensus is measured through aggregation similarity. By achieving the consensus, the code behavior difference can be reduced, so that the occurrence of chain errors is relieved, and the code generation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code generation, and in particular to a large language model code generation system and method oriented to complex requirements. Background Art

[0002] Since OpenAI launched ChatGPT in November 2022, large language models have quickly become the focus of technology and business. In the following two years, major technology companies around the world have successively released their own language model products, promoting a fierce model competition. In addition to OpenAI's GPT series, models from Baidu's Wenxin Yiyan, Alibaba's Tongyi Qianwen, and other domestic and foreign technology companies have also made their debut. Code pre-training has attracted widespread attention, and early models were mainly based on small language models (SLM). In recent years, with the development of large-scale pre-training technology, large language models (LLMs) in the code field have emerged and have shown significant performance in downstream code tasks. Although large language models can skillfully generate simple code snippets, when the code requirements become complex, chain errors will occur, leading to the problem of low accuracy of code generation. Summary of the invention

[0003] The purpose of the present invention is to provide a large language model code generation system and method for complex requirements, in view of the existing code generation method, which will cause chain errors to occur when the code requirements become complex, thereby leading to the problem of low accuracy of code generation.

[0004] The technical solution adopted by the present invention to solve the above technical problems is:

[0005] A large language model code generation system for complex requirements, the system comprising a problem segmentation module and a problem solving module;

[0006] The problem segmentation module specifically performs the following steps:

[0007] Step 1: Get the function header in text form and the function document describing the function header, and use the function header as the root node;

[0008] Step 2: Input the function head and function document into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as a child node;

[0009] Step 3: Input the function head corresponding to each child node and the function document corresponding to the function head into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as the child node of the previous parent node;

[0010] Step 4: Repeat step 3 until the large language model is no longer segmented, and use the function head obtained by the last segmentation as a leaf node;

[0011] The problem solving module specifically performs the following steps:

[0012] Step 1: Input the function header corresponding to each leaf node and the function document of the function header into the large language model to obtain the function implementation;

[0013] Step 2: Input the function headers, function documents, function implementations of all leaf nodes belonging to the same parent node, as well as the function headers and function documents corresponding to the parent node, into the large language model, and obtain the function implementation of the parent node based on the implementation prompt words;

[0014] Step 3: Take the parent node in step 2 as the child node and repeat step 2 until the function implementation of all nodes is obtained, that is, the complete implementation of the function header in text form in step 1.

[0015] Furthermore, the function is implemented by the following steps:

[0016] Step A: obtaining a function header and multiple function implementations of a function document corresponding to the function header, and obtaining multiple groups of function inputs;

[0017] Step B: Based on a set of function inputs, outputs of multiple function implementations are obtained, and similarity judgment is performed on the outputs of multiple function implementations;

[0018] Step C: Repeat step B to obtain the similarity judgment results corresponding to multiple groups of function inputs, and select the function with the highest similarity to implement f * Reaching consensus means the final function is realized.

[0019] Further, the similarity is expressed as:

[0020]

[0021] Where D(f) represents the input domain of function f, X represents a subset of the input domain, x represents a specific input, X~D(f) represents the sampling subset X approximating the input domain D(f), f(x) represents the output of function f when the input is x, g(x) represents the output of function g when the input is x, and sim(f,g) represents the similarity between functions f and g with the same input domain D(f)=D(g).

[0022] Furthermore, the function with the highest similarity achieves f * It is expressed as:

[0023]

[0024] Where F = {f (i)}, F indicates that functional consensus is required, that is, the function to obtain the function implementation, f (i) Represents different samples of the function F.

[0025] Furthermore, the function implementation is expressed as:

[0026]

[0027] Among them, f cur Indicates the current function. represents the current function after reaching functional consensus, that is, the final implementation of the current function, f' cur Indicates the partial implementation of the current function, CHILD(f cur ) indicates f cur The set of sub-functions of Indicates the sub-function f that has reached functional consensus i , i=1,2,...。

[0028] A large language model code generation method for complex requirements, the method comprising the following steps:

[0029] Step 1: Get the function header in text form and the function document describing the function header, and use the function header as the root node;

[0030] Step 2: Input the function head and function document into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as a child node;

[0031] Step 3: Input the function head corresponding to each child node and the function document corresponding to the function head into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as the child node of the previous parent node;

[0032] Step 4: Repeat step 3 until the large language model is no longer segmented, and use the function head obtained by the last segmentation as a leaf node;

[0033] Step 5: Input the function header corresponding to each leaf node and the function document of the function header into the large language model to obtain the function implementation;

[0034] Step 6: Input the function headers, function documents, function implementations of all leaf nodes belonging to the same parent node, as well as the function headers and function documents corresponding to the parent node, into the large language model, and obtain the function implementation of the parent node based on the implementation prompt words;

[0035] Step 7: Take the parent node in step 6 as the child node and repeat step 6 until the function implementation of all nodes is obtained, that is, the complete implementation of the function header in text form in step 1.

[0036] Furthermore, the function is implemented by the following steps:

[0037] Step A: obtaining a function header and multiple function implementations of a function document corresponding to the function header, and obtaining multiple groups of function inputs;

[0038] Step B: Based on a set of function inputs, outputs of multiple function implementations are obtained, and similarity judgment is performed on the outputs of multiple function implementations;

[0039] Step C: Repeat step B to obtain the similarity judgment results corresponding to multiple groups of function inputs, and select the function with the highest similarity to implement f * Reaching consensus means the final function is realized.

[0040] Further, the similarity is expressed as:

[0041]

[0042] Where D(f) represents the input domain of function f, X represents a subset of the input domain, x represents a specific input, X~D(f) represents the sampling subset X approximating the input domain D(f), f(x) represents the output of function f when the input is x, g(x) represents the output of function g when the input is x, and sim(f,g) represents the similarity between functions f and g with the same input domain D(f)=D(g).

[0043] Furthermore, the function with the highest similarity achieves f * It is expressed as:

[0044]

[0045] Where F = {f (i)}, F indicates that functional consensus is required, that is, the function to obtain the function implementation, f (i) Represents different samples of the function F.

[0046] Furthermore, the function implementation is expressed as:

[0047]

[0048] Among them, f cur Indicates the current function. represents the current function after reaching functional consensus, that is, the final implementation of the current function, f' cur Indicates the partial implementation of the current function, CHILD(f cur ) indicates fcur The set of sub-functions of Indicates the sub-function f that has reached functional consensus i , i=1,2,...。

[0049] The beneficial effects of the present invention are:

[0050] This application starts with the main problem and introduces new functions to handle specific sub-problems. These new functions will be recursively decomposed to eventually form a function tree. Subsequently, these functions are combined from the bottom up to achieve increasingly complex goals. By decomposing tasks into simpler sub-functions, the complexity can be gradually reduced. However, errors in sub-functions may propagate throughout the program, thereby compromising the overall reliability. The functional consensus mechanism proposed in this application samples multiple functions and selects functions that show consensus among candidate functions, and its consensus is measured by aggregate similarity. By reaching a consensus, this application can reduce differences in code behavior, thereby alleviating the occurrence of chain errors and improving the accuracy of code generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart for the framework of this application;

[0052] Figure 2 A comparison chart between the decomposition by planning and this application. DETAILED DESCRIPTION

[0053] It should be particularly noted that, in the absence of conflict, the various embodiments disclosed in this application can be combined with each other.

[0054] Specific implementation method 1: The large language model code generation system for complex requirements described in this implementation method includes a problem segmentation module and a problem solving module;

[0055] The problem segmentation module specifically performs the following steps:

[0056] Step 1: Get the function header in text form and the function document describing the function header, and use the function header as the root node;

[0057] Step 2: Input the function head and function document into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as a child node;

[0058] Step 3: Input the function head corresponding to each child node and the function document corresponding to the function head into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as the child node of the previous parent node;

[0059] Step 4: Repeat step 3 until the large language model is no longer segmented, and use the function head obtained by the last segmentation as a leaf node;

[0060] The problem solving module specifically performs the following steps:

[0061] Step 1: Input the function header corresponding to each leaf node and the function document of the function header into the large language model to obtain the function implementation;

[0062] Step 2: Input the function headers, function documents, function implementations of all leaf nodes belonging to the same parent node, as well as the function headers and function documents corresponding to the parent node, into the large language model, and obtain the function implementation of the parent node based on the implementation prompt words;

[0063] Step 3: Take the parent node in step 2 as the child node and repeat step 2 until the function implementation of all nodes is obtained, that is, the complete implementation of the function header in text form in step 1.

[0064] This application is used to assist large language models in completing code generation capabilities under complex requirements and improve code quality. In general, it adopts a divide-and-conquer strategy and a functional consensus mechanism to decompose complex problems. Starting from the main problem, the framework recursively introduces new functions to solve specific sub-problems, and finally forms a function tree structure. DCGRAMMER combines these functions in a bottom-up manner to achieve more complex task goals. In addition, through the functional consensus mechanism, the framework is also able to reach consensus between candidate functions, thereby reducing differences in code behavior and mitigating the cascading effects of errors.

[0065] This application is inspired by the programming habits of human programmers - human programmers tend to decompose tasks into clearly defined sub-functions and implement them recursively, making functions reusable and taking advantage of the divide and conquer principle. DCGRAMMER recursively decomposes requirements and solves function problems to form complex solutions, thereby unleashing the potential of large language models in code generation. In order to make the following description clearer, it is necessary to briefly describe the relevant representation of the function. A function is defined as a relationship between an input set and an output set, where each input corresponds to exactly one output, expressed as y=f(x). In computer programming, a function is represented by its head h f and subject b f Identification and usually accompanying documentation f To improve readability. Functions can be called from other programs, which allows large and complex requirements to be broken down into smaller structures that have higher understandability and quality.

[0066] The processing flow of the code generation framework DCGRAMMER proposed in this application is as follows Figure 1 The pseudo code algorithm description of the overall process is shown in Algorithm 1.

[0067]

[0068] For the first module: Divide and conquer strategy. In order to improve the programming ability of large language models under complex requirements, following the design and definition of programming problems in existing popular public datasets, this application focuses on code generation tasks in open fields.

[0069] Divide is a top-down process that iteratively decomposes the problem. Given a code generation problem, this process starts with the entry function f root Start. This application instructs the model to write the current function f cur When introducing a new function f i ∈CHILD(f cur ), these functions solve certain sub-goals. In order to reduce the complexity of each generation process, this application only requires the generation of the header of the new function and Documentation And their realization After completing the current function, the model starts processing those unimplemented sub-functions and completes to f' i This process continues until the model believes that the function is simple enough and does not need to be further decomposed, eventually forming a dependency tree T = TREE (f root ,CHILD(f root )). The partitioning process is similar to the search starting from the entry function, gradually involving new sub-functions while writing the current function, and recursively implementing them. This application guides the whole process through depth-first search.

[0070] Problem Solving Conquer. Conquer is a process of achieving a complex goal by aggregating smaller functions. This application notes that in the process of writing parent functions from top to bottom, the child functions have not yet been implemented. Therefore, these parent functions may not effectively utilize the child functions, and in the worst case, misuse them. DCGRAMMER solves this problem by regenerating functions in reverse topological order on the dependency tree T - starting from the leaf nodes, and handling complex goals by combining solved child functions, that is:

[0071]

[0072] The divide-and-conquer strategy naturally achieves decomposition and combination in the code generation process. Unlike the two-stage method and the agent-based method, this application dynamically introduces new functions in the process, which is much easier than making a complete plan at the beginning. In addition, while the planning or agent methods require chat capabilities, DCGRAMMER represents subtasks through functions (see Figure 2 ), which makes it more applicable in specialized code generation models.

[0073] For the second module: Functional Consensus. The decomposition of complex tasks benefits from solving simpler sub-goals, but this may introduce the risk of cascading errors. To address this problem, this application introduces functional consensus, which aims to reduce inconsistencies in program behavior. This is achieved by sampling multiple functions and selecting those that show consensus, which is measured by the aggregation of functional similarity between candidate functions, thereby reducing the impact of abnormal functions.

[0074] Functional similarity: A program specifies its functionality (or behavior) through the control flow defined by its code semantics. However, comparing the functionality between two programs based on semantics is challenging. By decomposing requirements into functions, our DCGRAMMER is able to treat function behavior as a black box that maps parameters to return values. Considering two functions f and g with the same input domain D(f) = D(g), we define the similarity sim(f, g) between them as the consistency of the output given the same input values.

[0075]

[0076] The similarity is 1 if and only if two functions output the same value for all inputs: In most cases, the input domain D(f) is unbounded, which makes it almost impossible to measure it in practice. Therefore, this application approximates this measure by sampling a subset X~D(f) from the possible inputs, which is done with the help of a large language model.

[0077] By implementing F = {f (i)} to sample and select the candidate function f with the highest similarity to other functions * Reach a consensus.

[0078]

[0079] By introducing a functional consensus mechanism, the functions generated by DCGRAMMER are more consistent and universal in functionality, while filtering out abnormal samples. This process is not only applied to the final program, but also covers every subtree in the bottom-up conquest phase, ensuring that step-by-step and comprehensive verification can be performed from the most basic functions to the entire program.

[0080] In general, DCGRAMMER is a code generator. This application designs DCGRAMMER as a problem-solving process whose input is the function signature f(x) and whose output is the final solution f * (x), such as Figure 1 Given a problem f(x), DCGRAMMER will partially implement it as a function f'(x), and refer to the unimplemented sub-functions g(y) and h(z). These sub-functions will be recursively handed over to DCGRAMMER for processing. Then, this application is based on the solved sub-function g * (y) and h * (z) Sample k realizations f' (i) (x). Functional consensus is calculated by evaluating candidate functions under possible inputs, and finally the functions with the most similar behavior to the sub-functions are selected and merged to generate the final solution.

[0081] To ensure fairness and quantifiable comparison, this application was tested on competition-level code generation and mathematical reasoning benchmarks. They cover the most advanced large language models, and in addition to the GPT series models, Llama3 is also used. 8b , StableCode 3b , CodeLlama 34b This application used instruction variants of these models and performed inference using vLLM on a single A100-80G at BF16 precision.

[0082] Code generation experiment: This application selected three benchmarks to evaluate the effect of code generation, namely HumanEval proposed by Chen et al. in 2021, which includes entry-level programming problems; MBPP proposed by Austin et al. in 2021, which includes standard library calls and programming basics; and xCodeEval proposed by Khan et al. in 2023, which includes algorithm challenges from the competitive programming platform CodeForces. In the specific experimental setting, this application selected the full test set of HumanEval (164 questions), sampled 200 questions from MBPP, and sampled 500 questions from xCodeEval, and followed the practice of EbTech to divide them into four subsets according to difficulty: Easy (≤1200), Mid (100-1599), Hard (1600-1999), Expert (≥2000), and the default evaluation indicator for code generation is Pass@1.

[0083] This experiment compares DCGRAMMER with standard prompts (Brown et al., 2020), the two-stage decomposition method Parsel (Zelikman et al., 2023), the self-testing method CodeT (Chen et al., 2023a), the self-improvement methods Reflexion and LDB (Shinn et al., 2023; Zhong et al., 2024), and the multi-agent development framework MetaGPT (Hong et al., 2024). An example demonstration is used in this experiment to implement standard prompts. CodeT samples 11 solutions using standard prompts and evaluates them on model-generated tests. The results of Reflexion are reproduced from the original code. In addition, DCGRAMMER uses 2 hints (2-shot) in the segmentation stage and 1 hint (1-shot) in the subfunction solving stage. The number of implementations sampled in the function consensus is set to 11, which is suitable for the code generation task.

[0084] Table 1 Experimental results on three code generation datasets

[0085]

[0086] Table 1 shows the code generation performance on advanced proprietary models GPT-3.5 (Ouyang et al., 2022) and GPT-4 (OpenAI, 2023). For basic programming problems such as HumanEval and MBPP, DCGRAMMER surpasses the previous state-of-the-art method by +3.3% on Pass@1 and reduces the error rate by 18.6%. In addition, DCGRAMMER shows significant improvements on competition-level problems, outperforming other methods by 10.4% when using GPT-4 and 35.3% when using GPT-3.5. It can be observed that DCGRAMMER is able to enhance the ability of large language models to solve more complex programming tasks, with an average accuracy improvement of 82.3% over the baseline on the medium and hard subsets of xCodeEval.

[0087] Table 2 Code generation performance of open source models on the HumanEval dataset

[0088]

[0089] Evaluations are also performed on open source large language models, including Llama3, StableCode, and CodeLlama, and the results are shown in Table 2. DCGRAMMER also improves the performance of smaller models in code generation, with an average improvement of +38.0% compared to standard prompts, and outperforms the previous best method CodeT by +14.6% on HumanEval. Experimental results show that the invention achieves cutting-edge performance on a variety of models, ranging from basic programming to competition questions.

[0090] Mathematical Reasoning Experiments: Programming can be viewed as a tool to enhance the reasoning capabilities of large language models (LLMs). Compared to text-based reasoning methods (such as "CoT"), programs offer unique advantages in iteration and computation. To test the generalization ability of DCGRAMMER beyond algorithmic challenges, this application conducts experiments on MATH (Hendrycks et al., 2021b), a competition-level mathematical reasoning benchmark.

[0091] The experiments are conducted on a subset of the MATH test set, containing 500 randomly selected questions, which can be divided into 7 non-overlapping topics or 5 difficulty levels. This application compares DCGRAMMER with text-based baseline models: Standard Prompting; Chain-of-Thought (Wei et al., 2022); Program-of-Thought (Chen et al., 2023b); Self-Refine (Madaan et al., 2023); Cumulative Reasoning (Zhang et al., 2024). The results of Cumulative Reasoning are from the original paper. Standard prompting and chain-of-thought reasoning use 7 example prompts constructed from the training set.

[0092] Program-of-Thought and Self-Refine use 1-shot prompts to generate a solution() function to solve the problem. In addition, Self-Refine iteratively optimizes the program based on runtime feedback. All baseline methods are repeated 5 times with self-consistency. DCGRAMMER adopts a program-assisted reasoning setting and writes

[0093] The `solution()` function is used to obtain the final prediction results by running the program. In the feature consensus, the number of sampled realizations |F| is set to 5 to match the baseline method.

[0094] Table 3 Experimental results on the competition-level mathematical reasoning benchmark MATH

[0095]

[0096] The experimental results of MATH are shown in Table 3. The results show that program-assisted reasoning is generally better than text-based reasoning. Based on GPT-4, DCGRAMMER is higher than the strongest baseline model CumulativeReasoning (6.0 / 8.3%), and surpasses the basic program-assisted baseline Program-of-Thought (PoT) (10.0 / 14.7%). When using GPT-3.5-turbo as the basis, DCGRAMMER is higher than the strongest baseline model (6.2 / 11.1%) and higher than PoT (13.0 / 31.7%), which shows that the method of this application has obvious advantages over text reasoning and other program-assisted reasoning methods.

[0097] On the open source model, DCGRAMMER using Llama3 outperforms PoT (12.4 / 38.0%), and even achieves similar performance when compared to the state-of-the-art method based on GPT-3.5 (45.0 vs. 48.6). When based on StableCode and CodeLLaMA, this application achieves significant improvements of (12.2 / 84.7%) and (9.2 / 60.5%) respectively. This improvement shows that this application can significantly improve the capabilities of small LLMs and give open source LLMs stronger complex reasoning capabilities through programming.

[0098] Example

[0099] Ablation experiment and token usage related to this application

[0100] Here, the present application observes the performance changes by adding, subtracting or replacing the implementation methods of some key steps, and the optimization contribution results of each key step are shown in Table 4.

[0101] Table 4 DCGRAMMER ablation study of GPT-3.5 on HumanEval

[0102]

[0103] In order to analyze the impact of the divide-and-conquer strategy and functional consensus in DCGRAMMER, this application conducted ablation experiments with different settings. The experiment also includes a study of replacing functional consensus with self-testing. The ablation experiment was conducted based on HumanEval using GPT-3.5, as shown in Table 4. This application observed that function decomposition and recombination brought cumulative performance improvements. In addition, functional consensus outperforms self-testing in performance. Overall, DCGRAMMER improves performance by 17.1% over the baseline model, while using 5.09 times the number of tokens as the baseline. Compared with the previous state-of-the-art LDB method (about 23KTokens), this application improved performance by 2.5% while reducing token usage by 76.5%.

[0104] The DCGRAMMER of this application can effectively reduce the complexity of problems in the complex code generation process and improve the reliability of the generated code. By recursively decomposing tasks and adopting a functional consensus mechanism, this application can not only reduce the risk of error propagation, but also ensure the consistency of the generated sub-functions in functional behavior. Compared with the prior art, this application shows stronger adaptability and robustness in dealing with complex problems, especially in tasks such as code generation and multi-level logic processing, and can significantly improve the quality and efficiency of code generation. Experimental results show that the results on the GPT model surpass the current most advanced methods by an average of 9.8%. Further experiments on the mathematics competition benchmark MATH showed that the performance was improved by 6.0% after using GPT-4, which shows that DCGRAMMER can also be extended to complex reasoning tasks.

[0105] It should be noted that the specific implementation is only an explanation and description of the technical solution of the present invention, and cannot be used to limit the scope of protection of the rights. Any partial changes made according to the claims and description of the present invention should still fall within the scope of protection of the present invention.

Claims

1. A large language model code generation system for complex requirements, characterized by The system includes a problem segmentation module and a problem solving module; The problem segmentation module specifically performs the following steps: Step 1: Get the function header in text form and the function document describing the function header, and use the function header as the root node; Step 2: Input the function head and function document into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as a child node; Step 3: Input the function head corresponding to each child node and the function document corresponding to the function head into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as the child node of the previous parent node; Step 4: Repeat step 3 until the large language model is no longer segmented, and use the function head obtained by the last segmentation as a leaf node; The problem solving module specifically performs the following steps: Step 1: Input the function header corresponding to each leaf node and the function document of the function header into the large language model to obtain the function implementation; Step 2: Input the function headers, function documents, function implementations of all leaf nodes belonging to the same parent node, as well as the function headers and function documents corresponding to the parent node, into the large language model, and obtain the function implementation of the parent node based on the implementation prompt words; Step 3: Take the parent node in step 2 as the child node and repeat step 2 until the function implementation of all nodes is obtained, that is, the complete implementation of the function header in text form in step 1.

2. The large language model code generation system for complex requirements according to claim 1 is characterized in that The function is implemented by the following steps: Step A: obtaining a function header and multiple function implementations of a function document corresponding to the function header, and obtaining multiple groups of function inputs; Step B: Based on a set of function inputs, outputs of multiple function implementations are obtained, and similarity judgment is performed on the outputs of multiple function implementations; Step C: Repeat step B to obtain the similarity judgment results corresponding to multiple groups of function inputs, and select the function with the highest similarity to implement f * Reaching consensus means the final function is realized.

3. The large language model code generation system for complex requirements according to claim 2 is characterized in that The similarity is expressed as: Where D(f) represents the input domain of function f, X represents a subset of the input domain, x represents a specific input, X~D(f) represents the sampling subset X approximating the input domain D(f), f(x) represents the output of function f when the input is x, g(x) represents the output of function g when the input is x, and sim(f,g) represents the similarity between functions f and g with the same input domain D(f)=D(g).

4. The large language model code generation system for complex requirements according to claim 3 is characterized in that The function with the highest similarity achieves f * It is expressed as: Where F = {f (i) }, F indicates that functional consensus is required, that is, the function to obtain the function implementation, f (i) Represents different samples of the function F.

5. The large language model code generation system for complex requirements according to claim 4 is characterized in that The function implementation is expressed as: Among them, f cur Indicates the current function. represents the current function after reaching functional consensus, that is, the final implementation of the current function, f' cur Indicates the partial implementation of the current function, CHILD(f cur ) indicates f cur The set of sub-functions of Indicates the sub-function f that has reached functional consensus i , i=1,2,...。 6. A large language model code generation method for complex requirements, characterized by The method comprises the following steps: Step 1: Get the function header in text form and the function document describing the function header, and use the function header as the root node; Step 2: Input the function head and function document into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as a child node; Step 3: Input the function head corresponding to each child node and the function document corresponding to the function head into the large language model, and obtain the segmented function head and the function document corresponding to each segmented function head according to the segmentation prompt word, and use the segmented function head as the child node of the previous parent node; Step 4: Repeat step 3 until the large language model is no longer segmented, and use the function head obtained by the last segmentation as a leaf node; Step 5: Input the function header corresponding to each leaf node and the function document of the function header into the large language model to obtain the function implementation; Step 6: Input the function headers, function documents, function implementations of all leaf nodes belonging to the same parent node, as well as the function headers and function documents corresponding to the parent node, into the large language model, and obtain the function implementation of the parent node based on the implementation prompt words; Step 7: Take the parent node in step 6 as the child node and repeat step 6 until the function implementation of all nodes is obtained, that is, the complete implementation of the function header in text form in step 1.

7. The method for generating large language model code for complex requirements according to claim 6, characterized in that The function is implemented by the following steps: Step A: obtaining a function header and multiple function implementations of a function document corresponding to the function header, and obtaining multiple groups of function inputs; Step B: Based on a set of function inputs, outputs of multiple function implementations are obtained, and similarity judgment is performed on the outputs of multiple function implementations; Step C: Repeat step B to obtain the similarity judgment results corresponding to multiple groups of function inputs, and select the function with the highest similarity to implement f * Reaching consensus means the final function is realized.

8. The method for generating large language model code for complex requirements according to claim 7, characterized in that The similarity is expressed as: Where D(f) represents the input domain of function f, X represents a subset of the input domain, x represents a specific input, X~D(f) represents the sampling subset X approximating the input domain D(f), f(x) represents the output of function f when the input is x, g(x) represents the output of function g when the input is x, and sim(f,g) represents the similarity between functions f and g with the same input domain D(f)=D(g).

9. The method for generating large language model code for complex requirements according to claim 8, characterized in that The function with the highest similarity achieves f * It is expressed as: Where F = {f (i) }, F indicates that functional consensus is required, that is, the function to obtain the function implementation, f (i) Represents different samples of the function F.

10. The method for generating large language model code for complex requirements according to claim 9, characterized in that The function implementation is expressed as: Among them, f cur Indicates the current function. represents the current function after reaching functional consensus, that is, the final implementation of the current function, f' cur Indicates the partial implementation of the current function, CHILD(f cur ) indicates f cur The set of sub-functions of Indicates the sub-function f that has reached functional consensus i , i=1,2,...。

Citation Information

Patent Citations

  • Zero-sample large model generation code detection method and system

    CN117608648A

  • Global optimization method for decomposition and execution of complex data query task

    CN118689899A

  • Test method and device for code generation model

    CN119003350A

  • Code completion method and system based on large language model

    CN119356686A

  • Code generation method and system for multi-source information fusion based on large language model

    CN119512524A