A code generation task adaptive reasoning method, device and equipment
By converting the code generation task into pseudocode and using the complexity evaluation model to dynamically adjust the reasoning mode, the overthinking problem of large language models in code generation is solved, and efficient and accurate code generation is achieved, which is suitable for actual software development.
Patent Information
- Application Number
- CN202511052757.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing code generation technologies based on large language models suffer from the overthinking problem, which leads to high consumption of computing resources and inaccurate generation results, limiting their application in actual software development.
By obtaining the problem description of the code generation task, converting it into pseudocode, and using the complexity assessment model to dynamically perceive the task complexity, different reasoning modes are matched for code generation, including simple, medium and difficult modes, to avoid unnecessary waste of computing resources.
It achieves adaptive adjustment of inference depth and resource allocation according to the actual needs of the task, significantly improves inference efficiency and accuracy, saves about 70% of inference length, and optimizes the efficiency and accuracy of code generation.
Smart Images

Figure CN120560664B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large model application, and in particular relates to a method, device and equipment for adaptive reasoning of code generation tasks. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Currently, the mainstream approach to intelligent code generation based on large models focuses on improving code generation performance through large-scale model training. However, building high-quality training datasets is time-consuming and labor-intensive. With the emergence of reasoning models such as GPT-4o and Deepseek-R1, test-time scaling has become a research hotspot. This approach eliminates the need for additional model training and instead guides the model to generate detailed chains of reasoning (CoTs) during testing to improve performance on complex logical problems, such as mathematics and code generation.
[0004] However, existing test-time expansion techniques suffer from significant overthinking. When handling even simple code generation tasks, models often generate reasoning processes lasting thousands of tokens, including lengthy and unnecessary reasoning. This significantly increases computing resource consumption and response time, severely limiting the application of large language model-assisted code generation techniques in real-world software development scenarios. Furthermore, redundant reasoning processes can turn originally correct reasoning into errors, compromising the accuracy of generated results. Summary of the Invention
[0005] In view of this, the present invention provides a code generation task adaptive reasoning method, device and equipment to realize the dynamic perception capability of large language models to task complexity.
[0006] One aspect of the present invention provides a code generation task adaptive reasoning method, comprising the following steps:
[0007] Obtain a problem description for a code generation task, and convert the problem description into pseudocode based on a large language model;
[0008] According to the problem description and the corresponding pseudocode, a complexity evaluation is performed on the code generation task based on a pre-trained complexity evaluation model; the complexity evaluation model is trained with the problem description and the pseudocode as input and the task complexity as output;
[0009] According to the complexity evaluation results, the reasoning mode is matched, and the set reasoning model is adopted to perform the code generation task based on the matched reasoning mode; among them, different complexities correspond to reasoning modes with different reasoning depths.
[0010] In some embodiments, after obtaining the problem description of the code generation task, subtask decomposition is performed based on the problem description to obtain a problem description of each subtask;
[0011] For each subtask, the corresponding problem description is converted into pseudocode based on the large language model; based on the problem description and corresponding pseudocode of each subtask, the complexity of the subtask is evaluated based on the pre-trained complexity evaluation model;
[0012] According to the complexity evaluation results, the reasoning pattern is matched for each subtask, and the code generation tasks are executed in sequence according to the corresponding reasoning pattern according to the execution order of the subtasks.
[0013] In some embodiments, the complexity assessment model training method includes:
[0014] Get the training data set. Each piece of training data includes a problem description, its corresponding code, and a complexity label.
[0015] Based on the large language model, the problem description in each training data is converted into pseudocode;
[0016] For each training data, we extract complexity indicators based on the problem description and pseudocode, and construct prompt words based on the complexity indicators.
[0017] The large language model is trained with problem description, pseudocode and prompt words as input and task complexity as output.
[0018] In some embodiments, a method for labeling a complexity tag includes:
[0019] For each piece of training data, based on the set large language model, multiple independent code snippets are repeatedly generated according to the problem description;
[0020] Input the test case input data into each code snippet, run the code, and record the output results. If the output results are consistent with expectations, the code snippet passes the test, and the pass rate of multiple code snippets corresponding to the problem description is obtained;
[0021] The complexity label of the problem description is determined according to the pass rate.
[0022] In some embodiments, the complexity evaluation model is obtained by fine-tuning a given large language model through LoRA.
[0023] In some embodiments, a cross-entropy loss function is used during the training of the complexity assessment model, and asymmetric penalty weights are introduced into the cross-entropy loss function to distinguish the costs of different types of misjudgments; among them, the weight is the largest when a difficult task is misjudged as a simple task, followed by a difficult task being misjudged as a medium task and a medium task being misjudged as a simple task.
[0024] In some embodiments, a set reasoning model is adopted to execute a code generation task based on a matching reasoning pattern, specifically including: generating a prompt word for the code generation task, wherein the prompt word includes a reasoning pattern label; taking the problem description and the prompt word as input, so that the reasoning model adopts the specified reasoning pattern to execute the code generation task.
[0025] In some embodiments, when the complexity assessment result of the code generation task is difficult, a test case is also generated based on a large language model; after the code is generated based on the difficult mode, the generated code is tested based on the test case. If the test fails, the test error information feedback is appended to the prompt word, requiring the model to analyze the cause of the error and regenerate the code until the generated code snippet passes the test or reaches the preset iteration limit.
[0026] A second aspect of the present invention provides a code generation task adaptive reasoning device, comprising:
[0027] a pseudocode generation module configured to obtain a problem description of a code generation task and convert the problem description into pseudocode based on a large language model;
[0028] a complexity evaluation module configured to perform a complexity evaluation on the code generation task based on the problem description and the corresponding pseudocode and a pre-trained complexity evaluation model; the complexity evaluation model is trained with the problem description and the pseudocode as input and the task complexity as output;
[0029] The adaptive reasoning module is configured to match the reasoning mode according to the complexity evaluation result, adopt the set reasoning model, and perform the code generation task based on the matched reasoning mode; among which, different complexities correspond to reasoning modes with different reasoning depths.
[0030] A third aspect of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the method described.
[0031] The one or more technical solutions above realize dynamic perception of task complexity of the large language model through a complexity evaluation model, and through a complexity mapping inference mode, different inference modes correspond to different inference depths, so that the inference depth and resource allocation can be adaptively adjusted according to actual requirements of a task, and unnecessary waste of computing resources is avoided. BRIEF DESCRIPTION OF DRAWINGS
[0032] The accompanying drawings, which form a part of the specification, are included to provide a further understanding of the application and are incorporated herein in their entirety. The embodiments illustrated in the drawings are provided merely as examples of the present application and therefore should not be construed as to limit the scope of the present application.
[0033] Figure 1 A schematic diagram of a computer system is shown according to an example embodiment of the present application;
[0034] Figure 2 A flowchart of a code generation task adaptive inference method is shown according to an example embodiment of the present application;
[0035] Figure 3 A simplified schematic diagram of a code generation task adaptive inference method is shown according to an example embodiment of the present application;
[0036] Figure 4 A flowchart of complexity evaluation model construction and training is shown according to an example embodiment of the present application;
[0037] Figure 5 A schematic diagram of an inference mode performing a code generation task based on complexity mapping is shown according to an example embodiment of the present application;
[0038] Figure 6 A flowchart of a code generation task adaptive inference method is shown according to another example embodiment of the present application;
[0039] Figure 7 A structural block diagram of a code generation task adaptive inference device is shown according to an example embodiment of the present application. DETAILED DESCRIPTION
[0040] Embodiments of the present application will be described in more detail by referring to the attached drawings. Although certain embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are for exemplary purposes only and are not intended to limit the scope of protection of the present application.
[0041] In the description of the embodiments of the present application, the term “including” and similar terms should be understood as open inclusion, that is, “including but not limited to.” The term “based on” should be understood as “at least partially based on.”
[0042] As described in the background, existing code generation tasks based on large language models suffer from significant overthinking, resulting in low reasoning efficiency and high error rates. To effectively alleviate this overthinking problem during the reasoning phase, one or more embodiments of the present invention propose an adaptive reasoning method based on dynamic perception of task complexity, which can adaptively adjust reasoning depth and resource allocation based on the actual needs of the task.
[0043] Figure 1 The following is a block diagram of a computer system 100 according to an exemplary embodiment of the present application. Computer system 100 includes a server 110 and a terminal 120. Server 110 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal 120 can be an electronic device such as a mobile phone, a tablet computer, or a personal computer (PC).
[0044] The server 110 and the terminal 120 may communicate with each other via a network, such as a wired or wireless network.
[0045] In some embodiments, when a large language model is deployed on server 110, terminal 120 can obtain a problem description for a code generation task and send the obtained information to server 110 via a wireless or wired network. Server 110 performs a complexity assessment on the code generation task, determines an appropriate reasoning model, and executes reasoning to obtain the code corresponding to the code generation task. Finally, server 110 sends the generated code to terminal 120, which receives and presents the code.
[0046] In some embodiments, the large language model is deployed on server 110 or terminal 120. The server 110 or terminal 120 obtains the problem description of the code generation task, performs a complexity assessment on the code generation task, determines an appropriate reasoning mode, and performs reasoning to obtain the code corresponding to the code generation task. Finally, the server 110 or terminal 120 presents the generated code.
[0047] Figure 2 A flowchart of a code generation task adaptive reasoning method provided by an exemplary embodiment of the present application is shown, wherein the method is executed by a server or a terminal, which may be Figure 1 The server or terminal shown. The method comprises the following steps:
[0048] S210: Obtain a problem description of a code generation task, and convert the problem description into pseudocode based on a large language model.
[0049] S220: According to the problem description and the corresponding pseudocode, based on a pre-trained complexity evaluation model, the complexity evaluation model is trained with the problem description and the pseudocode as input and the task complexity as output.
[0050] S230: Match the reasoning mode according to the complexity evaluation result, adopt the set reasoning model, and execute the code generation task based on the matched reasoning mode; wherein different complexities correspond to reasoning modes with different reasoning depths.
[0051] Based on this, the complexity evaluation model is used to realize the dynamic perception of task complexity by the large language model. Through the complexity mapping reasoning mode, different reasoning modes correspond to different reasoning depths, enabling it to adaptively adjust the reasoning depth and resource allocation according to the actual needs of the task, avoiding unnecessary waste of computing resources.
[0052] In step S210, the requirement description can be a code function requirement entered by the user, such as a few sentences of requirement description language containing the required function entered by the user through a dialog box, or a requirement description document entered by the user. First, the problem description is converted into corresponding pseudocode by a large model, for example, using the Qwen2.5-Coder-7B model. It is understood that the above process can be implemented using prompt words.
[0053] In step S220, the problem description and pseudocode are input into a trained complexity assessment model. The model is based on the Qwen2.5-Coder-7B model and is trained using a pre-built training dataset through LoRA fine-tuning. Complexity levels include easy, medium, and difficult.
[0054] The training of the complexity evaluation model is implemented through steps S221-S224.
[0055] S221: Obtain a training dataset. Each piece of training data corresponds to a code generation task and a complexity label. Each code generation task includes a problem description and its corresponding code. For example, the training dataset is constructed based on the open source datasets APPS and TACO. APPS contains 10,000 samples, covering a variety of programming tasks from entry-level to competition-level. TACO is aimed at more challenging programming competition topics and contains a total of 26,443 samples. The samples of both datasets contain information such as problem description, code, and test cases. The programming language is Python. Problem descriptions and their corresponding code test cases are collected from the open source datasets to construct a training dataset. For example, 5,000 samples are evenly collected from the above datasets to construct the training set.
[0056] Generate a complexity label for each sample in the training dataset. In some embodiments, the complexity is automatically labeled based on the code generation success rate of the large model on the sample. The specific method is as follows:
[0057] (1) For each piece of training data, based on the set large language model, multiple independent code snippets are repeatedly generated according to the problem description. For example, 5,000 collected problem descriptions are input into the GPT-4o model, and 9 independent code snippets are repeatedly generated for each problem description;
[0058] (2) Input the test case input data into each code snippet, run the code, and record the output results. If the output results are consistent with expectations, the code snippet passes the test, and the pass rate of the multiple code snippets corresponding to the problem description is obtained. For example, use the sample test case to test the 9 code snippets generated by the model, and calculate the pass rate: the number of code snippets that passed the test / 9.
[0059] (3) Determine the complexity label based on the pass rate. For example, set three levels of complexity labels: easy, medium, and difficult. A pass rate greater than 2 / 3 is labeled "easy"; a pass rate between 1 / 3 and 2 / 3 is labeled "medium"; and a pass rate less than 1 / 3 is labeled "difficult."
[0060] Based on this, a labeled training dataset is obtained, the input is the problem description, and the output is the task complexity.
[0061] However, there is a semantic gap between natural language descriptions and code, and judging complexity based solely on the problem description can easily lead to misjudgments. Therefore, in some embodiments, pseudocode corresponding to the problem description is added to the input of the training dataset, leveraging its structured information along with the unstructured information of the problem description to improve assessment accuracy.
[0062] S222: Based on the large language model, the problem description in each training data is converted into pseudocode. It should be noted that pseudocode is a form between natural language and code and has no strict definition. However, the more specific the pseudocode generated by the large language model is, the worse the generation accuracy will be. Therefore, by limiting the language mixing style in the prompt words and giving structural examples, the complexity of the pseudocode output is controlled to ensure the correctness of the pseudocode output. The language mixing style is, for example: using Chinese natural language to describe core operations, allowed basic programming keywords, and prohibited specific syntax; the structural specification is a basic structure containing basic programming keywords. By limiting the language mixing style and giving structural examples, the output pseudocode can be made closer to natural language, while retaining the logical framework of the algorithm, which can ensure the accuracy of the generated pseudocode.
[0063] For example, the problem description is converted into pseudocode using the Qwen2.5-Coder-7B-Instruct model through prompt word engineering. Based on this, the training dataset obtained has the problem description and pseudocode as input and the task complexity as output.
[0064] S223: For each piece of training data, complexity indicators are extracted based on the problem description and pseudocode, and prompt words are constructed based on the complexity indicators; wherein, the complexity indicators include the semantic complexity of the problem description, the number of pseudocode lines, the number of conditional branches, the number of loop structures, and the maximum number of loop nesting layers, etc., which are used to guide the model to analyze complexity according to key dimensions.
[0065] S224: Take the problem description, pseudocode and prompt words as input and the task complexity as output to train the set large language model.
[0066] The large language model uses Qwen2.5-Coder-7B-Instruct as the base model for complexity evaluation. The model is trained using a pre-built training dataset through LoRA fine-tuning. Specifically, a LoRA adapter is added to the base model. During training, all original parameters of the base model are frozen, and only the parameters of the LoRA adapter are updated. It is important to note that the same base model (Qwen2.5-Coder-7B-Instruct) can be used for pseudocode generation and complexity evaluation. The original model can be used for pseudocode generation; the trained LoRA adapter must be loaded for complexity evaluation.
[0067] Furthermore, different types of complexity misjudgments carry different costs: for example, misjudging a difficult problem as simple will cause the large model to directly output an error code, which has far more serious consequences than misjudging a simple problem as difficult. Therefore, some embodiments also optimize the loss function, setting asymmetric penalty weights for different misjudgment categories. Specifically, the loss function is:
[0068]
[0069] in, is the sample size; is the number of categories. For example, when the complexity is divided into difficult, medium, and easy, the value is 3; is the one-hot encoding of the sample label. When the true label of the sample is i, Take 1, otherwise Take 0; is the probability that the model predicts that the sample is category i; is the loss weight, which is determined by the true category and predicted category of the sample; j is the model prediction label, that is, The category corresponding to the maximum value in .
[0070] The true category of the sample is represented as k, then when i=k, , when i≠k, Will The value of is brought into the above loss function, and the formula can be simplified to:
[0071] ; ;
[0072] Where k is the true category of the sample and j is the model predicted category.
[0073] The cost is highest when a difficult problem is predicted to be an easy problem, and the weight is set to 5; when a difficult problem is predicted to be medium, and a medium problem is predicted to be easy, the weight is set to 3; in other cases, the weight is set to 1. That is,
[0074] .
[0075] By inputting the problem description as unstructured information and the pseudocode as structured information, supplemented by prompt words that integrate complexity indicators, the evaluation accuracy is improved; in addition, the use of Lora fine-tuning technology helps to efficiently train the complexity evaluation model, and by introducing asymmetric loss function weights, the probability of high-risk misjudgments is reduced.
[0076] In step S203, the reasoning mode is matched according to the complexity evaluation result, specifically: if the complexity evaluation result is simple, the reasoning mode corresponds to the no reasoning mode; if the complexity evaluation result is moderate, the reasoning mode corresponds to the normal reasoning mode; if the complexity evaluation result is complex, the reasoning mode corresponds to the detailed reasoning mode.
[0077] The code generation task is executed based on a specified reasoning model and a matching reasoning pattern. Specifically, a prompt word is generated for the code generation task, the prompt word including the reasoning pattern label; the problem description and the prompt word are used as input, causing the reasoning model to execute the code generation task using the specified reasoning pattern. It is understood that pseudocode can also be generated based on the problem description, using the problem description, pseudocode, and prompt word as input. The pseudocode generation process here is the same as in S222 and will not be further described here.
[0078] Specifically, the reasoning mode label corresponding to the no-reasoning mode is " / no_think", the normal reasoning mode label is " / think", and the detailed reasoning mode labels include " / think" and "problem description decomposition". For example, the hybrid reasoning model Qwen3-32B is used for code generation tasks. The hybrid reasoning mechanism of Qwen3 supports the soft switching mechanism of "thinking" and "non-thinking" modes, that is, adding / think or / no_think tags in the prompt words to switch the reasoning mode. When the / no_think tag is added to the prompt word, the "non-thinking" mode outputs an empty string. <think>< / think> Tags, skip the inference process. The specific configuration is as follows:
[0079] (1) Simple mode
[0080] Add the / no_think tag to the prompt to enable "no-think" mode, which skips the reasoning process and directly outputs the final code snippet to avoid wasting computing resources on simple problems. The maximum generation length of the model (max_new_tokens) is 2048.
[0081] (2) Medium mode
[0082] Add the / think tag to the prompt to enable "thinking" mode, which outputs a clear reasoning path and final code snippet. The maximum generated model length is 4096.
[0083] (3) Hard Mode
[0084] The prompt word adds the tag / think and explicitly requires the model to perform detailed problem decomposition in the prompt word; the maximum generated model length is: 16384.
[0085] When the complexity assessment result of the code generation task is difficult, test cases are generated based on the large language model, and the generated code is optimized through multiple rounds of feedback iteration. Specifically, after generating code based on the difficult mode, the generated code is tested based on the test cases. If the test fails, the test error information feedback is appended to the prompt word, requiring the model to analyze the cause of the error and regenerate the code until the generated code snippet passes the test or reaches the preset iteration limit.
[0086] The following is an example prompt for generating test cases based on a large model: You are a senior test engineer. Please design a test case for the following programming problem: [Problem Description]: {question}; [Pseudocode]: {pseudo_code}; Requirements: Include three normal test cases (covering typical scenarios), three edge case flows (covering empty input, extreme values, illegal data types, etc.), and output in JSON format: {{"tests":[{{"input":...,"output":...}}]}}. When testing the code, use the test case input as input. The test passes when the output matches the test case output.
[0087] One or more of the above-mentioned embodiments adopt an adaptive reasoning mechanism, analyze the complexity of the code problem description through a complexity assessment model, and divide it into three levels: simple, medium and difficult. The reasoning mode is then dynamically adjusted according to the evaluation results. While maintaining the code generation effect, the reasoning efficiency is significantly improved (saving approximately 70% of the reasoning length), achieving an optimal balance between accuracy and reasoning efficiency, and laying a theoretical and practical foundation for the efficient application of large reasoning models in actual software development.
[0088] It is understandable that sometimes the functions required by users can only be realized after other functions are realized. When the requirements are complex, it is necessary to parse the path required for function realization and decompose it into multiple subtasks to improve the accuracy of code generation. The implementation complexity of each subtask also varies. In order to realize the automatic code generation of complex tasks, some embodiments also provide the following alternative solutions, which can decompose complex tasks into subtasks and match appropriate reasoning modes for each subtask, such as Figure 3 As shown, specifically including:
[0089] S310: After obtaining the problem description of the code generation task, decompose the subtasks based on the problem description to obtain the problem description of each subtask;
[0090] S320: For each subtask, convert the corresponding problem description into pseudocode based on the large language model;
[0091] S330: According to the problem description and the corresponding pseudo code of each subtask, the complexity of the subtask is evaluated based on a pre-trained complexity evaluation model.
[0092] S340: According to the complexity evaluation result, each subtask is matched with an inference mode, and the code generation task is executed according to the corresponding inference mode according to the execution order of the subtasks; wherein different complexity corresponds to different inference modes with different inference depths.
[0093] Based on this, for a complex code generation task, the inference mode of each subtask can be dynamically adjusted according to the complexity of each subtask, thereby improving the overall execution efficiency of the code generation task.
[0094] The subtask decomposition based on the problem description in step S310 can divide the task by steps or by functional modules, and in addition, the dependency relationship between multiple tasks, i.e., the execution order of the subtasks, needs to be obtained.
[0095] The processes of processing and complexity evaluation of the subtasks, and inference mode matching in steps S320-S340 are the same as those in steps S210-S230, which will not be described in detail here.
[0096] Figure 5 A schematic diagram of an apparatus provided by one or more embodiments of the present application is shown. The apparatus includes: a pseudo code generation module 410 configured to obtain a problem description of a code generation task, and convert the problem description into a pseudo code based on a large language model; a complexity evaluation module 420 configured to evaluate the complexity of the code generation task based on a pre-trained complexity evaluation model according to the problem description and the corresponding pseudo code; the complexity evaluation model is trained to take the problem description and the pseudo code as input and the task complexity as output; an adaptive inference module 430 configured to match an inference mode according to the complexity evaluation result, adopt a set inference model, and execute the code generation task based on the matched inference mode; wherein different complexity corresponds to different inference modes with different inference depths.
[0097] One or more embodiments of the present application also provide an electronic device which can be used to implement the method in the above embodiments. The electronic device includes one or more processors, one or more memories coupled to the processors, and a communication module coupled to the processors.
[0098] The memory in the embodiments of the present application is used to store various types of data to support the execution of the methods as shown in Figure 1
[0099] It is understood that the memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The memory in the embodiment of the present invention can store the following: Figure 1 The computer programs corresponding to the steps in the method shown in . The operating system includes various system programs, such as a framework layer, a core library layer, and a driver layer, which are used to implement various basic services and handle hardware-based tasks. The application program can include various application programs.
[0100] As an example, a processor can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0101] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, the computer program including a computer program for executing Figure 1 In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion and / or installed from a removable medium. When the computer program is executed by the central processing unit, the various functions defined in the apparatus of the present application are performed.
[0102] in, Figure 1 The computer program instructions corresponding to the method shown can also be stored in a computer readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0103] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A code generation task adaptive reasoning method, characterized in that: The following steps are involved: Obtain a problem description for a code generation task, and convert the problem description into pseudocode based on a large language model; According to the problem description and the corresponding pseudocode, based on the pre-trained complexity evaluation model, the complexity evaluation of the code generation task is performed; The complexity assessment model is trained with problem description and pseudocode as input and task complexity as output; The complexity assessment model training method includes: Get the training data set. Each piece of training data includes a problem description, its corresponding code, and a complexity label. Based on the large language model, the problem description in each training data is converted into pseudocode; For each training data, we extract complexity indicators based on the problem description and pseudocode, and construct prompt words based on the complexity indicators. Take the problem description, pseudocode and prompt words as input and the task complexity as output to train a large language model; The complexity assessment model is trained using a cross-entropy loss function, which introduces asymmetric penalty weights to distinguish the costs of different types of misjudgments. The weight is highest when a difficult task is misjudged as an easy task, followed by a difficult task misjudged as a medium task and a medium task misjudged as an easy task. According to the complexity evaluation results, the reasoning mode is matched, and the set reasoning model is adopted to perform the code generation task based on the matched reasoning mode; among them, different complexities correspond to reasoning modes with different reasoning depths.
2. The code generation task adaptive reasoning method according to claim 1, characterized in that: After obtaining the problem description of the code generation task, subtask decomposition is performed based on the problem description to obtain a problem description of each subtask; For each subtask, based on the large language model, convert the corresponding problem description into pseudocode; Based on the problem description and corresponding pseudocode of each subtask, the complexity of the subtask is evaluated based on the pre-trained complexity evaluation model; According to the complexity evaluation results, the reasoning pattern is matched for each subtask, and the code generation tasks are executed in sequence according to the corresponding reasoning pattern according to the execution order of the subtasks.
3. The code generation task adaptive reasoning method according to claim 1, characterized in that: The annotation methods of complexity labels include: For each piece of training data, based on the set large language model, multiple independent code snippets are repeatedly generated according to the problem description; Input the test case input data into each code snippet, run the code, and record the output results. If the output results are consistent with expectations, the code snippet passes the test, and the pass rate of multiple code snippets corresponding to the problem description is obtained; The complexity label of the problem description is determined according to the pass rate.
4. The code generation task adaptive reasoning method according to claim 3, characterized in that: The complexity evaluation model is obtained by fine-tuning a given large language model through LoRA.
5. The code generation task adaptive reasoning method according to claim 1 or 2, characterized in that: A set reasoning model is adopted to execute a code generation task based on a matching reasoning pattern, specifically including: generating a prompt word for the code generation task, wherein the prompt word includes a reasoning pattern label; taking the problem description and the prompt word as input, so that the reasoning model adopts the specified reasoning pattern to execute the code generation task.
6. The code generation task adaptive reasoning method according to claim 1, characterized in that: When the complexity evaluation result of the code generation task is difficult, test cases are also generated based on the large language model; After generating code based on the difficult mode, the generated code is tested based on the test case. If the test fails, the test error information feedback is appended to the prompt word, requiring the model to analyze the cause of the error and regenerate the code until the generated code snippet passes the test or reaches the preset iteration limit.
7. A code generation task adaptive reasoning device, characterized in that include: a pseudocode generation module configured to obtain a problem description of a code generation task and convert the problem description into pseudocode based on a large language model; A complexity evaluation module is configured to perform a complexity evaluation on the code generation task based on the problem description and the corresponding pseudocode and based on a pre-trained complexity evaluation model; The complexity assessment model is trained with problem description and pseudocode as input and task complexity as output; The complexity assessment model training method includes: Get the training data set. Each piece of training data includes a problem description, its corresponding code, and a complexity label. Based on the large language model, the problem description in each training data is converted into pseudocode; For each training data, we extract complexity indicators based on the problem description and pseudocode, and construct prompt words based on the complexity indicators. Take the problem description, pseudocode and prompt words as input and the task complexity as output to train a large language model; The complexity assessment model is trained using a cross-entropy loss function, which introduces asymmetric penalty weights to distinguish the costs of different types of misjudgments. The weight is highest when a difficult task is misjudged as an easy task, followed by a difficult task misjudged as a medium task and a medium task misjudged as an easy task. The adaptive reasoning module is configured to match the reasoning mode according to the complexity evaluation result, adopt the set reasoning model, and perform the code generation task based on the matched reasoning mode; among which, different complexities correspond to reasoning modes with different reasoning depths.
8. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein computer instructions are stored in the memory. When the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Code generation method and related device
WO2025123711A1