Large language model multi-agent collaborative project-level code generation method and system
Patent Information
- Application Number
- CN202611316777.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0002]近年来,大语言模型推动代码生成技术快速发展,可依据自然语言描述生成程序代码,在函数、文件级代码生成场景得到有效应用,但难以适配包含多模块、跨文件依赖与长上下文约束的项目级代码生成需求
本发明通过任务规划智能体对项目需求进行递归拆解与依赖感知,能够将复杂编程任务自动转化为结构清晰、依赖明确任务依赖图,有效降低了人工分解项目的认知负担,提高了大规模代码生成的可管理性和准确性。同时,采用动态调度机制依据就绪任务优先级和执行成本灵活调整执行顺序,提升了多任务并行处理效率和资源利用率。在执行环节,通过本地缓存库检索相似历史任务并复用参考代码,避免了重复开发,大幅缩短了代码生成周期。另外,结合测试智能体的自动验证与迭代修复机制,能够精准定位错误来源,并自动触发重规划或重构,形成规划-生成-测试-优化的闭环反馈,有效减少了人工调试成本,提高了最终代码的正确性和稳定性。此外,引入复合奖励函数对拆解质量和执行效能进行持续优化,进一步保障了任务规划的逻辑合理性和整体执行效率,使得本方案在应对大型、多依赖项目时具备显著的智能化优势和高鲁棒性。
Smart Images

Figure CN122837809A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and software engineering technology, and in particular to a method and system for generating project-level code for multi-agent collaborative programming of large language models. Background Technology
[0002] In recent years, large language models have driven the rapid development of code generation technology, enabling the generation of program code based on natural language descriptions. This has proven effective in function and file-level code generation scenarios, but it struggles to meet the demands of project-level code generation, which often involves multiple modules, cross-file dependencies, and long context constraints. Existing solutions often employ single-agent or loosely collaborative architectures for requirement analysis, code generation, and defect fixing. As project scales, these approaches become increasingly limited by context length, leading to issues such as inadequate task planning, cross-module interface conflicts, and insufficient output stability.
[0003] While multi-agent frameworks enhance the ability to handle complex tasks through multi-agent collaboration, they still have significant shortcomings: they lack dependency-aware task planning strategies for project development, resulting in weak global coordination capabilities; they cannot effectively reuse historical development task experience, leading to repetitive generation and low efficiency; code repair is limited to local modifications, lacking backtracking of upper-level task logic, resulting in high trial-and-error costs; and the lack of standardized collaboration mechanisms among agents makes it difficult to efficiently link task planning, coding, and testing iteration processes. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a multi-agent collaborative project-level code generation method and system for large language models, which improves the correctness, stability, generation efficiency, and code style consistency of project-level code generation.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for multi-agent collaborative project-level code generation for large language models, comprising: Obtain project requirement information input by the user; A task planning agent is used to recursively decompose project requirements into subtasks, generating multi-level subtasks and constructing a task dependency graph. Ready subtasks are dynamically scheduled based on the task dependency graph to generate a task execution plan. A composite reward function is used to optimize the decomposition and scheduling actions. According to the task execution plan, the execution agent is invoked to extract the task feature information of the current subtask, retrieve and match similar historical tasks, and generate the target code corresponding to the current subtask. Based on the task dependency graph, the target code corresponding to each subtask is assembled to generate the complete project code; The test agent is used to test and verify the complete project code. If the preset requirements are not met, the task planning stage or the code implementation stage is refactored according to the source of the error, and the test and verification are carried out again until the test results meet the preset requirements and the final code is generated.
[0006] Secondly, the present invention provides a large language model multi-agent collaborative project-level code generation system, comprising: The task acquisition module is configured to acquire project requirement information input by the user. The information analysis module is configured to use a task planning agent to recursively decompose project requirement information into sub-tasks, generate multi-level sub-tasks, and construct a task dependency graph; dynamically schedule ready sub-tasks according to the task dependency graph to generate a task execution plan; and optimize the decomposition and scheduling actions using a composite reward function. The sub-code generation module is configured to, according to the task execution plan, call the execution agent to extract the task feature information of the current sub-task, search for matching similar historical tasks, and generate the target code corresponding to the current sub-task; The code assembly module is configured to assemble the target code corresponding to each subtask according to the task dependency graph, and generate complete project code. The testing and generation module is configured to use a testing agent to test and verify the complete project code. If the preset requirements are not met, the task planning stage or the code implementation stage will be refactored according to the source of the error, and the testing and verification will be carried out again until the test results meet the preset requirements and the final code is generated.
[0007] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for generating project-level code for a large language model multi-agent collaborative project as described in the first aspect.
[0008] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the method for generating project-level code for a large language model multi-agent collaborative project as described in the first aspect.
[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes a task planning agent to recursively decompose and dependency-awarely manage project requirements. This automatically transforms complex programming tasks into clearly structured, dependency-defined task dependency graphs, effectively reducing the cognitive burden of manual project decomposition and improving the manageability and accuracy of large-scale code generation. Simultaneously, a dynamic scheduling mechanism flexibly adjusts the execution order based on the priority and execution cost of ready tasks, enhancing the efficiency and resource utilization of multi-task parallel processing. During execution, similar historical tasks are retrieved from a local cache library, and reference code is reused, avoiding redundant development and significantly shortening the code generation cycle. Furthermore, combined with the automatic verification and iterative repair mechanism of the testing agent, the source of errors can be accurately located, automatically triggering replanning or refactoring, forming a closed-loop feedback loop of planning-generation-testing-optimization. This effectively reduces manual debugging costs and improves the correctness and stability of the final code. Moreover, the introduction of a composite reward function continuously optimizes the decomposition quality and execution efficiency, further ensuring the logical rationality of task planning and overall execution efficiency. This makes this solution possess significant intelligent advantages and high robustness when dealing with large, multi-dependency projects.
[0010] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.
[0011] Figure 1 The main flowchart of a multi-agent collaborative project-level code generation method for a large language model provided in this embodiment of the invention; Figure 2 A flowchart illustrating a method for multi-agent collaborative project-level code generation using a large language model, provided in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the task tree construction and dynamic scheduling process of a task planning agent provided in an embodiment of the present invention. Figure 4 A flowchart of task cache retrieval and similar code reuse for the execution agent provided in an embodiment of the present invention; Figure 5 The flowchart illustrates the logical replanning and code refactoring of the test agent provided in this embodiment of the invention. Detailed Implementation
[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0013] Example 1 like Figure 1As shown in the figure, this embodiment discloses a method for multi-agent collaborative project-level code generation for large language models, including the following steps: S1: Obtain project requirement information input by the user; S2: Utilize a task planning agent to recursively decompose project requirements into sub-tasks, generating multi-level sub-tasks and constructing a task dependency graph; dynamically schedule ready sub-tasks based on the task dependency graph to generate a task execution plan; wherein, a composite reward function is used to optimize the decomposition and scheduling actions; S3: Based on the task execution plan, call the execution agent to extract the task feature information of the current subtask, search for and match similar historical tasks, and generate the target code corresponding to the current subtask; S4: Assemble the target code corresponding to each subtask according to the task dependency graph to generate complete project code; S5: Use a test agent to test and verify the complete project code. If the preset requirements are not met, refactor the task planning stage or the code implementation stage according to the source of the error, and retest and verify until the test results meet the preset requirements and generate the final code.
[0014] Next, combined Figure 2 This embodiment provides a detailed description of a multi-agent collaborative project-level code generation method for large language models. This embodiment achieves automated generation and iterative optimization of complex software projects by constructing a collaborative framework consisting of a task planning agent, an execution agent, and a testing agent.
[0015] In step S1, the system receives project requirement information input by the user. This project requirement information can be software functional requirements, system development requirements, module development requirements, interface development requirements, or other forms of software development tasks. The project requirement information may include functional objectives, business rules, technology stack requirements, interface constraints, input / output formats, and existing project context information.
[0016] In step S2, in order to break down the solution granularity, clarify the execution sequence and data dependencies, the task planning agent is used to recursively decompose the project requirements and perform dependency-aware planning, transforming the complex overall project requirements into sub-tasks with dependency relationships that can be implemented independently.
[0017] Specifically, such as Figure 3As shown, the task planning agent first performs semantic parsing on the project requirement information input by the user, extracting requirement constraint information, including: functional objectives, business rules, technology stack constraints, input / output requirements, and interface constraints. Based on the extraction results, it determines whether the current project requirement type is a single subtask or a complex multitasking task according to preset judgment conditions. If it is a complex multitasking task, it performs recursive task decomposition. The complex multitasking task includes requirements to be decomposed or complex tasks.
[0018] The "requirements to be decomposed" refer to project requirements that include two or more functional objectives, involve multiple code files, have multiple interface call relationships, or require multiple execution steps to complete. The "composite task" refers to a development task that simultaneously includes at least two types of tasks: data processing, interface implementation, business logic implementation, file organization, module calling, or testing and verification. For example, when a project requirement simultaneously demands the implementation of a user login interface, a database access module, and permission verification logic, the task planning agent will classify this project requirement as a composite task.
[0019] The task planning agent identifies requirements to be broken down or complex tasks based on preset judgment conditions. These preset judgment conditions include: the project requirement contains multiple functional verbs or functional objectives; the project requirement involves multiple entity objects or data structures; the project requirement involves multiple code files, modules, or interfaces; the project requirement has a clear sequential execution relationship, input-output dependency relationship, or calling relationship; or the task planning agent determines that a single subtask cannot independently complete the project requirement. When at least one of the above judgment conditions is met, it is determined that the current project requirement has requirements to be broken down or complex tasks; otherwise, it is determined that the current project requirement can enter the dependency analysis process as a single subtask.
[0020] When there are requirements to be broken down or complex tasks, the task planning agent performs recursive task decomposition. Specifically, the task planning agent first breaks down the project requirements into several first-level subtasks, and generates a task identifier, task description, input information, output information, dependent objects, and execution constraints for each first-level subtask. Then, it determines whether each first-level subtask still belongs to the requirements to be broken down or complex tasks. If a certain first-level subtask still meets the conditions for being broken down, it continues to be broken down into the next level of subtasks. If a certain subtask can be directly generated by the execution agent, the further decomposition of that subtask is stopped, and it is designated as a leaf task node. Through the above recursive decomposition process, the task planning agent constructs a subtask tree containing a root task node, intermediate task nodes, and leaf task nodes.
[0021] Furthermore, the task planning agent analyzes the data dependencies, module call relationships, and execution constraints among the subtasks, and constructs a task dependency graph accordingly.
[0022] Among them, data dependency refers to the output data, data structure, configuration file or interface definition of one subtask being used as input by another subtask; module call relationship refers to the function, class, interface or service generated by one subtask being called by another subtask; execution constraint relationship refers to the sequential execution relationship caused by technology stack constraints, file generation order, environment configuration order or testing order.
[0023] The task dependency graph includes multiple task nodes and dependency edges connecting the task nodes. Task nodes are used to represent subtasks, and dependency edges are used to represent data dependencies, module call relationships, or execution constraint relationships between subtasks.
[0024] In this embodiment, by introducing three different types of constraint relationships, the logical connections between subtasks can be comprehensively depicted, overcoming the problem of missing relationships caused by simply dividing tasks according to their order. This clearly represents data flow, module calls, and execution timing constraints. It improves the completeness of task dependency identification and provides a reliable basis for ready task judgment and dynamic scheduling of subtasks.
[0025] Based on the task dependency graph, it is determined whether there are ready tasks that meet the execution conditions. When ready tasks exist, they are dynamically scheduled based on the dependency relationships to generate a task execution plan.
[0026] Specifically, the system first filters ready tasks that meet the execution conditions from the task dependency graph; then, it sorts the ready tasks according to subtask priority, dependency information, the number of unlockable downstream tasks, estimated execution cost, and historical task reuse; finally, it selects the subtasks with the highest ranking to generate task execution plans, which are then handed over to the execution agent for processing. When no ready tasks exist, the system waits for upstream tasks to complete, or updates task dependencies based on test feedback and reschedules.
[0027] The execution conditions are as follows: all prerequisite tasks of the current subtask have been completed, the input information, interface definitions, data structures, configuration files, or context information required by the current subtask have been generated, and there are no unresolved upstream blocking relationships for the current subtask. When all prerequisites of a task node are marked as completed, the task node is added to the ready task set.
[0028] It should be understood that prerequisite tasks are predefined tasks that must be executed sequentially, belonging to static explicit dependencies; upstream blocking relationships are implicit constraints dynamically generated at runtime (resource consumption, data verification anomalies, failure to meet external conditions, etc.), independent of whether the prerequisite tasks are completed. It should be understood that prerequisite tasks and upstream blocking relationships can be defined by those skilled in the art. Furthermore, subtask priorities, the number of unlockable downstream tasks, estimated execution costs, and historical task reuse information are readily available to those skilled in the art. For example, subtask priorities originate from task definition configurations or dynamic parsing of requirements; the number of unlockable downstream tasks is calculated through topological analysis of the task dependency graph; estimated execution costs are obtained from historical execution logs, performance monitoring, or resource evaluation models; and historical task reuse information is obtained from task fingerprint comparison results in the task caching system and execution record database. Simultaneously, the determination of priorities can be set by those skilled in the art according to actual needs.
[0029] In this embodiment, ready tasks are identified based on the task dependency graph, enabling reasonable dynamic scheduling of subtasks. This avoids errors such as executing downstream tasks before the preceding tasks are completed, reduces issues like undefined variables, missing interfaces, and disordered file order, ensures that subtasks are executed in an orderly manner according to reasonable logic, reduces the risk of logical breakage during code generation, and improves the reliability and execution efficiency of the overall project coding.
[0030] Furthermore, to facilitate modeling the project-level task planning process, as one implementation method, the task planning agent formalizes the project planning process into a task planning process. At any time t, the state of the task planning agent... Represented as: ; in, This represents the currently dynamically evolving task dependency graph; This represents a queue of demands to be broken down or processed. It is generated by the task planning agent during the recursive task breakdown process and is used to record task nodes that have not yet been broken down or are waiting to be executed. It represents project context information, which consists of project requirements, technology stack constraints, existing codebase summary, and historical task execution records.
[0031] The actions of a task planning agent include decomposing actions and scheduling actions, and the action set is represented as follows: ; in, This indicates a decomposition action, used to break down complex tasks or high-level requirements into multiple sub-tasks and their internal dependencies. This represents a scheduling action used to select the next task to be executed from the set of ready subtasks that meet the execution conditions, and hand it over to the executing agent for processing.
[0032] In some embodiments, the task planning agent is trained using a hybrid strategy optimization approach. This hybrid strategy optimization includes supervised fine-tuning, preference optimization, and online strategy optimization. Specifically, the supervised fine-tuning approach involves training the task planning agent during its initialization phase, enabling it to learn task decomposition patterns based on historical project task data. The preference optimization approach involves training the task planning agent during the task planning scheme optimization phase, based on the quality differences between different task decomposition schemes, to select task decomposition schemes with more complete coverage and more reasonable dependencies. The online strategy optimization approach involves dynamically adjusting the task decomposition and task scheduling strategies during the project code generation and execution phase, based on the code generation results returned by the execution agent and the error information reported by the testing agent.
[0033] Furthermore, in order to simultaneously optimize both task decomposition quality and task execution efficiency, the task planning agent employs a composite reward function for optimization, which is expressed as: ; in, Indicates the total reward. This indicates a reward for the quality of dismantling. γ represents the performance incentive, and γ represents the discount factor.
[0034] Disassembly quality bonus is represented as follows: ; in, This indicates the degree to which the set of subtasks covers the original project requirements. It is calculated based on the matching relationship between the number of functional objectives contained in the generated subtasks and the number of functional objectives in the original requirements. This represents the subtask granularity score, which is determined based on whether the subtask meets the independent execution condition, the completeness of the task description, and the task size. This indicates the degree of redundancy between subtasks, and it is calculated based on the degree of overlap in the functional descriptions and implementation goals of different subtasks. , and These are the weight parameters.
[0035] Performance-based rewards are expressed as follows: ; in, This represents the number of downstream tasks that can meet the execution conditions after the current task is completed. It is obtained by statistically analyzing the state changes of successor nodes in the task dependency graph. Indicates task execution efficiency. This represents the baseline execution time, which is obtained based on the execution time of similar historical tasks or the preset project development time. This indicates the actual execution time of the current task, which is obtained based on the runtime statistics during code generation and testing. This represents the cost of rework, which is determined based on the additional execution overhead incurred by replanning tasks after test failures and code refactoring. , and These are the weight parameters.
[0036] This embodiment sets up a composite reward function. On the one hand, the quality decomposition reward function quantifies and evaluates the requirement decomposition results from multiple dimensions, such as the coverage of original project requirements, subtask granularity, and subtask redundancy. This effectively suppresses planning defects such as subtask omissions, overly coarse or fine granularity, and task duplication and redundancy, improving the rationality of subtask decomposition. On the other hand, the execution efficiency reward function comprehensively considers the number of downstream tasks that can meet the execution conditions, task execution efficiency, and rework costs, balancing the executability of subsequent tasks with the overall operational cost, and reducing task rework caused by improper decomposition in the early stages. The composite reward combines decomposition quality with subsequent execution benefits, avoiding local optima caused by optimization in only a single dimension. It guides the agent to output subtask solutions that balance planning rationality and actual execution effect, reducing the probability of errors in the subsequent code generation stage and improving the overall project processing efficiency.
[0037] In step S3, the system schedules the execution agent to process the corresponding subtasks according to the task execution plan generated by the task planning agent. The task execution plan is determined by the task dependency graph, task node states, and dynamic scheduling results, and represents the execution order, prerequisites, execution status, and scheduling information of each subtask. The execution agent obtains the ready subtasks that currently meet the execution conditions according to the task execution plan and executes the corresponding code generation tasks in the scheduling order. After completing the processing of the current subtask, the execution agent returns the generation result, task execution status, and related feedback information to the system to update the task dependency graph and support subsequent task scheduling.
[0038] Specifically, such as Figure 4As shown, the executing agent performs task cache retrieval and code generation. After receiving the current subtask, the executing agent first extracts task feature information and retrieves historical successful tasks from the local task cache library. The task feature information includes at least one of task type features, entity features, technology stack features, and logical semantic features. Among them, the task type feature describes the functional category to which the current subtask belongs, including task types such as interface development, data processing, business logic implementation, and module expansion; the entity feature represents the data objects, business objects, interface objects, and module objects involved in the current task; the technology stack feature represents the programming language, development framework, database environment, and third-party component information used by the current task; and the logical semantic features represent the functional goals, input-output relationships, execution flow, and code implementation logic of the current task.
[0039] Furthermore, the local task cache library M is used to store historical successful tasks and their corresponding code information. For each historical task that has passed testing and verification, the system extracts the task type, involved entities, technology stack, logical semantic features, and a summary of the corresponding code path and design pattern, and persists this information as a cache entry. The historical task that has passed testing and verification refers to a task instance that has completed code generation and passed functional testing, interface testing, or runtime verification. After a task is completed, the system analyzes the task description information, generated code, test results, and code structure information to extract corresponding task features and form a cache entry. A cache entry... It can be represented as: ; in, Indicates the task type. Indicates that entities are involved. Indicates the technology stack. Representing logical semantic features, Indicates the code path, This represents a summary of the design pattern.
[0040] Furthermore, when the agent processes a new subtask, the system calculates the comprehensive similarity between the current subtask and historical tasks, and determines whether there are similar historical tasks that meet a preset threshold based on the comprehensive similarity. The comprehensive similarity can be expressed as: ; in, Indicates the currently pending subtask. This represents the i-th historical cached task entry; These represent the logical semantic feature vectors of the current subtask and the historical task, respectively. These represent the task types of the current subtask and the historical task, respectively. These represent the sets of entities involved in the current subtask and the historical tasks, respectively. Represents the cosine similarity between logical semantic feature vectors; This represents a task type matching function, which takes the value 1 when two task types match, and 0 otherwise. This indicates the similarity between sets of entities. , and These are the weighting parameters. Logical semantic similarity measures the degree of similarity between the current task and historical tasks in terms of functional goals and execution logic; task type matching determines whether two tasks belong to the same type of development task; and entity set similarity measures the degree of overlap between data objects, interface objects, or module objects involved in two tasks. By comprehensively calculating these multiple dimensions, the accuracy of historical task retrieval can be improved.
[0041] Furthermore, the system determines whether there are similar historical tasks that meet a preset threshold θ based on the comprehensive similarity; when At that time, determine the historical cache entries. The corresponding historical task is a similar historical task, and the threshold θ is determined based on the retrieval effect of the historical task or a preset empirical value. This is then used as the reference context information for the current sub-task; when... When this condition is met, it is determined that there are no similar historical tasks in the local task cache library M that meet the conditions.
[0042] When there is a condition that meets the preset threshold When encountering similar historical tasks, the executing agent loads the corresponding historical reference code, design pattern summary, and key logic fragments, constructs context-enhanced hints, and generates target code based on the current task requirements. The context-enhanced hints include a description of the current subtask, historical reference code fragments, design pattern summaries, key logic descriptions, and task constraints. The executing agent structures this information and inputs it into a large language model, enabling the model to generate target code based on past successful task experience and current task requirements. When no similar historical tasks satisfy a preset threshold θ, the executing agent directly calls the code generation model to generate target code. The generated subtask code is returned to the system and, after passing testing and verification, written to the local task cache library M.
[0043] In this embodiment, similarity matching is performed using multi-dimensional features. Compared to simple text matching, this improves the accuracy of historical task retrieval, fully reuses existing mature logic and design patterns, reduces the workload of repetitive reasoning by the agent, lowers the probability of hallucination generation, and avoids logical defects caused by coding from scratch. Historical task information is used to enhance the contextual clues, effectively improving code generation quality and output stability, shortening code generation time, and ensuring that the output code fits the same type of task.
[0044] In step S4, the system assembles the target code corresponding to each subtask. Specifically, based on the task tree structure and the dependencies between subtasks, the system combines the code files, interface definitions, module call relationships, and configuration files generated by each subtask to form complete project code.
[0045] In step S5, the system uses a test agent to test, verify, and iteratively optimize the generated code. Specifically, as follows... Figure 5 As shown, after receiving the generated code, the test agent executes the test verification and obtains the test results and error feedback. The test verification includes functional correctness verification, module integration verification, and runtime result verification.
[0046] When the test results meet the preset requirements, the system determines that the currently generated code has passed verification and outputs the final code. When the test results do not meet the preset requirements, the testing agent analyzes the error type, error location, and error cause, and generates a correction plan based on the error information, original task requirements, and current code state. The correction plan describes the cause of the error, the target function implementation logic, the repair path, and the control logic adjustment method.
[0047] Specifically, because project-level code involves functional collaboration between multiple modules, simply modifying code based on local error information can easily lead to effective local fixes but inconsistent overall logic. Therefore, the test agent first re-analyzes the task objectives and implementation logic based on error feedback and the current code state, generates a correction plan, and then executes code refactoring. The debugging process can be represented as a two-stage mapping including a logic replanning stage and a code refactoring stage: ; in, This indicates the current state of the code when the test failed, including the code file, module structure, and current execution result; This indicates error feedback information generated during the testing process, including functional errors, interface errors, runtime exceptions, and reasons for test failure; This represents the logical correction plan generated by the test agent based on error feedback; This indicates the result of regenerating the code based on the logical correction plan.
[0048] During the logic replanning phase, the test agent first completes error location and error source analysis, and sends the error feedback to the task planning agent; the task planning agent generates a correction plan based on the original task description, current code, error feedback, and historical fix information. The process is represented as follows: ; in, This represents the original task description. It displays historical repair information, including the types of historical test failures, the reasons for the errors, the repair strategies used, code modification records, and corresponding test results, which is used to assist the test agent in generating a more reasonable repair plan.
[0049] During the code refactoring phase, the task planning agent sends the generated revised plan to the execution agent, which then executes the revised plan accordingly. Current code and error feedback Calling the large language model to generate refactored code The test agent then re-verifies the refactored code and decides whether to continue iterating based on the verification results.
[0050] ; Through the two-stage approach of logical replanning and code refactoring described above, the model reviews the overall logic before modifying the code, avoiding patch-like modifications based solely on local error logs.
[0051] Furthermore, after generating the correction plan, the testing agent determines the source of the error. If the error originates from the task planning stage, such as unreasonable task decomposition, incorrect task dependencies, or abnormal task execution order, the task planning agent is returned to rebuild the task tree and task execution plan. If the error originates from the code implementation stage, code refactoring is performed according to the correction plan. Through this error source determination mechanism, planning errors and implementation errors during project-level code generation are distinguished and handled, avoiding mistaking task planning defects for code implementation problems. Testing is then re-executed for verification. The above task replanning and code refactoring processes can be iteratively executed until the generated results meet the preset test requirements or reach the preset number of iterations.
[0052] Through the above method, this invention utilizes the collaborative working mechanism among the task planning agent, the execution agent, and the testing agent to automate the execution of project requirements analysis, task planning, code generation, testing and verification, and iterative optimization, thereby improving the correctness, stability, and development efficiency of project-level code generation.
[0053] Example 2 This embodiment provides a large language model multi-agent collaborative project-level code generation system, including: The task acquisition module is configured to acquire project requirement information input by the user. The information analysis module is configured to use a task planning agent to recursively decompose project requirement information into sub-tasks, generate multi-level sub-tasks, and construct a task dependency graph; dynamically schedule ready sub-tasks according to the task dependency graph to generate a task execution plan; and optimize the decomposition and scheduling actions using a composite reward function. The sub-code generation module is configured to, according to the task execution plan, call the execution agent to extract the task feature information of the current sub-task, search for matching similar historical tasks, and generate the target code corresponding to the current sub-task; The code assembly module is configured to assemble the target code corresponding to each subtask according to the task dependency graph, and generate complete project code. The testing and generation module is configured to use a testing agent to test and verify the complete project code. If the preset requirements are not met, the task planning stage or the code implementation stage will be refactored according to the source of the error, and the testing and verification will be carried out again until the test results meet the preset requirements and the final code is generated.
[0054] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the method for generating project-level code for a large language model multi-agent collaborative system as described in Embodiment 1 above.
[0055] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for generating project-level code for a large language model multi-agent collaborative system as described in Embodiment 1 above.
[0056] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for multi-agent collaborative project-level code generation for large language models, characterized in that, include: Obtain project requirement information input by the user; The task planning agent is used to recursively decompose project requirements into tasks, generate multi-level sub-tasks, and construct a task dependency graph. The ready subtasks are dynamically scheduled based on the task dependency graph to generate a task execution plan; a composite reward function is used to optimize the decomposition and scheduling actions. According to the task execution plan, the execution agent is invoked to extract the task feature information of the current subtask, retrieve and match similar historical tasks, and generate the target code corresponding to the current subtask. Based on the task dependency graph, the target code corresponding to each subtask is assembled to generate the complete project code; The test agent is used to test and verify the complete project code. If the preset requirements are not met, the task planning stage or the code implementation stage is refactored according to the source of the error, and the test and verification are carried out again until the test results meet the preset requirements and the final code is generated.
2. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, Before the task planning agent performs recursive task decomposition of project requirements information, the method further includes: The project requirement information is semantically parsed to extract requirement constraint information, and the extraction results are used to determine whether there are any requirements to be broken down or complex tasks in the current project requirements. Specifically, when the project requirement information contains multiple functional verbs or functional goals, involves multiple entity objects or data structures, involves multiple code files or interfaces, has a clear sequential execution relationship or dependency relationship, or a single subtask cannot independently complete the project requirement, it is determined that there is a requirement to be decomposed or a compound task, and recursive task decomposition is performed; otherwise, the current project requirement is determined to be a single subtask, and the dependency analysis process is directly entered.
3. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, The construction task dependency graph specifically includes: The subtasks are used as task nodes, and the data dependencies, module call relationships, or execution constraint relationships between nodes are used as edges to construct a task dependency graph. The data dependency relationship refers to the output data, data structure, configuration file, or interface definition of one subtask being used as input by another subtask; the module call relationship refers to the function, class, interface, or service generated by one subtask being called by another subtask; and the execution constraint relationship refers to the sequential execution relationship caused by technology stack constraints, file generation order, environment configuration order, or testing order.
4. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, The step of dynamically scheduling ready subtasks based on the task dependency graph and generating a task execution plan specifically includes: The task dependency graph filters ready tasks that meet the execution conditions. The execution conditions are: all the prerequisites of the current subtask have been completed, the input information, interface definition, data structure, configuration file or context information required by the current subtask have been generated, and there are no unresolved upstream blocking relationships for the current subtask. When all the prerequisites of the task node are marked as completed, the current task node is added to the ready task set. The ready tasks are sorted according to subtask priority, dependency information, number of unlockable downstream tasks, estimated execution cost, and historical task reuse. Select the subtasks that appear at the top of the sorting results to generate a task execution plan; When no ready task exists, the task waits for the upstream task to complete, or it updates the task dependencies based on test feedback and then reschedules.
5. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, The composite reward function includes a decomposed quality reward function and an execution performance reward function; The decomposition quality reward function is constructed based on the coverage of the subtask set to the original project requirements, the subtask granularity score, and the redundancy between subtasks; The execution performance reward function is constructed based on the number of downstream tasks that can meet the execution conditions after the current task is completed, the task execution efficiency, and the rework cost.
6. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, The step of extracting the task feature information of the current subtask, retrieving and matching similar historical tasks, and generating the target code corresponding to the current subtask specifically includes: Extract the logical semantic feature vector, task type, and entity set of the current subtask as the features of the current task; In the local task cache library, the characteristics of the current task are compared with the corresponding characteristics of each historical cached task entry to calculate a comprehensive similarity value; specifically: ; in, Indicates the current subtask. This represents the i-th historical cached task entry; These represent the logical semantic feature vectors of the current subtask and the historical task, respectively. These represent the task types of the current subtask and the historical task, respectively. These represent the sets of entities involved in the current subtask and the historical tasks, respectively. Represents the cosine similarity between logical semantic feature vectors; This represents a task type matching function, which takes the value 1 when two task types match, and 0 otherwise. Indicates the similarity between sets of entities. , and These are weight parameters; When the overall similarity value exceeds a preset threshold, it is determined that there are similar historical tasks, and the reference code, design pattern summary and key logic fragments corresponding to the historical tasks are loaded as context enhancement information. The contextual enhancement information is combined with the requirements of the current subtask to construct enhanced prompts, which are then input into the executing agent to generate target code.
7. The method for generating project-level code for a large language model through multi-agent collaboration as described in claim 1, characterized in that, The refactoring of the task planning phase or the code implementation phase based on the source of the error specifically includes: If the error originates in the task planning phase, the test agent first completes the error location and error source analysis, and sends the error feedback to the task planning agent; the task planning agent generates a correction plan based on the original task description, current code, error feedback and historical repair information; If the error originates in the code implementation phase, the task planning agent will send the generated correction plan to the execution agent, which will then generate the refactored code based on the correction plan, the current code, and the error feedback.
8. A multi-agent collaborative project-level code generation system for large language models, characterized in that, include: The task acquisition module is configured to acquire project requirement information input by the user. The information analysis module is configured to use a task planning agent to recursively decompose project requirement information into sub-tasks, generate multi-level sub-tasks, and construct a task dependency graph; dynamically schedule ready sub-tasks according to the task dependency graph to generate a task execution plan; and optimize the decomposition and scheduling actions using a composite reward function. The sub-code generation module is configured to, according to the task execution plan, call the execution agent to extract the task feature information of the current sub-task, search for matching similar historical tasks, and generate the target code corresponding to the current sub-task; The code assembly module is configured to assemble the target code corresponding to each subtask according to the task dependency graph, and generate complete project code. The testing and generation module is configured to use a testing agent to test and verify the complete project code. If the preset requirements are not met, the task planning stage or the code implementation stage will be refactored according to the source of the error, and the testing and verification will be carried out again until the test results meet the preset requirements and the final code is generated.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the large language model multi-agent collaborative project-level code generation method as described in any one of claims 1-7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the large language model multi-agent collaborative project-level code generation method as described in any one of claims 1-7.