A test case driven multi-agent security code generation method
Patent Information
- Application Number
- CN202611340611.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-01
- Publication Date
- 2026-09-29
AI Technical Summary
一,采用了多智能体协同机制设计,以解决单个大语言模型的性能局限性问题
(1)显著提升漏洞检测综合能力。步骤302中通过多个智能体共同检测的漏洞检测策略,实现了功能性和安全性检测的有效兼顾,减少了漏洞检测时的误报和漏报情况。实验结果表明,采用本发明方法实施在SecodePLT数据集上,功能性指标达到85.4%,安全性指标达到95.2%,与基线方法相比分别提升了30.9%和18.2%。
Smart Images

Figure CN122839397A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent agent secure code generation technology, specifically involving a test case-driven multi-agent secure code generation method. Its direct application areas include: large language model-assisted programming, automated code security auditing, intelligent software engineering tools, etc. Background Technology
[0002] In recent years, large language models (LLMs) with generative capabilities have developed rapidly and have been widely applied in many engineering fields. In the field of code understanding and generation, large language models have demonstrated powerful application capabilities. However, the security issues of the code generated by large language models during the code generation process have become increasingly prominent. Multiple empirical studies have shown that code generated by large language models carries serious security vulnerabilities. As the application of large language models in code generation deepens, the security of model-generated code has received increasing attention, prompting research into the secure code generation problem of large language models. Secure code generation refers to improving the generation and reasoning capabilities of large language models through system design or model training, while meeting the functional requirements of the code, so that the large language models can generate highly secure code.
[0003] There is currently some research on secure code generation. However, existing work has the following shortcomings: From the perspective of vulnerability detection, existing solutions have weak comprehensive detection capabilities, which easily leads to false positives and false negatives, and cannot take into account both functional and security aspects of detection; From the perspective of code repair, existing solutions have weak ability to utilize vulnerability detection results, lack a basis for code repair, and often result in ineffective repairs; From the perspective of enhancing model generation capabilities, existing solutions rely on the capabilities of the model itself, have poor compatibility with diverse code generation tasks, and the proposed frameworks have poor scalability across different languages and application scenarios. Summary of the Invention
[0004] Objective: To address potential security vulnerabilities in existing code generation technologies, this invention provides a test case-driven multi-agent secure code generation method and system. By generating high-quality test cases, it enhances the vulnerability detection and code generation capabilities of intelligent agents, enabling iterative repair of secure code and generating highly secure target code.
[0005] Technical solution: A test case-driven method for generating secure multi-agent code, comprising the following: First, a multi-agent collaborative mechanism was adopted to address the performance limitations of a single large language model. The tasks of code generation, vulnerability detection, test case generation, and code repair were broken down and assigned to different agents. A multi-agent framework was used, employing two main types of agents: one responsible for code generation and repair, and the other for vulnerability detection and test case generation. A role-definition and role-playing approach was adopted, assigning specific keyword descriptions to both types of agents to enhance their professionalism.
[0006] Second, a test case-driven vulnerability detection method is adopted. Specifically, each participating agent is assigned a different detection task, which both distributes the vulnerability detection load across all agents and accommodates different focuses. For example, a security detection agent can be designed to enhance its own vulnerability detection capabilities based on potential security vulnerability targets, such as the CWE vulnerability list. This multi-agent collaborative vulnerability detection strategy significantly enhances vulnerability detection capabilities, achieving a balance between functionality and security, and reducing false positives and false negatives. Furthermore, at the code implementation level, vulnerability detection is implemented through subclass inheritance. Specifically, the instantiation port of the vulnerability detection agent class is retained, allowing for customized design of user-specific vulnerability detection agents through porting operations. The detection agents can also integrate other vulnerability detection technologies; for example, a fuzzing detection agent can be designed to guide the agent in performing fuzzing tests on target code, greatly enhancing the system's scalability.
[0007] Third, a data structuring design is employed to ensure the stability of the output structure. This mainly includes two aspects. First, by designing a communication protocol within the multi-agent framework, data consistency is ensured for the information transmitted between agents. Second, by maintaining the same data flow throughout execution, it is guaranteed that for the same task, key information in the original task description will not deviate significantly from its original meaning due to the uncertainty of agent-generated content after multiple iterations. Through these two designs, data consistency throughout the entire execution process can be achieved, thereby maintaining its stability. A stable execution process ensures the effectiveness of the secure code generation function.
[0008] The specific implementation steps of the method are as follows: Step 1, Input Processing; Step 101: Accept the original task description content and set preset keywords. These keywords correspond to the key information in the task description in terms of content, and the key information is extracted through character segmentation and matching. The keywords include information such as the function functional requirements description, security policy description, function return value policy, exception handling policy and function name in the task description. For example, the keyword "function name" corresponds to the specific function naming requirements in the task description.
[0009] Step 102: Perform preprocessing on the input original task description content, such as string splitting and keyword matching; Step 103: Perform data structuring on the preprocessing results; specifically, design a data structure, establish different fields, and fill the preprocessed content into the corresponding fields of the structure after filtering and confirmation.
[0010] Step 104: Initialize the data flow in the process and pass the results of data structuring processing into the running data flow. Input processing ends. After initialization, the data flow has the necessary information fields for the entire process. During initialization, the content of the fields is empty. The running data flow is responsible for interacting with each stage and maintaining the field information in the data flow unchanged to achieve data consistency.
[0011] Step 2, code generation; Step 201: Read the necessary information from the running data stream, including task description and key information.
[0012] Step 202: If it is the first iteration, initialize the agent Coder according to the settings information; the agent Coder is responsible for code generation and code repair tasks.
[0013] Step 203: The task description and key information are given to the agent Coder. The agent Coder generates an initial version of the target code and receives feedback from the agent Coder.
[0014] Step 204: The feedback results of the intelligent agent Coder are processed in a structured manner; the structured processing aims to store the information of generating target code into the running data stream in the form of a designed structure, and store the code results into the "code" field of the structure. Step 205: Add the structured processing results to the running data stream; code generation ends.
[0015] Step 3, vulnerability detection; Step 301: Read necessary information from the running data stream, including the task description and key information field content stored in the input processing stage, as well as the code generated in the current iteration round, i.e., the code field content.
[0016] Step 302: If it is the first iteration, the agent Detector is initialized according to the setting information. The agent Detector is composed of multiple agents, and each member agent is initialized in turn. In this step, a vulnerability detection strategy of multiple agents jointly detecting vulnerabilities is implemented. The agent that implements functional detection detects functional defects in the code and judges whether the specific functional requirements in the "key information" field are met one by one. The agent that implements security detection detects potential security risks in the code by retrieving and reinforcing learning knowledge of security vulnerabilities.
[0017] Step 303: Submit the generated code information of the current iteration to the agent Detector. The agent will perform vulnerability detection on the current version of the code based on the provided information. The detection results will include specific information about the vulnerabilities and the corresponding generated test cases. The agent will then provide feedback.
[0018] Step 304: Summarize the detection results of each member agent in the Agent Detector and perform structured processing. In this stage, the structured processing aims to store the summarized results in the form of a designed structure into the runtime data stream. In the "Vulnerability List" field of the structure, each vulnerability has a separate hierarchical structure, which includes the vulnerability's triggering conditions, potential impact, and corresponding test case information. In particular, a special result check is performed on the vulnerability detection results and the generated test case information. This check will remove vulnerability results that are queried repeatedly, vulnerability results with incomplete corresponding test cases, and vulnerability results with incorrect detection content.
[0019] Step 305: After removing false positive vulnerability detection information found in the above steps and deduplicating, add the vulnerability detection results to the running data stream, and the vulnerability detection ends.
[0020] Step 4, code repair; Step 401: Read necessary information from the runtime data stream, including the code generated in the current iteration round, as well as vulnerability detection results and test case information; Step 402: Re-enable the agent Coder; Step 403 requires the agent Coder to review the code generated in the previous iteration and, in conjunction with the vulnerability detection results and test case information stored in the data stream of the current iteration, to perform code repair. Based on each vulnerability result, the agent Coder locates the current version of the code for repair and executes test cases for detection. The agent Coder will return the repaired code, the result of whether the vulnerability repair was successful, and the result of whether the test cases passed, as the code repair result. Step 404: Obtain the code repair result from the agent Coder, and require the agent Coder to update the corresponding vulnerability field in the runtime data stream with the code repair result; Step 405: Based on the current iteration round and the code fix results, determine whether to proceed to the next iteration round. The determination method is as follows: if the preset maximum number of iterations is reached, the iteration stops; if all vulnerabilities are successfully fixed and test cases pass in the current fix round, the iteration stops; if there are cases where vulnerabilities are not fixed or test cases fail in the current fix round, the iteration continues. Step 406: If the next iteration is initiated, the repaired code will replace the code in the running data stream, and the process will proceed to vulnerability detection, executing step 301. Step 407: If the repair is successful or the maximum number of iterations is reached, the iteration ends, and the repaired code is output as the final output result, thus ending the code repair process.
[0021] Step 5, code evaluation; Step 501: Initialize the evaluator according to user settings; Step 502: The evaluator evaluates the code output after code repair and provides the evaluation results to the user. The evaluator includes multiple evaluation metrics, with the pass rate of test cases on the benchmark dataset as the primary evaluation metric (a higher pass rate indicates better performance). Other evaluation metrics include the evaluation functions corresponding to each dataset, semantic detection results, and large model-assisted judgment results, performing multiple evaluations.
[0022] A test case-driven multi-agent security code generation system includes an input processing module, a code generation module, a vulnerability detection module, a code repair module, and a code evaluation module. The implementation of each module is described below: The input processing module accepts the original task description and sets the system's preset keywords; preprocesses the input content; performs data structuring on the preprocessed results; initializes the data flow in the system process, which, after initialization, contains all the necessary information fields for the entire process. During initialization, the fields are empty, and the processed input data is passed into the system running data flow. The code generation module reads necessary information from the structured system runtime data stream. If it is in the first iteration, it initializes the intelligent agent Coder according to the system settings. This agent is responsible for code generation and code repair tasks throughout the system process. The task description and key information are given to the intelligent agent Coder, which generates an initial version of the target code based on the provided information. The system receives feedback from the intelligent agent Coder. The returned results from the intelligent agent Coder are then structured. The information of the generated target code is stored in the system runtime data stream in the form of a structure designed by this system, and the code result is stored in the "Code" field. The structured result is then added to the system runtime data stream. The vulnerability detection module reads necessary information from the structured system runtime data stream, including task descriptions and key information fields stored during the input processing phase, as well as the code generated by the code generation module in the current iteration. If it is in the first iteration, it initializes the agent Detector according to the system settings. This agent is composed of multiple agents, and each member agent is initialized sequentially. The code information of the current iteration is submitted to the agent Detector, which performs vulnerability detection on the current version of the code based on the provided information. The system obtains the feedback results from the agents. The detection results of each agent are summarized and structured. Finally, the vulnerability detection results are added to the system runtime data stream. The code repair module reads necessary information from the structured system runtime data stream, including the code generated by the code generation module and the vulnerability and test case information generated by the vulnerability detection module in the current iteration round; it reactivates the agent Coder; it requires the agent Coder to review the code generated in the previous iteration round and, in conjunction with the vulnerability detection and test case information stored in the data stream of the current iteration round, to perform code repair; it obtains the code repair results from the agent Coder and updates the corresponding vulnerability field in the system runtime data stream; based on the current iteration round and the code repair results, it determines whether to proceed to the next iteration round; if it proceeds to the next iteration round, it replaces the code content in the runtime data stream with the repaired code and transfers the data to the vulnerability detection module; if the repair is successful or the maximum number of iteration rounds has been reached, it ends the iteration and outputs the repaired code as the final output of the system. The code evaluation module initializes the system's evaluator based on user settings; it evaluates the code output after code repair and provides the evaluation results to the user.
[0023] The implementation process and methods of the system are the same and will not be described again.
[0024] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: (1) Significantly improves the overall vulnerability detection capability. The vulnerability detection strategy in step 302, which uses multiple agents for joint detection, effectively balances functional and security detection, reducing false positives and false negatives during vulnerability detection. Experimental results show that when the method of this invention is implemented on the SecodePLT dataset, the functional index reaches 85.4% and the security index reaches 95.2%, which are 30.9% and 18.2% higher than the baseline method, respectively.
[0025] (2) Improve the effectiveness and reliability of code repair. Through a test case-driven mechanism, this invention provides a reliable basis for code repair, with each vulnerability having a corresponding test case to verify the repair effect. The system ensures the effectiveness of the repair through iterative detection and repair processes.
[0026] (3) Enhanced system stability and robustness. Through the multi-agent framework design, this invention avoids the problem of system failure caused by a serious error in a single model, and overcomes the limitation of poor robustness of a single model. Ablation experiments demonstrate that data consistency design and data structuring processing significantly improve system performance.
[0027] (4) Improve the scalability and compatibility of the system. Based on the vulnerability detection method of this invention, at the code implementation level, the vulnerability detection design can be implemented through subclass inheritance and the instantiation port can be retained. Users can customize the detection agent according to their needs and integrate other vulnerability detection technologies. The system supports multiple programming languages (Python, C / C++) and different application scenarios (function-level code generation, backend application development).
[0028] (5) Reduce dependence on the capabilities of the model itself. Experimental results show that this method still performs well on lightweight models (such as 8B parameter models), is less affected by the limitations of model parameters and the number of tokens, and has higher code generation stability than the baseline method. Attached Figure Description
[0029] Figure 1 This is a diagram showing the execution sequence and data flow of the system of the present invention. Detailed Implementation
[0030] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0031] A test case-driven multi-agent safe code generation method is provided, in which a code task description in natural language form is accepted as input, and the final output is high-quality code that has both good functionality and security.
[0032] The secure code generation process primarily employs three core designs: a multi-agent collaborative mechanism, a test case-driven vulnerability detection method, and data structuring processing. For example... Figure 1 As shown, the multi-agent secure code generation system consists of five modules: input processing module, code generation module, vulnerability detection module, code repair module, and code evaluation module.
[0033] The method and system employ a multi-agent collaborative mechanism, primarily implementing two key agents: a Coder agent responsible for code generation and subsequent code repair, and a Detector agent, a combination of multiple agents, each with unique vulnerability detection capabilities. These agents collaborate to achieve vulnerability detection. Specifically, each detection agent possesses targeted detection capabilities; for example, the security inspection agent pre-learns common security vulnerabilities through retrieval enhancement techniques. The Detector simultaneously distributes the received current code information to all detection agents for individual detection. After validity checks, the detection results are normalized into structured vulnerability information and arranged by vulnerability type to form a vulnerability list, thus achieving the collaborative detection function of the entire agent combination.
[0034] The vulnerability detection method driven by test cases is designed as follows: (1) Stable generation of test cases. In vulnerability detection, the vulnerability detection results and generated test case information are structured and then output. The format of vulnerability detection and test case generation is specified in the preset prompt words. And detection is performed during output. Ensure that each detected vulnerability and each generated test case correspond to each other to achieve stable generation of test cases. The drawback of generating only a single test case for each vulnerability is compensated by implementing iterative detection and repair. (2) Generate efficient test cases. By adding fields in the data structure, the behavior of real people performing code testing is imitated. The agent is required to add additional information to the generated test cases, such as adding key information fields in steps 103 and 104, and adding vulnerability triggering conditions and test case information fields in steps 303 and 304, thereby improving the effectiveness of generating test cases. (3) The generated test cases have a wide coverage. In step 302, functional or security testing content is added to the structured design of vulnerability detection and test case generation. Combined with the strategy of multiple intelligent agents jointly conducting vulnerability detection, the generated test cases are avoided from being biased towards functional requirement testing and ignoring security vulnerability testing, thus achieving the generation of test cases with a wider coverage.
[0035] The data structure processing design adopted is shown in Table 1. In the input processing stage, the structured data includes two fields: task description and key information. The key information includes subfields such as function name. In the code generation stage, a generated code field is added to the structured data, storing the generated target code for the current iteration. In the vulnerability detection stage, a vulnerability list field is added to the structured data, with each vulnerability stored independently as a subfield. In the code repair stage, the structured data updates the code field and the vulnerability list field. The vulnerability list field contains vulnerability information and test case information. Vulnerability information includes subfields such as vulnerability number, vulnerability description, triggering conditions, code location, and repair status. Test case information, as a subfield of vulnerability information, includes test case content, expected output, and test case execution results. During code repair, the vulnerability repair status and test case execution results are updated.
[0036] Table 1
[0037] The achieved data consistency is mainly reflected in two characteristics: First, during operation, the fixed information added to the data stream remains unchanged, including task descriptions and key information. For information that needs to be updated, namely code fields and vulnerability list fields, the updated content is strictly controlled. Second, when various modules and system data streams interact, and when the agent receives information and returns results, all involved data undergoes structured processing and necessary checks. This design is reflected in the data-related steps 104, 201, 203, 205, 301, 303, 305, 401, 404, and 406.
[0038] like Figure 1The diagram illustrates the execution order and data flow of the system's five main modules. The system's intelligent agents are divided into two categories: Coder and Detector. The Coder agent is responsible for the code generation and code repair modules, while the Detector agent is responsible for the vulnerability detection module. The Detector class includes functional detection agents, security detection agents, and other detection agents. The entire process begins with the "input processing module," which processes data and provides the system with the original task description and key information for the data flow. In the code generation module, the Coder agent generates and provides the system with the original version of the code. In the vulnerability detection module, the Detector agent generates and provides the system with vulnerability detection and test case information. In the code repair module, the Coder agent generates and provides the system with the vulnerability repair status and the repaired code. After code repair, if all vulnerabilities are repaired or the maximum number of iterations is reached, code evaluation is performed and the final target code is generated; otherwise, iterative repair is performed, and the vulnerability detection and code repair modules are re-executed. The specific execution process of each module is as follows: First, the system accepts the original task description as input data. This task description mainly refers to the functional description of the code, such as "write a CRUD (Create, Read, Update, Delete) function for a certain dataset." It can also include key information, including but not limited to: specifying the code name, specifying the program's return value, specifying the parameter format accepted by the program, the expected security policy adopted by the program, and preset whitelists and blacklists. After input, the system enters the input processing stage. This part not only structures the code task description but also organizes key information using system-preset keywords, transferring it into the system data stream in a structured form. The main steps of input processing are as follows: The specific implementation steps of the method are as follows: Step 1, Input Processing; Step 101: Accept the original task description and set preset keywords; Step 102: Perform preprocessing on the input original task description content, such as string splitting and keyword matching; Step 103: Perform data structuring on the preprocessing results; design a data structure and set up different fields. After filtering and confirming, the preprocessed content is filled into the corresponding fields of the structure. For example, the original task description in natural language form is stored in the "Task Description" field, and the matched keywords are stored in the "Key Information" field, which includes secondary fields such as "Function Name" and "Parameter Settings". Step 104: Initialize the data flow in the process and pass the results of data structuring into the running data flow. Input processing ends. After initialization, the data flow has the necessary information fields for the entire process. During initialization, the content of the fields is empty. The running data flow is responsible for interacting with each stage and maintaining the field information in the data flow unchanged to achieve data consistency. The method implements a multi-round iteration mechanism. In the first round of iteration, code is generated based on the task description.
[0039] Step 2, code generation; Step 201: Read necessary information from the running data stream, including task description and key information; Step 202: If it is the first iteration, initialize the agent Coder according to the settings information; the agent Coder is responsible for code generation and code repair tasks. Step 203: The task description and key information are given to the agent Coder. The agent Coder generates the initial version of the target code and obtains the feedback results from the agent Coder. Step 204: The feedback results of the agent Coder are processed in a structured manner. In this stage, the structured processing aims to store the information of generating the target code into the running data stream in the form of a designed structure, and store the code results into the "code" field of the structure. Step 205: Add the structured processing results to the runtime data stream; code generation is now complete. This stage generates an initial version of the target code, which typically meets basic task requirements but contains functional defects and security risks, requiring subsequent processes for vulnerability detection and code patching.
[0040] In each iteration, two key pieces of information are focused on. First, the current code content: after the code is generated in the first iteration, the patched code in each subsequent iteration is updated in the system data stream. Second, the vulnerability detection results for the code. These results not only include information on potential vulnerabilities in the code, but also generate high-quality test cases for each vulnerability to ensure the effectiveness of vulnerability detection and code patching. Furthermore, the vulnerability detection results can contain various types of information. For example, if a vulnerability is already recorded in the CWE list as a high-risk vulnerability, the vulnerability name / code will be recorded simultaneously. If the vulnerability can be specifically located to a particular line of code, or if the specific triggering method can be detected, this information will also be recorded in the vulnerability detection results.
[0041] Step 3, vulnerability detection; Step 301: Read necessary information from the running data stream, including the task description and key information field content stored in the input processing stage, as well as the code generated in the current iteration round; Step 302: If it is the first iteration, initialize the agent Detector according to the settings. This agent Detector is composed of multiple agents, and each member agent is initialized sequentially. Specifically, this step implements a vulnerability detection strategy that uses multiple agents for joint detection. The agent implementing functional detection checks the implementation of functional requirements in the code. This functional detection agent uses the information in the "task description" and "key information" fields as a basis to generate and execute corresponding functional test cases. The success of the execution determines whether the corresponding function is implemented normally. This includes checking whether the specific functional requirements extracted from the original task description in the "key information" field are met. The agent implementing security detection checks for potential security vulnerabilities in the code. This security detection agent learns open-source security vulnerability knowledge through retrieval enhancement technology and uses the CWE vulnerability list as a knowledge base for learning. Step 303: Submit the generated code information of the current iteration to the agent Detector. Based on the provided information, the agent performs vulnerability detection on the current version of the code. The detection results will include specific vulnerability information and corresponding generated test cases. Obtain feedback results from the agent Detector. During the test case generation process, for functional testing, the specific requirements are used as the basis; for security testing, the specific security vulnerabilities and triggering methods are used as the basis. When generating test cases, the test case input and the expected correct output after the test case is executed will be generated. When executing the test case, whether the output content meets the expectations is used to determine whether the test case has passed. Step 304: Summarize the detection results of each member agent in the Agent Detector and perform structured processing. In this stage, the structured processing aims to store the summarized results in the form of a designed structure into the runtime data stream. In the "Vulnerability List" field of the structure, each vulnerability has a separate hierarchical structure, which includes the vulnerability's triggering conditions, potential impact, and corresponding test case information. In particular, a special result check is performed on the vulnerability detection results and the generated test case information. This check will remove duplicate vulnerability results, vulnerability results with incomplete corresponding test cases, and vulnerability results with incorrect detection content. Step 305: After removing false positive vulnerability detection information found in the above steps and deduplicating, add the vulnerability detection results to the running data stream, and the vulnerability detection ends.
[0042] For each round of vulnerability detection, a code fix will be performed. The code fix result includes not only the fixed code but also the vulnerability fix status, which corresponds one-to-one with the vulnerability detection result. Specifically, for each vulnerability, the code fix result will record whether the vulnerability has been fixed and whether the corresponding test cases have passed.
[0043] Step 4, code repair; Step 401: Read necessary information from the runtime data stream, including the code generated in the current iteration round, as well as vulnerability detection results and test case information; Step 402: Re-enable the agent Coder and enable the agent Coder's context memory capability; Step 403 requires the agent Coder to review the code generated in the previous iteration and, in conjunction with the vulnerability detection results and test case information stored in the data stream of the current iteration, to perform code repair. Specifically, the agent Coder locates the current version of the code for repair based on each vulnerability result, executes test cases for detection, and returns the repaired code, the result of whether the vulnerability repair was successful, and the result of whether the test cases passed, as the code repair result. Step 404: Obtain the code repair result from the agent Coder, and require the agent Coder to update the corresponding vulnerability field in the runtime data stream with the code repair result; Step 405: Based on the current iteration round and the code fix results, determine whether to proceed to the next iteration round. The determination method is as follows: if the preset maximum number of iterations is reached, the iteration stops; if all vulnerabilities are successfully fixed and test cases pass in the current fix round, the iteration stops; if there are cases where vulnerabilities are not fixed or test cases fail in the current fix round, the iteration continues. Step 406: If the next iteration is initiated, the repaired code will replace the code in the running data stream, and the process will proceed to vulnerability detection, executing step 301. Step 407: If the repair is successful or the maximum number of iterations is reached, the iteration ends, and the repaired code is output as the final output result, thus ending the code repair process.
[0044] Step 5, code evaluation; Step 501: Initialize the evaluator according to user settings; Step 502: The evaluator evaluates the code output after code repair and provides the evaluation results to the user. The evaluator includes multiple evaluation metrics, with the pass rate of test cases on the benchmark dataset as the primary evaluation metric (a higher pass rate indicates better performance). Other evaluation metrics include the evaluation functions corresponding to each dataset, semantic detection results, and large model-assisted judgment results, performing multiple evaluations.
[0045] In summary, this invention provides a test case-driven multi-agent secure code generation method, which can help improve the application of current intelligent agents in the field of code generation. The system described in this invention has been experimentally proven to perform well in different programming languages and real-world application scenarios, and achieves better performance compared to existing baseline methods in the same research field. Therefore, the system described in this invention has significant research and application value.
[0046] Obviously, those skilled in the art should understand that the steps of the test case-driven multi-agent secure code generation method or the modules of the test case-driven multi-agent secure code generation system described in the above embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
Claims
1. A test case-driven method for generating secure multi-agent code, characterized in that, Includes the following: Multi-agent collaboration mechanism: A multi-agent framework is adopted, which enables two types of agents. One type of agent is responsible for code generation and code repair, and the other type of agent is responsible for vulnerability detection and test case generation. Test case-driven vulnerability detection strategy: Assign each agent participating in vulnerability detection to perform different detection tasks; Data structuring: By designing a communication protocol for agents within a multi-agent framework, data consistency is designed for the information transmitted by agents. By designing data structures and setting different fields, the processed content is filled into the corresponding fields of the structure after being filtered and confirmed. By maintaining the same data flow during execution, that is, for the same task, the original task description remains unchanged after multiple iterations and multiple rounds of agent generation.
2. The test case-driven multi-agent security code generation method according to claim 1, characterized in that, The data structuring process is completed in the input processing stage of the multi-agent safe code generation method. In the input processing stage, the original task description content is first received, and preset keywords are set, corresponding to the key information in the task description. The keywords include the function functional requirements description, security policy description, function return value policy, exception handling policy, and function name. The original task description content is preprocessed by string splitting and keyword matching. Then, the preprocessed result is subjected to data structuring processing. By designing data structures and setting up different fields, the preprocessed content is filtered and confirmed before being filled into the corresponding fields of the structure; The input processing stage initializes the data flow in the process and passes the results of data structuring into the running data flow. The running data flow is responsible for interacting with each stage and keeping the field information in the data flow unchanged.
3. The test case-driven multi-agent safe code generation method according to claim 1, characterized in that, In the code generation phase, code generation is achieved using an agent responsible for both code generation and code repair tasks, including the following steps: Step 201: Read the task description and key information from the running data stream; Step 202: If it is the first iteration, initialize the agent Coder according to the settings information; the agent Coder is responsible for code generation and code repair tasks. Step 203: The task description and key information are given to the agent Coder. The agent Coder generates the initial version of the target code and obtains the feedback results from the agent Coder. Step 204: The feedback results of the intelligent agent Coder are processed in a structured manner; Structured processing aims to store the information that generates the target code into the runtime data stream in the form of a designed structure, and store the code results into the "code" field of the structure; Step 205: Add the structured processing results to the running data stream; code generation ends.
4. The test case-driven multi-agent security code generation method according to claim 1, characterized in that, The generated target code is then subjected to vulnerability detection using a test case-driven vulnerability detection strategy. The implementation steps for vulnerability detection are as follows: Step 301: Read the task description and key information fields stored in the input processing stage from the running data stream, as well as the code generated in the current iteration round; Step 302: If it is the first iteration, initialize the agent Detector according to the setting information. The agent Detector is used to complete the tasks of vulnerability detection and test case generation. The agent Detector is composed of multiple agents. Each member agent is initialized in turn. Step 303: Submit the generated code information of the current iteration to the agent Detector. The agent will perform vulnerability detection on the current version of the code based on the provided information. The detection results will include specific information about the vulnerabilities and the corresponding generated test cases. The agent will then provide feedback results. Step 304: Summarize the detection results of each component agent in the Agent Detector and perform structured processing; In this stage, the structured processing aims to store the summarized results into the runtime data stream in the form of a designed structure. In the "Vulnerability List" field of the structure, each vulnerability has a separate hierarchical structure, which includes the vulnerability's triggering conditions, potential impact, and corresponding test case information. For the vulnerability detection results and the generated test case information, a special result check is performed. This check will remove vulnerability results that are queried repeatedly, vulnerability results with incomplete corresponding test cases, and vulnerability results with incorrect detection content. Step 305: Add the vulnerability detection results to the running data stream; vulnerability detection ends.
5. The test case-driven multi-agent security code generation method according to claim 1, characterized in that, The code repair process includes the following steps: Step 401: Read necessary information from the runtime data stream, including the code generated in the current iteration round, as well as vulnerability detection results and test case information; Step 402: Re-enable the agent Coder; Step 403 requires the agent Coder to review the code generated in the previous iteration and, in conjunction with the vulnerability detection results and test case information stored in the data stream of the current iteration, to repair the code; Based on each vulnerability result, the intelligent agent Coder locates the current version of the code for repair and executes test cases for detection. The intelligent agent Coder will return the repaired code, the result of whether the vulnerability was successfully repaired, and the result of whether the test cases passed, as the code repair result. Step 404: Obtain the code repair result from the agent Coder, and require the agent Coder to update the corresponding vulnerability field in the runtime data stream with the code repair result; Step 405: Based on the current iteration round and the code fix results, determine whether to proceed to the next iteration round. The determination method is as follows: if the preset maximum number of iterations is reached, the iteration stops; if all vulnerabilities are successfully fixed and test cases pass in the current fix round, the iteration stops; if there are cases where vulnerabilities are not fixed or test cases fail in the current fix round, the iteration continues. Step 406: If the next iteration is initiated, the corrected code will replace the code in the running data stream, and the process will proceed to vulnerability detection. Step 407: If the repair is successful or the maximum number of iterations is reached, the iteration ends, and the repaired code is output as the final output result, thus ending the code repair process.
6. A test case-driven multi-agent safe code generation system, characterized in that, It includes an input processing module, a code generation module, a vulnerability detection module, a code repair module, and a code evaluation module; The input processing module accepts the original task description and sets preset keywords; it then preprocesses the input content. The preprocessed results are then subjected to data structuring. Initialize the system running data stream. The system running data stream includes the required fields. During initialization, the content of the corresponding fields is empty, and the processed input data is passed into the system running data stream. The code generation module reads necessary information from the structured system operation data stream. If it is in the first iteration, it initializes the intelligent agent Coder according to the system settings. The intelligent agent Coder is responsible for code generation and code repair tasks throughout the system process. The task description and key information are given to the intelligent agent Coder, which generates the initial version of the target code based on the provided information. The system receives feedback from the intelligent agent Coder and performs structured processing on the returned results. The information of generating target code is stored in the system running data stream in the form of a structure designed by this system, and the code result is stored in the "code" field; Add the structured results to the system's runtime data stream; The vulnerability detection module reads necessary information from the system's runtime data stream, including the task description and key information field content stored in the input processing stage, as well as the code generated by the code generation module in the current iteration round; If it is in the first iteration, the agent detector is initialized according to the system settings. The agent detector is composed of multiple agents, and each member agent is initialized in turn. The code information of the current iteration round is submitted to the agent detector, which performs vulnerability detection on the current version of the code based on the provided information, and the system obtains the feedback results from the agents. The detection results for each agent are summarized and then processed in a structured manner. Add the vulnerability detection results to the system runtime data stream; The code repair module reads necessary information from the system's runtime data stream, including the code generated by the code generation module in the current iteration and the vulnerability and test case information generated by the vulnerability detection module. Reactivate the agent Coder; require the agent Coder to review the code generated in the previous iteration and, in conjunction with the vulnerability detection and test case information stored in the data stream of the current iteration, perform code repair. Obtain the code repair results from the intelligent agent Coder and update the corresponding vulnerability field in the system runtime data stream; Based on the current iteration round and the code fix results, determine whether to proceed to the next iteration round; If the next iteration is initiated, the fixed code will replace the code in the running data stream and the process will be transferred to the vulnerability detection module. If the repair is determined to be successful or the maximum number of iterations has been reached, the iteration ends, and the repaired code is output as the final output of the system. The code evaluation module initializes the system's evaluator based on user settings; The code output after the code is fixed is evaluated, and the evaluation results are provided to the user.
7. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the test case-driven multi-agent safe code generation method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instruction is executed by the processor, it implements the steps of the test case-driven multi-agent safe code generation method as described in any one of claims 1-5.