Software attribute test script automatic generation method and system based on multi-agent cooperation

CN122817067APending Publication Date: 2026-09-25Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610726867.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]为解决现有技术难以从高层安全规约全自动生成、迭代优化高质量且具备漏洞探测能力的 PBT 脚本,同时存在弱断言和逻辑幻觉的技术问题;本发明提出一种基于多智能体协同的软件属性测试脚本自动生成方法及系统,通过多智能体协同与变异验证,实现从自然语言安全规约自动生成高漏洞探测能力、可执行的PBT 脚本,大幅降低编写门槛并保障测试有效性

Benefits of technology

[0036]1、本发明构建“生成-执行-评估-精化”的循环反馈机制,依托多智能体协同与大语言模型深度融合,可直接解析模糊、非结构化的自然语言安全规约,自动生成语法合规、逻辑严谨且可迭代修正的高质量PBT脚本,同时针对缺乏明确规约的代码,自动生成推荐安全规约,大幅降低了PBT脚本的编写门槛与技术难度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817067A_ABST
    Figure CN122817067A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of software security testing, in particular to a software property test script automatic generation method and system based on multi-agent cooperation. A central scheduling agent receives user input information; the central scheduling agent selects a regulation driving mode or a code auditing mode according to the user input information, and establishes a formal verification benchmark; a PBT generation agent generates an initial PBT script based on the formal verification benchmark and the code to be tested; an execution analysis agent executes the initial PBT script and the code to be tested in a sandbox environment, captures a structured execution log and feeds back to the central scheduling agent; an evaluation refinement agent diagnoses and analyzes the structured execution log, judges whether the initial PBT script has a basic defect, and executes a mutation test to calculate a mutation score to judge whether a test passing condition is met. The application realizes automatic generation of a PBT script with high vulnerability detection capability from a natural language security regulation, and significantly reduces the writing threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software security testing technology, specifically to a method and system for automatically generating software attribute test scripts based on multi-agent collaboration. Background Technology

[0002] Large Language Models (LLMs) are rapidly changing the software engineering (SE) paradigm, demonstrating exceptional capabilities in tasks such as code generation, code repair, and documentation. Models like CodeT5+ can efficiently convert natural language descriptions into executable code, while agent-based LLM systems further drive the automation of software development. However, behind this wave of automation, primarily focused on "functional correctness," lies a profound and significant "security gap." Numerous experiments have shown that LLM-generated code commonly exhibits security vulnerabilities. Research by Pearce et al. found a considerable proportion of security vulnerabilities in code generated by GitHubCopilot. Evaluations of C / C++ code also reveal systemic deficiencies in areas such as memory safety and encryption implementation. This predicament stems from LLM's emphasis on pattern matching and probabilistic reasoning during training, lacking deep internalization of security knowledge and systematic reasoning capabilities regarding attack vectors. Although some work has attempted to improve the security of LLM-generated code through methods such as knowledge injection, how to systematically verify and ensure the security compliance of code remains a pressing challenge.

[0003] To address software security challenges, academia and industry have developed various automated verification technologies. Static Application Security Testing (SAST) has a high false positive rate and cannot detect deep semantic vulnerabilities. Dynamic Application Security Testing (DAST) and fuzzing, while adept at discovering crashes such as memory corruption, still struggle to verify the correctness of specific security policies or business logic. Property-Based Testing (PBT) shifts the focus of testing from concrete input / output examples to abstract logical properties. Experiments show that this method can effectively discover defects missed by other testing methods, particularly in exploring logical errors caused by complex input combinations and edge cases, and has been successfully applied in multiple fields such as quantum programming, property syntax, and even security protocols.

[0004] Recognizing the respective advantages of LLM in understanding natural language and PBT in logical verification, researchers have begun to explore ways to combine the two. The first approach is LLM-assisted traditional test generation. Research by Schäfer et al. shows that LLM performs well in generating example-based unit tests. However, this approach inherently inherits the limitation of traditional unit tests, which struggle to uncover deep logical vulnerabilities caused by unexpected input combinations. The second approach is to directly generate PBTs from natural language specifications. Related research indicates that software defects largely stem from vague or incomplete requirement specifications, and existing work has attempted to use LLM to extract formal specifications from unstructured natural language requirements. Vikram et al. directly explored strategies for generating PBTs from API documents using LLM and analyzed its failure modes. These works validate the potential of LLM to understand specifications and generate PBTs, but limitations include: 1) the input is often well-structured API documents, rather than more challenging security specifications geared towards adversarial thinking; 2) the quality of generated tests varies, lacking an iterative refinement mechanism to ensure the final quality of the tests. Summary of the Invention

[0005] To address the challenges of existing technologies in automatically generating and iteratively optimizing high-quality PBT scripts with vulnerability detection capabilities from high-level security specifications, while also suffering from weak assertions and logical illusions, this invention proposes a method and system for automatically generating software attribute test scripts based on multi-agent collaboration. Through multi-agent collaboration and mutation verification, this method enables the automatic generation of executable PBT scripts with high vulnerability detection capabilities from natural language security specifications, significantly reducing the writing threshold and ensuring the effectiveness of testing.

[0006] To achieve the above objectives, the technical solution adopted is:

[0007] This invention provides a method for automatically generating software attribute test scripts based on multi-agent collaboration, comprising the following steps:

[0008] S1: The central scheduling agent receives user input information, which includes natural language security specifications and code to be tested, or only the code to be tested;

[0009] S2: The central scheduling agent selects either the specification-driven mode or the code audit mode based on the user input information and establishes a formal verification benchmark.

[0010] S3: The PBT agent generates an initial PBT script based on the formal verification benchmark and the code to be tested.

[0011] S4: The execution parsing agent executes the initial PBT script and the code to be tested in the sandbox environment, captures the structured execution log, and feeds it back to the central scheduling agent;

[0012] S5: Evaluate and refine the agent to perform diagnostic analysis on the structured execution log, determine whether there are fundamental defects in the initial PBT script, and perform mutation tests to calculate mutation scores;

[0013] S6: Evaluate and refine the agent to determine whether the test pass conditions are met. If they are met, output the final PBT script and test report. If not, generate refinement instructions, and the PBT-generated agent iteratively corrects the current PBT script. Repeat steps S4 to S6 until the test pass conditions are met or the preset iteration limit is reached.

[0014] According to the software attribute test script automatic generation method based on multi-agent collaboration of the present invention, in the specification-driven mode, the central scheduling agent receives the natural language security specification and the code to be tested, and maps the natural language security specification into formal boundary conditions; in the code audit mode, the central scheduling agent only receives the code to be tested, calls the logic diagnostic unit to infer the intent of the code to be tested, and generates recommended security specifications in combination with the built-in security knowledge base as formal verification benchmarks.

[0015] According to the present invention, the software attribute test script automatic generation method based on multi-agent collaboration is further described in that the PBT generation agent is constructed based on a pre-trained large language model fine-tuned by a large-scale code corpus, and has dual functions of initial code generation and iterative correction. In the initial generation stage, the PBT generation agent generates an initial PBT script according to the formal verification benchmark and the code to be tested. In the iterative correction stage, the PBT generation agent performs incremental logic repair on the current PBT script according to the refinement instructions.

[0016] According to the present invention, the method for automatically generating software attribute test scripts based on multi-agent collaboration further includes an integrated isolated Python sandbox environment for the execution parsing agent. This sandbox environment employs a multi-layered security isolation mechanism encompassing process isolation, resource limitations, and operation interception, including:

[0017] An independent child process is generated using a process creation module to isolate the test code from the memory space of the main process;

[0018] A resource management module is used to limit the maximum CPU usage time and memory quota of child processes.

[0019] The simulated interception module physically intercepts file system operations, preventing test code from writing or deleting data on the host system.

[0020] According to the software attribute test script automatic generation method based on multi-agent collaboration of the present invention, the execution parsing agent, when executing the test script in the sandbox environment, synchronously captures the standard output, standard error information and exception stack data during the execution process, and generates a structured execution log for subsequent defect diagnosis and script refinement.

[0021] According to the present invention, the method for automatically generating software attribute test scripts based on multi-agent collaboration further includes an evaluation and refinement agent constructed based on a large language model with long-chain reasoning capabilities, possessing log diagnosis, mutation verification, and report generation functions.

[0022] Log diagnostic function: Analyze the structured execution log to determine whether the PBT script has syntax errors or runtime defects, and generate corresponding refinement instructions;

[0023] Mutation verification function: Based on the key business nodes in the security specification, generate multiple variants of the code to be tested and calculate the mutation score;

[0024] Report generation function: Transforms the final test results into a human-readable test report.

[0025] According to the method for automatically generating software attribute test scripts based on multi-agent cooperation of the present invention, the formula for calculating the mutation score is further as follows: ,in, For the currently generated test script, This refers to a set of mutant programs generated after injecting logical perturbations according to the Natural Language Security Specification. This indicates that the test script successfully triggered an assertion error during execution; if the mutation score is lower than the preset threshold, the test script is deemed unqualified, and the iterative refinement process is triggered.

[0026] According to the software attribute test script automatic generation method based on multi-agent collaboration of the present invention, the mutation verification function further includes feedback and evolution: the evaluation and refinement agent extracts the source code of surviving mutants with defects not detected by the test script, identifies mutation points through white-box comparison, and generates refinement instructions to guide the PBT generation agent to perform logic hardening.

[0027] According to the software attribute test script automatic generation method based on multi-agent collaboration of the present invention, after the initial PBT script is generated, the central scheduling agent calls a preset cleaning function to perform purification processing. The purification processing includes: correcting spelling errors in library attributes, eliminating circular dependencies within the script, and stripping redundant and duplicate code definitions, and outputting a standardized initial execution script.

[0028] Furthermore, the present invention also provides an automatic software attribute test script generation system based on multi-agent collaboration, comprising:

[0029] The user input receiving module is used to receive user input information by the central scheduling intelligent agent. The user input information includes natural language security specifications and code to be tested, or only includes code to be tested.

[0030] The mode selection and specification establishment module is used by the central scheduling agent to select the specification-driven mode or the code audit mode based on the user input information, and to establish a formal verification benchmark.

[0031] The initial script generation module is used to generate an initial PBT script from the PBT-generated agent based on the formal verification benchmark and the code to be tested;

[0032] The sandbox execution and log capture module is used by the execution parsing agent to execute the initial PBT script and the code to be tested in the sandbox environment, capture structured execution logs and feed them back to the central scheduling agent;

[0033] The diagnosis and mutation assessment module is used by the evaluation and refinement agent to perform diagnostic analysis on the structured execution log, determine whether there are basic defects in the initial PBT script, and perform mutation tests to calculate the mutation score.

[0034] The condition judgment and iteration control module is used by the evaluation and refinement agent to determine whether the test pass conditions are met. If they are met, the final PBT script and test report are output. If they are not met, refinement instructions are generated, and the PBT generation agent iteratively corrects the current PBT script. The sandbox execution and log capture module, the diagnosis and mutation evaluation module, and the module itself are repeatedly triggered until the test pass conditions are met or the preset iteration limit is reached.

[0035] The beneficial effects achieved by adopting the above technical solution are:

[0036] 1. This invention constructs a "generation-execution-evaluation-refinement" cyclical feedback mechanism. Relying on multi-agent collaboration and deep integration with large language models, it can directly parse fuzzy and unstructured natural language security specifications, automatically generate high-quality PBT scripts that are grammatically compliant, logically rigorous, and iteratively correctable. At the same time, for code that lacks explicit specifications, it automatically generates recommended security specifications, which greatly reduces the threshold and technical difficulty of writing PBT scripts.

[0037] 2. This invention designs a dynamic quality gate mechanism based on mutation analysis. By automatically generating variants of the code to be tested, it reverse-verifies the effectiveness of the PBT script. This effectively suppresses problems such as weak assertions and logical illusions in the code generated by large models, improves the PBT script's ability to accurately capture real vulnerabilities, ensures that the generated test script has industrial-grade security verification strength, and enhances the security protection level of the software supply chain. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. The drawings are merely illustrative of some embodiments of the present invention and are not intended to limit the scope of the present invention to all embodiments.

[0039] Figure 1 This is a flowchart of a method for automatically generating software attribute test scripts based on multi-agent collaboration, according to an embodiment of the present invention. Detailed Implementation

[0040] The exemplary solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art.

[0041] This invention discloses an automatic generation method for software attribute test scripts based on multi-agent collaboration, aiming to achieve end-to-end generation and refinement from natural language security specifications to executable, high-quality attribute base tests (PBTs). In this embodiment, an "agent" is defined as a computational entity that is goal-oriented, capable of reasoning and decision-making using large language models, and able to interact with other components or external tools to complete specific sub-tasks. The method of this invention is based on this definition to construct a multi-agent system (MAS). Its core designer decomposes the PBT script generation task under their responsibility and ensures the quality of the final output through a feedback-driven iterative process.

[0042] This method follows a classic coordinator-worker design pattern, delegating tasks to four core agents with specialized roles: a central scheduling agent acting as the "command center," and three specialized worker agents responsible for code generation, sandbox execution, and logic evaluation, respectively. This design allows for the decoupling and collaboration of LLM-based cognitive capabilities (such as code generation and logical reasoning) with non-LLM tool execution capabilities (such as code compilation and execution). The method of this invention is a feedback-driven, highly automated closed-loop process, such as... Figure 1 As shown, the specific implementation process is as follows.

[0043] Step S1: The central scheduling agent receives user input information, which includes natural language security specifications and code to be tested, or only the code to be tested.

[0044] In this embodiment, the central scheduling agent acts as the "command center" of the entire system. It adopts a state machine architecture to schedule the workflow and is supported by a large language model with long-chain reasoning capabilities. It is responsible for parsing user input information, driving the workflow, and routing data and distributing instructions among various specialized agents. Data routing includes transmitting PBT scripts, execution logs, and refined instructions.

[0045] Step S2: The central scheduling agent selects either the specification-driven mode or the code audit mode based on the user input information, and establishes a formal verification benchmark.

[0046] The central scheduling agent selects its working mode based on the application scenario. In the specification-driven mode, the central scheduling agent receives the natural language security specification and the code to be tested input by the user, and directly maps the natural language security specification into formal boundary conditions. In the code audit mode, the central scheduling agent only receives the code to be tested, calls the logic diagnostic unit to infer the intent of the code to be tested, and generates recommended security specifications in combination with the built-in security knowledge base, thereby establishing a formal verification benchmark.

[0047] Step S3: The PBT agent generates an initial PBT script based on the formal verification benchmark and the code to be tested.

[0048] The PBT generator is built upon a pre-trained large language model fine-tuned from a massive code corpus. It possesses dual functions: initial code generation and iterative correction. In the initial generation phase, the PBT generator generates an initial PBT script containing data generation strategies and attribute assertions based on formal verification benchmarks (i.e., formal specifications) and the code to be tested. Then, a central scheduling agent calls pre-defined cleaning functions for purification, including correcting spelling errors in library attributes, eliminating circular dependencies within the script, and removing redundant and duplicate code definitions, ultimately outputting a standardized initial execution script. In subsequent iterative correction phases, the PBT generator incrementally repairs the current PBT script according to refinement instructions, generating a new version of test code.

[0049] Step S4: The execution parsing agent executes the initial PBT script and the code to be tested in the sandbox environment, captures the structured execution log, and feeds it back to the central scheduling agent.

[0050] The central scheduling agent forwards the initial PBT script and the code to be tested to the execution parsing agent. The execution parsing agent integrates an isolated Python sandbox environment and acts as an "external tool execution unit." It executes the test process in the sandbox, fully capturing the standard output, standard error, and exception stack during runtime, and then sends back structured execution log information to the central scheduling agent.

[0051] Specifically, the sandbox environment employs a multi-layered security isolation mechanism involving process isolation, resource restrictions, and operation interception:

[0052] The test code is separated from the main process's memory space by using the process creation module (Python's multiprocessing module), ensuring that the test code cannot directly access the main process's memory space.

[0053] A resource management module is used to limit the maximum CPU usage time and memory quota of child processes to prevent test code from excessively consuming system resources.

[0054] The unittest.mock module is used to physically intercept file system operations (such as open and os.remove), preventing test code from writing or deleting data on the host system, thus preventing test code from causing physical damage to the host system and ensuring host security.

[0055] Step S5: Evaluate and refine the agent to perform diagnostic analysis on the structured execution log, determine whether there are fundamental defects in the initial PBT script, and perform mutation tests to calculate mutation scores.

[0056] The evaluation refinement agent is built on a large language model with long-chain reasoning capabilities and has functions such as log diagnosis, mutation verification, and report generation.

[0057] Log diagnostic function: Analyzes structured execution logs to determine if there are defects such as syntax errors or runtime errors in PBT scripts, and generates corresponding refinement instructions for PBT generation agents to correct.

[0058] Mutation verification function: Based on the key business nodes in the security specification, generate multiple variants of the code to be tested (variants refer to malicious variant programs containing specific logical defects), calculate the mutation score, and verify the defect detection capability of the PBT code.

[0059] Report generation function: Transforms the final test results into a human-readable test report.

[0060] Step S6: Evaluate and refine the agent to determine whether the test pass conditions are met. If they are met, output the final PBT script and test report. If not, generate refinement instructions, and the PBT agent will iteratively correct the current PBT script. Repeat steps S4 to S6 until the test pass conditions are met or the preset iteration limit is reached.

[0061] The following explains the definition, purpose, and specific implementation process of mutation verification: Mutation verification refers to introducing minor logical deviations or vulnerability features into the original benign code to construct a series of controlled defect samples. These samples are then used to reverse engineer the automatically generated test scripts to verify whether they possess the expected security boundary detection capabilities. Its core purpose is to overcome the "weak assertion" and "logic illusion" problems generated by large language models during code generation, ensuring that the generated test code not only passes syntax-level compilation checks but also serves as a high-fidelity security verification oracle, accurately identifying deep-seated logical vulnerabilities. The specific implementation process of mutation verification includes the following four stages.

[0062] Mutant generation phase: The evaluation and refinement agent synthesizes a set of mutant programs based on key business nodes in the security specification (such as encryption strength, access control logic, etc.) and using preset perturbation operators (such as operator inversion, boundary value offset, etc.).

[0063] Mutation test execution phase: The execution parsing agent runs the test script in a physically hardened sandbox to fight against each mutation sample, and fully records the execution status and exit code of each sample.

[0064] Result Determination Phase: Based on the execution results of the mutation test, the refined agent is evaluated by statistically analyzing the proportion of mutants with defects detected by the PBT script and calculating the mutation score. The mutation score (MS) serves as the core quantitative indicator for evaluating the quality of the test script. This indicator is used to accurately measure the detection coverage of the generated test script for potential security logic vulnerabilities in the program. The formula for calculating the mutation score is as follows:

[0065]

[0066] in, For the currently generated test script, This refers to a set of mutant programs generated after injecting logical perturbations according to the Natural Language Security Specification. This indicates that the test script successfully triggered an assertion error during execution. If the mutation score of the test script is lower than the preset threshold, the test script is deemed to be unqualified in terms of logical rigor, and subsequent iterative refinement processes must be forcibly triggered. To ensure that the generated test script has industrial-grade defect detection capabilities, this embodiment sets the quantitative judgment threshold to a range between 0.5 and 0.9, which can be determined according to the specific application scenario.

[0067] Feedback and Evolution Phase: The evaluation and refinement agent extracts the source code of surviving variants with defects not detected by the tested scripts. It identifies the physical paths corresponding to the mutation points through white-box comparison and uses this information as key logical evidence to generate refinement instructions, guiding the PBT generator agent to perform targeted logic hardening. When generating refinement instructions, an evidence-driven structure is adopted, using the execution log captured by the execution parsing agent and the key code fragments of surviving variants as context. Through difference comparison, the PBT generator agent is guided to perform incremental logic patching repair on the existing test scripts, avoiding logic oscillations caused by full regeneration.

[0068] Corresponding to the above method, embodiments of the present invention also disclose an automatic software attribute test script generation system based on multi-agent collaboration, comprising:

[0069] The user input receiving module is used to receive user input information by the central scheduling intelligent agent. The user input information includes natural language security specifications and code to be tested, or only includes code to be tested.

[0070] The mode selection and specification establishment module is used by the central scheduling agent to select the specification-driven mode or the code audit mode based on the user input information, and to establish a formal verification benchmark.

[0071] The initial script generation module is used to generate an initial PBT script from the PBT-generated agent based on the formal verification benchmark and the code to be tested;

[0072] The sandbox execution and log capture module is used by the execution parsing agent to execute the initial PBT script and the code to be tested in the sandbox environment, capture structured execution logs and feed them back to the central scheduling agent;

[0073] The diagnosis and mutation assessment module is used by the evaluation and refinement agent to perform diagnostic analysis on the structured execution log, determine whether there are basic defects in the initial PBT script, and perform mutation tests to calculate the mutation score.

[0074] The condition judgment and iteration control module is used by the evaluation and refinement agent to determine whether the test pass conditions are met. If they are met, the final PBT script and test report are output. If they are not met, refinement instructions are generated, and the PBT generation agent iteratively corrects the current PBT script. The sandbox execution and log capture module, the diagnosis and mutation evaluation module, and the module itself are repeatedly triggered until the test pass conditions are met or the preset iteration limit is reached.

[0075] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for automatically generating software attribute test scripts based on multi-agent collaboration, characterized in that, Includes the following steps: S1: The central scheduling agent receives user input information, which includes natural language security specifications and code to be tested, or only the code to be tested; S2: The central scheduling agent selects either the specification-driven mode or the code audit mode based on the user input information, and establishes a formal verification benchmark. S3: The PBT agent generates an initial PBT script based on the formal verification benchmark and the code to be tested. S4: The execution parsing agent executes the initial PBT script and the code to be tested in the sandbox environment, captures the structured execution log, and feeds it back to the central scheduling agent; S5: Evaluate and refine the agent to perform diagnostic analysis on the structured execution log, determine whether there are fundamental defects in the initial PBT script, and perform mutation tests to calculate mutation scores; S6: Evaluate and refine the agent to determine whether the test pass conditions are met. If they are met, output the final PBT script and test report. If not, generate refinement instructions, and the PBT-generated agent iteratively corrects the current PBT script. Repeat steps S4 to S6 until the test pass conditions are met or the preset iteration limit is reached.

2. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 1, characterized in that, In the specification-driven mode, the central scheduling agent receives the natural language security specification and the code to be tested, and maps the natural language security specification into formal boundary conditions. In the code audit mode, the central scheduling agent only receives the code to be tested, calls the logic diagnostic unit to infer the intent of the code to be tested, and generates recommended security specifications in combination with the built-in security knowledge base as a formal verification benchmark.

3. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 1, characterized in that, The PBT-generated intelligent agent is constructed based on a pre-trained large language model that has been fine-tuned with a large-scale code corpus, and has the dual functions of initial code generation and iterative correction. In the initial generation phase, the PBT generation agent generates an initial PBT script based on the formal verification benchmark and the code to be tested. During the iterative correction phase, the PBT generating agent performs incremental logical repairs on the current PBT script based on the refinement instructions.

4. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 1, characterized in that, The execution parsing agent integrates an isolated Python sandbox environment. This sandbox environment employs a multi-layered security isolation mechanism, including process isolation, resource restrictions, and operation interception. An independent child process is generated using a process creation module to isolate the test code from the main process's memory space; A resource management module is used to limit the maximum CPU usage time and memory quota of child processes. The simulated interception module physically intercepts file system operations, preventing test code from writing or deleting data on the host system.

5. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 4, characterized in that, When the execution parsing agent executes the test script in the sandbox environment, it synchronously captures the standard output, standard error information and exception stack data during the execution process, and generates a structured execution log for subsequent defect diagnosis and script refinement.

6. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 1, characterized in that, The evaluation and refinement agent is built upon a large language model with long-chain reasoning capabilities and features log diagnostics, mutation verification, and report generation. Log diagnostic function: Analyze the structured execution log to determine whether the PBT script has syntax errors or runtime defects, and generate corresponding refinement instructions; Mutation verification function: Based on the key business nodes in the security specification, generate multiple variants of the code to be tested and calculate the mutation score; Report generation function: Transforms the final test results into a human-readable test report.

7. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 6, characterized in that, The formula for calculating the variation score is as follows: ,in, For the currently generated test script, This refers to a set of mutant programs generated after injecting logical perturbations according to the Natural Language Security Specification. This indicates that the test script successfully triggered an assertion error during execution; if the mutation score is lower than the preset threshold, the test script is deemed unqualified, and the iterative refinement process is triggered.

8. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 6, characterized in that, The mutation verification function also includes feedback and evolution: evaluating and refining the source code of surviving mutants whose defects were not detected by the untested scripts, identifying mutation points through white-box comparison, and generating refinement instructions to guide PBT to generate agents for logic hardening.

9. The method for automatically generating software attribute test scripts based on multi-agent collaboration according to claim 1, characterized in that, After the initial PBT script is generated, the central scheduling agent calls a preset cleaning function to perform purification processing. The purification processing includes: correcting spelling errors in library attributes, eliminating circular dependencies within the script, and removing redundant and duplicate code definitions, and outputting a standardized initial execution script.

10. A system for automatically generating software attribute test scripts based on multi-agent collaboration, characterized in that, include: The user input receiving module is used to receive user input information by the central scheduling intelligent agent. The user input information includes natural language security specifications and code to be tested, or only includes code to be tested. The mode selection and specification establishment module is used by the central scheduling agent to select the specification-driven mode or the code audit mode based on the user input information, and to establish a formal verification benchmark. The initial script generation module is used to generate an initial PBT script from the PBT-generated agent based on the formal verification benchmark and the code to be tested; The sandbox execution and log capture module is used by the execution parsing agent to execute the initial PBT script and the code to be tested in the sandbox environment, capture structured execution logs and feed them back to the central scheduling agent; The diagnosis and mutation assessment module is used by the evaluation and refinement agent to perform diagnostic analysis on the structured execution log, determine whether there are basic defects in the initial PBT script, and perform mutation tests to calculate the mutation score. The condition judgment and iteration control module is used by the evaluation and refinement agent to determine whether the test pass conditions are met. If they are met, the final PBT script and test report are output. If the conditions are not met, a refinement instruction is generated, and the PBT-generated agent iteratively corrects the current PBT script. The sandbox execution and log capture module, the diagnosis and mutation assessment module, and the module itself are repeatedly triggered until the test pass conditions are met or the preset iteration limit is reached.