API test case rapid generation method and system based on AI large model
By combining multimodal data parsing, hybrid reasoning, and reinforcement learning optimizer, the problems of low efficiency and insufficient coverage of API test case generation in existing technologies are solved, and efficient and accurate test case generation is achieved, which is suitable for CI/CD pipelines.
Patent Information
- Application Number
- CN202510858901.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, API test case generation efficiency is low, coverage is insufficient, and the human error rate is high. Traditional tools and AI methods are unable to handle complex business logic and lack context awareness, resulting in a high rate of code syntax errors.
A multimodal data parsing layer is used to jointly parse and semantically fuse API documents and requirement documents. A hybrid reasoning engine is used to generate test cases. A reinforcement learning optimizer is used to optimize parameter combinations. Static and dynamic verification is performed in combination with a multidimensional verification system to ultimately generate executable test code that can be integrated into the CI/CD pipeline.
It improves the efficiency and coverage of API test case generation, reduces the human error rate, can handle complex business logic and generate test cases that meet business needs, and ensures code quality.
Smart Images

Figure CN120803919A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of software engineering and artificial intelligence, in particular to an API test case rapid generation method and system based on an AI large model. BACKGROUND
[0002] Defects of prior art:
[0003] 1) Manual writing mode: relies on the experience of developers, has problems of low efficiency (more than 30 minutes is consumed for each test case on average), insufficient coverage (abnormal scene coverage rate is less than 60%) and high human error rate (about 15% of test cases need to be revised twice).
[0004] 2) Rule-driven tools: such as Swagger Codegen, Postman and the like, can only generate basic template test cases and cannot handle complex business logic (such as order state transition, payment exception simulation) and dynamic parameter combination.
[0005] Traditional AI method: the generation model based on RNN / decision tree has defects of high code syntax error rate (>25%) and lack of context awareness ability (such as inability to associate user authentication with business interface). SUMMARY
[0006] The application aims to provide an API test case rapid generation method and system based on an AI large model to solve the problems in the background.
[0007] To achieve the above-mentioned purpose, the application provides the following technical scheme: an API test case rapid generation method based on an AI large model, comprising the following steps:
[0008] The API document and the requirement document are jointly analyzed and semantically fused by a multi-modal data analysis layer to build a complete test context;
[0009] The test case is generated based on the pre-trained large model by using a hybrid inference engine;
[0010] The parameter combination strategy is optimized according to the preset reward function by using a reinforcement learning optimizer;
[0011] The generated test case is subjected to static verification and dynamic verification by a multi-dimensional verification system, and iterative optimization is carried out according to the verification result;
[0012] Finally, the executable test code directly integrated into the CI / CD pipeline is generated.
[0013] Preferably, the multi-modal data analysis layer comprises:
[0014] The structured analysis module extracts endpoints, request methods, parameter constraints and response Schema in the API document through the OpenAPI parser to form a structured data pool and supports OpenAPI 2.0 / 3.0 multi-version compatibility analysis.
[0015] The unstructured analysis module uses a BERT-BiLSTM model to extract business rules in the requirement document, makes up for the deficiency of API metadata, and constructs training data through artificial annotation and data enhancement;
[0016] The context modeling module constructs an API dependency graph, realizes semantic alignment and fusion of API parameters and business entities in the requirement document, and identifies explicit / implicit dependencies between APIs.
[0017] Preferably, the hybrid inference engine uses a pre-trained large model as a generation base, adopts a hierarchical fine-tuning strategy and a dynamic Prompt engine, and automatically generates guide instructions according to the API type to generate test cases that meet business requirements.
[0018] Preferably, the reinforcement learning optimizer defines the reward function as: reward value = coverage weight x branch coverage + exception capture weight x defect discovery rate, and optimizes the parameter combination strategy through the PPO algorithm to solve the combination explosion problem.
[0019] Preferably, the multi-dimensional verification system includes:
[0020] The static verification module checks the AST syntax tree, detects business rule conflicts, captures syntax errors and potential anti-patterns before compilation, and ensures that the test case parameters meet the business constraint consistency;
[0021] The dynamic verification module evaluates the effectiveness of the test case in the sandbox environment, the sandbox adopts a hierarchical isolation strategy including process isolation, network isolation, file isolation and behavior monitoring, and evaluates effectiveness through path coverage, defect triggering, execution efficiency and resource consumption indicators;
[0022] The iterative optimization module feeds the failed use cases back to the training data set to realize continuous optimization of the test cases.
[0023] An API test case rapid generation method based on an AI large model, comprising:
[0024] A multi-modal data analysis layer for joint analysis and semantic fusion of API documents and requirement documents to build a complete test context;
[0025] A hybrid inference engine based on a pre-trained large model to generate test cases;
[0026] A reinforcement learning optimizer for optimizing parameter combination strategies;
[0027] A multi-dimensional verification system for static verification and dynamic verification of generated test cases;
[0028] And an iterative optimization module for iterative optimization of test cases according to verification results, finally generating executable test codes that can be directly integrated into a CI / CD pipeline.
[0029] Preferably, the multi-modal data parsing layer comprises:
[0030] A structured parsing module for extracting interface metadata in API documents through an OpenAPI parser to form a structured data pool and supporting multi-version compatibility parsing;
[0031] An unstructured parsing module for extracting business rules from requirement documents using a BERT-BiLSTM model to make up for the deficiency of API metadata;
[0032] A context modeling module for constructing an API dependency graph to realize semantic alignment and fusion of API parameters and business entities in requirement documents, and to identify explicit / implicit dependencies between APIs.
[0033] Preferably, the hybrid inference engine uses a pre-trained large model as a generation base, and adopts a hierarchical fine-tuning strategy and a dynamic Prompt engine to automatically generate guide instructions according to API types to generate test cases that meet business requirements.
[0034] Preferably, the reinforcement learning optimizer defines a reward function as: reward value = coverage weight x branch coverage + exception capture weight x defect discovery rate, and optimizes parameter combination strategies through a PPO algorithm to solve the combination explosion problem and improve the coverage rate and defect discovery capability of test cases.
[0035] Preferably, the multi-dimensional verification system comprises:
[0036] A static verification module for checking through an AST syntax tree, detecting business rule conflicts, capturing syntax errors and potential anti-patterns before compilation, and ensuring that test case parameters meet business constraint consistency;
[0037] A dynamic verification module for evaluating the effectiveness of test cases in a sandbox environment, the sandbox adopts a hierarchical isolation strategy, and evaluates effectiveness through path coverage, defect triggering, execution efficiency, and resource consumption indicators; and an iterative optimization module for feeding execution failed cases to a training data set to realize continuous optimization and adaptive maintenance of test cases.
[0038] Compared with the prior art, the present application has the following advantages:
[0039] The application provides an API test case rapid generation method and system based on an AI large model, which automatically extracts the associated constraints of API documents and requirement documents through semantic understanding; solves the combination explosion problem by optimizing the parameter combination strategy through reinforcement learning; constructs an adaptive test case maintenance mechanism to respond to API version iteration; and generates executable test code that can be directly integrated into a CI / CD pipeline. BRIEF DESCRIPTION OF DRAWINGS
[0040] Fig. 1 It is a dynamic prompt engine architecture diagram of the application.
[0041] Fig. 2 It is a reinforcement learning system workflow diagram of the application.
[0042] Fig. 3 It is a business rule conflict detection implementation flowchart of the application.
[0043] Fig. 4 It is a multi-model collaborative hybrid expert architecture diagram of the application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical scheme of the application clear, complete and the advantages more clear and obvious, the embodiments of the application will be further described in detail below with reference to the drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the application, not all embodiments, and are only used to explain the embodiments of the application, and do not limit the embodiments of the application. All other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0045] Embodiment one, please refer to Figs. 1 to 4 The application provides a technical solution: an API test case rapid generation method based on an AI large model, comprising the following steps:
[0046] 1. Automatically extract the associated constraints of API documents and requirement documents through semantic understanding;
[0047] 2. Solve the combination explosion problem by optimizing the parameter combination strategy through reinforcement learning;
[0048] 3. Construct an adaptive test case maintenance mechanism to respond to API version iteration;
[0049] 4. Generate executable test code that can be directly integrated into a CI / CD pipeline.
[0050] The following core modules are included:
[0051] Multi-modal data parsing layer: realizes joint parsing and semantic fusion of API documents and requirement documents, and constructs a complete test context.
[0052] Structured parsing: Extract endpoints, request methods, parameter constraints through OpenAPI parser, extract standardized interface metadata from API documents, form structured data pool. Adopt mature tools (such as Swagger Parser, Prance), support OpenAPI 2.0 / 3.0 multi-version compatibility parsing, ensure syntax verification and automatic conversion of documents. Extract the following core elements: interface endpoint, i.e. URL path (such as / api / users / {id}); HTTP method, such as GET / POST / PUT / DELETE, etc.; parameter constraints, including parameter type (path, query, Body), data type (string / integer), validation rules (min / max, regular expression, mandatory); response Schema, including status code, return data field structure (JSON Schema). Convert the above information into standard JSON structure or object model (such as Python apispec object) for downstream processing.
[0053] Unstructured parsing: Use BERT-BiLSTM model to extract business rules from requirement documents, extract implicit business rules from natural language description requirement documents, make up for the deficiency of API metadata. Among them, the BERT layer uses pre-trained model (such as bert-base-uncased) to generate context-sensitive word vectors; BiLSTM layer is used to capture long-distance dependence and identify sequence patterns in business rules; CRF layer (optional) is used to optimize the sequence continuity of entity labels. Through defining the extraction target type (such as conditional constraints, state transition, permission rules, etc.), the business rule classification is realized. Through artificial annotation (entity (such as numerical value, role) and relationship annotation of rule sentences in requirement documents) and data enhancement (synthetic data is generated through synonym replacement and sentence transformation to improve generalization), training data construction is realized.
[0054] Context modeling: Construct API dependency graph (e.g., the output of service A is the input of service B), build a unified test context that covers interface metadata and business rules, and identify explicit / implicit dependencies between APIs. Semantic alignment and fusion, including: 1) Align API parameters (e.g., user_id) with business entities in the requirement document (e.g., "user number"), and use word vector similarity (Cosine) or knowledge graph (Wikidata) to disambiguate; 2) Attach business rules to corresponding API parameters (e.g., "balance ≥ 100 yuan" is bound to the amount parameter of POST / withdraw). API dependency graph construction, including: 1) Inferred data dependencies based on request / response Schema (e.g., service B's input.userId comes from service A's response.user.id); 2) Discover runtime call chains through log analysis or traffic recording (requires integration with tracking systems such as Jaeger); 3) Model rules through state machines (e.g., "order state transition" triggers API calls).
[0055] Hybrid inference engine: Use pre-trained large models (e.g., GPT-4, CodeLlama) as a generation base, and use a hierarchical fine-tuning strategy and a dynamic Prompt engine to automatically generate guiding instructions based on API type.
[0056] Reinforcement learning optimizer: The reward function is defined as follows: reward value = coverage weight x branch coverage + exception capture weight x defect discovery rate, and the parameter combination strategy is optimized through the PPO algorithm.
[0057] Multi-dimensional verification system: Static verification: Check through AST syntax tree, detect business rule conflicts, catch syntax errors and potential anti-patterns before compilation, and ensure that test case parameters meet business constraints for consistency. Key detection items include: 1) Security violations, such as non-parameterized SQL queries, hardcoded keys, etc.; 2) Test smells, such as redundant assertions, excessive mocking, etc.; 3) Resource leaks, such as unclosed files / connection operations; 4) Concurrency issues, such as unsynchronized shared resource access.
[0058] Dynamic verification: Evaluate the effectiveness of test cases in a sandbox environment. Sandbox isolation strategies include process isolation, network isolation, file isolation, and behavior monitoring. Effectiveness evaluation indicators include path coverage, defect triggering, execution efficiency, and resource consumption.
[0059] Iterative optimization: Feedback failed use cases to the training data set.
[0060] In embodiment two, based on embodiment one, a system for generating API test cases quickly based on AI large models is provided according to claim 5, comprising:
[0061] A multi-modal data parsing layer is configured to jointly parse and semantically fuse the API document and the requirement document to construct a complete test context. The multi-modal data parsing layer comprises: a structured parsing module configured to extract interface metadata in the API document by an OpenAPI parser to form a structured data pool and support multi-version compatibility parsing; an unstructured parsing module configured to extract business rules from the requirement document using a BERT-BiLSTM model to make up for the deficiency of the API metadata; and a context modeling module configured to construct an API dependency graph to realize semantic alignment and fusion of API parameters and business entities in the requirement document and identify explicit / implicit dependencies between APIs.
[0062] A hybrid inference engine is configured to generate test cases based on a pre-trained large model. The hybrid inference engine uses a pre-trained large model as a generation base and adopts a hierarchical fine-tuning strategy and a dynamic Prompt engine to automatically generate guide instructions according to API types to generate test cases that meet business requirements.
[0063] A reinforcement learning optimizer is configured to optimize a parameter combination strategy. The reinforcement learning optimizer defines a reward function as: reward value = coverage weight x branch coverage + exception capture weight x defect discovery rate, and optimizes the parameter combination strategy by a PPO algorithm to solve the combination explosion problem and improve the coverage rate and defect discovery capability of the test cases.
[0064] A multi-dimensional verification system is configured to perform static verification and dynamic verification on the generated test cases, and an iterative optimization module is configured to iteratively optimize the test cases according to the verification results to finally generate executable test code that can be directly integrated into a CI / CD pipeline.
[0065] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for rapidly generating API test cases based on an AI large model, characterized by: The following steps are involved: The multimodal data parsing layer performs joint parsing and semantic fusion of API documents and requirement documents to build a complete test context. Generate test cases based on pre-trained large models using a hybrid inference engine; Use reinforcement learning optimizer to optimize parameter combination strategy according to the preset reward function; Perform static and dynamic verification on the generated test cases through a multi-dimensional verification system, and perform iterative optimization based on the verification results; Finally, executable test code is generated that can be directly integrated into the CI / CD pipeline.
2. The method for rapidly generating API test cases based on an AI big model according to claim 1 is characterized by: The multimodal data parsing layer includes: The structured parsing module extracts endpoints, request methods, parameter constraints, and response schemas from API documents through the OpenAPI parser to form a structured data pool. It also supports OpenAPI 2.0 / 3.0 multi-version compatibility parsing. The unstructured parsing module uses the BERT-BiLSTM model to extract business rules from requirement documents, making up for the lack of API metadata, and constructs training data through manual annotation and data augmentation. The context modeling module builds an API dependency graph, achieves semantic alignment and integration of API parameters with business entities in the requirements document, and identifies explicit / implicit dependencies between APIs.
3. The method for rapidly generating API test cases based on an AI big model according to claim 2, characterized in that: The hybrid inference engine uses a pre-trained large model as the generation base, and adopts a hierarchical fine-tuning strategy and a dynamic prompt engine to automatically generate guidance instructions based on the API type to generate test cases that meet business needs.
4. The method for rapidly generating API test cases based on an AI large model according to claim 3 is characterized by: The reinforcement learning optimizer defines the reward function as: reward value = coverage weight × branch coverage + exception capture weight × defect discovery rate, and optimizes the parameter combination strategy through the PPO algorithm to solve the combinatorial explosion problem.
5. The method for rapidly generating API test cases based on an AI big model according to claim 4 is characterized in that: The multi-dimensional verification system includes: The static verification module captures syntax errors and potential anti-patterns before compilation through AST syntax tree checking and business rule conflict detection, and ensures that test case parameters comply with business constraints. The dynamic verification module evaluates the effectiveness of test cases in a sandbox environment. The sandbox adopts a layered isolation strategy, including process isolation, network isolation, file isolation, and behavior monitoring. The effectiveness is evaluated through path coverage, defect stimulation, execution efficiency, and resource consumption indicators. The iterative optimization module feeds back failed test cases to the training dataset to achieve continuous optimization of test cases.
6. A system for the method for rapidly generating API test cases based on an AI large model according to claim 5, characterized in that: include: Multimodal data parsing layer, used to jointly parse and semantically fuse API documents and requirement documents to build a complete test context; Hybrid inference engine, generating test cases based on pre-trained large models; Reinforcement learning optimizer, used to optimize parameter combination strategies; Multi-dimensional verification system for static and dynamic verification of generated test cases; And an iterative optimization module is used to iteratively optimize test cases based on verification results, and ultimately generate executable test code that can be directly integrated into the CI / CD pipeline.
7. A system according to claim 6, characterized in that: The multimodal data parsing layer includes: The structured parsing module is used to extract interface metadata from API documents through the OpenAPI parser, form a structured data pool, and support multi-version compatibility parsing; Unstructured parsing module, which uses the BERT-BiLSTM model to extract business rules from requirement documents, thus addressing the lack of API metadata. The context modeling module is used to build an API dependency graph, achieve semantic alignment and integration of API parameters with business entities in the requirements document, and identify explicit / implicit dependencies between APIs.
8. A system according to claim 7, characterized in that: The hybrid inference engine uses a pre-trained large model as the generation base, and adopts a hierarchical fine-tuning strategy and a dynamic prompt engine to automatically generate guidance instructions based on the API type to generate test cases that meet business needs.
9. A system according to claim 8, characterized in that: The reinforcement learning optimizer defines the reward function as: reward value = coverage weight × branch coverage + exception capture weight × defect detection rate, and optimizes the parameter combination strategy through the PPO algorithm to solve the combinatorial explosion problem and improve the coverage and defect detection capabilities of test cases.
10. A system according to claim 9, characterized in that: The multi-dimensional verification system includes: a static verification module that captures syntax errors and potential anti-patterns before compilation through AST syntax tree checking and business rule conflict detection, and ensures that test case parameters comply with business constraints; A dynamic verification module is used to evaluate the effectiveness of test cases in a sandbox environment. The sandbox adopts a hierarchical isolation strategy and evaluates effectiveness through path coverage, defect stimulation, execution efficiency, and resource consumption indicators; and an iterative optimization module is used to feed back failed test cases to the training dataset to achieve continuous optimization and adaptive maintenance of test cases.
Citation Information
Cited By
Model-driven algorithm test method and system, computer and storage medium
CN121051028A
Software function deep analysis method based on large model driving
CN121412096A
A large model driving-based software function depth analysis method
CN121412096B