Software development automation test case generation system based on artificial intelligence

By using artificial intelligence technology, a dynamically evolving test loop is constructed, which solves the problems of frequent test case failures and insufficient dynamic scenario coverage in automated software development testing, and realizes efficient testing of complex software systems and cross-project knowledge sharing.

CN120929376APending Publication Date: 2025-11-11SHANDONG BIAOFAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511047090.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for automated testing in software development suffer from problems such as frequent test case failures, insufficient coverage of dynamic scenarios, and fragmented multi-dimensional data, making it difficult to adapt to the complex changes and verification requirements of complex software systems.

Method used

Employing an AI-based immune heuristic test case self-repair module, a spatiotemporal coupled test scenario generation engine, a cross-dimensional holographic test case synthesis module, a user intent inversion test case generation module, a causal interpretability test case inference module, a chaotic edge scenario mining module, and a decentralized test knowledge ecosystem module, a dynamically evolving test closed loop is formed through immune heuristic self-repair, spatiotemporal convolutional networks, tensor decomposition, federated learning, generative adversarial networks, causal graph reasoning, and blockchain technology.

Benefits of technology

It enables self-healing and dynamic adaptation of test cases, comprehensively covers functions and performance in complex scenarios, improves testing efficiency and accuracy, supports cross-project knowledge sharing, and significantly reduces maintenance workload and testing blind spots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929376A_ABST
    Figure CN120929376A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of software development and testing, and discloses an artificial intelligence-based software development automatic test case generation system, which is characterized in that an immune heuristic case self-repairing module is adopted, defects are regarded as antigens, antibody cases capable of being self-updated are generated by using a clone selection algorithm, and a gene rearrangement mechanism is automatically triggered when an interface is changed, so that the test efficiency is improved. The details of the use case are adjusted while the core detection logic is reserved; compared with a traditional method, the mechanism can realize use case dynamic adaptation without manual intervention, the maintenance workload is remarkably reduced, and the method is particularly suitable for a complex software system with frequent iteration; the space-time coupling test scene generation engine fuses dynamic scenes such as interaction and state transition of a module and short-time operation after precise coverage login by using a space-time convolutional network; the cross-dimension holographic use case synthesis module integrates multi-source data such as codes, hardware and user behaviors through tensor decomposition to generate a composite use case; functions and performance of software in a complex scene can be comprehensively verified, and test blind areas are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of software development and testing technology, specifically an automated test case generation system for software development based on artificial intelligence. Background Technology

[0002] As software systems become increasingly complex, traditional static testing methods face several technical challenges: For example, an e-commerce platform might use script-based automated testing tools to execute fixed-process test cases. However, with frequent promotional activities causing interface parameter changes, significant weekly manpower is required to maintain the scripts, and many test cases still fail due to timing logic issues. Furthermore, insufficient coverage of dynamic scenarios makes it difficult to capture the coupling relationship between module interaction timing and system state changes; critical scenarios such as "paying within 10 seconds of logging in" are easily overlooked. Finally, fragmented multi-dimensional data, including ineffective integration of code logic, hardware performance, and user behavior information, makes it impossible to verify complex requirements.

[0003] A search revealed that invention patent CN115422074A discloses a method and apparatus for generating test cases. While generating end-to-end test cases through interface dependencies improves interface testing efficiency, it has significant limitations: it lacks self-healing capabilities, requiring manual updates for interface changes; it doesn't model spatiotemporal coupling characteristics, focusing only on the call order; and its data integration is simplistic, relying solely on interface parameters. This solution achieves breakthroughs through immune heuristic self-healing, spatiotemporal convolutional network modeling, and cross-dimensional data synthesis. Combining federated learning, causal reasoning, and blockchain technology, it forms a dynamically evolving test loop, addressing the shortcomings of existing technologies in adaptability, scenario coverage, and knowledge reuse. Summary of the Invention

[0004] The purpose of this invention is to provide an automated test case generation system for software development based on artificial intelligence, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence-based automated test case generation system for software development, the system comprising:

[0006] The immune-heuristic test case self-repair module simulates the biological immune mechanism, treating defects as "antigens" and using a clonal selection algorithm to generate "antibody test cases." Through gene rearrangement, it achieves test case self-repair, ensuring the continued effectiveness of test cases and providing a stable basic test case library for subsequent modules, thus avoiding test interruptions due to changes.

[0007] Spatiotemporal Coupled Test Scenario Generation Engine: Based on stable test cases of the immune module, it uses a spatiotemporal convolutional network to couple the spatial interaction and temporal state of the software operation to generate dynamic scene test cases with spatiotemporal tags, providing basic materials with spatiotemporal features for the cross-dimensional holographic test case synthesis module;

[0008] Cross-dimensional holographic test case synthesis module: After receiving data from the spatiotemporal coupled test scenario generation engine, it integrates multi-source information such as code, hardware, network, and user behavior, and generates cross-dimensional test cases through tensor decomposition and fusion, breaking through the limitations of a single dimension and providing a comprehensive perspective for capturing potential user needs;

[0009] User intent inversion test case generation module: Based on cross-dimensional data, federated learning and GAN are used to invert potential user operation paths, and sentiment computing is combined to transform user feedback into test targets to generate test cases that are close to real-world scenarios, providing samples for the causal interpretability test case inference module;

[0010] Causal Explainability Use Case Inference Module: For use cases in the User Intent Reversal Use Case Generation Module, a traceable link is constructed using causal graphs and counterfactual reasoning, and a causal visualization report is output to solve the "black box" problem of AI use cases and provide guidance for the chaotic edge scenario mining module;

[0011] Chaotic Edge Scene Mining Module: Based on causal relationships, boundary parameters are generated using chaotic mapping, and extreme scenes are identified by combining fractal analysis to mine the "butterfly effect" cascading failure modes, providing new defect materials for knowledge sharing;

[0012] The decentralized testing knowledge ecosystem module does not simply store the defect knowledge of the chaotic edge scenario mining module. Instead, it transforms edge scenario defect patterns into 'cross-project universal testing rules' through federated transfer learning: ① For the differences in technology stacks between different projects (such as e-commerce payment systems and medical data systems), feature alignment technology (such as Maximum Mean Difference MMD) is used to achieve defect pattern migration and adaptation; ② Combining the automatic execution characteristics of blockchain smart contracts, when a new project triggers a similar scenario, the test rules of historical defects are automatically invoked to generate adapted test cases, solving the problem of 'low scenario matching degree' in traditional knowledge reuse. The linkage between this module and the chaotic edge scenario mining module forms an innovative closed loop of 'extreme scenario mining - cross-domain knowledge transformation - front-end module adaptation', which cannot be achieved by isolated knowledge storage or transfer learning in existing technologies.

[0013] Preferably, the immune heuristic use case self-healing module includes:

[0014] (1) Biological mechanism simulation and use case generation: Simulate the principle of the biological immune system, define software defects as "antigens", train the identification model through clonal selection algorithm, accurately capture defect features and generate targeted "antibody use cases". These use cases not only cover known defects, but also have dynamic adaptation capabilities.

[0015] (2) Self-repair guarantee: When the software interface or parameters change, the module triggers the gene rearrangement mechanism to adjust the test case details while retaining the core detection logic, ensuring its continued effectiveness and providing a stable basic test case library for subsequent modules, avoiding interruption of the test process due to test case failure.

[0016] Preferably, the spatiotemporal coupling test scenario generation engine includes:

[0017] (1) Spatiotemporal Dimension Fusion Technology: Based on the stable use cases provided by the immune heuristic use case self-repair module, a spatiotemporal convolutional network (STCN) is adopted. Its network structure includes: an input layer (receiving the 128-dimensional use case feature vector output by the immune heuristic use case self-repair module), two temporal convolutional layers (convolutional kernel size 3×1, number of kernels 64 and 128 respectively, stride 1, activation function ReLU), two spatial convolutional layers (convolutional kernel size 1×3, number of kernels 128 and 64 respectively, stride 1, activation function ReLU), one residual connection layer (used to alleviate gradient vanishing), and an output layer (generating a 64-dimensional spatiotemporal coupled feature vector). Through this network, the module interaction (spatial dimension, extracting the interaction weights of 10 core modules) during software runtime is deeply coupled with the state transition (temporal dimension, collecting system state every 50ms) to construct a dynamic scene model.

[0018] Spatiotemporal convolution formula:

[0019]

[0020] In the formula: O i,j To output the coupling strength of the i-th module in the feature map at the j-th key timestamp (used to mark scenarios such as "payment 10 seconds after login");

[0021] W k,l 3x3 convolution kernel (e.g., W) 1,1 =0.2 indicates the weight of the top-left neighborhood;

[0022] I i+k,j+l The value of the input data at position (i+k,j+l);

[0023] b is the bias term (fixed value -0.1, used to suppress weak coupling noise);

[0024] Formula origin: The extension of two-dimensional convolution in the spatiotemporal dimension is used to extract features of module interaction (space) and state transition (time). This formula generates the basic features of dynamic scene models.

[0025] (2) Scene output: Generate scene use cases with timestamps and module status snapshots, fully record the system behavior of key nodes, provide basic data with spatiotemporal features for cross-dimensional holographic use case synthesis module, and facilitate the integration of multi-source information.

[0026] Preferably, the cross-dimensional holographic use case synthesis module includes:

[0027] (1) Multi-source data integration method: Taking the dynamic scene data output by the spatiotemporal engine, further incorporating heterogeneous information such as software code logic, hardware performance logs, network traffic, and user behavior trajectories, and decomposing high-dimensional data through tensor decomposition technology: ① Standardize the 4th order input tensor X (code branch / hardware indicator / network parameter / user operation) to the [0,1] interval; ② Initialize 3 sets of factor matrices A (code dimension, 100×5), B (hardware dimension, 50×5), and C (network-user dimension, 50×5), and assign random values; ③ Iteratively optimize using alternating least squares method: fix B and C, solve A to minimize the reconstruction error; fix A and C, solve B; fix A and B, solve C; iterate 500 times until the error is less than 1e-5; ④ Feature fusion and reorganization: assign weights to each factor vector through the attention mechanism (based on feature importance, code logic weight 0.3, hardware performance weight 0.2, network parameter weight 0.2, user behavior weight 0.3), and generate cross-dimensional use case features by weighted summation;

[0028] Tensor CP decomposition formula:

[0029]

[0030] In the formula: X is a 4th-order input tensor (dimensions 100X50X30X20), which corresponds to code branches / hardware metrics / network parameters / user operations respectively;

[0031] a r This is a code dimension factor vector (r = 1...5, such as a1 corresponding to the encryption algorithm branch feature);

[0032] b r This is a hardware-dimensional factor vector (e.g., b2 corresponds to CPU load characteristics);

[0033] c r This is a network-user joint factor vector (e.g., c3 corresponds to click features under a 4G network);

[0034] λ r Factor weights (∑λ) r =1, the three factors with the highest weights are used to synthesize core use cases;

[0035] ° represents the outer product operation;

[0036] Formula source: CP decomposition in tensor decomposition technology, used to decompose heterogeneous data such as code logic and hardware logs into low-dimensional subspaces to achieve cross-dimensional data fusion and recombination;

[0037] (2) Synthesis effect: The generated cross-dimensional test cases break the limitations of a single data type and can simultaneously verify the correctness of the code, the hardware carrying capacity and the smoothness of user operation, providing a comprehensive test perspective for the user intent inversion test case generation module to capture potential needs.

[0038] Preferably, the user intent inversion-based use case generation module includes:

[0039] (1) User behavior inversion mechanism: Based on multi-source data from cross-dimensional modules, federated learning is used to aggregate anonymous user behavior data, train a behavior preference model, and combine generative adversarial network (GAN) to invert the operation path that the user did not directly execute but may trigger;

[0040] GAN loss function:

[0041]

[0042] In the formula:

[0043] LG is the loss function of the generator, which is used to measure the effect of the generator in generating "pseudo-data" and guide the generator to optimize.

[0044] LD is the loss function of the discriminator, which measures the discriminator's ability to distinguish between "real data" and "generator-faked data", guiding the discriminator to optimize and make it as accurate as possible in identifying real / fake data.

[0045] E represents the mathematical expectation, which is the calculation of the "long-term average outcome" of the random variable in parentheses, and is used to integrate the statistical characteristics of a large number of samples.

[0046] G(z) is the output of the generator, namely the "pseudo-operation sequence". The generator receives noise z and generates data similar to real user operations through model mapping, such as fake user clicks and input behaviors.

[0047] z is a 128-dimensional Gaussian noise vector (mean 0, variance 1), simulating random perturbations from user operations;

[0048] D(x) is the discriminator's judgment result for "real user operation x". The output is a probability value between 0 and 1. When it is >0.8, it is judged as "real operation" and otherwise as "fake operation". However, in theory, real data should be identified with a high probability, so D(x) is usually close to 1.

[0049] Let x represent the probability that the discriminator determines x to be real data for all real user operations, and calculate the expected value.

[0050] Let G(z) be the pseudo data generated by all noise z. Calculate the logarithm of the probability that the discriminator determines G(z) to be fake data, and find its expectation.

[0051] x represents a sequence of real user actions (extracted from 100,000 anonymous log entries);

[0052] D(G(z)) is the discriminator's judgment result on the generated sample G(z) (between 0 and 1, >0.8 is judged as a real operation);

[0053] p z The noise distribution is fixed at N(0,1);

[0054] p data The true operating distribution is fitted using kernel density estimation.

[0055] Formula source: Standard loss function of generative adversarial networks, used to invert potential user operation paths and drive the model to learn user behavior patterns;

[0056] Federated learning aggregation formula:

[0057]

[0058] In the formula: θ global These are global model parameters (used to invert cross-terminal common operation paths);

[0059] θ i The parameters for the i-th terminal model (including 1000 user behavior prediction weights);

[0060] n i Let i be the sample size of the i-th terminal (e.g., i click data on a mobile device);

[0061] N = ∑n i Total sample size (fixed N=10) 6 (Add zeros if necessary);

[0062] K = 10 represents the number of terminals (including 6 mobile terminals + 4 server terminals);

[0063] Formula source: Weighted average aggregation in federated learning, used for privacy-preserving multi-terminal data fusion to achieve secure aggregation of anonymous user behavior data;

[0064] (2) Use case value: By introducing affective computing to transform user feedback into test targets, the generated use cases not only cover explicit needs but also uncover potential demands, providing analysis samples that are close to real-world scenarios for the causal interpretability use case inference module.

[0065] Preferably, the causal interpretability use case inference module includes:

[0066] (1) Causal Link Construction Technology: For scenario-based use cases generated by the user intent inversion use case generation module, the causal graph model and Do-Calculus algorithm are used to construct the link: ① Constructing the causal graph: Using 'input parameter (X)', 'module behavior (M)', and 'output result (Y)' as nodes, the directed edges between nodes are determined by conditional independence test (such as Pearson correlation coefficient) (P<0.05 indicates the existence of a causal relationship); ② Identifying confounding variables Z: Based on the Pearl backdoor criterion, variables that simultaneously affect X and Y (such as network latency and server load) are screened; ③ Applying the Do-Calculus algorithm: First, all edges pointing to X are removed to block confounding paths; second, X is fixed to a specific value (such as input field length = 256); third, the probability of the result after intervention is calculated by formula, and a traceable link 'X→M→Y' is constructed to accurately locate the direct cause of the result;

[0067] Do-Calculus intervention formula:

[0068]

[0069] In the formula: X is the intervention variable, such as the length of the input field;

[0070] Y is the outcome variable, such as payment success rate.

[0071] Z is a set of mixed variables (containing 3 categories: network latency z1, server load z2, and user equipment z3);

[0072] P(Y|X,Z) is a conditional probability table (obtained through statistical analysis of 5000 historical test data);

[0073] do(X) forces X to be set to a specific value (e.g., X = 256 characters), excluding the influence of other paths;

[0074] P(Z) represents the probability distribution of the set of confounding variables Z, which includes three types of confounding variables: network latency Z1, server load Z2, and user equipment Z3.

[0075] Formula source: Do-Calculus in causal inference, used to isolate interfering factors such as network fluctuations, locate the root cause of defects, and build a traceable "input-behavior-result" chain;

[0076] (2) Explainability: When generating test cases, a causal chain visualization report is output synchronously, clearly showing the impact of parameter adjustment on the results, solving the "black box" problem of AI test cases, providing clear causal guidance for the chaotic edge scenario mining module, and improving the test targeting.

[0077] Preferably, the chaotic edge scene mining module includes:

[0078] (1) Edge scene exploration technology: Based on the clear causal relationship provided by the causal module, the Lorenz system and other chaotic mapping algorithms are introduced to generate test parameters of random fluctuation at the normal threshold edge, construct a "chaotic disturbance field", and force the software to expose nonlinear response;

[0079] Lorenz system equations:

[0080]

[0081] In the formula: x is the normalized memory usage rate (0-1, 1 corresponds to the 80% memory usage threshold);

[0082] y represents the normalized API response time (0-1, 1 corresponds to a 1-second timeout threshold);

[0083] z represents the normalized number of concurrent users (0-1, 1 corresponds to a concurrent threshold of 1000 users);

[0084] σ = 10 (memory-response time coupling coefficient), ρ = 28 (critical state amplification factor),

[0085] β = 8 / 3 (attenuation coefficient);

[0086] t is the sampling step size (100ms / step, simulating the interval of real user operation);

[0087] Formula source: Differential equations of the Lorenz system, used to generate chaotic perturbation fields near the critical threshold, forcing the software to expose nonlinear response characteristics;

[0088] (2) Defect mining: Identify the “butterfly effect” cascading failure scenarios through fractal geometry analysis and mine new defect patterns. These results will serve as important materials to provide core content for the knowledge accumulation of the decentralized testing knowledge ecosystem module.

[0089] Preferably, the decentralized testing knowledge ecosystem module includes:

[0090] (1) Distributed knowledge architecture construction: Integrate edge scenario knowledge mined by the chaos module, and build a distributed knowledge node network based on blockchain technology. Each node stores verified test rules, defect patterns, etc., and realizes trusted sharing through smart contracts. The smart contracts include the following core rules: ① Node admission: New nodes need to submit digital certificates (signed by more than 3 authoritative nodes in the consortium chain) and obtain read and write permissions after verification; ② Knowledge storage: Test rules and defect patterns are stored in the format of 'hash value + triple (defect type - trigger condition - repair solution)', and the hash value is used for fast retrieval; ③ Knowledge verification: Before a new defect pattern is written, it needs to be simulated and verified by more than 5 nodes (if the reproduction rate is ≥90%, it passes); ④ Update mechanism: Adopt the 'timestamp priority' principle, new records of the same defect pattern overwrite old records, and historical versions are retained on the chain for traceability;

[0091] (2) Ecosystem Closed Loop: The specific mechanism for achieving cross-project knowledge reuse through federated transfer learning is as follows: ① The defect patterns of blockchain storage are abstracted into "domain-independent features" (such as the general triggering condition of "memory leaks under high concurrency") and "domain-related features" (such as the difference in concurrency thresholds between e-commerce and medical systems); ② Adversarial Domain Adaptation Network (ADAN) is used to minimize the domain differences between different projects, and the discriminator distinguishes whether the features come from the source project or the target project, driving the feature extractor to generate domain-invariant features; ③ When new use cases are generated, the domain-related feature weights are automatically matched based on the target project's technology stack (such as microservice architecture / monolithic architecture) to achieve accurate knowledge adaptation. This mechanism solves the problem of "negative transfer caused by direct transfer" in traditional transfer learning, and combines with the trusted storage of blockchain to form an innovative process of "knowledge abstraction - cross-domain adaptation - trusted reuse", which is not a simple superposition of federated learning and blockchain in existing technologies.

[0092] The beneficial effects of this invention are as follows:

[0093] 1. This invention uses an immune-heuristic self-healing test case module to treat defects as antigens and uses a clonal selection algorithm to generate self-updating antibody test cases. When the interface changes, a gene rearrangement mechanism is automatically triggered to adjust the test case details while retaining the core detection logic. Compared with traditional methods, this mechanism can achieve dynamic adaptation of test cases without manual intervention, solves the pain point of frequent test case failures, and significantly reduces maintenance workload. It is especially suitable for complex software systems with frequent iterations.

[0094] 2. This invention achieves a breakthrough through a spatiotemporal coupled test scenario generation engine and a cross-dimensional holographic test case synthesis module. The former utilizes a spatiotemporal convolutional network fusion module to accurately cover dynamic scenarios such as short-term operations after login, based on interaction (space) and state transitions (time). The latter integrates multi-source data such as code, hardware, and user behavior through tensor decomposition to generate composite test cases. Compared with the limitations of existing technologies that only focus on a single interface or dimension, this invention can comprehensively verify the functionality and performance of software in complex scenarios, significantly reducing test blind spots.

[0095] 3. This invention generates test cases with traceable links through a causal interpretability test case inference module, accumulates new defect patterns by combining a chaotic edge scenario mining module, and finally constructs a distributed knowledge network by a decentralized testing knowledge ecosystem module. This network realizes cross-project knowledge sharing through smart contracts, reuses testing experience by utilizing federated transfer learning, and forms a closed loop of "generation-verification-accumulation-reuse". Compared with isolated testing processes, it can significantly improve the testing efficiency of new projects, and is especially beneficial for the rapid identification and avoidance of common defects in the industry. Attached Figure Description

[0096] Figure 1 This is a flowchart of the AI-based automated test case generation system for software development according to the present invention. Detailed Implementation

[0097] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0098] like Figure 1 As shown, this embodiment of the invention provides an artificial intelligence-based automated test case generation system for software development, which includes:

[0099] The immune-heuristic test case self-repair module simulates the biological immune mechanism, treating defects as "antigens" and using a clonal selection algorithm to generate "antibody test cases." Through gene rearrangement, it achieves test case self-repair, ensuring the continued effectiveness of test cases and providing a stable basic test case library for subsequent modules, thus avoiding test interruptions due to changes.

[0100] Spatiotemporal Coupled Test Scenario Generation Engine: Based on stable test cases from the immune heuristic test case self-healing module, it utilizes a spatiotemporal convolutional network to couple the spatial interaction and temporal state of the software operation, generating dynamic scenario test cases with spatiotemporal markers, providing basic materials with spatiotemporal features for the cross-dimensional holographic test case synthesis module;

[0101] Cross-dimensional holographic test case synthesis module: After receiving data from the spatiotemporal coupled test scenario generation engine, it integrates multi-source information such as code, hardware, network, and user behavior, and generates cross-dimensional test cases through tensor decomposition and fusion, breaking through the limitations of a single dimension and providing a comprehensive perspective for capturing potential user needs;

[0102] User intent inversion test case generation module: Based on cross-dimensional data, federated learning and GAN are used to invert potential user operation paths, and sentiment computing is combined to transform user feedback into test targets to generate test cases that are close to real-world scenarios, providing samples for the causal interpretability test case inference module;

[0103] Causal Explainability Use Case Inference Module: For use cases in the User Intent Reversal Use Case Generation Module, a traceable link is constructed using causal graphs and counterfactual reasoning, and a causal visualization report is output to solve the "black box" problem of AI use cases and provide guidance for the chaotic edge scenario mining module;

[0104] Chaotic Edge Scene Mining Module: Based on causal relationships, boundary parameters are generated using chaotic mapping, and extreme scenes are identified by combining fractal analysis to mine the "butterfly effect" cascading failure modes, providing new defect materials for knowledge sharing;

[0105] Decentralized Testing Knowledge Ecosystem Module: Integrates knowledge from the Chaos Edge Scenarios Mining Module, constructs a blockchain knowledge node network, and achieves cross-project knowledge sharing through smart contracts and federated transfer learning, feeding back to the front-end module to form a "generation-reuse" process.

[0106] This invention employs an immune-heuristic self-healing test case module, treating defects as antigens and using a clonal selection algorithm to generate self-updating antibody test cases. When the interface changes, a gene rearrangement mechanism is automatically triggered, adjusting test case details while preserving core detection logic; this significantly reduces maintenance workload, making it particularly suitable for complex software systems with frequent iterations. A spatiotemporal coupled test scenario generation engine utilizes a spatiotemporal convolutional network to fuse module interactions and state transitions, accurately covering dynamic scenarios such as short-term operations after login. A cross-dimensional holographic test case synthesis module integrates multi-source data such as code, hardware, and user behavior through tensor decomposition to generate composite test cases, comprehensively verifying the software's functionality and performance in complex scenarios and significantly reducing testing blind spots. A causal interpretability test case inference module generates test cases with traceable links, and a chaotic edge scenario mining module accumulates novel defect patterns. Finally, a decentralized testing knowledge ecosystem module constructs a distributed knowledge network. This network achieves cross-project knowledge sharing through smart contracts and utilizes federated transfer learning to reuse testing experience, significantly improving the testing efficiency of new projects, especially facilitating the rapid identification and avoidance of common industry defects.

[0107] The immune heuristic use case self-healing module includes:

[0108] (1) Biological mechanism simulation and use case generation: Simulating the principle of the biological immune system, software defects are defined as "antigens", and the identification model is trained by clonal selection algorithm: ① Initialize the defect feature library (containing 1000+ historical defect samples); ② Use the roulette wheel selection operator to screen high affinity antibody use cases; ③ Generate candidate use cases by mutation operator (mutation rate 0.05); ④ Iteratively train until the identification accuracy is ≥90%; accurately capture defect features and generate targeted "antibody use cases", which not only cover known defects, but also have dynamic adaptation capabilities;

[0109] (2) Self-healing guarantee: When the software interface or parameters change, the module triggers the gene rearrangement mechanism through the following logic: ① Real-time monitoring of the change rate of interface metadata (such as parameter type, field length, and call frequency), and triggering rearrangement when the change rate exceeds the preset threshold (such as 30%); ② Extracting the core detection logic of the test cases (such as defect judgment rules and key verification nodes) through the abstract syntax tree (AST) and solidifying it; ③ Rearranging non-core details (such as parameter passing format and interface call order) using a genetic algorithm: using historical valid test cases as the parent generation, generating offspring test cases through crossover (exchanging parameter combination methods) and mutation (adjusting interface call timing), updating the test case library after validity verification (pass rate ≥ 85% through simulated execution test), providing a stable basic test case library for subsequent modules, and avoiding interruption of the test process due to test case failure.

[0110] The spatiotemporal coupling test scenario generation engine includes:

[0111] (1) Spatiotemporal Dimension Fusion Technology: Based on the stable use cases provided by the immune heuristic use case self-repair module, a spatiotemporal convolutional network (STCN) (containing 2 temporal convolutional layers + 2 spatial convolutional layers, with a temporal convolutional kernel size of 3×1 and a spatial convolutional kernel size of 1×3, and ReLU activation function) is adopted to deeply couple the module interaction (extracting 10 core module interaction features) and state transition (recording the state once every 50ms) during software runtime. Gradient propagation is optimized through 3 residual connections to construct a dynamic scene model;

[0112] Spatiotemporal convolution formula:

[0113]

[0114] In the formula: O i,j To output the coupling strength of the i-th module in the feature map at the j-th key timestamp (used to mark scenarios such as "payment 10 seconds after login");

[0115] W k,l 3x3 convolution kernel (e.g., W) 1,1 =0.2 indicates the weight of the top-left neighborhood;

[0116] Ii+k,j+l The value of the input data at position (i+k,j+l);

[0117] b is the bias term (fixed value -0.1, used to suppress weak coupling noise);

[0118] Formula origin: The extension of two-dimensional convolution in the spatiotemporal dimension is used to extract features of module interaction (space) and state transition (time). This formula generates the basic features of dynamic scene models.

[0119] (2) Scene output: Generate scene use cases with timestamps and module status snapshots, fully record the system behavior of key nodes, provide basic data with spatiotemporal features for cross-dimensional holographic use case synthesis module, and facilitate the integration of multi-source information.

[0120] The cross-dimensional holographic use case synthesis module includes:

[0121] (1) Multi-source data integration method: Taking the dynamic scene data output by the spatiotemporal engine, further incorporating heterogeneous information such as software code logic, hardware performance logs, network traffic, and user behavior trajectories, and decomposing high-dimensional data through tensor decomposition technology: ① Initialize factor matrix (randomly generate factor vectors of code, hardware, network-user dimensions); ② Iteratively optimize using alternating least squares method (500 iterations, convergence threshold 1e-5); ③ Fuse and recombine low-dimensional factor features through attention mechanism (weight coefficients are assigned based on feature importance);

[0122] Tensor CP decomposition formula:

[0123]

[0124] In the formula: X is a 4th-order input tensor (dimensions 100X50X30X20), which corresponds to code branches / hardware metrics / network parameters / user operations respectively;

[0125] a r This is a code dimension factor vector (r = 1...5, such as a1 corresponding to the encryption algorithm branch feature);

[0126] b r This is a hardware-dimensional factor vector (e.g., b2 corresponds to CPU load characteristics);

[0127] c r This is a network-user joint factor vector (e.g., c3 corresponds to click features under a 4G network);

[0128] λ r Factor weights (∑λ) r =1, the three factors with the highest weights are used to synthesize core use cases;

[0129] ° represents the outer product operation;

[0130] Formula source: CP decomposition in tensor decomposition technology, used to decompose heterogeneous data such as code logic and hardware logs into low-dimensional subspaces to achieve cross-dimensional data fusion and recombination;

[0131] (2) Synthesis effect: The generated cross-dimensional test cases break the limitations of a single data type and can simultaneously verify the correctness of the code, the hardware carrying capacity and the smoothness of user operation, providing a comprehensive test perspective for the user intent inversion test case generation module to capture potential needs.

[0132] The user intent inversion test case generation module includes:

[0133] (1) User behavior inversion mechanism: Based on multi-source data from the cross-dimensional holographic use case synthesis module, federated learning is used to aggregate anonymized user behavior data (using the FedAvg algorithm, 10 local training rounds, and an aggregation period of 5 minutes) to train a behavior preference model (input is a 128-dimensional user behavior feature vector); combined with a generative adversarial network (GAN, generator is a 3-layer LSTM, discriminator is a 2-layer CNN, training optimizer is Adam, learning rate is 0.0002) to invert the operation path that the user did not directly execute but may trigger; sentiment calculation uses TextCNN to extract user feedback text features (word vector dimension 300, convolution kernel size 2 / 3 / 4), and maps them to three test target weights of "satisfied / neutral / dissatisfied" through a softmax classifier;

[0134] The specific steps of sentiment computing are as follows: ① Preprocessing of user feedback text (removing stop words, segmenting words, and generating 300-dimensional word vectors using a Word2Vec pre-trained model; the training corpus consists of 100,000 software review texts); ② TextCNN structure: the input layer is a 128×300 word vector matrix (128 is the maximum text length), the convolutional layer contains 3 sets of convolutional kernels (sizes 2×300, 3×300, and 4×300, 32 kernels per set), activated by ReLU and then subjected to global max pooling to obtain a 128-dimensional feature vector; ③ The fully connected layer uses dropout (rate 0.5) to suppress overfitting, and the output layer uses a softmax classifier to obtain the probability distribution of 'satisfied / neutral / unsatisfactory', corresponding to test target weights of 0.2 / 0.5 / 0.8 respectively (the higher the weight, the more priority should be given to covering this type of feedback);

[0135] GAN loss function:

[0136]

[0137] In the formula:

[0138] LG is the loss function of the generator, which is used to measure the effect of the generator in generating "pseudo-data" and guide the generator to optimize.

[0139] LD is the loss function of the discriminator, which measures the discriminator's ability to distinguish between "real data" and "generator-faked data", guiding the discriminator to optimize and make it as accurate as possible in identifying real / fake data.

[0140] E represents the mathematical expectation, which is the calculation of the "long-term average outcome" of the random variable in parentheses, and is used to integrate the statistical characteristics of a large number of samples.

[0141] G(z) is the output of the generator, namely the "pseudo-operation sequence". The generator receives noise z and generates data similar to real user operations through model mapping, such as fake user clicks and input behaviors.

[0142] z is a 128-dimensional Gaussian noise vector (mean 0, variance 1), simulating random perturbations from user operations;

[0143] D(x) is the discriminator's judgment result for "real user operation x". The output is a probability value between 0 and 1. When it is >0.8, it is judged as "real operation" and otherwise as "fake operation". However, in theory, real data should be identified with a high probability, so D(x) is usually close to 1.

[0144] Let x represent the probability that the discriminator determines x to be real data for all real user operations, and calculate the expected value.

[0145] Let G(z) be the pseudo data generated by all noise z. Calculate the logarithm of the probability that the discriminator determines G(z) to be fake data, and find its expectation.

[0146] x represents a sequence of real user actions (extracted from 100,000 anonymous log entries);

[0147] D(G(z)) is the discriminator's judgment result on the generated sample G(z) (between 0 and 1, >0.8 is judged as a real operation);

[0148] p z The noise distribution is fixed at N(0,1);

[0149] p data The true operating distribution is fitted using kernel density estimation.

[0150] Formula source: Standard loss function of generative adversarial networks, used to invert potential user operation paths and drive the model to learn user behavior patterns;

[0151] Federated learning aggregation formula:

[0152]

[0153] In the formula: θglobal These are global model parameters (used to invert cross-terminal common operation paths);

[0154] θ i The parameters for the i-th terminal model (including 1000 user behavior prediction weights);

[0155] n i Let i be the sample size of the i-th terminal (e.g., i click data on a mobile device);

[0156] N = ∑n i Total sample size (fixed N=10) 6 (Add zeros if necessary);

[0157] K = 10 represents the number of terminals (including 6 mobile terminals + 4 server terminals);

[0158] Formula source: Weighted average aggregation in federated learning, used for privacy-preserving multi-terminal data fusion to achieve secure aggregation of anonymous user behavior data;

[0159] (2) Use case value: By introducing affective computing to transform user feedback into test targets, the generated use cases not only cover explicit needs but also uncover potential demands, providing analysis samples that are close to real-world scenarios for the causal interpretability use case inference module.

[0160] The causal interpretability use case inference module includes:

[0161] (1) Causal Link Construction Technology: For scenario-based use cases generated by the user intent inversion use case generation module, the causal graph model (nodes are defined as three categories: "input parameters - module behavior - output results", and edges represent direct causal relationships) and the Do-Calculus algorithm are used to: ① identify confounding variables (based on Pearl backdoor criteria); ② remove confounding factors through 3-step intervention (remove the edge pointing to X from the parent node → set X to a specific value → calculate conditional probability); ③ construct a traceable link of input-behavior-result to accurately locate the direct cause of the result;

[0162] Do-Calculus intervention formula:

[0163]

[0164] In the formula: X is the intervention variable, such as the length of the input field;

[0165] Y is the outcome variable, such as payment success rate.

[0166] Z is a set of mixed variables (containing 3 categories: network latency z1, server load z2, and user equipment z3);

[0167] P(Y|X,Z) is a conditional probability table (obtained through statistical analysis of 5000 historical test data);

[0168] do(X) forces X to be set to a specific value (e.g., X = 256 characters), excluding the influence of other paths;

[0169] P(Z) represents the probability distribution of the set of confounding variables Z, which includes three types of confounding variables: network latency Z1, server load Z2, and user equipment Z3.

[0170] Formula source: Do-Calculus in causal inference, used to isolate interfering factors such as network fluctuations, locate the root cause of defects, and build a traceable "input-behavior-result" chain;

[0171] (2) Interpretability: A causal chain visualization report is output synchronously when generating test cases. The specific format is as follows: ① Node layer: Circular nodes represent input parameters (such as "input field length"), square nodes represent module behaviors (such as "data encryption"), and triangular nodes represent output results (such as "payment success rate"). The size of the node is positively correlated with the influence weight; ② Edge layer: Solid lines represent direct causal relationships, and dashed lines represent indirect causal relationships. The thickness of the edges corresponds to the causal strength (probability value calculated based on Do-Calculus); ③ Interactive function: It supports clicking on nodes to view the specific parameter value range (such as "input field length ∈ [1,256]") and dragging parameters to simulate the result change curve after adjustment. This report is generated by the D3.js visualization library and the output format is SVG, which is convenient for embedding into the test report system.

[0172] The chaotic edge scene mining module includes:

[0173] (1) Edge Scene Exploration Technology: Based on the explicit causal relationship provided by the causal module, the Lorenz system chaotic mapping algorithm is introduced: ① Initialize parameters (x = 0.1, y = 0.1, z = 0.1); ② Iterate and solve the differential equation with a step size of 100ms (using the Runge-Kutta method); ③ Transform the iteration results (x, y, z) into actual test parameters through linear mapping (e.g., x∈[0.8,1.0] corresponds to a memory usage rate of 70%-80%), generate test parameters with random fluctuations at the normal threshold edge, construct a "chaotic perturbation field", and force the software to expose nonlinear responses;

[0174] Lorenz system equations:

[0175]

[0176] In the formula: x is the normalized memory usage rate (0-1, 1 corresponds to the 80% memory usage threshold);

[0177] y represents the normalized API response time (0-1, 1 corresponds to a 1-second timeout threshold);

[0178] z represents the normalized number of concurrent users (0-1, 1 corresponds to a concurrent threshold of 1000 users);

[0179] σ = 10 (memory-response time coupling coefficient), ρ = 28 (critical state amplification factor),

[0180] β = 8 / 3 (attenuation coefficient);

[0181] t is the sampling step size (100ms / step, simulating the interval of real user operation);

[0182] Formula source: Differential equations of the Lorenz system, used to generate chaotic perturbation fields near the critical threshold, forcing the software to expose nonlinear response characteristics;

[0183] (2) Defect Discovery: Identifying cascading failure scenarios with the 'butterfly effect' through fractal geometry analysis: ① Collect system state sequences (including 5 key indicators such as memory, response time, and concurrency, with a sampling frequency of 100ms / time and a sequence length ≥1000 points); ② Calculate the fractal dimension using box counting: Divide the state space into cubes with a side length of ε (ε ranges from 0.01 to 0.1, increasing in a geometric progression), count the number of cubes N(ε) covering all sequence points, and fit the linear relationship between lgN(ε) and lg(1 / ε) using a double logarithmic coordinate system. The slope is the fractal dimension; ③ When the fractal dimension > 1.5, it is determined to be a cascading failure scenario with the 'butterfly effect' characteristics (at this time, the system state exhibits nonlinear sensitivity to the initial disturbance); ④ Trace the failure chain through the fault propagation graph (nodes are modules, edges are call relationships) to discover new defect patterns. These results will serve as important materials to provide core content for the knowledge accumulation of the decentralized testing knowledge ecosystem module.

[0184] The decentralized testing knowledge ecosystem module includes:

[0185] (1) Distributed knowledge architecture construction: Integrate edge scenario knowledge mined by the chaos module, and build a distributed knowledge node network based on blockchain technology (adopting a consortium chain architecture and PBFT consensus mechanism): ① Each node contains a data layer (storing test rule hash values) and a contract layer (smart contracts define knowledge read and write permissions: ① Nodes need to submit a digital certificate to join; ② Read permissions are open, and writing requires verification by more than 3 nodes; ③ Update frequency ≤ 1 time / minute, and in case of conflict, the latest timestamp shall prevail); ② Node communication adopts the P2P protocol, and the data synchronization cycle is 1 minute; Each node stores verified test rules and defect patterns (stored in the triple format of "defect type-trigger condition-repair plan"), and achieves trusted sharing through smart contracts;

[0186] (2) Ecosystem Closed Loop: The specific mechanism for achieving cross-project knowledge reuse through federated transfer learning is as follows: ① The defect patterns of blockchain storage are abstracted into "domain-independent features" (such as the general triggering condition of "memory leaks under high concurrency") and "domain-related features" (such as the difference in concurrency thresholds between e-commerce and medical systems); ② Adversarial Domain Adaptation Network (ADAN) is used to minimize the domain differences between different projects, and the discriminator distinguishes whether the features come from the source project or the target project, driving the feature extractor to generate domain-invariant features; ③ When new use cases are generated, the domain-related feature weights are automatically matched based on the target project's technology stack (such as microservice architecture / monolithic architecture) to achieve accurate knowledge adaptation. This mechanism solves the problem of "negative transfer caused by direct transfer" in traditional transfer learning, and combines with the trusted storage of blockchain to form an innovative process of "knowledge abstraction - cross-domain adaptation - trusted reuse", which is not a simple superposition of federated learning and blockchain in existing technologies.

[0187] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0188] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An AI-based automated test case generation system for software development, characterized by: The system includes: Immune-inspired self-repair module for test cases: It simulates the biological immune mechanism, uses a clonal selection algorithm to generate antibody test cases, and achieves test case self-repair through gene rearrangement; Spatiotemporal Coupled Test Scenario Generation Engine: Based on antibody test cases, it uses a spatiotemporal convolutional network to couple the spatial interaction and temporal state of the software operation to generate dynamic scenario test cases with spatiotemporal tags; Cross-dimensional holographic test case synthesis module: Receives data from the spatiotemporal coupled test scenario generation engine, integrates code, hardware, network, and user behavior, and generates cross-dimensional test cases through tensor decomposition and fusion; User intent inversion test case generation module: Based on cross-dimensional data, federated learning and GAN are used to invert potential user operation paths, and sentiment computing is combined to transform user feedback into test targets to generate test cases for real-world scenarios; Causal Explainability Use Case Inference Module: For use cases generated by the User Intent Reversal Use Case Generation Module, a traceable link is constructed using causal graphs and counterfactual reasoning, and a causal visualization report is output. Chaotic edge scene mining module: Based on causal relationships, boundary parameters are generated using chaotic mapping, and extreme scenes are identified by combining fractal analysis; Decentralized Testing Knowledge Ecosystem Module: Integrates knowledge from the Chaos Edge Scenario Mining Module, constructs a blockchain knowledge node network, and achieves cross-project knowledge sharing through smart contracts and federated transfer learning.

2. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The immune heuristic use case self-healing module includes: (1) Biological mechanism simulation and use case generation: Simulate the principle of the biological immune system, define software defects as antigens, train the identification model through clonal selection algorithm, capture defect features and generate targeted antibody use cases; (2) Self-repair guarantee: When the software interface or parameters change, the gene rearrangement mechanism is triggered to adjust the test case details while retaining the core detection logic.

3. The automated test case generation system for software development based on artificial intelligence according to claim 1, characterized in that: The spatiotemporal coupling test scenario generation engine includes: (1) Spatiotemporal dimensional fusion technology: Based on the stable use cases provided by the immune module, a spatiotemporal convolutional network is used to deeply couple the module interaction and state transition during software runtime to construct a dynamic scene model; (2) Scenario output: Generate scenario use cases with timestamps and module status snapshots to fully record the system behavior of key nodes.

4. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The cross-dimensional holographic use case synthesis module includes: (1) Multi-source data integration method: After receiving the dynamic scene data output by the spatiotemporal coupling test scene generation engine, software code logic, hardware performance logs, network traffic, and user behavior trajectory are incorporated. High-dimensional data is decomposed through tensor decomposition technology and then recombined through feature fusion. (2) Synthesis effect: The generated cross-dimensional test cases can simultaneously verify code correctness, hardware carrying capacity and user operation smoothness.

5. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The user intent inversion-based use case generation module includes: (1) User behavior inversion mechanism: Based on multi-source data from the cross-dimensional holographic use case synthesis module, federated learning is used to aggregate anonymous user behavior data, train a behavior preference model, and combine it with a generative adversarial network; (2) Use case value: By introducing affective computing to transform user feedback into test targets, the generated use cases not only cover explicit needs but also uncover potential demands.

6. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The causal interpretability use case inference module includes: (1) Causal link construction technology: For the scenario-based use cases generated by the user intent inversion use case generation module, the causal graph model and Do-Calculus algorithm are used to remove confounding factors, construct a traceable link of input-behavior-result, and locate the direct cause of the result; (2) Explainability: When generating test cases, a causal chain visualization report is output synchronously, clearly showing the impact of parameter adjustments on the results.

7. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The chaotic edge scene mining module includes: (1) Edge scene exploration technology: Based on the explicit causal relationship provided by the causal interpretability use case inference module, the Lorenz system chaos mapping algorithm is introduced to generate test parameters with random fluctuations at the normal threshold edge, so that the software exposes nonlinear response; (2) Defect mining: Identify butterfly effect cascade failure scenarios through fractal geometry analysis and mine new defect patterns.

8. The AI-based automated test case generation system for software development according to claim 1, characterized in that: The decentralized testing knowledge ecosystem module includes: (1) Distributed knowledge architecture construction: integrate edge scenario knowledge mined by the chaotic edge scenario mining module, and build a distributed knowledge node network based on blockchain technology. Each node stores verified test rules and defect patterns. (2) Ecosystem closed loop: Utilize federated transfer learning to achieve cross-project knowledge reuse, automatically retrieve the validity of all network nodes when new use cases are generated, and write new defect patterns into the blockchain to complete the update.

Citation Information

Patent Citations

  • Method and device for generating test case

    CN115422074A

Cited By

  • Common cause failure risk analysis method and system based on software structure coupling network

    CN121233455A

  • Common cause failure risk analysis method and system based on software structure coupling network

    CN121233455B

  • Test field knowledge graph construction and intelligent use case generation method

    CN122633571A