Intelligent Software Test Data Generation Method and System Based on Internet Public Information

By using multi-objective regional Bayesian optimization technology, time enhancement and variational Bayesian method, natural language processing technology and reinforcement learning algorithm in software test data generation, the problems of low efficiency, insufficient authenticity and poor adaptability in the existing technology are solved, and high-quality, diverse and adaptable test data generation is achieved.

CN119473888BActive Publication Date: 2025-06-17ZHONGBEI UNIV

Patent Information

Application Number
CN202411549359.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-06-17
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

In the prior art, software test data generation efficiency is low, insufficient authenticity and poor adaptability, making it difficult to fully cover complex testing scenarios and changes in software requirements.

Method used

Using intelligent software test data generation method based on public information on the Internet, we generate high-quality, highly authentic and diverse test data through multi-objective regional Bayesian optimization technology, time enhancement and variational Bayesian method, natural language processing technology and reinforcement learning algorithms, and dynamically adjust the generation strategy to adapt to changes in software needs.

Benefits of technology

It improves the authenticity and diversity of test data, can better simulate practical application scenarios, improves the quality and efficiency of test data generation, and adapts to changes in software requirements and complex business logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119473888B_ABST
    Figure CN119473888B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent software test data generation method and system based on publicly available Internet information. The method includes: collecting data from the Internet based on software requirements and test scenarios and preprocessing it; enhancing the data using multi-objective region Bayesian optimization technology; processing time-series data using time enhancement and variational Bayesian methods; reading software development documents and extracting key information; performing semantic understanding and extension based on a hierarchical classification Gaussian extended semantic model; generating preliminary test data that conforms to business logic; applying the test data and collecting results; and dynamically adjusting the generation strategy using a reinforcement learning algorithm. The present invention also provides a corresponding system. The method and system can automatically generate high-quality test data, improving software test efficiency and quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software testing, and particularly to an intelligent software test data generation method and system based on publicly available Internet information. Background Art

[0002] Software testing is a key link to ensure software quality, and high-quality test data is crucial for effective software testing. Traditional software test data generation methods mainly rely on manual design and random generation, and these methods often have problems such as low efficiency and insufficient coverage.

[0003] Currently, some automated test data generation technologies have been developed, such as model-based test data generation and constraint-based test data generation. These methods have improved the generation efficiency and quality of test data to a certain extent. For example, the model-based method generates test data by establishing a formal model of the software system, while the constraint-based method generates test data that meets the requirements by parsing the constraint conditions in the software specification.

[0004] However, the existing automated test data generation technologies still have some limitations. First, these methods usually require a large amount of manual intervention, such as building an accurate model or defining detailed constraint conditions, which increases the workload of testers. Second, the generated test data often lacks authenticity and diversity, and it is difficult to comprehensively cover various complex test scenarios. Finally, the existing methods are difficult to adapt to the rapid changes in software requirements and complex business logics, resulting in the generated test data may deviate from the actual requirements.

[0005] Therefore, how to automatically generate high-quality, authentic and widely covered software test data, while being able to adapt to the changes in software requirements and complex business logics, is an urgent problem to be solved in the current software testing field. Summary of the Invention

[0006] In view of this, the present application provides an intelligent software test data generation method and system based on publicly available Internet information, which solves the problems of low efficiency, insufficient authenticity and poor adaptability in test data generation in the prior art.

[0007] The embodiment of the present application provides an intelligent software test data generation method based on Internet public information, including: collecting relevant data from Internet public information based on software requirements and test scenarios and performing preprocessing; using the multi-objective regional Bayesian optimization technology to enhance the preprocessed data; using the time enhancement and variational Bayesian method to process the time series data in the preprocessed data to generate enhanced time series data; reading the software development document, using natural language processing technology to extract key information, and performing semantic understanding and expansion on the extracted key information based on the hierarchical classification Gaussian extended semantic model; according to the results of semantic understanding and expansion, improving and adjusting the enhanced data and the enhanced time series data to generate preliminary test data that conforms to business logic; applying the generated preliminary test data to the software test process, collecting test results and performance indicators; based on the test results, using the reinforcement learning algorithm to dynamically adjust the test data generation strategy; repeating the above steps to continuously optimize the test data generation process and form a closed-loop intelligent test process.

[0008] Using the multi-objective regional Bayesian optimization technology to enhance the preprocessed data specifically includes: defining a set of optimization objective functions, modeling each objective function using a Gaussian process, selecting the next evaluation point by maximizing the expected improvement, and iteratively optimizing within a predefined search space until a Pareto optimal solution that meets the conditions is found.

[0009] Using the time enhancement and variational Bayesian method to process the time series data specifically includes: performing time warping and time interpolation processing on the original time series data, constructing a variational autoencoder model, learning the latent variables of the VAE by minimizing the reconstruction error and KL divergence, and generating new time series data using the learned VAE model.

[0010] Reading the software development document and using natural language processing technology to extract key information, including: preprocessing the document using lexical analysis and syntactic analysis technologies, identifying key entities and attributes using named entity recognition technology, extracting the relationships and constraint conditions between entities using dependency syntactic analysis technology, identifying the main themes and business rules using topic modeling technology, and calculating the semantic similarity between words using word vector technology.

[0011] Performing semantic understanding and expansion on the extracted information based on the hierarchical classification Gaussian extended semantic model specifically includes: constructing a hierarchical structure of semantic concepts, calculating the Mahalanobis distance between the input semantics and each concept node, determining the concept category to which the input semantics belongs, generating new semantics related to the input semantics, and expanding the original semantic space.

[0012] Based on the results of semantic understanding and expansion, improve and adjust the enhanced data and the enhanced time-series data, specifically including: performing consistency checks, adjusting outliers and boundary values, supplementing missing key attributes and relationships, generating diverse combinations of test data, and validating the generated test data.

[0013] Based on the test results, use a reinforcement learning algorithm to dynamically adjust the test data generation strategy, specifically including: modeling the test data generation process as a Markov decision process, defining state, action, and reward functions, using a deep Q-network to learn the optimal data generation strategy, and adopting experience replay and target network techniques to optimize the learning process.

[0014] An embodiment of this application also provides an intelligent software test data generation system based on Internet public information, including a data collection and preprocessing module, a data enhancement module, a semantic understanding module, a data improvement module, a test execution module, and a dynamic adjustment module. The system may also include a data storage module, a model training module, and a user interface module.

[0015] An embodiment of this application also provides a computer device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned intelligent software test data generation method based on Internet public information.

[0016] An embodiment of this application also provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the above-mentioned intelligent software test data generation method based on Internet public information.

[0017] An embodiment of this application also provides a computer program product, including computer instructions, which implement the steps of the above-mentioned intelligent software test data generation method based on Internet public information when executed by a processor.

[0018] This application has the following technical effects: By collecting data from Internet public information and performing intelligent processing, the authenticity and diversity of test data are improved, and it can better simulate actual application scenarios. By adopting a variety of advanced machine learning and natural language processing technologies, such as multi-objective region Bayesian optimization, variational Bayesian methods, and hierarchical classification Gaussian extended semantic models, etc., the quality and efficiency of test data generation are greatly improved. By introducing a reinforcement learning algorithm to dynamically adjust the test data generation strategy, continuous optimization and adaptability of the test process are achieved, and it can better cope with changes in software requirements and complex business logics. Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following accompanying drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.

[0020] Figure 1 is a flowchart of an intelligent software test data generation method based on Internet public information provided by an embodiment of the present invention;

[0021] Figure 2 is a structural block diagram of an intelligent software test data generation system based on Internet public information provided by an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of a multi-objective region Bayesian optimization technique in an embodiment of the present invention;

[0023] Figure 4 is a flowchart of a time series data processing method using time enhancement and variational Bayesian methods in an embodiment of the present invention;

[0024] Figure 5 is a structural schematic diagram of a hierarchical classification Gaussian extended semantic model in an embodiment of the present invention;

[0025] Figure 6 is a flowchart of a reinforcement learning algorithm for dynamically adjusting a test data generation strategy in an embodiment of the present invention. Detailed Embodiments

[0026] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0027] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present invention. It should be understood that when an element in an embodiment of the present invention is referred to as being "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0028] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0029] The core of the present invention is to provide an intelligent software test data generation method and system based on Internet public information. The following will further describe the present invention in detail with reference to the accompanying drawings and specific embodiments.

[0030] It should be understood that the terms "including" and "having" used herein and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0031] Embodiment 1

[0032] As Figure 1 shown, the intelligent software test data generation method based on Internet public information provided by the embodiments of the present invention includes the following steps:

[0033] S1: Data collection and preprocessing

[0034] S1.1: Collect relevant data from Internet public information based on software requirements and test scenarios

[0035] The purpose of this step is to obtain real data related to the software to be tested, so as to improve the authenticity and diversity of the generated test data. The specific implementation can include the following aspects:

[0036] Requirement Analysis: Analyze the software requirement document to extract keywords and themes.

[0037] Web Crawler: Based on the extracted keywords and themes, develop a targeted web crawler to collect data from relevant websites, forums, social media, and other channels.

[0038] API Call: Utilize public data API interfaces, such as government open data platforms, meteorological data interfaces, etc., to obtain structured data in related fields.

[0039] Data Filtering: According to the requirements of the test scenario, conduct preliminary screening and filtering on the collected data to remove obviously irrelevant or low-quality data.

[0040] S1.2: Preprocess the collected data

[0041] Data preprocessing is to prepare for subsequent processing steps, mainly including the following operations:

[0042] Data Cleaning: Remove duplicate data, handle missing values, correct outliers, etc.

[0043] Data Transformation: Unify the data format, such as date formatting, numerical normalization, etc.

[0044] Data Integration: Integrate data from different sources together to establish a unified data structure.

[0045] Feature Extraction: According to the requirements of software testing, extract useful features from the original data.

[0046] Data Classification: Classify the data according to different test scenarios or functional modules for subsequent targeted processing.

[0047] S2: Data Augmentation

[0048] S2.1: Use multi-objective regional Bayesian optimization technology to augment the preprocessed data

[0049] As Figure 3 shown, multi-objective regional Bayesian optimization (MORBO) is an efficient multi-objective optimization method, especially suitable for scenarios such as test data generation that require balancing multiple objectives. The specific implementation steps are as follows:

[0050] 1. Define the set of optimization objective functions {f1(x), f2(x),..., f k (x)}, where x represents the parameter set for data augmentation and k is the number of optimization objectives. These objective functions can include:

[0051] Data diversity: Measure the degree of variation in the generated data

[0052] Boundary value coverage: Evaluate the coverage of the generated data for boundary conditions

[0053] Business rule compliance: Measure the degree of compliance between the generated data and business rules

[0054] Test coverage: Evaluate the coverage of the generated data for software functions

[0055] 2. Model each objective function using Gaussian processes. Gaussian processes can capture the uncertainty of the function and help balance exploration and exploitation during the optimization process.

[0056] 3. Select the next evaluation point by maximizing the Expected Improvement. Expected Improvement takes into account the current optimal solution and the predicted uncertainty, and can effectively guide the search process.

[0057] 4. Iteratively optimize within the predefined search space until a Pareto optimal solution that meets the conditions is found. A Pareto optimal solution is a solution where no other objective can be improved without sacrificing any one objective.

[0058] Through the MORBO technique, the embodiments of the present invention can generate high-quality test data that meets multiple optimization objectives, ensuring both data diversity and good coverage of software functions and business rules.

[0059] S2.2: Process time series data using time augmentation and variational Bayesian methods

[0060] For test data containing time information, the embodiments of the present invention use time augmentation and variational Bayesian methods (such as Figure 4 ) for special processing to generate more realistic and diverse time series data.

[0061] The specific steps are as follows:

[0062] 1. Perform time stretching and time interpolation on the original time series data:

[0063] Time stretching warping: Adjust the local speed of the time series through non-linear transformation to simulate time stretching in real scenarios.

[0064] Time interpolation: Insert new data points between the original time points to increase the density and continuity of the time series data.

[0065] 2. Build a variational autoencoder (VAE) model, where both the encoder and decoder adopt a recurrent neural network (RNN) structure:

[0066] Encoder: Maps the input time series data to the latent space and captures the high-level features of the data.

[0067] Decoder: Reconstructs the time series data from the latent space and generates new data samples.

[0068] Using the RNN structure can effectively capture the long-term dependencies of time series data.

[0069] 3. Learn the latent variable z of the VAE by minimizing the reconstruction error and KL divergence:

[0070] Reconstruction error: Ensures that the generated data has a similar distribution to the original data.

[0071] KL divergence: Promotes the distribution in the latent space to be close to the standard normal distribution, facilitating subsequent sampling and generation.

[0072] 4. Use the learned VAE model to generate new time series data while maintaining the time characteristics of the data:

[0073] Sample the latent variable from the standard normal distribution. Input the sampled latent variable into the decoder to generate new time series data. Post-process the generated data to ensure the preservation of time characteristics (such as periodicity, trend).

[0074] Through this method, the embodiments of the present invention can generate new time series data that has similar statistical characteristics to the original data but is not exactly the same, greatly enriching the time dimension of the test data set.

[0075] Exemplarily, taking the order system of an e-commerce platform as an example to illustrate this process.

[0076] First, the embodiments of the present invention collect a batch of initial order data from the publicly available information on the Internet. This data includes basic information such as order amount, commodity category, order time, etc. However, there may be problems with these original data, such as uneven distribution and insufficient coverage of boundary cases. For example, the embodiments of the present invention may lack large-order or order data during specific festivals.

[0077] To solve this problem, the embodiments of the present invention first apply the multi-objective regional Bayesian optimization technology. The embodiments of the present invention define multiple optimization objectives, such as increasing the proportion of large orders, improving the coverage of festival orders, maintaining the diversity of commodity categories, etc. The system will generate new order data through iterative optimization according to these objectives. For example, it may generate some orders close to the maximum order amount limit of the platform, or high-frequency purchase behaviors during important festivals.

[0078] Next, the embodiments of the present invention focus on the processing of time series data. The data in the order system has obvious time characteristics, such as periodic fluctuations (increased order volume on weekends) and seasonal trends (sales peaks during holidays). The embodiments of the present invention use time augmentation and variational Bayesian methods to process these time series characteristics. Specifically, the embodiments of the present invention may stretch or compress the original time series to simulate the order distribution under different time scales. For example, the embodiments of the present invention can generate data simulating intensive order placement during a long holiday, or simulate sparse order placement during late-night hours.

[0079] In this way, the embodiments of the present invention not only increase the quantity of test data, but more importantly, improve the quality and coverage of the data. The enhanced dataset contains more boundary cases and abnormal scenarios, such as orders with extremely large amounts, frequently cancelled orders, international orders spanning multiple time zones, etc. These data can help the embodiments of the present invention more comprehensively test the various functions and performances of the order system.

[0080] It should be noted that throughout the augmentation process, the embodiments of the present invention always respect the original data distribution. Although the newly generated data expands the original distribution range, it still maintains reasonable statistical characteristics to ensure the authenticity and effectiveness of the test.

[0081] Through such a data augmentation process, the embodiments of the present invention finally obtain a more rich and diverse test dataset. This dataset not only contains regular order situations, but also covers various extreme situations and boundary conditions, laying a solid foundation for subsequent comprehensive testing. This method greatly increases the possibility for the embodiments of the present invention to discover potential bugs and system defects, thereby helping the embodiments of the present invention build a more robust and reliable order system.

[0082] S3: Semantic Understanding and Expansion

[0083] S3.1: Read the software development documentation and use natural language processing techniques to extract key information

[0084] This step aims to extract key information from the software development documentation that is helpful for generating high-quality test data. The specific implementation includes the following aspects:

[0085] 1. Preprocess the software development documentation using lexical analysis and syntactic analysis techniques:

[0086] Lexical analysis: Split the text into words or tokens, and identify the part of speech and word form. Syntactic analysis: Analyze the grammatical structure of the sentence and construct a syntactic tree.

[0087] 2. Adopt named entity recognition technology to identify key entities and attributes in the document:

[0088] Identify entities in specific domains, such as users, products, functional modules, etc. Extract attribute information of entities, such as data types, value ranges, etc.

[0089] 3. Use dependency syntactic analysis technology to extract relationships and constraints between entities:

[0090] Analyze the dependency relationships between entities, such as subject-predicate relationships, modification relationships, etc. Extract the constraints in the business logic, such as "the user age must be greater than 18 years old".

[0091] 4. Apply topic modeling technology to identify the main topics and business rules in the document:

[0092] Use algorithms such as LDA (Latent Dirichlet Allocation) to extract the topics of the document. Identify important business rules and functional descriptions based on the topic distribution.

[0093] 5. Use word vector technology to calculate the semantic similarity between words, assisting semantic understanding and expansion:

[0094] Train or use pre-trained word vector models (such as Word2Vec, GloVe). Calculate the cosine similarity between keywords to discover potential semantic associations.

[0095] Through these NLP technologies, the embodiments of the present invention can extract rich semantic information from software development documents, providing guidance and constraints for subsequent test data generation.

[0096] S3.2: Semantically understand and expand the extracted key information based on the hierarchical classified Gaussian expansion semantic model

[0097] As Figure 5 shown, the hierarchical classified Gaussian expansion semantic model (HCGESM) is a new semantic understanding and expansion model that can effectively capture the hierarchical relationships and semantic similarities between concepts. The specific implementation steps are as follows:

[0098] 1. Construct a hierarchical structure of semantic concepts, and each concept node i is represented by a Gaussian distribution N(μ i , Σ i ):

[0099] According to domain knowledge and document content, construct a concept tree. Assign a multi-dimensional Gaussian distribution to each concept node, where the mean μ i represents the central position of the concept, and the covariance matrix Σ i represents the uncertainty and coverage of the concept.

[0100] 2. For a given input semantics x, calculate its Mahalanobis distance from each concept node.

[0101] The Mahalanobis distance takes into account the uncertainty of concepts and is more suitable for measuring the semantic space than the Euclidean distance.

[0102] 3. Based on the calculated Mahalanobis distance, determine the concept category to which the input semantics x belongs:

[0103] Nearest neighbor classification or soft classification methods can be used.

[0104] Considering the hierarchical relationship of concepts, multi-level concept attribution is allowed.

[0105] 4. Using the determined concept category, based on the hierarchical relationship and similarity between concepts, generate new semantics related to the input semantics:

[0106] Sample in the Gaussian distribution of the belonging concept to generate semantic vectors that are similar but not exactly the same.

[0107] Considering parent-child concepts and sibling concepts, generate related but different semantics.

[0108] 5. Use the generated new semantics to expand the original semantic space, increasing the diversity and coverage of test data:

[0109] Map the generated new semantic vectors back to the original feature space.

[0110] Generate test data examples that conform to specific concepts or rules based on the new semantics.

[0111] Through HCGESM, the embodiments of the present invention can achieve in-depth understanding and intelligent expansion of the extracted key information, generate richer and more diverse test data, and at the same time maintain consistency with the original business semantics.

[0112] Take a smart home control system as an example:

[0113] In a smart home system, users may use various different language expressions to control the devices at home. For example, instructions such as "Turn on the lights in the living room", "Turn on the lights", "Illuminate the living room" etc. actually express the same intention. The goal of the embodiments of the present invention is to understand these different expressions and expand to more possible semantic variants in order to more comprehensively test the speech recognition and command processing functions of the system.

[0114] First, the embodiment of the present invention constructs a hierarchical structure of semantic concepts. In this structure, "device control" may be the top-level concept, which includes sub-concepts such as "lighting control", "temperature control", "security control", etc. "Lighting control" is further divided into "turn on the light", "turn off the light", "adjust the brightness", etc. Each concept node is represented by a Gaussian distribution, where the mean represents the central position of the concept, and the covariance matrix represents the uncertainty and coverage of the concept.

[0115] Suppose the embodiment of the present invention extracts the key information "turn on the light in the living room" from the software development document. The system will first calculate the Mahalanobis distance between this input and each concept node. In this example, it may be closest to the distances of "lighting control" and "turn on the light".

[0116] Based on this calculation result, the system determines that the input semantics belong to the concept category of "turn on the light". Then, it will use the Gaussian distribution of this concept to generate relevant new semantics. For example, it may generate variants such as "turn on the living room lighting", "light up the living room", "activate the ceiling light in the living room", etc.

[0117] In addition, the system will also consider the hierarchical relationship and similarity between concepts. For example, the parent concept of "turn on the light" is "lighting control", and the sibling concepts include "turn off the light" and "adjust the brightness". Based on these relationships, the system may further expand and generate related but slightly different semantics such as "adjust the living room light to the brightest" and "turn off the light in the living room".

[0118] In this way, the embodiment of the present invention not only understands the original instruction "turn on the light in the living room", but also expands a series of related semantic variants. These variants include synonymous expressions, similar instructions, and even some opposing instructions (such as "turn off the light"), which helps to test the robustness of the system.

[0119] It should be noted that when generating these semantic variants, the system will consider the language characteristics of the entire smart home field. For example, verbs such as "turn on", "open", "start" may be common in different device controls, and the system will learn and apply these common patterns to generate more natural language expressions.

[0120] This method of semantic understanding and expansion enables the embodiment of the present invention to generate more comprehensive and diverse test cases. For example, the embodiment of the present invention can test whether the system can correctly understand indirect expressions such as "illuminate the living room", or test how it processes composite instructions such as "turn on the light in the living room, but dim it a little". In this way, the embodiment of the present invention can more comprehensively evaluate the language understanding ability and command execution accuracy of the system.

[0121] Generally speaking, this method based on the hierarchical classification Gaussian extended semantic model enables the embodiments of the present invention to generate rich and diverse test data starting from limited initial information, thereby better verifying the functional integrity and user experience of the smart home control system. This not only improves the test coverage rate but also enhances the system's ability to handle various real user inputs.

[0122] S4: Test Data Generation and Improvement

[0123] In the test data generation and improvement stage, the present invention adopts a series of innovative technologies to ensure that the generated test data not only conforms to the business logic but also comprehensively covers various test scenarios. First, the system performs a consistency check on the preliminarily generated test data according to the business rules extracted previously. This step usually uses a rule engine or a constraint satisfaction problem (CSP) solver to verify whether the data meets all known business rules. For data that does not meet the rules, the system will make corresponding adjustments or regenerate it.

[0124] Secondly, based on the identified data constraint conditions, the present invention makes fine adjustments to the outliers and boundary values in the test data. This process includes identifying the valid value ranges and constraint conditions for each field, and then generating test data that covers the boundary conditions and abnormal situations. In this way, the embodiments of the present invention ensure that the generated data set not only contains normal values but also boundary values and outliers, thereby comprehensively testing the robustness of the software.

[0125] Thirdly, using the semantic understanding results obtained in the previous steps, the system supplements the key attributes and relationships that may be missing in the preliminary test data. This step mainly identifies the potentially missing attributes based on the semantic extension results generated by the HCGESM model. At the same time, the embodiments of the present invention also use association rule mining technology to discover the potential relationships between attributes and reasonably supplement the values of the missing attributes according to these discoveries.

[0126] In addition, to meet the requirements of different business scenarios, the present invention adopts a variety of methods to generate diverse combinations of test data. The embodiments of the present invention use the combinatorial testing method to generate test cases that cover key parameter combinations, apply the Monte Carlo method to simulate random scenarios with different probability distributions, and consider the time series characteristics to generate data changes reflecting different time scales (such as days, weeks, months, quarters).

[0127] Finally, a comprehensive verification of the generated test data is carried out to ensure that it meets the functional and performance requirements of the software. This includes using static analysis tools to check whether the format and structure of the data are correct, applying dynamic verification methods (such as verifying the validity of the data in a simulated test environment), and collecting feedback from domain experts to ensure the rationality of the generated data in terms of business logic.

[0128] Through this series of steps, the embodiments of the present invention finally obtain a set of high-quality preliminary test data that not only conforms to the business logic but also comprehensively covers various test scenarios. These data not only cover all functional modules and business scenarios of the software but also include data for normal, boundary, and abnormal situations, reflecting the real data distribution and business rules, and having timeliness and relevance, effectively simulating the actual business process.

[0129] S5: Test Execution and Result Collection

[0130] In the test execution and result collection phase, the present invention actually applies the generated test data to software testing and comprehensively collects test results and various performance indicators. First, the embodiments of the present invention need to prepare the test environment well. This includes building a test environment similar to the production environment and configuring necessary test tools and frameworks, such as automated test tools, performance monitoring tools, etc.

[0131] Secondly, based on the generated test data, the embodiments of the present invention design test cases covering various scenarios, including multiple dimensions such as functional testing, performance testing, and security testing. To improve test efficiency, the embodiments of the present invention use test automation tools (such as Selenium, JMeter, etc.) to write test scripts and embed the generated test data into these automated scripts.

[0132] During the test execution process, the embodiments of the present invention not only run automated test scripts and execute the designed test cases but also conduct manual testing on test scenarios that cannot be automated. At the same time, the embodiments of the present invention use monitoring tools to observe the behavior and performance of the software during the test in real time and record key indicators, such as response time, resource occupancy, error rate, etc.

[0133] The collection of test results is a comprehensive and detailed process. The embodiments of the present invention record the pass / fail status of each test case, capture and record the exceptions and error information that occur during the test, and collect test logs containing detailed execution steps and intermediate results. In addition, the embodiments of the present invention also collect key performance indicators, such as response time, throughput, CPU usage, memory occupancy, etc., as well as the results of load testing and stress testing to evaluate the extreme performance of the system.

[0134] In the embodiments of the present invention, the embodiments of the present invention also particularly focus on the collection of user experience data. If possible, the embodiments of the present invention will collect experience data simulating user operations, such as page loading time, operation fluency, etc., which are very important for evaluating the actual use experience of the software.

[0135] In addition, the embodiments of the present invention use a code coverage tool to analyze the code coverage of the tests and evaluate the coverage of the test data for the business scenarios. Finally, the embodiments of the present invention perform a preliminary analysis on the collected data, generate test reports and performance charts, and use data visualization tools to intuitively display the test results and performance trends.

[0136] Through this comprehensive test execution and result collection process, the embodiments of the present invention can comprehensively evaluate the quality and performance of the software, providing a solid data foundation for subsequent optimization.

[0137] S6: Dynamic adjustment and optimization

[0138] S6.1: Based on the test results, use a reinforcement learning algorithm to dynamically adjust the test data generation strategy

[0139] In the dynamic adjustment and optimization stage, the present invention introduces a reinforcement learning algorithm to continuously optimize the test data generation process. As Figure 6 shown, first, the embodiments of the present invention model the test data generation process as a Markov decision process (MDP). In this model, the state represents the characteristics of the current test data set, such as data distribution, coverage rate, etc.; the action represents possible data generation or modification operations, such as adding specific types of data, adjusting data distribution, etc.; and the reward function is defined based on the test results, reflecting the quality of the test data.

[0140] To better represent and process this complex decision-making problem, the embodiments of the present invention use a vectorization method to represent the characteristics of the test data set, such as the proportion of various types of data, the number of covered scenarios, etc. At the same time, the embodiments of the present invention define a series of possible operations as the action space, such as "adding boundary value data", "adjusting the periodicity of time series data", etc.

[0141] In the present invention, the design of the reward function is a key link. The embodiments of the present invention comprehensively consider factors such as test coverage rate, defect discovery rate, and test efficiency. The test coverage rate reflects the coverage of the test data for the software functions, the defect discovery rate measures the ability of the test data to discover software defects, and the test efficiency takes into account the test execution time and resource consumption. These factors are combined through weighting to form the final reward function, and the weights can be adjusted according to specific requirements.

[0142] To learn the optimal data generation strategy, the embodiments of the present invention adopt a Deep Q-Network (DQN). DQN is a reinforcement learning algorithm that combines deep learning and Q-learning, and is particularly suitable for dealing with problems in large-scale state spaces. In the implementation of the embodiments of the present invention, a deep neural network serves as an approximator of the Q function, with the state representation as the input and the Q values for each possible action as the output. The embodiments of the present invention use an ε-greedy strategy to balance exploration and exploitation, ensuring that the algorithm can explore potential better strategies while utilizing the known optimal strategy.

[0143] It should be noted that, to optimize the learning process of DQN, the embodiments of the present invention adopt the experience replay and target network techniques. Experience replay improves the sample utilization efficiency by storing and reusing past experiences. The target network, on the other hand, improves the training stability by using a separate network to calculate the target Q values. The combination of these two techniques greatly enhances the learning efficiency and effect.

[0144] In practical applications, the embodiments of the present invention continuously learn complex non-linear strategies through continuous iteration and optimization, achieving continuous improvement in the test data generation process. The system will regularly update the Q network to adapt to changes in software and test requirements. At the same time, experiences are continuously accumulated in practical applications to further optimize the decision-making strategy.

[0145] Through this dynamic adjustment mechanism based on reinforcement learning, the system of the embodiments of the present invention can continuously adjust and optimize the test data generation strategy according to the actual test effect, thereby continuously improving the test efficiency and quality.

[0146] S6.2: Repeat the above steps to continuously optimize the test data generation process and form a closed-loop intelligent test process

[0147] Based on the dynamic adjustment strategy of reinforcement learning, the present invention constructs a closed-loop intelligent test process. First, the embodiments of the present invention will regularly evaluate the quality of the current test data and the test effect, and analyze the changing trends of key indicators such as test coverage rate and error discovery rate. This periodic evaluation provides an important basis for subsequent optimization.

[0148] Secondly, based on the evaluation results, the embodiments of the present invention will update the reinforcement learning model. This includes adjusting the reward function to reflect the latest test requirements and priorities. In this way, the embodiments of the present invention ensure that the system can continuously adapt to the changing test environment and objectives.

[0149] After updating the strategy, the system will apply the new strategy to generate test data. In this process, the embodiments of the present invention will gradually eliminate inefficient test data and increase high-quality test cases. This dynamic adjustment ensures the continuous optimization of the test data set.

[0150] In addition, an important feature of the present invention is continuous learning. In the embodiments of the present invention, newly generated data and their test results are continuously added to the experience pool, constantly enriching and updating the knowledge base of the model. This mechanism enables the system to learn from each test and continuously improve its ability to generate high-quality test data.

[0151] It should be emphasized that the system of the embodiments of the present invention has strong adaptability. According to changes such as software version updates and new function additions, the system can dynamically adjust the goals and strategies for test data generation. At the same time, it maintains sensitivity to newly emerging boundary conditions and abnormal situations, ensuring the comprehensiveness and effectiveness of testing.

[0152] Through this closed-loop intelligent testing process, the present invention can continuously improve the quality and efficiency of test data, adapt to the dynamic changes in software development, and ultimately achieve the goal of continuous improvement of software quality. This method not only greatly improves the testing efficiency but also can effectively cope with various challenges in the software development process, providing strong support for improving software quality.

[0153] Generally speaking, the intelligent software test data generation method and system provided by the present invention, by combining a variety of advanced artificial intelligence technologies, realizes the automatic generation of high-quality, highly authentic, and wide-coverage test data. This method not only improves the testing efficiency but also can adapt to changes in software requirements and complex business logics, providing an innovative solution for improving software quality.

[0154] Embodiment 2

[0155] The embodiments of the present application also provide an intelligent software test data generation system based on publicly available Internet information, as Figure 2 shown, including:

[0156] A data collection and preprocessing module, configured to collect relevant data from publicly available Internet information based on software requirements and test scenarios and perform preprocessing;

[0157] A data enhancement module, configured to enhance the preprocessed data by using multi-objective region Bayesian optimization technology, and process the time-series data in the preprocessed data by using time enhancement and variational Bayesian methods to generate enhanced time-series data;

[0158] A semantic understanding module, configured to read software development documents, extract key information by using natural language processing technology, and perform semantic understanding and expansion on the extracted key information based on a hierarchical classification Gaussian extended semantic model;

[0159] A data improvement module, configured to improve and adjust the enhanced data and the enhanced time-series data according to the results of semantic understanding and expansion to generate preliminary test data that conforms to business logics;

[0160] A test execution module, which is used to apply the generated preliminary test data to the software test process and collect test results and performance metrics;

[0161] A dynamic adjustment module, which is used to dynamically adjust the generation strategy of test data based on the test results by using a reinforcement learning algorithm.

[0162] A data storage module, which is used to store the collected relevant data, the generated test data and the test results;

[0163] A model training module, which is used to train and update various machine learning models used in the system;

[0164] A user interface module, which is used to receive the software requirements and test scenarios input by the user and display the test results and performance metrics.

[0165] An embodiment of the present application further provides a computer device, which includes:

[0166] At least one processor; and,

[0167] A memory communicatively connected to the at least one processor; wherein,

[0168] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned intelligent software test data generation method based on publicly available Internet information.

[0169] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the above-mentioned intelligent software test data generation method based on publicly available Internet information.

[0170] An embodiment of the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the above-mentioned intelligent software test data generation method based on publicly available Internet information are implemented.

[0171] In the description and claims of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than that shown or described in words.

[0172] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0173] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A method for generating intelligent software test data based on Internet public information, characterized in that: The following steps are involved: Based on software requirements and test scenarios, relevant data is collected from public information on the Internet and preprocessed; Multi-objective regional Bayesian optimization techniques are used to enhance the preprocessed data; Use time enhancement and variational Bayesian methods to process the time series data in the preprocessed data to generate enhanced time series data; Read software development documents, extract key information using natural language processing technology, and perform semantic understanding and expansion of the extracted key information based on the hierarchical classification Gaussian extended semantic model; According to the results of semantic understanding and expansion, the enhanced data and enhanced time series data are improved and adjusted to generate preliminary test data that conforms to business logic; Apply the generated preliminary test data to the software testing process and collect test results and performance indicators; Based on the test results, the test data generation strategy is dynamically adjusted using reinforcement learning algorithms; Repeat the above steps to continuously optimize the test data generation process and form a closed-loop intelligent testing process; The semantic understanding and expansion of the extracted information based on the hierarchical classification Gaussian extended semantic model specifically includes: Construct a hierarchy of semantic concepts, and use Gaussian distribution N(μ i ,Σ i )express; For a given input semantics x, calculate its Mahalanobis distance with each concept node; Based on the calculated Mahalanobis distance, determine the concept category to which the input semantics x belongs; Using the determined concept categories, new semantics related to the input semantics are generated based on the hierarchical relationship and similarity between concepts; Use the generated new semantics to expand the original semantic space and increase the diversity and coverage of test data.

2. The method according to claim 1, characterized in that: The multi-objective regional Bayesian optimization technology is used to enhance the pre-processed data, specifically including: Define the optimization objective function set {f1(x), f2(x), ..., f k (x)}, where x represents the parameter set of data enhancement and k is the number of optimization objectives; Each objective function was modeled using a Gaussian process; Select the next evaluation point by maximizing expected improvement; Iterate the optimization within the predefined search space until a Pareto optimal solution that meets the conditions is found.

3. The method according to claim 1, characterized in that The time series data is processed using the time enhancement and variational Bayes method, specifically including: Perform time warping and time interpolation processing on the original time series data; Construct a variational autoencoder VAE model, in which both the encoder and decoder use a recurrent neural network RNN ​​structure; Learn the VAE’s hidden variable z by minimizing the reconstruction error and KL divergence; Utilize the learned VAE model to generate new time series data while maintaining the temporal characteristics of the data.

4. The method according to claim 1, characterized in that: The software development document is read and key information is extracted using natural language processing technology, including: Use lexical analysis and syntactic analysis techniques to preprocess software development documents; Use named entity recognition technology to identify key entities and attributes in documents; Use dependency parsing techniques to extract relationships and constraints between entities; Apply topic modeling techniques to identify major topics and business rules in documents; Use word vector technology to calculate the semantic similarity between words to assist semantic understanding and expansion.

5. The method according to claim 1, characterized in that According to the results of semantic understanding and expansion, the enhanced data and enhanced time series data are improved and adjusted, specifically including: Perform consistency check on the generated preliminary test data according to the extracted business rules; Adjust outliers and boundary values ​​in preliminary test data based on identified data constraints; Use semantic understanding results to supplement key attributes and relationships missing in preliminary test data; Generate a variety of test data combinations based on the business scenario requirements analyzed; The generated test data is verified to ensure that it meets the functional and performance requirements of the software.

6. The method according to claim 1, characterized in that Based on the test results, the test data generation strategy is dynamically adjusted using a reinforcement learning algorithm, specifically including: Model the generation process of test data as a Markov decision process MDP; Define the state in the MDP as the current test data set, and the action as the possible data generation or modification operation; Define a reward function based on test coverage, error detection rate, and test efficiency; Use the deep Q network DQN to learn the optimal data generation strategy; Use experience replay and target network technology to optimize the learning process of DQN; Through continuous iteration and optimization, complex nonlinear strategies are learned to achieve continuous improvement in the test data generation process.

7. An intelligent software test data generation system based on Internet public information, characterized in that: include: Data collection and preprocessing module, used to collect relevant data from public information on the Internet and perform preprocessing based on software requirements and test scenarios; A data enhancement module is used to enhance the preprocessed data using a multi-objective regional Bayesian optimization technique, and to process the time series data in the preprocessed data using a time enhancement and variational Bayesian method to generate enhanced time series data; The semantic understanding module is used to read software development documents, extract key information using natural language processing technology, and perform semantic understanding and expansion of the extracted key information based on the hierarchical classification Gaussian extended semantic model; The data improvement module is used to improve and adjust the enhanced data and enhanced time series data according to the results of semantic understanding and expansion, and generate preliminary test data that conforms to business logic; A test execution module, which is used to apply the generated preliminary test data to the software testing process and collect test results and performance indicators; A dynamic adjustment module is used to dynamically adjust the test data generation strategy based on the test results using a reinforcement learning algorithm; The semantic understanding and expansion of the extracted information based on the hierarchical classification Gaussian extended semantic model specifically includes: Construct a hierarchy of semantic concepts, and use Gaussian distribution N(μ i ,Σ i )express; For a given input semantics x, calculate its Mahalanobis distance with each concept node; Based on the calculated Mahalanobis distance, determine the concept category to which the input semantics x belongs; Using the determined concept categories, new semantics related to the input semantics are generated based on the hierarchical relationship and similarity between concepts; Use the generated new semantics to expand the original semantic space and increase the diversity and coverage of test data.

8. The system according to claim 7, characterized in that Also includes: A data storage module, used to store collected relevant data, generated test data and test results; Model training module, used to train and update various machine learning models used in the system; The user interface module is used to receive software requirements and test scenarios input by users and display test results and performance indicators.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Automatic testing method for intelligent software

    CN116594913A

  • Loop code test data generation method based on deep learning fuzzy test

    CN116680164A

Cited By

  • Multilayer architecture-oriented software development intelligent optimization deployment system, method and equipment

    CN121209891A