Synthetic Data Generator for Software Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing software testing methods face challenges with managing data for regression testing, including cumbersome and time-consuming processes for handling full production-sized data sets, lack of ability to create relationally consistent subsets, and difficulties in modifying data due to interdependencies within data structures.
Innovation Solution
A data generator system that produces relationally consistent data, capable of generating supersets and subsets, and allows for statistical influence, using relationship algorithms to control data generation, and parallel execution to reduce time, eliminating the need for archiving data for regression testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If full production-sized data sets are used for testing, then data completeness is improved, but data management complexity and time consumption increase significantly
Solution Approach 1:
The patent creates synthetic copies of production data that replicate the statistical properties, relationships, and patterns of real data without requiring actual production data. The system generates test data sets that are complete and realistic in structure but synthetic in content, eliminating the need to harvest, cleanse, and transfer large volumes of production data while maintaining data completeness for testing purposes
2Reliability
If production data is harvested and cleansed for testing, then data quality is improved, but data processing complexity increases
Solution Approach 1:
The synthetic data generation system is self-configuring and automatically generates high-quality test data without requiring manual data harvesting, cleansing, or preparation steps. The system uses relationship algorithms and statistical models to automatically create data that meets quality requirements, eliminating complex data processing workflows while maintaining data reliability
3Stability of the object's composition
If relationally consistent subsets of production data are created, then data consistency is improved, but data extraction complexity increases
Solution Approach 1:
The system creates synthetic data subsets that replicate the relational structure and consistency properties of production data through algorithmic generation. Instead of extracting and transforming actual production data subsets, the system generates new data that inherently maintains referential integrity and relational consistency through its generation process, eliminating complex extraction and transformation operations
4Adaptability or versatility
If data is archived for later regression testing steps, then data availability is improved, but storage costs and retrieval complexity increase
Solution Approach 1:
The system generates synthetic data on-demand for each testing step rather than archiving production data for later use. Test data is created fresh when needed, eliminating the need for data archiving, storage management, and retrieval operations. The synthetic data generation process ensures data availability without the overhead of data lifecycle management
5Adaptability or versatility
If changes are made to traditionally generated test data, then test adaptability is improved, but data structure interdependencies create modification difficulties
Solution Approach 1:
The synthetic data generation system is highly configurable and dynamic, allowing test data characteristics to be adjusted by modifying generation parameters and relationship algorithms. The system can adapt data volume, complexity, statistical properties, and relational structures through configuration changes rather than data manipulation, making test adaptability easy while avoiding the interdependency problems of modifying statically generated or extracted data
Data Source
AI summary
According to one aspect, it is appreciated that it may be useful and particularly advantageous to provide a data generator that creates more realistic data for testing purposes, especially in data systems where large volumes of data are necessary. In one implementation, a data generator is provided that produces relationally consistent data for testing purposes. For instance, a synthetic data generation process may be performed that produces any number of relationally consistent data table structures. Further, in another implementation, generation of the data can be statistically influenced so that the data generated can take on the “look and feel” of production data. Also, data may be produced as needed, and its generation may be performed in parallel, depending on interdependencies in the data.


