Synthetic Data Generator for Software Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing software testing methods face challenges with managing data for regression testing, including cumbersome and time-consuming processes for handling full production-sized data sets, lack of ability to create relationally consistent subsets, and difficulties in modifying data due to interdependencies within data structures.

Innovation Solution

A data generator system that produces relationally consistent data, capable of generating supersets and subsets, and allows for statistical influence, using relationship algorithms to control data generation, and parallel execution to reduce time, eliminating the need for archiving data for regression testing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If full production-sized data sets are used for testing, then data completeness is improved, but data management complexity and time consumption increase significantly

Engineering Contradiction:
Improvedata volumeVSAvoiddata management time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of production data that replicate the statistical properties, relationships, and patterns of real data without requiring actual production data. The system generates test data sets that are complete and realistic in structure but synthetic in content, eliminating the need to harvest, cleanse, and transfer large volumes of production data while maintaining data completeness for testing purposes

Inventive Principle:
Principle #26Copying

2Reliability

If production data is harvested and cleansed for testing, then data quality is improved, but data processing complexity increases

Engineering Contradiction:
Improvedata qualityVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The synthetic data generation system is self-configuring and automatically generates high-quality test data without requiring manual data harvesting, cleansing, or preparation steps. The system uses relationship algorithms and statistical models to automatically create data that meets quality requirements, eliminating complex data processing workflows while maintaining data reliability

Inventive Principle:
Principle #25Self-service

3Stability of the object's composition

If relationally consistent subsets of production data are created, then data consistency is improved, but data extraction complexity increases

Engineering Contradiction:
Improvedata consistencyVSAvoiddata extraction complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The system creates synthetic data subsets that replicate the relational structure and consistency properties of production data through algorithmic generation. Instead of extracting and transforming actual production data subsets, the system generates new data that inherently maintains referential integrity and relational consistency through its generation process, eliminating complex extraction and transformation operations

Inventive Principle:
Principle #26Copying

4Adaptability or versatility

If data is archived for later regression testing steps, then data availability is improved, but storage costs and retrieval complexity increase

Engineering Contradiction:
Improvedata availabilityVSAvoiddata retrieval complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system generates synthetic data on-demand for each testing step rather than archiving production data for later use. Test data is created fresh when needed, eliminating the need for data archiving, storage management, and retrieval operations. The synthetic data generation process ensures data availability without the overhead of data lifecycle management

Inventive Principle:
Principle #26Copying

5Adaptability or versatility

If changes are made to traditionally generated test data, then test adaptability is improved, but data structure interdependencies create modification difficulties

Engineering Contradiction:
Improvetest adaptabilityVSAvoiddata modification ease
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The synthetic data generation system is highly configurable and dynamic, allowing test data characteristics to be adjusted by modifying generation parameters and relationship algorithms. The system can adapt data volume, complexity, statistical properties, and relational structures through configuration changes rather than data manipulation, making test adaptability easy while avoiding the interdependency problems of modifying statically generated or extracted data

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9710365B2System and method for generating synthetic data for software testing purposes
Publication Date: 2017.07.18 WALMART APOLLO LLC
  • US9710365B2 patent drawing
  • US9710365B2 patent drawing
  • US9710365B2 patent drawing

AI summary

According to one aspect, it is appreciated that it may be useful and particularly advantageous to provide a data generator that creates more realistic data for testing purposes, especially in data systems where large volumes of data are necessary. In one implementation, a data generator is provided that produces relationally consistent data for testing purposes. For instance, a synthetic data generation process may be performed that produces any number of relationally consistent data table structures. Further, in another implementation, generation of the data can be statistically influenced so that the data generated can take on the “look and feel” of production data. Also, data may be produced as needed, and its generation may be performed in parallel, depending on interdependencies in the data.