Seed Data Generation Using Statistical Distribution Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating seed data are costly and time-consuming, often resulting in unrealistic data that lacks anomalies, which can lead to inadequate testing and demonstration of applications, as they do not accurately replicate the statistical and semantic distributions of real data.

Innovation Solution

A computer-implemented method that analyzes original data to determine its distribution characteristics, including statistical and semantic patterns, and generates seed data that mimics these characteristics, incorporating anomalies found in the original data, to create realistic data for testing and demonstration purposes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to generate seed data manually or using simple algorithms, then the data generation process is simple and fast, but the generated data is unrealistic and lacks statistical distributions and anomalies found in real data

Engineering Contradiction:
Improverealism of generated dataVSAvoidcomplexity of data generation process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent copies the statistical distribution characteristics and anomaly patterns from real data into the seed data generation process. By analyzing real data to extract distribution parameters and anomaly types, the system creates seed data that replicates the realistic properties of actual data without copying the data itself, thus achieving realism while maintaining data security and avoiding transformation costs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameters of data generation by incorporating statistical distribution parameters (mean, standard deviation, skewness, kurtosis) and anomaly parameters (frequency, types, severity) into the seed data creation process. This transforms simple uniform distribution generation into a sophisticated process that reproduces the complex statistical properties of real data, resolving the contradiction between simplicity and realism.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If real data is used as seed data, then the data is realistic and accurate, but it requires costly and time-consuming transformation to address proprietary and confidential information concerns

Engineering Contradiction:
Improveaccuracy of data distributionVSAvoidtime for data transformation
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts only the essential statistical characteristics and anomaly patterns from real data, separating these properties from the actual data values. By extracting distribution parameters (mean, standard deviation, etc.) and anomaly characteristics without using the original data values, the system creates seed data that maintains accuracy while eliminating the need for costly transformation and de-identification processes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates disposable seed data that replicates the statistical properties of real data without requiring the real data itself. This inexpensive copy can be freely transformed and adapted to different schemas without the legal and security constraints associated with using actual proprietary data, thus saving time and resources.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Productivity

If conventionally generated seed data is used for testing, then the testing process is fast and simple, but the testing effectiveness is reduced due to lack of anomalies and unrealistic data characteristics

Engineering Contradiction:
Improvespeed of testing processVSAvoideffectiveness of application testing
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary analysis of real data to identify and catalog anomaly types, frequencies, and patterns before generating seed data. This advance preparation allows the seed data generation process to automatically incorporate realistic anomalies without requiring manual intervention during testing, thus maintaining fast testing speeds while improving testing effectiveness through realistic data characteristics.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8805768B2Techniques for data generation
Publication Date: 2014.08.12 ORACLE INT CORP
  • US8805768B2 patent drawing
  • US8805768B2 patent drawing
  • US8805768B2 patent drawing

AI summary

Techniques, including systems and methods, for generating data are disclosed and suggested herein. Original data used in connection with one or more applications is analyzed in order to determine one or more distribution characteristics for the original data. The distribution characteristics are used to generate data that is similarly distributed. The generated data may be used as seed data for demonstrating, testing, or otherwise using one or more applications.