Synthetic Data Generation for Database Schema Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Database planning often inaccurately predicts future growth due to misinterpretation of sample data characteristics, leading to biased database and index designs.
Innovation Solution
Generating synthetic data based on stored numeric distribution information within a schema, which is compared to actual data to ensure statistical similarity and adjust the schema accordingly, improving database performance and fraud detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sample data is used to plan database growth, then database design can be created, but the design may be biased by individual characteristics and outliers in the sample data
Solution Approach 1:
The patent generates synthetic data that copies the statistical properties and distribution patterns of actual data without replicating individual records. This creates representative samples that eliminate bias from specific outliers while maintaining the underlying data characteristics needed for accurate database planning
Solution Approach 2:
The patent introduces synthetic data as an intermediary between actual data and database planning. This intermediary layer filters out individual anomalies and outliers while preserving the statistical patterns, providing a cleaner basis for growth predictions without losing essential data characteristics
2Productivity
If synthetic data is generated without distribution information, then data can be created quickly, but the synthetic data lacks statistical accuracy
Solution Approach 1:
The patent performs preliminary analysis to extract distribution information from actual data before generating synthetic data. By pre-calculating statistical parameters and distribution patterns, the system prepares the foundation for accurate synthetic data generation, ensuring both speed and statistical fidelity
Solution Approach 2:
The patent transforms distribution information into specific parameters that guide synthetic data generation. By converting statistical patterns into actionable parameters, the system maintains statistical accuracy while enabling efficient generation through parameterized processes rather than complex computations
3Measurement precision
If schema is repeatedly modified to match actual data, then synthetic data becomes statistically similar, but the process requires multiple iterations
Solution Approach 1:
The patent implements a feedback loop where synthetic data is continuously compared against actual data, and schema parameters are adjusted based on the comparison results. This iterative feedback process systematically reduces statistical differences until the synthetic data adequately represents the actual data distribution
Solution Approach 2:
The patent replaces manual, trial-and-error schema adjustment with an automated statistical comparison and adjustment system. By using computational methods to calculate optimal parameter values based on distribution matching, the system reduces the number of iterations needed compared to manual adjustment processes
Data Source
AI summary
A system, method, and computer-readable medium for generating synthetic data are described. Improved data models for databases may be achieved by improving the quality of synthetic data upon for modeling those databases and for checking the authenticity of existing numerical data. According to some aspects, these and other benefits may be achieved by using numeric distribution information in a schema describing one or more numeric fields and, based on that schema, distribution-appropriate numerical data may be generated. Also, another schema may be used to generate a second set of numerical data having a different distribution that is not expected for the one or more numeric fields. Actual data may be compared against the generated datasets. When the actual data is determined to be statistically similar to the second numerical dataset, an alert may be generated. A benefit includes finding potentially fraudulent datasets using an efficient approach.


