Synthetic Data Generation for Database Schema Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Database planning often inaccurately predicts future growth due to misinterpretation of sample data characteristics, leading to biased database and index designs.

Innovation Solution

Generating synthetic data based on stored numeric distribution information within a schema, which is compared to actual data to ensure statistical similarity and adjust the schema accordingly, improving database performance and fraud detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sample data is used to plan database growth, then database design can be created, but the design may be biased by individual characteristics and outliers in the sample data

Engineering Contradiction:
Improveaccuracy of database growth predictionVSAvoidstatistical representativeness of sample data
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent generates synthetic data that copies the statistical properties and distribution patterns of actual data without replicating individual records. This creates representative samples that eliminate bias from specific outliers while maintaining the underlying data characteristics needed for accurate database planning

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between actual data and database planning. This intermediary layer filters out individual anomalies and outliers while preserving the statistical patterns, providing a cleaner basis for growth predictions without losing essential data characteristics

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If synthetic data is generated without distribution information, then data can be created quickly, but the synthetic data lacks statistical accuracy

Engineering Contradiction:
Improvespeed of synthetic data generationVSAvoidstatistical accuracy of synthetic data
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary analysis to extract distribution information from actual data before generating synthetic data. By pre-calculating statistical parameters and distribution patterns, the system prepares the foundation for accurate synthetic data generation, ensuring both speed and statistical fidelity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms distribution information into specific parameters that guide synthetic data generation. By converting statistical patterns into actionable parameters, the system maintains statistical accuracy while enabling efficient generation through parameterized processes rather than complex computations

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If schema is repeatedly modified to match actual data, then synthetic data becomes statistically similar, but the process requires multiple iterations

Engineering Contradiction:
Improvestatistical similarity between synthetic and actual dataVSAvoidnumber of iterations for schema adjustment
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements a feedback loop where synthetic data is continuously compared against actual data, and schema parameters are adjusted based on the comparison results. This iterative feedback process systematically reduces statistical differences until the synthetic data adequately represents the actual data distribution

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual, trial-and-error schema adjustment with an automated statistical comparison and adjustment system. By using computational methods to calculate optimal parameter values based on distribution matching, the system reduces the number of iterations needed compared to manual adjustment processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10719490B1Forensic analysis using synthetic datasets
Publication Date: 2020.07.21 CAPITAL ONE SERVICES LLC
  • US10719490B1 patent drawing
  • US10719490B1 patent drawing
  • US10719490B1 patent drawing

AI summary

A system, method, and computer-readable medium for generating synthetic data are described. Improved data models for databases may be achieved by improving the quality of synthetic data upon for modeling those databases and for checking the authenticity of existing numerical data. According to some aspects, these and other benefits may be achieved by using numeric distribution information in a schema describing one or more numeric fields and, based on that schema, distribution-appropriate numerical data may be generated. Also, another schema may be used to generate a second set of numerical data having a different distribution that is not expected for the one or more numeric fields. Actual data may be compared against the generated datasets. When the actual data is determined to be statistically similar to the second numerical dataset, an alert may be generated. A benefit includes finding potentially fraudulent datasets using an efficient approach.