Verified Virtual Tabular Data Generation for Private AI Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Financial institutions face challenges in generating accurate training data for deep-learning models due to the need for sensitive user information, which cannot be used directly, leading to low relevance and performance issues in AI and machine learning applications.

Innovation Solution

A method for generating virtual tabular data using a deep-learning module, involving multiple prompts and verification processes to ensure accuracy, diversity, and adherence to user-defined schemas, replacing sensitive information with highly accurate data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If public data is used to replace sensitive user information, then data availability is improved, but data relevance and accuracy deteriorate

Engineering Contradiction:
Improvedata availabilityVSAvoiddata relevance
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent generates virtual tabular data that copies the statistical characteristics, data types, and structural patterns of sensitive user information without replicating actual personal data. This allows充足的数据供应 for training deep-learning models while maintaining data privacy and improving relevance compared to generic public data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms the parameters of data generation by using deep-learning models to synthesize virtual data with controlled statistical properties, data type distributions, and relational patterns that match the characteristics of sensitive user information, thereby improving data relevance while maintaining availability

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple verification operations are performed on generated data, then data accuracy is improved, but processing time increases

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs verification operations in a preliminary manner during the data generation process itself, checking data accuracy, schema compliance, and statistical characteristics before finalizing the virtual tabular data. This preliminary verification approach ensures high data accuracy while minimizing additional processing time by integrating checks into the generation workflow rather than adding separate post-processing steps

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual data collection processes are used, then data accuracy is improved, but productivity deteriorates

Engineering Contradiction:
Improvedata accuracyVSAvoiddata generation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical data collection processes with automated deep-learning models that generate virtual tabular data programmatically. This substitution maintains data accuracy by preserving statistical characteristics and structural patterns while dramatically improving productivity through automated generation without manual intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data generation where the deep-learning model automatically creates virtual tabular data with appropriate statistical properties, data types, and relational patterns without requiring manual data collection or curation. This self-service approach maintains accuracy while significantly boosting productivity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250284739A1Virtual tabular data generation method and server performing the same
Publication Date: 2025.09.11 KAKAOBANK CORP
  • US20250284739A1 patent drawing
  • US20250284739A1 patent drawing
  • US20250284739A1 patent drawing

AI summary

A method of generating virtual tabular data, performed on a server using a deep-learning module, comprising: generating a first prompt for generating a table schema, calibrating a table schema by comparing the table schema generated based on the first prompt with a predefined reference table schema, generating a second prompt by referring to the calibrated table schema, generating table condition data for first tabular data generated based on the second prompt, generating a third prompt by referring to the table condition data and the calibrated table schema, and deriving final tabular data through a verification operation on second tabular data generated based on the third prompt.