Neural Network Synthetic Data Generation for Sensitive Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating synthetic data for training AI systems are slow, error-prone, and fail to replicate the statistical characteristics of sensitive data, posing challenges for developers working with regulated data like customer financial records or patient healthcare data, while neural networks require large datasets for specific tasks and are inefficient in processing new log file types.

Innovation Solution

A platform that uses neural networks to automatically tokenize and generate synthetic data similar to sensitive datasets, employing classifiers to direct unstructured data to appropriate neural networks for parsing and training, even with insufficient initial training sets, and generating synthetic data to supplement training when necessary.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual desensitization methods are used, then data security is improved, but processing speed and accuracy deteriorate

Engineering Contradiction:
Improvedata securityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces manual mechanical desensitization processes with automated neural network-based systems. The neural networks automatically identify and desensitize sensitive data patterns, eliminating the need for manual expert intervention while maintaining or improving both security and processing speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service through automated neural networks that independently perform desensitization tasks. The networks learn from training data and automatically apply desensitization rules without requiring continuous human oversight, thereby improving processing efficiency while maintaining security standards.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual desensitization methods are used, then data security is improved, but error rate increases

Engineering Contradiction:
Improvedata securityVSAvoidaccuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent replaces error-prone manual desensitization with neural network-based automated systems. These networks consistently apply learned patterns for identifying sensitive data, eliminating human errors such as missed patterns or incorrect desensitization while maintaining security requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates feedback mechanisms where neural networks are trained on labeled data and continuously improve their desensitization accuracy. The feedback loop allows the system to learn from correct and incorrect examples, progressively reducing error rates while maintaining security standards.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If existing synthetic data generation methods are used, then data availability is improved, but data quality and statistical characteristics deteriorate

Engineering Contradiction:
Improvedata availabilityVSAvoidstatistical characteristics
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent replaces traditional synthetic data generation methods with neural network-based generation. These networks learn the statistical distributions and relationships in real data, then generate synthetic samples that preserve these characteristics, thereby improving both data availability and statistical fidelity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system uses parameter changes in the neural network training process to control the statistical characteristics of generated synthetic data. By adjusting training parameters and network architecture, the system can preserve specific statistical properties from real data while generating sufficient synthetic samples for training purposes.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If specialized neural networks are trained for each log file type, then parsing accuracy is improved, but system complexity and training data requirements increase

Engineering Contradiction:
Improveparsing accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal neural network architecture that can handle multiple log file types. Instead of training separate specialized networks for each log type, the system uses a single multi-functional network that learns to adapt to different formats, thereby reducing system complexity while maintaining parsing accuracy through transfer learning and fine-tuning capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220004878A1Systems and methods for synthetic document and data generation
Publication Date: 2022.01.06 CAPITAL ONE SERVICES LLC
  • US20220004878A1 patent drawing
  • US20220004878A1 patent drawing
  • US20220004878A1 patent drawing

AI summary

The present disclosure relates to systems and methods for determining synthetic information for documents. In one implementation, a system for determining synthetic information for documents may include at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: receiving a plurality of documents; determining a distribution of values for pixels of the documents; identifying at least one input field based on the determined distribution; extracting information from the at least one input field; calculate at least one statistic associated with the extracted information; generating a template having the at least one input field; and inserting data based on the calculated at least one statistic into the at least one input field of the template to generate a synthetic document.