Synthetic Data Pattern Generation for Secure Data Protection Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data simulation methods for data protection systems face challenges in generating accurate simulation data that reflect real user data patterns due to the reluctance of users to share sensitive data and the complexity of data patterns, leading to inefficient stress testing and competitive analysis.
Innovation Solution
A method and system that uses a Generative Adversarial Network (GAN) to generate simulation data based on first data pattern information obtained from real data operations, allowing for the creation of simulation data that mimics the characteristics of real data patterns without requiring actual user data, thus optimizing performance and enabling competitive analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real user data is used for simulation, then data pattern accuracy is improved, but data sensitivity and security risks worsen
Solution Approach 1:
The patent creates a copy of the data pattern through simulation data generation. Instead of using actual user data, the system generates synthetic data that replicates the statistical characteristics and patterns of real data. This copy allows analysis and testing while avoiding the use of sensitive original data, thus resolving the contradiction between accuracy and security.
Solution Approach 2:
The patent introduces simulation data as an intermediary between the real user data and the analysis/testing process. This intermediary layer preserves the necessary pattern information for accurate simulation while eliminating direct exposure to sensitive user data, enabling secure data utilization.
2Object-affected harmful factors
If manual data pattern generation is used, then data sensitivity is reduced, but generation efficiency and accuracy worsen
Solution Approach 1:
The patent implements automated data pattern generation where the system learns from real data patterns and autonomously generates simulation data. The automated learning and generation process eliminates manual intervention, significantly improving efficiency while maintaining data sensitivity protection through synthetic data creation.
Solution Approach 2:
The patent replaces manual data generation processes with automated computational systems. Instead of manual creation, the system uses algorithms and machine learning models to automatically generate simulation data,大幅提高 generation efficiency while maintaining security.
3Ease of operation
If simple data simulation methods are used, then ease of operation is improved, but simulation accuracy and reliability worsen
Solution Approach 1:
The patent dynamically adjusts generation parameters based on learned data patterns. The system adapts parameters such as data distribution, complexity, and characteristics to match real data patterns, ensuring high simulation accuracy while maintaining ease of operation through automated parameter optimization.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system continuously refines simulation data based on pattern recognition and validation. This feedback loop ensures that generated data maintains high accuracy and reliability while keeping the user interface simple and easy to operate.
Data Source
AI summary
According to example embodiments of the present disclosure, a method, device and computer program product for data simulation are proposed. The method for data simulation includes: obtaining first data pattern information that is associated with a first set of operations executed on real data in a data protection system; generating, based on the first data pattern information, second data pattern information that is associated with a second set of operations executable by the data protection system; and generating, based on the second data pattern information, simulation data different from the real data, for the data protection system to execute the second set of operations on the simulation data. Thereby, the present solution can simulate efficiently and reliably a data pattern of real data, and thus generating simulation data of a data pattern similar to that of the real data.


