Synthetic Sample Generation for Secure Data Detection Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in constructing a sample data set for specific data detection models efficiently and cost-effectively, as obtaining original specific data is difficult due to security concerns and manual labeling is costly and inefficient, especially for unstructured data where context semantic analysis is lacking.
Innovation Solution
A method for sample generation involves acquiring a reference data set, determining a target template based on information sample values, and generating sample information and data sets using this template, which characterizes the information structure and reduces the need for manual labeling and data leakage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If original specific data is obtained for training detection models, then model accuracy is improved, but data security risks and acquisition costs increase
Solution Approach 1:
The patent creates synthetic sample data that copies the structural and semantic characteristics of original specific data without using the actual sensitive information. Templates are generated based on the format and context patterns of target information types (like phone numbers, IDs), allowing model training with realistic-looking data that doesn't compromise security.
Solution Approach 2:
The patent generates inexpensive synthetic samples that can be freely created and discarded for training purposes. These artificial samples serve as disposable training data that replace expensive and risky original data, enabling multiple training iterations without data security concerns.
2Manufacturing precision
If manual labeling is performed for sample data construction, then data quality is improved, but time consumption and costs increase
Solution Approach 1:
The system performs automatic template generation and sample creation without human intervention. The algorithm independently analyzes data patterns, generates appropriate templates for different information types, and creates labeled training samples autonomously, eliminating the need for manual labeling while maintaining high data quality.
Solution Approach 2:
The patent transforms the data construction process from manual parameter setting to automated parameter generation. By changing the approach from human experts manually defining labels to algorithms automatically generating templates and samples based on data patterns, the system achieves both speed and quality improvements.
3Measurement precision
If context semantic analysis is performed on unstructured data, then detection accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent performs semantic analysis and template generation in advance during the sample creation phase. By pre-processing and structuring the data with appropriate templates before model training, the system reduces the computational burden during actual detection operations, as the complex semantic understanding work is already completed.
Solution Approach 2:
The patent divides unstructured data into structured templates based on different information types (phone numbers, IDs, addresses). This segmentation transforms complex unstructured data processing into manageable structured components, making semantic analysis more efficient and computationally feasible.
Data Source
AI summary
A method and apparatus for sample generation, a method and apparatus for information detection, a computer device, and a storage medium are provided, and the method for sample generation includes: acquiring a reference data set, where the reference data set includes first reference data and second reference data, the first reference data includes at least one information sample value, the information sample value belongs to a target information type, and the second reference data does not include an information sample value of the target information type; determining a target template corresponding to the first reference data based on the at least one information sample value in the first reference data; generating a plurality of sample information based on the target template corresponding to the first reference data; and generating a sample data set based on the plurality of sample information and the second reference data.


