Synthetic Sample Generation for Secure Data Detection Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in constructing a sample data set for specific data detection models efficiently and cost-effectively, as obtaining original specific data is difficult due to security concerns and manual labeling is costly and inefficient, especially for unstructured data where context semantic analysis is lacking.

Innovation Solution

A method for sample generation involves acquiring a reference data set, determining a target template based on information sample values, and generating sample information and data sets using this template, which characterizes the information structure and reduces the need for manual labeling and data leakage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If original specific data is obtained for training detection models, then model accuracy is improved, but data security risks and acquisition costs increase

Engineering Contradiction:
Improvedetection model accuracyVSAvoiddata security risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic sample data that copies the structural and semantic characteristics of original specific data without using the actual sensitive information. Templates are generated based on the format and context patterns of target information types (like phone numbers, IDs), allowing model training with realistic-looking data that doesn't compromise security.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent generates inexpensive synthetic samples that can be freely created and discarded for training purposes. These artificial samples serve as disposable training data that replace expensive and risky original data, enabling multiple training iterations without data security concerns.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Manufacturing precision

If manual labeling is performed for sample data construction, then data quality is improved, but time consumption and costs increase

Engineering Contradiction:
Improvesample data qualityVSAvoidlabeling time consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs automatic template generation and sample creation without human intervention. The algorithm independently analyzes data patterns, generates appropriate templates for different information types, and creates labeled training samples autonomously, eliminating the need for manual labeling while maintaining high data quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the data construction process from manual parameter setting to automated parameter generation. By changing the approach from human experts manually defining labels to algorithms automatically generating templates and samples based on data patterns, the system achieves both speed and quality improvements.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If context semantic analysis is performed on unstructured data, then detection accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs semantic analysis and template generation in advance during the sample creation phase. By pre-processing and structuring the data with appropriate templates before model training, the system reduces the computational burden during actual detection operations, as the complex semantic understanding work is already completed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides unstructured data into structured templates based on different information types (phone numbers, IDs, addresses). This segmentation transforms complex unstructured data processing into manageable structured components, making semantic analysis more efficient and computationally feasible.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240394478A1Method for sample generation, computer device, and storage medium
Publication Date: 2024.11.28 BEIJING VOLCANO ENGINE TECH CO LTD
  • US20240394478A1 patent drawing
  • US20240394478A1 patent drawing
  • US20240394478A1 patent drawing

AI summary

A method and apparatus for sample generation, a method and apparatus for information detection, a computer device, and a storage medium are provided, and the method for sample generation includes: acquiring a reference data set, where the reference data set includes first reference data and second reference data, the first reference data includes at least one information sample value, the information sample value belongs to a target information type, and the second reference data does not include an information sample value of the target information type; determining a target template corresponding to the first reference data based on the at least one information sample value in the first reference data; generating a plurality of sample information based on the target template corresponding to the first reference data; and generating a sample data set based on the plurality of sample information and the second reference data.