Sample Data Generation via Semantic and Grammatical Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in accurately and efficiently detecting specific data with security requirements from a large dataset, where traditional methods lack context semantic analysis, leading to inaccurate detection and high labeling costs due to the difficulty in acquiring and labeling specific data for training models.

Innovation Solution

A method and apparatus for generating sample data through semantic, lexical, and grammatical structure analysis of reference data to create positive and negative samples, which are then used to construct a sample dataset, reducing data leakage risks and labeling costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data detection methods are used, then the detection process is simple, but the detection accuracy is low and labeling costs are high

Engineering Contradiction:
Improvedetection accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data analysis process into three distinct components: semantic analysis (meaning extraction), lexical structure analysis (pattern recognition), and grammatical structure analysis (syntax validation). This segmentation allows each component to specialize in specific aspects of data detection, improving overall accuracy while making the complex process more manageable through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary analysis processing on reference data to extract features, keywords, and structural patterns before actual detection occurs. By pre-processing and storing these analytical results, the system can quickly generate accurate detection models without repeating the entire analysis process during actual data detection, thus improving accuracy while reducing real-time processing complexity

Inventive Principle:
Principle #10Preliminary action

2Productivity

If manual labeling of specific data is performed, then the labeling accuracy is high, but the labeling cost and time consumption are high

Engineering Contradiction:
Improvelabeling efficiencyVSAvoidlabeling accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements self-service labeling by automatically generating labels through the three-fold analysis process. The system analyzes reference data, extracts semantic meanings, lexical patterns, and grammatical structures, then uses these to automatically label new data without human intervention. This maintains high labeling accuracy through systematic analysis while dramatically improving productivity by eliminating manual labeling operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual labeling process with an automated computational system that performs semantic, lexical, and grammatical analysis to generate labels automatically. This substitution of manual mechanical operations with automated analytical processes achieves both high efficiency and maintained accuracy through multi-dimensional analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If specific data with security requirements is acquired for training, then the training data quality is high, but the data leakage risk and acquisition difficulty are high

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata leakage risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates copies of the analysis process and results from reference data to generate training samples. By copying the analytical framework and applying it to generate synthetic training data, the system maintains high training data quality while avoiding direct handling of sensitive original data, thus reducing data leakage risks

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces analysis results (semantic meanings, lexical patterns, grammatical structures) as intermediary elements between the reference data and training data generation. These intermediaries serve as a buffer that allows quality training data to be generated without directly exposing sensitive information from the original reference data, thereby reducing data leakage risks

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240394475A1Generation method for sample data, device and storage medium
Publication Date: 2024.11.28 BEIJING VOLCANO ENGINE TECH CO LTD
  • US20240394475A1 patent drawing
  • US20240394475A1 patent drawing
  • US20240394475A1 patent drawing

AI summary

A generation method and apparatus for sample data, a method and apparatus for information detection, a device and a storage medium are provided, and the generation method includes: acquiring first reference data, the first reference data including target information matching a target information type, and the target information type being a preset information type with a security requirement; performing analysis processing on the target information in the first reference data to generate an analysis result corresponding to the target information, the analysis processing including semantic analysis, lexical structure analysis and grammatical structure analysis; generating a plurality of positive sample information and a plurality of negative sample information based on the analysis result corresponding to the target information; generating a sample data set including positive sample data and negative sample data based on the plurality of positive sample information, the plurality of negative sample information, and second reference data.