Sensitive Data Rule Library for Scalable Industry Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sensitive data identification methods rely on customized approaches with limited scalability and accuracy, failing to meet industry standards and efficiently handle multiple data categories and large data volumes.

Innovation Solution

A universal, industry-level sensitive data identification method using text mining to create a sensitive data rule library from data security specification files, augmented with synonyms and context-aware techniques to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If customized approaches with expert-formulated sensitive word libraries and identification rules are used, then identification accuracy for specific enterprise needs is improved, but scalability and ability to handle multiple data categories deteriorate

Engineering Contradiction:
Improveidentification accuracyVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal sensitive data identification system that can handle multiple data categories and industries through a standardized rule library framework. The system uses industry-specific data security specification files to generate comprehensive sensitive data rules that cover various data types (personal information, financial data, health records, etc.), enabling one system to serve multiple purposes across different enterprises and sectors while maintaining high identification accuracy through structured rule formulations including sensitive words, regular expressions, and feature items.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If customized sensitive word libraries and identification rules are manually formulated by experts, then identification accuracy for specific needs is improved, but processing efficiency and ability to handle large data volumes deteriorate

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs preliminary actions by pre-generating comprehensive sensitive data rules from industry data security specification files before actual data identification tasks. The system extracts sensitive data rules, sensitive words, and feature items in advance and stores them in structured rule libraries. This preliminary rule generation and organization enables rapid processing of large data volumes during actual identification tasks while maintaining high accuracy through the pre-validated rule structures.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If comprehensive sensitive data rules covering multiple data categories are implemented, then identification coverage and compliance with industry standards are improved, but system complexity and difficulty of maintenance increase

Engineering Contradiction:
Improveidentification coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive sensitive data identification system into distinct modular components: industry-specific data security specification files, sensitive data rules with structured formulations, sensitive word libraries, feature item definitions, and rule library management modules. Each component handles specific aspects of identification (e.g., personal information, financial data, health records), allowing the system to achieve broad coverage while maintaining manageable complexity through clear separation of concerns and standardized interfaces between modules.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12585688B2Sensitive data identification method and apparatus, device, and computer storage medium
Publication Date: 2026.03.24 CHINA UNIONPAY
  • US12585688B2 patent drawing
  • US12585688B2 patent drawing
  • US12585688B2 patent drawing

AI summary

The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.