Sensitive Data Rule Library for Scalable Industry Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sensitive data identification methods rely on customized approaches with limited scalability and accuracy, failing to meet industry standards and efficiently handle multiple data categories and large data volumes.
Innovation Solution
A universal, industry-level sensitive data identification method using text mining to create a sensitive data rule library from data security specification files, augmented with synonyms and context-aware techniques to enhance recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If customized approaches with expert-formulated sensitive word libraries and identification rules are used, then identification accuracy for specific enterprise needs is improved, but scalability and ability to handle multiple data categories deteriorate
Solution Approach 1:
The patent creates a universal sensitive data identification system that can handle multiple data categories and industries through a standardized rule library framework. The system uses industry-specific data security specification files to generate comprehensive sensitive data rules that cover various data types (personal information, financial data, health records, etc.), enabling one system to serve multiple purposes across different enterprises and sectors while maintaining high identification accuracy through structured rule formulations including sensitive words, regular expressions, and feature items.
2Measurement precision
If customized sensitive word libraries and identification rules are manually formulated by experts, then identification accuracy for specific needs is improved, but processing efficiency and ability to handle large data volumes deteriorate
Solution Approach 1:
The patent performs preliminary actions by pre-generating comprehensive sensitive data rules from industry data security specification files before actual data identification tasks. The system extracts sensitive data rules, sensitive words, and feature items in advance and stores them in structured rule libraries. This preliminary rule generation and organization enables rapid processing of large data volumes during actual identification tasks while maintaining high accuracy through the pre-validated rule structures.
3Adaptability or versatility
If comprehensive sensitive data rules covering multiple data categories are implemented, then identification coverage and compliance with industry standards are improved, but system complexity and difficulty of maintenance increase
Solution Approach 1:
The patent segments the comprehensive sensitive data identification system into distinct modular components: industry-specific data security specification files, sensitive data rules with structured formulations, sensitive word libraries, feature item definitions, and rule library management modules. Each component handles specific aspects of identification (e.g., personal information, financial data, health records), allowing the system to achieve broad coverage while maintaining manageable complexity through clear separation of concerns and standardized interfaces between modules.
Data Source
AI summary
The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.


