Data security intelligent treatment method based on data identification
Through intelligent governance methods based on data identification and combined with automated security analysis and control strategies, the problems of insufficient fine-grainedness, cumbersome rules and insufficient systematization in the existing data security governance methods are solved, and efficient, intelligent and systematic data security governance is achieved, improving protection efficiency and analysis effect.
Patent Information
- Application Number
- CN202510113784.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
The existing data security governance methods have problems such as insufficient granularity of data objects, cumbersome rules, lack of adaptability, and insufficient systematization, which is difficult to support rich and meticulous data security management needs.
Through intelligent governance methods based on data identification, multi-dimensional data identification combined with automated security analysis and control strategies can achieve efficient, intelligent and systematic data security governance. Specific steps include identifying and identifying data sources, configuring data categories and levels, acquiring intelligent classification models, adding and managing data security protection equipment, configuring security protection policies, collecting logs and identification information, performing abnormal detection and fusion correlation analysis.
It realizes accurate identification and tracking of data objects, improves protection efficiency and analysis effects, reduces manual losses, and enhances the adaptability and systematization of data security governance.
Smart Images

Figure CN120012158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a data security intelligent management method based on data identification. Background Art
[0002] The essence of data security governance is to establish a systematic data security mechanism, clarify the goals, content and scope of data security, ensure the full utilization of existing achievements, and make up for missing capabilities in a planned manner, thereby forming a closed loop of data security capabilities. The current data security governance methods have the following main shortcomings:
[0003] (1) Insufficient granularity of data objects: The analysis of data objects is not detailed enough. Most of the data is identified based on metadata information, such as file name and ID, which makes it difficult to support rich and detailed data security management and control needs.
[0004] (2) Complicated rules and lack of adaptability: Many data security measures still rely on a large amount of manual configuration, which is not only inefficient but also prone to errors. The means and strategies of data security governance cannot adapt to new security threats and environmental changes and need to be manually updated and adjusted.
[0005] (3) Lack of systematization in data security governance: Current data security governance methods mostly involve piling up various types of security protection equipment, each of which reflects data security protection capabilities. The capabilities or effects of different devices are difficult to integrate and correlate. Summary of the invention
[0006] In view of this, the present invention provides a data security intelligent governance method based on data identification, which realizes efficient, intelligent and system-linked data security governance based on rich and multi-dimensional data identification and combined with automated security analysis and control strategies.
[0007] The present invention discloses a data security intelligent management method based on data identification, which includes:
[0008] Identify and label the data source; configure the data source information and authorize access to the data source;
[0009] Configure the categories and levels of data, obtain classification and grading retrieval rules for data content, conduct learning and training based on labeled samples in advance, and use the trained intelligent classification model to classify and predict data content; identify structured data in various databases and extract identification information; identify various unstructured data in various file servers and extract identification information;
[0010] Add and manage various data security protection devices, configure and issue security protection strategies based on data identification; the security protection strategies include anti-leakage strategies, desensitization strategies and security audit strategies;
[0011] Collect logs and identification information, obtain abnormal records or risk warnings through security audit strategies; obtain abnormal data access behavior in logs, record abnormal feature information, and generate abnormal detection results; perform fusion correlation analysis on the obtained abnormal records or risk warnings and abnormal detection results.
[0012] Furthermore, the method for obtaining the intelligent classification model includes:
[0013] Provide the pre-prepared labeled sample data to the device, which performs training based on the machine learning algorithm, extracts the common points of the sample data and generates an intelligent prediction model;
[0014] If it is impossible to clearly define the classification labels of the data and provide sample data classified by labels, clustering technology is used to utilize unsupervised machine learning to automatically cluster a large number of mixed files according to content similarity. After clustering is completed, the user confirms the labels for each type of data. If the clustering results do not meet the requirements, the parameters and configurations are adjusted to complete the data clustering again. After the clustering labels are confirmed, all labeled sample data are provided to the device so that the device can be trained and learned based on the machine learning algorithm.
[0015] Furthermore, the identifying of structured data in various databases and extracting identification information includes:
[0016] Conduct deep scans on various databases to identify the structured data stored therein, and use data fields as the smallest unit to extract various identification information including metadata, data HASH, fuzzy HASH, and classification and grading attributes.
[0017] Furthermore, the identifying of various unstructured data in various file servers and extracting identification information includes:
[0018] Perform deep scans on various file servers to identify various unstructured data stored in them. Taking a single file as the smallest unit, automatically extract various identification information including metadata, overall file HASH, file content fuzzy HASH, and classification and grading attributes; metadata is used to describe the basic information of the data, overall file HASH is used to identify the uniqueness of the data, file content fuzzy HASH is used to analyze the similarity of the data, and classification and grading attributes are used to determine the security factors of the data.
[0019] Furthermore, the adding and managing of various data security protection devices, and the configuration and issuance of security protection strategies based on data identification include:
[0020] Add various types of data security protection devices and configure them, manage the connection status of data security protection devices, and monitor the operating status of data security protection devices; formulate various security protection strategies based on data identification, and establish a sensitive data identification library; issue the sensitive data identification library and protection strategies to data security protection devices; the sensitive data identification library includes identifications in data identifications whose security levels are above the preset level.
[0021] Furthermore, the collection of logs and identification information, and obtaining abnormal records or risk warnings through security audit policies, include:
[0022] Collect data source access logs through traffic, collect third-party system logs through interfaces, and collect data identification information through interfaces; parse the collected logs into structured data based on field mapping; according to the stage in which the logs are generated, database access logs correspond to data query operations, and third-party system logs correspond to system type-associated operations. Each log traverses the policies under the operation behavior in the security audit policy list, and generates abnormal records or risk alerts based on the policy hits.
[0023] Furthermore, the obtaining of abnormal data access behavior in the log, recording abnormal feature information, and generating anomaly detection results include:
[0024] Traffic collection and preprocessing: Provide data access traffic log sample data, use unsupervised learning algorithms to train data access behavior clustering models, and learn the behavior patterns of different user groups;
[0025] Feature extraction: Select the required feature subset from the original features, extract features from the traffic data, and aggregate each feature after grouping and slicing;
[0026] Clustering: Automatically identify normal and abnormal patterns in traffic groups, and identify abnormal and normal group patterns through clustering;
[0027] Building a baseline: After obtaining the behavior labels of traffic logs through clustering, a baseline model is built to predict future traffic. The baseline model includes a statistical baseline and a memory baseline.
[0028] Real-time update and dynamic response: The memory baseline is trained incrementally according to fine-grained time intervals and directly called through the API;
[0029] Using the baseline model, based on the normal behavior feature range, feature extraction and baseline comparison are performed on the collected data access logs. Behaviors that exceed the feature range are defined as abnormal data access behaviors, abnormal feature information is recorded, and anomaly detection results are generated.
[0030] Furthermore, the obtained abnormal records or risk warnings and abnormal detection results are subjected to fusion correlation analysis, including:
[0031] Various data security devices perform data security monitoring and execution according to the data protection strategy based on data identification, and report the alarm and protection disposal data for unified summary;
[0032] Through the integration and correlation analysis of multiple data security logs within a specified time, the relevant data operation behaviors are linked into security events, and the occurrence chain of the events is restored; the events include data leakage and data abuse;
[0033] Based on the results of the fusion analysis, instructions are issued to various data security devices to achieve security management and control of the entire data life cycle.
[0034] Due to the adoption of the above technical solution, the present invention has the following advantages:
[0035] 1. This invention proposes a multi-dimensional identification method for data in view of the fact that traditional data security management methods cannot accurately describe data information, track data changes, and various security protection capabilities and effects have not formed an effective linkage, making it difficult to conduct fusion analysis. Based on data identification, various protection strategies and risk warning information are integrated, which can more comprehensively identify data objects, accurately track data changes, and comprehensively improve protection efficiency and analysis effects.
[0036] 2. The present invention aims to solve the problems in traditional data security governance methods, such as the need for frequent rule updates, low efficiency, and proneness to errors in achieving data classification and grading as well as data security auditing. This invention proposes a method based on labeled sample data and behavioral baseline data, combined with artificial intelligence machine learning algorithms, which can assist or replace traditional rule-based detection methods, reduce labor losses, and improve the accuracy of classification and anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0038] Figure 1 A schematic diagram of a data classification process according to an embodiment of the present invention;
[0039] Figure 2 This is a schematic diagram of multi-dimensional identification of structured data according to an embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of multi-dimensional identification of unstructured data according to an embodiment of the present invention;
[0041] Figure 4 A schematic diagram of the process of security device management and policy issuance according to an embodiment of the present invention;
[0042] Figure 5 A schematic diagram of a data identification-based security audit and fusion analysis process according to an embodiment of the present invention;
[0043] Figure 6 Schematic diagram of the data access behavior baseline calculation principle according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present invention is further described in conjunction with the accompanying drawings and embodiments, and the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present invention.
[0045] In order to solve the following technical problems: the identification and analysis of data security governance objects are inaccurate and incomplete, which makes it difficult to support the complex and changeable data environment and the needs of fine-grained control; the frequent update of rules, manual design of features, and the need to improve detection accuracy; the capabilities and effects of various security protection equipment cannot be effectively linked and integrated.
[0046] See also Figure 1 The present invention provides an embodiment of a data security intelligent management method based on data identification, which includes:
[0047] 1. Multi-dimensional identification of data security:
[0048] The first step of the intelligent data security governance method based on data identification is to accurately identify the object of data security governance, that is, the data assets themselves, and to enrich and multi-dimensionally identify them, so as to support the subsequent construction of systematic data security governance capabilities around data identification. The key steps of multi-dimensional data identification are as follows.
[0049] (1) Data source discovery and management:
[0050] 1) Data source scanning discovery
[0051] Based on scanning and analysis technology of different IP segments, different ports, different file transfer protocols (SM, TFTP, FTP, etc.), and database protocols (Oracle, MySql, DAMO, Mongodb, Hbase, etc.), it automatically identifies and discovers the data sources within the enterprise, and labels the discovered data sources to help enterprises understand the distribution of internal data resources.
[0052] 2) Data source information management
[0053] Clarify the various types of databases and file server resources and their ownership within the enterprise. By configuring the data source access account / password, department / owner and other information, authorize access to the data source to facilitate subsequent processing, analysis and storage of various data assets within it, and support the orderly development of data security governance.
[0054] (2) Classification and grading strategy configuration and model training
[0055] See also Figure 1 , the data classification process includes:
[0056] 1) Configuration of data types and data levels
[0057] First, configure the categories and levels of enterprise data (including library tables and file data), the corresponding relationships between categories and levels, and other information to associate with subsequent specific classification and grading implementation plans.
[0058] 2) Content-based retrieval matching rule configuration
[0059] Based on relevant requirements such as industry standards, corporate management specifications, business background, etc., we have sorted out classification and grading retrieval rules for data content, including but not limited to dictionaries, keywords, regular expressions, the frequency of occurrence of specific content, mutual inclusion or exclusion relationships, etc.
[0060] 3) Intelligent classification model training based on machine learning
[0061] Intelligent classification and grading based on machine learning first requires supervised learning training based on labeled samples in advance, and then using the trained model to automatically classify and predict the data content.
[0062] ① Intelligent machine learning
[0063] Prepare a certain amount of labeled sample data in advance and provide it to the device. The device performs training and learning based on the machine learning algorithm, extracts the common points of these sample data, and generates an intelligent prediction model.
[0064] ② Intelligent sample clustering
[0065] If the enterprise is unable to clearly define the data classification labels at the beginning and provide sample data classified by labels, it can use intelligent clustering technology and unsupervised machine learning to automatically cluster large numbers of mixed files according to content similarity, thereby greatly reducing the workload of security analysts and users in business sorting.
[0066] ③Sample clustering confirmation
[0067] After completing intelligent clustering, the user needs to cooperate in confirming the labels of each type of data. If you are not satisfied with the accuracy and granularity of the clustering results, you can adjust the relevant parameters and configurations to further complete the data clustering. After completing the clustering label confirmation, all labeled sample data are provided to the device to implement step ①.
[0068] (3) Multidimensional Data Identification Extraction
[0069] 1) Structured data identification
[0070] See also Figure 2 , conducts in-depth scans on various databases to identify the structured data stored in them. Taking the data field as the smallest unit, it automatically extracts various identification information including metadata, overall data HASH, fuzzy HASH of data content, and classification and grading attributes. The classification and grading attributes are the results obtained after sampling the database content according to the classification and grading strategy, performing content-based retrieval rule matching and artificial intelligence model-based prediction.
[0071] 2) Unstructured data identification
[0072] See also Figure 3 , conduct in-depth scans on various file servers to identify various unstructured data stored in them. Taking a single file as the smallest unit, automatically extract various identification information including metadata, file overall HASH, file content fuzzy HASH, and classification and grading attributes. Metadata is used to describe the basic information of the data, the file overall HASH is used to identify the uniqueness of the data, the file content fuzzy HASH is used to analyze the similarity of the data, and the classification and grading attributes are used to determine the security factors of the data. Among them, the classification and grading attributes are the results obtained after performing content-based retrieval rule matching and artificial intelligence model-based prediction on the file data according to the classification and grading strategy.
[0073] 2. Building security protection capabilities based on data identification:
[0074] The second step of the intelligent data security governance method based on data identification is to connect and manage various data security protection devices, configure and issue security protection strategies based on data identification, including anti-leakage strategies, desensitization strategies and security audit strategies; Figure 4 shown.
[0075] (1) Safety protection equipment management:
[0076] 1) Add various data security protection systems / equipment such as database auditing, data leakage prevention, data desensitization, database firewall, API monitoring, etc., and configure various information such as IP, type, identity authentication, interface parameter information, etc.
[0077] 2) Manage the connection status of the device and monitor the operating status of the device.
[0078] (2) Issuance of data protection policies based on data identification:
[0079] 1) Through the combined definition of multi-dimensional data identification, data path and other rules, formulate various security protection strategies based on data identification and establish a sensitive data identification library.
[0080] 2) Issue sensitive data identification library and protection strategies to data security protection systems / equipment such as database audit, data leakage prevention, API monitoring, database firewall, and data desensitization, and build data security protection capabilities with data identification as the core and multi-device linkage. The sensitive data identification library includes identifications with security levels above the preset level in the data identification.
[0081] 3. Data security audit and fusion analysis based on data identification:
[0082] See also Figure 5 The third step of the intelligent data security governance method based on data identification is to collect various logs of application systems and data sources, conduct fusion correlation analysis based on security audit strategies, intelligent baselines and correlation analysis models, and combine data identification libraries to timely discover and warn of risk events and quickly deal with them.
[0083] (1) Data security audit based on policy rules:
[0084] 1) Audit rule configuration
[0085] Customize various data security audit rules to define security risk warning rules. The general data security audit rules are as follows:
[0086] ① Data asset audit strategy. Data identification audit strategies such as non-compliant built-in security level identification format, inaccurate security level identification, non-compliant system classification, inaccurate manual classification, and non-compliant document data identification; data storage audit strategies such as high-sensitivity data with low storage, mismatch between personnel and data security level, and mismatch between personnel roles and data classification information. Used to audit and analyze data assets and discover risks in data identification and storage;
[0087] ② Database access log audit strategy. For example, abnormal SQL detection strategy. Used to analyze database access logs and discover database access risks;
[0088] ③Security audit strategy for the entire life cycle of data collection, transmission, storage, processing, exchange, and destruction. Customize the blacklist and whitelist strategies for different operation behaviors to analyze the operation behavior logs of the entire life cycle of data and discover abnormal behaviors and illegal operations.
[0089] 2) Audit rules implementation
[0090] ① Data collection: collect data source access logs through traffic, collect third-party system logs through interfaces, and collect data identification information through interfaces;
[0091] ②Log parsing: Parse the collected logs into structured data based on field mapping;
[0092] ③ According to the stage of log generation, the database access log corresponds to the data query operation behavior, and the third-party system log corresponds to the operation behavior associated with the system type (for example, the anti-leakage system corresponds to the data exchange behavior). Each log traverses the policy under the operation behavior in the security audit policy list, and generates abnormal records or risk alarms based on the policy hit situation.
[0093] (2) Security audit based on intelligent baseline of data access behavior
[0094] The calculation principle of data access behavior baseline is as follows Figure 6 shown.
[0095] 1) Traffic collection and preprocessing: Provide a certain amount of data access traffic log sample data, use unsupervised learning algorithms to train data access behavior clustering models, and automatically learn the behavior patterns of different user groups.
[0096] 2) Feature extraction: Select the most useful feature subset from the original features, extract meaningful features from the traffic data, and aggregate each feature after grouping and slicing.
[0097] 3) Clustering: Automatically identify normal and abnormal patterns in traffic groups. These group patterns reflect the natural structure in traffic data and can reveal the inherent laws and distribution of data. Through clustering, abnormal and normal group patterns can be automatically identified without knowing the distribution and structure of the data in advance, and this process is not interfered by human factors.
[0098] 4) Build a baseline: Building an accurate and effective baseline model is the core of analyzing abnormal database behavior. After obtaining the behavior labels of traffic logs through clustering, a baseline model can be built to predict future traffic, including statistical baseline and memory baseline modes.
[0099] 5) Real-time update and dynamic response: The memory baseline is trained incrementally according to fine-grained time intervals to update model parameters. The timing strategy is configurable, the whole process is automated, the service is independent, and can be called directly through the API.
[0100] 6) Using the data behavior intelligent baseline model, based on the normal behavior feature range, perform feature extraction and baseline comparison on the collected data access logs, define the behavior beyond the feature range as abnormal data access behavior, record the abnormal feature information, including feature name, feature value and the normal value range of the feature, and generate anomaly detection results.
[0101] (3) Fusion analysis based on association analysis model
[0102] 1) Various data security systems / equipment perform data security monitoring and execution according to the data protection strategy based on data identification, and report the alarm and protection disposal data for unified summary.
[0103] 2) Through the fusion and correlation analysis of multiple data security logs over a period of time, the relevant data operation behaviors are linked together into security events, and the occurrence links of events such as data leakage and data abuse are restored.
[0104] 3) Based on the results of the integrated analysis, blocking, releasing and other linkage disposal instructions are issued to various data security devices / systems to achieve security management and control of the entire life cycle of data.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A data security intelligent management method based on data identification, characterized in that: include: Identify data sources, configure data source information, and authorize access to data sources; Configure the categories and levels of data, obtain classification and grading retrieval rules for data content, conduct learning and training based on labeled samples in advance, and use the trained intelligent classification model to classify and predict data content; Identify structured data in various databases and extract identification information; identify various unstructured data in various file servers and extract identification information; Add and manage various data security protection devices, configure and issue security protection strategies based on data identification, including anti-leakage strategies, desensitization strategies and security audit strategies; Collect identification information of logs and data, obtain abnormal records or risk warnings through security audit strategies; obtain abnormal data access behavior in logs, record abnormal feature information, and generate abnormal detection results; The obtained abnormal records or risk warnings are fused and correlated with the abnormal detection results.
2. The method according to claim 1, characterized in that The method for obtaining the intelligent classification model includes: Provide the pre-prepared labeled sample data to the device, which performs training based on the machine learning algorithm, extracts the common points of the sample data and generates an intelligent prediction model; If it is impossible to clearly define the classification labels of the data and provide sample data classified by labels, clustering technology is used to utilize unsupervised machine learning to automatically cluster large amounts of unlabeled data according to content similarity. After clustering is completed, the user confirms the labels for each type of data. If the clustering results do not meet the requirements, the parameters and configurations are adjusted to complete the data clustering again. After the clustering labels are confirmed, all labeled sample data are provided to the device so that the device can be trained and learned based on the machine learning algorithm.
3. The method according to claim 1, characterized in that The identifying of structured data in various databases and extracting identification information includes: Conduct deep scans on various databases to identify the structured data stored therein, and use data fields as the smallest unit to extract various identification information including metadata, data HASH, fuzzy HASH, and classification and grading attributes.
4. The method according to claim 1, characterized in that: The method of identifying various unstructured data in various file servers and extracting identification information includes: Perform deep scans on various file servers to identify various unstructured data stored in them. Taking a single file as the smallest unit, automatically extract various identification information including metadata, overall file HASH, file content fuzzy HASH, and classification and grading attributes; metadata is used to describe the basic information of the data, overall file HASH is used to identify the uniqueness of the data, file content fuzzy HASH is used to analyze the similarity of the data, and classification and grading attributes are used to determine the security factors of the data.
5. The method according to claim 1, characterized in that The adding and managing of various data security protection devices, configuring and issuing security protection strategies based on data identification, includes: Add various types of data security protection devices and configure them, manage the connection status of data security protection devices, and monitor the operating status of data security protection devices; formulate various security protection strategies based on data identification, and establish a sensitive data identification library; issue the sensitive data identification library and protection strategies to data security protection devices; the sensitive data identification library includes identifications in data identifications whose security levels are above the preset level.
6. The method according to claim 1, characterized in that The collected logs and identification information are used to obtain abnormal records or risk warnings through security audit policies, including: Collect access logs of data sources through traffic, collect third-party system logs, and collect data identification information; parse the collected logs into structured data based on field mapping; according to the stage in which the logs are generated, database access logs correspond to data query operations, and third-party system logs correspond to system type-related operations. Each log traverses specific security audit policy rules, and generates abnormal records or risk alerts based on policy hits.
7. The method according to claim 1, characterized in that The abnormal data access behavior of the acquisition log, recording the abnormal feature information, and generating the abnormal detection result include: Traffic collection and preprocessing: Provide data access traffic log sample data, use unsupervised learning algorithms to train data access behavior clustering models, and learn the behavior patterns of different user groups; Feature extraction: Select the required feature subset from the original features, extract features from the traffic data, and aggregate each feature after grouping and slicing; Clustering: Automatically identify normal and abnormal patterns in traffic groups, and identify abnormal and normal group patterns through clustering; Building a baseline: After obtaining the behavior labels of traffic logs through clustering, a baseline model is built to predict future traffic. The baseline model includes a statistical baseline and a memory baseline. Real-time update and dynamic response: The memory baseline is trained incrementally according to fine-grained time intervals and directly called through the API; Using the baseline model, based on the normal behavior feature range, feature extraction and baseline comparison are performed on the collected data access logs. Behaviors that exceed the feature range are defined as abnormal data access behaviors, abnormal feature information is recorded, and anomaly detection results are generated.
8. The method according to claim 1, characterized in that: The obtained abnormal records or risk warnings and abnormal detection results are subjected to fusion correlation analysis, including: Various data security devices perform data security monitoring and execution according to the data protection strategy based on data identification, and report the alarm and protection disposal data for unified summary; Through the integration and correlation analysis of multiple data security logs within a specified time, the relevant data operation behaviors are linked into security events, and the occurrence chain of the events is restored; the events include data leakage and data abuse; Based on the results of the fusion analysis, instructions are issued to various data security devices to achieve security management and control of the entire data life cycle.
Citation Information
Cited By
Abnormal behavior risk early warning system, method and equipment based on autonomous learning
CN120342789A
Industrial templated modeling and dynamic access control method and system based on data lake
CN120639473A
Structured data security identifier intelligent extraction and matching method and model
CN121350027A