Data risk assessment method based on data concentration degree

By constructing a data classification grading tree and calculating data concentration index methods, the problems of strong subjectivity and ignoring data concentration in the existing technology are solved, and scientific and unified data classification and accurate risk assessment are realized, and data security protection capabilities are improved.

CN120197220AActive Publication Date: 2025-06-24HUAXIN CONSULTATING CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510670466.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing data classification and grading methods have strong subjectivity, large standards differences, and it is difficult to form an objective and unified classification system, and ignore the overall concentration of the data, resulting in limitations in risk assessment and data desensitization strategies.

Method used

By formulating a data classification grading rule table, building a data classification grading tree, and using the NLP model to parse new field data, dynamically update the classification grading tree. Then, the data concentration index is calculated through the leaf node weight formula, weighted distance formula and weighted distance mean formula, the data leakage risk is evaluated, and a predefined strategy library is built based on the risk level.

Benefits of technology

A scientific and unified data classification system has been realized, accurately measuring data concentration, improving the accuracy of risk assessment, providing targeted security strategies, and effectively improving data security protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197220A_ABST
    Figure CN120197220A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data security, and provides a data risk assessment method based on a data concentration ratio, which comprises the following steps of: formulating a data classification and grading rule table, constructing a data classification and grading tree for field data of a structured data table according to the data classification and grading rule table, and inserting new field data into the data classification and grading tree by utilizing a BERT model, according to the privacy level of field data in the data classification and grading tree, endowing the leaf nodes with weights to obtain a data weighted classification and grading tree, solving the shortest path distance of any two leaf nodes in the data weighted classification and grading tree by using a breadth-first search algorithm, calculating the weighted distance by combining the node weights, constructing a leaf pair weighted distance set, and obtaining a leaf pair weighted distance set; and calculating a weighted distance mean value and taking a reciprocal to obtain a data concentration index, comparing a risk assessment threshold with the data concentration index to determine a data risk level, and obtaining a security policy from a predefined policy library according to the data risk level. The method can scientifically evaluate the risk of data leakage, and provides powerful support for data security management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security, and particularly to a data risk assessment method based on data concentration. Background Art

[0002] In the big data era, data has become a key asset, and data security issues have become increasingly prominent. Data leakage may lead to serious consequences including user privacy leakage, corporate confidential information loss, and increased compliance risks.

[0003] Currently, the industry mainly uses traditional data classification and grading methods to reduce the risk of data leakage, but there are many problems. Existing methods mostly set data categories and levels based on rule matching or manual experience. There are large differences in standards between different industries and enterprises, and the subjectivity is strong. It is difficult to form an objective and unified classification system. Existing methods focus on data access control and encryption technologies, ignoring the effective measurement of the overall concentration of data. In the event of data leakage, the data concentration of data cannot be quantitatively analyzed, resulting in limitations in risk assessment and data desensitization strategy formulation, and it is easy to underestimate the real risk of the data set. At the same time, existing methods only focus on data levels and ignore the relevance between data categories. Therefore, how to achieve scientific and effective assessment of data leakage risk and provide strong support for data security management has become an urgent problem to be solved. For this reason, a data risk assessment method based on data concentration is proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a data risk assessment method based on data concentration to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: A data risk assessment method based on data concentration, comprising the following steps: S1. Formulate a data classification and grading rule table, and according to the data classification and grading rule table, construct a data classification and grading tree for the field data of the structured data table and store it in the database. At the same time, parse the new field data through the NLP model and insert it into the data classification and grading tree; S2. Obtain the privacy level corresponding to the field data through the data classification and grading tree, and accordingly assign weights to the leaf nodes of the data classification and grading tree to obtain a weighted data classification and grading tree; S3. For any two leaf nodes of the weighted data classification and grading tree, obtain the shortest path distance between the two leaf nodes through the breadth-first search algorithm, and obtain the weighted distance between the two leaf nodes according to the respective weights of the two leaf nodes, and accordingly construct a set of weighted distances of leaf pairs; S4. Calculate the weighted distance mean of leaf pairs based on the weighted distance set of leaf pairs, and obtain the reciprocal of the weighted distance mean of leaf pairs accordingly, which is the data concentration index of the structured data table; S5. Set a data risk assessment threshold, compare the data concentration index of the structured data table, and obtain the data risk level. At the same time, construct a predefined policy library, and obtain the security policy through the predefined policy library based on the data risk level.

[0006] Preferably, the method for formulating a data classification and grading rule table and constructing a data classification and grading tree for the field data of the structured data table according to the data classification and grading rule table: According to the "Data Classification and Grading Rules" (GB / T 43697-2024), and combined with the actual business needs of the industry or enterprise, formulate a data classification and grading rule table including data types, field instances, classification paths, and risk levels. The data types in the data classification and grading rule table macro-classify the field data according to business attributes or topics. The field instances in the data classification and grading rule table list the names of specific field data. The classification paths in the data classification and grading rule table describe the paths of field data in the classification and grading tree. The risk levels in the data classification and grading rule table assign the privacy levels of field data; Map the field data in the structured data table to the corresponding classification paths and privacy levels through the data classification and grading rule table. Take the set of all field data in the structured data table as the root node of the data classification and grading tree. Take the leaf nodes of the classification paths corresponding to the field data and the root node as the leaf nodes of the data classification and grading tree and the first-level nodes of the root node of the data classification and grading tree respectively. At the same time, merge the classification paths corresponding to the field data with the same path parts to obtain the data classification and grading tree; The data classification and grading tree is used to classify and manage the field data in the structured data table. The non-leaf nodes of the data classification and grading tree represent the classification levels of the field data. The leaf nodes of the data classification and grading tree store the field data and the corresponding privacy levels of the field data. All the nodes on the path from the root node of the data classification and grading tree to the leaf node storing the field data, excluding the root node of the data classification and grading tree, constitute the classification path corresponding to the field data, that is, the classification path in the data classification and grading rule table is the storage location of the field data in the data classification and grading tree; A node is a basic component unit of a tree-like data structure, used to store data and represent the relationships between nodes. The root node is the starting point of the data classification and grading tree, which is unique. The subsequent node levels are extended from the root node. The first-level node is the next-level node directly connected to the root node. The leaf node is the most terminal node in the data classification and grading tree, storing the specific field data and its corresponding privacy level, and there are no further sub-nodes extended. The non-leaf node is the node other than the leaf node; The structured data table is the basic data source for constructing the data classification and grading tree, and stores information on various business scenarios under different enterprises or industries.

[0007] Preferably, a method for parsing new field data through an NLP model and inserting it into the data classification and grading tree: Using the data classification and grading rule table as training data to complete the training of the NLP model. The trained NLP model can master the association between the semantic information contained in different field names and descriptions and the classification paths. Input the new field data into the trained NLP model to obtain the matching probabilities of the new field data with each classification path in the grading rule table, and select the classification path with the highest matching probability as the insertion position of the new field data in the data classification and grading tree. Insert the new field data into the data classification and grading tree to achieve dynamic update of the data classification and grading tree, enabling the data classification and grading tree to adapt to the constantly changing data environment.

[0008] Preferably, a method for obtaining the data weighted classification and grading tree: For the privacy level corresponding to the field data stored in the leaf nodes of the data classification and grading tree Obtain the weight corresponding to the field data through the leaf node weight formula , and store the weight corresponding to the field data in the leaf node storing the corresponding field data to obtain a data classification and grading tree in which the leaf nodes store field data, as well as the privacy level and weight corresponding to the field data, that is, the data weighted classification and grading tree; The leaf node weight formula is: ; where is the weight corresponding to the field data, is the privacy level corresponding to the field data, is the unique identifier of the leaf node of the data classification and grading tree, that is, the unique identifier of the field data.

[0009] Preferably, a method for obtaining the shortest path distance between two leaf nodes through the breadth-first search algorithm: Mark all nodes of the data weighted classification and grading tree as unvisited, which is used to clarify whether a node has been processed during the search process, and determine two leaf nodes in the data weighted classification and grading tree and , take one of the two leaf nodes as the starting node, mark it as visited, take the other leaf node as the destination node, set the path distance count value to 0, which is used to record the path length from the starting node to the destination node, that is, the shortest path distance between the two leaf nodes. Visit all the unvisited adjacent nodes of the starting node, and increment the path distance count value by 1. At the same time, determine whether all adjacent nodes include the destination node. If the destination node is included, the shortest path from the starting node to the destination node is found. At this time, the path distance count value is the shortest path distance between the two leaf nodes; If the destination node is not included, mark all the unvisited adjacent nodes of the starting node as visited and use them as the new starting node, and visit all the unvisited adjacent nodes of the new starting node; The adjacent node is a node directly connected to a specific node in a graph or tree data structure.

[0010] Preferably, the method for obtaining the weighted distance between two leaf nodes and constructing the leaf pair weighted distance set: The shortest path distance between two leaf nodes , as well as the weights of the two leaf nodes and , obtain the weighted distance between the two leaf nodes through the leaf node weighted distance formula ; The leaf node weighted distance formula is: ; Where is the weighted distance between two leaf nodes and , is the shortest path distance between two leaf nodes and , and are the weights corresponding to the leaf nodes and respectively; For all the leaf nodes of the data weighted classification and grading tree , perform non-repeating pairing, and calculate the weighted distance between two leaf nodes of all non-repeating pairings through the leaf node weighted distance formula to obtain the weighted distance between all two leaf nodes, and form a leaf pair weighted distance set accordingly ; The non-repeating pairing means that each pair of leaf nodes is only paired once.

[0011] Preferably, the method for calculating the mean value of the leaf pair weighted distance: For the leaf pair weighted distance set Obtain the weighted distance mean of leaf pairs through leaf nodes for the weighted distance mean formula; The weighted distance mean formula is: ; Wherein, is the weighted distance mean of leaf pairs, are two leaf nodes and 's weighted distance, i and j are index variables used to traverse leaf node pairs, and n is the number of field data in the structured data table, that is, the number of leaf nodes in the data classification and grading tree.

[0012] Preferably, set a data risk assessment threshold, compare the data concentration index of the structured data table, and obtain the method of data risk level: According to the weighted distance mean of leaf pairs, obtain the reciprocal of the weighted distance mean of leaf pairs , that is, the data concentration index of the structured data table , for the data concentration index of the structured data table Compare through the data risk assessment threshold; If is less than 0.1, then the data risk level of the structured data table is a low risk level, indicating that the correlation degree between the field data of the structured data table is relatively weak, and the possible impact when attacked or leaked is relatively small; If is greater than or equal to 0.1 and less than 0.2, then the data risk level of the structured data table is a medium risk level, indicating that there is a certain data correlation in the field data of the structured data table, and the data risk faced is also at a medium level, and corresponding protection measures need to be taken; If is greater than or equal to 0.2, then the data risk level of the structured data table is a high risk level, indicating that the correlation between the field data of the structured data table is very strong. Once a data leak occurs, attackers are likely to obtain more sensitive information through correlation analysis, causing serious security threats, and high-intensity security protection strategies must be taken immediately.

[0013] Preferably, construct a predefined security policy library, and obtain the security policy through the predefined security policy library according to the data risk level: Formulate corresponding security policies for different data risk levels, use the risk level field as the ID primary key, establish the mapping between the data risk level field and the security policy field, construct a predefined policy library containing the data risk level field and the security policy content field, match the data risk level with the data risk level field of the predefined security policy library, and map to obtain the security policy corresponding to the data risk level; The risk level field includes low risk level, medium risk level, and high risk level. The security policy field includes basic access control RBAC, log auditing, data backup, Advanced Encryption Standard 256-bit encryption algorithm AES-256, role-based fine-grained permission control, regular security assessment, multi-factor authentication MFA, real-time monitoring and alerting, full data encryption, dynamic data masking, data minimization principle, and network segmentation; The corresponding security policies formulated for different data risk levels are as follows: The security policies mapped to the low risk level are basic access control RBAC, log auditing, and data backup; The security policies mapped to the medium risk level are Advanced Encryption Standard 256-bit encryption algorithm AES-256, role-based fine-grained permission control, and regular security assessment; The security policies mapped to the high risk level are multi-factor authentication MFA, real-time monitoring and alerting, full data encryption, dynamic data masking, data minimization principle, and network segmentation.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention constructs a scientific and unified data classification and grading system, overcoming the traditional data classification and grading methods which are mostly based on rule matching or manual experience, with strong subjectivity, large differences in standards among different industries and enterprises, and it is difficult to form an objective and unified classification standard. The present invention formulates a rule table according to the "Rules for Data Classification and Grading" (GB / T 43697-2024), combines the actual business requirements of the industry or enterprise to construct a data classification and grading tree, clearly defines the data type, classification path, and risk level, realizes scientific classification and grading management of data, effectively avoids interference from subjective factors, and forms an objective and unified classification system.

[0015] 2. The present invention realizes accurate measurement of data concentration and improves the accuracy of risk assessment. Existing methods focus on data access control and encryption technology, ignoring the effective measurement of the overall data concentration, and there are limitations in risk assessment and data masking strategy formulation, which are prone to underestimating the true risk of the data set. The present invention obtains the data concentration index through the leaf node weight formula, leaf node weighted distance formula, and weighted distance mean formula, accurately quantifies the correlation degree between data, thereby accurately assessing the data leakage risk, providing a reliable basis for data masking and security protection, and making up for the deficiencies of traditional methods in risk assessment.

[0016] 3. The present invention takes into account the relevance of data categories and provides targeted security policies. The prior art only focuses on the data level and ignores the relevance between data categories. The present invention calculates the shortest path distance and weighted distance between leaf nodes in the data weighted classification and grading tree by means of the breadth-first search algorithm, constructs a set of weighted distances for leaf pairs, fully considers the data category association, comprehensively reflects potential risks. At the same time, it divides the risk levels according to the data concentration, constructs a predefined policy library, provides corresponding security policies for different risk levels, realizes the precise matching of security policies, and effectively improves the data security protection ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0018] Figure 1 It is a flowchart of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the present invention.

[0020] Embodiment, as Figure 1 shown, a data risk assessment method based on data concentration includes the following steps: S1. Formulate a data classification and grading rule table, and construct a data classification and grading tree for the field data of the structured data table according to the data classification and grading rule table, and store it in the database. At the same time, parse the new field data through the BERT model and insert it into the data classification and grading tree; S2. Obtain the privacy level corresponding to the field data through the data classification and grading tree, and accordingly assign weights to the leaf nodes of the data classification and grading tree to obtain a data weighted classification and grading tree; S3. For any two leaf nodes of the data weighted classification and grading tree, obtain the shortest path distance between the two leaf nodes through the breadth-first search algorithm, and obtain the weighted distance between the two leaf nodes according to the respective weights of the two leaf nodes, and accordingly construct a set of weighted distances for leaf pairs; S4. Calculate the weighted distance mean of leaf pairs based on the weighted distance set of leaf pairs, and obtain the reciprocal of the weighted distance mean of leaf pairs accordingly, which is the data concentration index of the structured data table; S5. Set a data risk assessment threshold, compare it with the data concentration index of the structured data table to obtain the data risk level. At the same time, construct a predefined policy library, and obtain the security policy through the predefined policy library based on the data risk level.

[0021] Furthermore, the working principle of the present invention will be described below through embodiments: Suppose there is an e-commerce enterprise. The structured data table of the e-commerce enterprise contains field data such as user information, order information, and payment information. User information includes name, ID number, and contact information. Order information includes order number, product name, and purchase amount. Payment information includes payment method and bank card number.

[0022] According to the "Rules for Data Classification and Grading" (GB / T 43697-2024) and the actual business requirements of the e-commerce enterprise, formulate a data classification and grading rule table. For example, classify the data into three categories: user-related data, business-related data, and payment-related data. Among them, user-related data is further divided into user basic information and user identity information, business-related data is divided into order information and product information, and payment-related data is divided into payment method information and bank card information. Map these data categories to the classification path and assign corresponding risk levels. For example, the ID number and bank card number are of high risk level, and the order number is of low risk level. Take the field data set as the root node, determine the leaf nodes and first-level nodes according to the classification path, merge the same path parts, construct a data classification and grading tree, and store it in the database. If new field data appears, such as the user's delivery address, use the data classification and grading rule table as training data to complete the training of the BERT model, input the user's delivery address, obtain the matching probability of the user's delivery address with each classification path, and select the classification path with the highest probability. For example: user-related data - user basic information. According to the classification path with the highest probability, insert the user's delivery address into the data classification and grading tree to realize the dynamic update of the data classification and grading tree.

[0023] Obtain the privacy level of each field data through the data classification and grading tree, and use the leaf node weight formula Calculate the weight. For example, the privacy level of the ID number is 4, and its weight is ; the privacy level of the order number is 2, and its weight is . Store the calculated weights in the corresponding leaf nodes to obtain a data weighted classification and grading tree.

[0024] Taking the field data of the two leaf nodes in the data weighted classification and grading tree as the ID card number and the bank card number as an example, mark all nodes as unvisited. Select the leaf node with the field data as the ID card number as the starting node and mark it as visited. The leaf node with the field data as the bank card number is used as the destination node. Set the path distance count value to 0. Visit all adjacent nodes of the starting node and count. Determine whether the destination node is included. If not, mark the adjacent node as visited and continue to visit it as the new starting node until the destination node is found to obtain the shortest path distance. Assume that the shortest path distance between the two leaf nodes with the field data as the ID card number and the bank card number is 3. According to the leaf node weighted distance formula , assuming the weight of the ID card number is , and the weight of the bank card number is , calculate the weighted distance between the two leaf nodes with the field data as the ID card number and the bank card number to be 12. Pair all leaf nodes of the data weighted classification and grading tree without repetition, calculate the weighted distance of all paired leaf nodes, and construct a set of weighted distances of leaf pairs.

[0025] For the set of weighted distances of leaf pairs, calculate the weighted distance mean through the leaf node pair weighted distance mean formula . Assume that there are 10 leaf nodes in the structured data table of the e-commerce enterprise. Calculate that the weighted distance mean is 8. Take the reciprocal of the weighted distance mean to obtain the data concentration index of 0.125.

[0026] Set the data risk assessment threshold and compare it with the data concentration index . When = 0.125, since 0.1 < 0.125 < 0.2, the data risk level of this structured data table is the medium risk level. Construct a predefined policy library, where the low risk level corresponds to basic access control RBAC, log auditing, and data backup. The medium risk level corresponds to the Advanced Encryption Standard 256-bit encryption algorithm AES-256, role-based fine-grained permission control, and regular security assessment. The high risk level corresponds to multi-factor authentication MFA, real-time monitoring and alerting, full data encryption, dynamic data masking, the data minimization principle, and network segmentation. Obtain the corresponding security policies from the predefined policy library according to the data risk level, that is, take security measures such as AES-256 encryption, role-based fine-grained permission control, and regular security assessment for the data of this e-commerce enterprise.

[0027] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above description of specific embodiments of this specification is given. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0028] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0029] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; any modification to the technical solutions recorded in the foregoing embodiments, or any equivalent replacement of some of the technical features, does not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of each embodiment of the present application, and shall all be included within the protection scope of the present application.

Claims

1. A data risk assessment method based on data concentration degree, characterized in that, It includes the following steps: S1. Formulate a data classification and grading rule table, and based on the data classification and grading rule table, construct a data classification and grading tree for the field data of the structured data table and store it in the database. At the same time, parse the new field data through the BERT model and insert it into the data classification and grading tree; S2. Through the data classification and grading tree, obtain the privacy level corresponding to the field data, and accordingly assign weights to the leaf nodes of the data classification and grading tree to obtain a data weighted classification and grading tree; S3. For any two leaf nodes of the data weighted classification and grading tree, obtain the shortest path distance between the two leaf nodes through the breadth-first search algorithm, and based on the respective weights of the two leaf nodes, obtain the weighted distance between the two leaf nodes, and accordingly construct a set of weighted distances for leaf pairs; S4. According to the set of weighted distances for leaf pairs, calculate the average value of the weighted distances for leaf pairs, and accordingly obtain the reciprocal of the average value of the weighted distances for leaf pairs, that is, the data concentration index of the structured data table; S5. Set a data risk assessment threshold, compare the data concentration index of the structured data table, obtain the data risk level, and at the same time, construct a predefined policy library, and obtain a security policy for the data risk level through the predefined policy library.

2. The data risk assessment method based on data concentration degree according to claim 1, wherein The method for formulating a data classification and grading rule table and constructing a data classification and grading tree for the field data of the structured data table according to the data classification and grading rule table: According to the data classification and grading rules and combined with the actual business needs of the industry or enterprise, formulate a data classification and grading rule table including data types, field instances, classification paths, and risk levels. The data type is used to macro-classify the field data according to business attributes or topics. The field instance is used to list the names of specific field data. The classification path is used to describe the path of the field data in the classification and grading tree. The risk level is used to assign the privacy level of the field data; Map the field data in the structured data table to the corresponding classification path and privacy level through the data classification and grading rule table. Take the set of all field data in the structured data table as the root node of the data classification and grading tree. Take the leaf nodes of the classification path corresponding to the field data and the root node as the leaf nodes of the data classification and grading tree and the first-level nodes of the root node of the data classification and grading tree respectively. At the same time, merge the classification paths corresponding to the field data with the same path part to obtain the data classification and grading tree; The data classification and grading tree is used to classify and grade the management of the field data in the structured data table. The non-leaf nodes of the data classification and grading tree represent the classification levels of the field data. The leaf nodes of the data classification and grading tree store the field data and the privacy level corresponding to the field data; The root node, first-level node, leaf node, and non-leaf node are components in the tree-like data structure.

3. The data risk assessment method based on data concentration degree according to claim 2, characterized in that, The method for parsing the new field data through the BERT model and inserting it into the data classification and grading tree: Use the data classification and grading rule table as training data to complete the training of the NLP model. Input the new field data into the trained BERT model to obtain the matching probabilities of the new field data with each classification path in the grading rule table. Select the classification path with the highest matching probability as the insertion position of the new field data in the data classification and grading tree, and insert the new field data into the data classification and grading tree; The BERT model is a natural language processing model.

4. The data risk assessment method based on data concentration degree according to claim 3, characterized in that, The method for obtaining the data weighted classification and grading tree: For the privacy level corresponding to the field data stored in the leaf nodes of the data classification and grading tree, obtain the weight corresponding to the field data through the leaf node weight formula, and store the weight corresponding to the field data into the leaf node storing the field data, thereby obtaining the data weighted classification and grading tree; The data weighted classification and grading tree is a data classification and grading tree in which the leaf nodes store field data, as well as the privacy level and weight corresponding to the field data.

5. The data risk assessment method based on data concentration degree according to claim 4, wherein The method for obtaining the shortest path distance between two leaf nodes through the breadth-first search algorithm: Mark all nodes of the data weighted classification and grading tree as unvisited. Determine two leaf nodes in the data weighted classification and grading tree. Take one of the two leaf nodes as the starting node and mark it as visited, and take the other leaf node as the destination node. Set the path distance count value to 0. Visit all unvisited adjacent nodes of the starting node, and increment the path distance count value by 1. At the same time, determine whether all adjacent nodes include the destination node. If the destination node is included, obtain the path distance count value, which is the shortest path distance; If the destination node is not included, mark all unvisited adjacent nodes of the starting node as visited and use them as the new starting node, and visit all unvisited adjacent nodes of the new starting node; The path distance count value is used to obtain the shortest path distance between two leaf nodes; The adjacent node is a node directly connected to a specific node in a graph or tree data structure.

6. The data risk assessment method based on data concentration degree according to claim 5, characterized in that The method for obtaining the weighted distance between two leaf nodes and constructing a set of weighted distances of leaf pairs: Obtain the weighted distance between two leaf nodes by using the shortest path distance between the two leaf nodes and the weights of the two leaf nodes through the leaf node weighted distance formula; For all leaf nodes of the data weighted classification and grading tree, perform non-repeating pairing, and calculate the weighted distances between all non-repeating paired two leaf nodes through the leaf node weighted distance formula to obtain the weighted distances between all two leaf nodes, thereby forming a set of weighted distances of leaf pairs; The non-repeating pairing means that each pair of two leaf nodes is only paired once.

7. The data risk assessment method based on data concentration degree according to claim 6, characterized in that, The method for calculating the mean value of the weighted distances of leaf pairs: Obtain the mean value of the weighted distances of leaf pairs from the set of weighted distances of leaf pairs through the leaf node pair weighted distance mean formula.

8. The data risk assessment method based on data concentration degree according to claim 7, wherein The method for setting a data risk assessment threshold and comparing the data concentration index of a structured data table to obtain a data risk level: Data concentration index for structured data tables Compare with the data risk assessment threshold; If is less than 0.1, the data risk level of the structured data table is a low risk level; If is greater than or equal to 0.1 and less than 0.2, the data risk level of the structured data table is medium risk level; If is greater than or equal to 0.2, the data risk level of the structured data table is a high risk level.

9. The data risk assessment method based on data concentration degree according to claim 8, characterized in that The method for constructing a predefined security policy library and obtaining a security policy based on the data risk level through the predefined security policy library: Formulate corresponding security policies for different data risk levels, use the risk level field as the ID primary key, establish the mapping between the data risk level field and the security policy field, construct a predefined policy library containing the data risk level field and the security policy content field, match the data risk level with the data risk level field of the predefined security policy library, and map to obtain the security policy corresponding to the data risk level; The risk level field includes low risk level, medium risk level, and high risk level, and the security policy field includes basic access control RBAC, log auditing, data backup, Advanced Encryption Standard 256-bit encryption algorithm AES-256, role-based fine-grained permission control, regular security assessment, multi-factor authentication MFA, real-time monitoring and alerting, full data encryption, dynamic data masking, data minimization principle, and network segmentation; The formulating corresponding security policies for different data risk levels is as follows: The security policies mapped by the low risk level are basic access control RBAC, log auditing, and data backup; The security policies mapped by the medium risk level are Advanced Encryption Standard 256-bit encryption algorithm AES-256, role-based fine-grained permission control, and regular security assessment; The security policies mapped by the high risk level are multi-factor authentication MFA, real-time monitoring and alerting, full data encryption, dynamic data masking, data minimization principle, and network segmentation.

Citation Information

Patent Citations

  • Gene signatures predictive of metastatic disease

    CN108513587A

  • Risk analysis method and system for third-party application of mobile internet operating system

    CN114547606A

  • Method for quantifying and protecting edge data

    CN118228300A

  • Network security management method and system for digital assets

    CN119250540A

  • Multi-source software supply chain intelligent analysis method and system

    CN119720225A