A method for building a data leakage prevention system

By identifying and integrating the key sensitivity characteristics of multi-source heterogeneous data, a data leakage risk prediction model is built, which solves the privacy leakage risk problem brought about by multi-source heterogeneous data combination in the big data era, and realizes the dual guarantee of data security and value.

CN119377995BActive Publication Date: 2025-06-27新疆国融信联大数据投资有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411522291.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-06-27
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In the era of big data, the combination of multi-source heterogeneous data may lead to unpredictable privacy leakage risks, and existing technologies are difficult to identify and prevent these risks while fully exploring the value of data.

Method used

By acquiring multiple data sets, data preprocessing and feature extraction are performed to identify key sensitivity features. A sensitivity fusion model for multi-source heterogeneous data was constructed, and principal component features were extracted through weighted fusion and dimensionality reduction. Based on these characteristics, a data leakage risk prediction model is constructed, a leakage risk score of the data combination is calculated, and early warning is triggered and desensitization is implemented based on preset thresholds.

Benefits of technology

Effectively identify and prevent potential leakage risks brought about by multi-source heterogeneous data combinations, improve data security, and ensure the rational use of data and the full release of value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377995B_ABST
    Figure CN119377995B_ABST
Patent Text Reader

Abstract

The present application provides a method for building a data leakage prevention system, including: constructing a data leakage risk prediction model according to the principal component features, training the model using the support vector machine algorithm to obtain the data leakage risk prediction model, inputting the multi-source heterogeneous data to be predicted into the risk prediction model, and calculating the leakage risk score of the data combination; presetting a sensitivity threshold, determining whether the leakage risk score exceeds the preset sensitivity threshold, if it exceeds the threshold, marking the data combination as highly sensitive data, triggering an early warning mechanism, and implementing desensitization and encryption on the data combination according to the preset desensitization rules and encryption methods; continuously monitoring the data update situation of different data sources, when new data is detected, automatically triggering the data fusion and leakage risk prediction process and obtaining the sensitivity feature representation of the new data, updating the data leakage risk prediction model, and dynamically adjusting the sensitivity level and protection measures of the data combination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for building a data leakage prevention system. Background Art

[0002] In the era of big data, enterprises and institutions often need to integrate and analyze heterogeneous data from multiple sources to explore the potential value contained therein. However, the combination of different data sets may generate unexpected risks of sensitive information leakage. Even if each individual data set has been desensitized, seemingly unrelated data combinations may still reveal highly sensitive information such as personal identity and privacy. This requires identifying which data combinations may produce high sensitivity before data fusion, and taking corresponding protection measures accordingly. The current technical contradiction is that the value of data fusion is often based on the full combination of multi-source heterogeneous data, but this combination may bring unpredictable privacy leakage risks. How to fully explore the value of data while identifying the sensitivity of various data combinations in advance and taking preventive measures is an urgent problem to be solved. This requires in-depth analysis of data characteristics and business rules in different industry scenarios, abstracting a set of scientific sensitivity assessment methods, and building a prediction model based on this to achieve early warning and precise prevention and control of high-risk data combinations. At the same time, this set of methods must be able to adapt to the dynamic changes in the data environment, continuously optimize the prediction effect when new data is continuously added, and ultimately achieve the full release of data value and strong protection of privacy security. Summary of the invention

[0003] The present invention provides a method for building a data leakage prevention system, which mainly includes:

[0004] Obtain multiple data sets from different data sources, perform data preprocessing on each data set, use data mining methods to extract features, obtain sensitivity features of multiple data sets, and screen out several key features that have a significant impact on data leakage risks based on preset sensitivity assessment criteria;

[0005] Aiming at several key characteristics, a sensitivity fusion model of multi-source heterogeneous data is constructed. By setting weight parameters, the key features of different data sources are weighted and fused to obtain the fused target sensitivity feature set. The principal component analysis method is used to reduce the dimension of the target sensitivity feature set and extract the principal component features that affect the contribution of data leakage risk.

[0006] According to the main component characteristics, a data leakage risk prediction model is constructed, and the support vector machine algorithm is used to train the model and obtain the data leakage risk prediction model. The multi-source heterogeneous data to be predicted is input into the risk prediction model, and the leakage risk score of the data combination is calculated;

[0007] Based on a preset sensitivity threshold, determine whether the leakage risk score exceeds the preset sensitivity threshold. If it exceeds the threshold, mark the data combination as highly sensitive data, trigger the warning mechanism, and perform desensitization and encryption on the data combination according to the preset desensitization rules and encryption methods.

[0008] Continuously monitor the data update situation of different data sources. When new data is detected, automatically trigger the data fusion and leakage risk prediction processes, obtain the sensitivity feature representation of the new data, update the data leakage risk prediction model, and dynamically adjust the sensitivity level and protection measures of the data combination.

[0009] Based on business requirements and data usage scenarios, formulate data access and sharing policies. For data combinations with low leakage risk but high sensitivity, strictly restrict data access and transfer to ensure data security. For data combinations with high leakage risk but low sensitivity, allow authorized users to use and analyze the data within a certain range according to user permissions and access approval rules.

[0010] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0011] The present invention discloses a method for building a data leakage prevention system. The method obtains multiple data sets, performs data preprocessing and feature extraction, and identifies key sensitivity features. Using association analysis technology, discovers potential associations between data sets and evaluates the risks after combination. Constructs a sensitivity fusion model, performs weighted fusion and dimensionality reduction on the features, and extracts the principal component features. Based on this, constructs a data leakage risk prediction model, calculates the leakage risk score of the data combination. When it exceeds the preset threshold, triggers a warning and performs desensitization and encryption. The present invention can also continuously monitor data updates, dynamically adjust the risk prediction model, and formulate differentiated data access policies according to business requirements. This method can effectively identify and prevent potential leakage risks brought by multi-source heterogeneous data combinations, improve data security, and at the same time ensure the reasonable use of data. Brief Description of the Drawings

[0012] Figure 1 It is a flowchart of a method for building a data leakage prevention system of the present invention.

[0013] Figure 2 It is a schematic diagram of a method for building a data leakage prevention system of the present invention.

[0014] Figure 3 It is another schematic diagram of a method for building a data leakage prevention system of the present invention. Detailed Embodiments

[0015] The technical solutions in the embodiments of the present invention will be clearly and detailedly described below in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0016] As Figures 1-3 , the method for building a data leakage prevention system in this embodiment may specifically include:

[0017] S101. Obtain multiple data sets from different data sources, perform data preprocessing on each data set, extract features using data mining methods to obtain the sensitivity features of the multiple data sets, and screen out several key features that have a significant impact on the data leakage risk according to the preset sensitivity evaluation criteria.

[0018] Obtain the data in the relational database according to the data source interface adapter, and read the data files on the file server through the FTP protocol; execute the data extraction, transformation, and loading process to obtain the original data set; perform data cleaning, standardization, and structured transformation on the original data set to obtain the preprocessed data set; reduce the dimension of the preprocessed data set to extract the main features, and calculate the influence degree of each feature on the data leakage risk; if the information gain ratio exceeds the preset threshold, include this feature in the candidate feature set; construct a feature importance evaluation model according to the candidate feature set, and iteratively eliminate the feature with the lowest importance through the recursive feature elimination method until the remaining feature quantity reaches the preset target; perform normalization processing on the remaining features to obtain the standardized key feature data set; use the Apriori algorithm to perform association rule mining on the standardized key feature data set, and calculate the support degree and confidence degree between feature pairs; draw a feature association network diagram according to the support degree and confidence degree, and use the Dijkstra shortest path algorithm to identify the risk propagation path.

[0019] Exemplarily, construct a data source interface adapter, use JDBC technology to connect to a relational database, adopt the FTP protocol to read data files on the file server, apply the Kettle tool to execute the data extraction, transformation, and loading process to extract raw data, perform data cleaning, standardization, and structural transformation according to the data type and format, and generate a preprocessed data set in a standard format. For the preprocessed data set, use the principal component analysis algorithm to reduce the dimension and extract the main features, adopt the information gain calculation formula in the ID3 algorithm to calculate the influence degree of each feature on the data leakage risk, set the information gain ratio threshold to 0.1, and filter out the candidate feature set that has a significant impact on the risk. Construct a feature importance evaluation model based on decision trees, input the candidate feature set and the preset sensitivity evaluation index, and iteratively eliminate redundant features through the recursive feature elimination method. Each iteration eliminates the 10% features with the lowest importance until the remaining feature quantity reaches the preset target, and obtain the final key feature subset. Normalize the selected key features, use the Min-Max normalization method to map the feature values to the [0,1] interval, and obtain the standardized key feature data set. Use the Apriori algorithm to perform association rule mining on the standardized key feature data set, set the minimum support to 0.05 and the minimum confidence to 0.6, calculate the support and confidence between feature pairs, draw a feature association network graph, use the Dijkstra shortest path algorithm to identify the risk propagation path, and use the PageRank algorithm to calculate the node centrality to determine the key nodes, forming the data leakage risk assessment result. When constructing the data source interface adapter, connect to the Oracle database through JDBC technology, and establish a connection using the connection string "jdbc:oracle:thin:@localhost:1521:orcl", username "admin", and password "password123". At the same time, connect to the file server using the FTP protocol, and access the " / data" directory on the server using the host address "ftp.example.com", port number 21, username "ftpuser", and password "ftppass456". Use the Kettle tool to create a data extraction, transformation, and loading job, set the data source connection parameters, and define the transformation steps including table input, text file input, field selection, filter, and table output. Uniformly convert the data from different sources into CSV format encoded in UTF-8 and store it in the preprocessed data set directory " / preprocessed_data". Apply the principal component analysis algorithm to the preprocessed data set, set the number of principal components to 10, and the cumulative explained variance ratio reaches 95%. Use the ID3 algorithm to calculate the information gain. For feature A, calculate the information gain G(D,A) = H(D) - H(D|A), where H(D) is the entropy of the data set D, and H(D|A) is the conditional entropy under the condition of feature A.Set the information gain ratio threshold to 0.1, and filter out the features with information gain ratio greater than the threshold as the candidate feature set. Construct a feature importance evaluation model based on decision trees, use the CART algorithm to generate decision trees, set the maximum depth of the tree to 10, and the minimum number of samples in the leaf nodes to 5. Iteratively eliminate redundant features through the recursive feature elimination method. In each iteration, remove the 10% features with the lowest importance scores until the number of remaining features is reduced to 50. Perform Min-Max normalization on the selected key features, map the feature value x to the interval [0, 1], and the calculation formula is the normalized value x'=(x - min) / (max - min), where min and max are the minimum and maximum values of this feature respectively. Use the Apriori algorithm to mine association rules from the normalized key feature data set, set the minimum support to 0.05 and the minimum confidence to 0.6, and generate an association rule set. Construct a feature association network graph based on the association rules, where the nodes represent features, the edges represent association relationships, and the weights of the edges are the confidence levels of the corresponding rules. Use the Dijkstra algorithm to calculate the shortest path between nodes and identify the risk propagation path. Use the PageRank algorithm to calculate the node centrality, set the number of iterations to 100, and the damping factor to 0.85. The top 10 nodes with the highest scores are determined as key nodes. The final data leakage risk assessment results include a list of key features, a feature association network graph, a risk propagation path, and key node information.

[0020] Analyze the data type, structure, and content of each data set, cross-compare the fields and information in different data sets, identify the associated data items, and use association rule mining or graphical analysis data association analysis methods to find the potential association relationships between different data sets, evaluate the new information or inferences generated after combining data sets, and identify potential risk points. Potential risk points include personal privacy leakage and sensitive information inference.

[0021] Obtain the metadata information of each data set, extract the data type, structure, and field information from the metadata information; calculate the Jaccard similarity of the field names according to the field information, and calculate the Pearson correlation coefficient of the data features for numerical fields; if the Jaccard similarity or Pearson correlation coefficient is greater than the preset similarity threshold, then determine the corresponding fields as potential associated field pairs; apply the Apriori algorithm to mine the frequent item sets between data sets; calculate the support and confidence of the association rules for the frequent item sets, if the support is greater than the preset minimum support and the confidence is greater than the preset confidence threshold, then determine it as a strongly associated data item combination; use the strongly associated data item combination to construct a data association network, and use the Louvain algorithm to detect the tightly associated data subsets in the data association network; based on the data association network, predict the information leakage risk and obtain a risk score; if the risk score is higher than the preset risk threshold, then mark the corresponding node as a high-risk node.

[0022] Exemplarily, the data type, structure, and field information of each dataset are extracted by a metadata parsing tool, and the pandas library in Python is used to calculate the field distribution characteristics, including mean, variance, and distribution type, to generate a dataset structure description file. The Jaccard similarity is used to calculate the field name similarity, and the Pearson correlation coefficient is used to calculate the data feature similarity of numerical fields. A similarity threshold of 0.8 is set to filter out potential associated field pairs and construct a field association graph between datasets. According to the field association graph, the Apriori algorithm is applied to mine the frequent item sets and association rules between datasets. A minimum support of 0.05 and a confidence threshold of 0.6 are set to identify strongly associated data item combinations. Semantic analysis is performed on the identified strongly associated data item combinations. The Word2Vec model is used to calculate the semantic similarity between fields to verify the rationality of the association. The semantic similarity threshold is set to 0.7. Based on the verified associated data items, a data association network is constructed. The Louvain algorithm is used to detect closely associated data subsets. Combining with the industry-standard personal information classification catalog and the sensitive information thesaurus automatically expanded by the Naive Bayes classifier, the nodes containing personal privacy and sensitive information are marked. Using the GraphSAGE graph neural network model, based on the constructed data association network, the potential information leakage risk is predicted, and a risk score and a warning report are generated. The scoring range is 0-100, and the nodes with scores higher than 80 are marked as high-risk, and the final list of potential risk points is output. During the implementation process, first, an automated metadata parsing tool is used to scan multiple datasets to extract field information including user ID, name, age, address, etc. Subsequently, the pandas library is used to calculate the statistical characteristics of these fields, such as the number of unique values of the user ID, the average length of the name is 10 characters, the mean age is 35 years old, and the variance is 15, etc. Then, the similarity of fields in different datasets is calculated. For example, the Jaccard similarity between "customer name" and "user name" is 0.85, and the Pearson correlation coefficient between "annual income" and "monthly salary" is 0.92, both higher than the set threshold of 0.8, so they are identified as potential associated fields. Then, the Apriori algorithm discovers frequent item sets among these associated fields. For example, the support of {user ID, mobile phone number, home address} is 0.06, and the confidence is 0.75, exceeding the preset threshold. Further semantic analysis is performed on these associated items. The pre-trained Word2Vec model is used to calculate the field semantic similarity. For example, the similarity between "address" and "location" is 0.85, verifying the rationality of their association. Subsequently, a data association network is constructed, where the nodes represent data fields and the edges represent association relationships. The Louvain algorithm is applied to detect communities and identify highly associated subgraphs, such as a closely associated subset containing {name, ID number, home address, work unit}. Combining with the industry-standard personal information classification catalog and the sensitive information thesaurus trained by the Naive Bayes classifier, potential sensitive information nodes are marked.Finally, use the GraphSAGE model to analyze this association network and predict the risk scores of each node. For example, a node containing detailed user consumption records received a high-risk score of 92 and was included in the final list of potential risk points, generating a detailed risk assessment report, including potential information leakage channels and possible impacts.

[0023] S102. For several key features, construct a sensitivity fusion model for multi-source heterogeneous data. By setting weight parameters, weightedly fuse the key features of different data sources to obtain a set of target sensitivity features after fusion. Use the principal component analysis method to perform dimensionality reduction processing on the set of target sensitivity features, and extract the principal component features that affect the contribution degree of data leakage risk.

[0024] Obtain the preset scoring criteria for multiple heterogeneous data sources. The preset scoring criteria include data update frequency, data integrity, and outlier ratio; calculate the reliability score of each data source according to the preset scoring criteria to obtain a data source weight vector; use the random forest algorithm to calculate the importance score of the key features in each data source in combination with the data source reliability score, and construct a weighted feature fusion matrix through the product of the data source weight and the feature importance score; use the weighted feature fusion matrix to perform a weighted summation operation on the key features of different data sources. If there are identical features, use the weighted average method for fusion to obtain a set of target sensitivity features after fusion; perform Z-score standardization processing on the set of target sensitivity features after fusion to eliminate the dimensional difference between different features; for the standardized set of target sensitivity features, calculate the feature covariance matrix and solve the eigenvalues and eigenvectors, select the first several principal components whose cumulative contribution rate reaches the preset threshold, and arrange them in descending order of eigenvalue magnitude to obtain the dimensionality-reduced set of principal component features.

[0025] Exemplarily, according to the preset data source reliability scoring criteria, including data update frequency, data integrity, and outlier ratio, multiple heterogeneous data sources are evaluated, the reliability score of each data source is calculated, and a data source weight vector is generated. The random forest algorithm is used to calculate the importance scores of key features in each data source in combination with the data source reliability scores, and a weighted feature fusion matrix is constructed through the product of the data source weights and the feature importance scores. Using the weighted feature fusion matrix, weighted summation operations are performed on the key features of different data sources. For features with the same name, a weighted average method is used for fusion to generate a fused target sensitivity feature set. Z-score normalization is performed on the fused feature set to eliminate the dimensional differences between different features. The principal component analysis algorithm is applied to the normalized target sensitivity feature set to calculate the feature covariance matrix, solve the eigenvalues and eigenvectors, select the first several principal components with a cumulative contribution rate reaching 85%, and arrange them in descending order of eigenvalue magnitude to obtain the dimension-reduced principal component feature set. In the actual implementation process, first, the reliability of three heterogeneous data sources (relational database, log file, and API data stream) is evaluated. The data update frequency is calculated by the number of updates per hour, the data integrity is represented by the ratio of non-empty fields, and the outlier ratio is identified by the 3-fold standard deviation method. For example, the update frequency of data source A is 24 times per day, the integrity is 98%, the outlier ratio is 0.5%, and the final reliability score is 0.9; the scores of data sources B and C are 0.8 and 0.7 respectively. Subsequently, the random forest algorithm is used to calculate feature importance. Suppose there are 10 key features, and the importance score of the "user behavior" feature in data source A is 0.8, which is multiplied by the data source weight 0.9 to obtain a weighted score of 0.72. This process is repeated for all features and data sources to construct a 10x3 weighted feature fusion matrix. Then, weighted averaging is performed on features with the same name in the matrix. For example, the weighted scores of "user ID" in the three data sources are 0.72, 0.64, and 0.56 respectively, and the fused score is 0.64. This operation is performed on all features to obtain a 10-dimensional fused feature vector. Then, Z-score normalization is performed on these 10 features, subtracting the mean value of each feature value and dividing by the standard deviation to make the mean value of all features 0 and the standard deviation 1. Finally, the principal component analysis is applied to the normalized feature set to calculate the feature covariance matrix and solve the eigenvalues and eigenvectors. Suppose the cumulative contribution rate of the first 4 principal components reaches 86%, exceeding the preset threshold of 85%, then these 4 principal components are selected as the final dimension reduction result to form a 4-dimensional principal component feature set, effectively extracting the main factors affecting the data leakage risk.

[0026] S103. Construct a data leakage risk prediction model based on the principal component features. Use the support vector machine algorithm to train the model and obtain the data leakage risk prediction model. Input the multi-source heterogeneous data to be predicted into the risk prediction model, and calculate the leakage risk score of this data combination.

[0027] Obtain the known risk level samples in the historical sample database and log files, and perform missing value filling and outlier detection processing on the sample data; according to the processed sample data, determine the principal component features as input variables and the risk level as the output variable to obtain a standardized training data matrix; use the grid search method to optimize the kernel function type, penalty parameter C, and kernel function parameter γ of the support vector machine. If the search range of the penalty parameter C is from a preset first value to a second value, and the search range of the kernel function parameter γ is from a preset third value to a fourth value, then determine the optimal parameter combination through cross-validation; according to the optimal parameter combination, train the support vector machine model, use the sequential minimal optimization algorithm to solve the dual problem, and obtain the support vectors and decision function; for the multi-source heterogeneous data to be predicted, through feature extraction and principal component transformation, obtain the principal component feature representation in the same format as the training data, and input the principal component feature representation into the trained support vector machine model to judge the leakage risk score of the data combination.

[0028] Exemplarily, a training data set is constructed based on the principal component features. Historical samples with known risk levels are extracted from the database and log files, and the data is filled with missing values and outliers are detected. The principal component features are used as input variables, and the risk level is used as the output variable to form a standardized training data matrix. The grid search method is used to optimize the kernel function type, penalty parameter C, and kernel function parameter γ of the support vector machine. The search range of C is set from 0.1 to 100, and the search range of γ is from 0.001 to 1, both growing exponentially. The optimal parameter combination is determined through 5-fold cross-validation. The support vector machine model is trained using the optimal parameter configuration, and the sequential minimal optimization algorithm is used to solve the dual problem. The specific implementation includes selecting variable pairs that violate the KKT conditions, updating the Lagrange multipliers and bias terms, and iterating until convergence to obtain the support vectors and decision function, thus constructing a data leakage risk prediction model. The performance of the trained model is evaluated, and metrics such as accuracy, precision, and recall are calculated using an independent test set. The risk warning threshold is set to 0.8. The multi-source heterogeneous data to be predicted is transformed into the principal component feature representation in the same format as the training data through feature extraction and principal component transformation, and is input into the trained support vector machine model. The leakage risk score of the data combination is calculated through the decision function. In the actual implementation process, first, 10,000 historical data samples are extracted from the enterprise database and log files, including 20 principal component features and corresponding risk level labels. The data is preprocessed, 5% of the missing values are filled using the mean value, and 2% of the outliers are detected and removed through the 3-sigma method, finally obtaining 9,800 valid samples. Then, the grid search is used to optimize the support vector machine parameters. The search range of C is set to [0.1, 1, 10, 100], and the search range of γ is [0.001, 0.01, 0.1, 1]. The 5-fold cross-validation is used to evaluate the performance of each group of parameters. After 200 iterations, the optimal parameter combination is determined as C = 10, γ = 0.01, and the RBF kernel is selected as the kernel function. The support vector machine model is trained using the sequential minimal optimization algorithm, and two samples that most violate the KKT conditions are selected in each iteration to update their Lagrange multipliers and the model bias term. After 5,000 iterations, the model converges, obtaining 1,500 support vectors. Subsequently, 2,000 independent test samples are used to evaluate the model performance, obtaining an accuracy of 94%, a precision of 92%, and a recall of 95%. The risk warning threshold is set to 0.8, that is, a risk score exceeding 0.8 is regarded as a high risk. Finally, a new set of multi-source heterogeneous data is predicted. First, the data from different sources is uniformly transformed into a 20-dimensional principal component feature representation and input into the trained model, obtaining a risk score of 0.75, which is lower than the warning threshold, and is determined to be at the medium risk level.

[0029] S104. Preset a sensitivity threshold, and determine whether the leakage risk score exceeds the preset sensitivity threshold. If it exceeds the threshold, mark the data combination as highly sensitive data, trigger the warning mechanism, and perform desensitization and encryption on the data combination according to the preset desensitization rules and encryption methods.

[0030] Obtain the leakage risk score of the data combination to be evaluated from the database, and determine whether the risk score exceeds the sensitivity threshold; if the risk score exceeds the sensitivity threshold, mark the data combination as highly sensitive data, and generate a warning message including the data ID, risk score, and trigger time; according to the type of sensitive information in the data combination and the highly sensitive data mark, select the corresponding desensitization algorithm from the preset desensitization rule library to perform desensitization processing on the data combination; for the data combination after desensitization processing, select an encryption algorithm from AES, RSA, or homomorphic encryption algorithm to perform encryption processing according to the data sensitivity level and processing performance requirements, and obtain the encrypted data combination; store the encrypted data combination in the security audit database, and record the information of desensitization processing and encryption processing.

[0031] Exemplarily, according to historical data statistics, the ROC curve analysis is used to select the best balance point, the sensitivity threshold is set to 0.8, the leakage risk scores of the data combinations to be evaluated are read from the database, and the comparator determines whether the risk scores exceed the preset threshold. If the risk score exceeds the threshold, the data marking module is called to mark the data combination as highly sensitive data, and at the same time, a warning signal is triggered to generate a warning message containing the data ID, risk score, and trigger time, which is written into the log system and pushed to relevant personnel. According to the different types of sensitive information and highly sensitive data markings in the data combination, the corresponding desensitization algorithms are automatically selected from the preset desensitization rule library. For example, the K-anonymity process is used for ID numbers, and the L-diversity algorithm is used to partially hide mobile phone numbers. The encryption module is called to select a suitable encryption algorithm from AES, RSA, and homomorphic encryption algorithms according to the data sensitivity level and processing performance requirements, and the desensitized data is encrypted to generate an encrypted data combination and stored. The data processing log is recorded, including the detailed information of desensitization and encryption operations, and is stored in the security audit database for subsequent data processing tracking and security auditing. In the actual implementation process, first, 10,000 historical data are used for ROC curve analysis to calculate the true positive rate and false positive rate at different thresholds, and the point 0.8 that maximizes the Youden index (true positive rate - false positive rate) is selected as the sensitivity threshold. Subsequently, the data combination to be evaluated is read from the Oracle database, and its leakage risk score is 0.85. The comparator determines that 0.85 exceeds the preset threshold of 0.8, automatically triggers the data marking module, marks the data combination as "highly sensitive", and at the same time generates a warning message as {data ID: DT20240415001, risk score: 0.85, trigger time: 2024-04-15 14:30:25}, which is written into the Elasticsearch log system and pushed to the security administrator through enterprise WeChat. Then, the system detects that the data combination contains two types of sensitive information, ID numbers and mobile phone numbers, and automatically selects the K-anonymity algorithm (K = 5) from the desensitization rule library to process the ID numbers and uses the L-diversity algorithm (L = 3) to partially hide the mobile phone numbers. Then, according to the high sensitivity of the data and real-time processing requirements, the AES-256 encryption algorithm is selected to encrypt the desensitized data, and the preset key "3F4528482B4D6251655468576D5A7134" is used for the encryption operation. Finally, the system generates a processing log as {operation ID: OP20240415001, desensitization algorithms: [K-anonymity, L-diversity], encryption algorithm: AES-256, processing time: 2024-04-15 14:30:26}, and stores it in the MongoDB security audit database.

[0032] S105. Continuously monitor the data update situation of different data sources. When new data is detected, automatically trigger the data fusion and leakage risk prediction processes, obtain the sensitivity feature representation of the new data, update the data leakage risk prediction model, and dynamically adjust the sensitivity level of the data combination and the protection measures.

[0033] Receive the update status information sent by the data source. The update status information includes the timestamp of the new data and the data source identifier; call the data fusion module according to the update status information. The data fusion module adopts an incremental update strategy based on timestamp comparison to integrate the new data with the existing data; perform feature extraction on the fusion data set generated by the data fusion module, and use the principal component analysis method to reduce the dimension of the features to obtain a new sensitivity feature representation; input the sensitivity feature representation into the data leakage risk prediction model to obtain the updated risk score; if there is a deviation between the updated risk score and the preset threshold, use the stochastic gradient descent algorithm to perform online learning and incremental training on the data leakage risk prediction model, and update the parameters of the data leakage risk prediction model; adjust the protection measures of the corresponding data combination according to the updated risk score, including the encryption intensity or the access control level; record the parameter update of the data leakage risk prediction model and the protection measure adjustment information in the security audit database.

[0034] Exemplarily, a data source monitoring program is deployed to detect the update status of each data source in real time through timed polling every 5 minutes or using the Apache Kafka message queue mechanism. When new data is found, the update timestamp and data source identifier are recorded. According to the detected new data, the data fusion module is called to integrate the new data with the existing data. An incremental update strategy based on timestamp comparison is adopted to update only the newly added or changed data, generating a fused data set. Feature extraction is performed on the fused data set, and principal component analysis is used for dimensionality reduction, retaining the principal components that explain 95% of the variance, obtaining a new sensitivity feature representation. The feature vector is input into the existing data leakage risk prediction model to calculate the updated risk score. Based on the new risk score, the parameters of the data leakage risk prediction model are updated, and the stochastic gradient descent algorithm is used for online learning and incremental training to dynamically adjust the model weights. At the same time, according to the new sensitivity level, the protection measures for the corresponding data combination are automatically updated, such as adjusting the encryption strength or modifying the access control level. Detailed logs of model updates and protection measure adjustments are recorded, including the update time, adjustment content, and operator, and the logs are stored in a security audit database to ensure the traceability of the entire process. In the actual implementation process, the system deploys a data source monitoring program, sets a timed task with a 5-minute interval using the Quartz scheduling framework, and configures the Apache Kafka message queue with the Topic "data_update". When 100 new records are detected in the database table "user_info", a message {source: "user_db", timestamp: "2024-04-15 15:30:00", count: 100} is immediately sent to Kafka. After receiving the message, the data fusion module uses JDBC to connect to the database and executes the SQL query "SELECT * FROM user_info WHERE update_time > '2024-04-15 15:25:00'" to obtain the new data. Subsequently, the merge function of the pandas library is used to merge the new data with the existing data set based on the "user_id" field, generating a fused data set containing 10,100 records. The PCA algorithm in the scikit-learn library is applied to the fused data set for dimensionality reduction, setting the n_components parameter to 0.95, reducing the original 50 features to 15 principal components. The new feature vector is input into the pre-trained support vector machine model, and the calculated updated average risk score is 0.72. Based on the new risk score, the stochastic gradient descent algorithm is implemented using SGDClassifier, setting the parameters loss = 'hinge', alpha = 0.0001, max_iter = 1000, and the model is incrementally trained.Meanwhile, according to the change of the risk score, the data encryption algorithm is upgraded from AES-128 to AES-256, and the access control level is adjusted from "Visible Internally" to "Confidential". Finally, the system generates an audit log {operation: "model_update", time: "2024-04-15 15:35:00", risk_score: 0.72, encryption: "AES-256", access_level: "Confidential"} and stores it in the "audit_log" collection of the MongoDB database.

[0035] S106. Based on business requirements and data usage scenarios, formulate data access and sharing policies. For data combinations with low leakage risk but high sensitivity, strictly restrict data access and transfer to ensure data security; for data combinations with high leakage risk but low sensitivity, allow authorized users to use and analyze data within a certain range according to user permissions and access approval rules.

[0036] Obtain information on data leakage risk and sensitivity, and construct a data risk matrix. The data risk matrix includes the horizontal axis representing leakage risk and the vertical axis representing sensitivity; map data combinations to several quadrants according to the data risk matrix, and formulate access and sharing policies for the quadrants. Among them, for the high-risk and low-sensitivity quadrant, adopt a temporary authorization and full-process monitoring strategy; if it is determined that the data combination belongs to the quadrant with low leakage risk but high sensitivity, implement a strong access control mechanism, which includes multi-factor authentication and fine-grained authorization; if it is determined that the data combination belongs to the quadrant with high leakage risk but low sensitivity, design a dynamic authorization mechanism, which includes granting temporary access rights after three-level approvals by the department head, data security officer, and system administrator; deploy an all-round log audit and anomaly monitoring system, which uses the isolation forest algorithm to analyze user behavior patterns in real time to obtain anomaly access detection results.

[0037] Exemplarily, a data risk matrix is constructed. The horizontal axis represents the leakage risk (low, medium, high), and the vertical axis represents the sensitivity level (low, medium, high). Data combinations are mapped to 9 quadrants, and corresponding access and sharing policies are formulated for each quadrant. For example, in the high-risk and low-sensitivity quadrant, a temporary authorization and full-process monitoring strategy is adopted. According to business requirements and data usage scenarios, data access templates are preset for different departments and roles. For example, the marketing department can access basic customer information but not transaction records. For data combinations with low leakage risk but high sensitivity, a strong access control mechanism is implemented, using multi-factor authentication such as fingerprint recognition combined with dynamic passwords and fine-grained authorization to limit the data access scope and usage duration. For data combinations with high leakage risk but low sensitivity, a dynamic authorization mechanism is designed. Based on user roles, business requirements, and access history, an access application workflow is automatically generated and granted temporary access rights after three-level approvals by the department head, data security officer, and system administrator. A full-range log audit and anomaly monitoring system is deployed to record all data access and operation behaviors. The isolation forest algorithm is used to analyze user behavior patterns in real time to detect abnormal access and potential data abuse situations. When abnormal behaviors are detected, an alarm is automatically triggered and the access rights of relevant users are suspended. In the actual implementation process, first, a 3x3 data risk matrix is constructed, and 100 data combinations are mapped to 9 quadrants. Among them, in the high-risk and low-sensitivity quadrant, for example, customer browsing records adopt a 24-hour temporary authorization and full-process behavior monitoring strategy. Subsequently, according to the business requirements of 5 core departments of the company, 15 data access templates are preset. For example, the marketing department can access basic customer information but is restricted from viewing transaction records in the most recent 3 months. For data combinations with a risk score lower than 0.3 but a sensitivity level higher than 0.8, such as employee salary information, a dual authentication mechanism is implemented, requiring users to provide fingerprint recognition and a 6-digit dynamic password, and limiting the access duration to within 30 minutes. For data combinations with a risk score higher than 0.7 but a sensitivity level lower than 0.4, such as anonymized user behavior data, a dynamic authorization process is designed. After the user submits an application, a workflow containing data description, usage purpose, and time range is automatically generated and undergoes three-level approvals in sequence: the department head approves within 4 hours, the data security officer approves within 12 hours, and the system administrator approves within 24 hours. After passing, a 72-hour temporary access right is granted. Finally, a log audit system based on the ELK (Elasticsearch, Logstash, Kibana) stack is deployed, processing 1000 access logs per second. The isolation forest algorithm (with a contamination rate set to 0.1) is used to analyze user behavior in real time. When abnormal behaviors are detected, such as a large amount of data being downloaded in a short period of time, the system automatically triggers an alarm within 5 seconds and immediately suspends the access rights of relevant users.

[0038] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present application. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A method for building a data leakage prevention system, characterized in that: The method comprises: Obtain multiple data sets from different data sources, perform data preprocessing on each data set, use data mining methods to extract features, obtain sensitivity features of multiple data sets, and screen out several key features that have a significant impact on data leakage risks based on preset sensitivity assessment criteria; Aiming at several key characteristics, a sensitivity fusion model of multi-source heterogeneous data is constructed. By setting weight parameters, the key features of different data sources are weighted and fused to obtain the fused target sensitivity feature set. The principal component analysis method is used to reduce the dimension of the target sensitivity feature set and extract the principal component features that affect the contribution of data leakage risk. According to the main component characteristics, a data leakage risk prediction model is constructed, and the support vector machine algorithm is used to train the model and obtain the data leakage risk prediction model. The multi-source heterogeneous data to be predicted is input into the risk prediction model to calculate the leakage risk score of the data combination; Preset a sensitivity threshold to determine whether the leakage risk score exceeds the preset sensitivity threshold. If it exceeds the threshold, the data combination is marked as highly sensitive data, triggering an early warning mechanism to desensitize and encrypt the data combination according to the preset desensitization rules and encryption methods; Continuously monitor the data updates of different data sources. When new data is detected, automatically trigger the data fusion and leakage risk prediction process and obtain the sensitivity feature representation of the new data, update the data leakage risk prediction model, and dynamically adjust the sensitivity and protection measures of the data combination; Based on business needs and data usage scenarios, formulate data access and sharing strategies. For data combinations with low leakage risk but high sensitivity, strictly limit data access and circulation to ensure data security. For data combinations with high leakage risk but low sensitivity, authorized users are allowed to use and analyze data within a certain scope based on user permissions and access approval rules.

2. The method according to claim 1, characterized in that The method obtains multiple data sets from different data sources, performs data preprocessing on each data set, and uses data mining methods to extract features to obtain sensitivity features of multiple data sets. According to preset sensitivity evaluation criteria, several key features that have a significant impact on data leakage risks are screened out, including: Acquire data from relational databases using data source interface adapters, and read data files on file servers using FTP protocol; Execute the data extraction, transformation and loading process to obtain the original data set; Perform data cleaning, standardization, and structural transformation on the original data set to obtain the preprocessed data set; Reduce the dimension of the preprocessed data set to extract the main features, and calculate the impact of each feature on the risk of data leakage; If the information gain ratio exceeds the preset threshold, the feature is included in the candidate feature set; A feature importance evaluation model is constructed based on the candidate feature set, and the features with the lowest importance are iteratively eliminated through a recursive feature elimination method until the number of remaining features reaches the preset target; Normalize the remaining features to obtain a standardized key feature data set; Use the Apriori algorithm to mine association rules on the standardized key feature data set and calculate the support and confidence between feature pairs; Draw a feature association network diagram based on support and confidence, and use the Dijkstra shortest path algorithm to identify the risk propagation path; It also includes: analyzing the data type, structure and content of each data set, cross-comparing the fields and information in different data sets, identifying associated data items, using association rule mining or graphical analysis data association analysis methods to find potential associations between different data sets, evaluating new information or inferences generated after combining data sets, and identifying potential risk points, including personal privacy leakage and sensitive information inference.

3. The method according to claim 2, characterized in that The data type, structure and content of each data set are analyzed, the fields and information in different data sets are cross-compared, the associated data items are identified, the potential associations between different data sets are found by using association rule mining or graphical analysis data association analysis methods, the new information or inferences generated after the combination of data sets are evaluated, and potential risk points are identified. Potential risk points include personal privacy leakage and sensitive information inference, including: Obtain metadata information of each data set, and extract data type, structure, and field information from the metadata information; Calculate the Jaccard similarity of field names based on field information, and calculate the Pearson correlation coefficient of data features for numeric fields; If the Jaccard similarity or the Pearson correlation coefficient is greater than the preset similarity threshold, the corresponding fields are determined to be potential associated field pairs; Apply Apriori algorithm to mine frequent itemsets between data sets; Calculate the support and confidence of the association rules for the frequent item sets. If the support is greater than the preset minimum support and the confidence is greater than the preset confidence threshold, it is determined to be a strongly associated data item combination. The data association network is constructed by combining strongly associated data items, and the Louvain algorithm is used to detect the closely associated data subsets in the data association network; Based on the data association network, predict the risk of information leakage and obtain the risk score; If the risk score is higher than the preset risk threshold, the corresponding node will be marked as a high-risk node.

4. The method according to claim 1, characterized in that: According to the above, a sensitivity fusion model of multi-source heterogeneous data is constructed for several key characteristics. By setting weight parameters, the key features of different data sources are weighted and fused to obtain a fused target sensitivity feature set. The principal component analysis method is used to reduce the dimension of the target sensitivity feature set and extract the principal component features that affect the contribution of data leakage risk, including: Obtain preset scoring criteria for multiple heterogeneous data sources, including data update frequency, data completeness, and outlier ratio; Calculate the reliability score of each data source according to the preset scoring criteria to obtain a data source weight vector; The random forest algorithm is used to calculate the importance scores of key features in each data source in combination with the data source reliability scores, and a weighted feature fusion matrix is ​​constructed by multiplying the data source weight and the feature importance score. The weighted feature fusion matrix is ​​used to perform weighted sum operation on the key features of different data sources. If there are features with the same name, the weighted average method is used for fusion to obtain the fused target sensitivity feature set. Perform Z-score normalization on the fused target sensitivity feature set to eliminate the dimensional differences between different features; For the standardized target sensitivity feature set, the feature covariance matrix is ​​calculated and the eigenvalues ​​and eigenvectors are solved. The first several principal components whose cumulative contribution rate reaches the preset threshold are selected and arranged in descending order according to the eigenvalue size to obtain the principal component feature set after dimensionality reduction.

5. The method according to claim 1, characterized in that The data leakage risk prediction model is constructed according to the principal component characteristics, the support vector machine algorithm is used to train the model and obtain the data leakage risk prediction model, the multi-source heterogeneous data to be predicted is input into the risk prediction model, and the leakage risk score of the data combination is calculated, including: Obtain samples of known risk levels from historical sample databases and log files, and perform missing value filling and outlier detection processing on sample data; According to the processed sample data, the principal component characteristics are determined as input variables and the risk level is determined as the output variable to obtain a standardized training data matrix; The kernel function type, penalty parameter C and kernel function parameter γ of the support vector machine are optimized by using a grid search method. If the search range of the penalty parameter C is from a preset first value to a second value, and the search range of the kernel function parameter γ is from a preset third value to a fourth value, the optimal parameter combination is determined by cross-validation. According to the optimal parameter combination, the support vector machine model is trained, and the sequential minimum optimization algorithm is used to solve the dual problem to obtain the support vector and decision function; For the multi-source heterogeneous data to be predicted, the principal component feature representation in the same format as the training data is obtained through feature extraction and principal component transformation. The principal component feature representation is input into the trained support vector machine model to determine the leakage risk score of the data combination.

6. The method according to claim 1, characterized in that The preset sensitivity threshold determines whether the leakage risk score exceeds the preset sensitivity threshold. If it exceeds the threshold, the data combination is marked as highly sensitive data, triggering the early warning mechanism, and desensitizing and encrypting the data combination according to the preset desensitization rules and encryption methods, including: Obtaining the leakage risk score of the data combination to be evaluated from the database, and determining whether the risk score exceeds the sensitivity threshold; If the risk score exceeds the sensitivity threshold, the data combination is marked as highly sensitive data, and an early warning message containing the data ID, risk score, and trigger time is generated; According to the sensitive information type and highly sensitive data mark in the data combination, select the corresponding desensitization algorithm from the preset desensitization rule library to desensitize the data combination; For the data combination after desensitization processing, according to the data sensitivity and processing performance requirements, an encryption algorithm is selected from AES and RSA or homomorphic encryption algorithm to perform encryption processing to obtain the encrypted data combination; The encrypted data combination is stored in the security audit database, and the information of the desensitization and encryption processing is recorded.

7. The method according to claim 1, characterized in that The continuous monitoring of data updates from different data sources, when new data is detected, automatically triggers the data fusion and leakage risk prediction process and obtains the sensitivity feature representation of the new data, updates the data leakage risk prediction model, and dynamically adjusts the sensitivity of the data combination and the protection measures, including: Receive update status information sent by the data source, the update status information includes the timestamp of the newly added data and the data source identifier; The data fusion module is called according to the update status information. The data fusion module adopts an incremental update strategy based on timestamp comparison to integrate the newly added data with the existing data. Extract features from the fused data set generated by the data fusion module, use the principal component analysis method to reduce the dimension of the features, and obtain a new sensitivity feature representation; Input the sensitivity feature representation into the data leakage risk prediction model to obtain an updated risk score; If there is a deviation between the updated risk score and the preset threshold, the stochastic gradient descent algorithm is used to perform online learning and incremental training on the data leakage risk prediction model to update the parameters of the data leakage risk prediction model; Adjust the protection measures for the corresponding data portfolio based on the updated risk score, including encryption strength or access control level; The parameter updates and protection measures adjustment information of the data leakage risk prediction model are recorded in the security audit database.

8. The method according to claim 1, characterized in that Based on business needs and data usage scenarios, data access and sharing strategies are formulated. For data combinations with low leakage risk but high sensitivity, data access and circulation are strictly restricted to ensure data security. For data combinations with high leakage risk but low sensitivity, authorized users are allowed to use and analyze data within a certain scope based on user permissions and access approval rules, including: Obtain data leakage risk and sensitivity information and construct a data risk matrix. The data risk matrix includes a horizontal axis representing leakage risk and a vertical axis representing sensitivity; Mapping data combinations to several quadrants according to the data risk matrix, defining access and sharing strategies for the quadrants, wherein a temporary authorization and full-process monitoring strategy is adopted for the high-risk and low-sensitivity quadrant; If the data combination is judged to be in the quadrant of low leakage risk but high sensitivity, a strong access control mechanism is implemented, which includes multi-factor authentication and fine-grained authorization; If the data combination is judged to belong to the quadrant with high leakage risk but low sensitivity, a dynamic authorization mechanism is designed. The dynamic authorization mechanism includes granting temporary access rights after approval by the department head, data security officer and system administrator at three levels; Deploy a comprehensive log audit and anomaly monitoring system. The comprehensive log audit and anomaly monitoring system uses the isolation forest algorithm to analyze user behavior patterns in real time and obtain abnormal access detection results.

Citation Information

Patent Citations

  • Data sensitivity identification method and apparatus

    CN107944283A

  • Government affair data authority management method and system based on big data analysis

    CN118396370A