Data leakage risk assessment method and system based on artificial intelligence
The AI-driven data leak risk assessment method automates and enhances the efficiency and accuracy of data security by processing database and user behavior logs to build a neural network model for real-time risk detection, addressing inefficiencies in manual analysis and improving data breach prevention.
Patent Information
- Application Number
- CN202510489111.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
AI Technical Summary
Existing data breach risk assessment methods rely on manual analysis, are inefficient and difficult to cope with complex and changeable network environments, and traditional methods are difficult to improve the efficiency and accuracy of data security protection.
Using artificial intelligence-based data leakage risk assessment method, we collect and process database operation logs and user behavior logs, crawl external intelligence, extract key features, build a neural network model for risk assessment, and generate a risk assessment report.
It realizes automated data leakage risk assessment, improves efficiency and accuracy, can monitor data security status in real time, provide decision support, and optimize data security management processes.
Smart Images

Figure CN120316802A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data security, and particularly relates to a method and system for evaluating data leakage risk based on artificial intelligence. Background Art
[0002] With the rapid development of information technology, data security issues have become increasingly prominent. In the existing data security protection system, although various security measures have been taken, data leakage incidents still occur frequently. Traditional methods for evaluating data leakage risk rely on manual analysis, which is inefficient and difficult to cope with complex and changing network environments. Therefore, an intelligent and automated method for evaluating data leakage risk is needed to improve the efficiency and accuracy of data security protection. Summary of the Invention
[0003] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:
[0004] A method for evaluating data leakage risk based on artificial intelligence, comprising the following steps:
[0005] Collect database operation logs and user behavior logs and extract abnormal behavior information;
[0006] Crawl external intelligence on data leakage in the abnormal behavior information and aggregate the external intelligence to obtain aggregated data;
[0007] Perform data processing on the aggregated data and extract key features of the aggregated data;
[0008] Build and train a data leakage risk assessment model based on the key features;
[0009] Collect real-time database operation information and real-time user behavior information, and perform leakage risk assessment based on the data leakage risk assessment model.
[0010] Preferably, the method for extracting the abnormal behavior information includes:
[0011] Perform format standardization processing on the database operation logs and the user behavior logs to obtain standardized logs;
[0012] Use information extraction scripts to extract key information from the standardized logs to obtain the abnormal behavior information; wherein, the abnormal behavior information includes: operation type, operation time, operation object, and user.
[0013] Preferably, the method for obtaining the aggregated data includes:
[0014] Design an intelligent crawler to traverse the abnormal behavior information using the intelligent crawler and crawl the external intelligence of data leakage in the abnormal behavior information;
[0015] Parse and deduplicate the external intelligence to obtain the deduplicated intelligence information;
[0016] Aggregate the deduplicated intelligence information using a data aggregation method to obtain the aggregated data.
[0017] Preferably, the method for extracting the key features includes:
[0018] Perform data cleaning on the aggregated data to remove irrelevant data and obtain the cleaned data;
[0019] Extract key features from the cleaned data using a feature selection algorithm, where the key features include: database data features, user behavior features, and traffic network features.
[0020] Preferably, the method for feature extraction includes:
[0021] Perform standardization processing on the cleaned data to obtain the standardized data:
[0022]
[0023] where x std represents the standardized data, x represents the cleaned data, μ represents the mean, and σ represents the standard deviation;
[0024] Calculate the covariance matrix of the standardized data:
[0025]
[0026] where C ij represents the covariance matrix, n represents the sample size of the standardized data, i, j, k represent natural numbers, x ki represents the i-th eigenvalue of the k-th sample, represents the mean of the i-th feature, x kj represents the j-th eigenvalue of the k-th sample, represents the mean of the j-th feature;
[0027] Solve the eigenvalues and eigenvectors of the covariance matrix:
[0028] Cv = λv where λ represents the eigenvalue and v represents the eigenvector;
[0029] Based on the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top m largest eigenvalues to construct a new feature space;
[0030] Project the cleaned data onto the selected principal components in the feature space to obtain a feature representation after dimensionality reduction, which is the key feature:
[0031] y = xW
[0032] Where y represents the key feature and W represents a matrix composed of feature vectors.
[0033] Preferably, the method for obtaining the data leakage risk assessment model includes:
[0034] Construct a neural network model, which includes an input layer, a hidden layer, and an output layer. The input layer is used to receive and process the key feature; the hidden layer contains multiple neurons and is used to perform non-linear transformation and feature extraction on the key feature; the output layer is used to output the result of data leakage risk assessment.
[0035] Train the neural network model based on the key feature to obtain a data leakage risk assessment model.
[0036] Preferably, the method for performing the leakage risk assessment includes:
[0037] Use the data leakage risk assessment model to analyze the real-time database operation information and the real-time user behavior information to obtain a risk assessment result;
[0038] Select risk factors based on the risk assessment result and calculate the total risk assessment value:
[0039]
[0040] Where TotalRisk represents the total risk assessment value, w i represents the weight of each risk factor, r i represents the risk value of each risk factor, and m represents the number of risk factors.
[0041] The present invention also provides an artificial intelligence-based data leakage risk assessment system. The assessment system applies the assessment method described in any one of the above, and includes: a log extraction module, a data aggregation module, a feature extraction module, a model construction module, and a risk assessment module;
[0042] The log extraction module is used to collect database operation logs and user behavior logs and extract abnormal behavior information;
[0043] The data aggregation module is used to crawl external intelligence on data leakage in the abnormal behavior information and aggregate the external intelligence to obtain aggregated data;
[0044] The feature extraction module is used to process the aggregated data and extract the key features of the aggregated data;
[0045] The model construction module constructs and trains a data leakage risk assessment model based on the key features;
[0046] The risk assessment module is used to collect real-time database operation information and real-time user behavior information, and perform leakage risk assessment based on the data leakage risk assessment model.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] Through automated data collection and processing, the present invention reduces manual intervention and significantly improves the efficiency of data leakage risk assessment; by using artificial intelligence technology to deeply analyze data features, the accuracy and reliability of risk assessment are improved; the present invention can monitor the data security status in real time, timely detect and respond to potential data leakage risks, and the generated risk assessment reports and countermeasures provide decision-making support for security managers, helping to optimize the data security management process. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0053] Embodiment 1
[0054] In this embodiment, as Figure 1 shown, a data leakage risk assessment method based on artificial intelligence includes the following steps:
[0055] S1. Collect database operation logs and user behavior logs and extract abnormal behavior information.
[0056] The methods for extracting abnormal behavior information include: performing format standardization processing on database operation logs and user behavior logs to obtain standardized logs; using information extraction scripts to extract key information from the standardized logs to obtain abnormal behavior information; where the abnormal behavior information includes: operation type, operation time, operation object, and user.
[0057] In this embodiment, it is mainly to collect database and user behavior logs. Standardization processing ensures unified log formats: First, use regular expressions to match and extract log data, remove irrelevant characters and redundant information, and retain key fields. Then, convert the time fields in the logs to the unified ISO 8601 standard time format for subsequent time series analysis and processing. Then, according to the format requirements, convert the fields in the log data to the corresponding data types. Finally, clean the log data to remove duplicate data, missing data, and abnormal data to ensure the integrity and accuracy of the data, obtaining standardized logs. Then use scripts to extract key information, including operation type, time, object, and user, and pay special attention to identifying abnormal behaviors. This step uses advanced log analysis techniques to automatically identify and mark abnormal behaviors, thus providing basic data support for subsequent data leakage risk assessment. Specifically, establish a normal behavior pattern library, and then compare the current logs to intelligently identify behaviors that deviate from the normal pattern and mark them as abnormal; the methods for establishing a normal behavior pattern library include: (1) Data representation and preprocessing: ① For time information time: Convert the timestamp to discrete time intervals, divide a day into 24 hours, map the timestamp to the corresponding hour number, and obtain an integer in the range of 0 to 23; ② For operation type operation type and object information object: Perform simple integer encoding on the operation type and object. Assume there are n operation types, map each operation type to a unique integer number i (i = 0, 1, 2,..., n - 1). For the object, if there are m types, mark it as the integer number j (j = 0, 1, 2,..., m - 1); number the users. If there are k users, each user is mapped to a unique integer number u (u = 0, 1, 2,..., k - 1); ③ For regional user information user: For each user behavior log entry, combine the processed time, operation type, object, and user information into a feature vector: z = [time, operation type, object, user]. (2) Use the K-means clustering algorithm to cluster the feature vectors of normal user behavior: First, collect a large amount of normal user behavior data and represent it in the form of the above-mentioned feature vectors. Select an appropriate number of clusters K according to the amount of data. The number is determined by methods such as the elbow method. Apply the K-means algorithm to cluster the normal behavior data:
[0058] G = K-means({z1, z2,..., z N},K)
[0059] where {z1, z2,..., z N} is the set of feature vectors of normal behavior data, and G is the clustering result, which contains K cluster centers and the cluster labels to which each data point belongs. (3) Establish a normal behavior pattern library: The cluster centers serve as the basis of the normal behavior pattern library. Each cluster center g k (k = 1, 2,..., K) = is a feature vector that represents the typical pattern of a group of similar user behaviors. The methods for identifying abnormal behaviors include: For new user behavior log data, convert it into a feature vector z new . Calculate the distance between this feature vector and each cluster center, using the Euclidean distance:
[0060]
[0061] where d represents the dimension of the feature vector, d = 4. Set a threshold θ. If the distance from the new data point z new to the nearest cluster center exceeds this threshold, mark it as an abnormal behavior; otherwise, consider it a normal behavior.
[0062] S2. Crawl the external intelligence of data leakage in the abnormal behavior information and aggregate the external intelligence to obtain the aggregated data.
[0063] The methods for obtaining the aggregated data include: Design an intelligent crawler, use the intelligent crawler to traverse the abnormal behavior information, and crawl the external intelligence of data leakage in the abnormal behavior information; Parse and deduplicate the external intelligence to obtain the deduplicated intelligence information; Use data aggregation methods to aggregate the deduplicated intelligence information to obtain the aggregated data.
[0064] In this embodiment, the intelligent crawler adopts intelligent path planning based on the ant colony algorithm. It can automatically adjust and optimize the crawling strategy according to the characteristics and requirements of abnormal behavior information, automatically identify and prioritize the crawling of intelligence closely related to the risk of data leakage, and at the same time avoid falling into the trap of irrelevant information, significantly improving the efficiency and quality of data collection. The methods for parsing and deduplication include: First, the external intelligence collected by the intelligent crawler is parsed to extract valuable information; then, the parsed intelligence is compared one by one using a similarity calculation deduplication algorithm to identify and eliminate duplicate information, ensuring that the aggregated dataset does not contain any duplicate content, thereby improving the efficiency and accuracy of subsequent data processing and analysis. The methods for data aggregation include: First, the deduplicated intelligence information is classified and sorted, and it is classified into different categories according to characteristics such as information source, type, and content; then, for each category, clustering analysis is used to effectively integrate similar information to form a more refined and representative data set, that is, the aggregated data, for subsequent data processing and feature extraction.
[0065] S3. Perform data processing on the aggregated data and extract the key features of the aggregated data.
[0066] The methods for extracting key features include: performing data cleaning on the aggregated data to remove irrelevant data and obtain the cleaned data; using a feature selection algorithm to extract features from the cleaned data to obtain the key features; among them, the key features include: database data features, user behavior features, and traffic network features.
[0067] In this embodiment, this step uses data cleaning technology and feature selection algorithms to ensure the quality and accuracy of the data. Feature extraction is one of the key steps in data leakage risk assessment, providing high-quality input for the risk assessment model.
[0068] The methods for data cleaning include: First, identify and remove data items irrelevant to data leakage risk assessment to reduce data redundancy; second, check and correct errors or outliers in the data, such as format errors and logical contradictions, to improve data quality; finally, handle missing data, such as through interpolation, filling, or deletion, to ensure data integrity.
[0069] The methods for feature extraction include: performing standardization processing on the cleaned data to eliminate the influence of different feature dimensions and obtain the standardized data:
[0070]
[0071] where x std represents the standardized data, x represents the cleaned data, μ represents the mean, σ represents the standard deviation; calculate the covariance matrix of the standardized data:
[0072]
[0073] Among them, C ij represents the covariance matrix, n represents the number of samples of the standardized data, i, j, k represent natural numbers, and x ki represents the i-th eigenvalue of the k-th sample, represents the mean value of the i-th feature, and x kj represents the j-th eigenvalue of the k-th sample, represents the mean value of the j-th feature; Solve the eigenvalues and eigenvectors of the covariance matrix:
[0074] Cv = λv
[0075] Among them, λ represents the eigenvalue, and v represents the eigenvector; Based on the magnitude of the eigenvalues, select the eigenvectors corresponding to the top m largest eigenvalues to construct a new feature space; Project the cleaned data onto the selected principal components in the feature space to obtain the dimensionality-reduced feature representation, that is, the key features:
[0076] y = xW. Among them, y represents the key features, and W represents a matrix composed of eigenvectors, which is used to achieve data dimensionality reduction and feature extraction. The columns of this matrix are the selected principal components (eigenvectors), and the number of rows is the same as the features of the original data. By multiplying the original data by this matrix "W", the dimensionality-reduced data (i.e., the key features "y") can be obtained.
[0077] S4. Build and train a data leakage risk assessment model based on the key features.
[0078] The method for obtaining the data leakage risk assessment model includes:
[0079] Build a neural network model. The neural network model includes: an input layer, a hidden layer, and an output layer. The input layer is used to receive and process the key features; the hidden layer contains multiple neurons, which are used to perform non-linear transformation and feature extraction on the key features; the output layer is used to output the results of the data leakage risk assessment;
[0080] In this embodiment, among them, the input layer is responsible for receiving and processing the key features; the hidden layer contains multiple neurons, which are used to perform non-linear transformation and feature extraction on the key features, and is the key part for the model to learn and express complex relationships; the output layer is used to output the results of the data leakage risk assessment, and the result is a probability value or a risk level, which is used to represent the possibility or risk degree of data leakage; Train the neural network model based on the key features to obtain the data leakage risk assessment model.
[0081] In this embodiment, the process of model training includes: First, divide the dataset of key features to obtain the training set data and the test dataset to ensure the robustness and accuracy of model training; Then, use the training set data to iteratively train the model, continuously adjust the model parameters to optimize the prediction performance of the model. At the same time, monitor the overfitting phenomenon by evaluating the performance of the model on the test dataset and make timely model adjustments. Finally, when the model meets the set evaluation criteria on the test dataset, it can be considered that the training is completed.
[0082] In this embodiment, a verification process for the model is also provided: First, verify the preliminarily trained model, and use the independent validation set data to evaluate the performance of the model, mainly including accuracy, recall rate, and F1-score metrics. Accuracy is an indicator to evaluate whether the prediction results of the model are correct, and the calculation formula is:
[0083]
[0084] The recall rate, also known as the recall ratio, is an indicator to evaluate the ability of the model to correctly identify positive samples, and the calculation formula is:
[0085]
[0086] The higher the recall rate, the more positive samples the model can identify, but it may also lead to more negative samples being misjudged as positive samples; The F1-score is the harmonic mean of accuracy and recall rate, which is used to comprehensively consider the accuracy and recall ability of the model, and the calculation formula is:
[0087]
[0088] Among them, Recall is the recall rate, which refers to the proportion of all true positive samples that are correctly predicted as positive samples by the model; Precision (precision rate) is the proportion of samples that are truly positive among the samples predicted as positive by the model, and the calculation formula is:
[0089]
[0090] Among them, TP (True Positive) represents the number of samples correctly predicted as positive, TN (True Negative) represents the number of samples correctly predicted as negative, FP (False Positive) represents the number of samples wrongly predicted as positive, and FN (False Negative) represents the number of samples wrongly predicted as negative.
[0091] By comparing and analyzing the differences between the model prediction results and the actual labels, potential biases or deficiencies in the model can be identified. Based on the verification results, improve the feature engineering method to extract more valuable features. After each optimization, re - verify to evaluate the improvement effect, and continuously iterate this process until the model performance reaches a satisfactory level. Finally, through continuous optimization and verification, ensure that the AI model has efficient and accurate capabilities in data leakage risk assessment.
[0092] S5. Collect real - time database operation information and real - time user behavior information, and conduct leakage risk assessment based on the data leakage risk assessment model.
[0093] The method for conducting leakage risk assessment includes: using the data leakage risk assessment model to analyze the real - time database operation information and real - time user behavior information to obtain the risk assessment results; selecting risk factors based on the risk assessment results, and calculating the total risk assessment value:
[0094]
[0095] where TotalRisk represents the total risk assessment value, w i represents the weight of each risk factor, r i represents the risk value of each risk factor, and m represents the number of risk factors.
[0096] In this embodiment, TotalRisk is a measure of the total risk, representing the overall risk level calculated based on each risk characteristic, its corresponding weight and risk value. w i This represents the weight of each risk factor, and the weight reflects the importance or influence degree of different risk factors in the overall risk assessment. r i This represents the risk value of each risk factor. The risk value can be quantitative (such as specific loss amounts, probability values, etc.) or qualitative (such as high, medium, low risk levels, etc.). C is a constant term, representing some fixed risk factors and the basic risk level of the system, which is determined by the system administrator. For each risk factor, calculate the product of its weight (w i ) and risk value (r i ), and sum these products and add the basic risk to obtain the total risk (TotalRisk).
[0097] In this embodiment, use the total risk assessment value to generate a risk assessment report, and feedback the risk assessment report and risk response strategies to the security management personnel to assist them in taking security measures, and continuously optimize the risk assessment model according to the feedback of the security management personnel. This step enables the security management personnel to take actions based on the risk assessment report, and at the same time the system adjusts the risk assessment model according to the effects of these actions, forming a closed - loop management.
[0098] Embodiment 2
[0099] In this embodiment, an artificial intelligence-based data leakage risk assessment system includes: a log extraction module, a data aggregation module, a feature extraction module, a model construction module, and a risk assessment module;
[0100] The log extraction module is used to collect database operation logs and user behavior logs and extract abnormal behavior information. The data aggregation module is used to crawl external intelligence on data leakage in the abnormal behavior information and aggregate the external intelligence to obtain aggregated data. The feature extraction module is used to process the aggregated data and extract key features of the aggregated data. The model construction module constructs and trains a data leakage risk assessment model based on the key features. The risk assessment module is used to collect real-time database operation information and real-time user behavior information and perform leakage risk assessment based on the data leakage risk assessment model.
[0101] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for assessing data leakage risk based on artificial intelligence, characterized in that, It includes the following steps: Collect database operation logs and user behavior logs and extract abnormal behavior information; Crawl the external intelligence of data leakage in the abnormal behavior information and aggregate the external intelligence to obtain aggregated data; Perform data processing on the aggregated data and extract the key features of the aggregated data; Construct and train a data leakage risk assessment model based on the key features; Collect real-time database operation information and real-time user behavior information and conduct leakage risk assessment based on the data leakage risk assessment model.
2. The method for evaluating the risk of data leakage based on artificial intelligence according to claim 1, characterized in that, The method for extracting the abnormal behavior information includes: Perform format standardization processing on the database operation logs and the user behavior logs to obtain standardized logs; Use an information extraction script to extract key information from the standardized logs to obtain the abnormal behavior information; wherein, the abnormal behavior information includes: operation type, operation time, operation object, and user.
3. The method for evaluating the risk of data leakage based on artificial intelligence according to claim 1, wherein The method for obtaining the aggregated data includes: Design an intelligent crawler, use the intelligent crawler to traverse the abnormal behavior information, and crawl the external intelligence of data leakage in the abnormal behavior information; Parse and de-duplicate the external intelligence to obtain de-duplicated intelligence information; Use a data aggregation method to aggregate the de-duplicated intelligence information to obtain the aggregated data.
4. The method for evaluating data leakage risk based on artificial intelligence according to claim 1, characterized in that The method for extracting the key features includes: Perform data cleaning on the aggregated data to remove irrelevant data and obtain cleaned data; Use a feature selection algorithm to extract features from the cleaned data to obtain the key features; wherein, the key features include: database data features, user behavior features, and traffic network features.
5. The method for evaluating the risk of data leakage based on artificial intelligence according to claim 4, wherein The method for feature extraction includes: Perform standardization processing on the cleaned data to obtain standardized data: where x std represents the standardized data, x represents the cleaned data, μ represents the mean, and σ represents the standard deviation; Calculate the covariance matrix of the standardized data: Among them, C ij represents the covariance matrix, n represents the number of samples of the standardized data, i, j, k represent natural numbers, and x ki represents the i-th eigenvalue of the k-th sample, represents the mean of the i-th feature, and x kj represents the j-th eigenvalue of the k-th sample, represents the mean of the j-th feature; Solve the eigenvalues and eigenvectors of the covariance matrix: Cv = λv where λ represents the eigenvalue and v represents the eigenvector; Based on the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top m largest eigenvalues to construct a new feature space; Project the cleaned data onto the selected principal components in the feature space to obtain a feature representation after dimensionality reduction, which is the key feature: y = xW where y represents the key feature and W represents the matrix composed of eigenvectors.
6. The method for evaluating data leakage risk based on artificial intelligence according to claim 1, wherein The method for obtaining a data leakage risk assessment model includes: Construct a neural network model, the neural network model includes: an input layer, a hidden layer, and an output layer, the input layer is used to receive and process the key features; the hidden layer contains multiple neurons and is used to perform non-linear transformation and feature extraction on the key features; the output layer is used to output the results of data leakage risk assessment; Train the neural network model based on the key features to obtain a data leakage risk assessment model.
7. The method for evaluating the risk of data leakage based on artificial intelligence according to claim 1, wherein The method for performing the leakage risk assessment includes: Use the data leakage risk assessment model to analyze the real-time database operation information and the real-time user behavior information to obtain a risk assessment result; Select risk factors based on the risk assessment result and calculate the total risk assessment value: Among them, TotalRisk represents the total risk assessment value, w i represents the weight of each risk factor, r i represents the risk value of each risk factor, and m represents the number of risk factors.
8. An artificial intelligence-based data leakage risk assessment system, wherein the assessment system applies the assessment method according to any one of claims 1-7, characterized in that Include: A log extraction module, a data aggregation module, a feature extraction module, a model construction module, and a risk assessment module; The log extraction module is used to collect database operation logs and user behavior logs and extract abnormal behavior information; The data aggregation module is used to crawl external intelligence on data leakage in the abnormal behavior information and aggregate the external intelligence to obtain aggregated data; The feature extraction module is used to perform data processing on the aggregated data and extract key features of the aggregated data; The model construction module constructs and trains a data leakage risk assessment model based on the key features; The risk assessment module is used to collect real-time database operation information and real-time user behavior information and perform leakage risk assessment based on the data leakage risk assessment model.
Citation Information
Patent Citations
Terminal data leakage event prediction method and system based on transverse federated learning
CN118468988A
Industrial chain risk information monitoring method based on network data and artificial intelligence technology
CN118798633A
Network data leakage monitoring system and method
CN119561794A
Method and system for tracing a source of leaked information
WO2010011182A2