Internal threat detection method and system based on user behavior representation technology
By using the user behavior representation technology of the clustering enhancement module and LSTM-VAE model in internal threat detection, the clustering center is dynamically adjusted, and the problems of data imbalance and low accuracy are solved, and more efficient internal threat detection is achieved.
Patent Information
- Application Number
- CN202510488485.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing internal threat detection methods have shortcomings in the problems of serious data imbalance and low accuracy, especially traditional machine learning methods are difficult to effectively model user behavior, and the detection results are not ideal.
The detection method based on user behavior representation technology is adopted, and the user behavior representation model composed of the clustering enhancement module and the LSTM-VAE model is trained using the time series data of normal user behavior characteristics, and the clustering center is dynamically adjusted during the training process, and the abnormal score of the user behavior data to be detected is calculated to judge the threat.
The user behavior representation model is trained through unsupervised learning methods, effectively dealing with data imbalance, improving the accuracy of internal threat detection, and effectively identifying user abnormal behaviors and detecting internal threats.
Smart Images

Figure CN120068090A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security technology, and particularly relates to an internal threat detection method and system based on user behavior characterization technology, which is applicable to the detection of abnormal user behaviors in the scenario of enterprise internal threats. Background Art
[0002] With the rapid development of information technology, especially the wide application of the Internet, cloud computing, big data, and artificial intelligence, the information systems within enterprises have become increasingly complex and large. Traditional security protection measures mainly focus on preventing external attacks. However, with the changes in the network environment and the increasing complexity of enterprise structures, internal threats have gradually become one of the most serious security challenges faced by enterprises.
[0003] Internal threats refer to attacks or improper behaviors initiated by internal personnel within an enterprise. Different from external threats, the characteristics of internal threats are that the attackers have legitimate access rights, usually enter the system for operations with an internal identity, and their activities usually occur during normal working hours. The attackers understand the enterprise's infrastructure and sensitive information, causing serious damage such as leakage, copying, and tampering of enterprise information.
[0004] Currently, the detection methods for internal threats are mainly divided into two types: one is the detection method based on rule-based pattern matching: by pre-defining a set of rules or signatures to match the threat behaviors and attack patterns occurring in the network or system to identify potential security threats. Since internal attackers often have knowledge of the system security mechanism and have a high degree of concealment, this method may have missed alarms and false alarms. Since the rules are defined based on known threat behaviors and attack patterns, this means that this method may not be able to detect new types of attack behaviors. Moreover, with the rapid evolution of user behaviors, this method requires frequent updates and maintenance of the rule library, which is very cumbersome and has limitations in practical applications.
[0005] The other is the detection method based on machine learning technology: using machine learning to learn the characteristics of normal behaviors or abnormal behaviors from a large amount of system historical data for detection. This method can learn effective behavior patterns from the data, avoiding the cumbersome process of manually defining rules. However, in real internal threat detection scenarios, the abnormal user behavior data is far less than the normal behavior data, and there is a serious data imbalance problem. Due to the lack of high-quality labeled user behavior data for training the model, traditional machine learning methods such as hidden Markov models and support vector machines often cannot effectively model user behaviors, and the detection effect is not ideal. Summary of the Invention
[0006] In order to solve the problems of serious data imbalance and low accuracy in the field of internal threat detection. The present invention provides an internal threat detection method and system based on user behavior characterization technology.
[0007] The present invention is achieved through the following technical solutions.
[0008] An internal threat detection method based on user behavior characterization technology, comprising the following steps: S1. Data preprocessing: Aggregate normal user behavior data from different logs according to the unique ID value of the user, extract user behavior features therefrom, and organize them into windowed normal user behavior feature time series data; S2. Construct a user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model, use the normal user behavior feature time series data to train the user behavior characterization model, and dynamically adjust the clustering centers during the training process; S3. Use the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combine the defined threshold to determine whether there is a threat to the user behavior.
[0009] Further preferably, in step S2, first transfer the user behavior features to the clustering enhancement module. The clustering enhancement module uses LSTM as an encoder to transform the user behavior features into latent variables and cluster them to obtain the best initial clustering centers, and then transfer the best initial clustering centers to the LSTM-VAE model. The LSTM-VAE model learns and characterizes the user behavior features, and dynamically adjusts the clustering centers during the training process of the LSTM-VAE model.
[0010] Further preferably, the process of clustering by the clustering enhancement module is as follows: The input data of the clustering enhancement module is user behavior features. The clustering enhancement module includes an LSTM, a linear layer, and a K-Means algorithm clustering module; The LSTM receives the input user behavior features and transforms them into hidden states The hidden state at the last moment of the LSTM passes through the fully connected layer to obtain the mean μ and logarithmic variance of the latent variable and then uses the reparameterization method to obtain the latent variable z; Use the K-Means algorithm to cluster the latent variables, and select the best number of clustering centers K through the silhouette coefficient method best ; Finally, use the best number of clustering centers K best to retrain the K-Means algorithm to obtain the best initial clustering centers.
[0011] Further preferably, the LSTM-VAE model uses a single layer of LSTM to replace the fully connected layer as the encoder and decoder of the variational autoencoder. The input data of the LSTM-VAE model is user behavior features; the encoder maps the input data to the latent variables in the latent space, and the decoder decodes the latent variables back to the input space to generate a reconstructed sequence.
[0012] Further preferably, the LSTM-VAE model is trained using a joint optimization strategy to optimize the reconstruction loss, KL divergence loss, and clustering loss simultaneously; ; where, is the total loss of the LSTM-VAE model, is the reconstruction loss, is the KL divergence loss, is the clustering loss, where and are hyperparameters.
[0013] Further preferably, the way to dynamically adjust the clustering center is: after setting every P batches, recalculate the clustering center based on the current latent variables, where is the momentum coefficient, which is used to control the information retention degree of the clustering center; ; where, is the clustering center before adjustment, is the clustering center after adjustment, where, is the k-th cluster, where is the number of latent variables in the k-th cluster; is the i-th latent variable, is the k-th clustering center.
[0014] Further preferably, the data preprocessing includes the following sub-steps: Step 1.1: Aggregate the normal user data from different logs according to the unique ID value of the user, extract the user's behavior data within a day with a daily time granularity, and then merge and sort it according to time; Step 1.2: Extract the user behavior features from the aggregated data and organize them into normal user behavior feature time series data; Step 1.3: Window the normal user behavior feature time series data.
[0015] The present invention also provides an internal threat detection system based on user behavior characterization technology, including: A data acquisition module for collecting user behavior data; Data preprocessing module: It is used to aggregate the normal user behavior data from different logs according to the unique ID value of the user, extract user behavior features therefrom, and organize them into windowed normal user behavior feature time series data; Threat detection module, which has a user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model built-in. It uses the normal user behavior feature time series data to train the user behavior characterization model and dynamically adjusts the clustering center during the training process; it uses the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combines the defined threshold to determine whether there is a threat in the user behavior.
[0016] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory. When the computer program is executed by the processor, it implements the internal threat detection method based on user behavior characterization technology as described above.
[0017] Advantages of the present invention: The present invention first transmits the user behavior features to the clustering enhancement module. The clustering enhancement module uses LSTM as an encoder to convert the user behavior features into latent variables and cluster them to obtain the best initial clustering center. Then, it transmits the best initial clustering center to the LSTM-VAE model. The LSTM-VAE model learns and characterizes the user behavior features, and dynamically adjusts the clustering center during the training process of the LSTM-VAE model, improving the accuracy of internal threat detection. The present invention uses an unsupervised learning method to train the user behavior characterization model, which can not only effectively cope with the data imbalance problem, but also effectively identify user abnormal behaviors and detect internal threats. Brief Description of the Drawings
[0018] Figure 1 It is a flowchart of the steps of the internal threat detection method based on user behavior characterization technology of the present invention; Figure 2 It is a structure diagram of the clustering enhancement module described in Embodiment 2; Figure 3 It is a structure diagram of the LSTM-VAE model described in Embodiment 2. Detailed Embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application. Embodiment 1
[0020] This embodiment proposes an internal threat detection method based on user behavior characterization technology, aiming to solve the problems of serious data imbalance and low accuracy in the field of internal threat detection. As Figure 1 shown, this method includes the following steps: S1. Data preprocessing: Aggregate the normal user behavior data from different logs according to the unique ID value of the user, extract the user behavior characteristics therefrom, and organize them into windowed normal user behavior characteristic time series data; S2. Construct a user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model, use the normal user behavior characteristic time series data to train the user behavior characterization model, and dynamically adjust the clustering center during the training process; S3. Use the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combine the defined threshold to determine whether the user behavior poses a threat. Embodiment 2
[0021] The following details the internal threat detection method of the user behavior characterization technology based on Embodiment 1. S1. Data preprocessing: Select the CERT dataset to verify the proposed internal threat detection method. This dataset belongs to a multi-source dataset and includes information such as system login and logout logs, device usage logs, email logs, network logs, and human resources. The data preprocessing includes the following steps: Step 1.1: Aggregate the normal user data from different logs according to the unique ID value of the user, extract the user's behavior data within a day with a daily time granularity, and then sort them in time order after merging.
[0022] Step 1.2: Extract the user behavior characteristics from the aggregated data and organize them into normal user behavior characteristic time series data. This embodiment counts 23 user behavior characteristics. Including: the total number of all behaviors, the total number of all behaviors during working hours, the total number of all behaviors during non-working hours, the number of logins, the number of logins during non-working hours, the number of USB device usages, the number of USB device usages during non-working hours, the number of file operations, the number of file operations during non-working hours, the number of picture accesses, the number of picture accesses during non-working hours, the number of document accesses, the number of document accesses during non-working hours, the number of text file accesses, the number of text file accesses during non-working hours, the number of executable file accesses, the number of executable file accesses during non-working hours, the number of emails sent or received, the number of emails sent or received during non-working hours, the number of web page browsings via the HTTP protocol, the number of web page browsings via the HTTP protocol during working hours, the number of web page browsings via the HTTP protocol during non-working hours.
[0023] Step 1.3: Window the time series data of normal user behavior characteristics. In this embodiment, a fixed-length sliding window method is adopted, and the window length is set to 5. Each window can be used as a training sample for the model, which contains the user behavior characteristics of the user within 5 days.
[0024] S2. Construct a user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model, use the time series data of normal user behavior characteristics to train the user behavior characterization model, and dynamically adjust the clustering center during the training process; First, pass the user behavior characteristics to the clustering enhancement module. The clustering enhancement module uses LSTM as an encoder to transform the user behavior characteristics into latent variables and cluster them to obtain an initial clustering center. Then, pass the initial clustering center to the LSTM-VAE model. The LSTM-VAE model learns and characterizes the user behavior characteristics, and dynamically adjusts the clustering center during the training process of the LSTM-VAE model. The step S2 includes the following steps: Step 2.1: The clustering enhancement module uses LSTM as an encoder to transform the user behavior characteristics into latent variables and cluster them to obtain the best initial clustering center. The input data for this step is the user behavior characteristics , is the user behavior characteristic at time t, and T is the total time length. The clustering enhancement module is as Figure 2 shown, including an LSTM, a linear layer, and a K-Means algorithm clustering module.
[0025] The task of the LSTM is to encode the user behavior characteristics and transform them into latent variables in the latent space of the variational autoencoder (VAE). The LSTM receives the input user behavior characteristics and transforms them into a hidden state , and the calculation process is shown in the following formula: (1); where, is the hidden state at time t-1, is the hidden state at time t.
[0026] The hidden state at the last moment of the LSTM will pass through a fully connected layer to obtain the mean μ and logarithmic variance of the latent variable, and the calculation formulas are shown in (2) and (3): (2); (3); where, is the trainable bias vector of the mean generation layer, is the trainable bias vector of the variance generation layer, is the trainable weight matrix of the mean generation layer, is the trainable weight matrix of the variance generation layer; The latent variable z is obtained using the reparameterization method, and the calculation formula is as shown in (4): (4); where, is the noise sampled from the standard normal distribution; Use the K-Means algorithm to cluster the latent variables z = {z 1 , z 2 , …, , …, z n}, is the i-th latent variable, i = {1, 2, …, n}, n is the number of latent variables, and the optimal number of cluster centers K is selected by the silhouette coefficient method best .
[0027] Select K data points as the initial cluster centers, and assign each latent variable to the nearest initial cluster center, that is, calculate the Euclidean distance, and the calculation formula is as shown in (5): (5); where, is the k-th cluster center, k = {1, 2, …, K}, K is the number of initial cluster centers; After the latent variable assignment is completed, the K-Means algorithm updates the cluster center of each cluster. The new cluster center is the mean of all latent variables within the cluster, and the specific calculation process is as shown in formula (6): (6); where, is the k-th cluster, where is the number of latent variables in the k-th cluster; The clustering process of the K-Means algorithm is an iterative process, and the above steps are repeated until the cluster centers converge.
[0028] Next, use the silhouette coefficient method to determine the optimal number of cluster centers K best , and its calculation process is as follows: Different latent variables represent different data points. Calculate the average distance from the latent variable to other latent variables in the same cluster (7); where, are other latent variables within the cluster.
[0029] Calculate the average distance of all latent variables to the nearest other cluster : (8); where are other clusters except the k-th cluster.
[0030] Calculate the silhouette coefficient of the latent variable : (9); For different numbers K of initial clustering centers, calculate the average value of the silhouette coefficients of all latent variables, and select K that maximizes the average silhouette coefficient as the optimal number of clustering centers K best : (10); where is the silhouette coefficient of the latent variable .
[0031] Finally, use the optimal number of clustering centers K best to retrain the K-Means algorithm to obtain the optimal initial clustering centers: (11); where is the set of optimal initial clustering centers, is the first optimal initial clustering center, is the K best -th optimal initial clustering center.
[0032] Step 2.2: Use the LSTM-VAE model with the optimal initial clustering centers added to learn and represent the user behavior features, and dynamically adjust the clustering centers during the training process of the LSTM-VAE model. As Figure 3 shown, the LSTM-VAE model uses a single-layer LSTM to replace the fully connected layer as the encoder and decoder of the variational autoencoder (VAE). In the encoder, after the LSTM processing, a linear layer is used for processing. LSTM can effectively capture the long-term and short-term dependencies in the time series data and can effectively address the problem of gradient vanishing or explosion existing in the simple recurrent network. The input data of this step is the user behavior features . The process of the encoder mapping the input data to the latent variable z in the latent space is the same as that described in Step 2.1. The decoder generates the reconstructed sequence by decoding the latent variable z back to the input space, is the reconstructed feature at time t, thus measuring the modeling ability of the LSTM-VAE model for user behavior patterns, as shown in Formulas (12) and (13): (12); (13); where is the hidden state generated by the decoder at time t, is the hidden state of the decoder at time t-1, is the trainable weight matrix of the decoder output layer, is the trainable bias vector of the decoder output layer, is the latent variable at time t.
[0033] The LSTM-VAE model is trained using a joint optimization strategy, simultaneously optimizing the reconstruction loss, KL divergence loss, and clustering loss. The reconstruction loss is used to measure the difference between the sequence reconstructed by the decoder based on the latent variable and the original input sequence. The reconstruction loss is calculated as shown in Formula (14): (14); The KL divergence loss is used to control the distribution in the latent space to make it close to the standard normal distribution. The KL divergence loss is calculated as shown in Formula (15): (15); where is the square of the mean of the latent variable , is the variance of the latent variable .
[0034] The clustering loss is defined as the average of the minimum Euclidean distances from the data points to the nearest cluster centers. Through the clustering loss, the model can separate normal and abnormal behaviors in the latent space more clearly, thereby improving the accuracy of anomaly detection. The clustering divergence loss is calculated as shown in Formula (16): (16); By jointly optimizing the reconstruction loss, KL divergence loss, and clustering loss, the reconstruction quality of the sequence, the distribution of the latent space, and the separability of the clustering structure are simultaneously optimized. The total loss is calculated as shown in Formula (17), where and are hyperparameters that control the regularization strength of the KL divergence and the influence of the clustering constraint on the LSTM-VAE model respectively.
[0035] (17); During the training process of the LSTM-VAE model, the clustering centers are dynamically adjusted. After every P batches, based on the current latent variables the clustering centers are recalculated as shown in formula (18), where is the momentum coefficient, which is used to control the degree of information retention of the clustering centers. (18); Based on the trained LSTM-VAE model, the anomaly score of the user behavior data to be detected is calculated. Combining with the defined threshold, it is judged whether the user behavior poses a threat. The anomaly score is obtained by calculating the reconstruction error and the clustering distance of the user behavior data to be detected, as shown in formula (19), where is the weight coefficient, which is used to control the importance of the clustering distance in the anomaly score.
[0036] (19); where, is the anomaly score. When the anomaly score of the user behavior data exceeds the threat threshold, it is determined that the user behavior poses a threat. Embodiment 3
[0037] An internal threat detection system based on user behavior characterization technology, comprising: A data acquisition module, which is used to acquire user behavior data; A data preprocessing module: which is used to aggregate the normal user behavior data from different logs according to the unique ID value of the user, extract the user behavior characteristics therefrom, and organize them into windowed normal user behavior characteristic time series data; A threat detection module, which has a user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model built in. It uses the normal user behavior characteristic time series data to train the user behavior characterization model and dynamically adjusts the clustering centers during the training process; it uses the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combines with the defined threshold to judge whether the user behavior poses a threat. Embodiment 4
[0038] An electronic device, comprising a processor, a memory, and a computer program stored in the memory. When the computer program is executed by the processor, it implements the internal threat detection method based on user behavior characterization technology as described above.
[0039] The above-disclosed are only some preferred embodiments of the present invention. Of course, the scope of rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. An internal threat detection method based on user behavior characterization technology, characterized in that: The following steps are involved: S1. Data preprocessing: Aggregate normal user behavior data from different logs according to the user's unique ID value, extract user behavior features from them, and organize them into windowed normal user behavior feature time series data; S2. Build a user behavior representation model consisting of a clustering enhancement module and an LSTM-VAE model, pass the user behavior features to the clustering enhancement module, the clustering enhancement module uses LSTM as an encoder to convert the user behavior features into latent variables and cluster them to obtain the best initial cluster center, and then pass the best initial cluster center to the LSTM-VAE model. The LSTM-VAE model learns and represents the user behavior features, and dynamically adjusts the cluster center during the LSTM-VAE model training process; S3. Use the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combine it with the defined threshold to determine whether the user behavior is a threat.
2. The internal threat detection method based on user behavior characterization technology according to claim 1 is characterized in that: The clustering process of the clustering enhancement module is as follows: The input data of the clustering enhancement module is user behavior features. The clustering enhancement module includes LSTM, linear layer and K-Means algorithm clustering module; LSTM receives input user behavior features , and convert it into a hidden state , the hidden state of the last moment of LSTM The mean μ and logarithmic variance of the latent variable are obtained through the fully connected layer , and then use the reparameterization method to get the latent variable z; Use the K-Means algorithm to cluster latent variables, and select the optimal number of cluster centers K by the silhouette coefficient method best ; Finally, the optimal number of cluster centers K is used best Retrain the K-Means algorithm to obtain the best initial clustering center.
3. The internal threat detection method based on user behavior characterization technology according to claim 1 is characterized in that: The LSTM-VAE model uses a single-layer LSTM instead of a fully connected layer as the encoder and decoder of the variational autoencoder. The input data of the LSTM-VAE model is user behavior features. The encoder maps the input data into latent variables in the latent space, and the decoder generates a reconstructed sequence by decoding the latent variables back to the input space.
4. The internal threat detection method based on user behavior characterization technology according to claim 1 is characterized in that: The LSTM-VAE model is trained using a joint optimization strategy to optimize reconstruction loss, KL divergence loss, and clustering loss simultaneously; ; in, is the total loss of the LSTM-VAE model, is the reconstruction loss, is the KL divergence loss, is the clustering loss, where and is a hyperparameter.
5. The internal threat detection method based on user behavior characterization technology according to claim 1 is characterized in that: The way to dynamically adjust the cluster center is: after each P batch, recalculate the cluster center based on the current hidden variables, where is the momentum coefficient, which is used to control the degree of information retention of the cluster center; ; in, is the cluster center before adjustment, is the adjusted cluster center, where is the kth cluster, where is the number of hidden variables in the kth cluster; is the ith hidden variable, is the kth cluster center.
6. The internal threat detection method based on user behavior characterization technology according to claim 1 is characterized in that: The data preprocessing includes the following sub-steps: Step 1.1: Aggregate normal user data from different logs based on the user's unique ID value, extract the user's behavior data within a day with the time granularity as the day, then merge them and sort them by time; Step 1.2: Extract user behavior features from the aggregated data and organize them into normal user behavior feature time series data; Step 1.3: Window the normal user behavior feature time series data.
7. A system for implementing the internal threat detection method based on user behavior characterization technology according to any one of claims 1 to 6, characterized in that: include: Data collection module, used to collect user behavior data; Data preprocessing module: used to aggregate normal user behavior data from different logs according to the user's unique ID value, extract user behavior features from them, and organize them into windowed normal user behavior feature time series data; The threat detection module has a built-in user behavior characterization model composed of a clustering enhancement module and an LSTM-VAE model. It uses normal user behavior feature time series data to train the user behavior characterization model and dynamically adjusts the cluster center during the training process. It uses the trained user behavior characterization model to calculate the anomaly score of the user behavior data to be detected, and combines the defined threshold to determine whether the user behavior poses a threat.
8. An electronic device comprising a processor, a memory, and a computer program stored in the memory, characterized in that: When the computer program is executed by the processor, the internal threat detection method based on user behavior characterization technology described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Internal threat detection method based on VAE and BPNN
CN111726350A
Microservice anomaly detection method based on call chain
CN115269357A
Low-voltage distribution network topological structure anomaly detection method, device and system based on VAE-LSTM
CN116011329A
Unsupervised internal threat detection method based on adversarial auto-encoder
CN116957049A
Power grid data anomaly detection method and system based on LSTM-VAE network
CN118861961A
Cited By
Internal threat detection method and system based on behavior trajectory modeling
CN122496329A