Abnormal login detection method based on adaptive threshold and multi-model integration
By using an adaptive threshold and multi-model integration method in abnormal login detection, using deep learning models to learn user login behavior and dynamically adjust thresholds, the problem that traditional methods cannot adapt to user behavior complexity is solved, and higher detection accuracy and system security are achieved.
Patent Information
- Application Number
- CN202311675836.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional abnormal login detection methods cannot adapt to the diversity and complexity of user behavior, resulting in high frequency of false alarms and missed reports, and the inability to effectively identify abnormal login behavior.
Anomaly login detection method based on adaptive thresholds and multi-model integration is adopted to learn user login behavior through deep learning models (such as LSTM, CNN, GRU), and dynamically adjust the threshold value in combination with adaptive threshold policies to adapt to changes in user behavior.
It effectively reduces the frequency of false alarms and missed reports, improves the accuracy and adaptability of abnormal login detection, enhances the security of the system, and provides solid protection for user data.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] This patent proposes an abnormal login detection method based on adaptive threshold and multi-model integration. This method learns the user's login behavior through a deep learning model, combines the adaptive threshold strategy to achieve dynamic threshold adjustment, effectively reduces the false alarm and missed alarm frequencies, improves the detection accuracy, further enhances the system security, and provides a solid guarantee for protecting user data. Background Art
[0002] With the rapid development of the Internet and information technology, people are increasingly relying on online services to handle daily affairs, such as shopping, socializing, finance, etc. However, network security issues have become increasingly prominent, especially the abnormal login behavior in the user identity authentication link. Abnormal login usually refers to the attacker using various means to steal the user's credentials and attempt illegal login, so as to carry out malicious behaviors such as stealing data and tampering with information. In order to prevent such behaviors, effective abnormal login detection technologies need to be adopted to ensure the security of user accounts and data.
[0003] Traditional abnormal login detection methods mainly rely on pre-set rules and fixed thresholds, such as login location, login time, device information, etc. Although these methods can achieve a certain degree of security protection in some cases, they often cannot adapt to the diversity and complexity of user behaviors. Too strict thresholds may lead to false alarms, while too loose thresholds may lead to missed alarms. In order to improve the accuracy and adaptability of abnormal login detection, new methods need to be explored to better capture the user's login behavior patterns. To solve this problem, this patent proposes an abnormal login detection method that can automatically learn the user's login behavior and dynamically adjust the threshold.
[0004] Deep learning technologies such as Long Short-Term Memory (LSTM) neural networks, Convolutional Neural Networks (CNN), and Gated Recurrent Units (GRU) have shown extremely high efficiency in processing time series data and capturing long-term dependencies. Each model has its unique advantages and characteristics. The LSTM model is good at processing time series data with long-term dependencies, the CNN model is good at identifying local patterns and spatial features, and the GRU model is between the two, being able to process time series data and better identify local patterns. Therefore, introducing these technologies into abnormal login detection can help us better understand and learn the user's login behavior patterns. We propose an abnormal login detection method based on adaptive threshold and multi-model integration. This system can not only learn the normal login behavior patterns of users, but also dynamically adjust the threshold to adapt to the changes in user behaviors, thus achieving more accurate abnormal login detection and ensuring the security of user accounts and data. Summary of the Invention
[0005] In view of the above problems, this patent proposes an abnormal login detection system based on adaptive threshold and multi-model integration. The system integrates LSTM, CNN, and GRU models to comprehensively understand and learn users' login behaviors, and dynamically adjusts the threshold according to users' behaviors. The system first collects a large amount of user login data to train the integrated model. During the training process, the system will automatically adjust the weights of the network model to best fit the normal login patterns of users. Then, new login requests are input into the trained integrated model to determine whether the login behavior is abnormal. When it is determined to be abnormal, the system will issue an alarm. At the same time, to adapt to the changes in users' login behaviors, the system will adjust the threshold accordingly based on new login data to optimize the performance of the system model. This method aims to overcome the limitations of traditional methods and improve the accuracy and adaptability of abnormal login detection.
[0006] The overall structure of the abnormal login detection system based on adaptive threshold and multi-model integration described in this patent is as Figure 1 shown and mainly consists of the following three parts: (1) Data collection and preprocessing module, which is used to collect user login data, perform preprocessing operations such as data cleaning and standardization, and extract key features from the original login data, such as login time, login location, etc. This module converts the original login data into a format suitable for input into the integrated model. (2) Adaptive neural network model module, which constructs and trains an integrated network containing multiple models of LSTM, CNN, and GRU to capture the temporal features and long-term dependencies of users' login behaviors in order to learn the normal login patterns of users. During the training process, the system will automatically adjust the weights of the network to best fit the normal login behaviors of users. At the same time, this module includes an adaptive threshold strategy. Based on users' historical login behaviors, this module will calculate a dynamic threshold that matches users' behaviors to determine whether a new login behavior is abnormal. The system incorporates new login data into the model to dynamically adjust the threshold to adapt to changes in users' login behaviors. The adaptive threshold strategy can effectively reduce false positives and false negatives and improve the accuracy of abnormal detection. (3) Abnormal login detection module. After obtaining the trained integrated model, the system preprocesses new login requests and inputs them into the integrated model. Each sub-model independently calculates the output results and conducts weighted voting, calculates the difference between the model output and the actual login data, and assigns an abnormal score to each login behavior. It determines whether the login behavior is abnormal according to the adaptive threshold. The abnormal score can help further analyze and process abnormal login behaviors. If the login request is abnormal, the system issues an alarm.
[0007] The flow of the abnormal login detection method based on adaptive threshold and multi-model integration described in this patent is as Figure 2 shown and is described in detail as follows:
[0008] Step 1: Collect a large amount of user login data, including login time, login location, login device, login result, etc. This data is used to train the deep learning ensemble model. Subsequently, perform data preprocessing on the original login data, extract the key features in the data, and convert the data into vectors so as to input them into the ensemble model.
[0009] Step 2: Build the deep learning ensemble model structure and use the preprocessed data to train this ensemble model. This ensemble model includes LSTM, CNN, and GRU sub-models. These sub-models will be trained in parallel, and their prediction results will be combined after the training ends. At the beginning stage, set an initial threshold, which can be set according to prior knowledge or data distribution. During the training process, the weights of the network will be automatically adjusted to fit the training data to capture the patterns of user login behaviors.
[0010] Step 3: For each login behavior, input the preprocessed data into the ensemble model, calculate the difference between the output of each sub-model and the actual login data. This difference value can be used as an anomaly score to determine whether the login behavior is abnormal. Use a fixed-size sliding window to select the user's historical login behavior data in the recent period. By statistically analyzing the distribution of these data's anomaly scores, calculate a new threshold for dynamic update.
[0011] Step 4: After obtaining the trained ensemble model and the threshold, input the new login request after preprocessing into the ensemble model, calculate the difference between the actual login data and the output of each sub-model, and perform weighted voting on the results of each sub-model to obtain the anomaly score of this login behavior. Compare it with the current threshold. If the anomaly score is higher than the threshold, determine that this login behavior is abnormal; otherwise, determine it as normal. For the login behaviors determined to be abnormal, the system will issue an alarm.
[0012] Step 5: To adapt to the changes in user login behaviors, the system will conduct regular evaluations and optimizations, incorporate the new login data into the current ensemble model, thereby continuously optimizing the model performance. At the same time, adjust the threshold accordingly based on the new user behavior data.
[0013] The abnormal login detection method based on adaptive threshold and multi-model integration described in this patent is mainly used to identify and prevent abnormal login behaviors in network security. Building and training a deep learning ensemble model to capture the temporal features and long-term dependencies of user login behaviors can effectively identify normal and abnormal login behaviors. The adaptive threshold strategy can effectively reduce the situations of false positives and false negatives, improving the accuracy and adaptability of detection. At the same time, the system can achieve online learning and real-time update to adapt to the changes in user login behaviors. Description of the Drawings
[0014] To more clearly illustrate the content of the present patent invention and the technical solutions in the embodiments, the attached drawings used will be briefly introduced below. The drawings in the following description are only some overall architectures and embodiments of the present patent. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0015] Figure 1 It is the overall architecture diagram of the abnormal login detection method provided by this patent;
[0016] Figure 2 It is the flow chart of the abnormal login detection method provided by this patent;
[0017] Figure 3 It is the adaptive threshold strategy diagram of the embodiment provided by this patent;
[0018] Figure 4 It is the abnormal login detection example diagram of the embodiment provided by this patent. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this patent clearer, the technical solutions in the embodiments of this patent will be clearly and completely described below with reference to the attached drawings in the embodiments of this patent. Obviously, the described embodiments are some, but not all, of the embodiments of this patent. Based on the embodiments in this patent, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this patent.
[0020] In the embodiments of this patent, the Linux Ubuntu 22.04 server is used as a container to detect whether there are abnormalities in the login behavior of the server. There are different degrees of differences in the business systems of different users, and the core applications and architecture complexities are also different. This example is only for illustration purposes. The required software packages and libraries need to be installed in the service area. First, ensure that Python3 is installed in the system for running python scripts. Then use pip to install the necessary python libraries, such as TensorFlow (for building an integrated model), Pandas (for data processing), and Numpy (for numerical calculations), etc. The commands are as follows:
[0021]
[0022] In this embodiment, user login data can be collected from " / var / log / auth.log", which contains the logs of successful and failed logins in the system. Information such as login time, login IP address, and login result can be extracted. In the written Python script, the Pandas library is used to clean and standardize the original login data and perform feature extraction. For non-text features such as login time and IP address, different methods will be used for vectorization. For example, for the login time, it will be converted into the number of seconds elapsed since a certain fixed time point to form a continuous numerical feature. For the IP address, it will be decomposed into four independent numerical features. Then, a deep learning ensemble model is used to learn and train the transformed feature vectors.
[0023] In this example, the TensorFlow library is used to build a deep learning ensemble model, which consists of three sub-models trained in parallel, namely LSTM, CNN, and GRU. Each sub-model is designed to learn and understand different characteristics of user login behavior. The LSTM model is used to capture the time series characteristics of login behavior and capture long-term dependency information; the CNN model is used to extract local features in the login data and capture spatial information; the GRU model combines the advantages of LSTM and CNN and can handle sequence data of different lengths while capturing long-term dependency information. These three sub-models share the preprocessed login data and are trained in parallel. Each model adjusts its weights automatically to fit the normal login behavior of users. After training, the prediction results of the three sub-models will be combined through weighted voting to form the final prediction result.
[0024] Before training, an initial threshold k is set according to prior knowledge or data distribution. For each login behavior vector input into the ensemble model, the difference between the model output and the actual login data is calculated, and this difference value is used as the anomaly score. The threshold will be continuously adjusted during the training process. Usually, the calculation of the threshold will refer to the login behavior data in the recent period. In the embodiment, a fixed window sliding size is adopted to select the latest 100 login data, and the threshold is updated and set to the mean of the abnormal distribution of these data plus one standard deviation. This method can dynamically adjust the threshold according to the changes in the data during the training process, enabling the model to better adapt to the changes in user behavior. After the update iteration during training, a trained deep learning ensemble model and an updated threshold k' can be obtained. The adaptive threshold strategy diagram is as Figure 3 shown. In the embodiment, the training process is encapsulated in a Python script, and this script is run on the Ubuntu server to train the model.
[0025] When a new login request is received, the system first preprocesses the login data and converts it into a vector form. These vectors are then input into a deep learning ensemble model. Each sub-model - LSTM, CNN, and GRU - will independently process the data and output corresponding prediction results. During the calculation process, the models (LSTM, CNN, and GRU) will compare the difference between the actual login data and the model output and assign an anomaly score to the login request. Then, a weighted vote is conducted on the prediction results of the three sub-models. In the example, a weight is assigned to each model, which is determined based on the accuracy of the model on the validation set. The model with higher accuracy will be given a greater weight. After that, the anomaly score of each model is multiplied by its corresponding weight, and these results are added together to obtain a comprehensive anomaly score.
[0026] Based on the threshold k’ obtained in the training phase, it can be determined whether the login behavior is abnormal. If the anomaly score exceeds the threshold, the login behavior is determined to be abnormal. In this case, the system will immediately send an alert to the administrator to detect and handle possible security risks as early as possible. The specific schematic diagram of abnormal login detection is as Figure 4 shown.
[0027] However, simply detecting and reporting abnormal login behaviors is not sufficient to address all network security threats. To further improve the security of the system, the model must be continuously updated and optimized. In this embodiment, the encapsulated script can run continuously and incorporate the new login data into the existing deep learning ensemble model when it is received. In this way, the model can continuously learn and improve its performance over time. In addition, the threshold will also be dynamically adjusted according to the new user behavior data to better adapt to the changes in user behavior.
[0028] To evaluate and improve the performance of the model, the model will be evaluated regularly. The system will divide all login data logs into a training set, a validation set, and a test set, which can help evaluate key metrics such as the accuracy and recall rate of the model. According to these evaluation results, the model structure, feature selection, and threshold update strategy can be further optimized to improve the accuracy and adaptability of abnormal login detection and better protect the security of the system.
[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this patent and are not intended to limit it; although this patent has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this patent.
Claims
1. An abnormal login detection method based on adaptive threshold and multi-model integration, characterized in that: Construct a deep learning integrated model (including LSTM, CNN, and GRU sub-models) to learn and train historical login data, and adopt an adaptive threshold strategy to achieve more accurate abnormal login detection.
2. The method according to claim 1, characterized in that: It consists of data collection and preprocessing, construction and training of a deep learning integrated model, and abnormal login detection.
3. The data collection and preprocessing module according to claim 2, characterized in that: Extract user login data from the login log and perform preprocessing and feature extraction, such as login time, login location, login IP address, etc.
4. The construction and training of the deep learning integrated model according to claim 2, characterized in that: Construct a deep learning integrated model, including LSTM, CNN, and GRU sub-models, and implement an adaptive threshold strategy.
5. The construction of the deep learning integrated model according to claim 4, characterized in that: Use the TensorFlow library to construct a deep learning integrated model and use the extracted data to train the model.
6. The adaptive threshold strategy according to claim 4, characterized in that: Use the login data within a fixed sliding window size to dynamically update the threshold to achieve the adaptability of the model.
7. The abnormal login detection step according to claim 2, characterized in that: The three sub-models assign an abnormal score to each login request behavior, and a comprehensive abnormal score is obtained after weighted voting on the three prediction results. Subsequently, it is judged whether the login behavior is abnormal according to the threshold to achieve abnormal login detection.
8. The method according to claim 1, wherein the performance of the deep learning integrated model can be continuously optimized, and the threshold can also be continuously adjusted to improve the accuracy and adaptability of abnormal login detection.
9. An abnormal login detection system based on adaptive threshold and multi-model integration for implementing the method according to claim 1, including a computer program, a data processing device, and a network connection device.