Abnormal behavior analysis method, storage medium and computer program product

Through the improved T-LSTM network and combined isolated forest anomaly detection algorithm and logistic regression algorithm, the problem of abnormal behavior analysis in the prior art is solved, and the accuracy of abnormal behavior analysis is low, achieving efficient abnormal behavior recognition and analysis.

CN120336137APending Publication Date: 2025-07-18中移信息技术有限公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510317721.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing abnormal behavior analysis methods cannot perform effective abnormality detection from the user's overall operational behavior sequence, and the detection accuracy is low.

Method used

The improved timing long short-term memory network (T-LSTM) is used to combine the isolated forest anomaly detection algorithm and logistic regression algorithm. By abnormal identification and analysis of user operation behavior feature vectors, timing learning is established on business processes to reduce the false positive rate.

Benefits of technology

The overall identification and detection of user operation behavior sequences is realized, the accuracy of behavior abnormality analysis is improved, and the false alarm rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336137A_ABST
    Figure CN120336137A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal behavior analysis method, a storage medium and a computer program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an operation behavior feature vector; based on an exception recognition model, performing exception recognition on the operation behavior feature vector to obtain an exception recognition result; and on the basis of an isolated forest anomaly detection algorithm and a logistic regression algorithm, obtaining an abnormal behavior analysis result according to the anomaly recognition result. In the application, the anomaly identification model is obtained by performing time perception improvement on the time sequence long-short term memory network, the anomaly identification is performed by time sequence learning based on the business process, and the false alarm rate is reduced through the isolated forest anomaly detection algorithm and the logistic regression algorithm. Therefore, the problems that an existing method cannot carry out anomaly detection on the whole operation behavior sequence and the detection accuracy is low are solved, overall recognition and detection on the user operation behavior sequence are achieved, and the accuracy of behavior anomaly analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to an abnormal behavior analysis method, a storage medium, and a computer program product. Background Art

[0002] With the deepening of digital transformation and the popularization of paperless office, a large number of daily tasks are completed through PCs. In daily office work, the weakening of employees' data security awareness, the theft of user accounts due to cyberattacks, and the malicious and illegal operations of employees have led to frequent data leakage incidents in enterprises. How to track users' operation behaviors, discover and predict high-risk behaviors during the operation process has become an important issue faced by current enterprise data security.

[0003] With the continuous evolution of cyberattack technologies, the analysis of users' abnormal behaviors has become increasingly prominent in network security defense. In addition, with the widespread application of cloud computing, the Internet of Things, and mobile devices, the network environment has become increasingly complex, and user behavior analysis also faces new challenges. Security experts need to continuously optimize analysis algorithms and technologies to adapt to different types and scales of network environments and improve the detection and defense capabilities against diverse attack means.

[0004] Traditional user behavior analysis mainly analyzes through rule-based strategies and machine learning algorithms. Although it can solve risk behaviors to a certain extent, the analysis time range is limited. As time goes by, the operation behavior patterns of users will also change over time, and the features are also changing, making it impossible to identify abnormal operation behaviors in the long-term sense. In addition, traditional methods have disadvantages such as high false alarm rates and low recall rates in practical applications.

[0005] The above content is only used to assist in understanding the technical solution of the present application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of the present application is to provide an abnormal behavior analysis method, a storage medium, and a computer program product, aiming to solve the technical problem that the existing abnormal behavior analysis methods cannot perform abnormal detection from the overall operation behavior sequence of users and have a low accuracy rate of abnormal detection.

[0007] To achieve the above object, the present application proposes an abnormal behavior analysis method, and the method includes:

[0008] Obtain an operation behavior feature vector;

[0009] Based on a pre-trained abnormal recognition model, perform abnormal recognition on the operation behavior feature vector to obtain an abnormal recognition result, where the abnormal recognition model is obtained by improving the time perception of a preset temporal long short-term memory network;

[0010] Based on a preset isolation forest anomaly detection algorithm and a pre - constructed logistic regression algorithm, an abnormal behavior analysis result is obtained according to the abnormal recognition result.

[0011] In one embodiment, before the step of performing abnormal recognition on the operation behavior feature vector based on a pre - trained abnormal recognition model to obtain an abnormal recognition result, the following steps are further included:

[0012] Obtain an initial training set and model initialization parameters, where the initial training set includes a number of operation sequence training sets;

[0013] Configure the time - series long short - term memory network according to the model initialization parameters to obtain a model to be trained;

[0014] Based on a preset forward propagation mechanism, input any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value;

[0015] Based on a preset loss function, calculate the error between the training set prediction value and the operation sequence label of the operation sequence training set to obtain a prediction error value;

[0016] Based on a preset backpropagation algorithm, optimize the parameters of the model to be trained according to the prediction error value to obtain a model to be trained with optimized parameters;

[0017] Return to execute the steps: Based on a preset forward propagation mechanism, input any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value, until all operation sequence training sets are traversed, and use the model to be trained with optimized parameters as the abnormal recognition model.

[0018] In one embodiment, the step of obtaining the initial training set includes:

[0019] Collect user behavior logs;

[0020] Pre - process the user behavior logs to obtain first - stage log data;

[0021] Perform vector conversion on the first - stage log data to obtain user behavior feature vectors;

[0022] Based on a preset clustering algorithm, perform time - series slicing on the user behavior feature vectors to obtain an operation feature data set;

[0023] Perform multi - dimensional array conversion on the operation feature data set to obtain the initial training set.

[0024] In one embodiment, the step of performing anomaly recognition on the operation behavior feature vector based on a pre-trained anomaly recognition model to obtain an anomaly recognition result includes:

[0025] Input the operation behavior feature vector into the anomaly recognition model to obtain a behavior prediction value;

[0026] Based on a preset cosine similarity calculation algorithm, calculate the similarity between the operation behavior feature vector and the behavior prediction value to obtain a user feature similarity;

[0027] Obtain the anomaly recognition result according to the user feature similarity.

[0028] In one embodiment, the step of obtaining an abnormal behavior analysis result based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the anomaly recognition result includes:

[0029] When the anomaly recognition result indicates that the user feature similarity is greater than a preset error threshold, based on a preset reverse conversion algorithm, convert the operation behavior feature vector to obtain a user behavior log;

[0030] Based on a preset risk rating dimension, perform risk assessment on the user behavior log to obtain a user risk rating data set;

[0031] Based on the isolation forest anomaly detection algorithm, perform anomaly monitoring on the user risk rating data set to obtain an anomaly detection result;

[0032] Obtain the user's abnormal operation according to the anomaly detection result;

[0033] Based on the logistic regression algorithm, calculate the risk coefficient of the user's abnormal operation to obtain an abnormal behavior analysis result.

[0034] In one embodiment, before the step of obtaining an abnormal behavior analysis result based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the anomaly recognition result, it further includes:

[0035] Obtain an original training data set;

[0036] Train a preset logistic regression model according to the original training data set to obtain the logistic regression algorithm. In one embodiment, the step of training the logistic regression algorithm according to the original training data set includes:

[0037] S1: Based on a preset random algorithm, extract any original training data in the original training data set as standard training data, and use the original training data set excluding the standard training data as the first training data set;

[0038] S2: Train the logistic regression model according to the first training dataset and the standard training data to obtain a second training dataset;

[0039] S3: Determine whether the determinant of the second training dataset is less than a preset determinant threshold:

[0040] S4: If not, define the second training dataset as the original training dataset;

[0041] Repeat steps S1 - S4 until the determinant of the second training dataset is less than the preset threshold, and obtain the logistic regression algorithm according to the trained logistic regression model.

[0042] In one embodiment, the step of training the logistic regression model according to the first training dataset and the standard training data to obtain a second training dataset includes:

[0043] Based on a preset cosine calculation algorithm, calculate the distance between the standard training data and the first training data in the first training dataset to obtain a set of training data distances;

[0044] Filter the first training dataset according to the set of training data distances to obtain a third training dataset;

[0045] Based on a preset gradient descent algorithm, train the logistic regression model according to the third training dataset;

[0046] Obtain a fourth training dataset according to the original training dataset and the third training dataset;

[0047] Discriminate the fourth training dataset through the trained logistic regression model to obtain a discrimination dataset;

[0048] Filter the discrimination dataset to obtain the second training dataset.

[0049] In addition, to achieve the above object, the present application also proposes an abnormal behavior analysis device, and the abnormal behavior analysis device includes:

[0050] An acquisition module, configured to acquire an operation behavior feature vector;

[0051] An identification module, configured to perform abnormal identification on the operation behavior feature vector based on a pre - trained abnormal identification model to obtain an abnormal identification result, and the abnormal identification model is obtained by improving the time perception of a preset temporal long - short - term memory network;

[0052] An analysis module, configured to obtain an abnormal behavior analysis result based on a preset Isolation Forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the anomaly recognition result.

[0053] In addition, to achieve the above object, the present application also provides an abnormal behavior analysis device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the abnormal behavior analysis method as described above.

[0054] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the abnormal behavior analysis method as described above.

[0055] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the abnormal behavior analysis method as described above.

[0056] The present application provides an abnormal behavior analysis method. The anomaly recognition model is obtained by improving the time perception of a preset time series long short-term memory network. When performing anomaly recognition, it will not break the complete business chain, so that the user behavior anomaly recognition is a time series learning based on the business process. And through the Isolation Forest anomaly detection algorithm and the logistic regression algorithm, the false alarm rate can be effectively reduced. Thus, the technical problem that the existing abnormal behavior analysis method cannot perform anomaly detection from the overall operation behavior sequence of the user and the accuracy of anomaly detection is low is solved, and the overall recognition and detection of the user operation behavior sequence are realized, and the accuracy of behavior anomaly analysis is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0058] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the abnormal behavior analysis method of the present application;

[0060] Figure 2It is a schematic flowchart provided for the second embodiment of the abnormal behavior analysis method of this application;

[0061] Figure 3 It is a schematic flowchart of the brief process of training and application of the abnormal recognition model provided for the first or second embodiment of this application;

[0062] Figure 4 It is a schematic diagram of the content of the cleaned log data provided for the first or second embodiment of this application;

[0063] Figure 5 It is a schematic diagram of the content of the encoded log data provided for the first or second embodiment of this application;

[0064] Figure 6 It is a schematic flowchart of the brief process of business slicing of the user operation sequence provided for the first or second embodiment of this application;

[0065] Figure 7 It is a schematic diagram of the content of the operation feature dataset after time series slicing provided for the first or second embodiment of this application;

[0066] Figure 8 It is a schematic diagram of the content of the initial training set provided for the first or second embodiment of this application;

[0067] Figure 9 It is a schematic diagram of the model structure of the time series long short-term memory network provided for the first or second embodiment of this application;

[0068] Figure 10 It is a schematic flowchart of the brief process of training and application of the abnormal recognition model based on blockchain evidence storage and update provided for the first or second embodiment of this application;

[0069] Figure 11 It is a schematic flowchart of the brief process of user abnormal analysis provided for the first or second embodiment of this application;

[0070] Figure 12 It is a schematic diagram of the module structure of the abnormal behavior analysis device in the embodiment of this application;

[0071] Figure 13 It is a schematic diagram of the device structure of the hardware operating environment involved in the abnormal behavior analysis method in the embodiment of this application.

[0072] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0073] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application. For a better understanding of the technical solutions of this application, the following will be described in detail in conjunction with the specification drawings and specific implementation manners.

[0074] The main solution of the embodiment of this application is: obtaining an operation behavior feature vector; based on a pre-trained anomaly recognition model, performing anomaly recognition on the operation behavior feature vector to obtain an anomaly recognition result, where the anomaly recognition model is obtained by improving the time perception of a preset temporal long short-term memory network; based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm, obtaining an abnormal behavior analysis result according to the anomaly recognition result.

[0075] Due to the traditional user behavior analysis related to the prior art, it is mainly analyzed by means of rule policies and machine learning algorithms. Although it can solve risk behaviors to a certain extent, the analysis time range is limited. As time goes by, the user operation behavior pattern also changes over time, and the features are also changing, and abnormal operation behaviors in the long-term sense cannot be recognized. The existing user behavior analysis methods can detect emerging abnormal operations. However, there are also several problems in its technical application.

[0076] First, the violation audit of basic policies: For different scenarios, a violation audit policy model is constructed in a rule-based manner. This is a basic audit method, mainly auditing the existing violation behaviors in daily life, such as "circumvention audit" commonly. In actual application, this method can only detect the violations of users at a certain point, but cannot discover the anomalies in the entire operation behavior sequence.

[0077] Second, the anomaly detection analysis of deviation from thresholds and baselines: Common processing methods such as setting a safety threshold and a safety baseline through expert experience to set the user safety behavior interval. This method mainly uses mathematical statistics to perform aggregated analysis on a single feature dimension of users. After the feature dimensions are aggregated, the abnormal patterns hidden in the user behavior are often masked. This method has disadvantages such as a high false alarm rate and a low recall rate in actual application.

[0078] Third, the time series analysis based on artificial feature extraction: Common processing methods such as slicing according to a fixed time and determining the feature dimensions according to expert experience for feature dimension aggregation. Then, linear regression is used for modeling. In actual application, the application effect of the model is greatly limited by the accuracy of the feature dimensions extracted by expert experience. And the modeling algorithm is too simple, and the expressiveness of the model is insufficient.

[0079] The present application provides a solution. The anomaly recognition model is obtained by improving the time perception of a preset temporal long short-term memory network. When performing anomaly recognition, it will not break the complete business chain, enabling the recognition of abnormal user behaviors to be a temporal learning based on the business process. Moreover, through the Isolation Forest anomaly detection algorithm and the Logistic Regression algorithm, the false alarm rate can be effectively reduced. Thus, the technical problem that the existing abnormal behavior analysis methods cannot perform anomaly detection from the overall operation behavior sequence of users and the accuracy of anomaly detection is relatively low is solved, realizing the overall recognition and detection of the user operation behavior sequence and improving the accuracy of behavior anomaly analysis.

[0080] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an abnormal behavior analysis system, etc. that can implement the above functions. Hereinafter, taking the abnormal behavior analysis system as an example, this embodiment and the following embodiments will be described.

[0081] Based on this, an embodiment of the present application provides an abnormal behavior analysis method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the abnormal behavior analysis method of the present application.

[0082] In this embodiment, the abnormal behavior analysis method includes steps S10 to S30:

[0083] Step S10, obtain an operation behavior feature vector;

[0084] The abnormal behavior analysis system obtains an operation behavior feature vector.

[0085] Step S20, based on a pre-trained anomaly recognition model, perform anomaly recognition on the operation behavior feature vector to obtain an anomaly recognition result, where the anomaly recognition model is obtained by improving the time perception of a preset temporal long short-term memory network;

[0086] The abnormal behavior analysis system performs anomaly recognition on the operation behavior feature vector based on a pre-trained anomaly recognition model to obtain an anomaly recognition result, where the anomaly recognition model is obtained by improving the time perception of a preset temporal long short-term memory network.

[0087] It should be noted that the anomaly recognition model is constructed based on deep learning technology, especially the long short-term memory network (T-LSTM). T-LSTM is a variant of the recurrent neural network (RNN), specifically designed to solve the problems of gradient vanishing and explosion encountered by ordinary RNNs when processing long sequence data. By introducing a gating mechanism (including an input gate, a forget gate, and an output gate), T-LSTM can effectively capture long-term dependencies in the time series.

[0088] In a feasible implementation manner, before step S20, steps S401 to S406 may further be included:

[0089] Step S401: Obtain an initial training set and model initialization parameters, where the initial training set includes a plurality of operation sequence training sets;

[0090] Step S402: Configure the temporal long short-term memory network according to the model initialization parameters to obtain a model to be trained;

[0091] Step S403: Based on a preset forward propagation mechanism, input any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value;

[0092] Step S404: Based on a preset loss function, calculate an error between the training set prediction value and the operation sequence label of the operation sequence training set to obtain a prediction error value;

[0093] Step S405: Based on a preset backpropagation algorithm, optimize the parameters of the model to be trained according to the prediction error value to obtain a model to be trained with optimized parameters;

[0094] Step S406: Return to execute the step: Based on a preset forward propagation mechanism, input any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value, until all the operation sequence training sets are traversed, and use the model to be trained with optimized parameters as the anomaly recognition model.

[0095] The abnormal behavior analysis system obtains an initial training set and model initialization parameters. The initial training set includes several operation sequence training sets. Then, according to the model initialization parameters, the temporal long short-term memory network is configured to obtain a model to be trained, which includes determining the number of network layers, the number of neurons in each layer, the activation function, the input and output dimensions, etc. Then, based on a preset forward propagation mechanism, any operation sequence training set is input into the model to be trained for prediction to obtain training set prediction values, so as to calculate prediction errors and optimize the model subsequently. Then, based on a preset loss function, an error calculation is performed on the training set prediction values and the operation sequence labels of the operation sequence training sets to obtain prediction error values. The role of the error calculation is to provide a feedback signal for model training. This signal indicates the accuracy of the model prediction and is used to calculate the direction and magnitude of parameter updates. Then, based on a preset backpropagation algorithm, according to the prediction error values, the parameters of the model to be trained are optimized to obtain the model to be trained with optimized parameters. Finally, the execution steps are returned: based on a preset forward propagation mechanism, any of the operation sequence training sets is input into the model to be trained for prediction to obtain training set prediction values until all operation sequence training sets are traversed, and the model to be trained with optimized parameters is used as the abnormal recognition model.

[0096] It should be noted that the initial training set contains a training data set of several operation sequences, which are used to train the model to enable it to identify normal and abnormal behaviors; the model initialization parameters are the parameters set before the start of model training, such as weight initialization strategies, learning rates, etc. These parameters have an important impact on the training and performance of the model; forward propagation is a basic process in deep learning, which involves passing input data through network layers until an output result is generated. In the T-LSTM network, forward propagation includes the calculation of input gates, forget gates, and output gates. These gates control the flow of information, enabling the network to capture long-term dependencies in time series data; the loss function is a function that measures the difference between the model prediction value and the actual value. During the training process, the loss function is used to evaluate the performance of the model and guide the optimization of model parameters; the backpropagation algorithm is the core algorithm for training neural networks in deep learning. It calculates the gradients of the loss function with respect to the model parameters and uses these gradients as the basis for updating the parameters to minimize the loss function and optimize the model performance.

[0097] In this embodiment, through parameter optimization and iterative training, the risks of overfitting and underfitting can be reduced, and the generalization ability of the model can be improved; iterative optimization improves the training efficiency and ensures that the model can converge quickly; the entire process improves the accuracy of anomaly recognition by continuously optimizing the model parameters and structure; by traversing all operation sequence training sets, the model can learn a wider range of patterns and improve its generalization ability on unknown data; error calculation provides necessary feedback for the training process, enabling dynamic adjustment during the training process. Thus, not only an anomaly recognition model based on T-LSTM is constructed, but also the performance and generalization ability of the model are improved through iterative training and parameter optimization, solving various technical problems that may be encountered in traditional anomaly detection.

[0098] In another feasible embodiment, step S401 may include steps S4011 to S4015:

[0099] Step S4011, collecting user behavior logs;

[0100] Step S4012, preprocessing the user behavior logs to obtain first log data;

[0101] Step S4013, performing vector transformation on the first log data to obtain user behavior feature vectors;

[0102] Step S4014, based on a preset clustering algorithm, performing time series slicing on the user behavior feature vectors to obtain an operation feature data set;

[0103] Step S4015, performing multi-dimensional array conversion on the operation feature data set to obtain the initial training set.

[0104] The abnormal behavior analysis system collects user behavior logs, then preprocesses the user behavior logs to obtain first log data. The purpose of preprocessing is to improve the data quality to make it more suitable for subsequent analysis and model training. Preprocessing can reduce errors during model training and improve the performance and accuracy of the model. Then, vector transformation is performed on the first log data to obtain user behavior feature vectors. The purpose is to extract useful information from the first log data and convert it into a format that can be understood and processed by machine learning models. Then, based on a preset clustering algorithm, time series slicing is performed on the user behavior feature vectors to obtain an operation feature data set. The purpose is to effectively segment the user behavior feature vectors according to the logic of the time series to construct a data set that can represent the user's operation behavior. Finally, multi-dimensional array conversion is performed on the operation feature data set to obtain the initial training set. The purpose is to convert the operation feature data set into a format that can be directly used by machine learning algorithms for effective model training.

[0105] It should be noted that the user behavior log is a detailed log recording various operations performed by users in the system, usually containing information such as timestamps, operation types, operation results, etc.; preprocessing is the process of data cleaning and transformation, aiming to convert the original user behavior log into a data format suitable for further analysis and model training. This process usually includes steps such as removing noise, filling missing values, formatting, and normalizing data; vector transformation is the process of converting structured or unstructured data into numerical feature vectors, which can represent the key information of the original data and are used as inputs for machine learning models. In the context of user behavior logs, this usually involves converting behavior data into a quantifiable format for analysis and prediction; the user behavior feature vector, as the input of the machine learning model, is crucial for the model to accurately capture user behavior patterns and identify abnormal behaviors; time series chunking is a method of splitting time series data into meaningful chunks or segments to better capture and analyze the time series patterns in the data. The clustering algorithm is used to group user behavior feature vectors with similar characteristics to form an operation feature data set; multi-dimensional array conversion is to convert the operation feature data set into a data structure suitable for machine learning model processing. In machine learning and data processing, multi-dimensional arrays (usually called tensors) are a basic data structure that can represent high-dimensional data such as vectors, matrices, and three-dimensional arrays.

[0106] In this embodiment, preprocessing solves the problems of data inconsistency and noise, and improves data quality; vector transformation solves the problem of unstructured data processing, enabling the model to process complex user behavior data; time series chunking solves the problems of time series data continuity and context loss, retaining the time series characteristics of user behavior; structured transformation solves the compatibility problem of data input into the model, enabling the data to be directly used in machine learning models; the entire process solves the problem of automated feature selection and engineering through automated feature extraction and transformation, reduces manual intervention, and improves efficiency; by converting user behavior logs into multi-dimensional arrays, the constructed data set has better scalability and flexibility, and can adapt to different model and analysis requirements. Thus, not only a high-quality initial training set is constructed, but also through a series of data processing and transformation steps, the usability of the data and the efficiency of model training are improved, and various technical problems that may be encountered in traditional data preprocessing and feature engineering are solved. This provides a solid foundation for the subsequent training of the anomaly detection model.

[0107] In another feasible embodiment, step S20 may include steps S201 to S203:

[0108] Step S201, input the operation behavior feature vector into the anomaly recognition model to obtain a behavior prediction value;

[0109] Step S202: Calculate the similarity between the operation behavior feature vector and the behavior prediction value based on a preset cosine similarity calculation algorithm to obtain the user feature similarity.

[0110] Step S203: Obtain the abnormal recognition result according to the user feature similarity.

[0111] The abnormal behavior analysis system inputs the operation behavior feature vector into the abnormal recognition model to obtain the behavior prediction value, then calculates the similarity between the operation behavior feature vector and the behavior prediction value based on a preset cosine similarity calculation algorithm to obtain the user feature similarity, and finally obtains the abnormal recognition result according to the user feature similarity.

[0112] It should be noted that cosine similarity is a similarity metric often used to measure the angle between two non-zero vectors, and its value ranges from -1 (completely dissimilar) to 1 (completely similar). In anomaly detection, cosine similarity can be used to measure the similarity between the operation behavior feature vector and the model prediction value, thereby evaluating the degree of abnormality of the behavior.

[0113] In this embodiment, the model can predict user behavior, solving the problem of how to identify potential abnormal behaviors from a large amount of behavior data; by calculating the similarity, it solves the problem of how to quantitatively compare the differences between the feature vector and the prediction value, which is crucial for understanding the deviation of the model prediction; it provides a clear abnormal recognition result, solving the problem of how to make accurate anomaly detection decisions in complex data; by calculating the similarity between the feature vector and the prediction value, it enhances the interpretability of the model prediction, making anomaly detection more than just a black box operation; the entire process evaluates the degree of abnormality of user behavior through a quantitative method, improving the objectivity and accuracy of anomaly detection; by comprehensively considering the similarity between the feature vector and the prediction value, it can reduce false positives and false negatives, improving the overall effectiveness of the anomaly detection system. Thus, it not only improves the accuracy and efficiency of anomaly detection, but also enhances the interpretability and reliability of the model prediction through a quantitative method, solving various technical problems that may be encountered in traditional anomaly detection. This provides a solid foundation for building a robust anomaly detection system.

[0114] Step S30: Based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm, obtain the abnormal behavior analysis result according to the abnormal recognition result.

[0115] The abnormal behavior analysis system obtains the abnormal behavior analysis result based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the abnormal recognition result.

[0116] It should be noted that the Isolation Forest anomaly detection algorithm is an unsupervised anomaly detection algorithm based on a tree model, which "isolates" data points by constructing a series of random trees. The core idea is that anomaly points are more likely to be isolated compared to normal points. The Logistic Regression algorithm is a statistical method widely used in binary classification problems, which maps the output of linear regression to the probability space through the logit function to predict the probability that a sample belongs to a certain category.

[0117] In a feasible implementation manner, step S30 may include steps S301 to S305:

[0118] Step S301, when the anomaly recognition result indicates that the user feature similarity is greater than a preset error threshold, based on a preset reverse conversion algorithm, convert the operation behavior feature vector to obtain a user behavior log;

[0119] Step S302, based on a preset risk rating dimension, perform a risk assessment on the user behavior log to obtain a user risk rating data set;

[0120] Step S303, based on the Isolation Forest anomaly detection algorithm, perform anomaly monitoring on the user risk rating data set to obtain an anomaly detection result;

[0121] Step S304, obtain the user's abnormal operations according to the anomaly detection result;

[0122] Step S305, based on the Logistic Regression algorithm, calculate the risk coefficient for the user's abnormal operations to obtain an analysis result of abnormal behavior.

[0123] When the anomaly recognition result indicates that the user feature similarity is greater than a preset error threshold, the abnormal behavior analysis system, based on a preset reverse conversion algorithm, converts the operation behavior feature vector to obtain a user behavior log, and then, based on a preset risk rating dimension, performs a risk assessment on the user behavior log to obtain a user risk rating data set. The purpose is to identify possible abnormal behaviors or high-risk activities by conducting a detailed risk assessment on the user behavior log. Then, based on the Isolation Forest anomaly detection algorithm, anomaly monitoring is performed on the user risk rating data set to obtain an anomaly detection result to identify possible abnormal behaviors or risk points. Then, according to the anomaly detection result, the user's abnormal operations are obtained. Finally, based on the Logistic Regression algorithm, the risk coefficient of the user's abnormal operations is calculated to evaluate the probability of these operations leading to security incidents and obtain an analysis result of abnormal behavior.

[0124] It should be noted that the reverse conversion algorithm is a process of mapping feature vectors back to their original data form. In machine learning, especially in the feature engineering stage, raw data is often converted into feature vectors for easier model processing. The purpose of the reverse conversion algorithm is to recover readable user behavior logs from these feature vectors; risk assessment is a systematic process that involves quantifying the potential risks in user behavior logs according to a series of predefined risk rating dimensions, which may include behavior frequency, time patterns, behavior types, user history, etc.

[0125] In this embodiment, the reverse conversion algorithm is used to convert the operation behavior feature vectors back to user behavior logs, which helps to recover the original behavior data from the model output and provides a basis for further analysis; the risk assessment of user behavior logs based on risk rating dimensions quantifies the risk level of user behavior and provides a basis for subsequent anomaly monitoring; the isolation forest algorithm is used to perform anomaly monitoring on the user risk rating data set, and this algorithm is particularly suitable for anomaly detection of high-dimensional data, improving the effectiveness of anomaly monitoring; identifying user abnormal operations according to the anomaly detection results helps to quickly locate and respond to potential security threats; using the logistic regression algorithm to calculate the risk coefficient of user abnormal operations provides a quantitative assessment of the probability of abnormal behavior occurrence and enhances the accuracy of anomaly detection. Thus, not only the accuracy and efficiency of anomaly detection are improved, but also the understanding and response ability to abnormal behavior are enhanced through quantitative risk assessment and anomaly monitoring.

[0126] This embodiment provides an abnormal behavior analysis method. The anomaly recognition model is obtained by improving the time perception of a preset time series long short-term memory network. When performing anomaly recognition, it will not break the complete business chain, so that the user behavior anomaly recognition is a time series learning based on the business process. And through the isolation forest anomaly detection algorithm and the logistic regression algorithm, the false alarm rate can be effectively reduced. Thus, the technical problem that the existing abnormal behavior analysis methods cannot perform anomaly detection from the overall operation behavior sequence of users and the accuracy of anomaly detection is relatively low is solved, and the overall recognition and detection of the user operation behavior sequence are realized, and the accuracy of behavior anomaly analysis is improved.

[0127] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 Before step S30, the abnormal behavior analysis method further includes steps S501 to S502:

[0128] Step S501, obtaining the original training data set;

[0129] Step S502: Train a preset logistic regression model based on the original training dataset to obtain the logistic regression algorithm.

[0130] The abnormal behavior analysis system obtains the original training dataset, and then trains a preset logistic regression model based on the original training dataset to obtain the logistic regression algorithm.

[0131] It should be noted that the original training dataset involves collecting sufficient data to train the model. These datasets usually contain historical records, which can be user behavior data, transaction records, system logs, etc. They will be used to train the model to identify normal and abnormal behaviors.

[0132] In a feasible implementation, step S502 may include steps S5021 to S5025:

[0133] Step S5021: Based on a preset random algorithm, extract any original training data from the original training dataset as the standard training data, and use the original training dataset excluding the standard training data as the first training dataset.

[0134] Step S5022: Train the logistic regression model based on the first training dataset and the standard training data to obtain the second training dataset.

[0135] Step S5023: Determine whether the determinant of the second training dataset is less than a preset determinant threshold.

[0136] Step S5024: If not, define the second training dataset as the original training dataset.

[0137] Step S5025: Repeat steps S5021 - S5024 until the determinant of the second training dataset is less than the preset threshold, and obtain the logistic regression algorithm based on the trained logistic regression model.

[0138] The abnormal behavior analysis system extracts any original training data from the original training dataset as the standard training data based on a preset random algorithm, and uses the original training dataset excluding the standard training data as the first training dataset. Then, based on the first training dataset and the standard training data, the logistic regression model is trained to obtain a second training dataset. Next, it is determined whether the determinant of the second training dataset is less than a preset determinant threshold, with the aim of evaluating whether the feature matrix of the second training dataset has sufficient linear independence to ensure the stability and accuracy of the logistic regression model. If the determinant of the second training dataset is greater than or equal to the preset determinant threshold, the system defines the second training dataset as the original training dataset. Finally, steps S5021 - S5024 are repeatedly executed until the determinant of the second training dataset is less than the preset threshold. Based on the trained logistic regression model, the logistic regression algorithm is obtained.

[0139] In this embodiment, a random algorithm is used to extract the standard training data from the original training dataset, and the remaining data is used as the first training dataset, which helps to ensure the generalization ability of model training. Training the logistic regression model using the first training dataset to obtain the second training dataset provides a basis for the initial training and evaluation of the model. Detecting the multicollinearity problem of the data by determining whether the determinant of the second training dataset is less than the preset threshold is crucial for controlling the feature correlation in model training. By iteratively optimizing the dataset until the determinant is less than the preset threshold, it helps to obtain a dataset with better linear independence, thereby improving the stability and accuracy of the model. After iterative optimization, the finally obtained logistic regression algorithm can more accurately predict abnormal behaviors, improving the practicality and reliability of the model. Thus, not only the quality and efficiency of logistic regression model training are improved, but also various technical problems that may be encountered in traditional model training are solved by controlling multicollinearity and optimizing dataset features.

[0140] In another feasible embodiment, step S5022 may include steps S50221 - S50226:

[0141] Step S50221, based on a preset cosine calculation algorithm, calculates the distance between the standard training data and the first training data in the first training dataset to obtain a training data distance set;

[0142] Step S50222, filters the first training dataset according to the training data distance set to obtain a third training dataset;

[0143] Step S50223, based on a preset gradient descent algorithm, trains the logistic regression model according to the third training dataset;

[0144] Step S50224: Obtain a fourth training dataset according to the original training dataset and the third training dataset;

[0145] Step S50225: Discriminate the fourth training dataset by using the trained logistic regression model to obtain a discrimination dataset;

[0146] Step S50226: Screen the discrimination dataset to obtain the second training dataset.

[0147] The abnormal behavior analysis system calculates the distance between the standard training data and the first training data in the first training dataset based on a preset cosine calculation algorithm to obtain a training data distance set, and then.

[0148] It should be noted that the cosine similarity is a similarity measurement method for measuring the angle between two vectors, usually used to evaluate the direction similarity of vectors in a multi-dimensional space. In step S50221, the cosine similarity is used to calculate the similarity between the standard training data and each data point in the first training dataset, thereby obtaining a distance set; according to the training data distance set, the first training dataset is screened to obtain a third training dataset, and then based on a preset gradient descent algorithm, the logistic regression model is trained according to the third training dataset, and then according to the original training dataset and the third training dataset, a fourth training dataset is obtained, and then the fourth training dataset is discriminated by using the trained logistic regression model to obtain a discrimination dataset, and finally the discrimination dataset is screened to obtain the second training dataset.

[0149] In this embodiment, the cosine calculation algorithm is used to accurately calculate the distance between the standard training data and the first training dataset, which helps to identify data points similar to the standard samples; the first training dataset is screened according to the distance set to obtain a more relevant third training dataset, improving the quality of the dataset; the gradient descent algorithm is used to train the logistic regression model, effectively optimizing the model parameters and improving the prediction accuracy of the model; the original training dataset and the third training dataset are integrated to obtain a more comprehensive fourth training dataset, enhancing the generalization ability of the model; the fourth training dataset is discriminated by using the trained logistic regression model to obtain a discrimination dataset, improving the accuracy of model prediction; the discrimination dataset is screened to obtain the final second training dataset, providing a refined training basis for the model. Thus, not only the quality and efficiency of the logistic regression model training are improved, but also various technical problems that may be encountered in traditional model training are solved through accurate feature matching and data screening.

[0150] In this embodiment, by obtaining the original training data set, a necessary data basis is provided for subsequent model training. These data sets usually contain user behaviors, transaction records, or other relevant information, which are the key to building the model; by training the original training data set with a logistic regression model, a logistic regression algorithm capable of identifying abnormal behaviors is generated. The training process enables the model to learn the patterns in the data, so as to make effective predictions in practical applications; the logistic regression model has good interpretability and can provide an intuitive understanding of the impact of features on the prediction results. This helps to analyze and understand the reasons for abnormal behaviors; the trained logistic regression model can provide risk assessment capabilities for subsequent abnormal behavior analysis and help identify potential risk behaviors.

[0151] Exemplarily, to help understand the implementation process of the abnormal behavior analysis method obtained by combining this embodiment with the above first embodiment, second embodiment, and third embodiment, please refer to Figures 3 - 11 , Figure 3 which is a schematic diagram of the brief process of training and applying the abnormal recognition model provided in the first or second embodiment of this application, Figure 4 which is a schematic diagram of the content of the cleaned log data provided in the first or second embodiment of this application, Figure 5 which is a schematic diagram of the content of the encoded log data provided in the first or second embodiment of this application, Figure 6 which is a schematic diagram of the brief process of slicing user operation sequences into business chunks provided in the first or second embodiment of this application, Figure 7 which is a schematic diagram of the content of the operation feature data set after slicing by time series provided in the first or second embodiment of this application, Figure 8 which is a schematic diagram of the content of the initial training set provided in the first or second embodiment of this application, Figure 9 which is a schematic diagram of the model structure of the time series long short-term memory network provided in the first or second embodiment of this application, Figure 10 which is a schematic diagram of the brief process of training and applying the abnormal recognition model based on blockchain evidence storage and update provided in the first or second embodiment of this application, Figure 11 which is a schematic diagram of the brief process of user abnormal analysis provided in the first or second embodiment of this application. This embodiment proposes a method that combines an improved T-LSTM network with user behavior analysis. Unsupervised learning is used for the user's historical operation behaviors, so as to learn the normal operation behavior patterns of the user. Furthermore, the behavior patterns that deviate from the normal operation behavior sequence are identified. Specifically, the following steps are included:

[0152] Step 1: Data collection and storage: Collect various types of user behavior logs, including application operation logs, 4A login logs, system operation logs, and other types of user behavior logs. And extract basic user behavior information such as "user ID", "operation type", "operation time", "source device", "target device", "operation result", etc. from the logs. At the same time, various conversions, cleaning, and data processing need to be performed on formatted and unformatted data to unify the data format and write it into the database.

[0153] Step 2: Construct feature vectors: Among the user behavior logs collected in Step 1, there is a large amount of character-type and categorical data, which needs to be converted into a data type that can be operated by the machine before inputting into the algorithm. For dimensions with many category values such as "operation type", "source device", and "target device", the Embedding method is used for processing (to avoid dimensional explosion). According to the collected data, methods such as digital encoding and One-Hot encoding are used to construct user feature vectors. The data encoding format used is as follows:

[0154] Operation result: Processed using 0 and 1 digital encoding. Among them, 0 represents "success" and 1 represents "failure".

[0155] Input of multi-category data such as operation type, source device, and target device: By constructing an Embedding network layer, map the "operation type", "source device", and "target device" in the user's operation logs to a low-dimensional dense vector space. At the same time, the Embedding vector can capture the similarities and relationships between the data. After mapping to the dense vector space, similar operations are mapped into vectors that are close to each other.

[0156] Time information: Perform time interval encoding, that is, for each user behavior event, calculate the time interval from the previous event. Embed it as an additional input into the feature vector.

[0157] Integrate the processed data above and process it into an ordered data set according to the operation timestamp.

[0158] Step 3: Temporal slicing: Temporal slicing is an important improvement in this embodiment. First, extract three dimensions: time interval, operation type, and target device from the user behavior feature vector data processed in Step 2, and perform unsupervised learning using a density-based clustering algorithm DBSCAN. DBSCAN realizes the slicing of adjacent operations into one block according to the method of calculating spatial density, and different blocks represent different operation services.

[0159] In this way, the long operation sequence of the user is segmented into different business blocks according to different business operation characteristics. Different business blocks are marked by a unique coding method, and the operation sequences of the same business block use the same coding.

[0160] Through the method of time-series segmentation, the long operation sequence of the user is segmented into short business blocks. Using the "time-series segmentation flag" as the grouping identifier, the operation feature data set of the user is integrated into a multi-dimensional array and provided to the T-LSTM for training.

[0161] Step 4: Set the model training parameters and initialize the model training process: In the original T-LSTM neural network structure, the "time step" is an important hyperparameter, which is used to segment the long operation sequence chain of the user into sequence blocks of the same length according to a fixed step length and input them into the T-LSTM for training. This fixed-step cutting method has two disadvantages: 1. Some operations with long business processes will be segmented into different sequence blocks for learning, resulting in incomplete information learned by the model; 2. Some operations with short business processes will introduce data that does not belong to this business process after being stretched by the time step, causing noise interference.

[0162] In this embodiment, a number of improvement measures are taken on the basis of the original T-LSTM neural network structure to achieve the purpose of optimizing the network structure.

[0163] First, in the improved T-LSTM neural network, an "Embedding" layer is introduced. The categorical data is processed into a dense vector with relevant relationships. Second, in the improved T-LSTM neural network, the DBSCAN clustering algorithm is introduced. Through the DBSCAN density clustering method, the long operation sequence of the user is segmented into different short sequence blocks in a scientific way. Finally, in the improved T-LSTM neural network, the hyperparameter "time step" is no longer a fixed parameter. After the long operation sequence of the user is segmented by DBSCAN, many small sequence blocks are formed. All the operation sequences of a single sequence block are used as a complete training data and input into the T-LSTM for training. Through these optimization measures, the neural network can learn more important information and reduce the influence of sequences outside the business process at the same time.

[0164] Based on the characteristics of user behavior features and historical experimental results, this embodiment proposes a training mode and parameter setting scheme for the user behavior features of an improved T-LSTM model to optimize the efficiency and accuracy of model training.

[0165] In the initialization stage of the model, appropriate learning rates, gradient descent algorithms, and batch sizes were selected to ensure that the model could converge quickly during the initial training process while avoiding overfitting and underfitting problems. In addition, considering the characteristics of time series data, a time window sliding strategy was adopted in this paper, enabling the model to effectively handle changes in user behavior patterns at different time scales. After multiple actual tests and optimizations, an improved T-LSTM model training scheme proposed in this embodiment demonstrated excellent performance in multiple experimental scenarios, significantly improving the accuracy and training efficiency of the model.

[0166] After long-term actual tests on the initialization parameters of model training, a set of stable and efficient parameter configurations are obtained as follows:

[0167] (1) Number of Hidden Units: The number of hidden units determines the complexity and learning ability of the model. To balance the expressive power and computational efficiency of the model, the number of hidden units is usually selected between 50 and 200. After experimental verification, for the user behavior feature dataset, setting it to 128 hidden units can effectively capture the time-dependent relationships and behavior patterns in the data.

[0168] (2) Number of Epochs: The setting of the number of epochs needs to find a balance between the convergence speed of model training and preventing overfitting. After repeated experiments, setting 68 epochs is optimal. At the same time, through the early stopping mechanism, the number of epochs is dynamically adjusted. When the validation set error no longer decreases within multiple consecutive epochs, the training is stopped in advance to avoid overfitting.

[0169] (3) Loss Function: For classification tasks or regression tasks, the choice of loss function affects the optimization objective of the model. In this model, due to the time perception characteristics, the cross-entropy loss function is selected as the loss function.

[0170] (4) Dropout: Dropout is a commonly used regularization method to prevent overfitting. For a complex T-LSTM model, an appropriate Dropout ratio can improve the generalization ability of the model. Through experiments, it is determined that setting a dropout rate of 0.25 can effectively reduce overfitting while maintaining a relatively high model accuracy.

[0171] (5) Time-Aware Parameters: The key feature of the T-LSTM model is its time-awareness. The Time Interval Embedding parameter is very important in this embodiment, as it can well reflect the impact of time intervals. In this embodiment, a normalized time interval embedding layer is used, combined with a time weight mechanism to dynamically adjust the impact of time steps.

[0172] (6) Batch Size: The batch size determines the amount of data processed each time the model is updated. A larger batch can improve the parallel computing efficiency of the GPU / TPU, but an overly large batch may lead to out-of-memory errors or a decrease in the convergence speed. After experimentation, for the T-LSTM model, a batch size of 64 achieves a good balance between efficiency and accuracy.

[0173] (7) Learning Rate: The learning rate is an important parameter that controls the step size of model weight updates. If the learning rate is too high, the model may not converge; if it is too low, the training process will be very slow. To ensure a stable convergence process, the initial learning rate is set to 0.001. At the same time, a learning rate decay strategy can be combined to gradually reduce the learning rate during training, in order to finely adjust the model weights when approaching the optimal point.

[0174] (8) Optimizer: The optimizer is used to update the model's parameters to minimize the loss function. For time series data and complex T-LSTM structures, choosing the Adam optimizer (Adaptive Moment Estimation) has achieved good results. In the experiments of the T-LSTM model, the Adam optimizer has shown excellent stability and accuracy. In addition, if the data noise is large, the AdamW optimizer can be used, which introduces a weight decay mechanism to further prevent overfitting.

[0175] Finally, the training parameters obtained from historical test experiments are executed for blockchain-based on-chain evidence storage. Subsequent model training parameters can be directly obtained on the chain, ensuring the credibility of the training parameters and avoiding being tampered with due to human or non-human reasons. At the same time, if more optimal training parameters are obtained subsequently, they can also be updated to the chain.

[0176] Step 5: T-LSTM Model Training: The user behavior data encoded as feature vectors is input into the T-LSTM model to update the model parameters and perform model training. T-LSTM processes the user behavior sequence data based on time intervals by introducing a time perception mechanism. In this embodiment, the time perception of user behavior is mainly reflected in the time-aware forget gate, time-aware input gate, etc. Then, the conduction calculation of user data is performed through forward propagation, and the user behavior model parameters are updated using the backpropagation algorithm for cyclic training.

[0177] Let Δt be the interval between the current time and the previous time step, x t be the input at the current time step t, h t-1 be the output of the hidden layer at the previous time step, c t-1 be the cell at the previous time step, W, U, V be the weight matrices related to the input, hidden state, and time, b be the bias term, σ represents the Sigmoid function, Θ represents the element-wise penalty operation, and the necessary network structures and components of the T-LSTM model used to define the user behavior time series are as follows.

[0178] (1) Time Gate: The Time Gate is used to process user behavior sequence data with unequal time intervals. The traditional T-LSTM model assumes that the input data is collected at equal time intervals, but in reality, user behavior data often has irregular timestamps. The role of the Time Gate is to incorporate the time interval factor into the state update process of the T-LSTM cell. It allows the model to adjust its internal state not only based on the content of the input data but also based on the time difference between data points. Specifically, the Time Gate enables the model to learn the importance of different time intervals for sequence data by introducing additional time-sensitive parameters. The state transition description of the Time Gate is defined as follows:

[0179] T t = σ(U t h t-1 + W t x t + V t Δt + b t ).

[0180] where W t , U t , V t are the weight matrices related to the input, hidden state, and time of the Time Gate, b t is the bias term of the Time Gate, T t is the output of the Time Gate at the current time step, Δt is the interval between the current time and the previous time step, x t is the input at the current time step t, and h t-1 is the output of the hidden layer at the previous time step.

[0181] (2) Forget Gate: The function of the Forget Gate is similar to that in the standard T-LSTM. However, in the T-LSTM of the user behavior time series in this embodiment, the operation of the Forget Gate incorporates the influence of time intervals. In the standard T-LSTM, the Forget Gate is a key component responsible for determining how much information is discarded from the cell state. The Forget Gate controls the degree to which the cell state information from the previous time step is retained through a value (ranging from 0 to 1) generated by a σ activation function. In the T-LSTM of this embodiment, the operation mechanism of the Forget Gate is extended to accommodate irregular time intervals. Additionally, the Forget Gate of the T-LSTM in this embodiment takes into account the time difference between two consecutive time steps, which is not present in the standard T-LSTM. The time interval can be used as an additional input to adjust the weights of the Forget Gate. This means that if the time interval between two time steps is long, the Forget Gate may be more inclined to forget old state information because the old information may have become outdated or less relevant. Conversely, if the time interval is short, the Forget Gate may retain more recent information because it may still be highly relevant to the current prediction or decision. The implementation of the Forget Gate in this embodiment not only determines which information should be forgotten from the cell state but also considers the timeliness of this information, thereby enhancing the model's ability to process irregularly sampled or uneven time interval data. Its state transition is described as follows:

[0182] f t = σ(U f h t-1 + W f x t + V f Δt + b f ).

[0183] Where W f , U f , V f are the input, hidden state, and time-related weight matrices of the Forget Gate, b f is the bias term of the Forget Gate, f t is the output of the Forget Gate at the current moment, Δt is the interval between the current time and the previous time step, x t is the input at the current moment t, and h t-1 is the output of the hidden layer at the previous moment.

[0184] (3) Input Gate: In the T-LSTM used in this embodiment, the operation of the input gate is also based on the Sigmoid activation function, but it also takes into account the time information additionally, that is, the interval between the current time step and the previous time step. That is, the input gate not only evaluates the importance of the current input, but also associates the time interval at which this input arrives, which has a very obvious improvement effect on processing the user behavior sequence data with irregular time intervals in this embodiment. Its state transition is described as follows:

[0185] i t = σ(U i h t-1 + W i x t + V i Δt + b i ).

[0186] Among them, W i , U i , V i are the weight matrices related to the input, hidden state, and time of the input gate, b i is the bias term of the input gate, i t is the output of the input gate at the current moment, Δt is the interval between the current time and the previous time, x t is the input at the current moment t, and h t-1 is the output of the hidden layer at the previous moment.

[0187] (4) Candidate Cell State: In the T-LSTM algorithm used in this embodiment, the candidate cell state is a key component for updating the cell state. In the standard T-LSTM, the candidate cell state is calculated from the input data at the current time step and the hidden state at the previous time step through a series of neural network layers. In the T-LSTM used in this embodiment, additional time interval information will be added to this process, so as to more accurately process data with irregular time intervals. Its state transition is described as follows:

[0188]

[0189] Among them, W c , U c , are the weight matrices related to the input and hidden state of the candidate cell state, b c is the bias term of the candidate cell state, C t is the output of the candidate cell state at the current moment, x t is the input at the current moment t, and h t-1 is the output of the hidden layer at the previous moment.

[0190] (5) Update Cell State: Updating the cell state is a crucial step in the T-LSTM algorithm used in this embodiment to ensure that the model can handle time series data. In the face of irregular time intervals in the user behavior sequence, this update mechanism enables the model to more accurately capture and utilize the temporal dependencies in the sequence. Its state transition is described as follows:

[0191]

[0192] Among them, f t is the output of the forget gate at the current time, i t is the output of the input gate at the current time, C t is the output of the cell state candidate value at the current time, c t-1 is the output of the updated cell state at the previous time, c t is the output of the updated cell state at the current time.

[0193] (6) Output Gate: In the T-LSTM algorithm used in this embodiment, the role of the output gate is similar to that in the standard T-LSTM, but includes an additional time perception mechanism. The main function of the output gate is to control the user behavior information extracted from the cell state and determine how this part of the information affects the hidden state at the next time step. Its state transition is described as follows:

[0194] o t = σ(U o h t-1 + W o x t + V o Δt + b o ).

[0195] Among them, W o , U o , V o are the input, hidden state, and time-related weight matrices of the output gate, b o is the bias term of the output gate, O t is the output of the output gate at the current time, Δt is the interval between the current time and the previous time, x t is the input at the current time t, h t-1 is the output of the hidden layer at the previous time.

[0196] (7) Current Hidden State: The hidden state is a very important component in the T-LSTM model used in this embodiment. It carries the information of the current time step and is passed as input to the next time step. The hidden state can be regarded as the "memory" of the model during the sequence processing. It not only contains the input information of the current user behavior but also integrates the information of previous time steps. The current hidden state in T-LSTM is the information transmitter, the direct source of output, the carrier of time-aware information, and a key part of the model's internal state update and feedback mechanism. The hidden state based on the cell state and the output gate can be expressed as:

[0197] h t = o t Θtanh(c t ).

[0199] Where, O t is the output of the output gate at the current moment, c t is the output of the updated cell state at the current moment, and h t is the output of the hidden layer at the current moment.

[0200] Then, the T-LSTM model of this embodiment is composed of the above components linked together. The user behavior data encoded as feature vectors is input into the T-LSTM model of this embodiment, and the model parameters are continuously updated. The T-LSTM model of this embodiment relies on its time-aware mechanism, which not only pays attention to the order of behaviors but also can capture the time intervals between behaviors. Therefore, the data of each time step contains specific time embedding features.

[0201] The feature vectors enter the T-LSTM cell through forward propagation, and the model extracts temporal dependency information through the memory cell and the gating mechanism. Next, the predicted value output by the model is compared with the true label, and the error is calculated through a loss function (such as cross-entropy or mean squared error). After the error calculation, the error is propagated layer by layer from the output layer back to the input layer using the backpropagation algorithm to optimize the weight and bias parameters of the model. With continuous batch training, the model gradually updates the parameters, learns and strengthens the ability to recognize user behavior patterns, and finally improves the prediction accuracy and generalization performance.

[0202] Finally, the model parameters obtained through training are stored on the blockchain for certification. Subsequently, the model can be directly obtained on the chain to ensure the credibility of the model and avoid being tampered with due to human or non-human reasons. At the same time, if a better model is obtained subsequently, it can also be updated to the chain.

[0203] Step 6: Abnormal behavior sequence detection: A model of historical behavior sequence patterns is trained according to the above steps. The newly generated operation behavior dataset of the user is integrated into the input data of the model in the way of "constructing feature vectors". When making a prediction, the T-LSTM will try to reconstruct the input operation sequence. If there is a large error between the output of the model and the actual input, it indicates that the operation sequence contains abnormal behavior. For the calculation of the reconstruction error, the "cosine similarity" method is used to calculate the similarity between the two. The cosine similarity is described as follows:

[0204] Suppose the vectors of users A and B are: U A =(A1, A2, …, A n ) and U B =(B1, B2, …, B n ), and the user feature similarity is:

[0205]

[0206] where the inner product Euclidean norm

[0207] Extract the operation sequences with large reconstruction errors (such as reconstruction error > 0.3), and convert the abnormal operation sequences into the user's actual operation logs by using the reverse conversion method.

[0208] Step 7: Abnormal risk rating: Extract the abnormal operation dataset output in Step 6, group it by "user ID", calculate indicators such as the number of abnormal operation sequences and the number of sensitive operation involved of the user, and finally use the risk rating method to rate the user's risk behavior.

[0209] (1) Risk rating dimensions: Use different dimensions such as "number of abnormal operation sequences", "number of abnormal operation sequence types", "number of sensitive operations involved in abnormal operations", and "proportion of abnormal operation sequences" to comprehensively evaluate the risk level of the user's operation behavior.

[0210] (2) iForest abnormal user detection: In the initial stage of model deployment, use the isolated forest abnormal detection method to conduct a secondary evaluation of the users who generate abnormal operation sequences, and accurately alarm abnormal users to a certain extent. Integrate the users alarmed by T-LSTM into a "user risk rating" dataset according to the four risk rating dimensions. Use the iForest algorithm for abnormal detection with the dataset as the input, and send the result set output by the model to the business side for verification.

[0211] (3) Risk rating method: The output result in step (2) will be presented in an HTML page. The business side can view the users with abnormal alerts on this page and mark the accuracy of the alerted users (accurate, false alarm) on this page. Relying on the business feedback information, black and white samples of users will be obtained. For the above data dimensions, the improved logistic regression (LR) method of this embodiment is used to calculate the risk coefficient of the abnormal operations of users (0-1). The higher the value, the more abnormal the user behavior. Define users with a risk value from 0.8 to 1 as high-risk users, 0.5-0.79 as medium-risk users, and below 0.5 as low-risk users.

[0212] The improved logistic regression method of this embodiment is as follows: This embodiment improves the method of model training applied to the logistic regression algorithm, combines the similarity of the training data itself and the feedback mechanism, and improves the accuracy of the training model. Assume that the original training data is The training process is as follows:

[0213] Step <1>: Based on the current block hash of the blockchain in X0 as the random number seed, randomly extract training data And calculate the distance from other training data according to the cosine similarity calculation method in "Step 6: Abnormal behavior sequence detection" above.

[0214] Step <2>: Take the k data with the largest distance from to form a training set

[0215] Step <3>: Use X1 to train the logistic regression model LR. The training uses the SGD algorithm, and the loss function is defined as follows:

[0216]

[0217] where k represents the number of training samples, l i represents the true label (0 or 1) of the i-th sample, p i represents the probability that the i-th sample is predicted to belong to the positive sample, λ represents the hyperparameter of the regularization strength, and ‖w‖ 2 represents the square of the L2 norm of the weight vector w.

[0218] Step <4>: Use the trained LR model to discriminate the data set X3 = X0 - X1, screen out the misjudged data, and form a set

[0219] Step <5>: Let X0 = X4, and go back to step <1> to continue training the model until ‖X4‖ < k.

[0220] Finally, the risk rating results are stored on the blockchain for subsequent verification through the blockchain to ensure the credibility of the model and prevent tampering due to human or non-human reasons.

[0221] Step 8: Anomaly analysis and display: Based on the above steps, the information of abnormal users is obtained, and then the specific abnormal operation sequences and corresponding risk levels are presented in the form of operation sequences.

[0222] Compared with the prior art, the abnormal behavior analysis method described in this embodiment has the following advantages:

[0223] First, an improved time-aware long short-term memory network (T-LSTM) is used to detect and analyze the abnormal behavior sequences of users. During the model training process of traditional T-LSTM, the "time step" is a fixed hyperparameter. This setting will break the original complete business chain during model training, resulting in the network remembering some information outside the business process and forgetting some important information within the business process. This proposal presents an improved algorithm for T-LSTM: using clustering to segment the user behavior sequences into different business blocks according to the user's behavior characteristics in a self-learning manner. The operation behaviors within each business block will be integrated into a time series data set and input into the T-LSTM algorithm.

[0224] Second, multiple algorithms are integrated, and a business feedback channel is provided to address the issues of early warning accuracy and false alarm rate at different stages of the model. First, at the initial stage of algorithm deployment, for the abnormal users output by the neural network, the unsupervised isolation forest algorithm (iForest) is used to detect the degree of abnormality of the user behavior characteristics, and users with a higher degree of abnormality are output for early warning, effectively reducing the number of abnormal result sets output by T-LSTM. Second, after the abnormal user alarm, the business side can feedback the accuracy of the model alarm by marking on the page, and the system will collect the feedback information from the business side and record it in the black and white sample library. Finally, after the business feedback information accumulates to a certain level, the supervised logistic regression algorithm (LR) is used to perform risk rating on the abnormal users, effectively reducing the false alarm rate of the model.

[0225] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the abnormal behavior analysis method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0226] This application also provides an abnormal behavior analysis device. Please refer to Figure 12 and the abnormal behavior analysis device includes:

[0227] An acquisition module 10, configured to acquire operation behavior feature vectors;

[0228] The recognition module 20 is configured to perform anomaly recognition on the operation behavior feature vector based on a pre-trained anomaly recognition model, and obtain an anomaly recognition result. The anomaly recognition model is obtained by improving the time perception of a preset time series long short-term memory network;

[0229] The analysis module 30 is configured to obtain an abnormal behavior analysis result based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the anomaly recognition result.

[0230] The abnormal behavior analysis device provided by this application adopts the abnormal behavior analysis method in the above embodiment, and can solve the technical problems that the existing abnormal behavior analysis method cannot perform anomaly detection from the overall operation behavior sequence of the user and the accuracy of anomaly detection is relatively low. Compared with the prior art, the beneficial effects of the abnormal behavior analysis device provided by this application are the same as those of the abnormal behavior analysis method provided by the above embodiment, and other technical features in the abnormal behavior analysis device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0231] This application provides an abnormal behavior analysis device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the abnormal behavior analysis method in Embodiment 1 above.

[0232] Reference is made below to Figure 13 , which shows a schematic structural diagram of an abnormal behavior analysis device suitable for implementing the embodiments of this application. The abnormal behavior analysis device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 13 The abnormal behavior analysis device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of this application.

[0233] As Figure 13As shown, the abnormal behavior analysis device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the abnormal behavior analysis device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display, a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the abnormal behavior analysis device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an abnormal behavior analysis device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.

[0234] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0235] The abnormal behavior analysis device provided by the present application adopts the abnormal behavior analysis method in the above embodiments, and can solve the technical problems that the existing abnormal behavior analysis methods cannot perform abnormal detection from the overall operation behavior sequence of users and the accuracy of abnormal detection is relatively low. Compared with the prior art, the beneficial effects of the abnormal behavior analysis device provided by the present application are the same as those of the abnormal behavior analysis method provided by the above embodiments, and other technical features in the abnormal behavior analysis device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0236] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0237] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0238] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal behavior analysis method in the above embodiments.

[0239] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0240] The above computer-readable storage medium can be included in the abnormal behavior analysis device; it can also exist separately and not be assembled into the abnormal behavior analysis device.

[0241] The above computer-readable storage medium carries one or more programs, which, when executed by an abnormal behavior analysis device, cause the abnormal behavior analysis device to: obtain an operation behavior feature vector; based on a pre-trained abnormal recognition model, perform abnormal recognition on the operation behavior feature vector to obtain an abnormal recognition result, where the abnormal recognition model is obtained by improving the time perception of a preset temporal long short-term memory network; based on a preset isolation forest abnormal detection algorithm and a pre-constructed logistic regression algorithm, obtain an abnormal behavior analysis result according to the abnormal recognition result.

[0242] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0243] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0244] The modules involved in the embodiments of the present application may be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0245] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal behavior analysis method, which can solve the technical problems that the existing abnormal behavior analysis methods cannot perform abnormal detection from the overall operation behavior sequence of users and the accuracy of abnormal detection is relatively low. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the abnormal behavior analysis method provided by the above embodiment, and will not be elaborated here.

[0246] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the abnormal behavior analysis method as described above. The computer program product provided by this application can solve the technical problems that the existing abnormal behavior analysis methods cannot perform abnormal detection from the overall operation behavior sequence of users and the accuracy of abnormal detection is relatively low. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the abnormal behavior analysis method provided by the above embodiment, and will not be elaborated here.

[0247] The above are only partial embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made by using the content of the specification and drawings of this application under the technical concept of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.

Claims

1. An abnormal behavior analysis method, characterized in that, The method includes: Obtaining an operation behavior feature vector; Based on a pre-trained anomaly recognition model, performing anomaly recognition on the operation behavior feature vector to obtain an anomaly recognition result, where the anomaly recognition model is obtained by improving the time perception of a preset temporal long short-term memory network; Based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm, obtaining an abnormal behavior analysis result according to the anomaly recognition result.

2. The method according to claim 1, wherein Before the step of performing anomaly recognition on the operation behavior feature vector based on a pre-trained anomaly recognition model to obtain an anomaly recognition result, it further includes: Obtaining an initial training set and model initialization parameters, where the initial training set includes a number of operation sequence training sets; Configuring the temporal long short-term memory network according to the model initialization parameters to obtain a model to be trained; Based on a preset forward propagation mechanism, inputting any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value; Based on a preset loss function, calculating an error between the training set prediction value and the operation sequence label of the operation sequence training set to obtain a prediction error value; Based on a preset backpropagation algorithm, optimizing the parameters of the model to be trained according to the prediction error value to obtain a model to be trained with optimized parameters; Return to execute the step: Based on a preset forward propagation mechanism, inputting any one of the operation sequence training sets into the model to be trained for prediction to obtain a training set prediction value, until all operation sequence training sets are traversed, and using the model to be trained with optimized parameters as the anomaly recognition model.

3. The method according to claim 2, characterized in that, The step of obtaining the initial training set includes: Collecting user behavior logs; Preprocessing the user behavior logs to obtain first log data; Performing vector conversion on the first log data to obtain a user behavior feature vector; Based on a preset clustering algorithm, performing temporal slicing on the user behavior feature vector method to obtain an operation feature data set; Converting the operation feature data set into a multi-dimensional array to obtain the initial training set.

4. The method according to claim 1, wherein The step of performing anomaly recognition on the operation behavior feature vector based on a pre-trained anomaly recognition model to obtain an anomaly recognition result includes: Inputting the operation behavior feature vector into the anomaly recognition model to obtain a behavior prediction value; Based on a preset cosine similarity calculation algorithm, calculating the similarity between the operation behavior feature vector and the behavior prediction value to obtain a user feature similarity; Obtaining the anomaly recognition result according to the user feature similarity.

5. The method according to claim 1, characterized in that, The step of obtaining an abnormal behavior analysis result based on a preset isolation forest anomaly detection algorithm and a pre-constructed logistic regression algorithm according to the anomaly recognition result includes: When the anomaly recognition result indicates that the user feature similarity is greater than a preset error threshold, based on a preset reverse conversion algorithm, converting the operation behavior feature vector to obtain a user behavior log; Based on a preset risk rating dimension, performing risk assessment on the user behavior log to obtain a user risk rating data set; Based on the isolated forest anomaly detection algorithm, perform anomaly monitoring on the user risk rating dataset to obtain an anomaly detection result; Obtain user abnormal operations according to the anomaly detection result; Based on the logistic regression algorithm, calculate the risk coefficient of the user abnormal operations to obtain an abnormal behavior analysis result.

6. The method according to claim 1, wherein Before the step of obtaining the abnormal behavior analysis result according to the anomaly recognition result based on the preset isolated forest anomaly detection algorithm and the pre-constructed logistic regression algorithm, the following steps are further included: Obtain the original training dataset; Train the preset logistic regression model according to the original training dataset to obtain the logistic regression algorithm.

7. The method according to claim 6, wherein The step of training the logistic regression algorithm according to the original training dataset includes: S1: Based on a preset random algorithm, extract any original training data in the original training dataset as the standard training data, and use the original training dataset excluding the standard training data as the first training dataset; S2: Train the logistic regression model according to the first training dataset and the standard training data to obtain a second training dataset; S3: Judge whether the determinant of the second training dataset is less than a preset determinant threshold: S4: If not, define the second training dataset as the original training dataset; Repeat steps S1 - S4 until the determinant of the second training dataset is less than the preset threshold, and obtain the logistic regression algorithm according to the trained logistic regression model.

8. The method according to claim 7, wherein The step of training the logistic regression model according to the first training dataset and the standard training data to obtain a second training dataset includes: Based on a preset cosine calculation algorithm, calculate the distance between the standard training data and the first training data in the first training dataset to obtain a training data distance set; Filter the first training dataset according to the training data distance set to obtain a third training dataset; Based on a preset gradient descent algorithm, train the logistic regression model according to the third training dataset; Obtain a fourth training dataset according to the original training dataset and the third training dataset; Discriminate the fourth training dataset through the trained logistic regression model to obtain a discrimination dataset; Filter the discrimination dataset to obtain the second training dataset.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the abnormal behavior analysis method according to any one of claims 1 to 8 are implemented.

10. A computer program product, characterized in that, The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the abnormal behavior analysis method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Online teaching platform-oriented service quality dynamic evaluation method

    CN121921153A