User abnormal behavior detection method and system based on isolated forest model

The anomaly detection system built through the isolated forest model solves the problems of low accuracy of malicious data crawling behavior recognition and slow calculation speed in the existing technology, and realizes efficient and accurate user abnormal behavior detection, reducing the system false alarm rate and computing resource occupation.

CN120415792APending Publication Date: 2025-08-01STATE GRID JIANGSU ELECTRIC POWER CO LTD MARKETING SERVICE CENT +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510494523.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing malicious data crawling behavior recognition method is based on a single data feature access frequency, resulting in low recognition accuracy and slow calculation speed, which cannot meet the needs of large business volumes.

Method used

The user abnormal behavior detection method based on the isolated forest model is adopted. By obtaining the historical log data, network traffic data and user identity verification data of the security center, pre-processing, feature screening, standardization and normalization processing is carried out, an abnormality detection model is constructed, effective features are screened using correlation coefficients, and training is carried out through the isolated forest algorithm to achieve real-time abnormality detection.

Benefits of technology

It improves the accuracy and speed of abnormal behavior detection, reduces system false alarms, reduces computing resource occupation and manual labeling costs, and enhances system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120415792A_ABST
    Figure CN120415792A_ABST
Patent Text Reader

Abstract

The invention discloses a user abnormal behavior detection method and system based on an isolated forest model. The method comprises the following steps: a system obtains information such as service operation log data, network flow data and user identity verification data of a security center; secondly, preprocessing the acquired business operation log data; and inputting the training sample set into an isolated forest model for training. And performing anomaly detection on the real-time user business operation log data based on the trained isolated forest model. According to the scheme, the business operation log data preprocessing mode is improved, the detection accuracy of the system is improved, the detection accuracy of the system is improved on the premise of occupying relatively low computing resources, the manual labeling cost is reduced, and the safety of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and particularly relates to a method and system for detecting user abnormal behavior based on an isolation forest model. Background Art

[0002] At present, there are strict control requirements for RPA process robots and applications used in information systems. All RPA and applications must be applied for and put on record, and cannot be used randomly. Therefore, when a terminal program similar to the program behavior pattern is found but not in the record ledger, it is necessary to promptly determine and process such programs to prevent malicious programs from stealing data.

[0003] The existing method for identifying malicious data crawling behavior is based on the usage frequency of interface calls, and a global unified threshold is set. The defect of this method is that it only distinguishes malicious data crawling behavior from normal data acquisition behavior based on a single data feature, that is, the interface access frequency, which will cause a large number of false alarms, and the calculation speed of ordinary abnormal data identification methods is relatively slow, unable to meet the requirements of a large volume of business. Based on this, there is an urgent need to provide an abnormal behavior detection system with a fast calculation speed and high identification accuracy. Summary of the Invention

[0004] In order to solve the deficiencies existing in the prior art, the present invention provides a method and system for detecting user abnormal behavior based on an isolation forest model to solve the technical problems of low accuracy in distinguishing abnormal behavior based on a single data feature, i.e., access frequency, and slow calculation speed.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions.

[0006] The present invention first discloses a method for detecting user abnormal behavior based on an isolation forest model, and the method includes the following steps:

[0007] Step 1: Obtain historical log data, network traffic data, and user identity authentication data of the security center;

[0008] Step 2: Preprocess the obtained log data, convert it into feature data corresponding to the log data, perform effective feature screening from the feature data using the correlation coefficient, perform standardization and normalization processing on the screened feature data, and label the feature data to obtain a training sample set;

[0009] Step 3: Input the training sample set into the isolation forest model for training to obtain an anomaly detection model;

[0010] Step 4: Perform real-time anomaly detection on the to-be-detected log data collected in real time based on the trained anomaly detection model, and output an anomaly detection result.

[0011] The present invention further includes the following preferred solutions:

[0012] The log data includes an operation ID, the account number of the operator, the organizational path of the operator, the operating system, the IP address, the operation time, the operation type, the operation content, the operation duration, the sensitivity level, and the operation certificate.

[0013] Preprocessing the obtained log data, converting it into feature data corresponding to the log data, screening effective features from the feature data using the correlation coefficient, performing standardization and normalization processing on the screened feature data, and labeling the feature data to obtain a training sample set, further including:

[0014] Cleaning the collected log data, including removing duplicate records, handling missing values, and correcting outlier data;

[0015] Performing feature extraction based on the cleaned log data set, extracting the behavioral features describing interface access in the set, including business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, to obtain an effective feature data set;

[0016] Performing standardization and normalization processing on the extracted effective feature data set to obtain a processed standardized feature data set;

[0017] Based on a sample set corresponding to the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, labeling the feature data in the sample set to obtain a training sample set with positive and negative sample information for abnormal detection model training.

[0018] The performing feature extraction based on the cleaned log data set, extracting the behavioral features describing interface access in the set, including business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, to obtain an effective feature data set, further including:

[0019] Parsing the cleaned log data set to obtain a structured log data set D of the system;

[0020] Performing feature value extraction on the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, i.e., the identity ID, corresponding to the behavioral features in the log data set D to obtain a feature data set; wherein, the feature value extraction includes the feature value extraction of the mean value, effective value, peak value, root mean square amplitude, waveform index, impulse index, and kurtosis index.

[0021] Screen out effective features from the feature data set using the correlation coefficient. Compare the calculated correlation coefficient with the correlation threshold. If it is higher than the correlation threshold, it is considered valid data; if it is lower than the correlation threshold, it is considered invalid data and is excluded to obtain an effective feature data set. The operation formula of the correlation coefficient C is as follows:

[0022]

[0023] where x i is a feature value of any type in the log data, is the arithmetic mean of any type of feature values x1, x2,..., x n in the log data, y i is the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index corresponding to any type of feature value in the log data; is the arithmetic mean of the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index y1, y2,..., y n corresponding to any type of feature value in the log data; N is the number of times of log data collection; x 1k is the estimated effective quantity of any type of feature value, x 2k is the overall feature value quantity of any type of feature value in this category, and p is the preset importance degree of the index type;

[0024] where any type can be any one of service interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information.

[0025] The parsing of the cleaned log data set further includes:

[0026] Adopt a data mining method based on clustering or a parsing method based on natural language processing for log parsing. The data mining method based on clustering includes SLCT, LogCluster, IPLoM, LKE, LogSig, and the parsing method based on natural language processing includes Logram, NLog.

[0027] The standardization and normalization processing of the extracted effective feature data set to obtain a processed standardized feature data set further includes:

[0028] Use the mean and mean absolute deviation of the original data to standardize the effective feature data set; the process of standardizing the original value X i of the data to X i ′ is expressed as:

[0029]

[0030]

[0031] Among them, AVG is the mean of the original data; STAD is the average absolute deviation of the original data. If the mean AVG is equal to 0 or the average absolute deviation is equal to 0, then X i ′ is equal to 0;

[0032] Based on the data in the data set E, the feature data is normalized, and each normalized value is normalized to the interval [0, 1]. Let X″ be the value after X ′ is normalized, X min is X ′ is the minimum value of X max is X ′ is the maximum value of X; n is the amount of data, and the sample set G is obtained; the normalization process is as follows:

[0033]

[0034] Inputting the training sample set into the isolation forest model for training to obtain an anomaly detection model further includes:

[0035] S31: Randomly select n points from the training sample set H as a sample subset and put them into the root node of the tree;

[0036] S32: Randomly specify a sample feature and randomly generate a cut point p in the current node data;

[0037] S33: Generate a hyperplane with this cut point p, and then divide the current node data space into 2 subspaces, specifically including putting the data less than p in the specified sample feature in the left child node of the current node, and putting the data greater than or equal to p in the right child node of the current node;

[0038] S34: Recursively perform step S32 and step S33 in the child nodes, and iteratively construct new child nodes until there is only one data in the child node or the child node has reached the specified height;

[0039] S35: Loop through step S31 to step S34 until T isolation trees are generated;

[0040] S36: After obtaining T isolation trees, complete the isolation forest training.

[0041] The present invention also discloses a user abnormal behavior detection system based on the isolation forest model using the foregoing user abnormal behavior detection method based on the isolation forest model, including:

[0042] A historical log data acquisition module for acquiring historical log data, network traffic data, and user authentication data of the security center;

[0043] A log preprocessing module for preprocessing the acquired log data, converting it into feature data corresponding to the log data, effectively screening features from the feature data using the correlation coefficient, performing standardization and normalization processing on the screened feature data, and annotating the feature data to obtain a training sample set;

[0044] A model training module for inputting the training sample set into an isolation forest model for training to obtain an anomaly detection model;

[0045] A log data anomaly detection module for performing real-time anomaly detection on the log data to be detected collected in real time based on the trained anomaly detection model and outputting an anomaly detection result.

[0046] Correspondingly, the present application also discloses a terminal, including a processor and a storage medium;

[0047] The storage medium is used for storing instructions;

[0048] The processor is used for operating according to the instructions to execute the steps of the user anomaly behavior detection method based on the isolation forest model described above.

[0049] Correspondingly, the present application also discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the user anomaly behavior detection method based on the isolation forest model described above.

[0050] The beneficial effects of the present invention are that, compared with the prior art, the present invention provides a user anomaly behavior detection method and system based on an isolation forest model, which improves the way of preprocessing business operation log data, including increasing the dimensions of feature extraction such as business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, etc. By using composite features, the system false alarms are reduced, the detection accuracy of the system is improved, and some easily identifiable abnormal logs can be screened out in advance, reducing the loss of the subsequent model and improving the detection speed of the system. The present invention introduces the isolation forest algorithm in machine learning into the anomaly behavior detection system, which can improve the detection accuracy of the system on the premise of occupying lower computing resources, and reduces the cost of manual annotation compared with supervised and semi-supervised machine learning, improving the security of the system. Description of the Drawings

[0051] Figure 1 It is a flowchart of the user anomaly behavior detection method based on the isolation forest model in the present invention.

[0052] Figure 2 It is a module diagram of the user abnormal behavior detection system based on the isolation forest model in the present invention.

[0053] Figure 3 It is an architecture diagram of a computer terminal for implementing the method of the present invention. Detailed implementation manners

[0054] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0055] The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0056] In recent years, with the continuous improvement of hardware performance, machine learning has become increasingly popular in the industrial community and has shown great potential in the field of abnormal log data detection. Machine learning learns the rules from data and predicts future data according to the obtained rules. It trains the model through big data, can extract effective features from the data, and automatically generates detection rules. These rules are saved in the trained model in the form of parameters and do not require manual intervention. Using machine learning not only improves the detection accuracy of the system, but also greatly reduces the development difficulty.

[0057] In view of the deficiencies of the prior art, the present invention draws on the isolation forest algorithm in the unsupervised machine learning model to implement a high-accuracy and real-time user log abnormal detection system. The system obtains information such as business operation log data, network traffic data, and user identity authentication data of the security center; secondly, preprocesses the obtained business operation log data; inputs the training sample set into the isolation forest model for training. Finally, abnormal detection is performed on the real-time user business operation log data based on the trained isolation forest model.

[0058] See Figure 1 As shown, the method for detecting user abnormal behavior based on the isolation forest model disclosed in the present invention includes the following steps:

[0059] Step 1: Obtain the historical log data, network traffic data, and user identity authentication data of the security center.

[0060] The log collection module of the user abnormal behavior detection system is responsible for receiving, integrating, and storing the historical log data of the security center. First, it obtains the log data of the security center. The content of the log data may include operation ID, operator's account, operator's organizational path, operating system, IP address, operation time, operation type, operation content, operation duration, sensitivity level, operation certificate, etc.

[0061] Step 2: Preprocess the obtained log data, convert it into feature data corresponding to the log data, use the correlation coefficient to screen effective features from the feature data, perform standardization and normalization processing on the screened feature data, and label the feature data to obtain a training sample set.

[0062] The log processing module of the user abnormal behavior detection system performs preprocessing on the obtained log data through data cleaning, including removing duplicate records, handling missing values, and correcting outlier data, and outputs feature data corresponding to features such as time period, frequency, and operation sequence. The specific steps are as follows:

[0063] S21: Clean the collected log data, including removing duplicate records, handling missing values, and correcting outlier data. The step S21 further includes:

[0064] S211: Remove duplicate records:

[0065] Remove duplicate records based on the fast log matching algorithm of secondary query to obtain the log data set A.

[0066] Generate a unique hash value as a key according to the combination of the log source IP address, source port, destination IP address, and destination port, key = hash(src_ip, src_port, dst_ip, dst_port). Where src_ip, src_port, dst_ip, and dst_port are the source IP address, source port, destination IP address, and destination port respectively.

[0067] If the hash value does not exist, the log must not be repeated and proceed directly to subsequent processing; if the hash value exists, further check the MD5 value of the log content and search in the MD5 linked list of the log content; if the MD5 value does not exist, the log is not repeated and proceed to subsequent processing; if the MD5 value exists, the log is a duplicate log and is discarded.

[0068] S212: Handle missing values:

[0069] Based on the log data set A, process the missing values in the set. When there are missing values in the data, methods such as filling and deletion are used for processing. The filling methods include mean, median, mode, interpolation, etc., to obtain the log data set B.

[0070] S213: Outlier data correction:

[0071] Based on the log data set B, correct the outlier data in the set, that is, calculate the sample standard deviation s and the population standard deviation σ. When s≥3σ, the standard deviation method is used to correct the data, that is, the log data exceeding 3 times the standard deviation is corrected to 3 times the standard deviation, to obtain the log data set C. The calculation formulas for the sample standard deviation and the population standard deviation are as follows:

[0072] Sample standard deviation:

[0073] Population standard deviation:

[0074] Using the above means, there is no need for complex algorithms or parameter tuning. After correction, the noise in the data set is reduced, the overall distribution is more concentrated, the sensitivity of the model to the data is reduced, the prediction results are more reliable, it is suitable for processing large-scale data, and the efficiency is improved.

[0075] S22: Feature extraction is performed based on the cleaned log data set, and the behavioral features that can describe interface access are extracted from the set, including business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature (identity ID) information, etc. The step S22 further includes:

[0076] S221: Parse the cleaned log data set to obtain the structured log data set D of the system.

[0077] Exemplarily, log parsing can adopt data mining methods based on clustering or parsing methods based on natural language processing. Data mining methods based on clustering include SLCT, LogCluster, IPLoM, LKE, LogSig, etc., and parsing methods based on natural language processing include Logram, NLog, etc.

[0078] S222: Extract the eigenvalue of the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, that is, identity ID, etc. corresponding to the behavioral features in the log data set D to obtain the feature data set.

[0079] Among them, the eigenvalue extraction includes the eigenvalue extraction of mean, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index.

[0080] S223: Screen out effective features from the feature data set using the correlation coefficient to obtain an effective feature data set.

[0081] Perform the operation of the correlation coefficient on the feature values extracted in step S222 and the original log data respectively. Compare the calculated correlation coefficient with the correlation threshold. If it is higher than the correlation threshold, it is considered valid data; if it is lower than the correlation threshold, it is considered invalid data and is excluded. Among them, the operation formula of the correlation coefficient C is as follows:

[0082]

[0083] Among them, x i is any type of feature value in the log data, is the arithmetic mean of any type of feature values x1, x2,..., x n in the log data, y i is the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index, etc. corresponding to any type of feature value in the log data; is the arithmetic mean of the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index y1, y2,..., y n corresponding to any type of feature value in the log data; N is the number of times of log data collection. x 1k is the estimated effective quantity of any type of feature value, x 2k is the overall feature value quantity of this type for any type of feature value, and p is the preset importance degree of the index type.

[0084] Features with high correlation may contain more prediction information. After excluding low-correlated or irrelevant features, the model training complexity is reduced. Using the correlation coefficient method can retain the signal with the strongest correlation with the target variable, reduce noise interference, improve the model prediction accuracy, reduce the tendency of overfitting to the training data, and improve the generalization ability.

[0085] Among them, any type can be any one of the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information (identity ID).

[0086] Compare the calculated correlation coefficient C with the preset empirical value threshold C'. The features that do not exceed the correlation coefficient threshold are effective features, and an effective feature data set E is obtained.

[0087] S23: Perform standardization and normalization processing on the extracted effective feature data set to obtain a processed standardized feature data set.

[0088] The log standardization and normalization module of the system obtains feature data. Through data standardization and normalization processing, it ensures the unified scale between different features and outputs a training sample set for the anomaly detection model. The step S23 further includes:

[0089] S231: Standardize the effective feature data set E using the mean and mean absolute deviation of the original data. The original value X of the data i is standardized to X i ′ The process is shown in the following formula:

[0090]

[0091] where AVG is the mean of the original data; STAD is the mean absolute deviation of the original data. When calculating, if the mean AVG is equal to 0 or the mean absolute deviation is equal to 0, then X i ′ is equal to 0.

[0092] S232: Based on the data in the data set E, perform normalization processing on the feature data to obtain the sample set G. Normalize each normalized value to the interval [0, 1]. Let X″ be the value after normalizing X ′ , X min is X ′ is the minimum value of X max is X ′ is the maximum value of X

[0093]

[0094] S24: Based on the sample set G that has a corresponding relationship with features such as the time period information, frequency information, operation sequence information, port network traffic information, and identity feature information of the service interface request, label the feature data in the sample set G to obtain a training sample set H with positive and negative sample information for anomaly detection model training.

[0095] Step 3: Input the training sample set into the isolation forest model for training to obtain an anomaly detection model.

[0096] According to a specific embodiment, step 3 may include the following steps:

[0097] S31: Randomly select n points from the training sample set H as a sample subset and put them into the root node of the tree;

[0098] S32: Randomly specify a sample feature and randomly generate a cut point p in the current node data;

[0099] S33: Generate a hyperplane at this cutting point, and then divide the current node data space into two subspaces. Specifically, place the data in the specified sample features that are less than p in the left child node of the current node, and place the data that is greater than or equal to p in the right child node of the current node;

[0100] S34: Recursively perform step S32 and step S33 in the child nodes, iteratively construct new child nodes until there is only one data in the child node or the child node has reached the limited height;

[0101] S35: Loop through step S31 to step S34 until T isolated trees are generated;

[0102] S36: After obtaining T isolated trees, the training of the isolated forest ends, and enter the prediction stage of the data.

[0103] The above model training process can efficiently process high-dimensional data. The construction of a single tree only needs to traverse part of the data, making the time complexity linear. It does not require labels and does not depend on the data obeying a specific distribution. It is more robust to skewed and non-uniformly distributed data, suitable for anomaly detection, and has strong scalability.

[0104] Step 4: Based on the trained anomaly detection model, perform real-time anomaly detection on the log data to be detected collected in real time, and output the anomaly detection results.

[0105] For example, the business operation logs can be collected every minute. Input the log data to be detected into the above constructed anomaly detection model, use this anomaly detection model for evaluation, score the relevant behavior features, and this score reflects the difference from the normal behavior. If this score exceeds the set threshold, it is determined that this behavior belongs to an abnormal behavior. In the case of detecting an abnormal behavior, send a warning notification to the relevant party. The warning notification can include detailed information about the abnormal behavior, such as the time, location, frequency, etc. of this behavior.

[0106] The beneficial effect of the present invention is that, compared with the prior art, the present invention provides a method and system for detecting user abnormal behaviors based on an isolated forest model, which improves the way of preprocessing business operation log data, including increasing the dimensions of feature extraction such as business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, etc. By compound features, the system false alarms are reduced, the detection accuracy of the system is improved, and some abnormal logs that are easily recognized can be screened out in advance, reducing the loss of the subsequent model and improving the detection speed of the system. The present invention introduces the isolated forest algorithm in machine learning into the anomaly behavior detection system, which can improve the detection accuracy of the system on the premise of occupying lower computing resources, and reduces the cost of manual annotation compared with supervised and semi-supervised machine learning, improving the security of the system.

[0107] The present invention may be a system, a method, and / or a computer program product. Refer to Figure 2 , the present invention also discloses a user abnormal behavior detection system based on the isolated forest model for the aforementioned user abnormal behavior detection method based on the isolated forest model, including:

[0108] A historical log data acquisition module, configured to acquire historical log data, network traffic data, and user authentication data of a security center;

[0109] A log preprocessing module, configured to preprocess the acquired log data, convert it into feature data corresponding to the log data, perform effective feature screening from the feature data by using a correlation coefficient, perform standardization and normalization processing on the screened feature data, and label the feature data to obtain a training sample set;

[0110] A model training module, configured to input the training sample set into an isolated forest model for training to obtain an anomaly detection model;

[0111] A log data anomaly detection module, configured to perform real-time anomaly detection on the to-be-detected log data collected in real time based on the trained anomaly detection model, and output an anomaly detection result.

[0112] Based on the spirit of the present invention, those skilled in the art can easily conceive that a computer program product can be obtained based on the aforementioned user abnormal behavior detection method based on the isolated forest model. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for causing a processor to implement various aspects of the present disclosure are uploaded. Refer to Figure 3 , that is, the present application also includes a terminal, including a processor and a storage medium; the storage medium is used for storing instructions; the processor is used for operating according to the instructions to execute the steps of the aforementioned user abnormal behavior detection method based on the isolated forest model.

[0113] Computer-readable storage medium can be the tangible device that can keep and store the instruction used by instruction execution device.Computer-readable storage medium can be, for example, but not limited to, electric storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or above-mentioned any suitable combination.The more specific example (non-exhaustive list) of computer-readable storage medium comprises: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical coding device, for example, punch card or the convex structure in the groove that stores instruction thereon and above-mentioned any suitable combination.Computer-readable storage medium used here is not interpreted as instantaneous signal itself, such as radio wave or other free propagating electromagnetic wave, electromagnetic wave (for example, by the light pulse of fiber optic cable) that waveguide or other transmission medium propagates or the electric signal that is transmitted by wire.

[0114] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0115] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for detecting user abnormal behavior based on the isolated forest model, characterized in that, It includes the following steps: Step 1: Obtain the historical log data, network traffic data, and user authentication data of the security center; Step 2: Preprocess the obtained log data, convert it into feature data corresponding to the log data, use the correlation coefficient to perform effective feature screening from the feature data, perform standardization and normalization processing on the screened feature data, and label the feature data to obtain a training sample set; Step 3: Input the training sample set into an isolation forest model for training to obtain an anomaly detection model; Step 4: Based on the trained anomaly detection model, perform real-time anomaly detection on the log data to be detected collected in real time, and output the anomaly detection result.

2. The user abnormal behavior detection method based on the isolation forest model according to claim 1, wherein The log data includes the operation ID, the account of the operator, the organizational path of the operator, the operating system, the IP address, the operation time, the operation type, the operation content, the operation duration, the sensitivity level, and the operation certificate.

3. The user abnormal behavior detection method based on the isolation forest model according to claim 2, characterized in that, The preprocessing of the obtained log data, converting it into feature data corresponding to the log data, using the correlation coefficient to perform effective feature screening from the feature data, performing standardization and normalization processing on the screened feature data, and labeling the feature data to obtain a training sample set further includes: Clean the collected log data, including removing duplicate records, handling missing values, and correcting outlier data; Based on the cleaned log data set, perform feature extraction, extract the behavioral features describing interface access in the set, including the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, to obtain an effective feature data set; Perform standardization and normalization processing on the extracted effective feature data set to obtain a processed standardized feature data set; Based on the sample set corresponding to the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, label the feature data in the sample set to obtain a training sample set with positive and negative sample information for anomaly detection model training.

4. The user abnormal behavior detection method based on the isolation forest model according to claim 3, wherein, The performing feature extraction based on the cleaned log data set, extracting the behavioral features describing interface access in the set, including the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, to obtain an effective feature data set further includes: Parse the cleaned log data set to obtain the structured log data set D of the system; Extract the feature values of the business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information, that is, the identity ID, corresponding to the behavioral features in the log data set D to obtain a feature data set; among them, the feature value extraction includes the feature value extraction of the mean value, effective value, peak value, root mean square amplitude, waveform index, impulse index, and kurtosis index. Screen out valid features from the set of feature data using the correlation coefficient. Compare the calculated correlation coefficient with the correlation threshold. If it is higher than the correlation threshold, it is considered valid data; if it is lower than the correlation threshold, it is considered invalid data and is excluded to obtain a set of valid feature data. The calculation formula for the correlation coefficient C is as follows: Among them, x i is an eigenvalue of any type in the log data, is any type of eigenvalue x1, x2,..., x in the log data n 's arithmetic mean, y i is the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index corresponding to any type of eigenvalue in the log data; is the mean value, effective value, peak value, root mean square amplitude, waveform index, pulse index, kurtosis index y1, y2,..., y corresponding to any type of eigenvalue in the log data n 's arithmetic mean; N is the number of times of log data collection; x 1k is the estimated effective quantity of any type of eigenvalue, x 2k is the overall eigenvalue quantity of this type for any type of eigenvalue, and p is the preset importance degree of the index type; Among them, any type can be any one of business interface request time period information, frequency information, operation sequence information, port network traffic information, and identity feature information.

5. The method for detecting abnormal user behavior based on the isolation forest model according to claim 4, wherein The parsing of the cleaned log data set further includes: Adopt a data mining method based on clustering or a parsing method based on natural language processing for log parsing. The data mining method based on clustering includes SLCT, LogCluster, IPLoM, LKE, LogSig, and the parsing method based on natural language processing includes Logram, NLog.

6. The user abnormal behavior detection system based on the isolation forest model according to claim 5, characterized in that, The normalization and standardization processing of the extracted set of valid feature data to obtain a processed set of standardized feature data further includes: Normalize the effective feature data set using the mean and mean absolute deviation of the original data; the process of normalizing the original value X of the data i to X i ′ is expressed as: Among them, AVG is the mean of the original data; STAD is the average absolute deviation of the original data. If the mean AVG is equal to 0 or the average absolute deviation is equal to 0, then X i ′ is equal to 0; Based on the data in data set E, perform normalization processing on the feature data, and normalize each standardized value to the interval [0, 1]. Let X″ be the value of X ′ after normalization, and X min is the minimum value of X ′ , and X max is the maximum value of X ′ ; n is the amount of data, and the sample set G is obtained; the normalization process is as follows:

7. The user abnormal behavior detection system based on the isolation forest model according to claim 6, characterized in that, The input of the training sample set into the isolation forest model for training to obtain an anomaly detection model further includes: S31: Randomly select n points from the training sample set H as a sample subset and place them at the root node of the tree; S32: Randomly specify a sample feature and randomly generate a cut point p in the current node data; S33: Generate a hyperplane with this cut point p, and then divide the current node data space into 2 subspaces. Specifically, it includes placing the data smaller than p in the specified sample feature in the left child node of the current node and placing the data greater than or equal to p in the right child node of the current node; S34: Recursively perform steps S32 and S33 in the child nodes, iteratively constructing new child nodes until there is only one data in the child node or the child node has reached the specified height; S35: Loop steps S31 to S34 until T isolation trees are generated; S36: After obtaining T isolation trees, complete the isolation forest training.

8. A user abnormal behavior detection system based on an isolation forest model, characterized in that It includes: A historical log data acquisition module for acquiring historical log data, network traffic data, and user identity authentication data of the security center; A log preprocessing module for preprocessing the acquired log data, converting it into feature data corresponding to the log data, screening valid features from the feature data using the correlation coefficient, performing normalization and standardization processing on the screened feature data, and annotating the feature data to obtain a training sample set; A model training module for inputting the training sample set into the isolation forest model for training to obtain an anomaly detection model; A log data anomaly detection module for performing real-time anomaly detection on the to-be-detected log data collected in real time based on the trained anomaly detection model and outputting an anomaly detection result.

9. A terminal, including a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the user abnormal behavior detection method based on the isolation forest model according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the user abnormal behavior detection method based on the isolation forest model according to any one of claims 1-7 are implemented.