Double-window adaptive semi-supervised network threat detection method

Through the dual-window adaptive semi-supervised network threat detection method, which utilizes semi-supervised learning combining multi-dimensional features and long and short time windows, the problems of insufficient real-time and noise resistance in existing technologies are solved, and efficient user abnormal behavior detection is achieved.

CN120639429APending Publication Date: 2025-09-12NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510944495.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies lack real-time performance in detecting abnormal user behavior, rely on predefined rules to cover new attack patterns, and machine learning methods have strong labeling dependence and insufficient noise resistance.

Method used

A dual-window adaptive semi-supervised network threat detection method is adopted. By defining multi-dimensional feature indicators, long and short time windows and a set of semi-supervised trained classifiers are used for anomaly detection. Combined with the tri-training machine learning method, pseudo labels are generated to improve detection flexibility and noise resistance.

Benefits of technology

It achieves efficient detection of abnormal network behavior, with a recall rate higher than 95% and an accuracy rate higher than 90%. It has strong noise resistance with small amounts of data and can identify risky users with similar behavior patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639429A_ABST
    Figure CN120639429A_ABST
Patent Text Reader

Abstract

The invention relates to a double-window adaptive semi-supervised network threat detection method, which comprises the following steps of: firstly, defining a multi-dimensional characteristic index, and respectively determining statistical thresholds in a long time window and a short time window according to a characteristic index distribution rule based on historical network flow data; and according to the statistical threshold value, the key concerned objects are screened. And carrying out anomaly detection on the focused object by using a semi-supervised trained classifier set. According to the method, the limitation of traditional single-sequence analysis can be broken through through multi-dimensional dynamic feature analysis; meanwhile, unsupervised primary screening is achieved through a double-time-window threshold model; and then in combination with semi-supervised integration, the implicit mode in the unlabeled data can be mined by using a small number of labeled samples, so that the problem of scarcity of labeled data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal behavior recognition, and more particularly to a dual-window adaptive semi-supervised network threat detection method. Background Art

[0002] In the field of information security, detecting abnormal user behavior is a key technology for preventing insider threats (such as "defector" attacks). Early research primarily relied on rule analysis and traditional modeling methods. For example, OKA et al. proposed modeling UNIX user command sequences using an eigenvalue co-occurrence matrix, identifying potentially malicious users by counting abnormal co-occurrence patterns of command combinations. Brdiczka et al. abstracted user behaviors into graph nodes, defining the relationships between behaviors as edges, and used graph topology anomalies to detect deviations in user behavior. Kandias et al., taking a social psychology perspective, constructed a dynamic psychological model using user operation logs, identifying risky users based on the anomalies of their behavioral motivations. Finally, Zhu et al. combined business logs to build a user behavior process model, detecting anomalies by detecting paths that deviate from the normal workflow.

[0003] However, the above methods lack real-time performance and either rely on predefined rules, making it difficult to cover new attack patterns; or they only focus on command sequences or static processes, ignoring collaborative anomalies of multi-dimensional behavioral characteristics.

[0004] With the development of machine learning technology, researchers have begun to turn to a data-driven detection paradigm: for example, Davison, Lane, and others identified anomalies by calculating the matching degree between user command sequences and historical patterns; to address the scarcity of abnormal samples, Zhou Qingsong and others used clustering algorithms to directly identify outliers in the behavioral feature space as anomalies; Liu Wei and others used data association rules to mine implicit abnormal patterns between entity behaviors, and Mo Fan and others combined machine learning to implement account anomaly detection.

[0005] Although machine learning methods have improved detection flexibility, they still have the drawbacks of strong labeling dependence and insufficient noise immunity. Therefore, how to overcome these shortcomings is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0006] In view of this, in order to at least partially solve the above problems, the present invention provides a dual-window adaptive semi-supervised network threat detection method. To achieve the above objectives, the present invention adopts the following technical solutions:

[0007] First, the present invention provides a dual-window adaptive semi-supervised network threat detection method, comprising:

[0008] Define multi-dimensional characteristic indicators, based on historical network flow data, and determine statistical thresholds in long and short time windows according to the distribution patterns of characteristic indicators;

[0009] Screening key targets based on statistical thresholds;

[0010] Anomaly detection is performed on key objects using an ensemble of semi-supervised trained classifiers.

[0011] In an optional embodiment, the multidimensional feature indicators include flow direction attributes (inflow / outflow), protocol type attributes (TCP / UDP / ICMP), service port number attributes (0-65535), and traffic indicator attributes (number of bytes / number of packets / number of IPs).

[0012] In an optional embodiment, the short time window is used to detect instantaneous traffic anomalies, and the long time window is used to detect persistent traffic anomalies, so as to reduce the false alarm rate by utilizing the periodic characteristics of traffic.

[0013] Preferably, the short time window is set to 1 hour, and the long time window is set to 1 day or 1 week.

[0014] In an optional embodiment, when the characteristic index obeys a normal distribution, a statistical threshold is determined based on the mean and the standard deviation;

[0015] When the characteristic index obeys the lognormal distribution, the statistical threshold is determined according to the logarithmic mean and logarithmic standard deviation.

[0016] In an optional embodiment, semi-supervised training includes:

[0017] The labeled abnormal sample set is randomly divided into three subsets;

[0018] Train three base classifiers using the same or different algorithms;

[0019] Generate pseudo labels for unlabeled samples through a majority voting mechanism and expand the training set until the model converges.

[0020] In an optional embodiment, the pseudo-label generation rule is: if the prediction results of two base classifiers for the same unlabeled sample are consistent, the current sample and the predicted label are added to the training set of the third classifier; and the training data of the three classifiers are dynamically updated during the iteration process.

[0021] In an optional embodiment, the three base classifiers include: integrated decision tree C4.5, naive Bayes classifier and random forest algorithm.

[0022] Second, the present invention provides a dual-window adaptive semi-supervised network threat detection system, comprising:

[0023] A statistical threshold determination unit is used to define multi-dimensional characteristic indicators and determine the statistical thresholds in long and short time windows respectively based on the distribution law of characteristic indicators based on historical network flow data;

[0024] A key object screening unit, used to screen key objects based on statistical thresholds;

[0025] The anomaly detection unit is used to perform anomaly detection on key objects using a set of semi-supervised trained classifiers.

[0026] Furthermore, the statistical threshold determination unit includes:

[0027] Feature index configuration module, used to define multi-dimensional feature indicators;

[0028] The time window configuration module is used to set and manage long and short time window parameters.

[0029] Third, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned dual-window adaptive semi-supervised network threat detection methods.

[0030] The present invention discloses a dual-window adaptive semi-supervised network threat detection method, which has the following advantages over the prior art:

[0031] 1) Multi-dimensional dynamic feature analysis: Simultaneously monitors four-dimensional features of flow direction, protocol type, port, and traffic indicators, breaking through the limitations of traditional single sequence analysis;

[0032] 2) Fusion of statistical modeling and semi-supervised learning: Unsupervised initial screening is achieved through a dual-time window threshold model (short windows capture transient anomalies, long windows suppress false positives). Combined with tri-training semi-supervised ensemble, this approach leverages a small number of labeled samples to mine implicit patterns in unlabeled data, addressing the lack of labeled data.

[0033] 3) Integrated decision-making noise resistance mechanism: Adopting the voting integration framework of C4.5+Naive Bayes+Random Forest to improve the robustness in complex noise environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0035] Figure 1 This is a flow chart of the dual-window adaptive semi-supervised network threat detection method of the present invention;

[0036] Figure 2 This is a diagram of the user behavior prediction process based on semi-supervised learning Tri-training in the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0039] This invention proposes an anomaly analysis method for network user behavior. It models existing threat behavior and third-party security situation data to identify all risky users with similar behavior patterns, enabling prediction of abnormal network behavior. Compared to existing host behavior threat detection methods, this application is more effective in detecting specific abnormal behaviors and similar behaviors, requires less data, and is highly resistant to noise.

[0040] Figure 1 Flow chart of the detection method of the present invention.

[0041] Reference Figure 1 The dual-window adaptive semi-supervised network threat detection method disclosed in the present invention comprises the following steps:

[0042] Define multi-dimensional characteristic indicators, based on historical network flow data, and determine statistical thresholds in long and short time windows according to the distribution patterns of characteristic indicators;

[0043] Screening key targets based on statistical thresholds;

[0044] Anomaly detection is performed on key objects using an ensemble of semi-supervised trained classifiers.

[0045] First, in one embodiment, the target flow record is defined in the following dimensions: flow direction, flow network protocol type, service port number, flow index, etc. The specific meanings are as follows:

[0046] (1) Flow direction: The direction of a flow is determined based on the direction of the data packets in the flow. That is, when the key target is the source address of the connection initiator, the direction of the flow is outflow; otherwise, the direction of the flow is inflow. This dimension contains two attributes: inflow and outflow; the introduction includes the ratio of the number of bytes flowing in to the number of outflows and the ratio of the number of packets flowing in to the number of outflows.

[0047] (2) Stream network protocol type: The stream in which the target communicates with the other end, the protocol running on the network layer, including three attributes: TCP, UDP, and ICMP;

[0048] (3) Service port number: The port number of the target for communication with the service peer. The attribute range is 0 to 65535. The introduction includes the ratio of TCP to UDP bytes and the ratio of TCP to UDP packets.

[0049] (4) Traffic indicators: These indicators mark the size of traffic. Currently, the mainstream indicators include the number of bytes, the number of network packets, the number of IP addresses, etc. Therefore, this dimension mainly includes the following attributes: the number of bytes, the number of network packets, and the number of peer IP addresses.

[0050] After forming the final multi-dimensional attribute combination feature set for key targets, the distribution pattern of each feature item is determined to calculate the statistical threshold;

[0051] This application conducts hypothesis tests on two distribution laws (normal distribution and log-normal distribution) for the statistical values ​​of each traffic feature in different windows.

[0052] For the feature items that obey the normal distribution law, their mean and standard deviation are calculated. For the feature items that obey the lognormal distribution law, their logarithmic mean and logarithmic standard deviation are calculated to obtain the statistical threshold.

[0053] Through the above steps, the statistical threshold range of historical network flow data on various combined features is obtained, and the construction of the multi-dimensional feature statistical threshold model is completed.

[0054] In this embodiment, the selection of the time domain window size is a key issue in applying statistical analysis methods to traffic anomaly detection. If the time domain window is too large, short bursts of abnormal traffic will be masked by longer periods of normal traffic, resulting in an increased false negative rate. Conversely, if the window is too small, the detection sensitivity will be too high, resulting in a high false positive rate and an inability to effectively detect long-term anomalies.

[0055] To this end, this application selects two time windows of different lengths, the short time window is used to detect instantaneous traffic anomalies, and the long time window is used to detect long-term traffic anomalies.

[0056] Specifically, the short-time window can detect explosive attacks in a short period of time, and has a certain real-time detection capability in subsequent anomaly detection; the long-time window is used to detect and discover long-term attack behaviors, and as a supplement to the short-time window detection, it reduces the false alarm rate.

[0057] By cross-combining the long and short windows, setting the long window to one day can take advantage of the daily cycle effect of traffic, while setting the long window to one week can take advantage of the weekly cycle effect of traffic. Preferably, in this embodiment, the long window is set to one day and the short window is set to one hour; alternatively, the long window can be set to one week and the short window to one day.

[0058] Second, after obtaining the target network traffic data, a statistical threshold model is used to make judgments, and users who do not meet the normal range of the threshold statistical model are set as key users.

[0059] Third, semi-supervised machine model construction and user abnormal behavior detection;

[0060] Regarding the construction of semi-supervised machine models, this application uses specific triggered abnormal behaviors as positive samples, utilizes the characteristics of machine learning to automatically learn the detection model, trains the rules of abnormal behaviors, and uses this model to find all risky users with similar behavior patterns among the focus objects.

[0061] In one embodiment, a tri-training machine learning approach is employed. A labeled sample set is randomly divided into three different sample subsets. These three subsets are then trained simultaneously using the same or different algorithms to generate three classifiers. Unlabeled samples are then classified using the three classifiers. In each round of classification, if two classifiers yield the same classification result for a particular unlabeled sample, that sample and its classification result are added to the labeled training set of the third classifier, generating a new training set. Training is repeated until convergence is achieved, concluding the training.

[0062] Figure 2 For the user behavior prediction process based on semi-supervised learning Tri-training, the labeled user abnormal behavior dataset is first repeatedly sampled to obtain three labeled user abnormal behavior training sets, and then a network abnormal behavior classifier is generated from each training set.

[0063] These three abnormal behavior classifiers are then used to generate pseudo-labeled samples in a "minority follows majority" fashion. For example, if two classifiers predict an unlabeled user behavior sample as positive, while a third classifier predicts it as negative, then that user behavior sample is provided as a pseudo-labeled positive sample to the third classifier for learning. Specifically, if two classifiers make the same prediction for the same unlabeled user behavior example, that example is considered to have a higher labeling confidence and, after labeling, is added to the labeled user behavior training set of the third classifier.

[0064] Preferably, the three base classifiers of the present application include: integrated decision tree C4.5, naive Bayes classifier and random forest algorithm. That is, after the final training is completed, the decision tree algorithm C4.5, naive Bayes classifier and random forest algorithm are used as a user behavior classifier integration through a voting mechanism. The voting mechanism here adopts the method of minority obeying majority; for example, if the decision tree algorithm C4.5 judges it as abnormal behavior, the naive Bayes classifier judges it as abnormal behavior, and the random forest algorithm judges it as normal behavior, then the comprehensive judgment is abnormal behavior, and an abnormal judgment report is generated.

[0065] In another embodiment, the present invention provides a dual-window adaptive semi-supervised network threat detection system and a computer-readable storage medium; both apply the dual-window adaptive semi-supervised network threat detection method as described above, so they are not repeated here.

[0066] Preferably, the system comprises:

[0067] A statistical threshold determination unit is used to define multi-dimensional characteristic indicators and determine the statistical thresholds in long and short time windows respectively based on the distribution law of characteristic indicators based on historical network flow data;

[0068] A key object screening unit, used to screen key objects based on statistical thresholds;

[0069] The anomaly detection unit is used to perform anomaly detection on key objects using a set of semi-supervised trained classifiers.

[0070] The statistical threshold determination unit includes:

[0071] Feature index configuration module, used to define multi-dimensional feature indicators;

[0072] The time window configuration module is used to set and manage long and short time window parameters.

[0073] Furthermore, this application uses the "CERT Insider Threat Tool" dataset (Carnegie Mellon Software Engineering Institute, Pittsburgh, Pennsylvania, USA) to conduct experiments to verify the effectiveness and superiority of the method proposed in this application.

[0074] The CERT dataset is an artificial dataset used to validate network threat detection frameworks. It includes employee computer usage logs (login, device, HTTP, file, and email), as well as some organizational information such as employee department and role. Each table contains columns related to user ID, timestamp, and activity. This example uses the sixth version of the CERT dataset for experiments. The experimental results are shown in the following table:

[0075] method Recall Accuracy Naive Bayes 78.6% 72.4% Support Vector Machine 82.0% 76.6% Cluster analysis 85.0% 82.0% Data association analysis 91.3% 88.7% Ensemble Learning 92.0% 90.1% The present invention 96% 92.6%

[0076] As can be seen from the table, compared with other existing user network behavior threat detection methods, this application has a higher recall rate, that is, it has the best detection effect on specific abnormal behaviors and similar behaviors, achieving a threat detection recall rate of more than 95% and a detection accuracy of more than 90%.

[0077] This invention combines a threshold model with machine learning methods to analyze user behavior data and identify abnormal user behavior. This ensures effective operation even with relatively small amounts of data. Furthermore, when using machine learning for detection, a more robust ensemble learning model is employed, rather than a single classification model. This improves the robustness of the system, ensuring strong noise immunity, increasing detection accuracy, and enhancing the system's ability to detect abnormal behavior.

[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0079] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A dual-window adaptive semi-supervised network threat detection method, characterized in that: Define multi-dimensional characteristic indicators, based on historical network flow data, and determine statistical thresholds in long and short time windows according to the distribution patterns of characteristic indicators; Screen key targets based on statistical thresholds; Anomaly detection is performed on key objects using an ensemble of semi-supervised trained classifiers.

2. The network threat detection method according to claim 1, characterized in that: The multi-dimensional characteristic indicators include flow direction attributes, protocol type attributes, service port number attributes, and traffic indicator attributes.

3. The network threat detection method according to claim 1, wherein: The short time window is used to detect instantaneous traffic anomalies, and the long time window is used to detect persistent traffic anomalies.

4. The network threat detection method according to claim 1, wherein: When the characteristic index obeys the normal distribution, the statistical threshold is determined according to the mean and standard deviation; When the characteristic index obeys the lognormal distribution, the statistical threshold is determined according to the logarithmic mean and logarithmic standard deviation.

5. The network threat detection method according to claim 1, wherein: Semi-supervised training includes: The labeled abnormal sample set is randomly divided into three subsets; Train at least three base classifiers; Generate pseudo labels for unlabeled samples based on multiple base classifiers through a majority voting mechanism, and expand the training set until the model converges.

6. The network threat detection method according to claim 5, characterized in that: The pseudo-label generation rule is: if the prediction results of two base classifiers for the same unlabeled sample are consistent, then the current sample and the predicted label are added to the training set of the third classifier; And the training data of the three classifiers are dynamically updated during the iteration process.

7. The network threat detection method according to claim 5, characterized in that: The three base classifiers include: integrated decision tree C4.5, naive Bayes classifier and random forest algorithm.

8. A dual-window adaptive semi-supervised network threat detection system, characterized in that: include: A statistical threshold determination unit is used to define multi-dimensional characteristic indicators and determine the statistical thresholds in long and short time windows respectively based on the distribution law of characteristic indicators based on historical network flow data; A key object screening unit, used to screen key objects based on statistical thresholds; The anomaly detection unit is used to perform anomaly detection on key objects using a set of semi-supervised trained classifiers.

9. The network threat detection system according to claim 8, characterized in that: The statistical threshold determination unit includes: Feature index configuration module, used to define multi-dimensional feature indicators; The time window configuration module is used to set and manage long and short time window parameters.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the dual-window adaptive semi-supervised network threat detection method according to any one of claims 1 to 7 is implemented.