A Dynamic Data Stream Clustering Method for Network Intrusion Behavior Detection
By integrating sparse constraints and dictionary initialization strategies in network intrusion detection, using sparse representation technology and sliding windows, the clustering problem of dynamic network traffic is solved, and efficient network intrusion behavior detection and identification is achieved.
Patent Information
- Application Number
- CN202310966366.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-02
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-02
AI Technical Summary
When facing dynamic network traffic, traditional network intrusion detection methods cannot effectively identify new or changed behavior patterns, and the computing resources and time consume too much, making it difficult to achieve real-time or near-real-time response.
Sparse constraints are used to integrate them into the representation of data objects, and through sparse matrix and dictionary initialization strategies, sparse representation technology and sliding windows are used to adaptively transfer knowledge, and cluster dynamic data flows with spectral clustering algorithms.
It improves the clustering performance of dynamic data flow, reduces the consumption of computing resources, and realizes efficient detection and identification of network intrusion behavior.
Smart Images

Figure CN117150322B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a dynamic data stream clustering method for network intrusion behavior detection. Background Art
[0002] In the field of network security, intrusion detection systems (IDS) are a key component. These systems typically use different methods to detect and prevent network attacks, including malware, botnets, denial-of-service attacks, and so on. IDS can work based on signatures or based on behavior. Signature-based IDS relies on predefined intrusion signatures, such as known malware signatures. However, this method may fail when faced with unknown threats. Behavior-based IDS, on the other hand, detects intrusions by learning the "normal" behavior of the network and looking for abnormal behavior.
[0003] Data clustering is a common machine learning technique used to group similar data points together. In the field of network security, data clustering can be used to identify abnormal or suspicious behavior patterns. For example, if a specific IP address sends a large number of requests in a short period of time, this may be identified as abnormal behavior by a clustering algorithm.
[0004] However, traditional clustering methods cannot effectively handle dynamic network traffic data. For example, if the network traffic pattern changes over time, static clustering methods may not be able to accurately identify new or changed behavior patterns. In addition, static clustering methods may require a large amount of computing resources and time, which may be impractical for network security applications that require real-time or near-real-time response.
[0005] In recent years, clustering techniques for analyzing network intrusion data have been widely studied. Several traditional clustering methods have been proposed for dynamic data streams, such as hierarchical-based, density-based, partition-based, and model-based algorithms. They usually adopt a local similarity strategy and use the Euclidean distance between network intrusion behavior data objects to preserve local neighborhood information. As new clusters emerge or other clusters disappear, the data stream will also change dynamically, which requires the clustering algorithm to be able to incrementally learn from large-scale dynamic data streams. Therefore, a new learning strategy is needed to adaptively transfer the previously learned knowledge to the current data stream over time. Summary of the Invention
[0006] The object of the present invention is to provide a dynamic data stream clustering method for network intrusion behavior detection, aiming to solve the problem that traditional methods need to incrementally learn from large-scale dynamic data streams when the data stream undergoes dynamic changes.
[0007] To achieve the above object, the present invention provides a dynamic data stream clustering method for network intrusion behavior detection, comprising the following steps:
[0008] S1 Prepare a data set, a large-scale dynamic data stream related to network intrusion behavior;
[0009] S2 Integrate sparse constraints into the representation learning of network intrusion behavior data objects to obtain a sparse matrix;
[0010] S3 Propose a dictionary initialization strategy;
[0011] S4 Based on the sparse matrix, adopt the dictionary initialization strategy to effectively transfer knowledge to the current sliding window;
[0012] S5 Cluster to obtain the detection result of network intrusion behavior.
[0013] Wherein, the sparse matrix is used as an affinity matrix for spectral clustering.
[0014] Wherein, the dictionary initialization strategy is used to retain the previously learned knowledge and provide cluster data for the current sliding window.
[0015] Wherein, the dictionary initialization strategy and sparse representation are iteratively updated by solving a convex optimization problem.
[0016] Wherein, the iterative update needs to introduce auxiliary variables.
[0017] The present invention proposes a dynamic data stream clustering method for network intrusion behavior detection. First, integrate sparse constraints into the representation learning of data objects to obtain a sparse matrix; then, propose a dictionary initialization strategy to effectively transfer knowledge to the current sliding window based on the sparse matrix; wherein, the dictionary initialization strategy and sparse representation are iteratively updated by solving a convex optimization problem, and auxiliary variables need to be introduced during the update process.
[0018] The input of the present invention is a large-scale dynamic data stream related to network intrusion behavior, and these data streams contain various network activity records, such as network traffic data, network connection logs, user behavior data, etc.
[0019] The output of the present invention is the detection result of network intrusion behavior. Through dynamic data stream clustering, potential intrusion patterns or abnormal behaviors and their positions in the data stream can be discovered.
[0020] The present invention utilizes sparse representation technology to exploit the intrinsic characteristics of data objects, automatically selects the number of adjacent data objects, ensures that highly correlated data objects are represented together, greatly enriches the relationships between data objects, and improves the clustering performance of data streams. In addition, through a reasonable dictionary and sliding window size, the computational cost can be effectively controlled. Compared with traditional methods, the method of the present invention does not require incremental learning from large-scale network intrusion dynamic data streams when dealing with dynamic changes in network intrusion data streams, thus solving the problems of computational resource consumption and detection efficiency. Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a flowchart of a dynamic data stream clustering method for network intrusion behavior detection provided by the present invention. Detailed Embodiments
[0023] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0024] Please refer to Figure 1 , the present invention provides a dynamic data stream clustering method for network intrusion behavior detection, including the following steps:
[0025] S1 Prepare a data set, a large-scale dynamic data stream related to network intrusion behavior;
[0026] Collect a data set related to network intrusion behavior detection. Use a network traffic monitoring tool to capture network traffic packets and save them in the pcap file format. Collect system log data generated by operating systems, applications, firewalls, etc. Utilize the log data generated by intrusion detection systems. According to the task requirements of network intrusion behavior detection, select features related to intrusion behavior for extraction. These features include network protocols, source IP addresses, destination IP addresses, port numbers, and transport protocols. Standardize the selected features to obtain DS.
[0027] S2 Integrate sparse constraints into the representation learning of network intrusion behavior data objects to obtain a sparse matrix;
[0028] Specifically, we consider a network intrusion data stream DS as a sequence of several network intrusion behavior data objects, and each sequence represents a sliding window, that is Let be the sequence of data objects in the t-th window of a network intrusion data stream. Each column of X t represents a network intrusion data object arriving in time t. In particular, let be a representative example of the data objects in the sliding window. TSRC makes full use of the sparse representation method to measure their relationships. For example, given a set of network intrusion data objects in a sliding window, we represent each data object as a linear combination of other objects, where the matrix coefficients should be sparse. The coefficients of the sparse representation of a single data object are adopted to measure the relationships between data objects. Therefore, designing an affinity matrix is a key step in evaluating the relationships between data objects.
[0029] Considering a given network intrusion data stream data object x i (1 ≤ i ≤ n), TSRC uses the following l0-norm minimization to achieve sparse representation:
[0030]
[0031] where ‖·‖ l , representing the noise term ε. The columns of matrix Z are composed of sparse coefficients, that is In addition, Z * is used to define an affinity matrix, that is Z * = |Z| + |Z T |, where the element value Z ij in the affinity matrix z * represents the similarity between the data objects x i and x j represented by the network intrusion behavior. The corresponding values of the sparse representation between each pair of data objects represent their similarity, and the number of non-zero sparse representation coefficients represents the potential relationship between the corresponding data objects. If any two data objects x i and x j are close in the intrinsic geometric structure of the data distribution, then the representations of these two data objects, that is z i and z j relative to the same basis X, are close to each other. For a larger similarity between X i and X j , the distance between z i and z j should be smaller. Therefore, it is reasonable to use sparse representation to measure the relationships between data objects in the network intrusion behavior detection data stream.
[0032] Clustering network intrusion data streams requires a process that can continuously cluster objects within a certain time limit. However, determining the sparsest representation of each data object representing a network intrusion individually in (4) leads to high computational costs. To accelerate the optimization, we apply the batch mode and obtain a sparse solution by solving the sparse representation optimization problem. Each element \(z_{ij}\) of matrix \(Z\) ij represents the similarity between data objects \(i\) and \(j\). To maintain the interpretability of the sparse representation results, the diagonal elements of \(Z\) must be zero. This indicates that there is no relationship between any data object and itself. Additionally, this also avoids the trivial solution where \(Z\) is the identity matrix from an optimization perspective. Specifically, we consider the following convex optimization problem to seek the sparse representation \(Z\).
[0033]
[0034] where \(\alpha\) is a scalar constant and \(D\) is a specific dictionary that linearly spans the network intrusion data space.
[0035] Given an initial dictionary \(D\), \(X\) can be represented as a linear combination:
[0036] \(X = DZ+E\), (6)
[0037] where \(D\) is closely related to \(X\).
[0038] Assuming that the result \(Z\) of the linear combination is sparse, the preliminary objective function is defined as:
[0039]
[0040] where \(\beta>0\) is a regularization parameter and \(D\) is the dictionary representing network intrusion behaviors to be learned.
[0041] The coefficients and dictionary of the sparse representation are iteratively updated by solving (7). We first transform this problem into the following equivalent problem by introducing an auxiliary variable \(J\):
[0042]
[0043] The augmented Lagrangian function of (8) is:
[0044]
[0045] where \(Y\) is the Lagrange multiplier and \(\mu>0\) is a penalty parameter. The above optimization problem can be formulated as follows:
[0046]
[0047]
[0048]
[0049] Equation (10) can be effectively solved using an inexact augmented Lagrangian multiplier (ALM) framework. The variables J, Z, and D can be updated alternately while the other two variables are fixed. Variable J has a closed-form solution, and the update scheme for J k+1 is:
[0050]
[0051] J k+1 = normalize (0,1] (J k+1 ).
[0052] J k+1 's columns are projected onto the unit circle by the second part of (12). This is because, under the assumption that ‖·‖2 < 1, the objective function value of the first equation in (12) with respect to variable Z is guaranteed not to increase. Given fixed J k+1 , Z k+1 is updated by the following scheme:
[0053]
[0054] Z k+1 = Z k+1 - diag(Z k+1 ).
[0055] The hard thresholding operator H is defined as follows:
[0056]
[0057] Using the operator H, a closed-form solution for the first part of (12) can be obtained:
[0058]
[0059] For fixed Z k+1 , at the (k + 1)-th iteration, D k+1 is updated by approximately minimizing the following surrogate function:
[0060]
[0061] The gradient of φ(D k ) can be computed as
[0062]
[0063] The update scheme for D k+1 is
[0064]
[0065] where δ is the gradient step and P represents the Euclidean projector onto the l2 norm. δ = A[j,j]. In practice, we usually compute the j-th column of D k+1 , where . For a fixed J k+1 and Z k+1 , the Lagrange multiplier Y k+1 is updated using the step size μ as:
[0066] Y k+1 = Y k + μ k (J k+1 - Z k+1 ). (18)
[0067] In Algorithm 1, convergence is achieved when ‖Z k - J k ‖ ∞ < ε, where ‖·‖ ∞ represents the infinity norm of each vector in the matrix. In practice, we usually set ε = 10 -2 . Spectral clustering techniques can be employed to obtain the clustering results of the current sliding window. For example, Normalized Cuts (NCuts), where the matrix Z k represents the final similarity of the network intrusion data objects in the current sliding window. The complete procedure for solving (7) is outlined in Algorithm 1.
[0068] S3 Proposed dictionary initialization strategy;
[0069] Specifically, selecting a suitable dictionary in (11) leads to an effective and fast algorithm for evaluating the sparse representation of network intrusions. An intuitive approach is to use the network intrusion data objects X as a dictionary; this is typically adopted in traditional clustering algorithms based on sparse representation. However, various degrees of noise are ubiquitous in the data stream, which may lead to unsatisfactory results. In addition, network intrusion data streams have special characteristics, such as being potentially infinite and dynamically changing, which makes them different from traditional static datasets. The data objects in the current window must be processed and discarded before the arrival of the next sliding window of the network intrusion data stream. It is necessary to develop an online dictionary learning strategy for sparse representation to adaptively transfer the previously learned knowledge to the current data stream over time. Therefore, we seek a more suitable dictionary rather than using the network intrusion data objects themselves as the dictionary. This reveals the true relationships between the data objects in the current sliding window.
[0070] To design a compact and discriminative dictionary, a crucial step is to retain valuable network intrusion data objects within a certain time and replace them with some new valuable network intrusion data objects later. Suppose we have representative network intrusion data objects These objects were learned from the network intrusion data stream at time t before. When t = 1, the problem of data stream clustering is considered a traditional clustering problem, i.e., D = X1. Then, X1 is considered as X of the current sliding window s . When X t arrives at time t > 1, we adopt and X = [X s , X t , which represent a set of data objects represented by network intrusion and its dictionary respectively, for (5). The sparse representation result Z and the corresponding dictionary D can be obtained using Algorithm 1
[0071] We further elaborate on the significance of the above sparse representation Z and the corresponding dictionary D for two key issues in network intrusion data stream clustering. First, the learned affinity matrix is based on Z, i.e., Z * = |Z| + |Z T |. Given the number of clusters c, TSRC considers Z * , as the affinity of the spectral clustering algorithm, such as NCuts. Thus, we obtain the final clustering result as X, which contains the clustering of X t in the current window. In addition, Z and D are crucial for accurately constructing the new X s of the current sliding window. The c sub-matrices are extracted from the sparse representation Z, i.e., Z1, Z2,..., Z c . Considering the data object x i from cluster j, we can easily select its sparse representation Z j from Z. Each coefficient matrix is a block matrix related to the j-th cluster. The dictionaries D for all c clusters can be regarded as D = [D1, D2,..., D c , where each D j is a sub-dictionary containing n j elements. Ideally, the non-zero entries in Z j will all be associated with the columns of D j . In other words, x i should be represented as a linear combination of the elements from D j , where A vector representing the sparse representation coefficients associated with the j-th cluster. However, missing and insufficient modeling accuracy may lead to small non-zero entries associated with some data objects in other clusters. By using only the coefficients associated with the corresponding cluster j, we can approximate the network intrusion data object X j Approximate as The definition of the approximate residual (AR) of each element of D, i.e., d k is given as follows:
[0072]
[0073] AR(d k ) represents the error of all dictionary elements of the j-th cluster when d k is removed, and is used to measure the importance of d k in the sparse representation. The higher the value of AR(d k ), the more important d k is. Then, the representative data object of each cluster C j is determined as the object with the largest m j residual. Algorithm 2 summarizes the complete network intrusion data stream clustering procedure.
[0074]
[0075]
[0076] In practice, there are often several data objects representing network intrusions that belong to more than one cluster. In traditional clustering problems, these can be considered outliers and are usually removed before clustering to prevent poor clustering performance due to insufficient representative data objects in the dictionary of the sparse representation. However, most such network intrusion objects may be valid in the data stream. TSRC overcomes the basic limitation of insufficient representative data objects in traditional sparse representation algorithms because Xs may contain representative data objects from all clusters.
[0077] S4 Based on the sparse matrix, adopt the dictionary initialization strategy to effectively transfer knowledge to the current sliding window.
[0078] Specifically, an online dictionary learning strategy is introduced to transfer the previously learned knowledge to the current sliding window. In particular, the exact number of clusters of the current sliding window can be determined over time, which greatly enriches the relationships between network intrusion data objects and improves the clustering performance of the network intrusion data stream.
[0079] S5 Clustering to obtain the detection result of network intrusion behavior
[0080] The clustering algorithm K-means is selected. By calculating the similarity between data objects, they are divided into different clusters or groups, making the data objects within the same cluster more similar and those between different clusters less similar. The clustering results are evaluated to measure the performance and accuracy of clustering. The evaluation metrics include the silhouette coefficient, mutual information, and adjusted Rand index. According to the clustering results, the network intrusion behavior detection results are classified. Each clustering cluster can represent a specific network intrusion behavior pattern. By analyzing the characteristics and behavior patterns of the data objects in each cluster, different network intrusion behaviors can be identified. Based on the detection results of network intrusion behaviors, further interpretation and analysis are carried out. According to the characteristics and behavior patterns of the clustering clusters, the potential intruder behaviors, vulnerabilities, or attack types can be inferred, and corresponding security measures can be taken.
[0081] The above-disclosed is only a preferred embodiment of a dynamic data stream clustering method for network intrusion behavior detection of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A dynamic data stream clustering method for network intrusion behavior detection, characterized in that Including the following steps: S1 Prepare a data set, a large-scale dynamic data stream related to network intrusion behavior; S2 Integrate sparse constraints into the representation learning of network intrusion behavior data objects to obtain a sparse matrix; S3 Propose a dictionary initialization strategy; S4 Based on the sparse matrix, use the dictionary initialization strategy to effectively transfer knowledge to the current sliding window; S5 Cluster to obtain the detection result of network intrusion behavior; The dictionary initialization strategy is used to retain the previously learned knowledge and provide cluster data for the current sliding window.
2. A dynamic data stream clustering method for network intrusion behavior detection according to claim 1, characterized in that The sparse matrix is used as the affinity matrix for spectral clustering.
3. A dynamic data stream clustering method for network intrusion behavior detection according to claim 1, characterized in that The dictionary initialization strategy and sparse representation are iteratively updated by solving a convex optimization problem.
4. A dynamic data stream clustering method for network intrusion behavior detection according to claim 3, characterized in that Auxiliary variables need to be introduced for the iterative update.
Citation Information
Patent Citations
Intrusion detection method based on improved dictionary learning
CN106991435A
Hyperspectral ground object automatic classification method and system based on sparse subspace clustering
CN112364730A