Intrusion detection method and system based on multi-scale feature extraction and variational clustering
By adopting multi-scale adaptive gating feature extraction and multi-dimensional index-driven clustering variational clustering method in the intrusion detection system, the problem of insufficient recognition ability of unknown attacks in the prior art is solved, and efficient known attack classification and unknown attack recognition are achieved.
Patent Information
- Application Number
- CN202510432941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively identify abnormal traffic when it is known that the classification error of attacks is large and the ability to identify unknown attacks is insufficient, especially when facing dynamic network traffic and noise interference.
A variational cluster intrusion detection method based on multi-scale adaptive gating feature extraction and multi-dimensional index-driven clustering is adopted, combined with deep learning and clustering algorithms, multi-scale features are extracted through multi-head attention mechanism and gating mechanism, and unknown attacks are identified using variational inference and density clustering.
It significantly improves the ability to identify unknown attacks, reduces the classification error of known attacks, and enhances the adaptability and security of intrusion detection systems.
Smart Images

Figure CN120151075A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security detection, and particularly relates to a variational clustering intrusion detection method and system based on multi-scale feature extraction and variational clustering. Background Art
[0002] In the field of network security, intrusion detection is of great importance. However, there are many problems in the existing technologies. For example, rule-based methods rely on static rule libraries, making it difficult to cope with constantly changing attack means and having insufficient detection capabilities for unknown attacks; machine learning-based methods are troubled by data imbalance and feature selection problems, resulting in large classification errors for known attacks. At the same time, the existing technologies perform poorly in identifying unknown attacks. Most are based on supervised learning and rely on a large amount of labeled data. Facing unknown attacks not covered by the training data, they lack generalization capabilities. In addition, characteristics such as multi-scale, high-dimensionality, and noise interference of network traffic data further increase the detection difficulty. Summary of the Invention
[0003] To solve the problems of large classification errors for known attacks and insufficient identification capabilities for unknown attacks in the existing technologies, the present invention proposes a variational clustering intrusion detection method and system based on multi-scale adaptive gating feature extraction and multi-dimensional index-driven clustering. This method combines the feature extraction ability of deep learning and the unsupervised learning ability of clustering algorithms, and can effectively identify unknown attacks on the basis of minimizing the classification errors of known attacks.
[0004] The technical solutions provided by the present invention are as follows:
[0005] On the one hand, a variational clustering intrusion detection method based on multi-scale feature extraction, comprising:
[0006] Step 1: Collect network traffic data, and extract features such as source IP, destination IP, port number, protocol type, traffic size, and connection duration to form a feature vector; perform standardization processing on the feature vector;
[0007] Step 2: Use the multi-head attention mechanism to extract features from the feature vector processed in Step 1, and perform a linear combination of the weighted outputs of all attention heads to generate a multi-scale adaptive gating feature corresponding to the final attention result;
[0008] Step 3: Use the similarity with the normal traffic feature vector to preliminarily determine whether the multi-scale adaptive gating feature extracted in Step 2 is abnormal. If it is preliminarily determined to be abnormal, go to Step 4; otherwise, directly go to Step 5;
[0009] Step 4: According to the preset clustering parameters, update the minimum number of points minPts and the neighborhood radius eps in the clustering parameters based on the silhouette coefficient, CH index, and k-distance graph;
[0010] Step 5: Using the latest clustering parameters, perform density algorithm clustering on the multi-scale adaptive gating features corresponding to the network traffic, and analyze the network traffic category based on the clustering results. If the feature vector data points corresponding to the network traffic are determined to be noise points or in an abnormal cluster, they are identified as suspicious network traffic;
[0011] Step 6: For the suspicious network traffic, make a final classification decision using the preset rules. If it is determined to be abnormal traffic, trigger an alarm and notify relevant personnel for handling; if it is determined to be normal traffic, allow the traffic to pass normally.
[0012] Make the final classification decision from three aspects: traffic time factor, traffic port factor, and clustering combination factor;
[0013] This solution breaks through the limitations of a single feature extraction method and constructs a detection framework that can adapt to dynamic traffic feature changes. Through the multi-head attention mechanism, capture traffic features at different time scales, and use the gating mechanism to filter out effective information to enhance the feature expression ability; combined with the optimized distance formula, multi-dimensional indicators drive clustering, making the clustering process more adaptive and improving the recognition accuracy of abnormal traffic; variational inference mechanism: introduce variational methods in the clustering process, enabling the system to more effectively model unseen abnormal traffic types and achieve the detection of unknown attacks.
[0014] Through the above innovations, the present invention not only realizes the efficient classification of known attacks, but also significantly improves the recognition ability of unknown attacks, solves the problem of insufficient response to new unknown attacks in the prior art, and enhances the adaptability and security of the intrusion detection system.
[0015] Further, the specific steps of the multi-scale adaptive gating feature extraction are as follows:
[0016] Step 2.1: Calculate the multi-scale attention weights:
[0017]
[0018] where, e ij represents the multi-scale attention weight between position i and position j in the feature vector corresponding to the network traffic. The larger the weight value, the stronger the association between the features of the two positions at multiple scales, and the greater the impact on subsequent feature extraction and classification decisions; M is the number of multi-scales; X i and X j are the feature vectors at position i and position j in the feature vector corresponding to the network traffic respectively, is a 128×128-dimensional weight matrix of the query matrix and the key matrix at the m-th scale; d (m) is The dimension for the scaling operation, and the superscript T represents the transpose of the matrix;
[0019] Step 2.2: Based on the multi-scale attention weights, calculate the attention output at each position on the feature vector according to the following formula;
[0020]
[0021] where h i represents the attention output at the i-th position of the feature vector, a ij is the original attention weight regarding positions i and j calculated based on the multi-scale attention weights, obtained by performing a softmax operation on the result e ij from Step 2.1, n represents the number of elements in the current input, which needs to be calculated according to the actual sequence length, X i and X j are the feature vectors at positions i and j in the input sequence, W 1 is a 4×4 dimensional context weight matrix for capturing the global relationships in the input sequence; α is a learnable parameter for adjusting the relative importance of the original attention weights and the context similarity; W 2 is a 128×128 dimensional weight matrix for linear transformation; β and δ are learnable parameters;
[0022] Step 2.3: Introduce a dynamic gating weight g j to the attention output obtained in Step 2.2, and the specific calculation formula is as follows;
[0023]
[0024] where, is the attention output result of h i at the Q-th time step and the z-th attention head obtained in Step 2.2, σ is the activation function sigmoid that converts the calculation result into the form of a gating weight, W 3 is a 4×1 dimensional gating weight matrix for performing a linear transformation on , H represents the number of attention heads, and different feature dimensions at the same time scale are captured by parallel calculation of multiple heads, b 1 is the bias term, which helps the model better fit the data;
[0025] Step 2.4: Perform a linear projection on the weighted output of the attention head after introducing the dynamic gating weight, and the formula is:
[0026]
[0027] where, is a matrix of 4×4 dimensions for output transformation of the features of the z-th attention head. is a matrix of 4×4 dimensions for weighted transformation of the features of the z-th attention head, W 4 is a global projection matrix of 128×128 dimensions that needs to be learned through backpropagation.
[0028] Furthermore, a position-aware function is introduced to enhance the features after linear projection, as follows:
[0029] The position-aware function is introduced to enhance the attention result after linear projection in combination with position encoding according to the following formula:
[0030]
[0031] where O is the output of the linear projection in step 2.4, represents the position-aware function,
[0032] W 5 is a learnable weight matrix of 128×128 dimensions, b 2 is a learnable bias vector, and λ is a hyperparameter used to adjust the influence of the position-aware part.
[0033] Furthermore, the output O enhanced by position awareness enhanced is used as the input Δ of the interactive position-enhanced feedforward network, and the features output by the interactive position-enhanced feedforward network are used as the final features;
[0034] The formula of the interactive position-enhanced feedforward network is as follows:
[0035]
[0036] where IPA represents "Interactive Position Enhancement", FFN is a standard feedforward network, represents the interaction term, represents the position enhancement term; W 6 is a learnable weight matrix of 128×512 dimensions for linear transformation of the input Δ; W 7 is a learnable weight matrix of 512×128 dimensions for linear transformation of the result of the first half; b 3 is a learnable bias vector, θ, is a learnable parameter used to adjust the importance of the interaction term and the position enhancement term; K is the number of other vectors that interact with the input vector, α ik is the weight related to the interaction between Δ and, Δ kis the k-th interaction feature representation of the input feature Δ, obtained by linear projection; Δ T is the transpose of the input feature Δ, is the learnable weight matrix of the k-th interaction channel of 128×128; m is the number of position encodings, θ il is the learnable weight, calculated according to the relationship between the input vector Δ and the position encoding P l The calculation method is P is the position encoding vector of the elements in the input sequence, using sine-cosine position encoding, is to project the input feature Δ and the position encoding P l into a 128×128-dimensional matrix in the same semantic space.
[0037] The high-order features extracted from the output of the interactive position enhancement feed-forward network enhance the clustering algorithm's ability to distinguish abnormal traffic, especially the detection of unknown attacks, and the output features improve the accuracy of cosine similarity calculation.
[0038] Furthermore, using the similarity with the normal traffic feature vector to make a preliminary judgment on whether the feature vector extracted by multi-scale adaptive gating is abnormal or not refers to calculating the similarity between the feature vector of the network traffic to be detected and the feature vector of the normal traffic in the database using cosine similarity.
[0039] Furthermore, in the clustering process, the following distance formula is used to calculate the distance between the feature vector ν 1 , λ 1 :
[0040]
[0041] where w i is the learnable weight of the i-th feature dimension, which can automatically adjust the importance of different dimensions in distance calculation according to the internal law of the data; μ is a learnable parameter used to adjust the overall importance of the new item; G is a preset integer representing the number of new distance-related features, determined according to the specific task and data characteristics; Z is the total number of dimensions of the feature vector judged to be preliminarily abnormal; F j (|v i -λ i |) is the j-th new feature, F j (|v i -λ i |) = sigmoid(W (1) ·|v i -λ i |+b (1) ), sigmoid is the variational activation function, which includes the probability modeling of the variational distribution for the feature difference, W(1) and b (1) Through variational posterior learning, it is used to enhance the sensitivity to abnormal data; ω j is the learnable weight related to the new feature j, and the calculation method is
[0042] where is the learnable weight matrix, φ i = |v i - λ i | is the absolute difference between the feature vector v 1 and λ 1 on the i-th feature dimension.
[0043] Furthermore, the process of determining the minimum number of points minPts and the neighborhood radius eps in the clustering parameters is as follows:
[0044] (1) Determine the minimum number of points minPts;
[0045] Step B1: Traverse the candidate parameters: For each value of minPts ∈ {2,..., 10}, perform initial clustering, and then calculate the silhouette coefficient s complex (z);
[0046]
[0047] where a(z) is the average distance from the sample z to other samples in the same cluster, reflecting the tightness of the sample i within its cluster; b(z) is the average distance from the sample z to all samples in the nearest cluster, used to measure the separation degree of the sample z from other clusters, and is the variational adjustment factor;
[0048] Step B2: Calculate the Calinski-Harabasz index:
[0049]
[0050] where B(k) is the between-cluster scatter matrix, obtained by calculating the covariance difference between the centroids of each cluster and the global centroid through the clustering result, used to measure the dispersion degree between different clusters; W(k) is the within-cluster scatter matrix, obtained by calculating the covariance difference between the samples within each cluster and the cluster centroid through the clustering result, reflecting the tightness of the samples within each cluster; tr(B(k) 3 ) represents taking the trace after cubing the between-cluster scatter matrix, highlighting the impact of the case with a large difference in between-cluster distances on the clustering quality; det(W(k)) is the determinant of the within-cluster scatter matrix, reflecting the overall characteristics of the sample distribution within the cluster;
[0051] Step B3: Traverse the changes in the silhouette coefficient and CH index for different minpts values, and select the value of minpts corresponding to the maximum value of s complex +CH(k) as the optimal value of minpts;
[0052] (2) Determine the neighborhood radius eps;
[0053] Step C1: Traverse the candidate parameters: For each eps ∈ {0.1, 0.2, …, 2}, select a value, perform initial clustering, then calculate the k-nearest neighbor distances of each sample, sort the k-nearest neighbor distances in reverse order, and find the inflection point of the sorted k-nearest neighbor distances;
[0054] Step C2: Calculate the silhouette coefficient s complex (z) of the eigenvector corresponding to the network traffic after initial clustering;
[0055] Step C3: Traverse the changes in the silhouette coefficient and distance inflection point for different eps values, select the inflection point of the k-nearest neighbor distance and the corresponding maximum silhouette coefficient, and obtain
[0056] Second aspect, a detection system based on the above multi-scale feature extraction and variational clustering intrusion detection method,
[0057] including:
[0058] Network traffic data acquisition unit: Extract source IP, destination IP, port number, protocol type, traffic size, and connection duration features, and form an eigenvector;
[0059] Standardization processing unit: Perform standardization processing on the eigenvector;
[0060] Feature acquisition unit: Use the multi-head attention mechanism to extract features from the eigenvector processed by the standardization processing unit, and perform a linear combination of the weighted outputs of all attention heads to generate a multi-scale adaptive gating feature corresponding to the final attention result;
[0061] Abnormal preliminary judgment unit: Use the similarity with the normal traffic eigenvector to preliminarily judge whether the eigenvector of the feature acquisition unit is abnormal;
[0062] Clustering parameter update unit: According to the preset clustering parameters, use the eigenvector corresponding to the network traffic preliminarily judged as abnormal by the abnormal preliminary judgment unit to update the minimum number of points minPts and neighborhood radius eps in the clustering parameters based on the silhouette coefficient, CH index, and k-distance graph;
[0063] Secondary judgment unit: Using the current latest clustering parameters, perform density algorithm clustering on the multi-scale adaptive gating features corresponding to network traffic, and analyze the network traffic category based on the clustering results. If the feature vector data points corresponding to the network traffic are determined to be noise points or in abnormal clusters, it is identified as suspicious network traffic;
[0064] Decision-making unit: For suspicious network traffic, perform final classification decisions using pre-set rules. If it is determined to be abnormal traffic, trigger an alarm and notify relevant personnel for handling; if it is determined to be normal traffic, allow the traffic to pass normally.
[0065] Third aspect, a computer device, including
[0066] One or more processors;
[0067] A memory storing one or more computer programs;
[0068] Wherein, the processor calls the computer program to implement:
[0069] The steps of the above-mentioned intrusion detection based on multi-scale feature extraction and variational clustering.
[0070] Fourth aspect, a computer-readable storage medium storing a computer program, and the computer program is called by a processor to implement:
[0071] The steps of the method for multi-scale feature extraction and variational clustering-based intrusion detection.
[0072] Beneficial effects
[0073] The technical solution provided by the present invention has the following advantages compared with the prior art:
[0074] The present invention proposes a variational clustering intrusion detection method based on multi-scale adaptive gating feature extraction and multi-dimensional index-driven clustering, which detects network traffic in two stages, respectively minimizing the empirical risk of known attacks and the open-set risk of unknown attacks, that is, dividing the known or unknown intrusion detection problem into a two-stage minimization problem. The first stage minimizes the empirical risk, and the second stage minimizes the open-set risk. This hierarchical nature of the problem formulation is beneficial to realizing the identification of unknown attacks while maintaining the classification accuracy of known attacks.
[0075] In the first stage, by introducing a multi-scale adaptive gating mechanism, the network traffic feature extraction process becomes more intelligent and can adaptively adjust the information weights at different time scales. This module can automatically adjust the extraction scale and weight, accurately identify known attack types, and significantly reduce classification errors.
[0076] In the second stage, a clustering algorithm driven by multi-dimensional indicators is combined to improve the recognition and classification capabilities of unknown attacks. Meanwhile, a variational clustering method is proposed, which combines density clustering and variational inference, enabling the clustering method to automatically adapt to different types of abnormal traffic and improving the detection capabilities of unknown attacks. The distance formula is optimized by combining learnable weights with newly added feature terms, enhancing the adaptability of the clustering algorithm in network traffic data and improving the recognition effect of abnormal traffic. Feature relationships and clustering strategies are fused. Traffic is initially screened through cosine similarity and then combined with variational clustering to reduce the false alarm rate and improve the detection accuracy, significantly enhancing the generalization ability and robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a schematic diagram of the intrusion detection process in the technical solution of the present invention;
[0078] Figure 2 It is a schematic diagram of the clustering effect on the test set using the technical solution of the present invention;
[0079] Figure 3 It is a schematic diagram of the clustering effect on the training set using the technical solution of the present invention;
[0080] Figure 4 It is a schematic diagram of the accuracy effect of correctly discriminating network traffic (attack traffic and benign traffic) on the test set and the training set using the technical solution of the present invention, and also includes a comparison effect schematic diagram with a traditional model; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The following further describes the embodiments of the present invention with reference to the drawings.
[0082] As Figure 1 shown, a multi-scale feature extraction and variational clustering-based intrusion detection method includes:
[0083] Step 1: Collect network traffic data, extract source IP, destination IP, port number, protocol type, traffic size, and connection duration features, and form a feature vector; perform standardization processing on the feature vector;
[0084] Perform standardization processing on these feature vectors to make them have a unified scale. For example, map the values to the interval [0, 1]. The formula is as follows:
[0085]
[0086] Among them, X represents the original feature value, which is the specific observation value on each feature dimension in the dataset. x min represents the minimum value of this feature in the dataset, and x max refers to the maximum value of this feature in the dataset, and x normis the standardized eigenvalue, that is, the result obtained through formula calculation. Such processing helps to improve the efficiency and stability of model training, and avoid certain features from having too much impact on the model due to too large or too small numerical ranges.
[0087] Step 2: Use the multi-head attention mechanism to extract features from the feature vectors processed in Step 1, and perform a linear combination of the weighted outputs of all attention heads to generate the multi-scale adaptive gating features corresponding to the final attention result;
[0088] The specific steps for extracting the multi-scale adaptive gating features are as follows:
[0089] Step 2.1: Calculate the multi-scale attention weights:
[0090]
[0091] where, e ij represents the multi-scale attention weight between position i and position j in the feature vector corresponding to the network traffic. The larger the weight value, the stronger the correlation between the features at the two positions at multiple scales, and the greater the impact on subsequent feature extraction and classification decisions; M is the number of multi-scales. In this example, M is set to 3, corresponding to long-term trends, short-term fluctuations, and emergencies; X i and X j are the feature vectors at position i and position j in the feature vector corresponding to the network traffic respectively, is a 128×128 dimensional weight matrix of the query matrix and the key matrix at the m-th scale; d (m) is 's dimension, used for scaling operation, and the superscript T represents the transpose of the matrix;
[0092] Step 2.2: Based on the multi-scale attention weights, calculate the attention output at each position on the feature vector according to the following formula;
[0093]
[0094] where, h i represents the attention output at the i-th position of the feature vector, a ij is the original attention weight regarding position i and position j calculated based on the multi-scale attention weights, obtained by performing a softmax operation on the result e ij from Step 2.1, n represents the number of elements included in the current input, which needs to be calculated according to the actual sequence length, X i and X j are the feature vectors at position i and position j in the input sequence, W 1is a 4×4 dimensional context weight matrix used to capture the global relationships in the input sequence; α is a learnable parameter used to adjust the relative importance of the original attention weights and context similarity; W 2 is a 128×128 dimensional weight matrix for linear transformation; β and δ are learnable parameters;
[0095] Step 2.3: Introduce a dynamic gating weight g j to the attention output obtained in Step 2.2, and the specific calculation formula is as follows;
[0096]
[0097] where, is the h obtained in Step 2.2 i at the Qth time step, the attention output result under the zth attention head, σ is the activation function sigmoid, which converts the calculation result into the form of a gating weight, W 3 is a 4×1 dimensional gating weight matrix used to perform a linear transformation, H represents the number of attention heads, and different heads are calculated in parallel to capture different feature dimensions at the same time scale, b 1 is a bias term that helps the model better fit the data; in the multi-head attention mechanism, different attention heads are responsible for capturing different levels of features. For example, some heads focus on long-term trends and some heads focus on short-term fluctuation values. The value of H is the same as the value of M.
[0098] Step 2.4: Perform a linear projection on the weighted output of the attention head after introducing the dynamic gating weight, and the formula is:
[0099]
[0100] where, is a 4×4 dimensional matrix for output transformation of the features of the zth attention head, is a 4×4 dimensional matrix for weighted transformation of the features of the zth attention head, W 4 is a 128×128 dimensional global projection matrix that needs to be learned through backpropagation.
[0101] Enhance the features after linear projection by introducing a position-aware function, as follows:
[0102] Introduce a position-aware function and enhance the attention result after linear projection in combination with position encoding according to the following formula:
[0103]
[0104] where, O is the output of the linear projection in Step 2.4 Represents the position-aware function, W 5 is a learnable weight matrix of dimension 128×128, and b 2 is a learnable bias vector, and λ is a hyperparameter used to adjust the influence of the position-aware part.
[0105] Since the network traffic data has temporal characteristics, the position-aware function incorporates position information into the attention output, helping the model better capture the temporal relationship and local features, improving the understanding and processing ability of the data, and thus more accurately judging the normality of the traffic.
[0106] The output O enhanced by position awareness enhanced serves as the input Δ to the interactive position-enhanced feedforward network, and the features output by the interactive position-enhanced feedforward network are used as the final features;
[0107] The formula of the interactive position-enhanced feedforward network is as follows:
[0108]
[0109] where IPA represents "Interactive Position Enhancement", FFN is a standard feedforward network, represents the interaction term, represents the position enhancement term; W 6 is a learnable weight matrix of dimension 128×512, which performs a linear transformation on the input Δ; W 7 is a learnable weight matrix of dimension 512×128 that performs a linear transformation on the result of the first half; b 3 is a learnable bias vector, θ, is a learnable parameter used to adjust the importance of the interaction term and the position enhancement term; K is the number of other vectors that interact with the input vector, and α ik is the weight related to the interaction with Δ, Δ k is the k-th interaction feature representation of the input feature Δ, obtained through linear projection; Δ T is the transpose of the input feature Δ, is a learnable weight matrix of 128×128 for the k-th interaction channel; m is the number of position encodings, and θ il is a learnable weight, calculated according to the relationship between the input vector Δ and the position encoding P l The calculation method is P is the position encoding vector of the elements in the input sequence, using sine-cosine position encoding, is a 128×128-dimensional matrix that projects the input feature Δ and the position encoding P l into the same semantic space.
[0110] The high-order features extracted from the output of the interactive position-enhanced feedforward network enhance the ability of the clustering algorithm to distinguish abnormal traffic, especially the detection of unknown attacks, and the output features improve the accuracy of cosine similarity calculation.
[0111] Step 3: Use the similarity with the normal traffic feature vector to preliminarily determine whether the multi-scale adaptive gating features extracted in Step 2 are abnormal. If it is preliminarily determined to be abnormal, go to Step 4; otherwise, directly go to Step 5.
[0112] Using the similarity with the normal traffic feature vector to preliminarily determine whether the feature vector extracted by the multi-scale adaptive gating feature means calculating the similarity between the feature vector of the network traffic to be detected and the feature vector of the total normal traffic in the database by using cosine similarity.
[0113] Calculate the cosine similarity between the feature vector of the traffic to be detected and the feature vector of the normal traffic in the database. The formula is:
[0114]
[0115] where is the feature vector of the traffic to be detected, is the feature vector of the normal traffic in the database.
[0116] Step 4: According to the preset clustering parameters, update the minimum number of points minPts and the neighborhood radius eps in the clustering parameters based on the silhouette coefficient, CH index, and k-distance graph.
[0117] During the clustering process, use the following distance formula to calculate the distance between the feature vector v 1 , λ 1 :
[0118]
[0119] where w i is the learnable weight of the i-th feature dimension, which can automatically adjust the importance of different dimensions in distance calculation according to the internal law of the data; μ is a learnable parameter used to adjust the overall importance of the new item; G is a preset integer representing the number of new distance-related features, which is determined according to the specific task and data characteristics; Z is the total number of dimensions of the feature vector determined to be preliminarily abnormal; F j (|v i -λ i |) is the j-th new feature, F j (|v i -λ i |) = sigmoid(W (1) ·|vi -λ i |+b (1) ) where sigmoid is a variational activation function that includes the probability modeling of the variational distribution for feature differences, W (1) and b (1) are learned through variational posterior learning to enhance the sensitivity to abnormal data; ω j is the learnable weight related to the new feature j, and the calculation method is
[0120] where is the learnable weight matrix, φ i =|v i -λ i | is the absolute difference between the feature vector v 1 and λ 1 on the i-th feature dimension.
[0121] In the feature space of network traffic data, features in different dimensions have different degrees of importance in judging the similarity of data points. For example, features such as traffic volume, connection duration, and protocol type contribute differently when measuring the similarity of two traffic data points. In the formula, v 1 and λ 1 are two network traffic data points for which the distance is to be calculated, and each data point corresponds to a network traffic sample (e.g., a TCP connection, an HTTP request, or a data packet); v i and λ i respectively represent the specific values of the data points v 1 and λ 1 on the i-th feature dimension. If i = 1 represents "traffic volume", then v 1 =1500 (bytes), λ 1 =80 (bytes), if i = 2 represents "protocol type", then v 2 =0 (TCP), λ 2 =1 (UDP), if i = 3 represents "connection duration", then ν 3 =200 (milliseconds), λ 3 =5000 (milliseconds);
[0122] The process of determining the minimum number of points minPts and the neighborhood radius eps in the clustering parameters is as follows:
[0123] (1) Determine the minimum number of points minPts;
[0124] Step B1: Traverse the candidate parameters: For each value of minPts ∈ {2,..., 10}, perform initial clustering and then calculate the silhouette coefficient s complex (z);
[0125]
[0126] Among them, a(z) is the average distance from sample z to other samples in the same cluster, reflecting the tightness of sample i within its belonging cluster; b(z) is the average distance from sample z to all samples in the nearest cluster, used to measure the separation degree of sample z from other clusters, and is a variational adjustment factor;
[0127] Generally speaking, the closer the value of the silhouette coefficient is to 1, the better the clustering effect; the closer the value is to -1, it indicates that the sample may be assigned to the wrong cluster; the value close to 0 means that the sample may be near the boundary of two clusters. When determining the minPts parameter, according to the change of the silhouette coefficient under different minPts values, the minPts value that maximizes the average silhouette coefficient can be selected as the final parameter to optimize the clustering process and results.
[0128] Step B2: Calculate the Calinski-Harabasz index:
[0129]
[0130] Among them, B(k) is the between-cluster scatter matrix, obtained by calculating the covariance difference between the centroids of each cluster and the global centroid through the clustering result, used to measure the dispersion degree between different clusters; W(k) is the within-cluster scatter matrix, obtained by calculating the covariance difference between the samples within each cluster and the cluster centroid through the clustering result, reflecting the tightness of the samples within each cluster; tr(B(k) 3 ) represents taking the trace after cubing the between-cluster scatter matrix, highlighting the influence of the situation with larger between-cluster distance differences on the clustering quality; det(W(k)) is the determinant of the within-cluster scatter matrix, reflecting the overall characteristics of the sample distribution within the cluster;
[0131] By enriching the description of the between-cluster scatter and within-cluster scatter from different angles through this index, it can more comprehensively and deeply evaluate the clustering quality, help determine appropriate parameters, and optimize the performance of the multi-dimensional index-driven clustering algorithm in network intrusion detection.
[0132] Step B3: Traverse the changes of the silhouette coefficient and CH index under different minPts values, and select the value of minpts corresponding to the value that maximizes s complex +CH(k) as the optimal value of minPts;
[0133] (2) Determine the neighborhood radius eps;
[0134] Step C1: Traverse the candidate parameters: For each value selection of eps ∈ {0.1, 0.2, …, 2}, perform initial clustering, then calculate the k-nearest neighbor distances of each sample, sort the k-nearest neighbor distances in reverse order, and find the inflection point of the sorted k-nearest neighbor distances;
[0135] In the feature space of network traffic data, by calculating the distance from each data point to its k-th nearest neighbor and sorting these distances to plot the k-distance graph, the density distribution of data points can be visually observed. In the k-distance graph, finding the "inflection point" is the key to determining eps. Before the inflection point, the data points are relatively close in distance and may belong to the same clustering cluster; after the inflection point, the distance increases rapidly, and the data points in these regions may belong to different clusters or be noise points. Selecting the distance corresponding to the inflection point as eps can ensure that there are enough samples around the core point to form a density-connected cluster and improve the accuracy of clustering.
[0136] Step C2: Calculate the silhouette coefficient s complex (z);
[0137] Step C3: Traverse the changes in the silhouette coefficient and the inflection point of the distance under different eps values, select the inflection point of the k-nearest neighbor distance and the corresponding maximum value of the silhouette coefficient, and obtain
[0138] Step 5: Using the latest clustering parameters, cluster the multi-scale adaptive gating features corresponding to the network traffic by the density algorithm, analyze the network traffic category based on the clustering results. If the data points of the feature vector corresponding to the network traffic are determined to be noise points or in abnormal clusters, they are identified as suspicious network traffic;
[0139] The clustering effects of the training set and the test set using the technical solution of the present invention are respectively presented in Figure 2 and Figure 3 . From Figure 2 it can be seen that the data is clustered into 4 clusters, namely Cluster 0 to Cluster 2 and Cluster -1. Among them, the number of sample points in Cluster 0 is the most considerable, mainly densely distributed in the area with the abscissa from -2 to 3 and the ordinate from 0 to 1; in contrast, the number of sample points in Cluster 1 and Cluster 2 is less and the distribution is more scattered; while the sample points in Cluster -1 are scattered in multiple positions in the figure.
[0140] Although there is a certain degree of distinguishability between clusters, the boundaries of some clusters, such as Cluster-1, are not very clear, and there is a high probability of a certain degree of overlap or misclassification. In view of this, the subsequent classification and recognition module will divide them into core points, boundary points, and noise points according to the distribution of points in the cluster. At the same time, the final classification decision link will combine the preset rules to make a comprehensive final classification judgment, thus effectively reducing the possibility of misjudgment.
[0141] Step 6: For the suspicious network traffic, make a final classification decision using the preset rules. If it is determined to be abnormal traffic, trigger an alarm and notify the relevant personnel for handling; if it is determined to be normal traffic, allow the traffic to pass normally.
[0142] The classification decision formula is as follows:
[0143]
[0144] Among them, T represents the time when the current traffic occurs (in hours, with a value range of 0-23), P represents the destination port number used by the current traffic, C represents the clustering label of the traffic data point by the multi-dimensional index-driven clustering algorithm (assuming that the normal traffic cluster label is 0, and others are abnormal-related labels), T normal is the proportion matrix of normal traffic in different time periods, T normal [i] represents the proportion of normal traffic in the historical data in the i-th hour (i = 0, 1,..., 23), P sensitive is the preset set of sensitive ports, which contains port numbers that may pose risks. is the weight coefficient, and
[0145] is used to adjust the importance of different factors in the judgment and can be adjusted according to the actual network environment and security requirements. Set a threshold θ (value range 0-1, which can be adjusted according to the actual situation). When f(T, P, C) ≥ θ, the traffic is determined to be abnormal traffic; when f(T, P, C) < θ, the traffic is determined to be normal traffic.
[0146] There is a time factor in the formula: This term considers the difference between the current time traffic and the proportion of historical normal traffic. If the proportion of normal traffic at the current time is lower, that is, the value of 1 - T normal [i] is larger, it means that the possibility of abnormal traffic appearing at this time is higher, and the value of this term is also larger.
[0147] Port factor:
[0148]
[0149] To determine whether the port for traffic usage is a sensitive port. If it is a sensitive port, the value of this item is indicating a certain abnormal risk; if not, the value is 0.
[0150] Clustering factor:
[0151]
[0152] According to the clustering result, if the clustering label is not the normal traffic cluster label (C≠0), it means that the clustering algorithm believes that this traffic has an abnormal tendency, and the value of this item is If it is the normal traffic cluster label (C = 0), the value is 0.
[0153] By adjusting the weight coefficient and the threshold θ, it can adapt to different network environments and security requirements, and more accurately classify and make decisions on traffic.
[0154] From Figure 4 the data, it can be seen that compared with the model corresponding to the method provided by the technical solution of the present invention, traditional models represented by the CNN model have gaps in the attack discrimination accuracy of both the training set and the test set, and the difference is more prominent in the test set stage.
[0155] This is mainly because the test set contains unknown attack samples, and traditional models are difficult to effectively identify them. In the training stage of the model corresponding to the technical solution provided by the present invention, through the multi-scale adaptive gating feature extraction module and the multi-dimensional index-driven clustering technology, it can fully capture and deeply analyze various features. These feature processing capabilities learned through training enable the model to accurately discriminate when facing unknown attacks in the test set, thus showing good discrimination ability for unknown attacks.
[0156] Embodiment 2
[0157] A detection system based on the above multi-scale feature extraction and variational clustering intrusion detection method, including:
[0158] Network traffic data collection unit: Extract the source IP, destination IP, port number, protocol type, traffic size, and connection duration features, and form a feature vector;
[0159] Standardization processing unit: Standardize the feature vector;
[0160] Feature acquisition unit: Use the multi-head attention mechanism to extract features from the feature vector processed by the standardization processing unit, and perform a linear combination of the weighted outputs of all attention heads to generate the multi-scale adaptive gating features corresponding to the final attention result;
[0161] Abnormal preliminary judgment unit: Based on the similarity with the normal traffic feature vector, preliminarily judge whether the feature vector of the feature acquisition unit is abnormal or not;
[0162] Clustering parameter update unit: According to the preset clustering parameters, use the feature vectors corresponding to the network traffic preliminarily judged as abnormal by the abnormal preliminary judgment unit, and update the minimum number of points minPts and the neighborhood radius eps in the clustering parameters based on the silhouette coefficient, CH index, and k-distance graph;
[0163] Secondary judgment unit: Use the current latest clustering parameters to cluster the multi-scale adaptive gating features corresponding to the network traffic by the density algorithm, analyze the network traffic category according to the clustering result. If the feature vector data point corresponding to the network traffic is determined to be a noise point or in an abnormal cluster, it is identified as suspicious network traffic;
[0164] Decision-making unit: For the suspicious network traffic, make a final classification decision using the preset rules. If it is determined to be abnormal traffic, trigger an alarm and notify relevant personnel for handling; if it is determined to be normal traffic, allow the traffic to pass normally.
[0165] It should be understood that for the specific implementation process of each module, please refer to the above method content, which will not be elaborated herein. And the above division of functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined, some functional modules can be split, and each functional module can be implemented in software, hardware, or a combination of software and hardware. Among them, the software and hardware devices include but are not limited to general computer devices, programmable gate arrays, digital signal processors, microprocessors, and their corresponding programming or burning software.
[0166] Embodiment 3
[0167] A computer device, including
[0168] One or more processors;
[0169] A memory storing one or more computer programs;
[0170] Wherein, the processor calls the computer program to implement:
[0171] The above steps of multi-scale feature extraction and variational clustering intrusion detection.
[0172] For the specific implementation process of each step, please refer to the description of the foregoing method.
[0173] Embodiment 4
[0174] A computer-readable storage medium storing a computer program, and the computer program is called by a processor to implement:
[0175] Steps of the intrusion detection method based on multi-scale feature extraction and variational clustering
[0176] For the specific implementation process of each step, please refer to the description of the foregoing method.
[0177] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the readable storage medium may also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store the data that has been output or is to be output.
[0178] Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing readable storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other media that can store program codes.
[0179] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The present application is a device that is used to implement the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram by the instructions executed by the processor according to the methods, devices (systems), and computer program products of the embodiments of the present application. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram.
[0180] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art according to the technical solutions of the present invention, whether modified or replaced, as long as they do not depart from the purpose and scope of the present invention, also belong to the protection scope of the present invention.
Claims
1. An intrusion detection method based on multi-scale feature extraction and variational clustering, characterized in that: include: Step 1: Collect network traffic data and extract source IP, destination IP, port number, protocol type, traffic size, and connection duration features to form a feature vector; standardize the feature vector; Step 2: Use the multi-head attention mechanism to extract features from the feature vector processed in step 1, and linearly combine the weighted outputs of all attention heads to generate multi-scale adaptive gating features corresponding to the final attention result; Step 3: Using the similarity between the feature vectors of normal traffic, make a preliminary judgment on whether the multi-scale adaptive gating feature extracted in step 2 is abnormal. If it is initially judged to be abnormal, proceed to step 4; otherwise, proceed directly to step 5. Step 4: According to the preset clustering parameters, the minimum number of points minPts and the neighborhood radius eps in the clustering parameters are updated based on the silhouette coefficient, CH index and k distance graph, and then go to step 5; Step 5: Using the latest clustering parameters, the multi-scale adaptive gating features corresponding to the network traffic are clustered using a density algorithm. The network traffic category is analyzed based on the clustering results. If the feature vector data point corresponding to the network traffic is determined to be a noise point or in an abnormal cluster, it is identified as suspicious network traffic. Step 6: For suspicious network traffic, use pre-set rules to make the final classification decision. If it is determined to be abnormal traffic, an alarm is triggered and relevant personnel are notified to handle it; if it is determined to be normal traffic, the traffic is allowed to pass normally.
2. The method according to claim 1, characterized in that The specific steps of multi-scale adaptive gating feature extraction are as follows: Step 2.1: Calculate multi-scale attention weights: Among them, e ij represents the multi-scale attention weights of position i and position j in the feature vector corresponding to the network traffic. The larger the weight value, the stronger the correlation between the features of the two positions at multiple scales, and the greater the impact on subsequent feature extraction and classification decisions; M is the number of multi-scales; X i and X j are the feature vectors at position i and position j in the feature vector corresponding to the network traffic, is the 128×128-dimensional weight matrix of the query matrix and the key matrix at the mth scale; d (m) yes The dimension is used for scaling operations, and the superscript T represents the transpose of the matrix; Step 2.2: Based on the multi-scale attention weights, calculate the attention output of each position on the feature vector according to the following formula; Among them, h i represents the attention output of the i-th position in the feature vector, a ij is the original attention weights for positions i and j calculated based on the multi-scale attention weights, which are obtained from the result e in step 2.1 ij After the softmax operation, n represents the number of elements contained in the current input, which needs to be calculated based on the actual sequence length. 1 is a 4×4 dimension context weight matrix used to capture the global relationship in the input sequence; α is a learnable parameter used to adjust the relative importance of the original attention weight and context similarity; W 2 is a 128×128 dimensional weight matrix for linear transformation; β and δ are learnable parameters; Step 2.3: Introduce dynamic gating weights g to the attention output obtained in step 2.2 j , the specific calculation formula is as follows; in, is the h obtained in step 2.2 i At the Qth time step, the attention output result under the zth attention head, σ is the activation function sigmoid, which converts the calculation result into the form of gated weights, W 3 is a 4×1 dimension gating weight matrix used for Perform linear transformation, H represents the number of attention heads, and multiple heads are used for parallel calculation to capture different feature dimensions at the same time scale. b1 is a bias term that helps the model fit the data better. Step 2.4: Linearly project the weighted output of the attention head after the dynamic gating weights are introduced. The formula is: in, is a 4×4 dimensional matrix that transforms the output features of the zth attention head. is a 4×4-dimensional matrix for weighted transformation of the features of the z-th attention head, W 4 is a 128×128 dimensional global projection matrix that needs to be learned through back-propagation.
3. The method according to claim 2, characterized in that The position-aware function is introduced to enhance the features after linear projection, as follows: The position-aware function is introduced and combined with the position encoding to enhance the attention result after linear projection according to the following formula: Where O is the output of the linear projection in step 2.4, represents the position-aware function, W 5 is a 128×128 dimensional learnable weight matrix, b2 is a learnable bias vector, and λ is a hyperparameter used to adjust the influence of the position-aware part.
4. The method according to claim 3, characterized in that Output O after location-aware enhancement enhanced As the input Δ of the interactive position enhanced feed-forward network, the features output by the interactive position enhanced feed-forward network are used as the final features; The interactive position enhanced feedforward network formula is as follows: Among them, IPA stands for "Interactive Position Augmentation", FFN is a standard feed-forward network, represents the interaction term, represents the position enhancement term; W6 is a 128×512-dimensional learnable weight matrix that performs a linear transformation on the input Δ; W7 is a 512×128-dimensional learnable weight matrix that performs a linear transformation on the result of the first half; b3 is a learnable bias vector, θ, is a learnable parameter used to adjust the importance of interaction terms and position enhancement terms; K is the number of other vectors that interact with the input vector, α ik is the weight associated with Δ and the interaction, Δ k is the kth interactive feature representation of the input feature Δ, obtained by linear projection; Δ T is the transpose of the input feature Δ, is the 128×128 learnable weight matrix of the kth interaction channel; m is the number of positional encodings, θ il is a learnable weight, based on the input vector Δ and the position code P l The relationship is calculated as follows: P is the position encoding vector of the elements in the input sequence using sine-cosine position encoding, To encode the input feature Δ and position P l Projected to a 128×128 matrix in the same semantic space.
5. The method according to claim 1, characterized in that By using the similarity with the normal traffic feature vector, a preliminary abnormality judgment is made on the feature vector extracted by the multi-scale adaptive gating feature, which means that the cosine similarity is used to calculate the similarity between the feature vector of the network traffic to be detected and the feature vector of the total normal traffic in the database.
6. The method according to claim 1, characterized in that In the clustering process, the following distance formula is used to calculate the distance between the feature vectors v1 and λ1 that are judged to be preliminary abnormalities: Among them, w i is the learnable weight of the i-th feature dimension, which can automatically adjust the importance of different dimensions in distance calculation according to the inherent laws of the data; Hong is a learnable parameter used to adjust the overall importance of the newly added items; G is a preset integer, which represents the number of newly added distance-related features, which is determined according to the specific task and data characteristics; Z is the total number of feature vector dimensions judged as preliminary abnormalities; F j (|v i -λ i |) is the jth newly added feature, F j (|v i -λ i |)=sigmoid(W (1) ·|v i -λ i |+b (1) ), sigmoid is a variational activation function, including the probability modeling of feature differences by variational distribution, W (1) and b (1) Through variational posterior learning, it is used to enhance the sensitivity to abnormal data; ω j is the learnable weight associated with the newly added feature j, calculated as in, is the learnable weight matrix, φ i =|v i -λ i | is the absolute difference between feature vectors v1 and λ1 in the i-th feature dimension.
7. The method according to claim 1, characterized in that The process of determining the minimum number of points minPts and the neighborhood radius eps in the clustering parameters is as follows: (1) Determine the minimum number of points minPts; Step B1: Traverse candidate parameters: perform initial clustering for each minPts∈{2,...,10} selected value, and then calculate the silhouette coefficient s of the eigenvector corresponding to the network traffic complex (z); Among them, a(z) is the average distance from sample z to other samples in the same cluster, reflecting the closeness of sample i in its cluster; b(z) is the average distance from sample z to all samples in the nearest cluster, which is used to measure the degree of separation between sample z and other clusters. and is the variation adjustment factor; Step B2: Calculate the Calinski-Harabasz index: Among them, B(k) is the inter-cluster dispersion matrix, which is obtained by calculating the covariance difference between the centroid of each cluster and the global centroid through the clustering results, and is used to measure the degree of dispersion between different clusters; W(k) is the intra-cluster dispersion matrix, which is obtained by calculating the covariance difference between the samples within each cluster and the cluster centroid through the clustering results, and reflects the compactness of the samples within each cluster; tr(B(k) 3 ) represents the trace of the inter-cluster dispersion matrix after cubic operation, highlighting the impact of large inter-cluster distance differences on clustering quality; det(W(k)) is the determinant of the intra-cluster dispersion matrix, reflecting the overall characteristics of the sample distribution within the cluster; Step B3: Traverse the changes of silhouette coefficient and CH index under different minPts values, and select complex +The minPts value corresponding to the maximum value of CH(k) is taken as the optimal value of minPts; (2) Determine the neighborhood radius eps; Step C1: Traverse candidate parameters: perform initial clustering for each eps∈{0.1, 0.2, …, 2} selected value and then calculate the k nearest neighbor distance of each sample, sort the k nearest neighbor distance in reverse order, and find the inflection point of the sorted k nearest neighbor distance; Step C2: Calculate the silhouette coefficient s of the eigenvector corresponding to the network traffic after initial clustering complex (z); Step C3: Traverse the changes of silhouette coefficient and distance inflection point under different eps values, select the k nearest neighbor distance inflection points and the corresponding maximum value of silhouette coefficient, and obtain 8. A detection system based on the multi-scale feature extraction and variational clustering intrusion detection method according to any one of claims 1 to 7, characterized in that: include: Network traffic data collection unit: extracts source IP, destination IP, port number, protocol type, traffic size, and connection duration features to form a feature vector; Standardization processing unit: standardize the feature vector; Feature acquisition unit: Use the multi-head attention mechanism to extract features from the feature vector processed by the normalization processing unit, and linearly combine the weighted outputs of all attention heads to generate multi-scale adaptive gating features corresponding to the final attention result; Anomaly preliminary judgment unit: uses the similarity between the feature vector and the normal traffic feature vector to make a preliminary judgment on whether the feature vector of the feature acquisition unit is abnormal; Clustering parameter updating unit: based on the preset clustering parameters, using the feature vector corresponding to the network traffic preliminarily judged as abnormal by the abnormal preliminary judgment unit, the minimum number of points minPts and the neighborhood radius eps in the clustering parameters are updated based on the silhouette coefficient, CH index and k distance graph; Secondary judgment unit: using the latest clustering parameters, the multi-scale adaptive gating features corresponding to the network traffic are clustered using a density algorithm, and the network traffic category is analyzed based on the clustering results. If the feature vector data point corresponding to the network traffic is judged to be a noise point or in an abnormal cluster, it is identified as suspicious network traffic; Decision-making unit: For suspicious network traffic, pre-set rules are used to make the final classification decision. If it is determined to be abnormal traffic, an alarm is triggered and relevant personnel are notified to handle it; if it is determined to be normal traffic, the traffic is allowed to pass normally.
9. A computer device, characterized in that: include one or more processors; a memory storing one or more computer programs; The processor calls the computer program to implement: The steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored, which is called by a processor to implement: The steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Network traffic anomaly classification method and system based on artificial intelligence
CN120512317A
Intelligent ship body part classification method and system based on internal and external contour features
CN121904466A