Network anomaly monitoring method and system for switch
By combining the switch-level model and central model architecture, multi-scale feature extraction and deep timing analysis of network traffic data is solved, and the problems of low network anomaly detection efficiency, high false alarm rate and lack of adaptive capabilities in the prior art are solved, and high accuracy and robust network anomaly detection and adaptive defense strategy generation are achieved.
Patent Information
- Application Number
- CN202411143610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-20
AI Technical Summary
The existing network anomaly detection methods are difficult to deal with complex and changeable network attacks, which are inefficient, have high false alarm rates, serious misreports, and lack adaptability and distributed detection architecture, which has single point of failure risk and privacy protection problems.
Using a architecture that combines switch-level model and central model, multi-scale feature extraction and feature dimensionality reduction of network traffic data is performed. The model is trained through the federated learning framework to realize distributed anomaly detection, and in-depth timing analysis is performed in combination with the LSTM network and attention mechanism to generate an adaptive defense strategy.
It improves the accuracy and robustness of network anomaly detection, reduces the risk of single point failure of the system, enhances detection efficiency and system scalability, protects user privacy, and improves the identification ability of complex attack patterns and the effectiveness of defense measures.
Smart Images

Figure CN119071052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of network anomaly monitoring, and in particular to a method and system for monitoring network anomaly of a switch. Background Art
[0002] As the scale of networks continues to expand and their complexity increases, network security faces unprecedented challenges. Traditional network anomaly detection methods often rely on fixed rules and static thresholds, which are difficult to cope with increasingly complex and changeable network attacks. These methods are inefficient when processing large-scale network traffic data and cannot detect hidden abnormal behaviors in a timely manner. At the same time, most existing anomaly detection systems adopt a centralized architecture, which is difficult to adapt to distributed network environments, and there are single point failure risks and privacy protection issues.
[0003] In addition, existing anomaly detection methods usually only focus on single-dimensional features, ignoring the multi-dimensional characteristics of network abnormal behavior, resulting in high false alarm rates and serious underreporting. In terms of anomaly analysis, there is a lack of in-depth explanation and correlation analysis of detection results, making it difficult to provide valuable decision support for network administrators. At the same time, most anomaly detection systems lack adaptive capabilities and cannot dynamically adjust detection strategies and defense measures according to changes in the network environment. Summary of the invention
[0004] The present application provides a method and system for monitoring network anomalies of a switch, thereby improving the accuracy of monitoring network anomalies of the switch.
[0005] A first aspect of the present application provides a network anomaly monitoring method for a switch, the network anomaly monitoring method for a switch comprising:
[0006] Collecting historical network traffic data of the network switch and preprocessing it to obtain preprocessed network traffic data;
[0007] Performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector;
[0008] Training a switch-level model and a central model based on the fused feature vector;
[0009] Performing anomaly detection on the real-time network traffic data of the network switch through the switch-level model to obtain preliminary anomaly detection results, and performing in-depth analysis in combination with the central model to obtain in-depth anomaly detection results;
[0010] Performing spatiotemporal correlation analysis and multi-source data fusion on the preliminary anomaly detection results and the deep anomaly detection results to obtain a comprehensive anomaly analysis report;
[0011] An adaptive defense strategy is generated based on the comprehensive anomaly analysis report, and the adaptive defense strategy is executed in the network switch.
[0012] A second aspect of the present application provides a network anomaly monitoring system for a switch, the network anomaly monitoring system for the switch comprising:
[0013] A collection module is used to collect historical network flow data of the network switch and pre-process it to obtain pre-processed network flow data;
[0014] A feature extraction module, used for performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector;
[0015] A training module, used for training a switch-level model and a central model based on the fused feature vector;
[0016] A detection module, configured to perform anomaly detection on the real-time network traffic data of the network switch through the switch-level model to obtain a preliminary anomaly detection result, and perform in-depth analysis in combination with the central model to obtain a deep anomaly detection result;
[0017] An analysis module, used to perform spatiotemporal correlation analysis and multi-source data fusion on the preliminary anomaly detection results and the deep anomaly detection results to obtain a comprehensive anomaly analysis report;
[0018] An execution module is used to generate an adaptive defense strategy based on the comprehensive anomaly analysis report and execute the adaptive defense strategy in the network switch.
[0019] The third aspect of the present application provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned network anomaly monitoring method for the switch.
[0020] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned network anomaly monitoring method for a switch.
[0021] Compared with the prior art, the present application has the following beneficial effects: by extracting multi-scale features from network traffic data, combining time, frequency domain, wavelet and graph features, the multi-dimensional features of network anomalies are fully captured, and the accuracy and robustness of anomaly detection are improved. The architecture combining the switch-level model and the central model is adopted to realize distributed anomaly detection, reduce the risk of single-point failure of the system, and improve the detection efficiency and system scalability. The switch-level model is trained using a federated learning framework to avoid direct sharing of raw data and effectively protect user privacy and sensitive information. The introduction of LSTM network and attention mechanism to perform deep time series analysis on network traffic can effectively capture long-term dependencies and key time steps, and improve the ability to identify complex attack patterns. Through dynamic baseline model and Bayesian inference, adaptive calculation of anomaly scores is realized, which improves the flexibility of detection and adaptability to environmental changes. Combining time series analysis and graph convolutional neural network, spatiotemporal correlation analysis of abnormal events is performed, which helps to discover potential attack paths and propagation patterns. The knowledge graph of abnormal events is constructed, and graph embedding and semantic analysis are performed to provide network administrators with intuitive and explainable anomaly analysis reports. By using reinforcement learning and deep Q-network, defense strategies are dynamically generated according to the network environment and attack characteristics, which improves the effectiveness and pertinence of defense measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] The structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantive technical significance. Any structural modification, change in proportion or adjustment of size, without affecting the effects and purposes that can be achieved by the present invention, should still fall within the scope of the technical contents disclosed by the present invention.
[0024] Figure 1 It is a flow chart of a method for monitoring network anomalies of a switch provided by an embodiment of the present invention;
[0025] Figure 2 The present invention is a schematic block diagram of the structure of a network anomaly monitoring system for a switch provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0028] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0029] It should be further understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Figure 1 , an embodiment of the network anomaly monitoring method of the switch in the embodiment of the present application includes:
[0030] Step 100: Collect historical network traffic data of the network switch and pre-process it to obtain pre-processed network traffic data;
[0031] It is understandable that the execution subject of the present application can be a network anomaly monitoring system of a switch, or a terminal or a server, which is not limited here. The present application embodiment is described by taking a server as the execution subject as an example.
[0032] Specifically, as a data communication device, the network switch records a large amount of network traffic data, including the start and end time of each network connection, the size of the data packet, the type of protocol transmitted, etc. Use network management protocols such as SNMP (Simple Network Management Protocol) or NetFlow technology to automatically collect the required data from the network switch. In the process of collecting data, set a scheduled task or use a network traffic monitoring tool to collect traffic data in real time or periodically. In order to avoid the collected data being too large or irrelevant, pre-set filtering conditions to only collect data related to network anomaly monitoring, such as traffic information of a specific port or a specific time period. Preprocess the raw data, including data cleaning, data conversion, and data standardization. In the data cleaning process, remove redundant information and invalid data, such as duplicate records, incomplete data packet information, or noise data. Convert the collected raw data into a format that is easier to analyze, such as converting the binary format of the data packet into a readable text format. Data standardization is to unify data from different sources so that they have consistent measurement units and formats. Aggregate the data by time period or traffic characteristics to reduce the data size and improve the efficiency of data processing to obtain preprocessed network traffic data.
[0033] Step 200: performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector;
[0034] Specifically, the pre-processed network traffic data is segmented into multi-scale sliding windows to capture the features of the data at different time scales. By setting different sliding window sizes, the statistical features in each window, such as mean, variance, kurtosis and skewness, are calculated. The statistical features reflect the time change trend and mutation characteristics of the network traffic, forming a time scale feature set. The time scale feature set is analyzed in the frequency domain by fast Fourier transform to extract frequency domain features. The features in the time domain are converted into frequency domain features to capture the periodicity and spectrum characteristics in the network traffic. The pre-processed network traffic data is subjected to discrete wavelet transform, and the data is decomposed and reconstructed at different resolutions to extract the multi-resolution features of the data. The generation of wavelet feature set can effectively capture the instantaneous changes and singularity features in the network traffic, especially when detecting short-term abnormal events, wavelet features have significant advantages. A network topology graph is constructed based on the pre-processed network traffic data. By calculating the graph structure features such as the degree, clustering coefficient, and path length of the nodes, a graph feature set is formed. These features reflect the distribution and correlation of the network traffic in the topological structure, which can help the system understand the complexity of the traffic and the network location of the abnormal points. The time scale feature set, frequency domain feature set, wavelet feature set and graph feature set are concatenated to form an initial feature vector. A multi-layer autoencoder network structure is introduced to perform feature dimension reduction and nonlinear transformation. By designing and training the autoencoder model, a more representative and compact fused feature vector is extracted from the initial feature vector. The encoder part of the autoencoder maps the high-dimensional input features to the low-dimensional latent space by learning the low-dimensional representation of the data, reducing the dimension of the features and maintaining the main information characteristics of the data. The encoder part of the autoencoder performs a nonlinear transformation on the initial feature vector to obtain a fused feature vector.
[0035] Step 300: training a switch-level model and a central model based on the fused feature vector;
[0036] It should be noted that the fused feature vector is segmented and divided into switch-level training data and central training data according to a preset ratio. The data partitioning strategy can ensure that sufficient data is obtained at the switch level and the central level for model training and optimization. The adaptive isolation forest model is initialized using the switch-level training data. The adaptive isolation forest model is an unsupervised learning algorithm for anomaly detection, which can effectively identify anomalies. The initial parameters and the number of trees are set to obtain the initial switch-level model. In order to perform distributed training on multiple switches, the federated learning framework is introduced. On each network switch participating in the training, the initial switch-level model is updated with local data to obtain a local updated model. The core idea of the federated learning framework is to allow each device to jointly train a global model without sharing local data, which can protect data privacy and make full use of distributed data resources. After the parameters of the local update model are updated, they are encrypted, and then the model updates from each network switch are aggregated on the central server using a secure aggregation algorithm to obtain a global switch-level model. The training results from various locations are integrated in a secure manner to ensure the performance of the global model. The global switch-level model is distributed to each network switch participating in the training for the next round of local training. This process is iterated continuously until the model converges and a stable and accurate switch-level model is obtained. In the process of building the central model, an LSTM network structure is constructed using the central training data to obtain the initial LSTM model. LSTM, or long short-term memory network, is a special recursive neural network that is good at processing sequence data and time series problems. The initial LSTM model usually includes an input layer, a multi-layer bidirectional LSTM layer, and an output layer. Such a structure helps capture the bidirectional dependencies in the time series. An attention layer is added after the LSTM layer to calculate the importance weights of different time steps to obtain an attention-enhanced LSTM model. The attention mechanism can help the model better focus on the key parts of the time series and improve the prediction accuracy of the model. The attention-enhanced LSTM model is trained using the back-propagation algorithm to minimize the prediction error and obtain the optimized LSTM model. The optimized LSTM model is combined with the fully connected layer to construct a hybrid model structure to form the initial central model. The initial central model is fine-tuned and verified using the central training data to ensure the generalization ability and accuracy of the model, and finally the central model is obtained.
[0037] Step 400: Perform anomaly detection on the real-time network traffic data of the network switch through the switch-level model to obtain preliminary anomaly detection results, and perform in-depth analysis in combination with the central model to obtain in-depth anomaly detection results;
[0038] Specifically, the real-time network traffic data of the network switch is preprocessed to ensure the integrity and consistency of the data. The preprocessing process includes steps such as data cleaning, denoising, and format conversion to obtain standardized real-time preprocessed data. Anomaly scores are calculated for the real-time preprocessed data based on the switch-level model. The switch-level model is usually trained by adaptive isolation forests or other unsupervised learning methods, which can effectively capture abnormal patterns in the data. Potential abnormal events are identified by calculating the original anomaly score. The original anomaly score is compared with a preset threshold to determine whether an abnormal alarm is triggered. The comparison process generates preliminary anomaly tags, marking those data points with scores exceeding the threshold as abnormal data points. The preliminary anomaly tags are quickly classified. According to the predefined rule set, the detected anomalies are divided into different types, which may include network attacks, equipment failures, traffic surges, etc. The preliminary anomaly tags, preliminary anomaly types, and corresponding context information are combined to obtain preliminary anomaly detection results. The context information may include the time when the anomaly occurred, information about related devices, the source and destination of the traffic, etc. The real-time preprocessed data is feature extracted and converted to conform to the input format of the central model. The central model is usually based on an LSTM network structure. The data after feature extraction and transformation is input into the central model for time series analysis through the LSTM network. The LSTM model can capture the long-term and short-term dependencies in the data, so as to better understand the temporal dynamics of the data. The attention mechanism is applied to analyze the time series features. The attention mechanism can help the model focus on the important parts of the time series, and obtain the weighted time series features by calculating the importance weights of different time steps. The anomaly scoring and classification are performed based on the weighted time series features to obtain the deep anomaly detection results.
[0039] The real-time preprocessed data is subjected to feature normalization, and each feature is scaled to the same range, usually zero mean and unit variance, to eliminate the influence between different dimensions and obtain a standardized feature vector. The sample space is constructed based on the standardized feature vector, and then the sample space is randomly subspaced. The entire sample space is divided into multiple subspaces to facilitate the construction of isolation trees. In each subspace, a feature subset is randomly selected, and these feature subsets are used to construct an isolation tree structure to form an isolation forest. Isolation forest is an unsupervised learning method specifically used for anomaly detection. It evaluates the isolation of data points through the structure of multiple isolation trees. For each data sample, its path length in the isolation forest is calculated to obtain a set of sample path lengths. Path length refers to the depth at which the sample is isolated in the tree structure. Generally, the path length of abnormal samples is shorter. The sample path length set is statistically analyzed, and the average path length and standard deviation are calculated to obtain the path length statistical characteristics. According to the path length statistical characteristics, an adaptive threshold function is designed, and the sample path length is normalized by the function to generate an initial anomaly score. The initial anomaly score is a preliminary abnormality assessment of each sample. The higher the score, the more likely the sample is abnormal. The initial anomaly score is exponentially smoothed to eliminate short-term random fluctuations, highlight long-term abnormal trends, and obtain a smoothed anomaly score. Based on the smoothed anomaly score, a dynamic baseline model is constructed. The model can track the abnormal distribution of historical data and calculate the abnormal distribution parameters. Bayesian inference is performed on the smoothed anomaly score and the abnormal distribution parameters, and the posterior probability distribution is calculated to obtain the corrected anomaly score. Bayesian inference provides a reasonable estimation method by combining prior knowledge and observed data, which can more accurately evaluate abnormality. A multi-level warning mechanism is designed based on the corrected anomaly score to quantitatively evaluate the degree of anomaly. The multi-level warning mechanism divides the anomaly scores into different levels, such as mild, moderate, and severe.
[0040] Step 500: Perform spatiotemporal correlation analysis and multi-source data fusion on the preliminary anomaly detection results and the deep anomaly detection results to obtain a comprehensive anomaly analysis report;
[0041] Specifically, a time series graph is constructed based on the preliminary anomaly detection results and the deep anomaly detection results, and the detected abnormal events are sorted in time to form an abnormal event timeline. Based on the abnormal event timeline, the time correlation between events is calculated to obtain a time correlation matrix. The time correlation matrix can reveal the time dependency between abnormal events and help identify whether the events are triggered or dependent in time. According to the network topology information and the network entities involved in the abnormal events, a spatial relationship graph is constructed to obtain a spatial distribution graph of abnormal events, showing the geographical location of the abnormal events in the network and the relationship between devices. The abnormal event spatial distribution graph is processed by a graph convolutional neural network to extract spatial features and propagation patterns from it to obtain a spatial correlation matrix. The graph convolutional neural network is a deep learning method for processing graph data, which can effectively extract the features of nodes in the graph and the relationship information of their neighbor nodes. The time correlation matrix and the spatial correlation matrix are tensor fused to construct a spatiotemporal correlation tensor to obtain a comprehensive correlation representation. Multi-source heterogeneous data such as network device logs and security device alarms are collected and standardized to obtain standardized multi-source data. Based on the standardized multi-source data and the comprehensive correlation representation, a knowledge graph of abnormal events is constructed to obtain an initial knowledge graph. Knowledge graph is a structured graph that can express entities and their relationships. In anomaly detection, it can effectively integrate information from multiple data sources. Perform graph embedding analysis on the initial knowledge graph, map nodes and relationships to a low-dimensional vector space, and obtain a graph embedding representation. Based on the graph embedding representation, group and associate abnormal events to form abnormal event clusters. Abnormal event clusters represent groups of abnormal events with similar characteristics or that are related to each other. Perform feature extraction and semantic analysis on abnormal event clusters to generate a comprehensive abnormal analysis report. The report includes a detailed description of each abnormal event, the scope of impact, the severity, and possible causes.
[0042] Step 600: Generate an adaptive defense strategy based on the comprehensive anomaly analysis report, and execute the adaptive defense strategy in the network switch.
[0043] Specifically, feature extraction is performed on the comprehensive anomaly analysis report. The abnormal events and their related features described in the report are converted into numerical representations to form an abnormal feature vector. Based on the feature vector, a set of candidate defense strategies that may be applicable to the current situation is screened from the predefined defense strategy library to form an initial strategy set. Based on the initial strategy set and historical defense effect data, a reinforcement learning environment is constructed. The environment includes a state space, an action space, and a reward function. The state space is described by the abnormal feature vector, the action space includes all optional defense strategies, and the reward function evaluates the effect of each defense strategy. The purpose of building a defense strategy learning environment is to provide a dynamic optimization platform for strategy selection. In this environment, a deep Q network algorithm is used for strategy training. The deep Q network is an algorithm that combines Q-learning and neural networks and can perform efficient learning in complex strategy spaces. The initial Q value function is obtained through training, which evaluates the expected return of selecting a specific action under different states. The initial Q value function is subjected to policy iteration and value iteration, and the Q value is updated through multiple iterations to make it more accurately reflect the effects of different defense strategies. This process continues until the Q value converges to obtain the optimized Q value function. Based on the optimized Q-value function, the ε-greedy strategy is used to select the optimal defense action. The ε-greedy strategy is a strategy that balances exploration and exploitation. It selects the current optimal action (defense strategy) in most cases, but also retains a certain probability to randomly select other strategies to prevent falling into the local optimum. The feasibility analysis is performed on the selected candidate defense strategies to evaluate the feasibility and effectiveness of each strategy in the actual environment. According to the feasibility score, the candidate defense strategies are adjusted and optimized as necessary to ensure that the selected adaptive defense strategy is not only feasible in theory, but also can be effectively executed in practice. The adaptive defense strategy is converted into specific configuration instructions to form a defense execution instruction set. These instruction sets include detailed configuration steps and parameter settings, and these instructions are sent to the corresponding network switches through network management protocols (such as SNMP or NetConf). The network switch adjusts the configuration in real time according to the instructions, so as to quickly respond to and resist the detected network anomalies and realize the function of adaptive defense.
[0044] In the embodiment of the present application, by extracting multi-scale features from network traffic data, combining time, frequency domain, wavelet and graph features, the multi-dimensional features of network anomalies are fully captured, and the accuracy and robustness of anomaly detection are improved. The architecture combining the switch-level model and the central model is adopted to realize distributed anomaly detection, reduce the risk of single-point failure of the system, and improve the detection efficiency and system scalability. The switch-level model is trained using a federated learning framework to avoid direct sharing of raw data and effectively protect user privacy and sensitive information. The introduction of LSTM network and attention mechanism to perform deep time series analysis on network traffic can effectively capture long-term dependencies and key time steps, and improve the ability to identify complex attack patterns. Through dynamic baseline model and Bayesian inference, adaptive calculation of anomaly scores is realized, which improves the flexibility of detection and adaptability to environmental changes. Combining time series analysis and graph convolutional neural network, spatiotemporal correlation analysis of abnormal events is performed, which helps to discover potential attack paths and propagation patterns. The knowledge graph of abnormal events is constructed, and graph embedding and semantic analysis are performed to provide network administrators with intuitive and explainable abnormal analysis reports. By using reinforcement learning and deep Q-network, defense strategies are dynamically generated according to the network environment and attack characteristics, which improves the effectiveness and pertinence of defense measures.
[0045] In a specific embodiment, the process of executing step 200 may specifically include the following steps:
[0046] Perform multi-scale sliding window segmentation on the pre-processed network traffic data, and calculate the statistical features of different time scales to obtain the time scale feature set;
[0047] According to the time scale feature set, the frequency domain features are extracted using fast Fourier transform to obtain the frequency domain feature set;
[0048] Perform discrete wavelet transform on the pre-processed network traffic data and extract multi-resolution features to obtain a wavelet feature set;
[0049] Construct a network topology graph based on the preprocessed network traffic data, and calculate the graph structure features to obtain a graph feature set;
[0050] The time scale feature set, frequency domain feature set, wavelet feature set and graph feature set are concatenated to obtain the initial feature vector;
[0051] A multi-layer autoencoder network structure is designed according to the initial feature vector, the autoencoder model is trained to obtain the encoder part, and the initial feature vector is nonlinearly transformed based on the encoder part to obtain a fused feature vector.
[0052] Specifically, the preprocessed network traffic data is segmented by multi-scale sliding windows. Sliding window segmentation is a commonly used data processing technology. By applying time windows of different sizes to time series data, the changes of data at different time scales can be captured. Set multiple different window sizes, and the size of the window represents different time scales. The data in each window will be used to calculate statistical features, such as mean, standard deviation, kurtosis, skewness, etc. These features can reflect the dynamic changes of network traffic at various time scales. Through this process, a set of time scale feature sets is obtained to reflect the characteristic changes of network traffic at different time scales. According to the time scale feature set, the frequency domain features are extracted using fast Fourier transform. Fast Fourier transform is an algorithm that converts time domain signals into frequency domain signals, which can reveal the periodic components and frequency characteristics in the signal. By performing fast Fourier transform processing on the time scale feature set, the frequency distribution at each time scale is obtained to form a frequency domain feature set. The frequency domain feature set can help identify periodic behaviors or abnormal periodic patterns in network traffic. Discrete wavelet transform is performed on the preprocessed network traffic data to extract multi-resolution features. Discrete wavelet transform is a multi-resolution analysis tool that captures local features and transient changes in signals by decomposing them into sub-signals of different frequency bands. Discrete wavelet transform decomposes signals through scale functions and wavelet functions to obtain a series of decomposition coefficients, which represent the characteristics of signals at different scales and positions. By analyzing these coefficients, a wavelet feature set is obtained, which reflects the multi-scale variation characteristics of network traffic. A network topology graph is constructed based on the preprocessed network traffic data, and the graph structural features are calculated. The construction of the network topology graph is based on the source and destination information of network traffic. The devices in the network are regarded as nodes in the graph, and the traffic is regarded as the edges connecting these nodes. The graph feature set is obtained by analyzing the structural features of the network topology graph, such as the degree of nodes, the weight of edges, and the connectivity of the network. The time scale feature set, frequency domain feature set, wavelet feature set, and graph feature set are concatenated to obtain the initial feature vector. A multi-layer autoencoder network structure is designed to process the initial feature vector. An autoencoder is an unsupervised neural network that maps input data to a low-dimensional space by learning an encoding function, and then restores it to the original data through a decoding function. The goal of the autoencoder is to minimize the reconstruction error between the input and output, thereby extracting the main features of the data. Let the input feature vector be x, and the encoding part of the autoencoder maps it to a low-dimensional latent space representation z through a series of nonlinear transformations, as follows:
[0053] z=f θ (x) = σ(Wx + b);
[0054] Among them, f θis the encoding function, W and b are the weight matrix and bias vector respectively, and σ is a nonlinear activation function (such as ReLU). The decoding part restores z to the reconstruction of the original space through the inverse transformation
[0055]
[0056] During training, minimize the reconstruction error:
[0057]
[0058] The network parameters θ and θ′ are continuously adjusted through the back propagation algorithm to minimize the reconstruction error. After the training is completed, the encoding part f θ This is the required encoder. The encoder performs a nonlinear transformation on the initial feature vector to obtain a fused feature vector. The fused feature vector contains streamlined features extracted from time, frequency, space and multi-scale information, which can more effectively represent the complex pattern of network traffic.
[0059] In a specific embodiment, the process of executing step 300 may specifically include the following steps:
[0060] Segment the fused feature vector into switch-level training data and central training data according to a preset ratio;
[0061] Initialize the adaptive isolation forest model according to the switch-level training data, set the initial parameters and the number of trees, and obtain the initial switch-level model;
[0062] Based on the federated learning framework, local data is used on each network switch participating in the training to update the initial switch-level model and obtain a local updated model.
[0063] Encrypt the parameters of the local update model, and use a secure aggregation algorithm to aggregate the model updates of each network switch on the central server to obtain a global switch-level model;
[0064] Distribute the global switch-level model to each network switch participating in the training, and conduct the next round of local training until the model converges to obtain the final switch-level model;
[0065] Construct an LSTM network structure based on the central training data to obtain an initial LSTM model, where the initial LSTM model includes an input layer, a multi-layer bidirectional LSTM layer, and an output layer;
[0066] Add an attention layer after the LSTM layer of the initial LSTM model, calculate the importance weights of different time steps, and obtain an attention-enhanced LSTM model;
[0067] The attention-enhanced LSTM model is trained using the back-propagation algorithm to minimize the prediction error and obtain the optimized LSTM model.
[0068] Combine the optimized LSTM model with the fully connected layer to build a hybrid model structure and obtain the initial central model;
[0069] The initial central model is fine-tuned and validated using the central training data to obtain the final central model.
[0070] Specifically, the fused feature vector is split into switch-level training data and central training data according to a preset ratio. A typical split ratio may be 80% for switch-level training and 20% for central training. The split method ensures that there is enough data at the switch level for local model training, while sufficient verification data is available on the central server. Use the switch-level training data to initialize the adaptive isolation forest model. Isolation forest is an unsupervised learning algorithm based on decision trees, suitable for anomaly detection. By randomly selecting features and randomly selecting split values to build trees, anomalies can be effectively isolated. When initializing the model, some initial parameters are set, such as the number of trees and the number of samples per tree. These parameters determine the complexity of the model and the depth of training. The formula is as follows:
[0071]
[0072] Where h(x) represents the path length of data point x in the tree, N is the number of samples, and c(N) is a normalization constant. Points with shorter path lengths are more likely to be outliers. After the model is initialized, the training phase based on the federated learning framework begins. Federated learning is a distributed machine learning method that allows all participants to jointly train a global model without sharing local data. On each network switch participating in the training, the initial switch-level model is updated using local data to generate a local update model, ensuring that the model can fully utilize the characteristics of local data and improve the local adaptability of the model. In order to protect the data privacy of each network switch, the parameters of the local update model are encrypted, and then these parameters are transmitted to the central server using a secure aggregation algorithm. On the central server, the encrypted parameters are aggregated to obtain a global switch-level model. Secure aggregation algorithms such as homomorphic encryption or secure multi-party computing can ensure that data is not leaked during the aggregation process. The aggregated global switch-level model is distributed back to each network switch participating in the training for the next round of local training. This process is an iterative process that continues until the performance indicators of the model (such as convergence or accuracy) meet the predetermined standards, resulting in a stable and efficient switch-level model. At the same time, the central training data is used to construct the LSTM network structure to obtain the initial LSTM model. LSTM (Long Short-Term Memory) is a recursive neural network that can capture long-term dependencies in time series. The initial LSTM model usually includes an input layer, a multi-layer bidirectional LSTM layer, and an output layer. The bidirectional LSTM layer can handle the forward and backward dependencies of the time series and improve the understanding of the sequence data. An attention layer is added to the LSTM layer to calculate the importance weights of different time steps. The attention mechanism enables the model to focus on the most critical part of the sequence by assigning different weights to different time steps. The output of the attention layer is combined with the output of the LSTM layer to form an attention-enhanced LSTM model. The formula is as follows:
[0073]
[0074] Among them, e t is the attention score at time step t, calculated by a feedforward neural network. The attention-enhanced LSTM model is trained using the backpropagation algorithm to minimize the prediction error. During the training process, the model parameters are continuously adjusted to obtain an optimized LSTM model. The optimized LSTM model is combined with the fully connected layer to form a hybrid model structure, namely the initial central model. The function of the fully connected layer is to map the output of the LSTM layer to the prediction result space. The initial central model is fine-tuned and verified using the central training data to ensure that the model performs well on different types of data and improve its generalization ability. During the fine-tuning process, the model parameters are further optimized to finally obtain the central model.
[0075] In a specific embodiment, the process of executing step 400 may specifically include the following steps:
[0076] Preprocessing the real-time network flow data of the network switch to obtain real-time preprocessed data;
[0077] Calculate the anomaly score of the real-time preprocessed data based on the switch-level model to obtain the original anomaly score;
[0078] Compare the original anomaly score with the preset threshold to determine whether an anomaly alarm is triggered and obtain a preliminary anomaly mark;
[0079] Quickly classify the preliminary anomaly marks, classify the anomalies into different types according to the predefined rule set, and obtain preliminary anomaly types;
[0080] Combine the preliminary anomaly mark, the preliminary anomaly type and the corresponding context information to obtain a preliminary anomaly detection result;
[0081] Extract and convert the real-time preprocessed data to make it conform to the input format of the central model, and obtain the central model input data;
[0082] Input the central model input data into the central model, perform time series analysis through the LSTM network, and obtain time series features;
[0083] Perform attention mechanism analysis on time series features and calculate the importance weights of different time steps to obtain weighted time series features;
[0084] Anomaly scoring and classification are performed based on weighted time series features to obtain deep anomaly detection results.
[0085] Specifically, raw traffic data is collected from network switches. This data may include information such as the size of the data packet, source and destination addresses, protocol type, and timestamp. The preprocessing steps usually include data cleaning, format conversion, and data standardization. The purpose of data cleaning is to remove noise and incomplete records, such as lost packets or malformed packets. Format conversion is to unify data in different formats into a processable structured format. Data standardization is to adjust the feature values to a uniform scale, such as scaling all numerical features to between 0 and 1 to eliminate the dimensional differences between different features. Use the switch-level model to calculate the anomaly score of the real-time preprocessed data. Use the previously trained model to evaluate the anomaly of each data point. Models such as adaptive isolation forests can determine the degree of anomaly by calculating the path length of the data point. The shorter the path length, the more likely the data point is to be anomaly. The calculated anomaly score is called the raw anomaly score. The anomaly score formula is as follows:
[0086]
[0087] Where h(x) represents the average path length of data point x in the forest, and c(N) is a normalization constant used to normalize the score. The original anomaly score and the preset threshold are compared to determine whether to trigger an anomaly alarm. If the anomaly score exceeds the preset threshold, an anomaly alarm is triggered, the data point is marked as anomaly, and a preliminary anomaly mark is obtained. The preliminary anomaly mark is a preliminary judgment on the abnormality of the data point. This mark can be binary (abnormal or normal) or multi-valued, indicating different levels of anomaly. The preliminary anomaly mark is quickly classified. According to the predefined rule set, the anomaly is divided into different types. For example, the rule set can classify the anomaly into traffic anomaly, protocol anomaly, source address anomaly, etc. according to the characteristics and behavior of the traffic. This classification helps to further analyze and process the abnormal event and obtain the preliminary anomaly type. The preliminary anomaly type is combined with the anomaly mark and context information (such as timestamp, source and destination address, etc.) to form a preliminary anomaly detection result. Feature extraction and transformation are performed on the real-time preprocessed data to make it conform to the input format of the central model. This may include extracting time series features, converting data formats, etc. After transformation, the central model input data is obtained. These data are fed into a central model, usually a LSTM (Long Short-Term Memory) based model for time series analysis. LSTM networks are able to capture the time dependency and order information in the data and are suitable for processing time series data. Through the LSTM network, the time series features of each data point are obtained, which can reflect the time dynamics of the data. The attention mechanism is applied to the LSTM network to analyze the time series features. The attention mechanism calculates the importance weights of different time steps, allowing the model to focus on important time points. By weighting these time series features, weighted time series features are obtained. Anomaly scoring and classification are performed based on the weighted time series features. By comprehensively considering the features and attention weights of different time steps, the abnormality of the data points can be more accurately evaluated to obtain deep anomaly detection results. The deep anomaly detection results not only include the degree of anomaly, but can also be subdivided into specific anomaly types, such as traffic surge, packet loss, etc.
[0088] In a specific embodiment, the execution step calculates an anomaly score for the real-time pre-processed data based on the switch-level model, and the process of obtaining the original anomaly score may specifically include the following steps:
[0089] Perform feature normalization on the real-time preprocessed data to obtain a standardized feature vector, construct a sample space based on the standardized feature vector, and perform random subspace division on the sample space to obtain multiple subspaces;
[0090] Randomly select a feature subset for each subspace, construct an isolation tree structure, obtain an isolation forest, and calculate the path length of each sample based on the isolation forest to obtain a set of sample path lengths;
[0091] Perform statistical analysis on the sample path length set, calculate the average path length and standard deviation, and obtain the path length statistical characteristics;
[0092] An adaptive threshold function is designed based on the statistical characteristics of the path length, and the sample path length is normalized to obtain the initial anomaly score;
[0093] The initial anomaly score is exponentially smoothed to obtain a smoothed anomaly score, and a dynamic baseline model is constructed based on the smoothed anomaly score to calculate the anomaly distribution of historical data and obtain the anomaly distribution parameters;
[0094] Perform Bayesian inference on the smoothed anomaly score and anomaly distribution parameters, calculate the posterior probability distribution, and obtain the corrected anomaly score;
[0095] A multi-level early warning mechanism is designed based on the corrected anomaly score, and the degree of anomaly is quantitatively evaluated to obtain the original anomaly score.
[0096] Specifically, relevant features are extracted from network traffic data. These features may include packet size, transmission rate, time interval, protocol type, etc. The real-time preprocessed data is subjected to feature standardization, and the values of different features are adjusted to the same scale. Usually, each feature is converted to a zero mean and unit variance form to obtain a standardized feature vector. Based on the standardized feature vector, a sample space is constructed, which contains all standardized data points. In order to improve the robustness of the model and prevent overfitting, the sample space is randomly subspaced. Feature subsets are randomly selected, and multiple isolation tree structures are constructed based on these subsets to form an isolation forest. Each tree in the isolation forest is constructed by randomly selecting features and split points, and its purpose is to isolate data points. The earlier a data point is isolated, the more likely it is to be an anomaly. In the isolation forest, the path length of each sample is calculated. The length refers to the depth of the path from the root node to the leaf node in the tree. The calculation formula for the path length is:
[0097]
[0098] Among them, E(h(x)) is the expected value of the path length of sample x, n is the number of samples, and c(n) is a constant used to standardize the path length. The shorter the path length, the more likely the sample is an outlier. The set of sample path lengths is obtained by isolation forest calculation. The sample path length set is statistically analyzed to calculate the average path length and standard deviation of each sample. Statistical features are used to evaluate the performance of samples in the isolation forest. According to the path length statistical characteristics, an adaptive threshold function is designed to normalize the sample path length to generate an initial anomaly score. The initial anomaly score reflects the degree of anomaly of the sample. The higher the score, the more likely the sample is an anomaly. In order to smooth the fluctuations of the anomaly score, the initial anomaly score is exponentially smoothed. Exponential smoothing is a commonly used time series data processing method that can reduce the impact of short-term fluctuations and highlight long-term trends. The formula for the smoothed anomaly score is as follows:
[0099] S t =α×X t +(1-α)×S t-1 ;
[0100] Among them, S t is the current smoothed value, X t is the current anomaly score, and α is the smoothing coefficient. Through smoothing, a smoothed anomaly score is obtained. Based on the smoothed anomaly score, a dynamic baseline model is constructed to better adapt to different network environments. The dynamic baseline model calculates the anomaly distribution of historical data and obtains the anomaly distribution parameters, which reflect the abnormal behavior pattern of the network under normal conditions. Bayesian inference is performed on the smoothed anomaly score and the anomaly distribution parameters to calculate the posterior probability distribution. Bayesian inference is a statistical inference method that combines prior knowledge and observed data to update the probability distribution. The posterior probability distribution formula is:
[0101] P(θ|D)∝P(D|θ)P(θ);
[0102] Among them, P(θ|D) is the posterior probability of parameter θ given data D, P(D|θ) is the likelihood function of the data, and P(θ) is the prior distribution of the parameter. Through Bayesian inference, the corrected anomaly scores are obtained, which more accurately reflect the probability of anomalies. A multi-level early warning mechanism is designed based on the corrected anomaly scores. The multi-level early warning mechanism sets different alarm levels according to the anomaly scores, such as low, medium, and high. Each level corresponds to different response measures. For example, low-level alarms may only require logging, medium-level alarms may require administrator attention, and high-level alarms may require immediate defensive measures.
[0103] In a specific embodiment, the process of executing step 500 may specifically include the following steps:
[0104] Build a time series diagram based on the preliminary anomaly detection results and the deep anomaly detection results, sort the abnormal events in time sequence, and obtain the abnormal event timeline;
[0105] Based on the abnormal event timeline, the time correlation between events is calculated to obtain the time correlation matrix;
[0106] According to the network topology information and the network entities involved in the abnormal events, a spatial relationship diagram is constructed to obtain the spatial distribution diagram of the abnormal events;
[0107] The spatial distribution map of abnormal events is processed by graph convolutional neural network to extract spatial features and propagation patterns, and obtain the spatial correlation matrix;
[0108] Perform tensor fusion of the temporal correlation matrix and the spatial correlation matrix to construct a temporal and spatial correlation tensor and obtain a comprehensive correlation representation;
[0109] Collect multi-source heterogeneous data of network equipment logs and security equipment alarms, and standardize the multi-source heterogeneous data to obtain standardized multi-source data;
[0110] Based on standardized multi-source data and comprehensive association representation, a knowledge graph of abnormal events is constructed to obtain an initial knowledge graph;
[0111] Perform graph embedding analysis on the initial knowledge graph, map nodes and relationships into a low-dimensional vector space, and obtain a graph embedding representation;
[0112] Based on graph embedding representation, abnormal events are grouped and associated with each other to obtain abnormal event clusters.
[0113] Perform feature extraction and semantic analysis on abnormal event clusters to generate a comprehensive abnormality analysis report containing abnormality description, impact scope, severity and possible causes.
[0114] Specifically, the timestamps and related features of abnormal events are collected. A time series graph is a graph of events arranged in chronological order. Each node represents an abnormal event, and the edge represents the chronological relationship between events. By sorting events in time, a timeline of abnormal events is formed to help understand the order in which events occur and possible causal relationships. Based on the timeline of abnormal events, the time correlation between events is calculated. By calculating the time difference between events, a time correlation matrix is formed. The elements in the time correlation matrix represent the time interval between two events. If two events are close in time, the value of the matrix element is larger. The method for calculating time correlation can use the standard correlation coefficient formula:
[0115]
[0116] Among them, r xy is the correlation coefficient between events x and y, x iand i are the time points of the events, and is the average time of the event. A spatial relationship graph is constructed based on the network topology information and the network entities involved in the abnormal events. The spatial relationship graph shows the geographical distribution of abnormal events in the network or the connection relationship between network nodes. Each node represents a network entity, such as a server or switch, and the edge represents the communication connection between these entities. By analyzing the spatial relationship graph, the spatial distribution graph of abnormal events is obtained, and the areas or key nodes where abnormal events are concentrated are identified. The spatial distribution graph of abnormal events is processed using a graph convolutional neural network. The graph convolutional neural network can process graph structured data. By performing convolution operations on the features of neighboring nodes, the local and global features of the nodes are extracted to obtain a spatial correlation matrix, which reveals the spatial dependencies and propagation paths of abnormal events in the network. The temporal correlation matrix and the spatial correlation matrix are tensor-fused to construct a spatiotemporal correlation tensor. The spatiotemporal correlation tensor is a comprehensive representation of temporal and spatial features, which can capture the spatiotemporal features of abnormal events and their interrelationships. Multi-source heterogeneous data such as network device logs and security device alarms are collected. These data come from different systems and devices and may include text logs, numerical data, alarm information, etc. Multi-source heterogeneous data are standardized to obtain standardized multi-source data. Based on standardized multi-source data and comprehensive association representation, a knowledge graph of abnormal events is constructed. A knowledge graph is a structured graph in which nodes represent entities (such as devices, events) and edges represent relationships between entities (such as communications, dependencies). Graph embedding analysis is performed on the initial knowledge graph. Graph embedding maps the nodes and relationships in the graph to a low-dimensional vector space so that the graph structure information can be used by machine learning models. The purpose of graph embedding is to preserve the topological structure and attribute information of the graph while simplifying the computational complexity. Based on the graph embedding representation, abnormal events are grouped and associated to form abnormal event clusters. Abnormal event clusters are collections of events with similar features or correlations. These clusters can help identify common features or patterns of abnormalities. Feature extraction and semantic analysis are performed on abnormal event clusters to generate a comprehensive abnormal analysis report containing abnormal descriptions, impact ranges, severity, and possible causes.
[0117] In a specific embodiment, the process of executing step 600 may specifically include the following steps:
[0118] Extract features from the comprehensive anomaly analysis report to obtain an anomaly feature vector, and based on the anomaly feature vector, select a candidate defense strategy set from a predefined defense strategy library to obtain an initial strategy set;
[0119] Based on the initial strategy set and historical defense effect data, a reinforcement learning environment is constructed, including state space, action space and reward function, to obtain a defense strategy learning environment;
[0120] In the defense strategy learning environment, the deep Q network algorithm is used to train the strategy and obtain the initial Q value function;
[0121] Perform strategy iteration and value iteration on the initial Q value function, continuously update the Q value, and obtain the optimized Q value function;
[0122] Based on the optimized Q-value function, the ε-greedy strategy is used to select the optimal defense action and obtain the candidate defense strategy;
[0123] Conduct feasibility analysis on candidate defense strategies to obtain feasibility scores, and adjust and optimize the candidate defense strategies based on the feasibility scores to obtain adaptive defense strategies;
[0124] The adaptive defense strategy is converted into configuration instructions to obtain a defense execution instruction set, and the defense execution instruction set is sent to the corresponding network switch through the network management protocol.
[0125] Specifically, the various abnormal events described in the analysis report and their contextual information, such as the attack type, target device, affected traffic characteristics, etc. are generated through feature extraction. Based on the abnormal feature vector, the potentially suitable defense strategies are screened out from the predefined defense strategy library to form a candidate defense strategy set. The defense strategy library contains various measures to deal with different types of network threats, such as blocking specific IPs, adjusting firewall rules, and traffic restrictions. A reinforcement learning environment is constructed based on the initial strategy set and historical defense effect data. The reinforcement learning environment requires the definition of state space, action space, and reward function. The state space is represented by an abnormal feature vector, and each state represents a specific abnormal situation of the network. The action space includes all possible defense strategies, that is, the combination of each strategy in the initial strategy set. The reward function is used to evaluate the effect of each defense strategy. For example, a positive reward can be obtained for successfully blocking an attack or reducing losses, while a defense failure or triggering a false alarm may result in a negative reward. The goal of reinforcement learning is to obtain a strategy through training so that the action selected in a given state can maximize the cumulative reward. In the defense strategy learning environment, a deep Q network algorithm is used for strategy training. The deep Q network is an algorithm that combines Q-learning and deep neural networks, which can effectively handle high-dimensional state and action spaces. The deep Q network approximates the Q value function, which is the expected cumulative reward for a given state and action, through a neural network. The training process of the initial Q value function is to update the Q value to reflect the effectiveness of the strategy by going through different combinations of states and actions. The update formula of the Q value function is as follows:
[0126]
[0127] Among them, Q(s,a) is the Q value of selecting action a in state s, α is the learning rate, r is the immediate reward, and γ is the discount factor max a′Q(s′,a′) is the maximum Q value in the next state s′. This formula describes the update rule of Q value. Through continuous iteration, the Q value function gradually converges and finally reflects the actual value of each state-action pair. The initial Q value function is subjected to policy iteration and value iteration, and the Q value is continuously updated until the Q value converges to obtain the optimized Q value function. The optimized Q value function can more accurately guide the selection of defense strategies under specific abnormal conditions. The feasibility and effectiveness of the defense strategy in the actual network environment are evaluated, for example, whether the strategy will have a negative impact on normal traffic, whether it complies with network security policies, etc. The actual feasibility and effect of each strategy are quantified through feasibility scoring. According to these scores, the candidate defense strategies are adjusted and optimized as necessary to obtain the adaptive defense strategy. The adaptive defense strategy not only takes into account the current abnormal situation, but also adapts to the changes in the network environment and the constraints of the strategy implementation. The adaptive defense strategy is converted into specific configuration instructions to form a defense execution instruction set. These instructions include specific operations required to implement the defense strategy, such as updating firewall rules, limiting specific traffic, isolating attacked devices, etc. Through network management protocols (such as SNMP or NetConf), these instructions are sent to the corresponding network switches to actually execute the defense measures.
[0128] The above describes the network anomaly monitoring method of the switch in the embodiment of the present application. The following describes the network anomaly monitoring system 10 of the switch in the embodiment of the present application. Figure 2 In one embodiment of the present application, a network anomaly monitoring system 10 for a switch includes:
[0129] The collection module 11 is used to collect historical network flow data of the network switch and pre-process it to obtain pre-processed network flow data;
[0130] A feature extraction module 12 is used to perform multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector;
[0131] A training module 13, used for training a switch-level model and a central model based on the fused feature vector;
[0132] The detection module 14 is used to perform anomaly detection on the real-time network traffic data of the network switch through the switch-level model to obtain preliminary anomaly detection results, and perform in-depth analysis in combination with the central model to obtain in-depth anomaly detection results;
[0133] The analysis module 15 is used to perform spatiotemporal correlation analysis and multi-source data fusion on the preliminary anomaly detection results and the deep anomaly detection results to obtain a comprehensive anomaly analysis report;
[0134] The execution module 16 is used to generate an adaptive defense strategy based on the comprehensive anomaly analysis report and execute the adaptive defense strategy in the network switch.
[0135] Through the collaboration of the above components, the multi-scale feature extraction of network traffic data, combined with time, frequency domain, wavelet and graph features, comprehensively captures the multi-dimensional features of network anomalies, and improves the accuracy and robustness of anomaly detection. The architecture combining switch-level models and central models is adopted to realize distributed anomaly detection, reduce the risk of single-point failure of the system, and improve detection efficiency and system scalability. The switch-level model is trained using a federated learning framework to avoid direct sharing of raw data and effectively protect user privacy and sensitive information. The introduction of LSTM network and attention mechanism to perform deep time series analysis on network traffic can effectively capture long-term dependencies and key time steps, and improve the ability to identify complex attack patterns. Through dynamic baseline model and Bayesian inference, adaptive calculation of anomaly scores is realized, which improves the flexibility of detection and adaptability to environmental changes. Combining time series analysis and graph convolutional neural network, the spatiotemporal correlation analysis of abnormal events is carried out, which helps to discover potential attack paths and propagation patterns. The knowledge graph of abnormal events is constructed, and graph embedding and semantic analysis are performed to provide network administrators with intuitive and explainable anomaly analysis reports. By using reinforcement learning and deep Q-network, defense strategies are dynamically generated according to the network environment and attack characteristics, which improves the effectiveness and pertinence of defense measures.
[0136] The present application also provides an electronic device, which includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the network anomaly monitoring method for the switch in the above-mentioned embodiments.
[0137] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the network anomaly monitoring method of the switch.
[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0139] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0140] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for monitoring network anomalies of a switch, characterized in that: The method comprises: Collecting historical network traffic data of the network switch and preprocessing it to obtain preprocessed network traffic data; Performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector; wherein the multi-scale feature extraction includes time scale feature extraction, frequency domain feature extraction, wavelet feature extraction and graph structure feature extraction; The switch-level model and the central model are trained based on the fused feature vector; the switch-level model is obtained by training the fused feature vector using an adaptive isolation forest model and a federated learning framework; the central model is obtained by training the fused feature vector using an LSTM network and an attention mechanism; Perform anomaly detection on the real-time preprocessed data after preprocessing the real-time network traffic data of the network switch through the switch-level model to obtain preliminary anomaly detection results; extract and convert features of the real-time preprocessed data to make it conform to the input format of the central model to obtain central model input data; input the central model input data into the central model, perform time series analysis through the LSTM network to obtain time series features; perform attention mechanism analysis on the time series features, and calculate the importance weights of different time steps to obtain weighted time series features; perform anomaly scoring and classification according to the weighted time series features to obtain deep anomaly detection results; Performing spatiotemporal correlation analysis on the preliminary anomaly detection results and the deep anomaly detection results, building an abnormal event knowledge graph based on the comprehensive correlation representation obtained from the spatiotemporal correlation analysis and the standardized multi-source heterogeneous data, performing graph embedding analysis on the knowledge graph, and obtaining a comprehensive anomaly analysis report; An adaptive defense strategy is generated based on the comprehensive anomaly analysis report, and the adaptive defense strategy is executed in the network switch.
2. The network anomaly monitoring method of a switch according to claim 1, characterized in that: The performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector includes: Performing multi-scale sliding window segmentation on the pre-processed network traffic data, and calculating statistical features of different time scales to obtain a time scale feature set; According to the time scale feature set, frequency domain features are extracted using fast Fourier transform to obtain a frequency domain feature set; Performing discrete wavelet transform on the pre-processed network traffic data and extracting multi-resolution features to obtain a wavelet feature set; Constructing a network topology graph based on the preprocessed network traffic data, and calculating graph structure features to obtain a graph feature set; The time scale feature set, the frequency domain feature set, the wavelet feature set and the graph feature set are concatenated to obtain an initial feature vector; A multi-layer autoencoder network structure is designed according to the initial feature vector, the autoencoder model is trained to obtain an encoder part, and a nonlinear transformation is performed on the initial feature vector based on the encoder part to obtain a fused feature vector.
3. The network anomaly monitoring method of a switch according to claim 1, characterized in that: The training of the switch-level model and the central model based on the fused feature vector comprises: Segmenting the fused feature vector into switch-level training data and central training data according to a preset ratio; Initialize the adaptive isolation forest model according to the switch-level training data, set initial parameters and the number of trees, and obtain an initial switch-level model; Based on the federated learning framework, the initial switch-level model is updated using local data on each network switch participating in the training to obtain a local updated model; Encrypting the parameters of the local update model, and then using a secure aggregation algorithm to transmit the encrypted parameters of the local update model to a central server, and aggregating the encrypted parameters of the local update model on the central server to obtain a global switch-level model; Distributing the global switch-level model to each network switch participating in the training, and performing the next round of local training until the model converges to obtain the final switch-level model; Constructing an LSTM network structure according to the central training data to obtain an initial LSTM model, wherein the initial LSTM model includes an input layer, a multi-layer bidirectional LSTM layer and an output layer; Adding an attention layer after the LSTM layer of the initial LSTM model, calculating the importance weights of different time steps, and obtaining an attention-enhanced LSTM model; The attention-enhanced LSTM model is trained using a back-propagation algorithm to minimize the prediction error and obtain an optimized LSTM model; Combining the optimized LSTM model with the fully connected layer to construct a hybrid model structure to obtain an initial central model; The initial central model is fine-tuned and verified using the central training data to obtain a final central model.
4. The network anomaly monitoring method of a switch according to claim 3, characterized in that: The obtaining of the preliminary anomaly detection result and the obtaining of the deep anomaly detection result include: Preprocessing the real-time network traffic data of the network switch to obtain real-time preprocessed data; Calculating an anomaly score for the real-time preprocessed data based on the switch-level model to obtain an original anomaly score; Compare the original anomaly score with a preset threshold to determine whether an anomaly alarm is triggered, and obtain a preliminary anomaly mark; Rapidly classify the preliminary anomaly marks, classify the anomalies into different types according to a predefined rule set, and obtain preliminary anomaly types; Combining the preliminary anomaly mark, the preliminary anomaly type and the corresponding context information to obtain a preliminary anomaly detection result; the context information includes the time when the anomaly occurred, the information of the relevant equipment, and the source and destination of the traffic; Extracting and converting the real-time preprocessed data to conform to the input format of the central model, thereby obtaining central model input data; Inputting the central model input data into the central model, performing time series analysis through an LSTM network, and obtaining time series features; Performing an attention mechanism analysis on the time series features, and calculating the importance weights of different time steps to obtain weighted time series features; Anomaly scoring and classification are performed according to the weighted time series features to obtain a deep anomaly detection result.
5. The method for monitoring network anomalies of a switch according to claim 4, characterized in that: The calculating anomaly score for the real-time pre-processed data based on the switch-level model to obtain an original anomaly score includes: Performing feature standardization processing on the real-time preprocessed data to obtain a standardized feature vector, constructing a sample space according to the standardized feature vector, and performing random subspace division on the sample space to obtain a plurality of subspaces; Randomly select a feature subset for each subspace, construct an isolation tree structure, obtain an isolation forest, and calculate the path length of each sample based on the isolation forest to obtain a sample path length set; Performing statistical analysis on the sample path length set, calculating the average path length and standard deviation, and obtaining path length statistical characteristics; Designing an adaptive threshold function according to the path length statistical characteristics, normalizing the sample path length, and obtaining an initial anomaly score; Performing exponential smoothing on the initial anomaly score to obtain a smoothed anomaly score, building a dynamic baseline model based on the smoothed anomaly score, calculating the anomaly distribution of historical data, and obtaining anomaly distribution parameters; Performing Bayesian inference on the smoothed anomaly score and the anomaly distribution parameter, calculating the posterior probability distribution, and obtaining a corrected anomaly score; A multi-level early warning mechanism is designed according to the corrected anomaly score, and the degree of anomaly is quantitatively evaluated to obtain an original anomaly score.
6. The network anomaly monitoring method of a switch according to claim 1, characterized in that: The comprehensive abnormal analysis report obtained includes: Constructing a time series diagram based on the preliminary anomaly detection result and the deep anomaly detection result, sorting the abnormal events in time sequence, and obtaining a timeline of the abnormal events; Based on the abnormal event timeline, calculate the time correlation between events to obtain a time correlation matrix; According to the network topology information and the network entities involved in the abnormal events, a spatial relationship diagram is constructed to obtain the spatial distribution diagram of the abnormal events; Performing graph convolutional neural network processing on the spatial distribution map of the abnormal events, extracting spatial features and propagation patterns, and obtaining a spatial correlation matrix; Performing tensor fusion on the time correlation matrix and the space correlation matrix to construct a time-space correlation tensor to obtain a comprehensive correlation representation; Collecting multi-source heterogeneous data of network equipment logs and security equipment alarms, and performing standardization processing on the multi-source heterogeneous data to obtain standardized multi-source data; Based on the standardized multi-source data and the comprehensive association representation, construct an abnormal event knowledge graph to obtain an initial knowledge graph; Performing graph embedding analysis on the initial knowledge graph, mapping nodes and relationships into a low-dimensional vector space, and obtaining a graph embedding representation; Based on the graph embedding representation, abnormal events are grouped and associated analyzed to obtain abnormal event clusters; Feature extraction and semantic analysis are performed on the abnormal event cluster to generate a comprehensive abnormality analysis report including abnormality description, impact scope, severity and possible causes.
7. The network anomaly monitoring method of a switch according to claim 1, characterized in that: The step of generating an adaptive defense strategy based on the comprehensive anomaly analysis report and executing the adaptive defense strategy in the network switch includes: Extracting features from the comprehensive anomaly analysis report to obtain an anomaly feature vector, and screening out a candidate defense strategy set from a predefined defense strategy library based on the anomaly feature vector to obtain an initial strategy set; Based on the initial strategy set and historical defense effect data, a reinforcement learning environment including a state space, an action space and a reward function is constructed to obtain a defense strategy learning environment; In the defense strategy learning environment, a deep Q network algorithm is used to perform strategy training to obtain an initial Q value function; Performing strategy iteration and value iteration on the initial Q value function, continuously updating the Q value, and obtaining an optimized Q value function; Based on the optimized Q-value function, an ε-greedy strategy is used to select an optimal defense action to obtain a candidate defense strategy; Performing a feasibility analysis on the candidate defense strategy to obtain a feasibility score, and adjusting and optimizing the candidate defense strategy according to the feasibility score to obtain an adaptive defense strategy; The adaptive defense strategy is converted into a configuration instruction to obtain a defense execution instruction set, and the defense execution instruction set is sent to a corresponding network switch through a network management protocol.
8. A network anomaly monitoring system for a switch, characterized in that: Used to execute the network anomaly monitoring method of a switch according to any one of claims 1 to 7, the network anomaly monitoring system of the switch comprises: A collection module is used to collect historical network flow data of the network switch and pre-process it to obtain pre-processed network flow data; A feature extraction module, used for performing multi-scale feature extraction and feature dimension reduction on the pre-processed network traffic data to obtain a fused feature vector; A training module, used for training a switch-level model and a central model based on the fused feature vector; The detection module is used to perform anomaly detection on the real-time preprocessed data after preprocessing the real-time network traffic data of the network switch through the switch-level model to obtain preliminary anomaly detection results; extract and convert the real-time preprocessed data to conform to the input format of the central model to obtain central model input data; input the central model input data into the central model, perform time series analysis through the LSTM network to obtain time series features; perform attention mechanism analysis on the time series features, and calculate the importance weights of different time steps to obtain weighted time series features; perform anomaly scoring and classification according to the weighted time series features to obtain deep anomaly detection results; An analysis module is used to perform spatiotemporal correlation analysis on the preliminary anomaly detection results and the deep anomaly detection results, and to construct an abnormal event knowledge graph based on the comprehensive correlation representation obtained by the spatiotemporal correlation analysis and the standardized multi-source heterogeneous data, and to perform graph embedding analysis on the knowledge graph to obtain a comprehensive anomaly analysis report; An execution module is used to generate an adaptive defense strategy based on the comprehensive anomaly analysis report and execute the adaptive defense strategy in the network switch.
9. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory so that the electronic device executes the network anomaly monitoring method for a switch according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the network anomaly monitoring method for a switch according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Network traffic intrusion detection method based on big data
CN116647374A
Network intrusion detection method based on federated learning
CN116708009A
Intrusion detection system and method based on intelligent network
CN118413406A
SD-WAN application identification method and system based on deep learning
CN118474043A
Cited By
Real-time detection system for abnormal flow of switch based on deep learning
CN121217612A
Deep learning based switch traffic anomaly real-time detection system
CN121217612B