A Network Security Situation Awareness Method, Computer Device and Storage Medium

By introducing BiLSTM network, improved ResNeXt network and dual attention mechanism into the network security situation awareness model, the dynamic snake convolution and dual attention mechanism are used to solve the problem of low accuracy in network security situation awareness in the existing technology, and more efficient network attack detection and prediction are achieved.

CN118590885BActive Publication Date: 2025-05-30HEBEI NORMAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410846710.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-05-30
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

Existing network security situation awareness methods are difficult to achieve high-precision prediction and detection when facing dynamically changing network data and complex attack patterns.

Method used

The network security situation awareness model based on BiLSTM network, improved ResNeXt network and dual attention mechanism is adopted. Through dynamic serpentine convolution and dual attention mechanism, the core characteristics and context information of network traffic data are captured to improve the prediction accuracy of the model.

Benefits of technology

It significantly improves the accuracy of network security situation awareness, can more effectively capture long-term dependency and suppress gradient problems, and enhances the detection ability of complex network attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118590885B_ABST
    Figure CN118590885B_ABST
Patent Text Reader

Abstract

The present invention discloses a network security situation awareness method, a computer device, and a storage medium, relating to the field of network security technologies. The method includes: obtaining current network traffic data; based on the current network traffic data and a network security situation awareness model, predicting the security situation of the current network and determining the category of the current network; wherein the category of the current network includes normal network communication and various types of network attacks; the network security situation awareness model is pre-trained using a training sample set; the network security situation awareness model includes a BiLSTM network, an improved ResNeXt network, a dual attention network, and a fully connected layer connected in sequence; the improved ResNeXt network is: replacing the ordinary convolution in the residual module of ResNeXt with a dynamic snake-shaped convolution. The present invention improves the accuracy of network security situation awareness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technologies, and in particular, to a network security situational awareness method, a computer device, and a storage medium. Background Art

[0002] Communication and network technologies, especially wireless networks, have developed rapidly in recent years. However, the network also faces increasing security risks, posing a major threat to users' privacy and security. Some common network security threats include data interception, cracking, transmission interference, configuration problems, freeloading, and denial-of-service attacks, etc. Traditionally, an Intrusion Detection System (IDS) has been used to detect attacks and provide security by identifying unauthorized use or abuse of computer systems. However, emerging wireless networks, such as the Internet of Things, due to the lack of standardization, limited resources, device vulnerabilities, and usually being produced by manufacturers who do not pay attention to security, face more security problems and threats than traditional networks, resulting in a large number of increasingly complex attacks, making it impossible for the IDS to detect these attacks in a timely manner. The mobility and flexibility of nodes also pose greater difficulties for incident handlers or network administrators to make appropriate and timely decisions. This contradiction has intensified the rapid development and application of various new network technologies.

[0003] Network Security Situational Awareness (NSSA) aims to monitor, analyze, and predict network security situations in real time and provide timely and effective countermeasures. An intelligent NSSA system can monitor and capture various types of threats, analyze them, and formulate plans to avoid further attacks. The comprehensive design of an intelligent NSSA system will help decision-makers understand the current and upcoming security status of the network. Network security situational awareness is the main component of an intelligent NSSA system, which is an important proactive defense against network risks and threats, involving evaluating the network security status based on the current situation and predicting the future network security status based on historical data, predicting the probability, severity, and impact of potential attacks and risks.

[0004] However, in actual situations, the cycle of network data may change dynamically, such as low-frequency and multi-stage attacks in industrial Internet of Things and hidden behaviors in the network, which makes it impossible for, for example, attack graphs, generalized Bayesian classification neural networks, Markov-related models, evidence reasoning rules, BiGRU neural networks, and other machine learning methods to meet the comprehensive situation prediction requirements, and the accuracy of network security situational awareness is low. Summary of the Invention

[0005] The object of the present invention is to provide a network security situation awareness method, a computer device and a storage medium, which can improve the accuracy of network security situation awareness.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A network security situation awareness method, including:

[0008] Obtain the current network traffic data;

[0009] According to the current network traffic data, based on the network security situation awareness model, predict the security situation of the current network and determine the category of the current network; wherein, the category of the current network includes normal network communication and various types of network attacks; the network security situation awareness model is pre-trained using a training sample set.

[0010] The network security situation awareness model includes a BiLSTM network, an improved ResNeXt network, a dual attention network and a fully connected layer connected in sequence; the improved ResNeXt network is: replacing the ordinary convolution in the residual module of ResNeXt with a dynamic snake-shaped convolution.

[0011] To achieve the above object, the present invention provides the following solutions:

[0012] A computer device, including: a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above network security situation awareness method.

[0013] To achieve the above object, the present invention provides the following solutions:

[0014] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above network security situation awareness method is implemented.

[0015] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The network security situation awareness model used in the present invention is based on a BiLSTM network, an improved ResNeXt network using dynamic snake-shaped convolution and a dual attention mechanism, pays more attention to the core structural features, can capture information in the spatial and channel dimensions, effectively capture long-term dependencies, and suppress the gradient problem, thereby improving the accuracy of network security situation awareness. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a flowchart of the network security situation awareness method provided by the present invention;

[0018] Figure 2 It is a structural schematic diagram of the network security situation awareness model;

[0019] Figure 3 It is a schematic diagram of the construction process of the network security situation awareness model;

[0020] Figure 4 It is a structural schematic diagram of the BiLSTM network;

[0021] Figure 5 It is a coordinate calculation diagram of the dynamic snake-shaped convolution;

[0022] Figure 6 It is a schematic diagram of the spatial attention module;

[0023] Figure 7 It is a schematic diagram of the channel attention module;

[0024] Figure 8 It is a structural schematic diagram of the residual module in the improved ResNeXt network. Specific Embodiments

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0026] The purpose of the present invention is to provide a network security situation awareness method, a computer device, and a storage medium. By using the improved model to extract the core features of network traffic data and paying more attention to context information, the accuracy of network security situation awareness can be improved.

[0027] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Embodiment 1

[0029] AsFigure 1 As shown in the figure, this embodiment provides a network security situation awareness method, including:

[0030] Step 1: Obtain the current network traffic data.

[0031] Step 2: According to the current network traffic data, based on the network security situation awareness model, predict the security situation of the current network and determine the category of the current network. Among them, the categories of the current network include normal network communication and various types of network attacks. The network security situation awareness model is pre-trained using a training sample set.

[0032] As Figure 2 shown in the figure, the network security situation awareness model includes a BiLSTM network, an improved ResNeXt network, a dual attention network, and a fully connected layer connected in sequence. The improved ResNeXt network is: replacing the ordinary convolution in the residual module of ResNeXt with a dynamic snake-shaped convolution.

[0033] The BiLSTM network can effectively process sequence data and capture its important features; the improved ResNeXt network can extract core features through dynamic snake-shaped convolution operations; finally, the features are weighted through a dual attention mechanism to enhance the attention of the network security situation awareness model to important features. Therefore, the network security situation awareness model can improve the accuracy of network security situation awareness.

[0034] Specifically, as Figure 3 shown in the figure, the construction process of the network security situation awareness model includes:

[0035] (1) Obtain a network security data set. The network security data set includes multiple network traffic data and the category label of each network traffic data.

[0036] In this embodiment, the UNSW-NB15 data set is used as the network security data set. The network traffic data in the UNSW-NB15 data set is presented in the form of network packets, and each packet contains multiple feature values. The following are some specific features of the network traffic data in the UNSW-NB15 data set:

[0037] Source IP address and destination IP address: Represent the source and destination IP addresses of the packet.

[0038] Source port number and destination port number: Represent the source and destination port numbers of the packet.

[0039] Protocol: indicates the network protocol used by the data packet, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or Internet Control Message Protocol (ICMP), etc.

[0040] Flow identifier: used to identify the flow to which the data packet belongs, that is, a set of associated data packets.

[0041] Data packet length: indicates the length of the data packet, usually in bytes.

[0042] Timestamp: indicates the capture timestamp of the data packet.

[0043] Flag bits: used to indicate the status or characteristics of the data packet, such as the Synchronize Sequence Numbers (SYN), Acknowledge character (ACK), and Finish (FIN) flags of the TCP connection, etc.

[0044] Characteristic values of the protocols at each layer of the data packet: such as IP header information, TCP header information, or UDP header information, etc.

[0045] The types of the network attacks include denial of service, distributed denial of service, and malware.

[0046] (2) Preprocess the network traffic data in the network security dataset to obtain a training sample set. The preprocessing includes data cleaning, digitization, standardization, and normalization. And divide the training sample set into a training set and a test set.

[0047] (3) Use the training sample set to iteratively train the network security situation awareness model until the maximum number of iterations is reached or the loss function converges. In the present invention, the network security situation awareness model is called the BDSCResNeXt model.

[0048] Specifically, the training sample set is input into the BiLSTM network. The BiLSTM network is used to process network traffic data and extract its features. The data processed by the BiLSTM network is input into the improved ResNeXt network. The improved ResNeXt network is used to perform deep feature extraction and classification on the input data. Then, the outputs of the BiLSTM network and the improved ResNeXt network are concatenated together to form a comprehensive feature vector. The comprehensive feature vector is weighted by the dual attention network to enhance the network security situation awareness model's attention to important features. Finally, the processed features are unfolded into a one-dimensional tensor and used as the output of the network security situation awareness model.

[0049] In the forward propagation, after the input data is processed through a series of layers, if the include_top parameter is True, the final classification is performed through global average pooling and a fully connected layer.

[0050] The output of the network security situation awareness model is compared with the class labels. According to the number of correctly predicted samples, the loss function (cross-entropy loss is selected in this embodiment) is backpropagated, the gradients are calculated, and the model's parameters are updated through an optimizer (AdaBelief is selected in this embodiment). The training step count and the total training step count are updated. At the end of each epoch, the accuracy of the current epoch is printed.

[0051] Specifically, first, the dataset is loaded, and training and test data loaders are created. Then, the network security situation awareness model is defined and moved to the appropriate device (GPU or CPU). Next, the loss function and the optimizer are set, and the number of training epochs is specified. In each training epoch, the training data loader is iterated, the loss and accuracy of the model are calculated, and the model's parameters are updated through backpropagation. In each test epoch, the test data loader is iterated, and the accuracy of the model on the test set is calculated. Finally, the training and test accuracies are output.

[0052] Furthermore, step 2 specifically includes:

[0053] Step 21: Perform preliminary feature extraction on the current network traffic data through the BiLSTM network to obtain preliminary feature data.

[0054] The BiLSTM network processes the sequence data in the UNSW-NB15 dataset. Network traffic data is usually generated in chronological order, so each network connection or packet can be regarded as a sequence. The BiLSTM network can perform temporal modeling on network traffic data to capture the relationships and context information between connections.

[0055] As Figure 4As shown, the BiLSTM network is used to process the input data and extract its features. The inputs of the BiLSTM are features, the initial hidden state, and the memory cell. After internal calculations, new memory cells and new hidden states are obtained. The hidden state can either be connected to the softmax output or continue to be calculated with the features of the next round. The BiLSTM network includes a forget gate, an input gate, and an output gate. The forget gate, input gate, and output gate all generate a weight, which is multiplied by an output to control the size of the output. That is, the forget gate determines the forgetting value of the memory cell of this main line according to the input features and the hidden layer, the input gate determines the replenishment value of the memory cell of this main line according to the input features and the hidden layer, and the output gate determines the value taken from the memory cell as the output of the new hidden state according to the input features and the hidden layer. The specific formulas are as follows:

[0056] Forget gate part:

[0057] Forget gate: f t = σ(U f h t-1 + W f x t );

[0058] Main line forgetting: k t = c t-1 ⊙ f t ;

[0059] Input gate part:

[0060] Input gate: i t = σ(U i h t-1 + W i x t );

[0061] Replenishment source: g t = tanh(U g h t-1 + W g x t );

[0062] Gate control replenishment size: j t = g t ⊙ i t ;

[0063] Main line replenishment: c t = j t + k t ;

[0064] Output gate part:

[0065] Output gate: o t = σ(U o h t-1 + Wo x t );

[0066] Main line generates output: h t = tanh(c t ) ⊙ o t ;

[0067] Where f t is the value of the forget gate at time t, i t is the value of the input gate at time t, o t is the value of the output gate at time t, x t is the input at time t, h t-1 is the hidden layer state value at time t-1, h t is the hidden layer state value at time t, g t is the initial feature extraction of h t-1 and x t at time t, k t is the calculation result of the forget gate at time t, j t is the calculation result of the input gate at time t, c t is the cell state at time t, c t-1 is the cell state at time t-1, W f is the weight coefficient of h t-1 in the forget gate, W i is the weight coefficient of h t-1 in the input gate, W o is the weight coefficient of h t-1 in the output gate, W g is the weight coefficient of h t-1 in the feature extraction process, U f is the weight coefficient of x t in the forget gate, U i is the weight coefficient of x t in the input gate, U o is the weight coefficient of x t in the output gate, U g is the weight coefficient of x t in the feature extraction process, tanh represents the hyperbolic tangent function, σ represents the activation function Sigmoid, and ⊙ is the Hadamard product.

[0068] The network parameters of the BiLSTM network include: the dimension of the input features, i.e., the length of the feature vector input at each time step; the dimension of the hidden layer, i.e., the output dimension of the LSTM cell; the length of the feature vector input in the LSTM network; the dimension of the hidden layer in the LSTM network; the number of layers in the LSTM network; a flag indicating whether the LSTM network is bidirectional, which is set to True here, indicating the construction of a bidirectional LSTM; the dimension of the input features of the fully connected layer, i.e., the dimension of the hidden state output by the LSTM multiplied by 2 (because it is bidirectional); the dimension of the output features of the fully connected layer, i.e., the number of classes of the classifier.

[0069] Step 22: Perform deep feature extraction on the preliminary feature data through the improved ResNeXt network to obtain deep feature data.

[0070] The improved ResNeXt network processes other features in the UNSW-NB15 dataset, such as numerical features. It can be used to extract feature representations of non-sequential data. It can capture the relationships between features through convolutional operations and model and extract numerical features.

[0071] The improved ResNeXt network specifically includes operations such as dynamic snake-shaped convolution, batch normalization, activation function, max pooling, etc., and constructs residual modules of different layers by calling the _make_layer method.

[0072] The dynamic snake-shaped convolution adds continuity constraints to the design of the convolutional kernel; each convolutional position takes its previous position as a reference and freely selects the swing direction, thus ensuring the continuity of perception while making free choices. On the one hand, the dynamic snake-shaped convolution can freely fit the structure to learn features, and on the other hand, it can not deviate too far from the target structure under the constraint conditions, thus paying more attention to the core structural features.

[0073] For a standard 3×3 two-dimensional convolutional kernel K, it is expressed as:

[0074] K = {(x - 1, y - 1), (x - 1, y),..., (x + 1, y + 1)}.

[0075] The center coordinates of the two-dimensional convolutional kernel K are K m = (x m , y m ).

[0076] To enable the convolutional kernel to more flexibly focus on the complex geometric features of the target, a deformation offset Δ is introduced, and an iterative strategy is adopted, as Figure 5 shown, successively select the next observation position of each target to be processed, ensure the continuity of attention, and prevent the perception field from spreading too far due to large deformation offsets.

[0077] The dynamic snake convolution kernel specifically includes: all biases that control the deformation of a single convolution kernel are learned in the network at one time, and there is only one range constraint for this bias, namely the receptive field range; controlling the deformation of all convolutions depends on the final loss constraint feedback of the entire network, and the change process is free; adding continuity constraints to the design of the convolution kernel, each convolution position is based on its previous position, and the swing direction can be freely selected, thereby ensuring the continuity of the perception while freely selecting.

[0078] In dynamic snake convolution, each grid position represents the sampling position of the convolution kernel on the input feature map. By accumulating these positions, a linear morphological structure can be formed, which helps to extract directional features.

[0079] The difference between dynamic snake convolution and ordinary convolution is as follows:

[0080] Convolution kernel shape: Ordinary convolution uses a fixed-shape convolution kernel, usually a rectangle or square. The size and shape of the convolution kernel remain unchanged throughout the convolution process. Dynamic snake convolution forms a linear morphological structure by accumulating offset parameters, which can realize convolution kernels of different shapes and directions. The convolution kernel of dynamic snake convolution can adaptively change shape according to task requirements, which has greater flexibility.

[0081] Sampling method: Ordinary convolution uses a fixed sampling method at each position, usually sampling the surrounding pixels symmetrically around the central pixel. This sampling method is suitable for extracting local features. Dynamic snake convolution accumulates offset parameters and samples along a straight line or curve at each position, which can extract features with more directional and shape perception. Dynamic snake convolution can better capture long-distance contextual information and global features.

[0082] Number of parameters: The number of parameters of ordinary convolution is fixed, which is related to the size of the convolution kernel and the number of channels of the input feature map. The number of parameters of dynamic snake convolution is variable, depending on the shape of the convolution kernel and the sampling method. Since the shape of the convolution kernel of dynamic snake convolution is dynamically adjusted, the number of parameters can vary depending on the input data.

[0083] In dynamic snake convolution, the standard convolution kernel is linearized along the x-axis and y-axis. Taking the x-axis as an example, the specific position of each grid in the convolution kernel is expressed as: K m±c =(x m±c ,y m±c), where \(c = \{0, 1, 2, 3, 4\}\) represents the horizontal distance from the central grid. Each grid position represents the sampling position of the convolutional kernel on the input feature map. By accumulating these positions, a linear morphological structure can be formed, which helps to extract directional features. Each grid position \(K\) in the two-dimensional convolutional kernel \(K\) m±c is selected through an accumulation process. Starting from the central position \(K\) m , the positions far from the central grid depend on the position of the previous grid: compared with \(K\) m , \(K\) m+1 increases by an offset \(\Delta=\{\delta|\delta\in[-1, 1]\}\). Therefore, the offsets need to be accumulated to ensure that the convolutional kernel conforms to the linear morphological structure.

[0084] Figure 5 The change in the x-axis direction in \(\) is represented as follows:

[0085]

[0086] The change in the y-axis direction is represented as follows:

[0087]

[0088] where \(\Delta y\) is the y-axis offset and \(\Delta x\) is the x-axis offset.

[0089] Since the offset \(\Delta\) is usually a fraction, bilinear interpolation is implemented as follows:

[0090] \(K\) 1 =\sum K 'B(K',K 1 )\cdot K';

[0091] where \(K\) 1 represents the fractional position, \(K'\) represents enumerating all integer space positions, \(B(\cdot)\) is the bilinear interpolation kernel, which is divided into two one-dimensional kernels, \(B(K 1 ,K') = b(K 1x ,K' x )\cdot b(K 1y ,K' y ), \(b(\cdot)\) represents the one-dimensional interpolation kernel function for calculating the weights in a single dimension. In this formula, \(b(K 1x ,K' x ) and \(b(K 1y ,K' y ) represent the one-dimensional interpolation kernel functions on the x-axis and y-axis respectively, and \(K 1x , K' x , K 1y , K' y represent the position relationships of the fractional positions relative to the nearest neighbor data points on the corresponding axes.

[0092] Bilinear interpolation is used to convert fractional positions to positive positions. When it is necessary to locate fractional positions on a discrete pixel grid, bilinear interpolation can estimate the pixel values of fractional positions by performing a weighted average of the nearest integer positions. This interpolation method can make the pixel values of fractional positions more continuous and smooth, thereby achieving more accurate position localization and transformation effects.

[0093] Step 23: Concatenate the preliminary feature data and the depth feature data to obtain a comprehensive feature vector.

[0094] Fuse the outputs of the BiLSTM network and the improved ResNeXt network to comprehensively utilize different types of feature information they have learned. By concatenating their feature vectors in dimension 1, sequence and non-sequence features can be combined into a comprehensive feature vector. The fusion strategy can comprehensively consider the temporal relationship in the data and the relationship between other features.

[0095] Step 24: Respectively learn the spatial dependence relationship and the channel dependence relationship of the comprehensive feature vector through a dual attention network to obtain spatial dependence features and channel dependence features.

[0096] Specifically, the dual attention network includes a spatial attention module (PositionAttention Module, PAM) and a channel attention module (ChannelAttentionModule, CAM). The dual attention network is divided into two dimensions and can collect rich context information.

[0097] Learn the spatial dependence relationship of the comprehensive feature vector through the spatial attention module to obtain spatial dependence features.

[0098] As Figure 6 shown, first, the comprehensive feature vector A (C×H×W) is respectively subjected to convolution operations through 3 convolutional layers to obtain a first feature vector B, a second feature vector C, and a third feature vector D. Then, the first feature vector B, the second feature vector C, and the third feature vector D are respectively reshaped into two-dimensional features (C×N), where N = H×W, to obtain a first two-dimensional feature, a second two-dimensional feature, and a third two-dimensional feature. Then, after transposing the first two-dimensional feature (N×C) and multiplying it by the second two-dimensional feature (C×N), and passing through the Softmax function, a two-dimensional spatial attention map S (N×N) is obtained. Then, perform a matrix multiplication operation on the third two-dimensional feature (C×N) and the transpose of the two-dimensional spatial attention map S (N×N), multiply by the spatial scale coefficient α, and reshape it into a three-dimensional feature (C×H×W) to obtain a three-dimensional spatial attention map. Finally, add the three-dimensional spatial attention map to the comprehensive feature vector A to obtain spatial dependence features E1 Among them, C is the number of channels of the comprehensive feature vector, H is the height of the comprehensive feature vector, and W is the width of the comprehensive feature vector. α is initialized to 0 and gradually learns to obtain a larger weight.

[0099]

[0100] Among them, is the feature at the j-th position in the spatial dependence feature, that is, the weighted sum of the features at all positions and on the original features, S ji is the influence of the i-th position on the j-th position in the two-dimensional spatial attention map, N is the total number of positions, A j is the feature at the j-th position in the comprehensive feature vector, D i is the feature at the j-th position in the third feature vector.

[0101] It can be seen from the above formula that the value of each point in the finally output spatial dependence feature E is the result obtained by calculating the weighted sum of each position of the original feature (comprehensive feature vector).

[0102] Learn the channel dependence relationship of the comprehensive feature vector through the channel attention module to obtain the channel dependence feature.

[0103] As Figure 7 shown, first reshape the comprehensive feature vector A into a two-dimensional feature to obtain the fourth two-dimensional feature (C×N). Then transpose the fourth two-dimensional feature set (N×C) and multiply it by the fourth two-dimensional feature, and then pass through the Softmax function to obtain the two-dimensional channel attention map X (C×C). Then, after transposing the two-dimensional channel attention map, perform matrix multiplication with the fourth two-dimensional feature (C×N), multiply by the channel scale coefficient β, and reshape it into a three-dimensional feature (C×H×W) to obtain the three-dimensional channel attention map. Finally, add the three-dimensional channel attention map to the comprehensive feature vector to obtain the channel dependence feature E 2 Among them, β is initialized to 0 and gradually learns to obtain a larger weight.

[0104]

[0105] Among them, is the feature at the j-th position in the channel dependence feature, A i is the feature at the i-th position in the comprehensive feature vector, A j is the feature at the j-th position in the comprehensive feature vector.

[0106] Step 25: Perform feature fusion processing on the spatial dependence feature and the channel dependence feature to obtain a dual attention feature.

[0107] Specifically, the spatial dependence features and channel dependence features are first summed element-wise to complete feature fusion, and then a convolution is performed once to generate the final prediction map, that is, the dual attention feature.

[0108] Step 26: Unfold the dual attention feature into a one-dimensional tensor and then classify it through a fully connected layer to determine the category of the current network.

[0109] The network security situation awareness model used in the present invention is based on the advantages of the BiLSTM network, the improved ResNeXt network, and the dual attention mechanism. It pays more attention to the core structural features, can capture information in the spatial and channel dimensions, effectively capture long-term dependencies, and suppress the gradient problem, thereby improving the accuracy of network security situation awareness.

[0110] To verify the network security situation awareness method in Embodiment 1, the CICIDS2017 dataset is used for training and testing below. The specific process is as follows:

[0111] S1: Obtain the publicly available network security dataset on the network and perform data preprocessing on the dataset. In the present invention, the CICIDS2017 (Canadian Institute for Cybersecurity Intrusion Detection Systems 2017) network intrusion detection dataset is used to evaluate and study the performance of network intrusion detection systems. This dataset contains a large amount of simulated network traffic data for simulating various network attacks and normal network communications. The CICIDS2017 dataset contains various types of network traffic, including normal network communications and various types of network attacks, such as DoS (Denial of Service), DDoS (Distributed Denial of Service), malware, etc. Each data sample in the dataset has a corresponding label indicating whether the data stream belongs to normal communication or different types of network attacks, which helps with supervised learning and evaluating model performance.

[0112] ① Data cleaning: Handling missing values: Detect and handle the missing values in the dataset, and fill or delete them; Handling duplicate values: Detect and handle the duplicate samples in the dataset; Handling outliers: Detect and handle the outliers in the dataset, identify outliers using statistical methods or rule-based methods, and perform appropriate processing.

[0113] ② Digitization: Convert the non-numerical data (such as categorical data) in the dataset into numerical data for subsequent modeling and analysis. The one-hot encoding method is adopted.

[0114] ③ Standardization: Perform standardization processing on the numerical features so that they follow the standard normal distribution, which helps to improve the convergence speed and stability of model training.

[0115] ④ Normalization: Scale the numerical features to the same range to avoid the influence of the dimension between different features on the modeling result. The min-max normalization method is used to standardize the feature values between 0 and 1.

[0116] S2: Build the BDSCResNeXt model, input the preprocessed data into the model for training and testing, which specifically includes:

[0117] ① First, input it into the BiLSTM network. The shape of the input data is [batch_size, max_len, n_class], where batch_size represents the number of samples in each batch, max_len represents the sequence length of each sample, and n_class represents the feature dimension of each time step. In the forward calculation process, the input data passes through multiple layers of bidirectional LSTM (BiLSTM) networks. First, the dimension of the input data is transposed through the transpose method to make it conform to the input format of the LSTM network. Then, the sequence data is processed through the LSTM layer, including the initialization of the hidden state, the operation of the LSTM unit, and the output of the hidden state. Finally, the output is obtained through the fully connected layer, and the hidden state is mapped to the target class space. Finally, the BiLSTM network returns the result of the forward calculation, and the dimension of the output result is [batch_size, n_class], representing the class probability distribution corresponding to each sample.

[0118] ② Input the output features obtained above into the improved ResNeXt network. The input data first performs feature extraction through a dynamic snake-shaped convolution, followed by a batch normalization layer and a ReLU activation function, then a max pooling operation, and then input into the first Layer layer, the second Layer layer, and the third Layer layer in sequence. The first Layer layer, the second Layer layer, and the third Layer layer each contain a residual module, and the first Layer layer, the second Layer layer, and the third Layer layer are respectively used to extract features at different levels. Then, global pooling of the features is performed through the average pooling operation, and the features are flattened into a vector through the flatten layer, and finally classification is performed through the fully connected layer.

[0119] As Figure 8 shown, in each residual module, the result output by the last layer is added to the initial input, and then ReLu is used to obtain the final output. Each residual module contains three consecutive snake-shaped dynamic convolution layers, and each snake-shaped dynamic convolution layer is followed by a batch normalization layer and a ReLU activation function. The input of the residual module passes through an optional downsampling layer to match the dimension of the main path. Finally, the output of the residual module is added to the input, and the final output is generated through another ReLU activation function.

[0120] ③ Concatenate the outputs of the BiLSTM network and the improved ResNeXt network to form a comprehensive feature vector. Subsequently, perform weighted processing on the comprehensive feature vector through a dual attention network to enhance the model's attention to important features. Finally, expand the processed features into a one-dimensional tensor and use it as the output of the model.

[0121] S3: Conduct training and testing. Load the dataset and create training and testing data loaders. Then define the model and move it to the appropriate device. Next, set the loss function and optimizer and specify the number of training epochs. In each training epoch, iterate through the training data loader, calculate the loss and accuracy of the model, and update the model parameters through backpropagation. In each testing epoch, iterate through the testing data loader and calculate the accuracy of the model on the test set. Finally, output the training and testing accuracies.

[0122] Example 2

[0123] A computer device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the network security situation awareness method in Example 1.

[0124] Example 3

[0125] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the network security situation awareness method in Example 1.

[0126] Example 4

[0127] A computer program product, comprising a computer program, and when the computer program is executed by a processor, it implements the network security situation awareness method in Example 1.

[0128] Example 5

[0129] A computer device, which can be a database. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store transactions to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the network security situation awareness method in Embodiment 1.

[0130] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the object or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0131] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present invention can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present invention can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0132] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0133] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A network security situation awareness method, characterized in that: The network security situation awareness method comprises: Get current network traffic data; According to the current network traffic data, based on the network security situation awareness model, the security situation of the current network is predicted to determine the category of the current network; wherein the category of the current network includes normal network communication and various types of network attacks; the network security situation awareness model is pre-trained using a training sample set; The network security situation awareness model includes a BiLSTM network, an improved ResNeXt network, a dual attention network and a fully connected layer connected in sequence; the improved ResNeXt network is: a dynamic snake convolution is used to replace the ordinary convolution of the residual module in ResNeXt; In the dynamic snake convolution, the standard convolution kernel is linearized along the x-axis and y-axis. For the x-axis direction, the specific position of each grid in the convolution kernel is expressed as: K m±c =(x m±c ,y m±c ), where c = {0, 1, 2, 3, 4} represents the horizontal distance from the center grid; each grid position represents the sampling position of the convolution kernel on the input feature map, and by accumulating these positions, a linear morphological structure is formed to extract directional features; the center coordinate of the two-dimensional convolution kernel K is K m =(x m ,y m ); Each grid position K in the two-dimensional convolution kernel K m±c The selection of is a cumulative process; from the central position K m Initially, the position of the grid away from the center depends on the position of the previous grid: m In comparison, K m+1 An offset Δ={δ|δ∈[-1,1]} is added; therefore, the offset needs to be accumulated to ensure that the convolution kernel conforms to the linear morphology structure; The change in the x-axis direction is expressed as follows: The change in the y-axis direction is expressed as follows: Among them, Δy is the y-axis offset, and Δx is the x-axis offset.

2. The network security situation awareness method according to claim 1 is characterized in that: The training process of the network security situation awareness model includes: Acquire a network security data set; the network security data set includes a plurality of network flow data and a category label for each network flow data; Preprocessing the network traffic data in the network security data set to obtain a training sample set; the preprocessing includes data cleaning, digitization, standardization and normalization processing; The network security situation awareness model is iteratively trained using the training sample set until a maximum number of iterations is reached or the loss function converges.

3. The network security situation awareness method according to claim 2 is characterized in that: The network traffic data includes source IP address, destination IP address, source port number, destination port number, protocol, flow identifier, data packet length, timestamp, flag bit and characteristic values ​​of each layer protocol of the data packet.

4. The network security situation awareness method according to claim 1 is characterized in that: The types of network attacks include denial of service, distributed denial of service and malware.

5. The network security situation awareness method according to claim 1 is characterized in that: According to the current network traffic data, based on the network security situation awareness model, the security situation of the current network is predicted to determine the category of the current network, specifically including: Performing preliminary feature extraction on the current network traffic data through a BiLSTM network to obtain preliminary feature data; Performing deep feature extraction on the preliminary feature data through an improved ResNeXt network to obtain deep feature data; Concatenating the preliminary feature data and the deep feature data to obtain a comprehensive feature vector; The spatial dependency and channel dependency of the comprehensive feature vector are respectively learned through a dual attention network to obtain spatial dependency features and channel dependency features; Performing feature fusion processing on the space-dependent feature and the channel-dependent feature to obtain a dual attention feature; The dual attention features are expanded into a one-dimensional tensor and classified through a fully connected layer to determine the category of the current network.

6. The network security situation awareness method according to claim 5 is characterized in that: The dual attention network includes a spatial attention module and a channel attention module; The spatial dependency and channel dependency of the comprehensive feature vector are learned through a dual attention network to obtain spatial dependency features and channel dependency features, including: Learning the spatial dependency of the comprehensive feature vector through a spatial attention module to obtain a spatial dependency feature; The channel-dependent feature is obtained by learning the dependency of the comprehensive feature vector on the channel through the channel attention module.

7. The network security situation awareness method according to claim 6 is characterized in that: The spatial dependency relationship of the comprehensive feature vector is learned through the spatial attention module to obtain spatial dependency features, which specifically include: The comprehensive feature vector is convolved through three convolutional layers to obtain a first feature vector, a second feature vector and a third feature vector; Reshape the first feature vector, the second feature vector and the third feature vector into two-dimensional features respectively to obtain a first two-dimensional feature, a second two-dimensional feature and a third two-dimensional feature; Transpose the first two-dimensional feature and multiply it with the second two-dimensional feature, and then pass it through the Softmax function to obtain a two-dimensional spatial attention map; Performing a matrix multiplication operation on the third two-dimensional feature and the transpose of the two-dimensional spatial attention map, multiplying the result by a spatial scale coefficient, and reshaping the result into a three-dimensional feature to obtain a three-dimensional spatial attention map; The three-dimensional spatial attention map is added to the comprehensive feature vector to obtain a spatial dependency feature.

8. The network security situation awareness method according to claim 6 is characterized in that: The channel attention module is used to learn the dependency of the comprehensive feature vector on the channel to obtain the channel dependency feature, which specifically includes: Reshaping the comprehensive feature vector into a two-dimensional feature to obtain a fourth two-dimensional feature; Transposing the fourth two-dimensional feature and multiplying it with the fourth two-dimensional feature, and then passing it through a Softmax function to obtain a two-dimensional channel attention map; After transposing the two-dimensional channel attention map, performing matrix multiplication operation with the fourth two-dimensional feature, multiplying it by the channel scale coefficient, and then reshaping it into a three-dimensional feature to obtain a three-dimensional channel attention map; The three-dimensional channel attention map is added to the comprehensive feature vector to obtain a channel-dependent feature.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the network security situation awareness method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the network security situation awareness method described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Information processing method and device, equipment and storage medium

    CN116915511A

  • Lightweight multi-modal medical image classification method for improving ResNeXt neural network

    CN117274662A