Network security defense method and system based on incremental network attack analysis learning

Through the method based on incremental network attack analysis and learning, dynamically identify and adapt to new attack characteristics and quickly respond to unknown attacks, solving the problem of difficult to identify new attacks and lack of dynamic response in the existing technology, and achieving efficient and flexible network security defense.

CN120200810AActive Publication Date: 2025-06-24JIANGSU SIJI TECH SERVICE CO LTD

Patent Information

Application Number
CN202510358590.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing network security defense systems are difficult to effectively identify new and variant cyber attacks, and defense models based on traditional machine learning lack dynamic response capabilities. Deep learning models rely on a large amount of pre-trained data and are not flexible enough to update, resulting in delays in responding to burst attacks.

Method used

Using an incremental network attack analysis and learning method, the network traffic characteristics are reduced dimensionality through a deep autoencoder, an initial network attack feature library is built, and real-time traffic data is processed in segments through a sliding time window, timing-related features are extracted, feature weights are dynamically adjusted, and deep reinforcement learning detection model is input for updates to generate a network security defense strategy.

Benefits of technology

It realizes dynamic identification and adaptation to new attack characteristics, quickly detects and responds to unknown attacks, improves the flexibility and adaptability of network defense, and improves the accuracy and timeliness of defense effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200810A_ABST
    Figure CN120200810A_ABST
Patent Text Reader

Abstract

The invention discloses a network security defense method and system based on incremental network attack analysis learning. The method comprises the following steps: collecting initial network flow data, and extracting a feature vector; and collecting real-time network flow data, performing segmentation processing based on a sliding time window, extracting time sequence association features from the segmented data, and matching the time sequence association features with the feature library to identify potential attacks or abnormal behaviors. And when the time sequence correlation feature is not matched with the feature library, marking the time sequence correlation feature as a candidate novel attack feature, calculating a mahalanobis distance between the time sequence correlation feature and a known attack feature to determine the attack variability, and dynamically adjusting the weight of the time sequence correlation feature. And inputting the adjusted feature weight and the real-time flow feature into a deep reinforcement learning detection model, and generating and updating a network security defense strategy. The scheme of the invention can effectively identify novel attack behaviors, dynamically respond and optimize defense strategies, and improve network security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data communication network security, and particularly relates to a network security defense method and system based on incremental network attack analysis and learning. Background Art

[0002] Network security defense usually relies on rule-based detection systems or predefined attack feature libraries to identify and intercept potential network attacks. Traditional methods mainly use technologies such as static feature matching, intrusion detection systems (IDSs), and firewalls to deal with known network attack patterns. In recent years, with the continuous evolution and complexity of network attack means, more and more methods have started to introduce machine learning and deep learning technologies to improve the intelligence level and adaptability of attack detection. These technologies establish models by analyzing a large amount of historical traffic data so as to identify and respond when attack features change.

[0003] However, there are some obvious limitations in the existing technologies. First, rule-based and static feature-based detection systems are difficult to effectively identify new and variant attacks. Especially in the case of continuously changing attack means, their detection accuracy will significantly decrease. Second, defense models based on traditional machine learning often lack sufficient dynamic response capabilities when facing real-time network traffic and cannot adapt to changes in attack features in a timely manner. In addition, although deep learning models can be used for attack detection, they usually rely on a large amount of pre-trained data and the model update is not flexible enough, making it difficult to achieve incremental learning and real-time adjustment, resulting in a certain delay in dealing with sudden attacks. Summary of the Invention

[0004] In order to solve the deficiencies existing in the prior art, the present invention provides a network security defense method and system based on incremental network attack analysis and learning to solve the technical problem of improving network security.

[0005] To solve the above technical problem, the present invention adopts the following technical solutions.

[0006] The present invention first discloses a network security defense method based on incremental network attack analysis and learning, and the method includes the following steps:

[0007] Collect initial network traffic data and extract traffic feature vectors; use a deep autoencoder to perform dimensionality reduction processing on the traffic feature vectors to obtain compressed feature vectors; construct an initial network attack feature library based on the compressed feature vectors;

[0008] Collect real-time network traffic data, and segment the real-time network traffic data based on a preset sliding time window; extract temporal correlation features from the segmented real-time traffic data; match the temporal correlation features with the features in the initial network attack feature library to identify potential attack behaviors or abnormal behaviors;

[0009] When the temporal correlation features do not match the features in the initial network attack feature library, mark the temporal correlation features as candidate new attack features; calculate the Mahalanobis distance between the temporal correlation features and the attack features stored in the initial network attack feature library to determine the attack variability of the new attack features;

[0010] Based on the determined attack variability, dynamically adjust the feature weights of the temporal correlation features;

[0011] Input the feature weights of the dynamically adjusted temporal correlation features and the temporal correlation features of the real-time network traffic into a deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits to update the deep reinforcement learning detection model;

[0012] Input the temporal correlation features of the real-time network traffic into the updated deep reinforcement learning detection model to obtain a network security defense strategy, where the network security defense strategy includes traffic restriction, access control, and data isolation.

[0013] The present invention further includes the following preferred solutions:

[0014] The collection of the initial network traffic data and the extraction of the traffic feature vectors further include:

[0015] Collect initial network traffic data from multiple network ingress nodes and egress nodes, covering different traffic sources and paths in the network. The collected initial network traffic data includes packet header information, packet size, traffic direction, and timestamp;

[0016] Preprocess the collected traffic data to remove redundant data, noise data, and irrelevant traffic data, and retain the key data that may contain attack features to generate preprocessed traffic data;

[0017] Segment the preprocessed traffic data, and extract multi-dimensional feature information from each segment of data to form traffic feature vectors.

[0018] The construction of the initial network attack feature library based on the compressed feature vectors further includes:

[0019] Cluster the compressed feature vectors according to the similarity of the traffic features, and group the vectors with similar features to reveal potential attack patterns;

[0020] Analyze the feature vectors in each cluster, extract representative features as the feature identifiers for each attack pattern, and record the feature identifiers in the initial network attack feature library;

[0021] By counting the occurrence frequency and distribution range of each feature identifier, add weight information to each attack pattern in the feature library. The weight information is used to provide a priority reference for subsequent real-time traffic matching;

[0022] Store the completed initial network attack feature library.

[0023] The step of collecting real-time network traffic data and segmenting the real-time network traffic data based on a preset sliding time window further includes:

[0024] Collect real-time network traffic data and transfer the collected real-time traffic data to the sliding window processing module;

[0025] In the sliding window processing module, segment the real-time network traffic data according to the preset sliding time window. Each sliding window represents a fixed time interval, and within this time interval, intercept and store the traffic data to form a traffic data segment corresponding to the window;

[0026] At the end of the window, encapsulate the data segment within the current time window as a whole, and record the start time and end time of the window;

[0027] At the start of the next time step, the sliding window automatically moves forward by one step, adds the newly incoming real-time traffic data to the new window, and forms a new traffic data segment, thereby iteratively achieving continuous segmentation processing of real-time network traffic.

[0028] The step of extracting temporal correlation features from the segmented real-time traffic data further includes:

[0029] Obtain the real-time traffic data segments after being segmented by the sliding time window;

[0030] Perform temporal feature analysis on each data segment, extract basic temporal features including packet transmission rate, packet time interval, traffic direction, and traffic peak, and obtain the dynamic traffic change features within the data segment;

[0031] Analyze the correlation between consecutive data segments, identify cross-time segment feature patterns including periodic traffic changes, burst traffic, and abnormal delays, and obtain the traffic change trend features;

[0032] Combine the extracted basic timing features and traffic change trend features to form a timing correlation feature vector for subsequent attack detection and feature matching.

[0033] The deep autoencoder includes an input layer, an encoding layer, a bottleneck layer, a decoding layer, and a feature selection layer;

[0034] The input layer is used to receive the traffic feature vector extracted from the initial network traffic data; perform normalization processing on the received data to obtain a normalized feature vector;

[0035] The encoding layer is used to receive the normalized feature vector from the input layer and extract features layer by layer through a stacked multi-layer convolutional network; each convolutional layer in the encoding layer applies a convolutional kernel of a specific size to capture the high-order feature relationships in the feature vector through a non-linear activation function; among them, the size of the convolutional kernel is adaptively adjusted layer by layer according to the distribution of the input features and the feature dimension of the network traffic; the output of the encoding layer is a low-dimensional encoded feature representation;

[0036] The bottleneck layer is used for the low-dimensional encoded feature representation, further compressing the feature dimension through a sparse activation function to generate a sparse compressed feature vector;

[0037] The decoding layer is used to receive the sparse compressed feature vector from the bottleneck layer, and perform layer-by-layer decoding and reconstruction on the compressed features through a multi-layer deconvolutional network symmetric to the encoding layer to obtain a high-dimensional feature representation; the output of the decoding layer is the reconstructed feature vector, which is used to calculate the reconstruction error with the original input during the training process to optimize the parameters of each layer of the autoencoder;

[0038] The feature selection layer is used to receive the sparse compressed feature vector from the bottleneck layer, weight each feature dimension through an attention mechanism, further screen out important features, and obtain a compressed feature vector.

[0039] Calculating the Mahalanobis distance between the timing correlation feature and the attack features stored in the initial network attack feature library to determine the attack variability of the new attack feature further includes:

[0040] Use formula (1) to calculate the attack variability of the new attack feature:

[0041]

[0042] Among them, D represents the attack variability of the timing correlation feature; N represents the number of attack features stored in the initial network attack feature library; d i represents the Mahalanobis distance between the timing correlation feature and the i-th known attack feature; σ iis the standard deviation of the feature values of the i-th known attack feature on different sample data in the initial network attack feature library; β i is the risk coefficient of the i-th known attack feature, determined according to the attack severity in historical data; det(Σ i ) is the determinant of the covariance matrix of the i-th known attack feature; w i represents the weight coefficient associated with the i-th known attack feature; f(d i ) is a regularization factor;

[0043] where the Mahalanobis distance d i is calculated by formula (2):

[0044]

[0045] where x represents the feature vector of the time-series correlation feature; μ i represents the mean vector of the i-th known attack feature; represents the inverse matrix of the covariance matrix of the i-th known attack feature;

[0046] where the weight coefficient w i is calculated by formula (3):

[0047]

[0048] where α is a tuning parameter; r i is the historical occurrence frequency of the i-th known attack feature; r threshold is a preset frequency threshold for distinguishing high-frequency and low-frequency attack features;

[0049] where the regularization factor f(d i ) is calculated by formula (4):

[0050]

[0051] where γ is a tuning parameter for controlling the amplification effect of the regularization factor; p is an exponential parameter for adjusting the nonlinear degree of regularization.

[0052] The dynamic adjustment of the feature weights of the time-series correlation features based on the determined attack variability further includes:

[0053] Determine the adjusted feature weight W through formula (5):

[0054] W = W0 × (1 + α0·log(1 + β0·(D - D threshold )))(5) p0 )

[0055] Among them, W0 represents the initial feature weight of the timing correlation feature; α0 is an adjustment coefficient used to control the increase amplitude of the feature weight; β0 is a non-linear adjustment parameter; D threshold is a preset threshold for the attack mutation degree; p0 is an exponential parameter.

[0056] The objective function of the deep reinforcement learning detection model is defined as:

[0057]

[0058] Among them, J(θ) represents the objective function of the deep reinforcement learning model, and this objective function is to maximize the cumulative reward based on the deep reinforcement learning model parameters θ; represents the mathematical expectation of all states and actions; T is the maximum number of time steps within the training period; t represents the current time step of training; γ is a discount factor used to balance the immediate reward and the long-term reward; R(s t , a t ) represents the immediate reward obtained by executing the action a t in the state s t ; λ is a regularization parameter used to control the weights of different reward terms; K is the total number of regularization features; k is the index of the regularization feature; φ k (s t , a t ) is a penalty function;

[0059] Among them, the penalty function φ k (s t , a t ) is implemented using formula (7):

[0060] φ k (s t , a t ) = α k log(1 + β k d k (s t , a t ) p1 )(7)

[0061] Among them, α k and β k are adjustment parameters related to the k-th feature, used to control the weight and non-linearity degree of the penalty term; p1 is an exponential parameter; d k (s t , a t ) is the deviation between the k-th feature and the expected feature under the current state s t and the action a t , and is calculated using formula (8):

[0062]

[0063] Among them, f k (s t , a t ) means in state s t Execute action a t The kth eigenvalue when ; is the target value of the kth feature, which represents the feature level under normal or expected behavior mode;

[0064] The immediate reward function is implemented using formula (9):

[0065] R(s t , a t )=κ·ρ·tanh(γ1·(e t -e min ))-ν·Ω(a t ) (9)

[0066] Among them, R(s t , a t ) is the immediate reward function, used to evaluate the state s t Take action a t The instant benefit; κ is the reward coefficient for successful defense, which represents the reward intensity obtained by the model when it correctly detects an attack or effectively defends; ρ is the threat level coefficient, which is dynamically determined according to the currently detected attack type and threat level; γ1 is the amplification parameter, which is used to control the sensitivity of the reward to the deviation of the defense effect; e t Indicates the current action a t In status t The defensive effect score under e min is the minimum effect benchmark, which is used to ensure that the reward function only gives significant rewards after the effect exceeds a certain level, thereby suppressing low-effect defense behaviors; v represents the defense cost adjustment coefficient, which is used to control the penalty intensity of the action cost on the reward; Ω(a t ) is the defense strategy a t The execution cost function is determined according to the amount of system resources occupied by the currently selected defense strategy.

[0067] The deep reinforcement learning detection model includes a first input layer, a feature extraction layer, a strategy generation layer and an output layer, and operates through the following steps:

[0068] The first input layer is used to receive the dynamically adjusted feature weights of the timing correlation features and the timing correlation features of the real-time network traffic;

[0069] The feature extraction layer includes multiple hidden layers, and processes the time series correlation features through a recursive neural network to extract the time series pattern and potential pattern of the real-time network traffic;

[0070] The policy generation layer uses the obtained temporal patterns and feature weights to generate corresponding defense policies through a deep reinforcement learning algorithm; the policy generation layer scores each defense policy based on a deep Q-network and selects the policy with the highest score for application; the policy generation layer iteratively optimizes the effect of the selected policy during the training process to maximize network security benefits;

[0071] The output layer generates network security defense policies, including traffic restriction, access control, or data isolation.

[0072] The present invention also discloses a network security defense system based on incremental network attack analysis and learning using the aforementioned network security defense method based on incremental network attack analysis and learning, further comprising:

[0073] An initial network attack feature library construction module, configured to collect initial network traffic data, extract traffic feature vectors; use a deep autoencoder to perform dimensionality reduction processing on the traffic feature vectors to obtain compressed feature vectors; construct an initial network attack feature library based on the compressed feature vectors;

[0074] A temporal correlation feature matching module, configured to collect real-time network traffic data and perform segmentation processing on the real-time network traffic data based on a preset sliding time window; extract temporal correlation features from the segmented real-time traffic data; match the temporal correlation features with the features in the initial network attack feature library to identify potential attack behaviors or abnormal behaviors;

[0075] An attack variability calculation module, configured to mark the temporal correlation features as candidate new attack features when the temporal correlation features do not match the features in the initial network attack feature library; calculate the Mahalanobis distance between the temporal correlation features and the attack features stored in the initial network attack feature library to determine the attack variability of the new attack features;

[0076] A feature weight adjustment module, configured to dynamically adjust the feature weights of the temporal correlation features based on the determined attack variability;

[0077] A deep reinforcement learning detection module, configured to input the feature weights of the dynamically adjusted temporal correlation features and the temporal correlation features of real-time network traffic into a deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits and updates the deep reinforcement learning detection model;

[0078] A network security defense policy acquisition module, configured to input the temporal correlation features of real-time network traffic into the updated deep reinforcement learning detection model to obtain network security defense policies, and the network security defense policies include traffic restriction, access control, and data isolation.

[0079] Correspondingly, the present application also discloses a terminal, including a processor and a storage medium;

[0080] The storage medium is used for storing instructions;

[0081] The processor is used for operating according to the instructions to execute the steps of the network security defense method based on incremental network attack analysis and learning as described above.

[0082] Correspondingly, the present application also discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the network security defense method based on incremental network attack analysis and learning as described above.

[0083] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention provides a network security defense method and system based on incremental network attack analysis and learning. By matching the time-series correlation features with the initial network attack feature library and calculating the attack mutation degree, it can dynamically identify and adapt to new attack features, quickly detect and respond to unknown attacks, and improve the flexibility and adaptability of network defense. Based on the deep reinforcement learning model, by inputting the dynamically adjusted time-series correlation feature weights and time-series correlation features, the optimization of network security benefits is realized, and the defense strategy is updated in real time, so that the defense measures can be automatically adjusted with the change of attack features, effectively improving the accuracy and timeliness of the defense effect. Using the incremental learning mechanism, after new attack features are identified and the weights are adjusted, the deep reinforcement learning detection model is iteratively optimized. Compared with the traditional model that needs to be retrained, incremental learning improves the adaptability of the model, enabling it to always maintain high detection performance in a changing network environment. By using a deep autoencoder to perform dimensionality reduction processing on the traffic feature vector and constructing an initial network attack feature library, the core feature information of network traffic is retained. Combining with the dynamically adjusted feature weights, it effectively reduces the misjudgment of normal traffic, reduces the false alarm rate, and improves the accuracy and robustness of detection. Description of the Drawings

[0084] Figure 1 is a flowchart of the network security defense method based on incremental network attack analysis and learning in the present invention. Detailed Embodiments

[0085] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0086] The embodiments described in this application are only a part of the embodiments of the present invention, not all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0087] In view of the deficiencies of the prior art, the present invention proposes a network security defense method and system based on incremental network attack analysis and learning. Refer to Figure 1 as shown, the method includes the following steps:

[0088] Step S101: Collect initial network traffic data, extract traffic feature vectors; use a deep autoencoder to perform dimensionality reduction processing on the traffic feature vectors to obtain compressed feature vectors; construct an initial network attack feature library based on the compressed feature vectors.

[0089] In step S101, first collect initial network traffic data. When collecting, obtain diverse traffic in the network environment through multiple network entry nodes and exit nodes to ensure that the collected data can fully cover various normal traffic and known attack traffic in the network. When collecting data, pay attention to the basic attributes of network packets, including source IP address, destination IP address, protocol type, packet size, port information, and timestamp, etc., so as to provide a comprehensive data basis for subsequent feature extraction and analysis. The data collection can be carried out at fixed time intervals to ensure the time distribution and integrity of the sample data.

[0090] After the data collection is completed, next extract traffic feature vectors to convert the network traffic data into feature vectors available for analysis. For this purpose, first preprocess the collected data, including data denoising and duplicate removal, to ensure that the extracted features accurately reflect the actual characteristics of the network traffic. Then extract key information from each network packet to form a feature vector. The construction of the feature vector can cover a variety of network traffic features, such as the protocol type of the packet, transmission direction, packet size, transmission time interval, connection duration, number of packets per second, distribution characteristics of packet arrival intervals, etc. The processed feature vector can comprehensively describe the behavior characteristics of the network traffic and provide a basis for subsequent pattern recognition.

[0091] The extracted feature vectors are then input into a deep autoencoder for dimensionality reduction. A deep autoencoder is a non-linear dimensionality reduction method based on neural networks. By encoding and decoding the feature vectors, it can effectively remove redundant information and extract the core features. The purpose of dimensionality reduction is to reduce the complexity of the data while retaining the key information related to attack features. The training of the deep autoencoder can be based on an unsupervised learning approach, using the initial traffic dataset for adaptive learning to extract the core patterns of various network traffic. The feature vectors after dimensionality reduction are called compressed feature vectors, which are simplified representations of the original features and contain the most representative features in the initial traffic data, facilitating subsequent efficient attack detection.

[0092] After obtaining the compressed feature vectors, they are used to construct an initial network attack feature library. In the process of constructing the initial network attack feature library, through the clustering analysis of the compressed feature vectors, the vectors with similar features are grouped into the same category to form the basic units of attack patterns. For the attack pattern features obtained from each clustering, record their key attributes, including the core information and occurrence frequency of the feature vectors, etc., and store them in the initial network attack feature library. This feature library can be used for feature matching and recognition in subsequent real-time traffic analysis. The establishment of the initial network attack feature library aims to construct a benchmark covering common attack features to facilitate the rapid identification of potential attacks and abnormal behaviors in subsequent real-time detection.

[0093] Through step S101, a feature library based on the initial network traffic is obtained, which can cover the typical traffic features in the network environment and provide a reference standard and initial data support for the subsequent detection and defense processes.

[0094] Furthermore, the collecting of the initial network traffic data and the extracting of the traffic feature vectors further include:

[0095] Collect the initial network traffic data from multiple network entry nodes and exit nodes, covering different traffic sources and paths in the network. The collected initial network traffic data includes packet header information, packet size, traffic direction, and timestamp;

[0096] Preprocess the collected traffic data to remove redundant data, noise data, and irrelevant traffic data, and retain the key data that may contain attack features to generate preprocessed traffic data;

[0097] Segment the preprocessed traffic data, and extract multi-dimensional feature information from each segment of data to form traffic feature vectors.

[0098] The process of collecting initial network traffic data and extracting traffic feature vectors includes collecting data from multiple network ingress and egress nodes, preprocessing the data to remove redundancy, and extracting features from the processed data. First, initial network traffic data is collected through multi-ingress nodes and egress nodes distributed throughout the network to comprehensively cover possible different traffic sources and paths in the network, thereby ensuring that the collected data is representative and extensive. The collected data should cover typical traffic information, including packet header information, packet size, traffic direction, and timestamp. These basic data provide the core information of network traffic and help in detailed traffic feature analysis in subsequent steps.

[0099] After data collection is completed, the preprocessing stage will be entered. In this stage, the traffic data is cleaned to remove redundant data, noise data, and irrelevant traffic data. For example, broadcast data or regular heartbeat signals in certain network activities do not contain effective information in security detection, so they can be removed during preprocessing to focus on retaining key data that may contain attack features. The preprocessing stage ensures the purity and efficiency of the traffic data, reduces data interference in subsequent analysis, and generates preprocessed traffic data, laying the foundation for feature extraction.

[0100] Next, the preprocessed traffic data is segmented to ensure the continuity and relevance of data analysis. During the segmentation process, each segment of data is divided by a specific time or the number of packets, and each segment of data forms an independent data block. The purpose of segmentation is to ensure that correlated features in time series can be captured during analysis, thereby more effectively extracting attack features in network traffic. Each data segment is used to extract multi-dimensional feature information, which includes but is not limited to protocol type, source and destination ports, packet size, time interval, connection duration, etc. By extracting multi-dimensional feature information from each segment of traffic data, the system can form a comprehensive traffic feature vector that contains the core behavioral features of network traffic and provides detailed feature data support for further attack detection.

[0101] Furthermore, constructing the initial network attack feature library based on the compressed feature vector further includes:

[0102] Clustering the compressed feature vectors according to the similarity of traffic features, grouping vectors with similar features, thereby revealing potential attack patterns;

[0103] Analyzing the feature vectors in each cluster, extracting representative features as the feature identifiers for each attack pattern, and recording the feature identifiers in the initial network attack feature library;

[0104] By counting the occurrence frequency and distribution range of each feature identifier, weight information is added to each attack pattern in the feature library to reflect the importance of this pattern in network traffic. The weight information is used to provide a priority reference in subsequent real-time traffic matching;

[0105] Store the initially constructed network attack feature library.

[0106] First, cluster the compressed feature vectors according to the similarity of traffic features. The purpose of clustering is to group feature vectors with similar features in order to identify and reveal potential attack patterns. During the clustering process, the system calculates the similarity of feature vectors based on their multi-dimensional features, groups vectors with higher similarity into the same group, and forms a clustering group representing a specific traffic feature or attack pattern. The purpose of this process is to identify possible attack feature patterns in the traffic data and provide a basis for subsequent feature library construction.

[0107] After clustering is completed, perform a detailed analysis on the feature vectors within each clustering group to extract the representative features of this clustering group. These representative features are used to form the feature identifiers of the attack patterns. The system extracts the features that can best reflect the characteristics of the group by analyzing the commonalities of the feature vectors in each clustering group and uses them as the unique identifiers of the attack patterns. Subsequently, record this feature identifier into the initially constructed network attack feature library so that the feature library can cover known typical attack features or traffic patterns.

[0108] After recording the feature identifiers, assign weight information to each attack pattern in the feature library by counting the occurrence frequency and distribution range of each feature identifier in the historical network traffic data. This weight information is used to reflect the importance of each attack pattern in network traffic. The system adjusts the weights based on the frequency of the attack pattern and its typicality in the traffic, assigns higher weight values to high-frequency attack patterns, and thus provides a priority reference in real-time detection. The weight information ensures that the system can quickly identify and preferentially detect frequently occurring or high-risk attack patterns when analyzing real-time traffic, thereby improving the detection efficiency.

[0109] After completing the above steps, store the constructed initially network attack feature library as a long-term reference data source so that feature matching can be performed quickly and accurately during subsequent real-time traffic matching. The storage of the feature library enables the system to call these analyzed and weighted attack patterns at any time during subsequent operations, thus playing a role in real-time network security defense and providing accurate detection and identification support for potential attack behaviors in network traffic.

[0110] Furthermore, the deep autoencoder includes an input layer, an encoding layer, a bottleneck layer, a decoding layer, and a feature selection layer;

[0111] Among them, the input layer is used to receive the traffic feature vectors extracted from the initial network traffic data; perform normalization processing on the received data to obtain the normalized feature vectors;

[0112] The encoding layer is used to receive the normalized feature vectors from the input layer, and extract features layer by layer through stacking multiple convolutional networks; each convolutional layer in the encoding layer applies a convolutional kernel of a specific size, and captures the high-order feature relationships in the feature vectors through a non-linear activation function; among them, the size of the convolutional kernel is adaptively adjusted according to the distribution of the input features and the feature dimensions of the network traffic in each layer; the output of the encoding layer is a low-dimensional encoded feature representation;

[0113] The bottleneck layer is used for the low-dimensional encoded feature representation, and further compresses the feature dimensions through a sparse activation function to generate a sparse compressed feature vector;

[0114] The decoding layer is used to receive the sparse compressed feature vectors from the bottleneck layer, and perform layer-by-layer decoding and reconstruction on the compressed features through a multi-layer deconvolutional network symmetric to the encoding layer to restore a high-dimensional feature representation close to the input feature vectors; the output of the decoding layer is the reconstructed feature vector, which is used to calculate the reconstruction error with the original input during the training process to optimize the parameters of each layer of the autoencoder;

[0115] The feature selection layer is used to receive the sparse compressed feature vectors from the bottleneck layer, weight each feature dimension through an attention mechanism, and further screen out important features to obtain a compressed feature vector.

[0116] The deep autoencoder includes an input layer, an encoding layer, a bottleneck layer, a decoding layer, and a feature selection layer. First, the input layer receives the traffic feature vectors extracted from the initial network traffic data. To make these feature vectors adapt to the processing of subsequent layers, the input layer performs normalization processing on the received feature data and converts it into normalized feature vectors. This normalization process ensures that the features have a unified scale in different dimensions, reducing the impact of scale differences between different features on subsequent processing.

[0117] After the normalization process is completed, the feature vector enters the encoding layer. The encoding layer consists of multiple stacked convolutional networks. Each convolutional layer applies a convolutional kernel of a specific size to extract the high-order feature relationships in the traffic features layer by layer. Each convolutional layer also activates the extracted features through a non-linear activation function (such as ReLU) to ensure that complex non-linear patterns in the feature vector can be captured. The size of the convolutional kernel is adaptively adjusted in each layer according to the distribution of the input features and the dimensions of the network traffic features. This dynamic adjustment mechanism enables the encoding layer to effectively extract feature patterns at different levels. The encoding layer finally outputs a low-dimensional encoded feature representation, which retains the key information in the original feature vector but removes the redundant part through dimensionality reduction.

[0118] The low-dimensional encoded feature representation is then passed to the bottleneck layer. The bottleneck layer further compresses the low-dimensional features through a sparse activation function, forming a sparse compressed feature vector. The sparse compression not only further reduces the feature dimensions but also ensures that only the most important features are retained, reducing the demand for the model's computing resources while highlighting the significant features in the traffic data, thereby enhancing the effectiveness of subsequent analysis.

[0119] The sparse compressed feature vector output from the bottleneck layer then enters the decoding layer. In the decoding layer, through a multi-layer transposed convolutional network symmetric to the encoding layer, the compressed features are decoded layer by layer to gradually restore a high-dimensional feature representation close to the feature vector of the input layer. This decoding process can retain the most critical feature information in the data while restoring a structure close to the original data, enabling the model to optimize the parameters of each layer of the autoencoder by calculating the error between the reconstructed feature vector and the original input during the training process. Through the feedback of the reconstruction error, the model continuously adjusts the parameters in each iteration, making the encoding and decoding processes more accurately retain the important information in the original data.

[0120] Finally, the decoded sparse compressed feature vector is passed to the feature selection layer. The feature selection layer uses the attention mechanism to weight each feature dimension, suppressing irrelevant features by identifying and enhancing the key features in the data. This layer dynamically calculates the weight of each feature and preferentially retains the features with larger weights as important features. The final compressed feature vector output by the feature selection layer is an optimized vector containing the key features of the network traffic, providing an efficient and accurate input for subsequent network attack detection.

[0121] Step S102: Collect real-time network traffic data, and segment the real-time network traffic data based on a preset sliding time window; extract temporal correlation features from the segmented real-time traffic data; match the temporal correlation features with the features in the initial network attack feature library to identify potential attack behaviors or abnormal behaviors.

[0122] The step S102 aims to detect potential attack behaviors or abnormal behaviors in the current network traffic. First, the system collects real-time network traffic data through multiple ingress and egress nodes to obtain comprehensive and real-time traffic information in the network environment. This collection process ensures that various types of traffic in the network can be covered, especially high-traffic areas, enhancing the representativeness and integrity of the data.

[0123] After collecting the real-time traffic data, the data is segmented according to a preset sliding time window. The sliding time window is a time frame for dividing data, and the data is segmented at fixed time intervals. This time window slides step by step at a fixed step size, so that each segmentation contains new traffic data and maintains a certain continuity, thus forming a series of traffic data segments with time correlation. In this way, multiple time window data segments with partial overlap can be formed in the time series, effectively retaining the changing trend and short-term dynamic information of the network traffic over time.

[0124] At the end of the time window, the system takes the real-time traffic data within the current window as a complete data segment and marks the start and end times of this segment to ensure the time order and traffic integrity of each data segment. At the beginning of the next time window, the sliding window automatically moves forward by one step size, adds the newly collected traffic data to the new window, and repeats the segmentation process. In this way, each data segment can cover the traffic data within a specific time period, enabling the system to perform accurate time series analysis in the subsequent processing stage.

[0125] From the segmented real-time traffic data, the system extracts time series correlation features. Specifically, the system analyzes information such as the arrival frequency of data packets, the time interval between data packets, the transmission rate, the traffic direction, the packet size, and the peak traffic in each data segment, so as to construct time series correlation features reflecting the characteristics of network traffic. By analyzing these features, the system can capture changes in traffic patterns and potential abnormal behaviors. These time series correlation features provide basic data for subsequent attack detection, especially for identifying abnormal behaviors in normal traffic, such as burst traffic, abnormal delays, etc.

[0126] The extracted time series correlation features are then matched with the features in the initial network attack feature library. The feature matching process aims to determine whether the behavior of the current real-time traffic conforms to known attack features or normal traffic patterns. By comparing with the features in the feature library, the system can identify potential attack behaviors or abnormal patterns. This matching process combines the correlation information of time series features, making the recognition results more accurate and reliable.

[0127] Furthermore, the collecting of the real-time network traffic data and the segmenting of the real-time network traffic data based on a preset sliding time window further include:

[0128] Collect real-time network traffic data and transfer the collected real-time traffic data to the sliding window processing module;

[0129] In the sliding window processing module, segment the real-time network traffic data according to a preset sliding time window. Each sliding window represents a fixed time interval, and within this time interval, intercept and store the traffic data to form a traffic data segment corresponding to the window;

[0130] At the end of the window, encapsulate the data segment within the current time window as a whole and record the start time and end time of the window;

[0131] At the start of the next time step, the sliding window automatically moves forward by one step, adds the newly incoming real-time traffic data to the new window, and forms a new traffic data segment, thereby iteratively achieving continuous segmentation processing of real-time network traffic.

[0132] First, collect real-time traffic data through the network traffic monitoring system and transfer the collected real-time data to the sliding window processing module for continuous time period division. In the sliding window processing module, the system uses a preset sliding time window to segment the data. The sliding time window is a fixed time interval, and each time window intercepts the real-time traffic data within this time period and stores the data within this time period to form a corresponding traffic data segment.

[0133] At the end of each sliding window, the system encapsulates the traffic data collected within the current time window into a complete data segment and marks the start and end times of the window. The purpose of recording the time information is to ensure that the time sequence of the data segments can be accurately traced during subsequent analysis, thereby ensuring the integrity and accuracy of the time series analysis.

[0134] When a time window ends and the data segment is encapsulated, the sliding window module automatically moves forward by one step, thereby starting the data collection for the next time period. This step movement method ensures the continuity of the sliding window. The system automatically adds the real-time traffic data to the new window to form a new traffic data segment. The sliding window gradually covers the real-time traffic in this way, forming a series of continuous traffic data segments, achieving seamless segmentation processing of real-time network traffic.

[0135] Through the sliding window processing method, it is possible to perform fine-grained time segmentation on network traffic while ensuring the time continuity of the data, which is convenient for subsequent feature extraction and analysis processing. The design of the sliding time window ensures the integrity of the traffic data within each time period and provides a reliable data basis for the dynamic analysis of real-time traffic.

[0136] Furthermore, extract temporal correlation features from the segmented real-time traffic data, which further includes:

[0137] Obtain the real-time traffic data segments after segmented processing by a sliding time window;

[0138] Perform temporal feature analysis on each data segment, extract basic temporal features including packet transmission rate, packet time interval, traffic direction, and traffic peak, and obtain the traffic dynamic change features within the data segment;

[0139] Analyze the correlation between consecutive data segments, identify the feature patterns across time periods including periodic traffic changes, burst traffic, and abnormal delays, and obtain the traffic change trend features;

[0140] Combine the extracted basic temporal features and traffic change trend features to form a temporal correlation feature vector for subsequent attack detection and feature matching.

[0141] First, obtain the real-time traffic data segments after segmented processing by a sliding time window. These data segments contain the network traffic information within a specific time window, facilitating a detailed analysis of the traffic characteristics in each time period. After obtaining these segmented data, start performing temporal feature analysis on each data segment, and extract the basic temporal features including packet transmission rate, the time interval between packets, traffic direction, and traffic peak. By extracting these basic features, the system can obtain the traffic dynamic change features within that time period, thereby reflecting the change trend of network traffic in a short time.

[0142] After extracting the basic temporal features of each data segment, it is also necessary to analyze the correlation between consecutive data segments. By observing the feature changes in adjacent time periods, the system can identify the feature patterns across time periods, such as periodic traffic changes, sudden traffic peaks, and abnormal delay situations. These cross-time period features can reveal the overall trend of network traffic, such as whether there are regular traffic peaks or abnormal traffic behaviors, enabling the system to more accurately capture potential threats in the network environment.

[0143] Finally, the system combines the extracted basic temporal features and traffic change trend features to form a complete temporal correlation feature vector. This feature vector contains both the dynamic change features within a single time period and the change trends between multiple time periods, and can comprehensively describe the temporal correlation of network traffic. This temporal correlation feature vector is then used for subsequent attack detection and feature matching, so that the system can more accurately identify potential attack behaviors and improve the detection ability for abnormal traffic.

[0144] Step S103: When the timing correlation feature does not match the features in the initial network attack feature library, mark the timing correlation feature as a candidate new attack feature; calculate the Mahalanobis distance between the timing correlation feature and the attack features stored in the initial network attack feature library to determine the attack variability of the new attack feature.

[0145] In step S103, first, when the timing correlation feature in the real-time network traffic data does not match the known features in the initial network attack feature library, the system regards the detected feature as a potential new attack feature. The situation of non - matching usually indicates that there are significant differences between the current traffic feature and the normal or known attack patterns stored in the system, thus triggering further analysis. After detecting such non - matching, the system immediately marks the feature so that subsequent steps can pay further attention to and evaluate this feature.

[0146] After marking the timing correlation feature, then calculate the Mahalanobis distance between this feature and the attack features stored in the initial network attack feature library. The Mahalanobis distance is a statistical method used to measure the similarity or deviation degree between the new feature and the existing features. By calculating the Mahalanobis distance between the timing correlation feature and each known attack feature in the feature library, the deviation degree of this new feature relative to the known features can be quantified, thus providing a basis for determining its attack variability. When calculating the Mahalanobis distance, multi - dimensional features in the feature vector are comprehensively considered, and the distance calculation is adjusted according to the distribution characteristics of these features to ensure that the calculation result of the distance can accurately reflect the difference between the new feature and the known features.

[0147] After obtaining the Mahalanobis distance, convert this distance into an "attack variability" value. The attack variability is a quantified index used to represent the deviation degree between the timing correlation feature and the existing attack features. The higher the variability, the greater the difference between this feature and the known attack features, indicating that it is more likely to be a new or variant attack. The calculation process of the variability ensures that the system can distinguish different degrees of new attacks, thus providing a reference basis for subsequent defense strategies.

[0148] Through the above steps, the system can timely identify candidate new attack features, and generate an attack variability according to the deviation degree between the new feature and the existing features, which is convenient for subsequent dynamic adjustment and the generation of defense strategies. This process ensures that the system has a stronger response ability when facing new attacks, and at the same time provides a flexible and accurate analysis tool for identifying and marking new network attacks.

[0149] Furthermore, the calculating the Mahalanobis distance between the timing correlation feature and the attack features stored in the initial network attack feature library to determine the attack variability of the new attack feature further includes:

[0150] Using formula (1), calculate the attack variability of the new attack feature:

[0151]

[0152] Among them, D represents the attack variability of the time-series correlation feature; this value measures the similarity and potential threat of the time-series correlation feature with the attack features stored in the initial network attack feature library. A higher D value usually indicates that the time-series correlation feature has higher aggressiveness or variability.

[0153] N represents the number of attack features stored in the initial network attack feature library; this number refers to the total number of typical attack patterns stored in the feature library.

[0154] d i represents the Mahalanobis distance between the time-series correlation feature and the i-th known attack feature;

[0155] σ i is the standard deviation of the feature values of the i-th known attack feature on different sample data in the initial network attack feature library, reflecting the feature distribution of the i-th known attack feature. For example, if the feature distribution of the attack feature is relatively concentrated, the standard deviation σ i will be smaller, thus amplifying the contribution of d i in the formula and increasing the attack variability.

[0156] β i is the risk coefficient of the i-th known attack feature, determined according to the attack severity in historical data; features with high risk coefficients often correspond to more severe attack types, such as attacks that cause service interruptions.

[0157] det(Σ i ) is the determinant of the covariance matrix of the i-th known attack feature, reflecting the distribution of this feature in the multi-dimensional feature space. The covariance matrix describes the correlation between each feature dimension, and the determinant value is used to adjust the weight in the calculation of the attack variability.

[0158] w i represents the weight coefficient associated with the i-th known attack feature;

[0159] f(d i ) is a regularization factor;

[0160] Among them, the Mahalanobis distance d i is calculated by formula (2):

[0161]

[0162] Among them, x represents the feature vector of the time-series correlation feature; μ irepresents the mean vector of the i-th known attack feature; represents the inverse matrix of the covariance matrix of the i-th known attack feature; The Mahalanobis distance is used to measure the similarity between the time-series correlation feature and each known attack feature, and can consider the distribution and correlation of each feature.

[0163] weight coefficient w i is calculated by formula (3):

[0164]

[0165] where α is a tuning parameter that controls the sensitivity of weight adjustment;

[0166] r i is the historical occurrence frequency of the i-th known attack feature; r threshold is a preset frequency threshold used to distinguish high-frequency and low-frequency attack features; For example, for high-frequency attack features, the weight w i will be larger, thus giving higher influence when calculating the attack variability.

[0167] regularization factor f(d i ) is calculated by formula (4):

[0168]

[0169] where γ is a tuning parameter used to control the amplification effect of the regularization factor; p is an exponential parameter used to adjust the non-linearity of regularization. Through the f(d i ) function, the contributions of different Mahalanobis distance values can be adjusted to ensure appropriate amplification or reduction of the attack variability when the distance is large.

[0170] In this network security defense method, the tuning parameters (such as α, γ, p, etc. in the formula) are usually determined through experiments and optimization processes to ensure the effectiveness and adaptability of the model under different attack scenarios. The acquisition of these parameters can generally adopt one of the following methods:

[0171] (1) Historical data analysis and initial setting: The initial values of the tuning parameters can be obtained by analyzing a large amount of historical network traffic data. During the analysis process, researchers will observe factors such as the distribution, frequency, and severity of different attack features in the feature library. Based on these statistical information, reasonable initial parameter values are set for subsequent optimization.

[0172] (2) Training and validation: Use a part of the training data with known attack features to train the network defense model. During the training process, the model adjusts the calculation results of the attack variability according to different adjustment parameter values. Then, use the validation sets of different attack types to evaluate the model performance to ensure that the parameters can handle the distribution changes, risk weights, and relationships between features of various attack features.

[0173] (3) Expert adjustment: In some complex network environments, it may be necessary to fine-tune these parameters in combination with the judgment of security experts. Security experts can manually adjust the parameters according to the actual severity, frequency, and other characteristics of the attack patterns to meet special network defense requirements.

[0174] Step S104: Dynamically adjust the feature weights of the time-series correlation features based on the determined attack variability.

[0175] The step S104 is used to improve the response sensitivity and detection effect of the system to new attack features. First, obtain the attack variability of the time-series correlation features in the previous step, and this variability represents the deviation degree between the time-series correlation features and the known attack features. After determining the attack variability, dynamically adjust the weights of the time-series correlation features according to the variability value, so as to give higher attention to this new feature in the subsequent detection model.

[0176] In the specific adjustment process, according to the preset weight adjustment rule, assign higher feature weights to the time-series correlation features with higher attack variability. Such a dynamic adjustment method ensures that the attack features with higher variability can occupy a more important position in the feature weights, enabling the detection model to give priority to these high-variability features when analyzing traffic data. This weight adjustment process may be based on a non-linear function or an adaptive strategy. While the system assigns higher weights to the attack features with high variability, it also controls the amplitude of the weight adjustment to avoid interfering with the normal traffic features.

[0177] Through this mechanism of dynamically adjusting feature weights, the system can enhance the detection sensitivity to the attack features with higher variability, thereby improving the response ability to new attacks. This dynamic adjustment process does not require retraining the model, but only instantaneously updates the parameters of the weights of the time-series correlation features, enabling the system to quickly adapt to and identify potential new threats. After this weight adjustment is completed, the new weights will be recorded and stored for reference in subsequent detections. This mechanism ensures that the network security defense system can respond in a timely manner when facing new attack features, thereby enhancing the overall security and defense effect of the system.

[0178] Furthermore, the dynamically adjusting the feature weights of the time-series correlation features based on the determined attack variability further includes:

[0179] The adjusted feature weight W is determined by formula (5):

[0180] W=W0×(1+α0·log(1+β0·(DD threshold ) p0 ))(5)

[0181] Wherein, W0 represents the initial feature weight of the time series correlation feature; α0 is the adjustment coefficient used to control the increase of the feature weight; β0 is the nonlinear adjustment parameter; D threshold is the preset threshold of attack variability; p0 is the exponential parameter.

[0182] Wherein, W represents the adjusted feature weight, which is the weight value used by the model to analyze the time series correlation feature. The larger the weight, the more attention the model pays to the feature. This adjustment can enhance the model's sensitivity to high-variability features, allowing it to prioritize identifying attack features with higher risks in subsequent detections.

[0183] W0 represents the initial feature weight of the time series correlation feature. When the feature is first detected, the system sets the initial weight value for it. Usually, this value can be determined based on the importance of the feature in historical data or the system default value. For example, if the feature is a common attack type in the system, the initial weight can be set to a higher value to ensure its high priority in the initial analysis.

[0184] α0 is an adjustment coefficient used to control the increase in feature weights. This coefficient determines the degree of influence of attack variability on feature weights. A larger α0 will cause feature weights to grow rapidly as variability increases, causing the model to give greater weight to high-variability attack features. For example, if α0 is set to 2, when attack variability increases, the increase in weight will be more significant than when α0 is set to 1.

[0185] β0 is a nonlinear adjustment parameter that controls the degree of nonlinearity in the weight adjustment process. By adjusting β0, the system can control the nonlinear effect of weight gain. A larger β0 will amplify the impact of attack variation on the weight, making the weight change more drastic; on the contrary, a smaller β0 will smooth the weight adjustment and make it more stable.

[0186] D threshold is the preset threshold of attack variability, which is used to distinguish high-variability and low-variability attack features. threshold The logarithmic term in the formula will only make a significant adjustment to the weight when the attack threshold is set. This threshold can be set based on historical attack data or system configuration. For example, if D thresholdSet to 1.5, then only when D is greater than 1.5 will the system significantly increase the weight of the feature in response to possible new attack behaviors.

[0187] p0 is an exponential parameter used to control the non - linearity of weight adjustment. By taking the power operation on the difference of (D - D threshold ), p0 can affect the increase amplitude of the attack mutation degree on the weight adjustment. For example, when p0 is 2, the increase in the attack mutation degree will cause the weight to increase exponentially, thus giving stronger emphasis to the attack features with high mutation degrees.

[0188] Formula (5) combines the initial weight W0 with the adjustment factor, so that in the case of a higher attack mutation degree, the system can assign higher weights to the temporal correlation features, so as to give priority to these high - risk features in subsequent analysis. This weight adjustment method ensures the flexibility of the system, enabling it to quickly adapt to and respond to changes in attack features, and effectively improving the sensitivity and accuracy of detecting new attacks.

[0189] Step S105: Input the feature weights of the dynamically adjusted temporal correlation features and the temporal correlation features of the real - time network traffic into the deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits, and updates the deep reinforcement learning detection model.

[0190] In step S105, first, after completing the dynamic adjustment of the feature weights, the updated feature weights are passed together with the temporal correlation features of the real - time network traffic extracted from the current time window to the input end of the deep reinforcement learning detection model. This input contains the enhanced feature weights for new attacks and the correlation features that capture the traffic temporal relationship, and can more effectively reflect the abnormal behaviors in the real - time network environment.

[0191] After receiving this input, the deep reinforcement learning detection model will use this information for feature analysis and attack recognition. The model will evaluate the combined effect of the feature weights and the temporal correlation features to identify potential attack features in the traffic and score different feature patterns. The deep reinforcement learning model optimizes these scores through a preset objective function to maximize network security benefits. This objective function can quantify the success or failure of the model in detecting attacks through a reward mechanism, so that the model can continuously adjust and optimize its parameters based on the reward feedback, thereby improving the overall detection performance.

[0192] In this process, the model gradually learns to identify the association between specific features and attack behaviors through multiple rounds of training and feedback loops, while dynamically adjusting the detection strategy. Through this optimization process, the deep reinforcement learning model can more sensitively capture new attack features and adapt to changes in the network environment. The optimized deep reinforcement learning detection model will save the updated parameters so that they can be directly used in subsequent real-time detection, ensuring that the model can respond to changes in network traffic in real time and detect potential security threats.

[0193] This optimization process not only improves the accuracy of the detection model, but also enhances its adaptability to new attack patterns, enabling the network security system to maintain a high level of defense effectiveness when facing complex and changing threats.

[0194] Furthermore, the objective function of the deep reinforcement learning detection model is defined by formula (6):

[0195]

[0196] Among them, J(θ) represents the objective function of the deep reinforcement learning model, which represents the expected total benefit of the deep reinforcement learning model under the parameter θ; this objective function enables the model to effectively respond to network threats in the long term by maximizing the cumulative benefits. Represents the mathematical expectation of all states and actions. By calculating the expected value, the model can evaluate the long-term effect of the strategy under different combinations of states and actions. T is the maximum number of time steps in the training cycle, that is, the total number of steps of the model in one training, which is used to limit the time range of training. t represents the current time step of training, which is used to mark the specific time position of the model in training. γ is a discount factor used to balance immediate rewards and long-term benefits; the value of the discount factor is between 0 and 1. The larger the value, the more the model attaches importance to future rewards; the smaller the value, the more the model is biased towards current immediate rewards. R(s t , a t ) means in state s t Execute action a t The instant reward obtained; λ is the regularization parameter, which is used to control the weights of different reward items, determined by experimental data or set according to expert knowledge; K is the total number of regularization features, indicating the number of all feature items involved in the detection process. k is the index of the regularization feature;

[0197] φ k (s t , a t ) is the penalty function, indicating that the model is in state s t Next, perform action a t The penalty value of the kth feature item is used to suppress the high expected value of a specific feature and ensure the stability of the model behavior.

[0198] Penalty function φ k (s t , a t ) is implemented using Equation (7):

[0199] φ k (s t , a t ) = α k log(1 + β k d k (s t , a t ))(7) p1

[0200] where α k is a tuning parameter that controls the weight of the penalty term. By adjusting this value, the system can increase or decrease the penalty for a specific feature term.

[0201] β k is a non - linear tuning parameter used to control the non - linear growth of the penalty term, in order to enhance the penalty for larger deviations.

[0202] p1 is an exponential parameter used to adjust the non - linearity of the penalty term. A larger value of p1 makes the influence of the deviation value more significant.

[0203] d k (s t , a t ) is the deviation between the k - th feature and the expected feature under the current state s t and action a t , and its calculation formula is Equation (8):

[0204]

[0205] where f k (s t , a t ) represents the k - th feature value when executing action a t in state s t , and actually reflects the performance of this feature when the current model executes the defense strategy.

[0206] is the target value of the k - th feature, representing the feature level under normal or expected behavior. If the current feature value f k (s t , a t ) deviates from , then a deviation d k (s t , a t ) will be generated and controlled by the penalty function. ​

[0207] The immediate reward function is implemented using formula (9):

[0208] R(s t , a t ) = κ·ρ·tanh(γ1·(e t -e min )) - ν·Ω(a t )(9)

[0209] where R(s t , a t ) is the immediate reward function, used to evaluate the immediate benefit of taking action a t under state s t ;

[0210] κ is the reward coefficient for successful defense, representing the reward intensity obtained by the model when correctly detecting an attack or effectively defending. The larger this value, the higher the reward for successful defense by the model.

[0211] ρ is the threat level coefficient, dynamically determined according to the currently detected attack type and threat level. If the detected attack type has a high threat level, the value of this coefficient is higher, thereby increasing the immediate reward.

[0212] γ1 is the amplification parameter, used to control the sensitivity of the reward to the deviation of the defense effect. A larger value of γ1 will amplify the impact of the effect score on the immediate reward.

[0213] e t represents the defense effect score of the current action a t under state s t , indicating the effectiveness of the defense measures taken by the model under the current conditions.

[0214] The defense effect score e t reflects the defense effect after the model executes action a t under state s t , that is, the effectiveness of the defense measures taken in detecting and preventing network attacks. To determine this score, usually the following several indicators are comprehensively considered, and based on the specific performance of these indicators, an overall score is calculated. For example:

[0215] Successful interception rate: If in a certain state s t , the system detects a potential attack and takes corresponding defense action a t , the successful interception rate indicates how many attacks this action has successfully blocked. Suppose that after executing a t , 95% of the attacks are successfully intercepted within a certain period of time, then this indicator will contribute a relatively high score. For example, a weight of 90 points can be assigned to this successful interception rate.

[0216] False Alarm Rate: The false alarm rate of a defense measure is also an important factor in evaluating its effectiveness. Suppose the system performs action a t and the false alarm rate (the frequency of misclassifying normal traffic as an attack) drops to 5%. A lower false alarm rate indicates a higher accuracy of the defense measure, which can improve the score of e t . For example, if the false alarm rate is very low, a weight of 80 points can be assigned to this part.

[0217] Response Speed: The response speed of the system in taking defense actions after detecting abnormal or attack behavior also affects the defense effectiveness score. For example, suppose the system responds and executes the corresponding defense measure a t within 1 second after detecting a threat. This indicates that the system responds quickly and effectively reduces the damage that an attack may cause. For a rapid response speed, a weight of 85 points can be contributed to it.

[0218] Resource Occupation: An effective defense measure usually needs to balance the control of resource consumption. For example, suppose after executing a t , it only occupies 20% of the resource load. This means that the defense measure provides high security without affecting the overall system performance, thus increasing the score. For example, low resource occupation can add 75 points to e t .

[0219] Attack Prevention Effect: If after executing action a t , the attack is effectively prevented and the system returns to the normal state, then this indicator gets a high score. For example, if the system successfully returns to a stable network state after detection and defense, it can contribute 80 points to the defense effectiveness score.

[0220] Based on these indicators, the scores of each item are aggregated or weighted averaged to obtain the defense effectiveness score e t . In this example, if the successful interception rate is 90 points, the false alarm rate is 80 points, the response speed is 85 points, the resource occupation is 75 points, and the attack prevention effect is 80 points, then the overall defense effectiveness score e t can be expressed as the average of these scores:

[0221]

[0222] In this example, the score e t = 82 indicates that the current defense measure a t has high effectiveness in state s t , thus having a positive impact on the immediate reward. This score helps the model adjust future strategies according to the actual defense effectiveness and optimize its response ability to new threats.

[0223] e minIt serves as the lowest effect benchmark, which is used to ensure that the reward function only gives significant rewards after the effect exceeds a certain level, thereby suppressing low-effect defense behaviors. It can be obtained through experimental data or directly set using expert knowledge.

[0224] ν represents the defense cost adjustment coefficient, which is used to control the penalty intensity of the action cost on the reward; a higher value of v will increase the penalty effect of the defense cost in the immediate reward. This coefficient can be obtained through experimental data or directly set using expert knowledge.

[0225] Ω(a t ) is the execution cost function of the defense strategy a t and is determined according to the occupancy of system resources by the currently selected defense strategy.

[0226] The execution cost function Ω(a t ) is used to evaluate the consumption of system resources by the defense strategy a t during the execution process. Its specific implementation usually involves the quantification of the following main resource indicators:

[0227] Computing resource consumption: including the occupancy rates of CPU and memory. When the defense strategy a t is executed, the system monitors the occupancy of CPU and memory, and assigns a higher cost value to the strategy with high consumption. For example, if a t occupies more than 50% of the CPU, the cost function will increase the corresponding value.

[0228] Network bandwidth usage: Some defense strategies may require more network bandwidth to transmit detection data or perform isolation operations. The system calculates its cost according to the bandwidth usage of the strategy a t . For example, a strategy that uses a large amount of bandwidth will increase the execution cost.

[0229] Storage resource occupancy: Some defense operations need to temporarily store a large amount of detection data or logs. The system calculates the cost according to the storage usage. For example, if the storage consumption exceeds the preset threshold, the cost will increase.

[0230] Latency and response time: If the defense strategy a t causes an increase in system latency or affects the response time, the system will regard this increased latency as part of the cost. A strategy with a longer latency corresponds to a higher execution cost.

[0231] By weighted aggregation or comprehensive evaluation of the above resource indicators, the system obtains the specific value of the execution cost function Ω(a t ). In this way, the model can balance the defense effect and resource consumption to ensure that the defense strategy is both effective and economical.

[0232] Furthermore, the deep reinforcement learning detection model includes a first input layer, a feature extraction layer, a policy generation layer, and an output layer, and operates through the following steps:

[0233] The first input layer is used to receive the feature weights of the dynamically adjusted temporal correlation features and the temporal correlation features of the real-time network traffic;

[0234] The feature extraction layer includes multiple hidden layers, processes the temporal correlation features through a recurrent neural network, and extracts the temporal patterns and potential patterns of the real-time network traffic;

[0235] The policy generation layer uses the temporal patterns and feature weights obtained from the feature extraction layer to generate corresponding defense policies through a deep reinforcement learning algorithm; the policy generation layer scores each defense policy based on a deep Q network and selects the policy with the highest score for application; the policy generation layer iteratively optimizes the effect of the selected policy during the training process to maximize the network security benefit;

[0236] The output layer generates network security defense policies, including traffic restriction, access control, or data isolation.

[0237] The deep reinforcement learning detection model includes a first input layer, a feature extraction layer, a policy generation layer, and an output layer, and realizes dynamic defense against real-time network attacks through these layers. The whole process starts from inputting real-time features and sequentially passes through feature extraction, policy generation, and defense policy output to ensure that the network security defense policy can adapt to the continuously changing attack features.

[0238] At the beginning of the operation, the first input layer receives the feature weights of the dynamically adjusted temporal correlation features and the temporal correlation features of the real-time network traffic. The feature weights reflect the importance of the temporal correlation features, while the temporal correlation features capture the time-varying patterns of the network traffic, ensuring that the system not only considers the current traffic features during the defense process but also pays attention to the dynamic feature evolution of the attacks, providing basic information for subsequent feature extraction and policy generation.

[0239] After entering the feature extraction layer, the input temporal correlation features are deeply analyzed through multiple hidden layers. This layer uses a recurrent neural network (RNN), which is specifically used for processing and analyzing temporal data. Through layer-by-layer processing, the RNN can extract the temporal patterns and potential patterns that may be contained in the real-time network traffic. This structure can capture the long-term dependencies of attack behaviors, thereby identifying attack features that are not obvious in the short term, and then improving the model's ability to identify continuous attacks. For example, by analyzing features such as the packet frequency, delay, and traffic peak changes over a period of time, the RNN may identify a gradually increasing packet traffic pattern, thus revealing a potential DDoS attack.

[0240] After obtaining the temporal patterns and feature weights passed from the feature extraction layer, the policy generation layer generates corresponding defense policies based on the deep reinforcement learning algorithm. In this process, the policy generation layer uses the Deep Q-Network (DQN) algorithm to evaluate different defense policies. DQN helps the system find the optimal solution among multiple defense options by assigning scores to each policy. The score of each defense policy is calculated based on its effectiveness in the current attack environment, and the policy with the highest score is preferentially selected. Through this scoring and selection mechanism, the policy generation layer can quickly identify and select the best defense policy. In addition, during the training process of the model, the policy generation layer iteratively optimizes its scoring mechanism and policy selection logic to maximize the network security benefits. This means that in the long-term training, the system can gradually improve the defense effect, making the scoring mechanism and policy selection more accurate.

[0241] After the policy generation layer selects the best defense policy, the output layer generates specific network security defense operations. This output layer implements the generated defense policies in forms including traffic restriction, access control, and data isolation. For example, if the output layer generates a traffic restriction policy, the system will reduce the bandwidth allocation for suspicious traffic to prevent the spread of potential attack activities. If an access control policy is generated, the access rights of the attack source to critical network areas will be restricted; if it is a data isolation policy, the system will isolate sensitive data areas to protect the network from possible intrusion threats. Through this multi-level defense output, the system can dynamically respond to different types of attack features to ensure the security and stability of the network.

[0242] Overall, this deep reinforcement learning detection model enables the network security system to have dynamic defense capabilities when dealing with complex attacks and effectively resist changing network threats through the feature reception of the input layer, pattern analysis of the feature extraction layer, scoring and selection mechanism of the policy generation layer, and defense policy output of the output layer.

[0243] Step S106: Input the temporal correlation features of the real-time network traffic into the updated deep reinforcement learning detection model to obtain network security defense policies, where the network security defense policies include traffic restriction, access control, and data isolation.

[0244] First, after the model parameters are optimized, the temporal correlation features of the real-time network traffic extracted within the current time window are input into the updated deep reinforcement learning detection model. At this time, the model already has adaptability to new attack features and the current network environment, so it can more accurately analyze the temporal correlation features and identify potential abnormal behaviors and attack patterns from them.

[0245] After receiving these time - series related features, the updated deep reinforcement learning detection model analyzes the features and generates logic based on specific strategies to formulate the best network security defense strategy. This defense strategy provides different countermeasures according to the types of detected attack features and the threat levels of attack behaviors. The generation process of the defense strategy is based on the multi - layer network structure and reinforcement learning mechanism of the model, and can be dynamically adjusted in different situations to ensure that the system can react in a timely manner and take corresponding defense measures when facing various types of network attacks.

[0246] This defense strategy usually includes three main measures: traffic restriction, access control, and data isolation. The traffic restriction strategy prevents the excessive consumption of network resources by attack traffic by controlling the transmission rate and total amount of malicious traffic. The access control strategy adjusts the access permissions for specific sources according to the abnormally detected behaviors in real - time, thus effectively preventing attackers from further accessing key network resources. The data isolation strategy is activated when a serious threat is detected, and isolates the attacked area from other normal areas to prevent the spread of the attack to a wider network area.

[0247] The whole process enables the system to dynamically generate and deploy network security defense strategies according to the changes in real - time traffic features and the results of model analysis, so as to make an effective response at the first time of an attack. Finally, this defense strategy will be applied to the network environment in real - time to ensure the continuous security and stability of the network.

[0248] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention provides a network security defense method and system based on incremental network attack analysis and learning. By matching the time - series related features with the initial network attack feature library and calculating the attack mutation degree, it can dynamically identify and adapt to new attack features, quickly detect and respond to unknown attacks, and improve the flexibility and adaptability of network defense. Based on the deep reinforcement learning model, by inputting the dynamically adjusted weights and time - series related features of the time - series related features, the optimization of network security benefits is realized, and the defense strategy is updated in real - time, so that the defense measures can be automatically adjusted with the changes of attack features, effectively improving the accuracy and timeliness of the defense effect. Using the incremental learning mechanism, after new attack features are identified and the weights are adjusted, the deep reinforcement learning detection model is iteratively optimized. Compared with the traditional model that needs to be retrained, incremental learning improves the adaptability of the model, enabling it to always maintain high - efficient detection performance in the constantly changing network environment. By using the deep auto - encoder to reduce the dimensionality of the traffic feature vector and construct the initial network attack feature library, the core feature information of network traffic is retained. Combined with the dynamically adjusted feature weights, the misjudgment of normal traffic is effectively reduced, the false alarm rate is lowered, and the accuracy and robustness of detection are improved.

[0249] The present invention can be a system, a method, and / or a computer program product. The present invention also discloses a network security defense system based on the aforementioned network security defense method for incremental network attack analysis and learning, further comprising:

[0250] An initial network attack feature library construction module, configured to collect initial network traffic data, extract traffic feature vectors; use a deep autoencoder to perform dimensionality reduction processing on the traffic feature vectors to obtain compressed feature vectors; construct an initial network attack feature library based on the compressed feature vectors;

[0251] A time-series correlation feature matching module, configured to collect real-time network traffic data, and perform segmentation processing on the real-time network traffic data based on a preset sliding time window; extract time-series correlation features from the segmented real-time traffic data; match the time-series correlation features with the features in the initial network attack feature library to identify potential attack behaviors or abnormal behaviors;

[0252] An attack mutation degree calculation module, configured to, when the time-series correlation features do not match the features in the initial network attack feature library, mark the time-series correlation features as candidate new attack features; calculate the Mahalanobis distance between the time-series correlation features and the attack features stored in the initial network attack feature library to determine the attack mutation degree of the new attack features;

[0253] A feature weight adjustment module, configured to dynamically adjust the feature weights of the time-series correlation features based on the determined attack mutation degree;

[0254] A deep reinforcement learning detection module, configured to input the feature weights of the dynamically adjusted time-series correlation features and the time-series correlation features of real-time network traffic into a deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits, and updates the deep reinforcement learning detection model;

[0255] A network security defense strategy acquisition module, configured to input the time-series correlation features of real-time network traffic into the updated deep reinforcement learning detection model to obtain a network security defense strategy, where the network security defense strategy includes traffic restriction, access control, and data isolation.

[0256] In the spirit of the present invention, those skilled in the art can easily conceive that a computer program product can be obtained based on the aforementioned network security defense method based on incremental network attack analysis and learning. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure. That is, the present application also includes a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the aforementioned network security defense method based on incremental network attack analysis and learning.

[0257] A computer-readable storage medium can be a tangible device that can retain and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example - but not limited to - an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of a computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the above. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0258] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0259] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0260] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A network security defense method based on incremental network attack analysis and learning, characterized in that: The following steps are involved: Collect initial network traffic data and extract traffic feature vectors; use The deep autoencoder performs dimensionality reduction processing on the traffic feature vector to obtain a compressed feature vector; Constructing an initial network attack feature library based on the compressed feature vector; Collecting real-time network traffic data, and segmenting the real-time network traffic data based on a preset sliding time window; extracting time series correlation features from the segmented real-time traffic data; Matching the time series correlation features with features in an initial network attack feature library to identify potential attack behaviors or abnormal behaviors; When the time series correlation feature does not match the feature in the initial network attack feature library, marking the time series correlation feature as a candidate new attack feature; Calculating the Mahalanobis distance between the time series correlation feature and the attack feature stored in the initial network attack feature library to determine the attack variability of the new attack feature; Based on the determined attack variability, dynamically adjusting the feature weight of the temporal correlation feature; Inputting the dynamically adjusted feature weights of the timing correlation features and the timing correlation features of the real-time network traffic into a deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits, and updates the deep reinforcement learning detection model; The time series correlation features of real-time network traffic are input into the updated deep reinforcement learning detection model to obtain a network security defense strategy, which includes traffic restriction, access control and data isolation.

2. The network security defense method based on incremental network attack analysis and learning according to claim 1 is characterized in that: The collecting of initial network traffic data and extracting traffic feature vectors further includes: Collecting initial network traffic data from multiple network ingress nodes and egress nodes, covering different traffic sources and paths in the network, the collected initial network traffic data including packet header information, packet size, traffic direction and timestamp; Preprocess the collected traffic data to remove redundant data, noise data, and irrelevant traffic data, retain key data that may contain attack features, and generate preprocessed traffic data; The preprocessed traffic data is segmented, and multi-dimensional feature information is extracted from each segment of data to form a traffic feature vector.

3. The network security defense method based on incremental network attack analysis and learning according to claim 2 is characterized in that: The step of constructing an initial network attack feature library based on the compressed feature vector further includes: Clustering the compressed feature vectors according to the similarity of traffic features, grouping vectors with similar features, thereby revealing potential attack patterns; Analyze the feature vectors in each cluster, extract representative features as feature identifiers of each attack mode, and record the feature identifiers in the initial network attack feature library; By counting the occurrence frequency and distribution range of each feature identifier, weight information is added to each attack mode in the feature library. The weight information is used to provide a priority reference in subsequent real-time traffic matching; The constructed initial network attack feature library is stored.

4. The network security defense method based on incremental network attack analysis and learning according to claim 3 is characterized in that: The collecting of real-time network traffic data and segmenting the real-time network traffic data based on a preset sliding time window further includes: Collect real-time network traffic data and pass the collected real-time traffic data to the sliding window processing module; In the sliding window processing module, the real-time network traffic data is segmented according to the preset sliding time window. Each sliding window represents a fixed time interval. The traffic data is intercepted and stored within this time interval to form the traffic data segment corresponding to the window. At the end of the window, the data segments in the current time window are encapsulated as a whole, and the start and end times of the window are recorded; At the beginning of the next time step, the sliding window automatically moves forward one step, adding the newly entered real-time traffic data to the new window and forming a new traffic data segment, thereby iteratively realizing continuous segmentation processing of real-time network traffic.

5. The network security defense method based on incremental network attack analysis and learning according to claim 4 is characterized in that: The step of extracting time series correlation features from the segmented real-time traffic data further includes: Obtain real-time traffic data segments after segmentation processing in sliding time windows; Perform time series feature analysis on each data segment, extract basic time series features including data packet transmission rate, data packet time interval, traffic direction and traffic peak value, and obtain dynamic traffic change characteristics within the data segment; Analyze the correlation between continuous data segments, identify cross-time characteristic patterns including periodic traffic changes, burst traffic, and abnormal delays, and obtain traffic change trend characteristics; The extracted basic time series features and traffic change trend features are combined to form a time series correlation feature vector for subsequent attack detection and feature matching.

6. The network security defense method based on incremental network attack analysis and learning according to claim 5 is characterized in that: The deep autoencoder includes an input layer, an encoding layer, a bottleneck layer, a decoding layer and a feature selection layer; The input layer is used to receive the traffic feature vector extracted from the initial network traffic data; perform standardization on the received data to obtain a standardized feature vector; The encoding layer is used to receive the standardized feature vector from the input layer and extract features layer by layer by stacking multiple layers of convolutional networks; Each convolution layer in the coding layer applies a convolution kernel of a specific size to capture high-order feature relationships in the feature vector through a nonlinear activation function; wherein the size of the convolution kernel is adaptively adjusted at each layer according to the distribution of input features and the feature dimension of network traffic; the output of the coding layer is a low-dimensional coded feature representation; The bottleneck layer is used for the low-dimensional coding feature representation, and the feature dimension is further compressed through a sparse activation function to generate a sparse compressed feature vector; The decoding layer is used to receive the sparse compressed feature vector from the bottleneck layer, and decode and reconstruct the compressed features layer by layer through a multi-layer deconvolution network symmetrical to the encoding layer to obtain a high-dimensional feature representation; the output of the decoding layer is the reconstructed feature vector, which is used to calculate the reconstruction error with the original input during the training process to optimize the parameters of each layer of the autoencoder; The feature selection layer is used to receive the sparse compressed feature vector from the bottleneck layer, weight each feature dimension through the attention mechanism, further screen out important features, and obtain a compressed feature vector.

7. The network security defense method based on incremental network attack analysis and learning according to claim 6 is characterized in that: The calculating of the Mahalanobis distance between the time series correlation feature and the attack feature stored in the initial network attack feature library to determine the attack variability of the new attack feature further includes: The attack variability of the new attack feature is calculated using formula (1): Wherein, D represents the attack variability of the time series correlation feature; N represents the number of attack features stored in the initial network attack feature library; d i represents the Mahalanobis distance between the temporal correlation feature and the i-th known attack feature; σ i is the standard deviation of the characteristic value of the i-th known attack feature on different sample data in the initial network attack feature library; β i is the risk factor of the i-th known attack feature, determined according to the severity of the attack in historical data; det(Σ i ) is the determinant of the covariance matrix of the i-th known attack feature; w i represents the weight coefficient associated with the i-th known attack feature; f(d i ) is a regularization factor; The Mahalanobis distance d i Calculated by formula (2): Wherein, x represents the feature vector of the time series correlation feature; μ i represents the mean vector of the i-th known attack feature; represents the inverse matrix of the covariance matrix of the i-th known attack feature; The weight coefficient w i Calculated by formula (3): Among them, α is the adjustment parameter; r i is the historical occurrence frequency of the i-th known attack feature; r threshold It is the preset frequency interval value, used to distinguish high-frequency and low-frequency attack characteristics; The regularization factor f(d i ) is calculated by formula (4): Among them, γ is a regulation parameter used to control the amplification effect of the regularization factor; p is an exponential parameter used to adjust the nonlinear degree of regularization.

8. The network security defense method based on incremental network attack analysis and learning according to claim 7 is characterized in that: The dynamically adjusting the feature weight of the time series correlation feature based on the determined attack variability further includes: The adjusted feature weight W is determined by formula (5): W=W0×(1+α0·log(1+β0·(D-D threshold ) p0 ))(5) Wherein, W0 represents the initial feature weight of the time series correlation feature; α0 is the adjustment coefficient used to control the increase of the feature weight; β0 is the nonlinear adjustment parameter; D threshold is the preset threshold of attack variability; p0 is the exponential parameter.

9. The network security defense method based on incremental network attack analysis and learning according to claim 8 is characterized in that: The objective function of the deep reinforcement learning detection model is defined as: Where J(θ) represents the objective function of the deep reinforcement learning model, which is to maximize the cumulative benefit based on the deep reinforcement learning model parameter θ; represents the mathematical expectation of all states and actions; T is the maximum number of time steps in the training cycle; t represents the current time step of training; γ is the discount factor used to balance immediate rewards and long-term benefits; R(s t , a t ) means in state s t Execute action a t The instant reward obtained; λ is the regularization parameter used to control the weights of different reward items; K is the total number of regularized features; k is the index of the regularized feature; φ k (s t , a t ) is the penalty function; The penalty function φ k (s t , a t ) is implemented using formula (7): f k (s t ,a t )=a k log(1+β k d k (s t ,a t ) p1 ) (7) Among them, α k and β k is the adjustment parameter related to the kth feature, which is used to control the weight and nonlinearity of the penalty term; p1 is the exponential parameter; d k (s t , a t ) is the current state s t and action a t The deviation between the kth feature and the expected feature is calculated using formula (8): Among them, f k (s t , a t ) means in state s t Execute action a t The kth eigenvalue when ; is the target value of the kth feature, which represents the feature level under normal or expected behavior mode; The immediate reward function is implemented using formula (9): R(s t ,are t )=κ·ρ·tanh(γ1·(e t -that min ))–v·Ω(a t ) (9) Among them, R(s t , a t ) is the immediate reward function, used to evaluate the state s t Take action a t The instant benefit; κ is the reward coefficient for successful defense, which represents the reward intensity obtained by the model when it correctly detects an attack or effectively defends; ρ is the threat level coefficient, which is dynamically determined according to the currently detected attack type and threat level; γ1 is the amplification parameter, which is used to control the sensitivity of the reward to the deviation of the defense effect; e t Indicates the current action a t In status t The defensive effect score under e min is the minimum effect benchmark, which is used to ensure that the reward function only gives significant rewards after the effect exceeds a certain level, thereby suppressing low-effect defense behaviors; v represents the defense cost adjustment coefficient, which is used to control the penalty intensity of the action cost on the reward; Ω(a t ) is the defense strategy a t The execution cost function is determined according to the amount of system resources occupied by the currently selected defense strategy.

10. The network security defense method based on incremental network attack analysis and learning according to claim 9 is characterized in that: The deep reinforcement learning detection model includes a first input layer, a feature extraction layer, a strategy generation layer and an output layer, and operates through the following steps: The first input layer is used to receive the dynamically adjusted feature weights of the timing correlation features and the timing correlation features of the real-time network traffic; The feature extraction layer includes multiple hidden layers, and processes the time series correlation features through a recursive neural network to extract the time series pattern and potential pattern of the real-time network traffic; The strategy generation layer uses the obtained timing patterns and feature weights to generate corresponding defense strategies through a deep reinforcement learning algorithm; the strategy generation layer scores each defense strategy based on a deep Q network and selects the strategy with the highest score for application; The strategy generation layer iteratively optimizes the effect of the selected strategy during the training process to maximize network security benefits; The output layer generates network security defense strategies, including traffic restriction, access control or data isolation.

11. A network security defense system based on incremental network attack analysis and learning, characterized in that: Further including: The initial network attack feature library construction module is used to collect initial network traffic data and extract traffic feature vectors; use The deep autoencoder performs dimensionality reduction processing on the traffic feature vector to obtain a compressed feature vector; Constructing an initial network attack feature library based on the compressed feature vector; A time series correlation feature matching module is used to collect real-time network traffic data and segment the real-time network traffic data based on a preset sliding time window; and extract time series correlation features from the segmented real-time traffic data; Matching the time series correlation features with features in an initial network attack feature library to identify potential attack behaviors or abnormal behaviors; an attack variation calculation module, configured to mark the time series correlation feature as a candidate new attack feature when the time series correlation feature does not match the feature in the initial network attack feature library; Calculating the Mahalanobis distance between the time series correlation feature and the attack feature stored in the initial network attack feature library to determine the attack variability of the new attack feature; A feature weight adjustment module, used for dynamically adjusting the feature weight of the time series correlation feature based on the determined attack variability; A deep reinforcement learning detection module, used for inputting the dynamically adjusted feature weights of the timing correlation features and the timing correlation features of the real-time network traffic into a deep reinforcement learning detection model; the deep reinforcement learning detection model optimizes its parameters through an objective function that maximizes network security benefits, and updates the deep reinforcement learning detection model; The network security defense strategy acquisition module is used to input the time-series correlation features of real-time network traffic into the updated deep reinforcement learning detection model to obtain a network security defense strategy, which includes traffic restriction, access control and data isolation.

12. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the network security defense method based on incremental network attack analysis and learning according to any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the network security defense method based on incremental network attack analysis and learning described in any one of claims 1-10 are implemented.

Citation Information

Patent Citations

  • Attack chain information analysis method and system

    CN111917793A

  • Big language model fusion method and device based on multi-modal data and medium

    CN118965283A

  • Method and system for detecting and analyzing malicious network traffic

    CN119094243A

  • Deep learning based intrusion prediction model

    IN202021056713A

  • System and process for detecting anomalous network traffic

    US20100138919A1

Cited By

  • Power system data security analysis method, device, equipment and medium

    CN120281570A

  • A method, device, equipment and medium for power system data security analysis

    CN120281570B

  • Wind speed monitoring identification method and system based on pattern matching

    CN120492951A

  • Method for determining security state of gateway request behavior, computing device and storage medium

    CN120528705A

  • Method, computing device, and storage medium for determining a security status of a gateway request behavior

    CN120528705B