Network attack mitigation method and apparatus, and electronic device
By performing feature extraction and encoding mapping processing on traffic data, and using pre-trained traffic attack detection model, the problem of insufficient efficiency and accuracy of traditional methods in multi-source heterogeneous network data environment is solved, and rapid detection and efficient mitigation of network attacks are achieved.
Patent Information
- Application Number
- CN202510546466.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-25
AI Technical Summary
The existing traditional attack behavior detection methods rely on manual feature extraction and rule matching, making it difficult to adapt to complex and changeable network attacks, especially in multi-source heterogeneous network data environments, and lack of effective mitigation measures.
By obtaining traffic data for feature extraction, using coding mapping processing to generate target vector sequences, and input a pre-trained traffic attack detection model, output detection results and mitigation measures, reduce the dependence of artificial feature engineering, and improve the generalization ability of model.
It realizes rapid detection and efficient mitigation of network attacks, provides targeted strategies, provides network administrators with flexible solutions, and improves the adaptability and accuracy of the model in multi-source heterogeneous network data environment.
Smart Images

Figure CN120378171A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of network security technologies, and in particular, to a network attack mitigation method, apparatus, and electronic device. Background Art
[0002] In the field of network security, traditional attack behavior detection methods are facing unprecedented challenges. These methods rely on manual feature extraction and rule matching. Although they have achieved results in early research, with the increasing complexity of network attack means, the adaptability and efficiency of these methods are gradually insufficient. Especially when facing multi-source heterogeneous network data, traditional methods have significant difficulties in processing data formats, feature scales, etc.
[0003] For existing traditional attack behavior detection and mitigation methods, first of all, these methods mainly rely on manually designed and extracted features, and identify abnormal traffic through predefined rules or pattern matching. This method has some obvious limitations when dealing with network attacks. Summary of the Invention
[0004] In view of this, an object of the present disclosure is to provide a network attack mitigation method, apparatus, and electronic device to solve or partially solve the above problems.
[0005] Based on the above object, a first aspect of the present disclosure provides a network attack mitigation method, the method including:
[0006] Obtain traffic data, perform feature extraction on the traffic data to obtain a target feature set corresponding to the traffic data;
[0007] Perform encoding mapping processing on the target feature set to obtain a target vector sequence;
[0008] Input the target vector sequence into a pre-trained traffic attack detection model, and after being processed by the traffic attack detection model, output a detection result;
[0009] In response to the detection result indicating an attack, determine the target attack type, and determine an attack mitigation measure corresponding to the target attack type according to the target attack type.
[0010] Based on the same inventive concept, a second aspect of the present disclosure proposes a network attack mitigation apparatus, the apparatus including:
[0011] A data acquisition module configured to obtain traffic data, perform feature extraction on the traffic data to obtain a target feature set corresponding to the traffic data;
[0012] An encoding mapping module configured to perform encoding mapping processing on the target feature set to obtain a target vector sequence;
[0013] An attack detection module, configured to input the target vector sequence into a pre-trained traffic attack detection model, process it through the traffic attack detection model, and output a detection result;
[0014] A mitigation measure determination module, configured to, in response to the detection result indicating an attack, determine a target attack type, and determine an attack mitigation measure corresponding to the target attack type according to the target attack type.
[0015] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the network attack mitigation method as described above is implemented.
[0016] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the network attack mitigation method as described above.
[0017] As can be seen from the above, the present disclosure provides a network attack mitigation method, device, and electronic device. Traffic data is acquired, feature extraction is performed on the traffic data to obtain a target feature set corresponding to the traffic data, reducing the dependence on manual feature engineering and significantly improving the model generalization ability. Encoding mapping processing is performed on the target feature set to obtain a target vector sequence, and the target vector sequence is input into a pre-trained traffic attack detection model. After being processed by the traffic attack detection model, a detection result is output. Detection of traffic attacks is achieved through the trained traffic attack detection model, and the detection is more accurate. In response to the detection result indicating an attack, a target attack type is determined, and an attack mitigation measure corresponding to the target attack type is determined according to the target attack type. By determining the corresponding attack mitigation measure through the target attack type, a targeted strategy can be quickly generated according to the real-time attack situation, providing an efficient and flexible solution for network administrators. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of the network attack mitigation method for the embodiments of the present disclosure;
[0020] Figure 2 The structural block diagram of the network attack mitigation device according to an embodiment of the present disclosure;
[0021] Figure 3 The structural schematic diagram of the electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0023] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second" and similar words used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0024] The following are the noun explanations involved in the present disclosure:
[0025] DDoS: Distributed Denial of Service (DDoS), that is, using a large number of legitimate distributed servers to send requests to a target, resulting in normal legitimate users being unable to obtain services.
[0026] GBT: Gradient Boosting Tree, which uses the negative gradient of the loss function at the value of the current model as an approximation of the residual. Essentially, it performs a first-order Taylor expansion on the loss function to fit a regression tree.
[0027] SVM: Support Vector Machine (SVM) is a classic supervised learning algorithm used to solve binary classification and multi-classification problems. Its core idea is to find an optimal hyperplane in the feature space for classification, and the margin is the largest.
[0028] UDP: User Datagram Protocol (UDP) is a protocol that operates at the transport layer in the OSI (Open Systems Interconnection) model. It uses IP as the underlying protocol and provides a protocol mechanism for applications to send messages to other programs with the least amount of protocol overhead.
[0029] TCP: Transmission Control Protocol (TCP) is a connection-oriented, reliable, byte-stream-based transport layer communication protocol defined by RFC 793 of the IETF.
[0030] ACL: Access Control List (ACL) consists of multiple "deny|permit" statements, each of which is a rule used to control the ingress and egress of network traffic. ACL is one of the important components of a firewall, and it filters data packets based on certain rules.
[0031] LoRA: LoRA model (Low-Rank Adaptation of Large Language Models) is a low-rank adaptation technique for fine-tuning large language models.
[0032] In the field of network security, traditional attack behavior detection methods are facing unprecedented challenges. These methods rely on manual feature extraction and rule matching, which were effective in early research but are gradually showing insufficient adaptability and efficiency as network attack methods become increasingly complex. Especially when dealing with multi-source heterogeneous network data, traditional methods have significant difficulties in handling data formats, feature scales, etc. In recent years, with the rapid development of machine learning and deep learning, more and more research has started to adopt automated feature learning and model optimization strategies to improve the accuracy and efficiency of attack behavior detection.
[0033] The prior art includes an attack behavior detection model called HAGRU (Hierarchical Attention Gated Recurrent Unit), which uses its hierarchical attention mechanism to integrate multi-level feature information to enhance its ability to detect attack behaviors. This model optimizes the usage efficiency of computing resources by focusing on attack behaviors in the data stream. Although HAGRU is effective in detecting attack behaviors, this model needs to rely on a dataset with manually screened and designed features for training, which is not only time-consuming and laborious but also may not guarantee the universality and effectiveness of the selected features for various types of attacks. In addition, manual feature engineering may limit the model's ability to identify new types of attacks because these new attacks may not be included in the feature set used for training. This dependence on manual feature selection also increases the difficulty of model update and maintenance. When network attack strategies change, it may be necessary to re-perform feature selection and model training. Therefore, the HAGRU model has limitations in terms of automation and adaptability, which may affect its long-term effectiveness in a rapidly evolving network threat environment. There is also an attack behavior detection model called MFFusion (Multi-Level Feature Fusion Model) based on deep learning proposed by Kunda Lin et al. This model extracts effective information from multiple perspectives by combining the time, byte, and statistical features of the data to obtain more efficient and robust detection performance. The MFFusion model adopts an adaptive balance training method and designs a new loss function - the attention loss function to address the data imbalance problem and improve the model performance. The MFFusion model has achieved excellent performance in network attack behavior detection, but it has not fully addressed the challenges of multi-source heterogeneous network data, which may affect its adaptability and accuracy in diverse network environments and devices. Multi-source heterogeneous network data usually involves data from different vendors, different protocols, and different formats. The integration and analysis of this data are crucial for improving the effectiveness of network security protection. Although the MFFusion model extracts feature information at different levels through the multi-level feature fusion method, it may face problems such as inconsistent data formats, non-uniform standards, and difficulties in integration when processing multiple heterogeneous source network data. These problems may limit the performance of the model in complex network environments, especially when dealing with continuously changing diverse network traffic. There is also an attack behavior detection model called the SSL / TLS (Secure Sockets Layer / Transport Layer Security Protocol) based on graph neural network proposed by Tang Ying et al., aiming to achieve accurate detection of encrypted attack behaviors. This model addresses the problem that traditional machine learning methods overly rely on expert experience. By analyzing SSL / TLS encrypted sessions, it uses a graph structure to represent the traffic session interaction information and transforms the attack behavior detection problem into a graph classification problem. The proposed model is based on a hierarchical graph pooling architecture. Through the aggregation of multi-layer convolutional pooling and combined with the attention mechanism, it fully mines the node features and graph structure information in the graph and realizes an end-to-end attack behavior detection method.Although this model has demonstrated excellent performance in the detection of cyber - attack behaviors, it fails to provide a set of cyber - attack mitigation strategies, which is an obvious shortcoming in cybersecurity practice. In cybersecurity management, detecting attack behaviors is only the first step. More crucially, based on these detection results, formulating and implementing effective mitigation measures to reduce or prevent the potential damage caused by attacks. Although this model effectively identifies cyber - attack behaviors, it lacks the ability to convert these detection results into specific defense actions, failing to achieve a complete closed - loop from threat detection to taking defense measures. How to effectively handle complex network environments and dynamically changing attack patterns remains a difficult point in current research.
[0034] For existing traditional cyber - attack detection methods, first of all, these methods mainly rely on manually designed and extracted features, and identify abnormal traffic through predefined rules or pattern matching. There are some obvious limitations in dealing with cyber - attacks using this method. These limitations include: Manually designed features require a large amount of human and time resources. Moreover, in the face of continuously changing and evolving cyber - attack means, these features are easily rendered inapplicable and difficult to effectively cope with new types of attacks. In addition, the data content generated by multi - source heterogeneous network devices varies greatly and lacks a unified standard, which makes the process of manual feature extraction more complex and inefficient. The data formats, semantics, and granularities of different devices are different, resulting in data being difficult to directly use for analysis and modeling, further limiting the efficiency and accuracy of the detection system. Furthermore, after traditional models detect attack behaviors, they often lack effective mitigation measures and cannot timely prevent the further development of attacks, limiting their value in practical applications. This requires not only advanced technical processing capabilities to deal with continuously changing cyber - attack characteristics but also analysts to have profound cybersecurity expertise. Only in this way can they effectively interpret complex and ever - changing attack patterns and formulate corresponding mitigation measures. For many organizations, meeting these requirements is undoubtedly a major challenge.
[0035] Based on the above description, this embodiment proposes a cyber - attack mitigation method, as Figure 1 shown, the method includes:
[0036] Step 101: Obtain traffic data, extract features from the traffic data to obtain a target feature set corresponding to the traffic data;
[0037] Step 102: Perform encoding and mapping processing on the target feature set to obtain a target vector sequence;
[0038] Step 103: Input the target vector sequence into a pre - trained traffic attack detection model, and after being processed by the traffic attack detection model, output a detection result;
[0039] Step 104, in response to the detection result indicating an attack, determine the target attack type, and determine the attack mitigation measures corresponding to the target attack type according to the target attack type.
[0040] In specific implementation, obtain traffic data, where the traffic data is multi-source heterogeneous network data. Extract features from the traffic data to obtain a target feature set corresponding to the traffic data. Among them, the target feature set includes at least one basic feature, and the basic feature includes at least one of the following: IP address, port number, protocol type, packet length, traffic rate, etc.
[0041] Perform encoding mapping processing on the target feature set. Specifically, convert the target feature set into a natural language description, and combine it with context-aware vector mapping to obtain a token sequence, and finally map the token sequence to a target vector sequence.
[0042] Obtain a pre-trained traffic attack detection model. In this embodiment, the traffic attack detection model is the TrafficLLM model. Input the target vector sequence into the pre-trained traffic attack detection model, and after being processed by the traffic attack detection model, output a detection result. Among them, the detection result is whether there is an attack in the traffic data corresponding to the target vector sequence.
[0043] If the detection result indicates an attack, when the traffic attack detection model outputs the detection result, it also outputs the target attack type. Furthermore, the attack mitigation measures corresponding to the target attack type can be determined according to the target attack type.
[0044] Through the above solution, obtain traffic data, extract features from the traffic data to obtain a target feature set corresponding to the traffic data, reduce the dependence on manual feature engineering, and significantly improve the model generalization ability. Perform encoding mapping processing on the target feature set to obtain a target vector sequence, input the target vector sequence into the pre-trained traffic attack detection model, and after being processed by the traffic attack detection model, output a detection result. The traffic attack detection model obtained through training is used to detect traffic attacks, and the detection is more accurate. In response to the detection result indicating an attack, determine the target attack type, and determine the attack mitigation measures corresponding to the target attack type according to the target attack type. Determine the corresponding attack mitigation measures through the target attack type, realize the rapid generation of targeted strategies according to the real-time attack situation, and provide an efficient and flexible solution for network administrators.
[0045] In some embodiments, step 101 specifically includes:
[0046] Step 1011, obtain traffic data, and preprocess the traffic data to obtain traffic data to be extracted;
[0047] Step 1012: Input the traffic data to be extracted into a pre-trained feature extraction model. After being processed by the feature extraction model, output the target feature set corresponding to the traffic data.
[0048] In specific implementation, obtain traffic data, preprocess the traffic to obtain the traffic data to be extracted. Among them, the preprocessing method is specifically cleaning and formatting processing. Through the preprocessing operation, invalid or duplicate data is removed.
[0049] Obtain a pre-trained feature extraction model, input the traffic data to be extracted into the pre-trained feature extraction model. After being processed by the feature extraction model, output the target feature set corresponding to the traffic data.
[0050] In this embodiment, the feature extraction model is a composite model. The feature extraction model includes a random forest model, a gradient boosting tree model, and a support vector machine model. Furthermore, in step 1012, the determination method of the feature extraction model specifically includes:
[0051] Step A: Obtain training traffic data, an initial traffic extraction model, and a candidate feature set. Among them, the candidate feature set includes multiple candidate features.
[0052] Step B: Input the training traffic data into the initial traffic extraction model, and use the initial traffic extraction model to extract features from the training traffic data based on the candidate feature set to obtain a training feature set.
[0053] Step C: Determine the extraction performance data of the initial traffic extraction model according to the candidate feature set and the training feature set.
[0054] Step D: In response to the extraction performance data not meeting the preset end condition, update the candidate feature set to obtain a new candidate feature set. Use the initial traffic extraction model to extract features from the training traffic data based on the new candidate feature set, determine the new extraction performance data, and repeat until the new extraction performance data meets the preset end condition.
[0055] Step E: In response to the extraction performance data meeting the preset end condition, obtain the feature extraction model.
[0056] In specific implementation, obtain training traffic data, an initial traffic extraction model, and a preset candidate feature set. Among them, the candidate feature set includes multiple candidate features, and the candidate feature set includes basic features such as IP address, port number, protocol type, packet length, and traffic rate.
[0057] In this embodiment
[0058] Input the training traffic data into the initial traffic extraction model, and use the initial traffic extraction model to extract features from the training traffic data based on the candidate feature set to obtain a training feature set. That is, the most ideal result output by the initial traffic extraction model is that the training feature set is exactly the same as the candidate feature set.
[0059] Determine the extraction performance data of the initial traffic extraction model according to the candidate feature set and the training feature set. The extraction performance data represents the performance of the traffic extraction model during feature extraction, and the extraction performance data includes at least one of the following: accuracy rate, running speed, and central processing unit usage rate.
[0060] In response to the extraction performance data not meeting the preset end condition, perform an update process on the candidate feature set to obtain a new candidate feature set. Use the initial traffic extraction model to extract features from the training traffic data based on the new candidate feature set, determine the new extraction performance data, and repeat the process until the new extraction performance data meets the preset end condition.
[0061] In response to the extraction performance data meeting the preset end condition, obtain a feature extraction model. And when using the feature extraction model subsequently, perform feature extraction based on the candidate feature set before the update of the corresponding new candidate feature set when the extraction performance data meets the preset end condition, that is, use the candidate feature set of the previous time when the extraction performance data meets the preset end condition.
[0062] Taking the preset end condition as an example that the accuracy rate of feature extraction based on the new candidate feature set is more than the preset value lower than the accuracy rate of feature extraction based on the candidate feature set before the previous update, specifically:
[0063] If the preset value is 10%, determine that the accuracy rate of feature extraction based on the candidate feature set is 90%. Perform an update process on the candidate feature set to obtain a new candidate feature set, which is called the first candidate feature set here. The accuracy rate of feature extraction based on the first candidate feature set is 89%, that is, the accuracy rate drops by 1% at this time. Continue to perform an update process on the first candidate feature set to obtain a second candidate feature set. The accuracy rate of feature extraction based on the second candidate feature set is 75%, that is, the accuracy rate drops by 14% at this time, which is more than the preset value of 10%. At this time, the preset end condition is met, obtain a feature extraction model, and perform feature extraction based on the first candidate feature set subsequently.
[0064] Specifically, the process of performing an update process on the candidate feature set to obtain a new candidate feature set specifically includes:
[0065] For each candidate feature in the candidate feature set: input the candidate feature into a random forest model to obtain a first feature importance value corresponding to the candidate feature; input the candidate feature into a gradient boosting tree model to obtain a second feature importance value corresponding to the candidate feature; input the candidate feature into a support vector machine model to obtain a third feature importance value corresponding to the candidate feature; perform weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value to obtain a target feature importance value corresponding to the candidate feature;
[0066] Sort the target feature importance values corresponding to all candidate features, and delete the candidate feature with the smallest target feature importance value to obtain a new candidate feature set.
[0067] In specific implementation, for each candidate feature in the candidate feature set, input the candidate feature into a random forest model to obtain a first feature importance value corresponding to the candidate feature. Among them, the random forest model is a collective learning method that improves the model performance by constructing multiple decision trees and integrating their prediction results. This algorithm uses the bootstrap sampling technique to generate a training set during the training process of each tree, and only considers a random subset of features when splitting nodes, so as to increase the diversity of the trees and reduce overfitting. The random forest selects the optimal splitting feature by calculating the information gain, and the random forest algorithm is represented by the formula:
[0068] IG(D,f i )=H(D)-H(D)|f i
[0069] where, IG(D,f i ) is the information gain, indicating the amount of reduction in the entropy of the data set D under the condition of the given feature f i , H(D) is the entropy of the data set, and H(D)|f i is the entropy under the condition of the feature f i , f i represents the i-th feature in the data set, and i is a subscript indicating the index of the feature.
[0070] Input the candidate feature into a gradient boosting tree model to obtain a second feature importance value corresponding to the candidate feature. The gradient boosting tree model is an ensemble learning algorithm that sequentially trains a series of weak prediction models, and each tree tries to correct the errors of the previous tree. The initial model may be just a constant value, and then a new decision tree is trained at each step to fit the residuals of the previous step.
[0071] The gradient boosting tree model combines the prediction results of these trees in a weighted manner, where the weights are controlled by the learning rate parameter, which is optimized in each iteration step. The goal of the algorithm is to improve the overall prediction performance by gradually reducing the error of the model. The final model is the weighted sum of all weak learners. Mathematically, the update rule of GBT can be expressed as follows. At the m-th iteration step, the prediction value of the model is updated using the formula:
[0072] F m (x) = F m-1 (x) + y m h m (x)
[0073] where, F m (x) represents the prediction function of the model after the m-th iteration step, F m-1 (x) represents the prediction function of the model after the m - 1-th iteration step, y m is the learning rate, and h m (x) is the prediction function of the m-th tree. In this way, GBT can effectively handle various data patterns and provide accurate prediction results.
[0074] Input the candidate features into the support vector machine model to obtain the third feature importance value corresponding to the candidate features. The support vector machine model is a supervised learning algorithm used for classification and regression analysis. It classifies by finding a hyperplane that maximizes the margin in the feature space, and this hyperplane is called the optimal separating hyperplane. The key idea of SVM is to maximize the width of the classification margin, thereby improving the generalization ability of the model.
[0075] For linearly separable data, SVM finds a hyperplane that can correctly classify all training samples and maximizes the margin between the two classes. For non-linearly separable data, SVM maps the data to a higher-dimensional space through the kernel trick to make it linearly separable.
[0076] The optimization objective function of SVM:
[0077]
[0078] where, w is the normal vector of the hyperplane, ξ is the slack variable, ξ i is the slack variable of the i-th sample, C is the penalty parameter, b is the bias term, which controls the trade-off between the error term and margin maximization, represents minimizing the weight vector w, the bias term b, and the slack variable ξ, represents half of the L2 norm of the weight vector w, represents the sum of all slack variables ξ.
[0079] In this embodiment, three different machine learning models are trained respectively: Random Forest (RF), Gradient Boosting Tree (GBT), and Support Vector Machine (SVM). Random Forest and GBT perform well in processing high-dimensional data and feature selection, while SVM has advantages in dealing with small samples and high-dimensional data. The combination of the three can capture the characteristics of network traffic data from different perspectives.
[0080] Perform weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value to obtain the target feature importance value corresponding to the candidate feature.
[0081] Sort the target feature importance values corresponding to all candidate features, delete the candidate feature with the smallest target feature importance value, and obtain a new candidate feature set.
[0082] Use the new candidate feature set to retrain the initial feature extraction model, and at the same time monitor the changes in model accuracy, running time, and CPU (Central Processing Unit) usage rate to find the best balance point that reduces computational resource consumption while maintaining high accuracy, and obtain the feature extraction model.
[0083] Specifically, in this embodiment, when performing weighted processing, the determination methods of the weights corresponding to the Random Forest model, the Gradient Boosting Tree model, and the Support Vector Machine model are specifically as follows:
[0084] Step a, obtain a validation data set, where the validation data set contains validation features and the actual importance values corresponding to the validation features;
[0085] Step b, input the validation features into the Random Forest model, the Gradient Boosting Tree model, and the Support Vector Machine model respectively to obtain the first candidate feature importance value, the second candidate feature importance value, and the third candidate feature importance value corresponding to the validation features;
[0086] Step c, according to the first candidate feature importance value, the second candidate feature importance value, the third candidate feature importance value, and the actual importance value, use the least squares method to determine the first weight corresponding to the Random Forest model, the second weight corresponding to the Gradient Boosting Tree model, and the third weight corresponding to the Support Vector Machine model.
[0087] In specific implementation, obtain a validation data set, where the validation data set contains validation features and the actual importance values corresponding to the validation features. Input the validation features into the Random Forest model, the Gradient Boosting Tree model, and the Support Vector Machine model respectively to obtain the first candidate feature importance value, the second candidate feature importance value, and the third candidate feature importance value corresponding to the validation features.
[0088] Exemplarily, the first candidate feature importance value, the second candidate feature importance value, and the third candidate feature importance value corresponding to the verification feature are expressed by the formula as follows:
[0089]
[0090] Wherein, is the first candidate feature importance value, is the second candidate feature importance value, is the third candidate feature importance value.
[0091] According to the first candidate feature importance value, the second candidate feature importance value, the third candidate feature importance value, and the actual importance value, the first weight corresponding to the random forest model, the second weight corresponding to the gradient boosting tree model, and the third weight corresponding to the support vector machine model are determined by using the least squares method. Among them, the least squares method is expressed by the formula as follows:
[0092] β = (X T X) -1 X T y true
[0093] Wherein, y true is the actual importance value.
[0094] In some embodiments, step 102 specifically includes:
[0095] Step 1021, using a domain-adapted tokenizer to perform a conversion process on the target feature set to obtain a natural language sequence corresponding to the target feature set;
[0096] Step 1022, using the byte pair encoding technique to perform an encoding process on the natural language sequence to obtain an initial token sequence;
[0097] Step 1023, obtaining a preset marker, and adding the preset marker to the initial token sequence to obtain a token sequence;
[0098] Step 1024, for each token in the token sequence: performing a mapping process on the token according to the position of the token in the token sequence to obtain an initial vector corresponding to the token;
[0099] Step 1025, counting all the initial vectors, and performing a normalization process on all the initial vectors to obtain a target vector sequence.
[0100] Specifically in implementation, a domain-adapted tokenizer is used to perform a conversion process on the target feature set to obtain a natural language sequence corresponding to the target feature set, where the domain-adapted tokenizer is the Tokenizer of TrafficLLM.
[0101] Exemplarily, the domain adaptation tokenizer defines a domain-specific vocabulary that can convert the original traffic data: {"source_ip":"192.168.1.1","destination_ip":"192.168.1.2","port":80,"protocol":"TCP"} into the natural language description "The source IP is 192.168.1.1, the destination IP is 192.168.1.2, the port number is 80, and the protocol type is TCP".
[0102] The natural language sequence is encoded using the byte pair encoding technique to obtain an initial token sequence, and the byte pair encoding technique adopts a byte pair encoding strategy based on the characteristics of the network domain.
[0103] Based on the above example, an integrity preservation mechanism is implemented for network entities such as IP addresses and port numbers. "192.168.1.1" is used as an independent token rather than a discrete digital sequence, and the semantic descriptor (such as "the protocol type is") is segmented by conventional BPE to generate the initial token sequence ["The source IP is", "192.168.1.1", ", the protocol type is", "TCP"].
[0104] To help the model understand the structure of the traffic data, a preset marker is obtained and added to the initial token sequence to obtain a token sequence, and the preset marker is a classification start and separator marker.
[0105] Based on the above example, by inserting special markers [CLS] (classification start) and [SEP] (separator) into the initial token sequence, the final token sequence ["[CLS]", "The source IP is", "192.168.1.1", ", the port number is", "80", ", the protocol type is", "TCP", "[SEP]"] is formed to clearly distinguish the boundaries of different traffic segments.
[0106] Through the vector table predefined by the embedding layer of TrafficLLM, for each token in the token sequence: the token is mapped according to its position in the token sequence to obtain the initial vector corresponding to the token.
[0107] Exemplarily, the token sequence is mapped to the initial vectors [v_CLS, v_The source IP is, v_192.168.1.1, v_The port number is, v_80, v_The protocol type is, v_TCP, v_SEP], where v_XXX represents the high-dimensional vector corresponding to each token, enabling the computer to process this data.
[0108] Since the vectors of each token are generated independently, the model itself cannot directly perceive the order of the tokens in the sequence from these vectors. To help the model understand the position information of the tokens in the sequence, a context-aware vector mapping technique is introduced, by introducing position encoding enhanced with traffic features during the embedding layer transformation:
[0109]
[0110] To strengthen the spatio-temporal correlation modeling ability of the keyword fields, a sequence of token vectors with weight coefficients is finally generated: [v_[CLS], v_source IP is (α = 1.2), v_192.168.1.1(α = 1.2), v_port number is (α = 1.1), v_80(α = 1.1), v_protocol type is (α = 1.0), v_TCP(α = 1.0), v_SEP], making the model pay more attention to these fields during modeling. Among them, pos is the position index, determined according to the order of the tokens in the sequence, representing the absolute position of the current token in the sequence; i is the dimension index, representing the serial number of the dimension in the embedding vector; d model is the model dimension, determined during the model design phase, representing the total dimension of the embedding vector, controlling the frequency distribution of the position encoding. The higher the dimension, the finer the position features that can be encoded; α f is the field type weight factor, a coefficient for dynamically adjusting the position encoding intensity according to different field types.
[0111] Statistically analyze all the initial vectors and perform normalization on all the initial vectors to obtain the target vector sequence. Among them, the normalization methods include layer normalization and scaling. The specific steps of the normalization process are as follows:
[0112] Calculate the mean and variance. For the input vector [x1, x2,..., x n , calculate its mean μ, variance σ 2 and standard deviation σ. Standardization: Standardize each element, and the formula is:
[0113] z i =(x i -μ) / σ
[0114] At this time, the data distribution becomes a mean of 0 and a standard deviation of 1.
[0115] Perform Min-Max scaling. The range of the standardized data is [z min , z max , and the scaling formula is:
[0116] scaled i =(z i -z min ) / (z max -zmin )
[0117] Through the above process, all data is adjusted to a unified range of 0 - 1, avoiding bias in the model when processing this data due to differences in numerical magnitudes. In this way, the input distribution is ensured to be consistent and adapted to the requirements of model pre-training.
[0118] In some embodiments, before using the domain adaptation tokenizer to perform transformation processing on the target feature set, the target feature set is first subjected to format unification processing. For traffic data from different sources, its protocol structure is parsed, unified fields are extracted, such as five-tuples, timestamps, payload lengths, etc., and converted into an intermediate representation in JSON (a lightweight data interchange format) format to eliminate format differences.
[0119] In some embodiments, the process of determining the traffic attack detection model specifically includes:
[0120] Step 10A, obtaining an initial traffic attack detection model and a training vector sequence, where the initial traffic attack detection model is a traffic attack detection model that has been pre-trained;
[0121] Step 10B, obtaining a preset instruction semantic vector, and using the preset attention weights to combine the training vector sequence and the instruction semantic vector to obtain a training instruction;
[0122] Step 10C, using the training instruction to perform language instruction fine-tuning on the self-attention module in the initial traffic attack model until the first preset training end condition is met, obtaining a first traffic attack model;
[0123] Step 10D, freezing the model parameters of the self-attention module and the feed-forward network in the first traffic attack model, inputting the training vector sequence into the first traffic attack model, and using the training vector sequence to fine-tune the first traffic attack model until the second preset training end condition is met, obtaining a traffic attack model.
[0124] Specifically in implementation, an initial traffic attack detection model and a training vector sequence are obtained, where the initial traffic attack detection model is a traffic attack detection model that has been pre-trained. Specifically, the pre-training process includes:
[0125] Pre-train TrafficLLM using massive network traffic data. Through the masked language modeling task, predict the masked token vectors. Specifically, the model randomly masks certain token vectors. For example, it masks the token vector v_80 representing "the port number is 80", and then predicts these masked tokens by analyzing the context. This process enables the model to learn the context relevance of traffic protocols, such as the transmission control protocol handshake rules, the burst traffic characteristics of DDoS attacks, etc., thereby capturing the general patterns of traffic behavior.
[0126] First, fine-tune the initial traffic attack detection model with natural language instructions. By injecting instruction-response pairs related to network security, learn the mapping relationship between task descriptions and traffic analysis. Obtain the preset instruction semantic vector. Based on the cross-modal semantic fusion strategy, encode the natural language instructions into semantic vectors to further enhance the model's semantic understanding ability of natural language instructions, enabling the model to more accurately capture the key information in the instructions and correspond it to the internal logic and data processing flow of the model.
[0127] Use the preset attention weights to combine the training vector sequence and the instruction semantic vector to obtain training instructions, making the model pay more attention to traffic features related to the instructions. Use the training instructions to fine-tune the self-attention module in the initial traffic attack model until the first preset training end condition is met to obtain the first traffic attack model.
[0128] In this embodiment, by learning a large number of instruction-response pairs, these instruction-response pairs include not only normal traffic but also various attack traffic samples, and the annotations are complete and accurate. For example, the instruction is "detect whether there is a UDP flood attack in the current traffic", and the response is "UDP flood attack detected, source IP is 192.168.1.100, target port is 80", etc. Through training with these data, the model has a preliminary understanding of the characteristics of network attack traffic, enabling the model to adapt to the semantic expression of natural language and establish a bridge between instructions and model behavior.
[0129] Use Sentence-BERT (a sentence embedding technology based on the BERT model, and BERT is a bidirectional encoder representation based on Transformer) to encode natural language instructions into semantic vectors v_instruction of a fixed dimension. This step converts the instructions into a digital form that is easier for the model to understand and process, while retaining the semantic information of the instructions.
[0130] By calculating the attention weights, the instruction semantic vector is combined with the training vector sequence. For example, if the instruction is "Detect UDP attacks", the model will calculate the association weights between the semantic vector v_instruction and each training vector (such as v_protocol_type and v_UDP). If the training vector corresponding to v_protocol_type is v_TCP, the weight is lower; if the training vector is v_UDP, the weight is higher.
[0131] By dynamically adjusting the attention weights, the model can pay more attention to the parts related to the instruction, improving the model's ability to understand and respond to instructions. The formula for calculating the attention weights is as follows:
[0132]
[0133] Among them, Q is the query matrix, representing the query requirements of the current token for context information, which is generated by linearly transforming the input vector. K is the key matrix, representing the retrievable features of the context tokens, which is generated by linearly transforming the input vector through another linear transformation. V is the value matrix, storing the actual content information of the context tokens, which is generated by linearly transforming the input vector through a third linear transformation. d k is the key / query vector dimension, used to control the computational complexity of the attention weights. The smaller the dimension, the lower the computational amount, but the expressive ability may be lost. λ is the instruction fusion coefficient, a learnable scalar parameter, used to adjust the contribution intensity of the instruction semantic vector E inst to the attention weights. E inst is the instruction semantic vector, representing the semantic encoding of the natural language instruction, which is generated by the pre-trained model. softmax is the activation function, used to convert the sum of the scaled dot product and the bias term into a probability distribution.
[0134] Through a well-designed prompt template, the instruction is transformed into a form that is easier for the model to understand. For example, the original instruction "Detect DDoS attacks" is formatted according to the prompt template as "Detect whether there is a DDoS attack in the current traffic and output the attack type and source IP address", etc. By repeatedly training the model to make accurate and coherent responses to the instructions, the model can accumulate sufficient training experience in instruction understanding, further consolidating its performance in the field of natural language instructions.
[0135] In this embodiment, the first preset training end condition is that the instruction understanding accuracy rate on the training vector sequence is greater than or equal to 99.5% (for example, among 100 instructions, at least 99 are correctly parsed), and the decrease amplitude of the cross-entropy loss for 3 consecutive epochs is less than 0.1%. Then it is considered that the fine-tuning part of the natural language instruction reaches the end condition. Among them, the cross-entropy loss is a commonly used loss function, used to measure the accuracy of the model's understanding of the instruction. When the decrease amplitude of the cross-entropy loss on the training vector sequence for 3 consecutive epochs is less than 0.1%, it indicates that the loss function has stabilized and the fine-tuning can end.
[0136] By performing natural language instruction fine-tuning on the initial traffic attack detection model, the model can accurately understand various instructions related to attack behavior detection, including identifying different types of network attacks (such as TCP flooding attacks, UDP flooding attacks, etc.), parsing key features in traffic data (such as IP addresses, port numbers, protocol types, etc.), and generating corresponding detection actions and results according to instructions, ensuring that the mapping relationship between instructions and model tasks is clear and accurate. For example, when the model receives an instruction of "detect whether there is a UDP flooding attack in the current traffic", it can clearly identify that the core intention of this instruction is to detect UDP-type flooding attacks and can correspond it to the relevant detection logic and data processing flow in the model. Secondly, the model can not only accurately understand specific instructions but also has good generalization ability. For various instructions with different expressions related to the attack behavior detection task, such as "monitor abnormal behaviors in network traffic" and "identify the sources and types of malicious traffic", it can correctly interpret and convert them into corresponding model operations, thus achieving flexible responses to attack detection tasks under different scenarios and requirements.
[0137] Perform fine-tuning on the initial traffic attack detection model for the attack behavior detection task, freeze the model parameters of the self-attention module and the feed-forward network in the first traffic attack model, input the training vector sequence into the first traffic attack model, and use the training vector sequence to fine-tune the first traffic attack model until the second preset training end condition is met to obtain the traffic attack model.
[0138] Specifically, the LoRA (Low-Rank Adaptor) technology is used for efficient fine-tuning, where the field weights in the training vector sequence affect the adjustment direction of LoRA. First, define the configuration parameters of LoRA, including setting the rank r of the low-rank matrix and determining the target module where the adaptor layer will be inserted. Subsequently, accurately embed the adaptor layer into the specified modules of the backbone network, which include the self-attention module for fine-tuning the query (Q), key (K), and value (V) projection matrices, and the feed-forward network for specifically adjusting the two linear layers it contains. Among them, the low-rank matrix is updated during fine-tuning to make the model pay more attention to high-weight features. If the weight coefficient of v_source IP is high, LoRA will enhance the model's sensitivity to abnormal IPs; if the weight coefficient of v_protocol type is low, the model will reduce its dependence on the protocol type.
[0139] After completing the embedding of the adapter layer, freeze all the original parameters of the self-attention module and the feed-forward network in the backbone network to ensure that these pre-trained parameters remain unchanged. At the same time, only train the parameters of the adapter layer. Through low-rank factorization, the adapter layer adjusts the weights of the backbone network, enabling the model to accurately adapt to the specific task of attack behavior detection. During this process, use the labeled attack traffic data (including normal traffic and attack samples) for supervised training to optimize the classification accuracy of the model for attack patterns.
[0140] In this embodiment, the second preset end condition is that the F1 score (the harmonic mean of precision and recall) of the training vector sequence ≥ 99.3%, and the performance of the validation set has no obvious improvement for 5 consecutive epochs (for example, the accuracy fluctuation is less than 0.05%), then it is considered that the fine-tuning part of the attack detection task reaches the end condition.
[0141] In some embodiments, step 104 specifically includes:
[0142] Step 1041, search the database according to the target attack type to determine the attack features corresponding to the target attack type;
[0143] Step 1042, determine the device name of the target device, and input the device name, the attack features, and the target attack type into a preset mitigation strategy template to obtain a semi-structured instruction;
[0144] Step 1043, perform format conversion on the semi-structured instruction to obtain an attack mitigation measure.
[0145] Specifically in implementation, search the database according to the target attack type to determine the attack features corresponding to the target attack type. Specifically, extract the key attack features according to the target attack type, including attack protocol, target port, packet rate, source IP distribution, traffic peak, etc., and perform pattern matching by associating with the historical attack pattern database. The historical attack pattern database stores the features and patterns of various attack behaviors detected in the past. By comparing the newly detected attack features with the historical patterns, the type and behavior pattern of the attack can be quickly identified, providing a more accurate basis for the generation of mitigation strategies.
[0146] Exemplarily, if the attack features currently detected are highly similar to the pattern of a certain large-scale UDP flood attack recorded in the database, then the system can quickly generate effective mitigation actions for this type of attack.
[0147] Determine the device name of the target device, and input the device name, the attack features, and the target attack type into a preset mitigation strategy template to obtain semi-structured instructions. The attack features include data features and specific attacks. The mitigation strategy template is a template that includes the attack features and the attack mitigation actions.
[0148] Exemplarily, the mitigation strategy template is {data features}: traffic analysis features and initial packet samples for {specific attacks}. Now, please develop a set of defense measures for {device name} to reduce the threat posed by {attack type}. Please elaborate on your strategy and configuration steps to ensure that these measures can be effectively executed on the device and reduce the negative impact of the attack.
[0149] In this embodiment, the main purpose of using semi-structured instructions is to transform complex attack features and defense requirements into a form that is easy to understand and execute. It combines the intuitiveness of natural language and the precision of structured data, making the generated instructions both quickly understandable by human network administrators and directly parsable and executable by automated systems.
[0150] Perform format conversion on the semi-structured instructions to obtain attack mitigation measures. The attack mitigation measures are device-specific configuration commands. After verifying the effectiveness of the commands through a simulation environment, they are pushed to the target network device for execution, and the execution status is fed back to the management platform in real time.
[0151] In this embodiment, the specific process of performing format conversion on the semi-structured instructions to obtain attack mitigation measures includes:
[0152] Extract specific measures from the semi-structured instructions, including key information such as attack type, attack source IP, specific device, and defense measures. Then, according to the extracted specific measures, look up the mapping rules for device-specific commands to find the corresponding mapping rule "access-list <id>"deny udp host <source IP> any", and fill the extracted key information into the mapping rule to generate the device's dedicated command "access-list 100 deny udp host 192.168.1.100 any". This step ensures that the command matches the device's configuration syntax, laying the foundation for subsequent verification and execution.
[0153] In this embodiment, the specific steps of the validity verification include:
[0154] First, set up a virtual network environment in the simulation environment, which includes a simulated attack source (IP address: 192.168.1.100) and a target device (such as an H3C router). Then, send data packets from the attack source to the target device by simulating UDP traffic, and observe whether the target device can correctly restrict UDP traffic from this IP address. If the command is effective, the target device will reject all UDP packets from 192.168.1.100, thus verifying the functionality of the command. When testing in the simulation environment, a network packet capture tool can be used to monitor network traffic and observe whether there is still UDP traffic from the attack source passing through. At the same time, the device's logs and statistical information can also be used to confirm whether the ACL is correctly applied and whether the attack traffic is successfully blocked. If it is found during the test that the command fails to achieve the expected effect, the command needs to be adjusted, such as checking whether the rule number of the ACL is correct, whether there are conflicting rules, or confirming whether the command is correctly applied to the target interface, etc. By performing functional verification in the simulation environment, it can be ensured that the generated command can effectively perform the defense task in actual application, thus providing reliable protection measures for network devices.
[0155] In this embodiment, a composite model is constructed based on three base learners: random forest, gradient boosting tree, and support vector machine. These base learners evaluate features from different perspectives, calculate the importance of each feature, and select key features according to the importance scores, reducing the time complexity of model training while enhancing the interpretability of the model.
[0156] In this embodiment, a normalization method for multi-source heterogeneous network abnormal traffic data is introduced. Based on the traffic domain token generation mechanism, through hybrid BPE tokenization (preserving the integrity of IP addresses) and positional encoding with enhanced traffic features, the heterogeneous data is converted into a standardized token sequence, solving the problems of inconsistent formats and large differences in feature scales of massive multi-source heterogeneous network abnormal traffic data.
[0157] In this embodiment, based on the detection results of the model, TrafficLLM is used to generate specific configuration instructions for network attack mitigation. These instructions can dynamically generate corresponding mitigation strategies according to the detected specific types of network attacks, so as to achieve rapid response and effective containment of the attacks.
[0158] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In this case of the distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0159] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order from those in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0160] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a network attack mitigation device.
[0161] Refer to Figure 2 , Figure 2 For the network attack mitigation device of the embodiment, it specifically includes:
[0162] A data acquisition module 201, configured to acquire traffic data, extract features from the traffic data, and obtain a target feature set corresponding to the traffic data;
[0163] An encoding and mapping module 202, configured to perform encoding and mapping processing on the target feature set to obtain a target vector sequence;
[0164] An attack detection module 203, configured to input the target vector sequence into a pre-trained traffic attack detection model, and output a detection result after being processed by the traffic attack detection model;
[0165] A mitigation measure determination module 204, configured to, in response to the detection result indicating an attack, determine the target attack type, and determine an attack mitigation measure corresponding to the target attack type according to the target attack type.
[0166] In some embodiments, the data acquisition module 201 is specifically configured to:
[0167] Obtain traffic data, preprocess the traffic data to obtain traffic data to be extracted;
[0168] Input the traffic data to be extracted into a pre-trained feature extraction model, and after being processed by the feature extraction model, output a target feature set corresponding to the traffic data.
[0169] In some embodiments, the device includes a traffic extraction model training module, and the traffic extraction model training module is specifically configured to:
[0170] Obtain training traffic data, an initial traffic extraction model, and a candidate feature set, where the candidate feature set includes multiple candidate features;
[0171] Input the training traffic data into the initial traffic extraction model, and use the initial traffic extraction model to extract features from the training traffic data based on the candidate feature set to obtain a training feature set;
[0172] Determine the extraction performance data of the initial traffic extraction model according to the candidate feature set and the training feature set;
[0173] In response to the extraction performance data not meeting the preset end condition, perform an update process on the candidate feature set to obtain a new candidate feature set, use the initial traffic extraction model to extract features from the training traffic data based on the new candidate feature set, determine new extraction performance data, and repeat the execution until the new extraction performance data meets the preset end condition;
[0174] In response to the extraction performance data meeting the preset end condition, obtain a feature extraction model;
[0175] The update process on the candidate feature set to obtain a new candidate feature set includes:
[0176] For each candidate feature in the candidate feature set: input the candidate feature into a random forest model to obtain a first feature importance value corresponding to the candidate feature, input the candidate feature into a gradient boosting tree model to obtain a second feature importance value corresponding to the candidate feature, input the candidate feature into a support vector machine model to obtain a third feature importance value corresponding to the candidate feature, and perform a weighted process on the first feature importance value, the second feature importance value, and the third feature importance value to obtain a target feature importance value corresponding to the candidate feature;
[0177] Sort the target feature importance values corresponding to all candidate features, delete the candidate feature with the smallest target feature importance value, and obtain a new candidate feature set.
[0178] In some embodiments, the traffic extraction model training module is further specifically configured to:
[0179] Obtain a validation data set, where the validation data set contains validation features and actual importance values corresponding to the validation features;
[0180] Input the validation features into a random forest model, a gradient boosting tree model, and a support vector machine model respectively, to obtain a first candidate feature importance value, a second candidate feature importance value, and a third candidate feature importance value corresponding to the validation features;
[0181] According to the first candidate feature importance value, the second candidate feature importance value, the third candidate feature importance value, and the actual importance value, use the least squares method to determine a first weight corresponding to the random forest model, a second weight corresponding to the gradient boosting tree model, and a third weight corresponding to the support vector machine model;
[0182] Use the first weight, the second weight, and the third weight to perform weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value, to obtain a target feature importance value corresponding to the candidate feature.
[0183] In some embodiments, the encoding and mapping module 202 is configured to:
[0184] Use a domain adaptation tokenizer to perform conversion processing on the target feature set, to obtain a natural language sequence corresponding to the target feature set;
[0185] Use byte pair encoding technology to perform encoding processing on the natural language sequence, to obtain an initial token sequence;
[0186] Obtain a preset marker, and add the preset marker to the initial token sequence, to obtain a token sequence;
[0187] For each token in the token sequence: perform mapping processing on the token according to the position of the token in the token sequence, to obtain an initial vector corresponding to the token;
[0188] Statistically analyze all the initial vectors, and perform normalization processing on all the initial vectors, to obtain a target vector sequence.
[0189] In some embodiments, the device includes a traffic attack detection model training module, and the traffic attack detection model training module is specifically configured to:
[0190] Obtain an initial traffic attack detection model and a training vector sequence, where the initial traffic attack detection model is a pre-trained traffic attack detection model;
[0191] Obtain a preset instruction semantic vector, and use the preset attention weight to combine the training vector sequence and the instruction semantic vector to obtain a training instruction;
[0192] Use the training instruction to perform language instruction fine-tuning on the self-attention module in the initial traffic attack model until the first preset training end condition is met, and obtain the first traffic attack model;
[0193] Freeze the model parameters of the self-attention module and the feed-forward network in the first traffic attack model, input the training vector sequence into the first traffic attack model, and use the training vector sequence to fine-tune the first traffic attack model until the second preset training end condition is met, and obtain the traffic attack model.
[0194] In some embodiments, the traffic attack detection model training module is further specifically configured to:
[0195] Obtain a preset instruction semantic vector and a preset attention weight, where the attention weight is represented by the formula:
[0196]
[0197] where Attention(Q, K, V) is the attention weight, Q is the query matrix, K is the key matrix, V is the value matrix, d k is the key query vector dimension, λ is the instruction fusion coefficient, E inst is the instruction semantic vector, and softmax() is the activation function;
[0198] Use the preset attention weight to combine the training vector sequence and the instruction semantic vector to obtain a training instruction.
[0199] In some embodiments, the mitigation measure determination module 204 is further specifically configured to:
[0200] Search the database according to the target attack type to determine the attack features corresponding to the target attack type;
[0201] Determine the device name of the target device, and input the device name, the attack features, and the target attack type into a preset mitigation strategy template to obtain a semi-structured instruction;
[0202] Perform format conversion on the semi-structured instruction to obtain an attack mitigation measure.
[0203] For the convenience of description, the above device is described by function as various modules respectively. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0204] The device of the above embodiment is used to implement the corresponding network attack mitigation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0205] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the network attack mitigation method described in any one of the above embodiments.
[0206] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0207] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0208] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0209] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0210] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0211] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0212] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0213] The electronic device of the above embodiment is used to implement the corresponding network attack mitigation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0214] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the network attack mitigation method described in any of the above embodiments.
[0215] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0216] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the network attack mitigation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0217] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0218] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0219] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0220] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0221] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0222] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions are to be regarded as illustrative rather than restrictive.
[0223] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0224] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.< / id>
Claims
1. A method for mitigating network attacks, characterized in that, Including: Obtain the traffic data of the target device, perform feature extraction on the traffic data to obtain a target feature set corresponding to the traffic data; Perform encoding mapping processing on the target feature set to obtain a target vector sequence; Input the target vector sequence into a pre-trained traffic attack detection model, and after being processed by the traffic attack detection model, output a detection result; In response to the detection result indicating an attack, determine the target attack type, and determine an attack mitigation measure corresponding to the target attack type according to the target attack type.
2. The method according to claim 1, wherein The obtaining traffic data, performing feature extraction on the traffic data to obtain a target feature set corresponding to the traffic data includes: Obtain traffic data, perform preprocessing on the traffic data to obtain traffic data to be extracted; Input the traffic data to be extracted into a pre-trained feature extraction model, and after being processed by the feature extraction model, output a target feature set corresponding to the traffic data.
3. The method according to claim 2, wherein The process of determining the feature extraction model specifically includes: Obtain training traffic data, an initial traffic extraction model, and a candidate feature set, where the candidate feature set includes multiple candidate features; Input the training traffic data into the initial traffic extraction model, and based on the candidate feature set, use the initial traffic extraction model to perform feature extraction on the training traffic data to obtain a training feature set; Determine the extraction performance data of the initial traffic extraction model according to the candidate feature set and the training feature set; In response to the extraction performance data not meeting the preset end condition, perform update processing on the candidate feature set to obtain a new candidate feature set, and based on the new candidate feature set, use the initial traffic extraction model to perform feature extraction on the training traffic data to determine new extraction performance data, and repeat the execution until the new extraction performance data meets the preset end condition; In response to the extraction performance data meeting the preset end condition, obtain a feature extraction model; The performing update processing on the candidate feature set to obtain a new candidate feature set includes: For each candidate feature in the candidate feature set: input the candidate feature into a random forest model to obtain a first feature importance value corresponding to the candidate feature, input the candidate feature into a gradient boosting tree model to obtain a second feature importance value corresponding to the candidate feature, input the candidate feature into a support vector machine model to obtain a third feature importance value corresponding to the candidate feature, and perform weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value to obtain a target feature importance value corresponding to the candidate feature; Perform sorting processing on the target feature importance values corresponding to all candidate features, delete the candidate feature with the smallest target feature importance value, and obtain a new candidate feature set.
4. The method according to claim 3, wherein The performing weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value to obtain a target feature importance value corresponding to the candidate feature includes: Obtain a validation data set, where the validation data set contains validation features and actual importance values corresponding to the validation features; Input the verification features into a random forest model, a gradient boosting tree model, and a support vector machine model respectively to obtain the first candidate feature importance value, the second candidate feature importance value, and the third candidate feature importance value corresponding to the verification features; According to the first candidate feature importance value, the second candidate feature importance value, the third candidate feature importance value, and the actual importance value, use the least squares method to determine the first weight corresponding to the random forest model, the second weight corresponding to the gradient boosting tree model, and the third weight corresponding to the support vector machine model; Use the first weight, the second weight, and the third weight to perform weighted processing on the first feature importance value, the second feature importance value, and the third feature importance value to obtain the target feature importance value corresponding to the candidate feature.
5. The method according to claim 1, characterized in that, The encoding and mapping process of the target feature set to obtain a target vector sequence includes: Use a domain-adapted tokenizer to perform transformation processing on the target feature set to obtain a natural language sequence corresponding to the target feature set; Use byte pair encoding technology to encode the natural language sequence to obtain an initial token sequence; Obtain a preset marker and add the preset marker to the initial token sequence to obtain a token sequence; For each token in the token sequence: perform mapping processing on the token according to the position of the token in the token sequence to obtain an initial vector corresponding to the token; Count all the initial vectors and perform normalization processing on all the initial vectors to obtain a target vector sequence.
6. The method according to claim 1, characterized in that, The determination process of the traffic attack detection model specifically includes: Obtain an initial traffic attack detection model and a training vector sequence, where the initial traffic attack detection model is a pre-trained traffic attack detection model; Obtain a preset instruction semantic vector, and use a preset attention weight to combine the training vector sequence and the instruction semantic vector to obtain a training instruction; Use the training instruction to perform language instruction fine-tuning on the self-attention module in the initial traffic attack model until the first preset training end condition is met to obtain a first traffic attack model; Freeze the model parameters of the self-attention module and the feed-forward network in the first traffic attack model, input the training vector sequence into the first traffic attack model, and use the training vector sequence to fine-tune the first traffic attack model until the second preset training end condition is met to obtain a traffic attack model.
7. The method according to claim 6, characterized in that, The obtaining of the preset instruction semantic vector, and using the preset attention weight to combine the training vector sequence and the instruction semantic vector to obtain a training instruction includes: Obtain a preset instruction semantic vector and a preset attention weight, where the attention weight is represented by the formula: Among them, Attention(Q, K, V) is the attention weight, Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key query vector, λ is the instruction fusion coefficient, E inst is the instruction semantic vector, and softmax() is the activation function; Use the preset attention weight to combine the training vector sequence and the instruction semantic vector to obtain a training instruction.
8. The method according to claim 1, wherein The determining of the attack mitigation measure corresponding to the target attack type according to the target attack type includes: Search the database according to the target attack type to determine the attack features corresponding to the target attack type; Determine the device name of the target device, and input the device name, the attack feature, and the target attack type into a preset mitigation strategy template to obtain a semi-structured instruction; Perform format conversion on the semi-structured instruction to obtain an attack mitigation measure.
9. A network attack mitigation device, characterized in that, Including: A data acquisition module, configured to acquire traffic data, perform feature extraction on the traffic data, and obtain a target feature set corresponding to the traffic data; An encoding mapping module, configured to perform encoding mapping processing on the target feature set to obtain a target vector sequence; An attack detection module, configured to input the target vector sequence into a pre-trained traffic attack detection model, and output a detection result after being processed by the traffic attack detection model; A mitigation measure determination module, configured to, in response to the detection result indicating an attack, determine the target attack type, and determine an attack mitigation measure corresponding to the target attack type according to the target attack type.
10. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Network security attack detection method, device, equipment and product
CN121509065A