Network attack identification method and device, equipment, storage medium and program product

Through multi-head attention mechanism and association rule algorithm, feature extraction of network traffic data is solved to identify network attack behavior, and the problems of weak protection capabilities of unknown types and great influence of interference information in the existing technology are solved, and the accuracy and protection performance of network attack recognition are improved.

CN119945788APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117472.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing cyber attack identification methods have weak protection capabilities when facing unknown types of attacks and are easily affected by interfering information, resulting in insufficient identification accuracy.

Method used

The multi-head attention mechanism is used to extract the overall feature of the network traffic data, and the local feature extraction is performed through the association rule algorithm to obtain the soft label feature vector, thereby identifying the network attack behavior.

Benefits of technology

Effectively extract the overall key information of network traffic, reduce the impact of interference information, expand the accuracy of network attack identification, and improve the protection performance of the protection module.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945788A_ABST
    Figure CN119945788A_ABST
Patent Text Reader

Abstract

The invention discloses a network attack identification method and device, equipment, a storage medium and a program product, and the method comprises the steps: determining a network traffic vector corresponding to network traffic data; performing overall feature extraction on the network flow vector based on a multi-head attention mechanism to obtain a first feature vector; performing local feature extraction on the network flow vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector; and based on the first feature vector and the second feature vector, identifying a network attack behavior to obtain a target network attack category corresponding to the network traffic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security and data mining, and in particular to a network attack identification method, device, equipment, storage medium and program product. Background Art

[0002] In terms of defense against network attacks, it can specifically include two steps: detection and protection. The premise of protecting against network attacks is to detect attacks, identify specific attack types, and then apply specific protection strategies. Among them, intrusion detection systems (IDS) are used to detect possible threats and prevent unauthorized access. However, most IDSs in related technologies use rule bases combined with feature-based intrusion detection methods to identify network attacks. This method has weak protection capabilities against unknown types of attacks and is easily affected by interference information. It still has limitations in identifying network attacks. Summary of the invention

[0003] To solve the above technical problems, embodiments of the present invention provide a network attack identification method, apparatus, device, storage medium and program product.

[0004] The network attack identification method provided in the embodiment of the present application includes:

[0005] Determine a network flow vector corresponding to the network flow data;

[0006] Performing overall feature extraction on the network traffic vector based on a multi-head attention mechanism to obtain a first feature vector;

[0007] Performing local feature extraction on the network traffic vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector;

[0008] Based on the first feature vector and the second feature vector, the network attack behavior is identified to obtain the target network attack category corresponding to the network traffic data.

[0009] The network attack identification device provided in the embodiment of the present application includes:

[0010] A determination unit, used to determine a network flow vector corresponding to the network flow data;

[0011] A first feature extraction unit, configured to perform overall feature extraction on the network traffic vector based on a multi-head attention mechanism to obtain a first feature vector;

[0012] A second feature extraction unit is used to extract local features of the network traffic vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector;

[0013] An identification unit is used to identify the network attack behavior based on the first feature vector and the second feature vector to obtain a target network attack category corresponding to the network traffic data.

[0014] The processing device provided in the embodiment of the present application includes: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute any one of the above-mentioned network attack identification methods.

[0015] The computer-readable storage medium provided in the embodiment of the present application is used to store a computer program, and the computer program enables a computer to execute any one of the above-mentioned network attack identification methods.

[0016] The computer program product provided in the embodiments of the present application includes computer program instructions, which enable a computer to execute any one of the above-mentioned network attack identification methods.

[0017] In the technical solution of the embodiment of the present application, by determining the network traffic vector corresponding to the network traffic data, and extracting the overall features of the network traffic vector based on the multi-head attention mechanism, a first feature vector is obtained, and by extracting the local features of the network traffic vector based on the association rule algorithm, a second feature vector is obtained, and the second feature vector is a soft label feature vector, so that based on the first feature vector and the second feature vector, the network attack behavior is identified, and the target network attack category corresponding to the network traffic data is obtained. In this way, from the overall and local perspectives, by using the multi-head attention mechanism for the network traffic, multiple self-attention heads are configured to extract features from different angles of the network traffic, which can effectively extract the overall key information of the network traffic, and reduce the influence of the interference information by reducing the weight of the interference information. At the same time, the column features of the network traffic are mined by association rules, and the clustered soft labels of each layer of iteration are used as new features of the network traffic. The characteristics of the network traffic are characterized from the perspective of the attack cluster, thereby introducing clustering features, which can expand the accuracy of network attack identification and improve the protection performance of the protection module. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of a network attack identification method provided in an embodiment of the present application;

[0019] Figure 2 It is a flowchart of a network attack identification method based on multi-head attention and multi-dimensional soft label clustering provided in an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of the structure of a network attack identification device provided in an embodiment of the present application;

[0021] Figure 4 It is a schematic diagram of the structure of the processing device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0023] In the description of the embodiments of the present application, the term "corresponding" may indicate a direct or indirect correspondence between two items, or an association relationship between the two items, or a relationship between indication and being indicated, configuration and being configured, and the like.

[0024] To facilitate understanding of the technical solutions of the embodiments of the present application, the relevant technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all belong to the protection scope of the embodiments of the present application.

[0025] In terms of defending against network attacks, there are two specific steps: detection and protection. The premise of protecting against network attacks is to detect the attack, identify the specific attack type, and then apply specific protection strategies. Among them, IDS is used to detect possible threats, prevent unauthorized access, and report attacks to security administrators. This research direction is mainly divided into two basic directions:

[0026] (1) Signature-based IDS, including signature matching and state matching methods. Signature matching methods rely on specific features in network traffic or system behavior, including protocol type, communication mode, request frequency, etc. State matching methods rely on the state transition rules in network traffic or system behavior, and detect potential attacks by analyzing the state of network connections, such as TCP / IP establishment, maintenance, and termination.

[0027] (2) IDS based on abnormal behavior detection can detect unknown attacks by learning network traffic behavior to classify traffic. The anomaly-based method, combined with machine learning and deep learning, collects data on network access requests and trains models for attack identification, which significantly improves the protection capability. Anomaly-based models include four types of models: anomaly statistics, protocol anomalies, traffic anomalies, and rule sets. Anomaly models have a high attack detection rate in their respective targeted areas.

[0028] Currently, most IDSs use a rule base combined with feature-based intrusion detection methods to identify network attacks. However, this technical solution still has the following problems:

[0029] (1) The rule base is limited. The rule base is excellent in terms of detection time and cost, but in the face of increasingly complex and intelligent network attacks, it is unable to identify attack variants that do not exist in the rule base, and its ability to protect against unknown types of attacks is weak.

[0030] (2) The selection of features in feature-based intrusion detection methods can be based on prior knowledge or automatically extracted through machine learning algorithms. Prior features depend on the expert's experience knowledge. If the experience knowledge cannot well characterize the characteristics of the attack, it will lead to a high false alarm rate. At present, feature extraction based on machine learning algorithms focuses on the characteristics of the data itself. Extracting data features from a holistic perspective will be subject to more interference information, and it also ignores the correlation between unknown attacks and known attacks that may exist in attack clusters.

[0031] To solve the above technical problems, this application proposes a network attack identification method based on multi-head attention combined with multi-dimensional soft label clustering, which mainly solves the following problems:

[0032] (1) A multi-head attention mechanism is used for network traffic. The attention weight corresponding to each attention head is calculated separately to extract different features of network traffic. By reducing the weight of interference information and thus reducing the impact of interference information, this method can more effectively extract the overall key information of network traffic.

[0033] (2) Using multi-dimensional soft label features, clustering features are introduced, and the apriori algorithm is used to mine association rules on the column features of the network traffic vector. The clustering soft labels of each layer of iteration are used as new features of the network traffic. The characteristics of the network traffic are characterized from the perspective of attack clusters, thereby introducing clustering features, expanding the field of view of the network attack identification system and improving the identification accuracy.

[0034] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The above related technologies can be combined arbitrarily with the technical solutions of the embodiments of the present application as optional solutions, and they all belong to the protection scope of the embodiments of the present application. The embodiments of the present application include at least part of the following contents.

[0035] This application embodiment proposes a network attack identification method. Figure 1 is a flow chart of a network attack identification method provided in an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:

[0036] Step 101: Determine a network traffic vector corresponding to network traffic data.

[0037] In an embodiment of the present application, network traffic data can be obtained by real-time capture of data packets of the interface, or network traffic data can be obtained through other traffic collection methods such as Wireshark, tcpdump and other network traffic collection tools; after obtaining the network traffic data, the network traffic data is cleaned to remove invalid or abnormal data to ensure the quality and consistency of the data, and useful features are extracted from the cleaned data, including source IP address, destination PI address, source port, destination port, protocol type, packet length, timestamp, etc., and the extracted features are then standardized, such as normalized or standardized, so that they have a unified dimension and range to facilitate subsequent model processing; the extracted features are then combined into an initial feature vector as a network traffic vector, for example, the source IP address, destination PI address, source port, destination port, protocol type, etc. are used as vector elements.

[0038] Here, if the time series characteristics of the traffic need to be considered, the network traffic data can also be arranged in chronological order to construct a network traffic vector, for example, the traffic characteristics in each time window are used as an element in the network traffic vector.

[0039] Here, if the network traffic vector contains a large number of protocol fields or specific identifiers (such as IP addresses, port numbers, etc.), these discrete features can be converted into continuous vector representations through word embedding, so as to better capture the correlation between features. When performing anomaly detection on network traffic data, word embedding can be used to help understand the semantic relationship between different traffic features and improve the accuracy of classification detection.

[0040] Here, if the network traffic vector has time series characteristics, such as the temporal changes in traffic, the order of protocol interactions, etc., a position vector can be assigned to each element in the network traffic vector through position embedding to represent the position information of each element in the network traffic vector, thereby more effectively capturing the time sequence information in the network traffic data. When performing anomaly detection on network traffic data, position embedding can be used to help understand the time dependency and sequential relationship of network traffic data.

[0041] Step 102: Perform overall feature extraction on the network traffic vector based on the multi-head attention mechanism to obtain a first feature vector.

[0042] In the embodiment of the present application, in order to avoid the self-attention mechanism focusing on a certain feature, multi-head attention can be used to extract the overall features of the network traffic vector, that is, by using multiple self-attention modules, each self-attention module extracts features from different angles in the network traffic vector, and then splices these features from different angles and linearly projects them to obtain the final output, thereby obtaining the overall feature vector of the extracted network traffic vector, that is, the first feature vector. Among them, the multi-head attention mechanism captures the diversity of network traffic data from different angles by processing the results of multiple attention heads in parallel, further enhancing the learning and expression capabilities of the model.

[0043] In some implementations, step 102 may specifically include:

[0044] Perform linear transformation on the network traffic vector to obtain the corresponding query vector, key vector and value vector;

[0045] Based on the query vector, the key vector, and the value vector, a first feature vector is determined.

[0046] Here, the network traffic vector can be linearly transformed through three different linear transformation layers (fully connected layers) to obtain a query vector, a key vector and a value vector, respectively. The query vector, the key vector and the value vector are processed separately to obtain multiple attention heads (multiple subspaces), each head having different linear transformation parameters, so that features of the network traffic vector can be extracted from different angles through multiple attention heads, and the feature vector obtained by merging the extracted features from different angles is the first feature vector.

[0047] In some implementations, determining the first feature vector based on the query vector, the key vector, and the value vector includes:

[0048] Divide the query vector, key vector, and value vector into multiple attention modules;

[0049] Based on each of the multiple attention modules, performing attention calculation on the query vector, the key vector, and the value vector in each attention module to obtain a first sub-feature vector output by each attention module;

[0050] Based on the first sub-feature vector corresponding to each attention module, a first feature vector is determined.

[0051] Here, multiple attention heads, i.e., multiple attention modules, process the query vector, key vector, and value vector separately to obtain multiple attention heads, i.e., divide the query vector, key vector, and value vector into multiple attention modules.

[0052] Among them, for each attention module, a scaled dot product attention operation is performed on the network traffic vector. Specifically, the dot product result of the query vector and the key vector in each attention module is first calculated, and then the dot product result is scaled and normalized using the softmax function to obtain the attention weight corresponding to each attention module, and then the attention weight corresponding to each attention module is weightedly operated with the value vector to obtain the first sub-feature vector output by each attention module; after obtaining the first sub-feature vector corresponding to each attention module, the first sub-feature vector corresponding to each attention module can be concatenated and merged to form a long vector, and the long vector is linearly transformed to integrate information from different attention modules to obtain the final multi-head attention output, namely the first eigenvector.

[0053] Step 103: extract local features from the network traffic vector based on an association rule algorithm to obtain a second feature vector.

[0054] In the embodiment of the present application, inspired by the idea of ​​"label is representation" and "matrix column is the representation space of the corresponding category" in contrast clustering, the column in the feature matrix is ​​a special category representation, corresponding to the probability that a certain instance belongs to a certain category. Based on this, the clustering soft label feature vector corresponding to each column vector in the network traffic vector can be mined in combination with the association rule algorithm and the clustering algorithm, and the clustering soft label feature vector that meets the support standard can be extracted from it, thereby obtaining the second feature vector, which is the soft label feature vector. Among them, the association rule algorithm includes the Apriori algorithm, the association rule tree algorithm, the association rule network algorithm, etc.

[0055] Here, soft labels use probability distribution to represent the possibility that a data sample belongs to each category. For example, in a three-category task, the soft labels of a data sample can be 0.3, 0.5, and 0.2, indicating that the probability that the data sample belongs to the three categories is 30%, 50%, and 20%, respectively. The soft label feature vector is a way of integrating soft label information into the feature vector. Specifically, the feature vector of each data point contains not only the original feature value, but also its corresponding soft label information. For example, the original feature vector of a data point is [x1,x2,x3,...,x n ], whose soft labels are [p1,p2,p3,...,p k ] (indicates the probability of belonging to k categories), then its soft label feature vector can be expressed as [x1,x2,x3,...,x n ,p1,p2,p3,...,p k ].

[0056] In some implementations, step 103 may specifically include:

[0057] Determine the frequent itemsets based on the column subscripts corresponding to each column vector in the network traffic vector;

[0058] Based on the association rule algorithm, local features of frequent itemsets are extracted to obtain the second feature vector.

[0059] Here, the column subscript corresponding to each column vector in the network traffic vector can be used as a frequent item set, and local features of the frequent item set are extracted based on the association rule algorithm. Specifically, the clustering evaluation index value and the clustering soft label feature vector corresponding to each column vector in the frequent item set are first determined, and then the support of the evaluation index value corresponding to each column vector is evaluated, and the column vectors that meet the support standard (clustering evaluation index threshold) are screened out, and then the clustering soft label feature vectors corresponding to the column vectors that meet the support standard are combined, so that the second feature vector can be finally obtained.

[0060] In some implementations, performing local feature extraction on frequent item sets based on an association rule algorithm to obtain a second feature vector includes:

[0061] Determine the frequent item set as the frequent k-item set, and the initial value of k is 0;

[0062] Iteratively executing the first process based on the frequent k-item set to obtain the second eigenvector; the iterative termination condition of the first process is: the frequent k-item set is empty, or the frequent k-item set includes a column subscript;

[0063] The first process includes:

[0064] Cluster the column vectors corresponding to each column subscript in the frequent k-item set to obtain the clustering evaluation index value and clustering soft label feature vector corresponding to each column vector;

[0065] For the column vector corresponding to each column subscript in the frequent k-item set, if it is judged that the clustering evaluation index value corresponding to the column vector is less than the clustering evaluation index threshold, the column subscript corresponding to the column vector is removed from the frequent k-item set to obtain the frequent k+1-item set;

[0066] Determine the soft label k+1 set based on the cluster soft label feature vector corresponding to the column vector corresponding to each column subscript in the frequent k+1 item set;

[0067] Increase the value of k by 1.

[0068] Specifically, firstly, the frequent item set is determined as a frequent k-item set, and the initial value of k is 0. The column vector corresponding to each column subscript in the frequent k-item set is clustered by the clustering algorithm to obtain the probability distribution of each data point in each column vector belonging to each cluster. Based on the probability distribution of each data point in each column vector belonging to each cluster, the clustering evaluation index value and the clustering soft label feature vector corresponding to each column vector can be determined; secondly, the clustering evaluation index threshold is set as the support threshold, and it is judged whether the clustering evaluation index value corresponding to each column vector meets the clustering evaluation index threshold. If not, the column subscript corresponding to the column vector that does not meet the support threshold is Eliminate from the frequent k-item set to obtain the frequent k+1-item set; then determine the clustering soft label feature vector corresponding to the column vector corresponding to each column subscript in the frequent k+1-item set as the soft label k+1 set, and add 1 to the value of k; and so on, continuously iterate the soft label extraction of the frequent k-item set through the above steps until the frequent k-item set is empty, or the frequent k-item set includes only one number of column subscripts. If the frequent k-item set is empty, the soft label set obtained by the previous iteration is determined as the second feature vector. If the frequent k-item set includes only one number of column subscripts, the soft label set obtained in this iteration is determined as the second feature vector.

[0069] Among them, the clustering soft label feature vector corresponding to the column vector includes the data points in the column vector and the probability information of the clusters to which the data points belong. The probability information of the clusters to which the data points belong refers to generating a vector for the data points in the column vector in the process of clustering the column vector. The vector represents the probability distribution of the data points belonging to each cluster.

[0070] Step 104: Based on the first feature vector and the second feature vector, the network attack behavior is identified to obtain a target network attack category corresponding to the network traffic data.

[0071] In some implementations, step 104 may specifically include:

[0072] Performing principal component analysis on the first eigenvector to obtain a preset number of principal components in the first eigenvector;

[0073] Determine a projection matrix based on a preset number of principal components, and project the first eigenvector based on the projection matrix to obtain a third eigenvector;

[0074] Perform regression fitting analysis on the second eigenvector to obtain the fourth eigenvector;

[0075] Based on the third eigenvector and the fourth eigenvector, the network attack behavior is identified to obtain the target network attack category.

[0076] Here, the principal component analysis (PCA) algorithm can be used to perform principal component analysis on the first eigenvector. Specifically, the covariance matrix of the first eigenvector is first calculated to represent the linear relationship between the features. Secondly, the covariance matrix is ​​decomposed to obtain a set of eigenvalues ​​and eigenvectors corresponding to the covariance matrix, wherein each eigenvalue and its corresponding eigenvector define a principal component in the first eigenvector, the size of the eigenvalue represents the proportion of the data variance that the principal component can explain, and the eigenvector represents the direction of the principal component. Then, based on the size sorting of this group of eigenvalues, the eigenvectors corresponding to the largest eigenvalues ​​of the first preset number (for example, the first k) are selected as the principal component, wherein the first principal component The eigenvector corresponding to the first principal component has the largest eigenvalue, which can capture the maximum variance of the network traffic data. The second principal component is orthogonal to the other principal components and can capture the second-level size variance of the network traffic data. By analogy, the kth principal component is orthogonal to the other principal components and can capture the kth size variance of the network traffic data. Finally, the eigenvectors corresponding to the preset number of principal components are used as column vectors to construct a projection matrix, and the first eigenvector is multiplied by the projection matrix to project the first eigenvector into the principal component space, thereby obtaining the projection vector of the first eigenvector in the principal component space. The projection vector is a k-dimensional vector, which represents the value of the first eigenvector on the k principal components. The projection vector is the third eigenvector.

[0077] Optionally, the feature importance of the first eigenvector can also be evaluated through a machine learning model (such as a decision tree model, a random forest model, etc.) to obtain the feature importance score of each sub-vector in the first eigenvector, and the feature importance score of each sub-vector in the first eigenvector is sorted in order from large to small, and the first preset number (for example, the first k) largest (with the largest feature importance score) sub-vectors are selected therefrom, thereby determining the third eigenvector based on the preset number of largest sub-vectors.

[0078] Here, the partial least squares regression (PLS) algorithm can be used to perform regression fitting analysis on the second eigenvector to extract a low-dimensional representation of the second eigenvector. Specifically, since the second eigenvector is a soft label eigenvector, the PLS algorithm will find a set of latent variables (or PLS characteristic factors) by maximizing the covariance (correlation) between the original eigenvector and the soft label in the second eigenvector. This set of latent variables is the fourth eigenvector.

[0079] Optionally, the fourth eigenvector can also be determined by calculating the correlation between the original eigenvector and the soft label vector in the second eigenvector through correlation analysis (such as Pearson correlation coefficient). Specifically, the first mean corresponding to each sub-vector in the original eigenvector and the second mean corresponding to the soft label vector can be calculated based on the sample data corresponding to the original eigenvector and the soft label vector in the second eigenvector, and the first deviation of each sample data in each sub-vector and the second deviation of each sample data in the soft label vector can be calculated based on the first mean and the second mean. Then, the fourth eigenvector is calculated based on the first deviation of each sample data in each sub-vector and the second deviation of each sample data in the soft label vector. Deviation, calculate the Pearson correlation coefficient between each sub-vector and the soft label vector; wherein, the size range of the Pearson correlation coefficient is [-1, 1]. If the coefficient is close to -1, it indicates that there is a strong positive correlation between the sub-vector and the soft label vector, otherwise there is a strong negative correlation. If the coefficient is close to 0, it indicates that there is no obvious linear relationship between the sub-vector and the soft label vector. After obtaining the Pearson correlation coefficient between each sub-vector and the soft label vector in the original feature vector, the sub-vectors with the largest (largest Pearson correlation coefficient) number of the first preset number (for example, the first k) can be selected therefrom, thereby determining the fourth feature vector based on the preset number of sub-vectors and the soft label vector.

[0080] After obtaining the third eigenvector and the fourth eigenvector, the third eigenvector and the fourth eigenvector can be input into the classifier as input data, and the network attack behavior corresponding to the network traffic data can be identified by the classifier to obtain the target network attack category.

[0081] In some implementations, based on the third feature vector and the fourth feature vector, the network attack behavior is identified, and the target network attack category is obtained, including:

[0082] Merging the third eigenvector and the fourth eigenvector to obtain a fused eigenvector;

[0083] Based on the classification model, the fused feature vector is classified and calculated to obtain the probability that the fused feature vector belongs to each network attack category in multiple network attack categories;

[0084] Based on the probability that the fused feature vector belongs to each network attack category, the network attack category corresponding to the maximum probability value is selected as the target network attack category.

[0085] Specifically, the third eigenvector and the fourth eigenvector can be expanded and merged in the column direction to obtain a fused eigenvector, and the fused eigenvector is input into the classifier as input data. The classifier identifies the network attack category of the network traffic data based on the fused eigenvector, and outputs the probability that the fused eigenvector belongs to each network attack category in multiple network attack categories. The network attack category corresponding to the maximum probability value is selected from the probability that the fused eigenvector belongs to each network attack category as the target network attack category, and the target network attack category is the predicted network attack category of the network traffic data. Among them, the classifier is a pre-trained model, which can identify the probability of belonging to each category from the input data after training, and uses the category corresponding to the highest probability as the output category. Specifically, the classifier can be a classification model such as a neural network, a random forest, a logistic regression, or a hybrid model composed of different models, which is not limited here.

[0086] In the technical solution of the embodiment of the present application, by determining the network traffic vector corresponding to the network traffic data, and extracting the overall features of the network traffic vector based on the multi-head attention mechanism, a first feature vector is obtained, and by extracting the local features of the network traffic vector based on the association rule algorithm, a second feature vector is obtained, and the second feature vector is a soft label feature vector, so that based on the first feature vector and the second feature vector, the network attack behavior is identified, and the target network attack category corresponding to the network traffic data is obtained. In this way, from the overall and local perspectives, by using the multi-head attention mechanism for the network traffic, multiple self-attention heads are configured to extract features from different angles of the network traffic, which can effectively extract the overall key information of the network traffic, and reduce the influence of the interference information by reducing the weight of the interference information. At the same time, the column features of the network traffic are mined by association rules, and the clustered soft labels of each layer of iteration are used as new features of the network traffic. The characteristics of the network traffic are characterized from the perspective of the attack cluster, thereby introducing clustering features, which can expand the accuracy of network attack identification and improve the protection performance of the protection module.

[0087] The embodiment of the present application also proposes a network attack identification method based on multi-head attention and multi-dimensional soft label clustering, which mainly includes three key parts: multi-head attention extraction, multi-soft label feature construction and classification identification. Figure 2 is a flow chart of a network attack identification method based on multi-head attention and multi-dimensional soft label clustering provided in an embodiment of the present application, such as Figure 2 As shown, the method comprises the following steps:

[0088] Step 201: Obtain a network traffic vector.

[0089] The network traffic data is obtained by real-time capture of data packets of the interface, or by other traffic collection methods such as Wireshark, tcpdump and other network traffic collection tools; after obtaining the network traffic data, the network traffic data is cleaned to remove invalid or abnormal data to ensure the quality and consistency of the data, and useful features are extracted from the cleaned data, including source IP address, destination PI address, source port, destination port, protocol type, packet length, timestamp, etc., and the extracted features are standardized, such as normalization or standardization, so that they have a unified dimension and range to facilitate subsequent model processing; the extracted features are then combined into an initial feature vector as the network traffic vector, for example, the source IP address, destination PI address, source port, destination port, protocol type, etc. are used as vector elements.

[0090] If it is necessary to consider the time characteristics of the traffic, the network traffic data can be arranged in chronological order based on the time series characteristics of the traffic to construct a network traffic vector, for example, the traffic characteristics in each time window can be used as an element in the network traffic vector.

[0091] If the network traffic vector contains a large number of protocol fields or specific identifiers (such as IP addresses, port numbers, etc.), these discrete features can be converted into continuous vector representations through word embedding, so as to better capture the correlation between features. When performing anomaly detection on network traffic data, word embedding can be used to help understand the semantic relationship between different traffic features and improve the accuracy of classification detection.

[0092] If the network traffic vector has time series characteristics, such as the temporal changes in traffic, the order of protocol interactions, etc., a position vector can be assigned to each element in the network traffic vector through position embedding to represent the position information of each element in the network traffic vector, thereby more effectively capturing the time sequence information in the network traffic data. When performing anomaly detection on network traffic data, position embedding can be used to help understand the time dependency and sequential relationship of network traffic data.

[0093] Step 202: Extract features of the network traffic vector through a multi-attention mechanism to obtain a multi-head attention feature vector.

[0094] The self-attention mechanism is to obtain the attention weight of each element in the sequence through calculation, and then add each element to the corresponding attention weight to obtain a new sequence extracted by the self-attention mechanism. In order to avoid the self-attention mechanism focusing on a certain feature, multi-head attention is used to extract the features of the network traffic vector. Multiple self-attention modules are used, and each self-attention module extracts different features. These features are then spliced ​​and linearly projected to obtain the final output, which is used as the feature vector of the network traffic vector. From an overall perspective, the multi-head attention mechanism is used to obtain data features from different angles of network traffic.

[0095] Assume that the input network traffic vector X = [x1, x2, x3, ..., x n ] T , whose feature vector X output by the multi-head attention mechanism mh =[x1′,x2′,x3′,...,x n ′] T , x i and x i ′ respectively represent the input value and output value corresponding to the i-th feature item in the network traffic vector. Among them, the output feature vector is formalized as follows:

[0096] For the i-th attention head i The extracted features can be obtained through Q i ,K i ,V i The matrix is ​​obtained according to the following formula:

[0097] Q=W q X,K=W k X,V=W v X(1)

[0098]

[0099]

[0100] in, represents the learned projection matrix, Q, K, V represent query, key, and value vectors respectively, and W q , W k , W v represents the linear transformation matrix, d k Indicates Q i and K i The matrix dimensions of .

[0101] Through the above formulas (1), (2), and (3), we can get multiple attention heads i Finally, multiple attention heads are merged and linearly transformed to obtain the final output, and its formal formula is as follows:

[0102] X mh =Concat(head1,...,head h )W O (4)

[0103] Among them, Concat(·) refers to the concatenation operation, W O Represents a linear transformation matrix.

[0104] Step 203: Use PCA to extract the principal components in the multi-head attention feature vector, and project the multi-head attention feature vector into the extracted principal components to obtain a projection vector.

[0105] For the multi-head attention feature vector, the features extracted by different attention heads express key information to different degrees. Based on this, the PCA principal component analysis method can be used to calculate the correlation between the features in the multi-head attention feature vector and extract the principal components in the multi-head attention feature vector.

[0106] Then, a projection matrix is ​​constructed based on the extracted principal components, and the multi-head attention feature vector is multiplied by the projection matrix to obtain the projection vector of the multi-head attention feature vector in the principal component space. The calculation formula is as follows:

[0107] W PCA = PCA(X mh )(5)

[0108] X mp =W PCA ·X mh (6)

[0109] Where PCA(·) represents the PCA algorithm, W PCA represents the projection matrix constructed by the principal component of the multi-head attention feature vector obtained by the PCA algorithm, X mp Represents the projection vector obtained by projecting the multi-head attention feature vector on the projection matrix, and the projection vector represents the value of the multi-head attention feature vector on the principal component.

[0110] Step 204: construct a multi-dimensional soft label feature for the column vector in the network traffic vector using the Apriori algorithm to obtain a soft label feature vector.

[0111] Inspired by the idea of ​​"labels are representations" and "matrix columns are representation spaces of corresponding categories" in contrastive clustering, the columns in the feature matrix are special category representations, corresponding to the probability that a certain instance belongs to a certain category. Based on this, a method of constructing feature vectors with clustering soft labels is proposed. The column subscripts of the column vectors in the network traffic vector are used as frequent itemsets. The clustering algorithm is used to cluster the column features corresponding to the column subscripts in the frequent itemsets to obtain the corresponding clustering evaluation index value, and the clustering evaluation index threshold is set as the support threshold.

[0112] The specific construction method is as follows:

[0113] (1) For the input network traffic vector X = [x1, x2, x3, ..., x n ] T , set the frequent item set D0 = [1,2,3...,n] as all column subscripts in X, the clustering evaluation index threshold is α0, and the soft label set is T k , determine the frequent item set D0 as the frequent k-item set D k , the initial value of k is 0.

[0114] (2) For each column subscript corresponding to the frequent k-item set, i Apply the clustering algorithm to obtain the corresponding clustering evaluation index value α i And the corresponding cluster soft label feature vector, if α i If the α0 condition is not met, the column subscript i is removed from the frequent k-item set D k We eliminate the frequent k+1 item set D k+1 If the frequent k-item set is empty, the soft label k-1 set T is returned directly. k-1 As a result of the algorithm, if there is only one frequent k-item set, the soft label k set T is directly returned. k As the algorithm result.

[0115] (3) The frequent k+1 item set D k+1 The cluster soft label feature vector corresponding to each column subscript in is determined as the soft label k+1 set T k+1 .

[0116] (4) k = k + 1, go to steps (2) and (3).

[0117] Through the above algorithm steps, the soft label feature vector corresponding to the network traffic vector can be obtained. The number of features increases from small to large, and the soft label features are gradually constructed to continuously enrich the feature dimensions of the network traffic vector. Among them, for the clustering algorithm in the above step (2), a suitable clustering algorithm can be selected according to the characteristics of the data, or the K-means clustering algorithm can be used.

[0118] Step 205: Perform PLS regression analysis on the soft label feature vector to extract the principal component feature vector in the soft label feature vector.

[0119] Since the soft label feature vector is obtained by clustering the feature columns from few to many, there is also correlation between the soft label features. Based on this, in order to further improve the representativeness of the features, the PLS algorithm can be used to perform regression fitting analysis on the soft label feature vector to extract the low-dimensional representation of the soft label feature vector. The low-dimensional representation is the principal component feature vector in the soft label feature vector, and its calculation formula is as follows:

[0120] X pls =PLS(T)(7)

[0121] Where PLS(·) represents the PLS algorithm, T represents the soft label feature vector, X pls Represents the PLS feature factor vector after data dimensionality reduction by the PLS algorithm. The PLS algorithm can solve the multicollinearity problem in the soft label feature vector, avoid the overfitting phenomenon caused by the correlation between features, and make the classification model more generalizable.

[0122] Step 206: Merge the multi-head attention weighted feature vector and the principal component feature vector to obtain a fused feature vector.

[0123] The multi-head attention weighted feature vector and the principal component feature vector in the soft label feature vector are expanded and merged column-wise to obtain the fused feature vector as the input data X in , and its calculation formula is as follows:

[0124] X in =[X mp X pls ](8)

[0125] Step 207: Input the fused feature vector into the classifier, and use the classifier to identify network attacks on the fused feature vector.

[0126] Step 208: The classifier outputs the probability that the fused feature vector belongs to different network attack categories.

[0127] Input data X in Input the classifier, identify the network attack through the classifier, calculate the probability of the corresponding category of the fused feature vector, and output the probability of belonging to different network attack categories. The network attack category corresponding to the highest probability is used as the network attack category corresponding to the network traffic. The probability calculation formula for each network attack category is as follows:

[0128]

[0129] label = argmax yi (10)

[0130] Among them, x i Represents the output data of the upper layer neurons, y i It represents the probability of belonging to each category, and label represents the category label corresponding to the highest probability.

[0131] In the technical solution of the embodiment of the present application, starting from the overall and local perspectives respectively, by using a multi-head attention mechanism on the network traffic, multiple self-attention heads are configured to extract features from different angles of the network traffic, which can effectively extract the overall key information of the network traffic and reduce the impact of interference information by reducing the weight of the interference information. At the same time, the column features of the network traffic are mined for association rules, and the clustering soft labels of each layer of iteration are used as new features of the network traffic. The characteristics of the network traffic are characterized from the perspective of attack clusters, thereby introducing clustering features, which can expand the accuracy of network attack identification and improve the protection performance of the protection module.

[0132] The present application also provides a network attack identification device. Figure 3 is a schematic diagram of the structure of a network attack identification device provided in an embodiment of the present application, such as Figure 3 As shown, the device comprises:

[0133] The determining unit 301 is used to determine a network flow vector corresponding to the network flow data.

[0134] The first feature extraction unit 302 is used to perform overall feature extraction on the network traffic vector based on a multi-head attention mechanism to obtain a first feature vector.

[0135] The second feature extraction unit 303 is used to extract local features from the network traffic vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector.

[0136] The identification unit 304 is used to identify the network attack behavior based on the first feature vector and the second feature vector to obtain the target network attack category corresponding to the network traffic data.

[0137] In some implementations, the first feature extraction unit 302 is specifically used to: perform a linear transformation on the network traffic vector to obtain a corresponding query vector, a key vector, and a value vector; and determine a first feature vector based on the query vector, the key vector, and the value vector.

[0138] In some embodiments, the first feature extraction unit 302 is further specifically used to: divide the query vector, key vector, and value vector into multiple attention modules; based on each attention module in the multiple attention modules, perform attention calculation on the query vector, key vector, and value vector in each attention module to obtain the first sub-feature vector output by each attention module; based on the first sub-feature vector corresponding to each attention module, determine the first feature vector.

[0139] In some implementations, the second feature extraction unit 301 is specifically used to: determine a frequent item set based on a column subscript corresponding to each column vector in the network traffic vector; and extract local features of the frequent item set based on an association rule algorithm to obtain a second feature vector.

[0140] In some implementations, the second feature extraction unit 301 is further specifically configured to: determine the frequent item set as a frequent k-item set, where the initial value of k is 0; iteratively execute the first process based on the frequent k-item set to obtain a second feature vector; and the iteration termination condition of the first process is: the frequent k-item set is empty, or the frequent k-item set includes a column subscript;

[0141] Among them, the first process includes: clustering the column vectors corresponding to each column subscript in the frequent k-item set to obtain the clustering evaluation index value and clustering soft label feature vector corresponding to each column vector; for the column vector corresponding to each column subscript in the frequent k-item set, if it is judged that the clustering evaluation index value corresponding to the column vector is less than the clustering evaluation index threshold, then the column subscript corresponding to the column vector is removed from the frequent k-item set to obtain the frequent k+1-item set; based on the clustering soft label feature vector corresponding to the column vector corresponding to each column subscript in the frequent k+1-item set, determine the soft label k+1 set; add 1 to the value of k.

[0142] In some embodiments, the identification unit 304 is specifically used to: perform principal component analysis on the first eigenvector to obtain a preset number of principal components in the first eigenvector; determine a projection matrix based on the preset number of principal components, and project the first eigenvector based on the projection matrix to obtain a third eigenvector; perform regression fitting analysis on the second eigenvector to obtain a fourth eigenvector; and identify network attack behavior based on the third eigenvector and the fourth eigenvector to obtain a target network attack category.

[0143] In some embodiments, the identification unit 304 is further specifically used to: merge the third feature vector and the fourth feature vector to obtain a fused feature vector; perform classification calculations on the fused feature vector based on a classification model to obtain the probability that the fused feature vector belongs to each network attack category in multiple network attack categories; based on the probability that the fused feature vector belongs to each network attack category, select the network attack category corresponding to the maximum probability value as the target network attack category.

[0144] In the technical solution of the embodiment of the present application, by determining the network traffic vector corresponding to the network traffic data, and extracting the overall features of the network traffic vector based on the multi-head attention mechanism, a first feature vector is obtained, and by extracting the local features of the network traffic vector based on the association rule algorithm, a second feature vector is obtained, and the second feature vector is a soft label feature vector, so that based on the first feature vector and the second feature vector, the network attack behavior is identified, and the target network attack category corresponding to the network traffic data is obtained. In this way, from the overall and local perspectives, by using the multi-head attention mechanism for the network traffic, multiple self-attention heads are configured to extract features from different angles of the network traffic, which can effectively extract the overall key information of the network traffic, and reduce the influence of the interference information by reducing the weight of the interference information. At the same time, the column features of the network traffic are mined by association rules, and the clustered soft labels of each layer of iteration are used as new features of the network traffic. The characteristics of the network traffic are characterized from the perspective of the attack cluster, thereby introducing clustering features, which can expand the accuracy of network attack identification and improve the protection performance of the protection module.

[0145] Those skilled in the art should understand that Figure 3 The implementation functions of each unit in the network attack identification device shown can be understood by referring to the relevant description of the aforementioned method. Figure 3 The functions of each unit in the network attack identification device shown can be implemented by a program running on a processor, or by a specific logic circuit.

[0146] Figure 4 is a schematic diagram of the structure of a processing device provided in an embodiment of the present application. The processing device may be a terminal device or a network device. Figure 4 The processing device shown includes a processor 401, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.

[0147] Alternatively, if Figure 4 As shown, the processing device may further include a memory 402. The processor 401 may call and run a computer program from the memory 402 to implement the method in the embodiment of the present application.

[0148] The memory 402 may be a separate device independent of the processor 401 , or may be integrated into the processor 401 .

[0149] Alternatively, if Figure 4 As shown, the processing device may further include a transceiver 403, and the processor 401 may control the transceiver 403 to communicate with other devices, specifically, may send information or data to other devices, or receive information or data sent by other devices.

[0150] The transceiver 403 may include a transmitter and a receiver. The transceiver 403 may further include an antenna, and the number of the antennas may be one or more.

[0151] The processing device may specifically be a network attack identification device of an embodiment of the present application, and the processing device may implement the corresponding processes implemented by each method of the embodiment of the present application, which will not be described in detail here for the sake of brevity.

[0152] It should be understood that the processor of the embodiment of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by the hardware integrated logic circuit or software instructions in the processor. The above processor can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and performed. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0153] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0154] It should be understood that the above-mentioned memory is exemplary but not restrictive. For example, the memory in the embodiments of the present application may also be static random access memory (static RAM, SRAM), dynamic random access memory (dynamic RAM, DRAM), synchronous dynamic random access memory (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (synch link DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DR RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not limited to these and any other suitable types of memory.

[0155] The embodiment of the present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to the processing device in the embodiment of the present application, and the computer program enables the computer to execute the corresponding process of each method implemented in the embodiment of the present application, which will not be described here for the sake of brevity.

[0156] The embodiment of the present application also provides a computer program product, including computer program instructions. The computer program product can be applied to the processing device in the embodiment of the present application, and the computer program instructions enable the computer to execute the corresponding process of each method implemented in the embodiment of the present application, which will not be described here for the sake of brevity.

[0157] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0158] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0159] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0160] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0162] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory,) ROM, random access memory (RandomAccess Memory, RAM), disk or optical disk and other media that can store program codes.

[0163] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A network attack identification method, characterized in that: The method comprises: Determine a network flow vector corresponding to the network flow data; Performing overall feature extraction on the network traffic vector based on a multi-head attention mechanism to obtain a first feature vector; Performing local feature extraction on the network traffic vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector; Based on the first feature vector and the second feature vector, the network attack behavior is identified to obtain the target network attack category corresponding to the network traffic data.

2. The method according to claim 1, characterized in that The overall feature extraction of the network traffic vector based on the multi-head attention mechanism to obtain a first feature vector includes: Performing a linear transformation on the network traffic vector to obtain a corresponding query vector, a key vector, and a value vector; The first feature vector is determined based on the query vector, the key vector, and the value vector.

3. The method according to claim 2, characterized in that The determining the first feature vector based on the query vector, the key vector, and the value vector includes: dividing the query vector, the key vector, and the value vector into a plurality of attention modules; Based on each of the multiple attention modules, perform attention calculation on the query vector, the key vector, and the value vector in each of the attention modules to obtain a first sub-feature vector output by each of the attention modules; Based on the first sub-feature vector corresponding to each attention module, the first feature vector is determined.

4. The method according to claim 1, characterized in that: The extracting local features of the network traffic vector based on the association rule algorithm to obtain a second feature vector includes: Determining frequent itemsets based on the column subscript corresponding to each column vector in the network traffic vector; Perform local feature extraction on the frequent item set based on an association rule algorithm to obtain the second feature vector.

5. The method according to claim 4, characterized in that The extracting local features of the frequent item set based on the association rule algorithm to obtain the second feature vector includes: Determine the frequent item set as a frequent k-item set, where the initial value of k is 0; Iteratively executing a first process based on the frequent k-item set to obtain the second feature vector; an iterative termination condition of the first process is: the frequent k-item set is empty, or the frequent k-item set includes a column subscript; The first process includes: Clustering the column vectors corresponding to each column subscript in the frequent k-item set to obtain a clustering evaluation index value and a clustering soft label feature vector corresponding to each column vector; For the column vector corresponding to each column subscript in the frequent k-item set, if it is determined that the clustering evaluation index value corresponding to the column vector is less than the clustering evaluation index threshold, the column subscript corresponding to the column vector is removed from the frequent k-item set to obtain a frequent k+1-item set; Determine a soft label k+1 set based on the cluster soft label feature vector corresponding to the column vector corresponding to each column subscript in the frequent k+1 item set; The value of k is increased by 1.

6. The method according to claim 1, characterized in that The identifying the network attack behavior based on the first feature vector and the second feature vector to obtain the target network attack category corresponding to the network traffic data includes: Performing principal component analysis on the first eigenvector to obtain a preset number of principal components in the first eigenvector; Determine a projection matrix based on the preset number of principal components, and project the first eigenvector based on the projection matrix to obtain a third eigenvector; Performing regression fitting analysis on the second eigenvector to obtain a fourth eigenvector; Based on the third feature vector and the fourth feature vector, the network attack behavior is identified to obtain the target network attack category.

7. The method according to claim 6, characterized in that The identifying the network attack behavior based on the third feature vector and the fourth feature vector to obtain the target network attack category includes: Merging the third eigenvector and the fourth eigenvector to obtain a fused eigenvector; Performing classification calculation on the fused feature vector based on a classification model to obtain the probability that the fused feature vector belongs to each network attack category in multiple network attack categories; Based on the probability that the fused feature vector belongs to each network attack category, the network attack category corresponding to the maximum probability value is selected as the target network attack category.

8. A network attack identification device, characterized in that: The device comprises: A determination unit, used to determine a network flow vector corresponding to the network flow data; A first feature extraction unit, configured to perform overall feature extraction on the network traffic vector based on a multi-head attention mechanism to obtain a first feature vector; A second feature extraction unit is used to extract local features of the network traffic vector based on an association rule algorithm to obtain a second feature vector; the second feature vector is a soft label feature vector; An identification unit is used to identify the network attack behavior based on the first feature vector and the second feature vector to obtain a target network attack category corresponding to the network traffic data.

9. A processing device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The method comprises computer program instructions which cause a computer to execute the method as claimed in any one of claims 1 to 7.