Intrusion detection model based on double-flow network architecture
Through the intrusion detection model based on the dual-stream network architecture, combined with feature extraction of local feature sequences and network topology relationships, and using the adaptive sparse attention mechanism, the insufficient detection of the existing model in complex attack scenarios is solved, and more efficient intrusion detection effect is achieved.
Patent Information
- Application Number
- CN202510500120.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-01
AI Technical Summary
The existing intrusion detection model is difficult to take into account the topological relationship between devices and the local mode of traffic data when dealing with complex attack scenarios, resulting in insufficient detection accuracy and efficiency.
The intrusion detection model based on the dual-stream network architecture is adopted, and the local feature sequence and heterogeneous graph structure data are obtained through the preprocessing module, feature extraction is performed in combination with CNN and GNN branches, and attention weight is generated using the adaptive sparse attention mechanism to realize global and local collaborative decision-making.
The detection capability of long-term attack mode is improved, the detection efficiency and accuracy in complex environments are enhanced, and the coverage capability of intrusion detection is optimized.
Smart Images

Figure CN120238362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and more specifically, to an intrusion detection model based on a dual-stream network architecture. Background Art
[0002] Currently, in the era of the Internet of Everything, Internet of Things (IoT) technology is changing our lives and work at an alarming rate. However, due to its inherent resource constraints (such as computing power, storage space, and energy consumption constraints), IoT devices face severe challenges in terms of security protection [1]. Intrusion detection, as an important part of network security, is crucial for timely detecting and responding to network attacks. Although deep learning technology has been widely applied to network traffic analysis, existing methods still have significant limitations in dealing with complex attack scenarios.
[0003] Currently, the mainstream intrusion detection models are mostly constructed based on a single data modality (such as traffic statistical features or packet payload content), and it is difficult to take into account both the local temporal characteristics and global topological relevance of attack behaviors. For example, methods based on convolutional neural networks can capture local anomalies in traffic sequences, but they cannot model the collaborative attack patterns between devices; while graph neural networks can represent network topological relationships, but they are not sensitive enough to short-term traffic mutations.
[0004] In addition, traditional Transformer models have high computational complexity when dealing with long sequences and rely on fixed threshold sparsification strategies, making it difficult to dynamically adapt to the temporal feature distributions of different attack scenarios, resulting in a trade-off dilemma between real-time detection efficiency and accuracy.
[0005] Therefore, how to improve the accuracy and efficiency of intrusion detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] In view of this, the present invention provides an intrusion detection model based on a dual-stream network architecture, which can capture the topological relationships between devices and the local patterns of traffic data, realize global-local collaborative decision-making, and thus effectively cope with long-term attack patterns and complex environments.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An intrusion detection model based on a dual-stream network architecture, characterized by comprising a preprocessing module, a dual-stream feature extraction module, and a classification module.
[0009] The preprocessing module is used to obtain the original network traffic data for data transformation, and obtain local feature sequences and heterogeneous graph structure data; the local feature sequences are used to represent the traffic statistical information of the original network traffic data in the time dimension; the heterogeneous graph structure data models the network topology through nodes and edges, and is used to capture distributed attack patterns and packet-level malicious behaviors.
[0010] The dual-stream feature extraction module respectively obtains the local feature sequences and the heterogeneous graph structure data for feature extraction, and correspondingly obtains a first feature vector representing local spatio-temporal correlation and a second feature vector representing global topological correlation.
[0011] The classification module is used to analyze the attack type probability according to the first feature vector and the second feature vector, and confirm the intrusion type.
[0012] Preferably, it further includes a feature enhancement module, and the feature enhancement module uses an adaptive sparse attention mechanism to generate attention weights to weight the first feature vector.
[0013] Preferably, the feature enhancement module specifically includes:
[0014] Dynamically generate a sparse ratio:
[0015] α = Sigmoid(W g · GlobalAvgPool(F CNN ) + b g )
[0016] where α is the sparse ratio, F CNN is the local feature sequence, W g is the weight matrix, and b g is the bias term;
[0017] Generate a mask matrix according to the sparse ratio, and obtain a sparse attention output under the mask matrix:
[0018] Attention(Q, K, V) = Softmax(S ⊙ M)V
[0019] where Q, K, and V are the query vector, key vector, and value vector respectively, S is the attention score matrix, and M is the mask matrix.
[0020] Preferably, the data transformation step of the preprocessing module includes:
[0021] Extract basic features from the original network traffic data and perform normalization processing to obtain the local feature sequence.
[0022] Construct nodes and edges based on the original network traffic data, and generate the heterogeneous graph structure data according to the nodes and the edges.
[0023] Preferably, constructing nodes and edges based on the original network traffic data, and generating the heterogeneous graph structure data according to the nodes and the edges specifically includes:
[0024] Define the devices in the network as device nodes and extract the device node features corresponding to each device node; define each traffic packet as a packet node, and encode the packet payloads corresponding to the respective traffic packets as corresponding packet features; define the edges connecting the device nodes and the packet nodes as inclusion edges, and extract the inclusion edge features corresponding to the inclusion edges; define the consecutive packet nodes under the same device node as consecutive edges, and extract the consecutive edge features corresponding to the consecutive edges.
[0025] Generate the heterogeneous graph structure data:
[0026]
[0027] Where G is the heterogeneous graph structure data, is the device node, is the packet node; is the inclusion edge, is the temporal edge; H is the feature matrix.
[0028] Preferably, the dual-stream feature extraction module includes a CNN branch and a GNN branch.
[0029] The CNN branch is used to perform convolution on the local feature sequence to obtain the first feature vector; the GNN branch is used to reason about the heterogeneous graph structure data and perform pooling based on the reasoning result to output the second feature vector.
[0030] Preferably, the CNN branch performing convolution on the local feature sequence specifically includes:
[0031] C (l) = ReLU(Conv1D(C (l-1) , W (l) ) + b (l) )
[0032] Where, is the output feature map of the l-th convolutional layer, n l is the sequence length, c l is the number of channels; is the convolutional kernel weight, is the bias term, and Conv1D represents a one-dimensional convolution operation.
[0033] Preferably, the GNN branch performs inference on the graph structure data, specifically including:
[0034]
[0035] Among them, is the updated device node feature; α ij is the attention coefficient between the device node and the packet node; W is the shared weight matrix; is in the l-th layer, the message passing function from the neighbor node v j to the target node v i ; is the set of node indices pointing to the node v i . is the updated packet node feature; α kj is the attention coefficient, weighting the importance of the neighbor v j to v k ; is in the l-th layer, the message passing function from the neighbor node v j to the packet node v k ; is the set of node indices pointing to the node v k .
[0036] Preferably, the calculation steps of the attention coefficient include:
[0037]
[0038] Among them, W is the shared weight matrix, a is the learnable parameter vector of the attention mechanism, h i represents the feature vector of the device node v i , h j represents the feature vector of the neighbor node v i of v j , and k is the neighbor node variable.
[0039] Preferably, the message passing function specifically includes:
[0040]
[0041] Among them, is the device node feature, a contain is the edge feature, || represents the concatenation operation, MLP dev is a two-layer fully connected network, is the packet node feature; MLP pkt and MLP time are both two-layer fully connected networks, adapting to the high dimension of the packet node feature; when updating the node, according to the target node v i , neighbor node vj and edge type e ij Select the corresponding message passing function; when the device node is updated, i.e., v i ∈V dev , and v j ∈V pkt , and the edge is an inclusion edge, use When the packet node is updated, i.e., v i ∈V pkt , if v j ∈V dev , use If v j ∈V pkt , and the edge is a timing edge, use For each node v i , obtain its neighbors and edge type, and call the corresponding MLP function according to (v i , v j , e ij ).
[0042] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an intrusion detection model based on a dual-stream network architecture, which can capture the topological relationship between devices and the local patterns of traffic data, realize global and local collaborative decision-making, and thus effectively cope with long-time sequence attack patterns and complex environments. The present invention introduces an adaptive sparse attention and multi-modal feature fusion technology to optimize the efficiency and coverage ability of intrusion detection: this mechanism adjusts the attention sparsity through a dynamically learnable gating function, focuses on key timing features to improve the detection ability of burst attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0044] Figure 1 The drawings are schematic diagrams of the structure of an intrusion detection model based on a dual-stream network architecture provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0046] As Figure 1 , an embodiment of the present invention discloses an intrusion detection model based on a dual-stream network architecture, including a preprocessing module, a dual-stream feature extraction module, and a classification module.
[0047] The preprocessing module is used to obtain the original network traffic data for data conversion, and obtain a local feature sequence and heterogeneous graph structure data;
[0048] The local feature sequence is used to represent the traffic statistical information of the original network traffic data in the time dimension; the heterogeneous graph structure data models the network topology through nodes and edges, and is used to capture distributed attack patterns and packet-level malicious behaviors.
[0049] The dual-stream feature extraction module respectively obtains the local feature sequence and the heterogeneous graph structure data for feature extraction, and correspondingly obtains a first feature vector representing local spatio-temporal correlation and a second feature vector representing global topological correlation.
[0050] The classification module is used to analyze the attack type probability according to the first feature vector and the second feature vector, and confirm the intrusion type.
[0051] In one embodiment, the classification module uses a cross-modal attention mechanism to deeply fuse the first feature vector and the second feature vector extracted by the dual-stream features.
[0052] Specifically, first, taking the first feature vector F CNN as the Query, the second feature vector F GNN as the Key and Value, calculate the cross-attention weight.
[0053]
[0054] Then, splice the cross-attention output with the original features into a joint feature vector, and classify through two layers of MLP (Multilayer Perceptron).
[0055] F fusion = [CrossAttn||F CNN ||F GNN
[0056] P = Softmax(W c ·ReLU(W f F fusion + b f ) + b c )
[0057] Among them, is the joint feature vector; P is the class probability, W f is the fully-connected layer weight matrix of the joint feature vector F fusion , and b f is the bias term of this layer. W c is the weight matrix of the classification layer, and b c is the bias term of this layer. m is the number of attack categories. The class weights are dynamically adjusted through the EQL v2 loss function, and the model strengthens the attention to the minority classes during training, and finally outputs a multi-class probability distribution. This can enable the model to jointly utilize local anomalies and global topological patterns, and significantly improve the detection ability for hybrid attacks.
[0058] In one implementation, the data transformation step of the preprocessing module includes: extracting basic features from the original network traffic data and performing normalization processing to obtain the local feature sequence; constructing nodes and edges based on the original network traffic data, and generating the heterogeneous graph structure data according to the nodes and the edges.
[0059] Among them, the core task of the preprocessing module is to convert the original network traffic data into structured inputs that can be processed by the model. To improve the accuracy of intrusion detection, in addition to obtaining the local features of the traffic sequence, the preprocessing module further models to obtain the network global topological relationship between devices.
[0060] To further implement the above technical solution, to obtain the local features of the traffic sequence, the traffic data is standardized to eliminate the feature scale difference. Specifically, for each traffic feature (such as packet length, protocol type, port number, etc.), Z-score normalization is used:
[0061]
[0062] where x' is the normalized local feature, and μ and σ are the mean and standard deviation of the feature respectively. The normalized traffic data is divided into local feature sequences X seq ∈R n×d , where n is the time step and d is the feature dimension.
[0063] In addition, to capture the network topological relationship, the data is further converted into a heterogeneous graph structure.
[0064] Two types of nodes are defined in the graph: device nodes and packet nodes. Among them, the device nodes that is, each device or IP address is defined as a node in the graph, and the node features include its traffic statistics information;
[0065] The packet nodes each traffic packet is defined as a sub-node, and the features are encoded by the byte sequence of the packet payload, and the insufficient part is padded with zeros.
[0066] The design of the edges is also divided into two categories: inclusion edges and temporal edges. The inclusion edges connect the device node to the packet node it sends, and the edge features include information such as packet direction, protocol type, etc.; the temporal edges connect consecutive packet nodes under the same device node, and the edge feature is the time difference t between packets δ . Finally, a heterogeneous graph is generated where H is the node feature matrix.
[0067] is composed of the concatenation of the features of device nodes and packet nodes, where |V| = |V dev | + |V pkt |:
[0068]
[0069] H (0) represents the initial feature matrix of all nodes in the heterogeneous graph. represents the initial feature vector of node v i . If v i is a device node, the traffic statistics feature h dev,i is directly used; if v i is a packet node, the reduced-dimensional packet payload feature h pkt,i is used.
[0070] The GNN branch uses two layers of HGAT to update the feature matrix H. Specifically: for the first-layer HGAT update, starting from H (0) , the message attention coefficient α ij is calculated,
[0071] and all node features are updated to generate H (1) :
[0072]
[0073] is the weight matrix of the node type. H (2) is the same by analogy.
[0074] With this design, the model can simultaneously model the interaction patterns between devices and the temporal dependence of traffic packets, providing a structured input for subsequent feature extraction.
[0075] In order to further implement the above technical solution, the dual-stream feature extraction module includes a CNN branch and a GNN branch. The CNN branch is used to convolve the local feature sequence to obtain the first feature vector; the GNN branch is used to infer the heterogeneous graph structure data, and pool based on the inference results to output the second feature vector.
[0076] Furthermore, the CNN branch adopts a one-dimensional convolutional architecture, but its input is adjusted to the local feature sequence X seq This branch contains five layers of convolution and pooling operations, and the parameters of each convolution kernel are shown in Table 1. Specifically, the convolution operation of the first layer can be expressed as:
[0077] C (l) =ReLU(Conv1D(C (l-1) ,W (l) )+b (l) )
[0078] in, is the output feature map of the lth layer, n l is the sequence length, c l is the number of channels. is the convolution kernel weight, is the bias term, and Conv1D represents a one-dimensional convolution operation.
[0079] For example, the first convolution layer uses 32 kernels of size 7, with a stride of 1, and the output feature size is 61×32 (. Each convolution layer is followed by an average pooling operation to further compress the feature dimension. Finally, the five-layer output is flattened into a 512-dimensional feature vector through global average pooling. It encodes the local statistical regularities of traffic data.
[0080] Table 1: CNN convolutional layer parameter table
[0081]
[0082] The GNN branch is based on the Heterogeneous Graph Attention Network (HGAT), which aims to model the complex interaction relationship between devices and traffic packets and supports feature aggregation of multiple types of nodes and edges.
[0083] Specifically, for the device node v dev and package node v pkt , define message functions separately, treat inclusion edges and timing edges differently, and enhance the multimodal modeling capabilities of heterogeneous graphs:
[0084]
[0085] in, It is the upper layer feature of the device node, including device traffic statistics; a containIt contains edge features, including protocol type and packet direction information; || represents the feature vector concatenation operation; MLP dev It is a two-layer fully connected network for device-to-packet message transformation; It is the feature of the upper layer of the packet node, encoding the packet payload information; MLP pkt It is a two-layer fully connected network to adapt to the high-dimensional input of the packet node; a time It is the temporal edge feature, representing the time difference between packets; MLP time It is a two-layer fully connected network designed specifically for temporal edges.
[0086] Calculate the attention coefficient α between the device node and the packet node according to the edge type and node type ij :
[0087]
[0088] a type It is the learnable parameter vector of the edge type (including edge or temporal edge). If e ij ∈E contain , a type =a contain , if e ij ∈E time , a type =a time , and are the feature vectors of the upper layer of the target node v i and the neighbor node v j respectively.
[0089] Finally, weighted aggregate the neighbor messages and update the node features: For the first-layer HGAT update, starting from H (0) , calculate the message attention coefficient α ij , update all node features, and generate H (1) :
[0090]
[0091] is the weight matrix of the node type. H (2) Similarly.
[0092] After two node updates using two-layer HGAT stacking, output the topological feature vector through global max pooling, which encodes the collaborative attack pattern of the device group and the propagation path of malicious packets.
[0093] Furthermore, select the corresponding transfer function according to different node transfer situations. In addition to the message transfer from the device node to the packet node, it also includes:
[0094]
[0095] Among them, MLP pkt and MLP time are both two-layer fully connected networks, adapting to the high dimension of the packet node features; when updating the node, according to the target node v i , neighbor node v j and edge type e ij select the corresponding message passing function.
[0096] When the device node is updated, that is, v i ∈V dev , and v j ∈V pkt , and the edge is an inclusion edge, use When the packet node is updated, that is, v i ∈V pkt , if v j ∈V dev , use If v j ∈V pkt , and the edge is a time-series edge, use For each node v i , obtain its neighbors and edge type, and call the corresponding MLP function according to (v i , v j , e ij ).
[0097] In another implementation, it further includes a feature enhancement module, and the feature enhancement module is an adaptive sparse Transformer module, which uses an adaptive sparse attention mechanism to generate attention weights to weight the first feature vector.
[0098] Specifically, first, for the feature sequence F CNN output by the CNN branch, generate Q (Query), K (Key), and V (Value):
[0099] Q = F CNN W Q
[0100] K = F CNN W K
[0101] V = F CNN W V
[0102] Calculate the attention score matrix
[0103] Subsequently, dynamically generate a sparse ratio α ∈ [0, 1] through a lightweight MLP controller:
[0104] α = Sigmoid(W g · GlobalAvgPool(F CNN ) + b g )
[0105] where α is the sparse ratio, F CNN is the local feature sequence, W g is the weight matrix, and b g is the bias term. The MLP controller uses the Adam optimizer (learning rate 0.001), dynamically adjusts the class weights with the EQL v2 loss, trains for 50 epochs, and the batch size is 32. During training, the range of α is constrained to [0.1, 0.9] through post-processing to avoid extreme sparsity or full retention.
[0106] Then, a mask matrix is generated according to the sparse ratio, and a sparse attention output is obtained under the mask matrix:
[0107] Attention(Q, K, V) = Softmax(S ⊙ M)V
[0108] where Q, K, and V are the query vector Query, the key vector Key, and the value vector Value respectively, S is the attention score matrix; M is the mask matrix, which determines the number of weights retained or masked in each row of the matrix according to the sparse ratio. For example, when the sparse ratio is 0.3, for the weights at n time steps, the 0.3n weights with the highest weights are retained.
[0109] Finally, according to the attention weights and the message passing function, the update inference of the nodes is completed:
[0110]
[0111] where, is the updated device node feature; α ij is the attention coefficient in different directions between the device node and the packet node; W is the shared weight matrix; in the l-th layer, is the message passing function from the neighbor node v j to the target node v i , is the set of node indices pointing to the node v i ; is the updated packet node feature; α kj is the attention coefficient, weighting the importance of the neighbor v j to v k ; in the l-th layer, is the message passing function from the neighbor node v j to the packet node v kThe message passing function, is a set of node indices pointing to node v k .
[0112] In this embodiment, this mechanism enables the model to autonomously adjust the attention range according to the input traffic characteristics. For example, in a DDoS attack, the sparse ratio α is increased to capture the burst traffic peak, while in steady traffic, α is decreased to reduce the computational overhead.
[0113] The input traffic characteristics refer to the dynamic statistical characteristics of network traffic (such as packet length distribution, protocol type) and temporal patterns (such as burst peaks, periodic communications). The model dynamically generates the sparse ratio α from the temporal global features extracted by the lightweight MLP controller from the CNN branch, and calculates the TopK attention weights to be retained for each row according to α. The binary mask matrix M masks the non-critical time steps and only retains the top k highest weights for each row. This design enhances the sensitivity while reducing redundant calculations, improving the detection ability for burst attacks.
[0114] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0115] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intrusion detection model based on a dual-stream network architecture, characterized in that: It includes a preprocessing module, a dual-stream feature extraction module, and a classification module; The preprocessing module is used to obtain the original network traffic data for data conversion to obtain local feature sequences and heterogeneous graph structure data; the local feature sequence is used to characterize the traffic statistics information of the original network traffic data in the time dimension; the heterogeneous graph structure data is used to model the network topology through nodes and edges to capture distributed attack patterns and packet-level malicious behaviors; the dual-stream feature extraction module respectively obtains the local feature sequence and the heterogeneous graph structure data for feature extraction, and correspondingly obtains a first feature vector characterizing local spatiotemporal correlation and a second feature vector characterizing global topological correlation; The classification module is used to analyze the attack type probability according to the first feature vector and the second feature vector, and confirm the intrusion type.
2. According to claim 1, an intrusion detection model based on a dual-stream network architecture is characterized in that: It also includes a feature enhancement module, which uses an adaptive sparse attention mechanism to generate attention weights to weight the first feature vector.
3. According to claim 2, an intrusion detection model based on a dual-stream network architecture is characterized in that: The feature enhancement module specifically includes: Dynamically generate sparse ratios: α=Sigmoid(W g ·GlobalAvgPool(F CNN )+b g ) Among them, α is the sparse ratio, F CNN is the local feature sequence, W g is the weight matrix, b g is the bias term; Generate a mask matrix according to the sparse ratio, and get the sparse attention output under the mask matrix: Attention(Q,K,V)=Softmax(S⊙M)V Among them, Q, K, V are query vector, key vector and value vector respectively, S is the attention score matrix, and M is the mask matrix.
4. According to claim 1, an intrusion detection model based on a dual-stream network architecture is characterized in that: The data transformation step of the preprocessing module includes: Extracting basic features from the original network traffic data and performing normalization processing to obtain the local feature sequence; Nodes and edges are constructed according to the original network traffic data, and the heterogeneous graph structure data is generated according to the nodes and the edges.
5. According to claim 4, an intrusion detection model based on a dual-stream network architecture is characterized in that: Constructing nodes and edges according to the original network traffic data, and generating the heterogeneous graph structure data according to the nodes and the edges, specifically includes: Define the devices in the network as device nodes and extract the device node features corresponding to each device node; Each traffic packet is defined as a packet node, and the packet payload corresponding to each traffic packet is encoded as the corresponding packet feature; defining an edge connecting the device node and the package node as an inclusion edge, and extracting an inclusion edge feature corresponding to the inclusion edge; Define the continuous packet nodes connected to the same device node as a continuous edge, and extract the continuous edge features corresponding to the continuous edge; Generate the heterogeneous graph structure data: Among them, G is heterogeneous graph structure data, is the device node, is the packet node; contain is the included edge, ε time is the time series edge; H is the feature matrix.
6. The intrusion detection model based on the dual-stream network architecture according to claim 1 is characterized in that: The dual-stream feature extraction module includes a CNN branch and a GNN branch; The CNN branch is used to convolve the local feature sequence to obtain the first feature vector; The GNN branch is used to infer the heterogeneous graph structure data, perform pooling based on the inference result, and output the second feature vector.
7. The intrusion detection model based on the dual-stream network architecture according to claim 6 is characterized in that ,The CNN branch performs convolution on the local feature sequence, specifically including: C (l) =ReLU(Conv1D(C (l-1) ,W (l) )+b (l) ) in, is the output feature map of the lth convolutional layer, n l is the sequence length, c l is the number of channels; is the convolution kernel weight, is the bias term, and Conv1D represents a one-dimensional convolution operation.
8. The intrusion detection model based on a dual-stream network architecture according to claim 6, characterized in that: The GNN branch performs reasoning on the graph structure data, specifically including: in, is the updated device node feature; α ij is the attention coefficient between the device node and the package node; W is the shared weight matrix; is the number of nodes in the first layer from the neighbor node v j Pass to the target node v i The message passing function, Points to node v i The node index collection of is the updated package node feature; α kj is the attention coefficient, weighted neighbor v j v k The importance of is the number of nodes in the first layer from the neighbor node v j Passed to package node v k The message passing function, Points to node v k The node index collection.
9. The intrusion detection model based on the dual-stream network architecture according to claim 8, characterized in that: The calculation steps of the attention coefficient include: Among them, W is the shared weight matrix, a is the learnable parameter vector of the attention mechanism, and h i Represents the device node v i The eigenvector of j Indicates v i Neighbor node v j The characteristic vector of , k is the neighbor node variable.
10. The intrusion detection model based on dual-stream network architecture according to claim 8, characterized in that: The message passing function specifically includes: in, is the device node feature, a contain To include edge features, || indicates the concatenation operation, MLP dev It is a two-layer fully connected network. is the package node feature; MLP pkt and MLP time Both are two-layer fully connected networks, adapting to the high dimensionality of packet node features; when updating a node, according to the target node v i 、Neighbor node v j and edge type e ij Select the corresponding message passing function; when the device node is updated, that is, v i ∈V dev , and v j ∈V pkt , the edge is included, use When the packet node is updated, that is, v i ∈V pkt , if v j ∈V dev ,use If v j ∈V pkt , the edge is a sequential edge, using For each node v i , get its neighbors and edge type, according to (v i ,v j ,e ij ) calls the corresponding MLP function.