A network traffic intrusion detection method based on a knowledge tracking model

By combining a knowledge tracing model-based network traffic intrusion detection method with temporal convolutional attention blocks and an encoder to extract global and local features of network traffic, the shortcomings of existing models in feature extraction and pattern mining are addressed, resulting in more efficient intrusion detection.

CN120811776BActive Publication Date: 2025-12-05ZHONGNAN UNIVERSITY OF ECONOMICS AND LAW
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511284590.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-12-05
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing network traffic intrusion detection models are insufficient in feature extraction and mining of network traffic behavior patterns. They fail to effectively combine global, local and temporal feature information, resulting in insufficient information utilization and difficulty in dealing with complex network attacks.

Method used

We employ a knowledge-based tracking model to generate embedded and joint embedded representations of network traffic through the data input layer. We combine temporal convolutional attention blocks and encoders to extract local and global features, and utilize temporal convolutional networks and multi-head attention layers to capture the temporal features and potential long-term development patterns of traffic data.

Benefits of technology

It improves the performance of network traffic data prediction and intrusion detection, enabling a more comprehensive understanding of network traffic behavior patterns and enhancing the ability to detect complex network attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811776B_ABST
    Figure CN120811776B_ABST
Patent Text Reader

Abstract

The application discloses a network traffic intrusion detection method based on a knowledge tracking model, belongs to the technical field of network traffic intrusion detection, and can effectively capture the time sequence characteristics and potential long-term development mode of traffic data, and improve the network traffic data prediction and intrusion detection performance; comprises the following steps: acquiring a network traffic sequence and a real label sequence; inputting the acquired network traffic sequence and real label sequence into a pre-trained knowledge tracking-based intrusion detection model to detect network traffic intrusion behaviors; the knowledge tracking-based intrusion detection model comprises a data input layer, a knowledge state extraction layer and a prediction layer; the data input layer is used for processing the network traffic sequence and the real label sequence, generating embedding representation and joint embedding representation of the network traffic; the knowledge state extraction layer is used for extracting a knowledge state based on the embedding representation and the joint embedding representation of the network traffic; and the prediction layer is used for outputting a prediction result of the network traffic intrusion behavior based on the knowledge state and the embedding representation of the network traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network traffic intrusion detection technology, specifically a network traffic intrusion detection method based on a knowledge tracing model. Background Technology

[0002] With the rapid development of the internet and digital technologies, cybersecurity has become increasingly important in the information society. One of the core protective measures of traditional cybersecurity is the firewall. However, as attack techniques continue to evolve, the limitations of firewalls have become increasingly apparent. The emergence of Intrusion Detection Systems (IDS) has filled the gaps in traditional network protection technologies. Unlike the passive defense approach of firewalls, intrusion detection technology can not only detect existing attacks but also identify potential threats. This signifies a shift in network defense strategies from a single-minded reliance on perimeter defense to a more comprehensive proactive monitoring and response mechanism, greatly enhancing the overall security of information systems.

[0003] In recent years, machine learning algorithms have been widely applied in intrusion detection research. Almomani et al. proposed an intrusion detection system based on stacked ensemble learning, which effectively improves detection performance using random forests, decision trees, and K-nearest neighbors as its base models. Azhar et al. proposed an advanced random forest intrusion detection method, IDRandom-Forest, which introduces an accuracy sliding window based on hierarchical feature sampling and a feature weighting mechanism to determine the optimal subset in the classic random forest, achieving high accuracy while significantly reducing detection time. Kim et al. proposed a packet-based machine learning model that achieves high detection accuracy and real-time intrusion detection capabilities by minimizing detection latency.

[0004] While machine learning-based intrusion detection methods have achieved some success in detecting network attacks and abnormal behavior, they still suffer from problems such as time-consuming feature extraction and weak generalization ability.

[0005] For the reasons mentioned above, deep learning-based intrusion detection methods have gradually attracted researchers' attention. Unlike traditional machine learning methods, deep learning models perform better when processing large-scale data and have stronger generalization ability for detecting unknown attacks. For example, Hu et al. proposed an intrusion detection model (SAG-BiGRU) based on a self-attention gating mechanism and a bidirectional gated recurrent unit. This model can fully extract high-dimensional information from data samples and improve the detection rate of network attacks. Kim et al. proposed a packet feature-based detection method that uses LSTM-DNN and GAN models to detect intrusion behavior as early as possible before the session ends, while maintaining detection performance comparable to related methods. Ren et al. proposed a hierarchical CNN-attention network (CANET), which combines CNN with an attention mechanism to form a CA module for the extraction of local spatiotemporal features, effectively improving the detection rate of minority classes.

[0006] Existing intrusion detection models have shortcomings in network traffic data feature extraction. They fail to effectively combine global, local, and temporal features for multi-dimensional comprehensive modeling, resulting in insufficient information utilization. Furthermore, although some studies have focused on the temporal characteristics of traffic data, they have failed to further explore the potential long-term development patterns that traffic sequences may exhibit over time, leading to an incomplete understanding of network traffic behavior patterns in the models.

[0007] The above background information is provided only to aid in understanding the concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0008] This application provides a network traffic intrusion detection method based on a knowledge tracing model, which can effectively capture the temporal characteristics of traffic data and its potential long-term development patterns, mine key information from multiple dimensions, and improve the performance of network traffic data prediction and intrusion detection.

[0009] To achieve the above objectives, the embodiments of this application disclose the following technical solutions:

[0010] A network traffic intrusion detection method based on a knowledge tracing model includes the following steps:

[0011] Obtain network traffic sequences and real label sequences;

[0012] The acquired network traffic sequence and real label sequence are input into a pre-trained knowledge-tracking-based intrusion detection model to detect network traffic intrusion behavior;

[0013] The knowledge-based intrusion detection model consists of a data input layer, a knowledge state extraction layer, and a prediction layer. The data input layer processes network traffic sequences and real label sequences to generate embedded representations and joint embedded representations of network traffic. The knowledge state extraction layer extracts knowledge states based on the embedded representations and joint embedded representations of network traffic. The prediction layer outputs prediction results of network traffic intrusion behavior based on the knowledge state and the embedded representations of network traffic.

[0014] In some possible implementations, the specific steps of the data input layer in generating the embedded representation and joint embedded representation of network traffic include:

[0015] Network traffic sequences are processed through embedding operations to generate embedded representations of the network traffic. , The calculation method is as follows:

[0016] ,

[0017] in, Represents a network traffic sequence. Represents the learnable weight matrix; It is the first An embedded representation of network traffic The upper limit of the number of elements is consistent with that of the network traffic sequence Z;

[0018] The actual label sequence is processed through embedding operations to generate a label embedding representation. The calculation method is as follows:

[0019] ,

[0020] in, Indicates the number of categories. Representing feature dimension, Represents the actual label sequence;

[0021] Embedded representation of converged network traffic and tag embedding representation Generate joint embedding representation The calculation method is as follows:

[0022] ;

[0023] in, It is the first A joint embedding representation.

[0024] In some possible implementations, the knowledge state extraction layer includes at least one temporal convolutional attention block; the temporal convolutional attention block includes a temporal convolutional network and an encoder;

[0025] The knowledge state extraction layer is used for network traffic-based embedding representations and joint embedding representations. The steps for extracting knowledge state include:

[0026] Temporal convolutional networks are used to extract local feature representations of network traffic based on joint embedding representations;

[0027] Global feature representations of network traffic are extracted by the encoder based on the embedding representation and local feature representation of network traffic;

[0028] Knowledge state is generated by fusing global and local feature representations by stacking multiple temporal convolutional attention blocks.

[0029] In some possible implementations, the specific steps for temporal convolutional networks to extract local feature representations include:

[0030] Perform dilated causal convolution on the joint embedding representation to obtain the result of the first dilated causal convolution.

[0031] The result of the first dilated causal convolution is first normalized by weights, then nonlinearly transformed by the ReLU activation function, and finally the transformed output is regularized by Dropout.

[0032] The dilated causal convolution is repeated, followed by weight normalization, ReLU activation, and Dropout regularization to obtain the result of the second dilated causal convolution.

[0033] The joint embedding representation of the result of the second dilated causal convolution and the initial input is fused through residual connections;

[0034] The results of the residual connections are processed by the ReLU activation function to generate local feature representations. ;in, :yes This local feature represents the first element in the set. Each element.

[0035] In some possible implementations, the specific steps of the encoder extracting global feature representations include:

[0036] The embedded representation of network traffic is used as the query and key, and the local feature representation is used as the value, which are then input into a multi-head attention layer with a masking mechanism. The attention score of the multi-head attention layer is calculated as follows:

[0037] ,

[0038] in, , and These represent the query, key, and value, respectively.

[0039] The masking mechanism is implemented using an upper triangular matrix, the mask matrix. The calculation method is as follows:

[0040] ,

[0041] in, and Indicates the position index. Indicates the first Is it possible to observe the [number]th [position]? Information about the current position, where 0 indicates that the information at the current position is not obscured. This indicates that the information about the current location is being concealed;

[0042] The output of the multi-head attention layer is fed into the feedforward neural network to generate a global feature representation. The specific calculation formula is as follows:

[0043] ;

[0044] In the formula, Used to stitch together each head in the multi-head attention. The output; , , , and Both represent learnable weight matrices; This represents the channel dimension of the multi-head attention module. The number of heads. ; This represents a feedforward neural network. For the encoder output, Representing global feature representation Global feature representation of the j-th network traffic.

[0045] In some possible implementations, the specific steps for generating a knowledge state include:

[0046] Deleting local feature representations The last time step features And by padding the first and second vectors with zeros, an adjusted local feature representation is generated. ;

[0047] Fusion of global feature representations and adjusted local feature representation Generate knowledge state The calculation method is as follows:

[0048] ;

[0049] in, Representing knowledge state The first element in the middle.

[0050] In some possible implementations, the prediction layer includes a first layer, a second layer, and a third layer, and the specific steps for the prediction layer to output the prediction result include:

[0051] splicing knowledge status Embedded representation of network traffic ;

[0052] The concatenation result is sequentially input into the first linear layer of the first layer, the ReLU activation function, and the Dropout regularization layer;

[0053] The output of the first layer is fed into the second linear layer, ReLU activation function, and Dropout regularization layer of the second layer.

[0054] The output of the second layer is fed into the third linear layer of the third layer and the Softmax activation function to generate probability distributions for each category.

[0055] The category with the highest probability value is taken as the prediction result, calculated as follows:

[0056] ,

[0057] in, , , For learnable parameter matrix, , , This is a bias term.

[0058] In some possible implementations, the training steps for a knowledge-tracking-based intrusion detection model include:

[0059] The binary cross-entropy loss function is used to handle binary classification tasks, and the calculation method is as follows:

[0060] ,

[0061] in, The total number of samples, For real labels, To predict probabilities;

[0062] The focus loss function is used to handle multi-class classification tasks, and the calculation method is as follows:

[0063] ,

[0064] in, The predicted probability for the target category. As a balance factor, It is a one-hot vector. This is the focusing factor.

[0065] Compared with the prior art, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0066] The knowledge-tracking-based intrusion detection model, which combines attention mechanisms and temporal convolutional networks, can effectively capture the temporal features of traffic data and their potential long-term development patterns. It can deeply mine key information in network traffic from multiple perspectives, including global and local ones, and overcome the limitations of existing intrusion detection models in dealing with complex network attacks.

[0067] Specifically, to address the issue of insufficient information extraction, the data input layer performs embedding operations on the network traffic sequence and the real label sequence separately, generating embedded representations of network traffic and labels, and then merges them into a joint embedding representation. This allows the model to integrate traffic features and label information at the initial stage, providing a more comprehensive input for subsequent feature extraction. The knowledge state extraction layer combines a temporal convolutional network and an encoder through temporal convolutional attention blocks. The temporal convolutional network uses dilated causal convolution to extract local features of network traffic, while the encoder extracts global features through a multi-head attention layer with a masking mechanism. The combination of the two achieves multi-angle comprehensive modeling of local and global features, avoiding the shortcomings of single-dimensional feature extraction, and thus making full use of the features of network traffic data. To address the issue of failing to uncover long-term development patterns, the knowledge state extraction layer continuously integrates global and local features of network traffic by stacking multiple temporal convolutional attention blocks. The generated knowledge state is essentially an accumulation of hidden representations of historical traffic patterns. Furthermore, when generating the knowledge state, the last time step feature of the local feature representation is deleted and zero vectors are filled in at the beginning to ensure that the knowledge state relies solely on historical information. This allows the layer to capture the potential development trends of traffic sequences over time, enabling in-depth mining of long-term network traffic behavior patterns and improving the model's ability to understand network traffic behavior patterns. Attached Figure Description

[0068] Figure 1 A flowchart illustrating a network traffic intrusion detection method based on a knowledge tracing model, provided for some embodiments of this application;

[0069] Figure 2 This is a schematic diagram illustrating the working principle of a knowledge-based intrusion detection model.

[0070] Figure 3 This is a diagram of the architecture of the knowledge-based intrusion detection (TCFormer) model.

[0071] Figure 4 This is a diagram of the prediction layer structure of a knowledge-based intrusion detection model.

[0072] Figure 5 Four configuration methods for position encoding;

[0073] Figure 6 These are two input configuration methods for the attention mechanism;

[0074] Figure 7 The impact of different parameter values ​​on the performance of the TCFormer model in binary classification tasks;

[0075] Figure 8 The impact of different parameter values ​​on the performance of the TCFormer model in multi-class classification tasks;

[0076] Figure 9 The confusion matrix diagram for the binary classification of the TCFormer model;

[0077] Figure 10 This is a multi-class confusion matrix diagram for the TCFormer model. Detailed Implementation

[0078] Specific embodiments of the invention will now be described in detail. Although the invention is described in conjunction with these specific embodiments, it should be understood that the invention is not intended to be limited to these specific embodiments. Rather, these embodiments are intended to cover alternative, modified, or equivalent embodiments that may be included within the spirit and scope of the invention as defined by the claims. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. The invention may be practiced without some or all of these specific details. In other instances, well-known processes have not been described in detail so as not to unnecessarily obscure the invention.

[0079] When used in conjunction with the terms "comprising," "method comprising," or similar language in this specification and appended claims, the singular forms "a," "some," and "the" include plural references unless the context clearly indicates otherwise. Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0080] Please see Figure 1-4 This application provides a network traffic intrusion detection method based on a knowledge tracing model, comprising the following steps:

[0081] S101: Obtain network traffic sequence and real label sequence;

[0082] S102: Input the acquired network traffic sequence and real label sequence into a pre-trained knowledge-tracking-based intrusion detection model to detect network traffic intrusion behavior;

[0083] The knowledge-based intrusion detection model (TCFormer model) includes a data input layer, a knowledge state extraction layer, and a prediction layer. The data input layer processes network traffic sequences and real label sequences to generate embedded representations and joint embedded representations of network traffic. The knowledge state extraction layer extracts knowledge states based on the embedded representations and joint embedded representations of network traffic. The prediction layer outputs prediction results of network traffic intrusion behavior based on the knowledge state and the embedded representations of network traffic.

[0084] Preferably, in some embodiments, the specific steps of the data input layer generating the embedded representation and joint embedded representation of network traffic include:

[0085] The first step is to process the network traffic sequence through embedding operations to generate an embedded representation of the network traffic. , The calculation method is as follows:

[0086] (1);

[0087] in, Represents a network traffic sequence. Represents the learnable weight matrix; It is the first An embedded representation of network traffic The upper limit of the number of elements is consistent with that of the network traffic sequence Z;

[0088] The second step involves processing the real label sequence through embedding operations to generate label embedding representations. The calculation method is as follows:

[0089] (2);

[0090] in, Indicates the number of categories. Representing feature dimension, Represents the actual label sequence;

[0091] The third step is to integrate the embedded representation of network traffic. and tag embedding representation Generate joint embedding representation The calculation method is as follows:

[0092] (3);

[0093] in, It is the first A joint embedding representation.

[0094] Preferably, in some embodiments, the knowledge state extraction layer includes at least one temporal convolutional attention block; the temporal convolutional attention block (TCA block) includes a temporal convolutional network (TCN) and an encoder.

[0095] The knowledge state extraction layer is used for network traffic-based embedding representations and joint embedding representations. The steps for extracting knowledge state include:

[0096] The first step is to extract local feature representations of network traffic based on joint embedding representation using a temporal convolutional network;

[0097] The second step is to extract the global feature representation of network traffic by using the encoder based on the embedding representation and local feature representation of network traffic;

[0098] The third step is to generate a knowledge state by stacking multiple temporal convolutional attention blocks to fuse global and local feature representations.

[0099] Preferably, in some embodiments, the specific steps for the temporal convolutional network to extract local feature representations include:

[0100] The first step is to perform a dilated causal convolution operation on the joint embedding representation to obtain the result of the first dilated causal convolution.

[0101] The second step is to first normalize the weights of the result of the first dilated causal convolution, then perform a nonlinear transformation using the ReLU activation function, and finally perform Dropout regularization on the transformed output.

[0102] The third step involves repeating the dilated causal convolution, weight normalization, ReLU activation, and Dropout regularization functions to obtain the result of the second dilated causal convolution.

[0103] The fourth step is to fuse the joint embedding representation of the result of the second dilated causal convolution with the initial input through residual connections;

[0104] The fifth step is to process the residual connection results using the ReLU activation function to generate local feature representations. ;in, :yes This local feature represents the first element in the set. Each element.

[0105] The above process is represented by the following formula 4:

[0106] (4)

[0107] in, Represents a temporal convolutional network. This represents the output of the TCN portion within the TCA block, i.e., a local characteristic representation of network traffic.

[0108] Preferably, in some embodiments, the specific steps for the encoder to extract global feature representations include:

[0109] The first step is to use the embedded representation of network traffic as the query and key, and the local feature representation as the value, and input them into a multi-head attention layer with a masking mechanism;

[0110] The attention score for a multi-head attention layer is calculated as follows:

[0111] (5);

[0112] in, , and These represent the query, key, and value, respectively.

[0113] The second step involves implementing the masking mechanism using an upper triangular matrix. The calculation method is as follows:

[0114] (6);

[0115] in, and Indicates the position index. Indicates the first Is it possible to observe the [number]th [position]? Information about the current position, where 0 indicates that the information at the current position is not obscured. This indicates that the information about the current location is being concealed;

[0116] It's important to note that the masking mechanism is used because the current knowledge state relies solely on information from the historical sequence. Therefore, when calculating attention weights, it's necessary to avoid accessing information from the current or future moments to prevent information leakage. This necessitates the introduction of a masking mechanism.

[0117] Through Add a mask matrix This can give the masked location a very large negative value, so that after the Softmax operation, the attention weight of these locations will approach zero.

[0118] The third step is to input the output of the multi-head attention layer into the feedforward neural network to generate a global feature representation. The specific calculation formula is as follows:

[0119] (7);

[0120] In the formula, Used to stitch together each head in the multi-head attention. The output;

[0121] (8);

[0122] , , and Both represent learnable weight matrices; This represents the channel dimension of the multi-head attention module. The number of heads. ; This represents a feedforward neural network. For the encoder output, Representing global feature representation Global feature representation of the j-th network traffic.

[0123] Preferably, in some embodiments, the specific steps for generating a knowledge state include:

[0124] The first step is to delete local feature representations. The last time step features And by padding the first and second vectors with zeros, an adjusted local feature representation is generated. ;

[0125] The second step is to fuse global feature representations. and adjusted local feature representation Generate knowledge state The calculation method is as follows:

[0126] (9);

[0127] in, Representing knowledge state The Middle Each element.

[0128] Preferably, in some embodiments, the prediction layer includes a first layer, a second layer, and a third layer, and the specific steps for the prediction layer to output the prediction result include:

[0129] The first step is to piece together the knowledge status. Embedded representation of network traffic ;

[0130] The second step is to input the splicing result sequentially into the first linear layer, the ReLU activation function, and the Dropout regularization layer of the first layer;

[0131] The third step is to input the output of the first layer into the second linear layer, the ReLU activation function, and the Dropout regularization layer of the second layer;

[0132] The fourth step is to input the output of the second layer into the third linear layer and the Softmax activation function of the third layer to generate the probability distribution of each category;

[0133] The category with the highest probability value is taken as the prediction result, calculated as follows:

[0134] (10);

[0135] in, , , For learnable parameter matrix, , , This is a bias term.

[0136] Preferably, in some embodiments, the training steps of the knowledge tracing-based intrusion detection model include:

[0137] The first step is to use the binary cross-entropy loss function to handle the binary classification task. The calculation method is as follows:

[0138] (11);

[0139] in, The total number of samples, For real labels, To predict probabilities;

[0140] The second step is to use the focus loss function to handle multi-class classification tasks. The calculation method is as follows:

[0141] (12),

[0142] in, The predicted probability for the target category. As a balance factor, It is a one-hot vector. This is the focusing factor.

[0143] The detailed model optimization and training are as follows:

[0144] (a) Location coding optimization strategy

[0145] In the design process of the TCFormer model, in order to explore and determine the optimal model structure, this study explored four different positional encoding configuration strategies.

[0146] (1) No position coding strategy (No-PE), that is, directly embedding network traffic data. The input is fed into the encoder section of the TCA block.

[0147] (2) Left position coding strategy (L-PE), which is to add position coding before the encoder.

[0148] (3) Right Position Encoding Strategy (R-PE), which is to apply position encoding before the temporal convolutional network.

[0149] (4) Left-Right Position Encoding Strategy (LR-PE), which adds position encoding before both the encoder and the temporal convolutional network.

[0150] Four configuration methods for position encoding, as follows: Figure 5 As shown.

[0151] This application employs the TCFormer model with a No-PE strategy. Specific results will be presented in the subsequent experimental section.

[0152] Attention mechanism optimization strategy

[0153] To further optimize the performance of the attention mechanism, this application explores different configurations of its input data. The attention mechanism has two input data sets. and Queries can be assigned in different ways. ,key ,value Three matrices. This application attempts two configuration methods: (1) ... As and , As (2) will As , As and Two configuration methods are as follows: Figure 6 As shown.

[0154] The first configuration method embeds a global representation of network traffic. As a query s and keys Local feature representation As a value It guides the allocation of attention weights through global information.

[0155] The second configuration method will As a query , Simultaneously serving as a key Sum The aim is to determine the attention weight allocation through local features, while enriching the model's feature representation ability by utilizing the fine-grained information of local features.

[0156] This application adopts configuration option one as the optimal structure for the TCFormer model. Detailed experimental results and analysis will be presented and explained in subsequent experimental sections.

[0157] The model training method is as follows:

[0158] (a) Binary cross-entropy loss

[0159] For the binary classification task, the binary cross-entropy loss is used as the loss function in this experiment, as shown in Equation 11.

[0160] (11)

[0161] in, Represents the total number of samples. Indicates the first The true label of each sample, with a value of either 0 or 1. Indicates the first The probability that a sample is predicted as positive, with a value between 0 and 1.

[0162] (ii) Focus Loss

[0163] In multi-class classification tasks, class imbalance is a frequent problem, with the number of normal traffic samples far exceeding that of abnormal traffic. Abnormal traffic is further divided into various attack categories, such as DoS, DDoS, and bots, and the sample distribution of these attack categories is also extremely uneven. This causes the model to favor predicting the larger number of normal traffic samples, making it difficult to identify the smaller number of abnormal attack samples. Focal Loss adds an adjustable parameter to the cross-entropy loss to increase attention to hard-to-classify samples while suppressing the influence on easily classified samples. Its definition is shown in Equation 12.

[0164] (12)

[0165] in, This represents the model's predicted probability for the target class. This represents a balancing factor, which assigns different weights to each category. This represents the one-hot vector of the current sample. This represents the focus factor, which controls the degree of attention given to difficult samples. The larger the value, the more attention is paid to difficult-to-classify samples.

[0166] The training process of the model is shown in the TCFormer algorithm:

[0167]

[0168] To verify the beneficial effects of the technical solution of this application, the following experiments were conducted:

[0169] I. Feasibility Verification of the Knowledge Tracking Model

[0170] 1. Introduction to Knowledge Tracking Datasets

[0171] This application selects six publicly available datasets widely used in knowledge tracing tasks to verify the effectiveness and rationality of the TCFormer model. Detailed information for each dataset is shown in Table 1.

[0172] Table 1 Knowledge Tracking Dataset

[0173]

[0174] 2. Knowledge Tracking Data Preprocessing

[0175] Data preprocessing includes the following three steps:

[0176] (1) Data filtering: If any element of the triplet interaction is missing, or if the learner has fewer than 3 interaction sequences, then the interaction is filtered.

[0177] (2) Data splitting: Five-fold cross-validation was used. 20% of the interaction sequences were randomly selected as the test set, and the remaining 80% of the data were randomly and evenly divided into 5 parts, 4 parts as the training set and 1 part as the validation set.

[0178] (3) Generating knowledge point subsequences: When a question contains multiple knowledge points, it is split into multiple knowledge points, expanding it from the "question-answer" sequence into a "knowledge point-answer" sequence. The answers corresponding to these knowledge points are the original answers to the question. Subsequently, the expanded "knowledge point-answer" sequence is truncated into a sequence of length 200. Sequences with a length less than 200 are padded with -1 at the end for subsequent model training.

[0179] 3. Knowledge Tracking Comparison Model

[0180] This application selected four commonly used and advanced comparative models to conduct comparative experiments with the TCFormer model. The specific descriptions of the comparative models are as follows:

[0181] (1) DKT (Deep Knowledge Tracing): The DKT model was the first to use a recurrent neural network to process knowledge tracing tasks and is a classic model widely used in the field of knowledge tracing.

[0182] (2) DKVMN (Dynamic Key-value Memory Networks): This model uses dynamic key-value memory networks to store and update students’ knowledge status for each knowledge point, which enhances the interpretability of the knowledge tracing model.

[0183] (3) SAKT (Self-Attentive Knowledge Tracing): The SAKT model was the first to apply the multi-head attention mechanism in the Transformer architecture for modeling in the field of knowledge tracing.

[0184] (4) SAINT (Separated Self-Attentive Neural Knowledge Tracing): SAINT is a knowledge tracing model based on the Transformer architecture, which has a complete "encoder-decoder" structure.

[0185] 4. Experimental Results

[0186] During training, the TCFormer model uses Bayesian search to find the optimal combination of hyperparameters on different datasets. The optimizer chosen is the Adam optimizer, and the loss function is the binary cross-entropy loss function. The hyperparameter search space is set as shown in Table 2.

[0187] Table 2 Hyperparameter Search Space

[0188]

[0189] This application uses AUC and ACC as performance evaluation metrics for the knowledge tracing model. The comparative experimental results are shown in Tables 3 and 4.

[0190] Table 3 AUC Experimental Results

[0191]

[0192] Table 4 ACC Experiment Results

[0193]

[0194] Experimental results show that the SAINT model, based on the Transformer structure, does not perform ideally on different knowledge tracing datasets. In contrast, the TCFormer model proposed in this application achieves excellent performance on all six datasets, demonstrating good versatility and stability. This also verifies the effectiveness and rationale of the model. This result lays the experimental foundation for applying the TCFormer model to the field of intrusion detection.

[0195] II. Experiment and Performance Analysis

[0196] 1. Intrusion Detection Dataset and Preprocessing

[0197] The CIC-IDS-2017 dataset was generated based on real-world network behavior patterns using the B-Profile system proposed by Sharafaldin et al. Data collection began at 9:00 AM on Monday, July 3, 2017, and ended at 5:00 PM on Friday, July 7, 2017, lasting a total of five days. The CIC-IDS-2017 dataset includes not only normal traffic data but also covers various common attack types, such as DoS, DDoS, web attacks, bots, heartbleed attacks, and brute-force attacks. The number of samples corresponding to different network traffic categories is shown in Table 5.

[0198] Table 5. Attack categories and their corresponding quantities in the CIC-IDS-2017 dataset.

[0199]

[0200] Raw network traffic data typically contains a large amount of noise, missing values, and redundant information. Therefore, data preprocessing is required before model training. The specific steps are as follows:

[0201] (1) Data integration: The CSV files from Monday to Friday were integrated into a complete dataset containing network traffic data for five days, which facilitates subsequent experimental processing and analysis.

[0202] (2) Feature Removal: Each sample in the CIC-IDS-2017 dataset contains 84 feature fields and 1 label field. Among them, “Flow ID”, “Source IP”, “Destination IP”, “Protocol”, and “Time stamp” are only used as network identification fields and have limited value for attack behavior analysis, so they need to be removed. In the “Flow Bytes / s” and “Flow Packets / s” fields, NaN and Infinity values ​​are also removed due to their presence. After processing, the feature dimension of each network traffic sample is reduced from 84 to 77.

[0203] (3) Numericalization of discrete variables: This application adopts an ordered encoding method for discrete variables, that is, each category is encoded with integers in sequence.

[0204] (4) Data standardization: Make the data conform to a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0205] (5) Generate time series: Assuming the original sequence contains N samples, divide the samples equally according to the time series length L, and discard the parts with a length less than L.

[0206] (6) Data segmentation: The original network traffic data is divided into training set, validation set and test set in a ratio of 6:2:2.

[0207] In the CIC-IDS-2017 dataset, some attack categories have very few samples; for example, Heartbleed has only 11 samples in the entire dataset. Directly using these categories in multi-class classification tasks would make it difficult for the model to learn effective features from such a small number of samples. Therefore, some attack categories need to be merged, resulting in six main categories, as shown in Table 6.

[0208] Table 6. Composition of the six categories after merging and their corresponding sample sizes

[0209]

[0210] In binary classification tasks, network traffic is directly categorized into two types: normal and abnormal, as shown in Table 7.

[0211] Table 7. The two categories after merging and their respective sample sizes.

[0212]

[0213] 2. Comparison of Models and Evaluation Indicators

[0214] This application selects eight models as comparison models for the TCFormer model, of which three are machine learning models and the remaining five are deep learning models: (1) Logistic Regression; (2) Bayesian Algorithm; (3) SVM, with gamma set to 0.5, C set to 0.1, and radial basis function (RBF) kernel function selected; (4) CNN: composed of four convolutional layers. The number of input feature channels are 128, 256, 512 and 1024 respectively, and the kernel size of each layer is 2; (5) LSTM: composed of three stacked LSTM layers. The hidden unit size in each LSTM layer is 128; (6) GRU: composed of one GRU layer. The hidden unit size in each GRU layer is 128; (7) CNN-LSTM: composed of two convolutional layers and two LSTM layers. The number of channels for input features in the convolutional layers are 128 and 256, respectively. The kernel size of each layer is 2, and the hidden unit size in each LSTM layer is 512. (8) CNN-GRU: It consists of 2 convolutional layers and 2 GRU layers. The number of channels for input features in the convolutional layers are 128 and 256, respectively. The kernel size of each layer is 2, and the hidden unit size in each GRU layer is 512.

[0215] This application uses Accuracy, Precision, Recall, F1, and Macro-F1 as evaluation metrics for this intrusion detection experiment.

[0216] 3. Parameter Settings

[0217] In this experiment, the controlled variable method was used to adjust the parameter combination of the TCFormer model in binary and multi-class classification tasks.

[0218] The initial parameter settings for the TCFormer model in binary and multi-class classification tasks are shown in Table 8.

[0219] Table 8 Initial parameter settings for the TCFormer model

[0220]

[0221] In binary and multi-class classification tasks, the model parameter tuning process is as follows: Figure 7 and Figure 8 As shown.

[0222] 4. Experimental Results and Analysis

[0223] 4.1 Optimization Experiment

[0224] To verify the effectiveness of the optimization scheme, this application conducted a series of experiments to compare and analyze the impact of different optimization schemes on the performance of the TCFormer model.

[0225] (I) Location Coding Optimization Experiment

[0226] The performance of the TCFormer model under the four positional encoding strategies is shown in Table 9.

[0227] Table 9 Performance of the TCFormer model under different positional encoding strategies

[0228]

[0229] Experimental results show that the No-PE strategy outperforms the other three location-based encoding strategies in accuracy, precision, recall, and F1 score. From a data characteristics perspective, network traffic data relies more on relative time intervals, while absolute location has limited effectiveness in identifying behavioral patterns; introducing location encoding would actually increase redundant information. From a model structure perspective, the combination of temporal convolutional networks and attention mechanisms is sufficient to fully exploit the temporal characteristics of traffic sequences, reducing the need for explicit location encoding. Furthermore, introducing location encoding may increase computational complexity or noise, negatively impacting the model's generalization ability.

[0230] (II) Optimization Experiment of Attention Mechanism

[0231] The performance of the TCFormer model under the two input configuration methods is shown in Table 10.

[0232] Table 10 Performance of the TCFormer model under different input configurations in the attention mechanism.

[0233]

[0234] The experimental results show that the attention mechanism of the TCFormer model employs... As and , As Configuration 1 outperforms Configuration 2 across all four evaluation metrics, with a significantly better performance on multi-class classification tasks. This is primarily because multi-class classification requires the model to distinguish between multiple classes, demanding the ability to capture subtle differences between them. Configuration 1 more effectively utilizes fusion embeddings and local features, allowing the attention mechanism to focus more fully on key differences between multi-class features, thereby improving the model's classification ability.

[0235] 4.2 Comparative Experiment

[0236] The performance evaluation results of different models in the CIC-IDS-2017 dataset are shown in Table 11.

[0237] Table 11 Comparison of Experimental Results of Different Intrusion Detection Models

[0238]

[0239] The experimental results show that, in binary classification tasks, the TCFormer model outperforms other comparative models in all evaluation metrics. Its accuracy reaches 0.9992, while its precision and recall reach 0.9985 and 0.9976 respectively, indicating that the TCFormer model effectively reduces false positives and false negatives. Furthermore, the model's F1 score of 0.9981 further validates its excellent performance in identifying attack behaviors. In multi-class classification tasks, the TCFormer model achieves an accuracy of 0.9990. Its precision is 0.9645, which is 3.84% higher than the CNN model and 0.73% higher than the suboptimal CNN-GRU model; its recall is 0.9810, a 4.79% improvement over the CNN model and a 2.07% improvement over CNN-LSTM; and its F1 score is 0.9724, a 4.30% higher than the CNN model and a 1.41% improvement over CNN-LSTM.

[0240] To more intuitively evaluate the performance of the TCFormer model, its corresponding confusion matrix is ​​shown below, such as... Figure 9 and Figure 10 As shown.

[0241] Furthermore, by comparing the F1 scores of different intrusion detection models across various attack categories, the performance advantages of the TCFormer model in complex attack detection tasks can be more clearly demonstrated, as shown in Table 12.

[0242] Table 12 Comparison of F1 scores of the model across different attack categories

[0243]

[0244] Experimental results show that the TCFormer model comprehensively outperforms other comparative models in detecting various types of attack traffic. In the Benign, DoS, and DDoS categories, the TCFormer model achieved the highest F1 score. Its detection advantage is particularly significant in PortScan, Brute Force, and Other traffic categories. Compared to other models, TCFormer achieved an F1 score of 0.9992 in the PortScan category and 0.9864 in the Brute Force category. Especially in the Other category, TCFormer's F1 score is 4.78% higher than the second-best comparative model and 18.15% higher than the CNN model. This indicates that the TCFormer model not only possesses strong identification capabilities for difficult-to-distinguish anomalous traffic but also exhibits high accuracy and strong robustness in detecting easily distinguishable traffic.

[0245] 4.2 Ablation Experiment

[0246] This application presents ablation experiments on the proposed TCFormer model, as shown in Table 13.

[0247] Table 13 F1 score results of ablation experiments

[0248]

[0249] TCFormer-NoKT eliminates the need for a knowledge tracking mechanism. Instead, the embedded representations of network traffic are input into both the encoder and the TCN, with the fused output features serving as the sole input to the prediction layer. Experimental results show that removing the knowledge tracking mechanism reduces the model's F1 score in both binary and multi-class classification tasks, particularly in multi-class classification where the F1 score drops by 6.19%. This validates the crucial role of knowledge tracking in improving model performance.

[0250] TCFormer-NoEC indicates that no encoder structure is used. Without an encoder structure, the model's F1 score on the multi-class classification task decreased by 3.26%. This demonstrates that the encoder structure can effectively capture and integrate global information from network traffic.

[0251] TCFormer-NoTC indicates that temporal convolutional networks are not used. Results show that TCN outperforms binary classification tasks in multi-class classification. The main reason is that multi-class classification involves more attack types, each with different temporal characteristics. TCN can capture dependencies over long time spans, helping to distinguish subtle differences between classes. In binary classification, there are fewer attack types, and the temporal differences between classes are relatively simple; therefore, removing TCN has a smaller impact on model performance.

[0252] To verify the effectiveness of focus loss in improving model performance, this application compares and analyzes the performance of the TCFormer model before and after applying focus loss. The experimental results are shown in Table 14.

[0253] Table 14. Focal loss ablation experiments using the TCFormer model.

[0254]

[0255] The experimental results show that without focus loss, the model's precision, recall, and F1 score for the "Other" class are all 0, indicating that the model failed to learn effectively for this class. However, after applying focus loss, the model's precision, recall, and F1 score for the "Other" class reached 0.8073, 0.8993, and 0.8509, respectively. This demonstrates that focus loss effectively alleviates the class imbalance problem and improves the model's ability to identify minority classes.

[0256] In summary, the network traffic intrusion detection method based on the knowledge tracing model effectively integrates global, local, and temporal information in network traffic data by introducing knowledge tracing technology, thus achieving comprehensive modeling of multi-dimensional features. Especially in the processing of time-series data, the model can capture the long-term development pattern of traffic sequences and adaptively adjust the understanding of network traffic status, thereby improving the ability to identify complex attack patterns and detection accuracy. It effectively solves two key problems existing in intrusion detection models: (1) failure to effectively combine features from multiple levels such as global, local, and temporal aspects for comprehensive modeling; (2) failure to deeply explore the long-term dynamic change patterns that network traffic sequences may exhibit over time.

[0257] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation methods of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A network traffic intrusion detection method based on a knowledge pursuit model, characterized in that, The method comprises the following steps: obtaining a network traffic sequence and a real label sequence; inputting the obtained network traffic sequence and real label sequence into a pre-trained knowledge tracking-based intrusion detection model to detect network traffic intrusion behavior; the knowledge tracking-based intrusion detection model comprises a data input layer, a knowledge state extraction layer and a prediction layer; the data input layer is used for processing the network traffic sequence and the real label sequence to generate an embedding representation of network traffic and a joint embedding representation; the knowledge state extraction layer is used for extracting a knowledge state based on the embedding representation of network traffic and the joint embedding representation; and the prediction layer is used for outputting a prediction result of network traffic intrusion behavior based on the knowledge state and the embedding representation of network traffic; the knowledge state extraction layer comprises at least one time convolution attention block; the time convolution attention block comprises a time convolution network and an encoder; the step of extracting a knowledge state based on the embedding representation of network traffic and the joint embedding representation by the knowledge state extraction layer comprises: extracting a local feature representation of network traffic based on the joint embedding representation by the time convolution network; extracting a global feature representation of network traffic based on the embedding representation of network traffic and the local feature representation by the encoder; generating the knowledge state by fusing the global feature representation and the local feature representation through stacking multiple time convolution attention blocks; the specific steps of generating the knowledge state comprise: deleting the local feature representation the last time step feature and filling the first zero vector to generate an adjusted local feature representation ; fusing the global feature representation and the adjusted local feature representation generating the knowledge state in a manner that ; wherein, representing a knowledge state in the first element. 2.The network traffic intrusion detection method based on knowledge tracing model according to claim 1, wherein, the specific steps of generating the embedding representation of network traffic and the joint embedding representation by the data input layer comprise: processing the sequence of network traffic by embedding operation to generate an embedded representation of the network traffic , The calculation is as follows: , wherein, denotes the sequence of network traffic, denotes a learnable weight matrix; is an embedding representation of the th network traffic, is consistent with the upper limit of the number of elements of the sequence of network traffic Z; The real label sequence is processed by an embedding operation to generate a label embedding representation , the calculation mode is: , wherein, represents the number of classes, represents the feature dimension, represents the true label sequence; fused embedding representation of the network traffic and the label embedding representation generating the joint embedding representation in a manner that ; wherein, is the th joint embedding representation. 3.The network traffic intrusion detection method based on knowledge tracing model according to claim 1, wherein, the specific steps of extracting the local feature representation by the time convolution network comprise: performing a cavity causal convolution operation on the joint embedding representation to obtain a first cavity causal convolution result; performing weight normalization processing on the first cavity causal convolution result, then performing nonlinear transformation on the first cavity causal convolution result through a ReLU activation function, and finally performing Dropout regularization processing on the transformed output; repeating the cavity causal convolution, weight normalization, ReLU activation and Dropout regularization processing functions to obtain a second cavity causal convolution result; fusing the second cavity causal convolution result and the initial input joint embedding representation through residual connection; The results of the residual connection are processed by the ReLU activation function to generate the local feature representation. ;in, yes This local feature represents the first element in the set. Each element.

4. The network traffic intrusion detection method based on knowledge tracing model according to claim 3, wherein, the specific steps of extracting the global feature representation by the encoder comprise: inputting the embedding representation of network traffic as a query and a key, and inputting the local feature representation as a value into a multi-head attention layer with a mask mechanism; the attention score of the multi-head attention layer is calculated in the following manner: , wherein, , and represent query, key and value, respectively; The mask mechanism is implemented by an upper triangular matrix, and the mask matrix The calculation manner is as follows: , wherein, and represents a position index, represents whether the information of the th position can be observed at the th position, 0 represents not masking the information of the current position, represents masking the information of the current position; inputting the output of the multi-headed attention layer to a feedforward neural network to generate the global feature representation The specific calculation formula is as follows: ; wherein the output of each head in the multi-headed attention is concatenated; all represent learnable weight matrices; represents the channel dimension of the multi-headed attention module, is the number of heads, represents a feed-forward neural network, is the output of the encoder, represents the global feature representation of the j-th network traffic.​​​​​​ 5.The network traffic intrusion detection method based on knowledge tracing model according to claim 1, wherein, the prediction layer comprises a first layer, a second layer and a third layer; and the specific steps of outputting the prediction result by the prediction layer comprise: stitching the knowledge state and the embedded representation of the network traffic ; inputting the splicing result into a first linear layer, a ReLU activation function and a Dropout regularization layer of the first layer in sequence; inputting the output of the first layer into a second linear layer, a ReLU activation function and a Dropout regularization layer of the second layer; inputting the output of the second layer into a third linear layer and a Softmax activation function of the third layer to generate a probability distribution of each category; and Taking the category with the highest probability value as the prediction result, the calculation method is: , wherein, , , is a learnable parameter matrix, , , is a bias term.

6. The network traffic intrusion detection method based on knowledge tracing model according to any one of claims 1-5, characterized in that, The training step of the knowledge tracking-based intrusion detection model comprises: A binary cross-entropy loss function is used to process a binary classification task, and the calculation method is: , wherein, is the total number of samples, is the true label, is the predicted probability; A focal loss function is used to process a multi-classification task, and the calculation method is: , wherein, is a predicted probability of the target class, is a balancing factor, is a one-hot vector, is a focus factor.

Citation Information

Patent Citations

  • Abnormal network flow detection method based on bidirectional time convolutional neural network and multi-head self-attention mechanism

    CN115941281A

  • Industrial Internet of Things intrusion detection method based on time sequence clustering

    CN120614155A