A method for designing a multi-task incremental pre-training model architecture for network traffic
By separating the pre-trained model from the common knowledge representation layer and introducing a specific task representation layer, knowledge sharing and feature extraction are optimized, the adaptability problem of the multi-task deep neural network framework in the dynamic changes of the network environment is solved, and efficient network traffic detection is achieved.
Patent Information
- Application Number
- CN202411751420.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing multi-task deep neural network frameworks find it difficult to maintain high-precision discrimination capabilities for new tasks when the network environment changes dynamically. They lack module independence and insufficient optimization of knowledge sharing between tasks, resulting in performance degradation of old tasks and inaccurate feature extraction, making them unable to adapt to complex and changing network environments.
A multi-task incremental pre-training model architecture for network traffic is designed. By separating the pre-training model from the common knowledge representation layer and introducing a task-specific representation layer, principal component analysis and attention mechanism are used to optimize knowledge sharing and task-specific feature extraction, thereby achieving model flexibility and scalability.
It significantly improves the generalization and adaptability of network traffic analysis tasks, can maintain the discrimination ability of old tasks when adding new tasks, reduce training costs, adapt to complex and changing network environments, and improve detection efficiency and accuracy.
Smart Images

Figure CN119544354B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for designing a network traffic multi-task incremental pre-training model architecture, and belongs to the technical field of network security detection. Background Art
[0002] Existing network security detection methods have the following shortcomings when dealing with the ever-increasing amount of network data and increasingly complex network attack patterns: (1) Traditional detection methods such as feature matching-based intrusion detection systems (IDS), port scanning, and deep packet inspection (DPI) rely on predefined feature engineering and rule libraries, which require a lot of manpower for feature design and maintenance, resulting in high costs; (2) Faced with increasingly complex and dynamically changing network attack behaviors, traditional detection methods have poor adaptability and are difficult to capture the nonlinear patterns implicit in network traffic. Especially when facing new attacks and unknown threats, the detection effect is obviously insufficient; (3) Existing network detection systems usually rely on large-scale labeled data for training. However, in practical applications, it is extremely difficult to obtain large-scale, accurate, and widely covered labeled network traffic data, especially in complex and changing network environments. The scarcity of labeled data sets further limits the performance improvement of traditional methods; (4) Existing technologies cannot maintain the detection capabilities of old tasks when the network environment changes dynamically, and are difficult to adapt to new network tasks and application scenarios.
[0003] In recent years, the application of deep learning in various fields has made great progress, especially the successful application of pre-trained models in natural language processing (NLP), which has significantly improved feature extraction and model generalization capabilities. Pre-trained models such as BERT and GPT can automatically learn complex feature representations from data through self-supervised learning on large-scale unlabeled data. This not only reduces the dependence on labeled data, but also enhances the model's ability to handle different tasks. In the field of NLP, pre-trained models can be applied to downstream tasks such as text classification and sentiment analysis through fine-tuning on a small amount of labeled data, greatly improving the performance of the model.
[0004] With the development of pre-training models, multi-task learning (MTL) has become an effective means to further enhance model capabilities. Multi-task learning allows a model to learn multiple related tasks at the same time, which can share feature representations and improve the generalization ability of the model. An important application of multi-task learning is in deep neural networks. By sharing the weights of the underlying network, knowledge can be transferred between different tasks, thereby improving the learning effect of each task. Among them, the BERT-based multi-task deep neural network framework (MT-DNN, Multi-Task Deep Neural Network) is a typical combination of pre-training models and multi-task learning. MT-DNN performs multi-task learning through a shared underlying pre-training model, allowing different tasks to share features and obtain high-precision task results with a small amount of fine-tuning. Therefore, MT-DNN has the ability to handle multiple related tasks and can achieve higher prediction performance when there is less labeled data.
[0005] MT-DNN mainly includes the following key parts: First, MT-DNN is pre-trained on large-scale unlabeled text data, and uses self-supervised learning methods such as masked language model (MLM) and next sentence prediction tasks to capture complex language feature relationships from unlabeled data. The pre-training stage provides the model with extensive context understanding capabilities; in the MT-DNN framework, all tasks share an underlying pre-trained model (such as BERT). The shared model can generate a common feature representation for each task. The shared representation layer extracts context-related deep features from the input data through the multi-layer Transformer network of the BERT model; based on the shared representation layer, MT-DNN designs a task-specific output layer for each task. The task-specific layer is usually a fully connected neural network responsible for converting the shared representation into the specific output required for the task. The above design enables different tasks to share low-level features. hierarchical features, while high-level features are fine-tuned through their respective task-specific layers; in the process of multi-task learning, MT-DNN improves the overall performance by fine-tuning its performance in downstream tasks. Through a small amount of labeled data, the shared model can quickly adapt to the needs of specific tasks while maintaining efficient processing capabilities for other tasks; but MT-DNN still has shortcomings. The design of MT-DNN is mainly aimed at natural language processing tasks. Although the pre-trained model can share features between multiple tasks, it is not optimized for the special characteristics of network traffic data. Therefore, directly applying MT-DNN to the field of network traffic may not be able to effectively capture the nonlinear and implicit pattern relationships in network data. When MT-DNN adapts to new tasks, it is easy for the performance of old tasks to degrade, especially in the field of network security detection. As the network environment continues to change, the addition of new tasks may lead to a decrease in the ability to discriminate against previous tasks.
[0006] In summary, the existing technologies have the following shortcomings: (1) The module independence of the model is insufficient. The existing multi-task learning pre-training model usually integrates the pre-training model and the knowledge sharing representation layer of all tasks into an overall framework, resulting in poor structural flexibility of the model. When training downstream tasks, since the pre-training model and the task-specific layer are not clearly separated, the basic knowledge learned in the pre-training stage may be overwritten by the training of subsequent tasks, resulting in poor performance of the model when processing new tasks. The design of insufficient module independence limits the adaptability and scalability of the model to tasks, making it difficult to respond quickly and adjust flexibly in dynamic task scenarios; (2) The knowledge sharing between tasks is not optimized enough. In the existing multi-task learning framework, the shared representation layer is usually used to process multiple tasks at the same time, and fails to fully distinguish the specific features between different tasks. Although the design of excessive sharing of features between tasks can save resources, it may cause important feature information of old tasks to be overwritten or lost when new tasks are added, thereby reducing the model's ability to discriminate against old tasks, resulting in the so-called The problem of forgetting old tasks, especially in application scenarios such as network traffic detection, different types of network traffic and security events often have significant feature differences. Unoptimized task sharing will lead to a decrease in the generalization ability of the model when processing key tasks, and it will not be able to effectively distinguish the characteristics of new and old tasks; (3) Task-specific feature extraction is not accurate enough. The existing multi-task pre-training model lacks the ability to specifically model the specific knowledge of each task when processing multiple tasks. Especially when processing complex network traffic detection tasks, different types of network traffic show significant feature differences. Relying solely on shared pre-training models cannot accurately capture these feature differences. Since the existing technology fails to effectively distinguish and learn the unique features of different tasks, the model shows a decrease in accuracy and detection ability when processing new types of network traffic, making it difficult to meet the precise detection needs in practical applications. Therefore, there is a need for a network traffic multi-task incremental pre-training model architecture design method that improves the module independence of the model, optimizes knowledge sharing between tasks, and improves the accuracy of task-specific feature extraction. Summary of the Invention
[0007] A brief overview of the present invention is provided below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important aspects of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is simply to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.
[0008] In view of this, in order to solve the problem in the prior art that the traditional multi-task deep neural network framework cannot maintain high-precision discrimination capabilities for new tasks as the network environment changes dynamically, the present invention provides a method for designing a network traffic multi-task incremental pre-training model architecture.
[0009] The technical solution is as follows: A method for designing a network traffic multi-task incremental pre-training model architecture includes the following steps:
[0010] S1. Construct a multi-task incremental pre-training model architecture that separates the pre-training model from the common knowledge representation layer and introduces a task-specific representation layer;
[0011] Specifically, the multi-task incremental pre-training model architecture includes a pre-training model, a common knowledge representation layer, and a task-specific knowledge layer. The pre-training model and the common knowledge representation layer are separated, and the pre-training model is used as an independent module, and the common knowledge representation layer is connected to the task-specific knowledge layer.
[0012] S2. Train the pre-trained model in the multi-task incremental pre-trained model architecture to obtain a trained multi-task incremental pre-trained model architecture;
[0013] S3. Fine-tune the trained multi-task incremental pre-training model architecture, and obtain the final multi-task incremental pre-training model architecture based on the set common knowledge representation layer architecture and specific task representation layer architecture.
[0014] Furthermore, the step S2 includes the following steps:
[0015] S21. Calculate and rank the discriminative power of the features extracted by CICFlowMeter using principal component analysis;
[0016] S22. Block processing is performed on the features of different distinguishing strengths extracted by CICFlowMeter to achieve feature form conversion;
[0017] S23. According to the features after the form conversion, two pre-training tasks, namely, the masking task and the service prediction task, are set for the pre-training model for training;
[0018] In the step S21, calculating and sorting the distinguishing strength of the features extracted by CICFlowMeter includes data preprocessing, calculating the covariance matrix, calculating the eigenvalues and eigenvectors to select the principal components, calculating the distinguishing strength of the features, and normalizing the distinguishing strength;
[0019] In the data preprocessing process, the input data of the pre-trained model, that is, the network dataset, is cleaned, and the network dataset is checked for missing values. The missing values are filled with 0, and the input data is standardized by the standardization formula to obtain the standardized data matrix X std ;
[0020] The standardized data matrix X std Expressed as:
[0021]
[0022] Where X is the original data matrix, μ is the mean vector of the feature, and σ is the standard deviation vector of the feature;
[0023] In the process of calculating the covariance matrix, according to the standardized data matrix X std , calculate the covariance matrix C;
[0024] The covariance matrix C is expressed as:
[0025]
[0026] Where T represents the matrix transpose and n represents the number of features;
[0027] Perform eigenvalue decomposition on the covariance matrix C and calculate the eigenvalues and eigenvectors of the covariance matrix C;
[0028] The decomposition process is expressed as:
[0029] Cv i =λ i v i
[0030] Among them, λ i is the i-th eigenvalue of the covariance matrix, v i is the corresponding eigenvector, and the integration results in a set of eigenvalues {λ1,λ2,...,λ n} and the corresponding feature vector set {v1,v2,...,v n};
[0031] Sort the principal components according to the size of the eigenvalue, select the principal component whose eigenvalue is greater than the set contribution value, and obtain the eigenvectors {v1, v2, ..., v k} and combined into a dimensionality reduction matrix W k , where k is the number of principal components, the standardized data matrix X std Project it onto the selected principal component to obtain the dimension-reduced data Z;
[0032] Z=X std
[0033] In the process of calculating the distinguishing power of features, the importance of each feature is measured by its contribution to each principal component. The importance of each feature is obtained by calculating the contribution rate, that is, the distinguishing power of CICFlowMeter features is obtained and ranked;
[0034] Contribution rate is expressed as:
[0035]
[0036] Among them, Variance explained by the feature represents the variance explained by the feature, and Totalvariance explained by all features represents all variance explained by all features;
[0037] In the S22, a vocabulary mapping with a range of [0, 65535] is used to divide the features extracted by CICFlowMeter into blocks of 16 bytes each, and feature values exceeding 16 bytes, i.e., greater than 65535, are processed in blocks and separated by a delimiter token[sep] to indicate that the separated blocks belong to the same feature value. After block processing, each feature is converted into two forms: token[val] and [token[sep], token[val_1], ..., token[val_n], token[sep]].
[0038] Furthermore, the step S3 includes the following steps:
[0039] S31. Set up a common knowledge representation layer architecture;
[0040] S32. Set up a specific task knowledge layer architecture;
[0041] In S31, each network flow of the extracted data collection is processed by the pre-trained model to generate a feature vector of a specific network flow with a size of 1×768, which is recorded as the first feature vector Input bert , the common knowledge representation layer represents a matrix of size 192×768, and the first eigenvector Input bert Add to the end of the common knowledge representation layer to realize the first feature vector Input bert Fusion with the common knowledge representation layer to obtain the first fusion matrix temp net , whose size is 193×768, where the first 192 rows are the rows of the matrix represented by the original common knowledge representation layer, and the last row is the first eigenvector Input bert ;
[0042] Use the attention mechanism to the first fusion matrix temp net Perform attention calculation to generate the first query vector Q, the first key vector K, and the first value vector V;
[0043] The first query vector Q is expressed as:
[0044] Q=W q Input bert
[0045] Among them, W q is the weight matrix of the first query vector Q;
[0046] The first key vector K is expressed as:
[0047] K=W k tmp net
[0048] Among them, W k is the weight matrix of the first key vector K;
[0049] The first value vector V is represented as:
[0050] V=W v @tmp net
[0051] Among them, W v is the weight matrix of the first value vector V;
[0052] According to the first query vector Q, the first key vector K and the first value vector V, the first attention matrix A is calculated and multiplied by the first value vector V to obtain the output vector Output of the shared knowledge representation layer, whose size is 1×768;
[0053] The first attention matrix A is expressed as:
[0054]
[0055] Where √D is the scaling factor, D represents the dimension of the hidden layer vector, and the softmax function represents the normalization process;
[0056] The output vector Output of the shared knowledge representation layer is expressed as:
[0057] Output=W o AV+b o
[0058] Among them, W o is the output linear transformation weight matrix, b o is the bias term;
[0059] In S32, each network flow of the extracted data collection is processed by the pre-training model and combined with the common knowledge representation layer to generate a feature vector of size 1×768, which is recorded as the second feature vector Output G , the task-specific knowledge layer represents a matrix of size n×768, where n represents the number of current tasks, and the second eigenvector Output G Splice with the task-specific knowledge layer on the first dimension to obtain the matrix tmp of spliced specific tasks and shared knowledge process , whose dimension is (n+1)×768. For each task, extract the one-dimensional task-specific knowledge vector corresponding to the task, which is recorded as the specific knowledge vector tmp input , using specific knowledge vector tmp input Perform attention calculation with the shared knowledge to generate the second query vector Q', the second key vector K' and the second value vector V';
[0060] The second query vector Q' is expressed as:
[0061] Q'=W q 'tmp input
[0062] Among them, W q ' is the weight matrix of the second query vector Q';
[0063] The second key vector K' is expressed as:
[0064] K'=W k 'tmp process
[0065] Among them, W k ' is the weight matrix of the second key vector K';
[0066] The second value vector V' is represented as:
[0067] V'=W v 'tmp process
[0068] Among them, W v ' is the weight matrix of the second value vector V'.
[0069] The beneficial effects of the present invention are as follows: a network traffic multi-task incremental pre-training model architecture design method proposed in the present invention not only realizes the effective separation of general knowledge and specific task knowledge, but also ensures the performance of the pre-training model in multi-task and incremental learning scenarios through synergy, significantly improves the generalization ability and adaptability in network traffic analysis tasks, and provides an effective solution for processing complex multi-task scenarios. Through the dynamic update mechanism, the combination of the common knowledge representation layer and the specific task knowledge layer realizes the perfect balance between knowledge sharing and task specificity; through modular separation of the pre-training model and the common knowledge representation layer, the flexibility and scalability of the model are significantly improved. Compared with the existing technology, the present invention can flexibly adjust the task-specific layer when adding new tasks, avoid covering the basic knowledge learned in the pre-training stage, greatly improve the adaptability of the pre-training model to downstream tasks, and enable the pre-training model to be efficiently expanded and adjusted to meet the needs of complex and changeable network environments, while reducing training costs; the present invention optimizes knowledge sharing between tasks through the double-layer structure design of the common knowledge layer and the specific task knowledge layer, ensures that the model's discrimination ability of old tasks will not decrease when adding new tasks, solves the problem of forgetting old tasks, and concludes The present invention combines an incremental learning mechanism with a small amount of labeled data to quickly adapt to new tasks through fine-tuning while maintaining high-performance discrimination capabilities for old tasks, thereby significantly improving the adaptability, stability and accuracy of the model in dynamic network traffic detection. This makes the present invention not only have higher application value in network security detection, but also greatly improves the efficiency and accuracy of processing complex network environments. By introducing the technical framework of the pre-trained model, the present invention uses massive unlabeled data to learn the complex nonlinear characteristics in network traffic, thereby effectively alleviating the problem of difficulty in obtaining labeled data. Through the characteristics of the pre-trained model, the present invention can complete the fine-tuning of network traffic feature modeling with less labeled data and training iterations, and adapt to specific network detection and management tasks. By combining multi-task learning with the incremental learning mechanism, the present invention achieves that the model can still maintain the discrimination performance of old tasks when adapting to new tasks. Therefore, the present invention not only solves the defects of traditional network security detection methods in terms of high labor cost, poor adaptability and dependence on large-scale labeled data, but also significantly improves the generalization ability, stability and detection effect of network traffic detection through the introduction of pre-trained models and incremental learning, making it suitable for complex and changing network environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0071] Figure 1A flowchart of a method for designing a multi-task incremental pre-training model architecture for network traffic;
[0072] Figure 2 Schematic diagram of the multi-task incremental pre-training model architecture;
[0073] Figure 3 This is a pseudo code diagram for block processing;
[0074] Figure 4 A schematic diagram of the shared knowledge representation layer architecture and dimensional changes;
[0075] Figure 5 Schematic diagram of the knowledge layer architecture and dimensional changes for a specific task. DETAILED DESCRIPTION
[0076] To make the technical solutions and advantages of the embodiments of the present invention more clearly understood, exemplary embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be noted that the embodiments described are only a portion of the embodiments of the present invention, and are not an exhaustive list of all embodiments. It should be noted that the embodiments of the present invention and the features thereof may be combined with each other unless they conflict.
[0077] refer to Figure 1-Figure 5 This embodiment is a method for designing a multi-task incremental pre-training model architecture for network traffic, specifically comprising the following steps:
[0078] S1. Construct a multi-task incremental pre-training model architecture that separates the pre-training model from the common knowledge representation layer and introduces a task-specific representation layer;
[0079] Specifically, the multi-task incremental pre-training model architecture includes a pre-training model, a common knowledge representation layer, and a task-specific knowledge layer. The pre-training model and the common knowledge representation layer are separated, and the pre-training model is used as an independent module, and the common knowledge representation layer is connected to the task-specific knowledge layer.
[0080] S2. Train the pre-trained model in the multi-task incremental pre-trained model architecture to obtain a trained multi-task incremental pre-trained model architecture;
[0081] S3. Fine-tune the trained multi-task incremental pre-training model architecture to obtain the final multi-task incremental pre-training model architecture based on the set common knowledge representation layer architecture and specific task representation layer architecture;
[0082] Specifically, the improved multi-task incremental pre-training model architecture has multiple advantages. On the one hand, the independent shared knowledge layer can improve the model's performance on multiple tasks. On the other hand, the scalability of the pre-trained model is enhanced. Since the pre-trained model and the shared knowledge layer are structurally separated, the pre-trained model can be flexibly expanded and adjusted.
[0083] A task-specific knowledge layer is also designed in the multi-task incremental pre-training model architecture to ensure that the model can effectively learn the unique features related to the task when faced with a new task. At the same time, the task-specific knowledge layer retains important parameters closely related to the performance of specific tasks, thereby avoiding the problem of the model losing key task information during multi-task training. Through this structured design, the pre-training model can find a balance between sharing task knowledge and maintaining task-specific performance, greatly improving the adaptability and performance of the task.
[0084] refer to Figure 2 The multi-task incremental pre-training model architecture constructed by the present invention not only fully utilizes the powerful pre-training capability of the pre-training model, but also improves the flexibility, scalability and adaptability of the model in multi-task scenarios through the separation design of the common knowledge layer and the specific task knowledge layer, showing broad application prospects in the field of network traffic analysis. Among them, BERT-Base is a pre-training model, and GS-knowledge is a network traffic multi-task pre-training model architecture proposed by the present invention, which consists of a pre-training model, a common knowledge representation layer, and a specific task representation layer.
[0085] The core concept of the multi-task incremental pre-training model architecture is to achieve the separation and coordination of general knowledge and task-specific knowledge. The main function of the shared knowledge representation layer is to extract basic features from network traffic data. Through continuous updates and iterations, the model's ability to express general network knowledge is gradually enhanced. The independence of the shared knowledge representation layer enables different tasks to share basic knowledge, thereby avoiding repeated learning of basic information and improving learning efficiency.
[0086] At the same time, the task-specific knowledge layer focuses on retaining knowledge closely related to the specified task, ensuring that the model does not forget the key features of the old task when dealing with new tasks. Through this knowledge separation strategy, the pre-trained model can handle different task requirements more flexibly, while ensuring that its performance in new tasks does not have a negative impact on the old tasks. The dynamic task-specific knowledge layer enables the model to retain its understanding and performance of old tasks when facing new tasks, effectively solving the task conflict problem in multi-task learning.
[0087] Furthermore, the step S2 includes the following steps:
[0088] S21. Calculate and rank the discriminative power of the features extracted by CICFlowMeter using principal component analysis (PCA);
[0089] S22. Block processing is performed on the features of different distinguishing strengths extracted by CICFlowMeter to achieve feature form conversion;
[0090] S23. According to the features after form conversion, two pre-training tasks, namely, a masking task and a service prediction task, are set for the pre-training model;
[0091] Specifically, in step S23, the purpose of designing the pre-training task is to enable the pre-training model to understand the relevant knowledge in a certain field, so as to utilize the learned basic knowledge in various downstream tasks and perform well. The present invention sets two pre-training tasks: the cover mask task and the service prediction task;
[0092] The masking task is similar to the pre-training task of the BERT model in natural language processing (NLP), but in the present invention, the input data is related to network traffic features. Its purpose is to enable the pre-trained model to understand and predict the complex relationship between network traffic features, extract the core information of the traffic features, and map it to a high-dimensional space so that it can play a role in different network application fields later. The application of the masking task to network traffic features enables the pre-trained model to learn the intrinsic correlation and dependency between features by masking and predicting some feature values, thereby improving the model's generalization ability for unseen network data patterns. Through the masking task, the pre-trained model can predict the masked features from a given feature combination and capture the statistical relationship and dependency structure between features. In addition, the masking task can also help the model enhance the robustness of feature representation by processing incomplete network traffic features.
[0093] The service prediction task requires the pre-trained model to predict whether two network flows belong to the same service type, such as HTTP, HTTPS, FTP, SMTP, etc., based on the extracted features. Through the service prediction task, the pre-trained model can not only capture the relationship between different traffic features, but also enhance its understanding of network context. At the same time, the service prediction task can help the pre-trained model improve its generalization ability when facing different service types.
[0094] By setting up two pre-training tasks, the pre-trained model can learn a deeper knowledge structure when understanding the complex relationships between network traffic characteristics and the context of service types, ensuring its better adaptability and expressiveness in subsequent specialized applications. This pre-training strategy will significantly improve the model's performance in various network traffic analysis tasks.
[0095] In the step S21, calculating and sorting the distinguishing strength of the features extracted by CICFlowMeter includes data preprocessing, calculating the covariance matrix, calculating the eigenvalues and eigenvectors to select the principal components, calculating the distinguishing strength of the features, and normalizing the distinguishing strength;
[0096] In the data preprocessing process, the input data of the pre-trained model, that is, the network dataset, is cleaned, and the network dataset is checked for missing values. The missing values are filled with 0, and the input data is standardized by the standardization formula to obtain the standardized data matrix X std ;
[0097] The standardized data matrix X std Expressed as:
[0098]
[0099] Where X is the original data matrix, each row represents a sample, each column represents a feature, μ is the mean vector of the feature, and σ is the standard deviation vector of the feature;
[0100] In the process of calculating the covariance matrix, according to the standardized data matrix X std , calculate the covariance matrix C;
[0101] The covariance matrix C is expressed as:
[0102]
[0103] Where T represents the matrix transpose and n represents the number of features;
[0104] Perform eigenvalue decomposition on the covariance matrix C and calculate the eigenvalues and eigenvectors of the covariance matrix C;
[0105] The decomposition process is expressed as:
[0106] Cv i =λ i v i
[0107] Among them, λ i is the i-th eigenvalue of the covariance matrix, v i is the corresponding eigenvector, eigenvalue λ i Represents the data in the feature vector v i The variance of the corresponding direction is integrated to obtain a set of eigenvalues {λ1,λ2,...,λ n} and the corresponding feature vector set {v1,v2,...,v n};
[0108] Sort the eigenvalues from large to small. Larger eigenvalues correspond to more important eigenvectors. Select the eigenvalues with a contribution rate of 85% as the important principal component. Get the eigenvectors corresponding to the principal components {v1, v2, ..., v k} and combined into a dimensionality reduction matrix W k , where k is the number of principal components, the standardized data matrix X std Project it onto the selected principal component to obtain the dimension-reduced data Z;
[0109] Z=X std
[0110] In the process of calculating the distinguishing power of features, the importance of each feature is measured by its contribution to each principal component. The importance of each feature is obtained by calculating the contribution rate, that is, the distinguishing power of CICFlowMeter features is obtained and ranked;
[0111] Contribution rate is expressed as:
[0112]
[0113] Among them, Variance explained by the feature represents the variance explained by the feature, and Totalvariance explained by all features represents all variance explained by all features;
[0114] Specifically, the covariance matrix reflects the linear correlation between the features in the data set. The key step of principal component analysis is to calculate the eigenvalues and eigenvectors of the covariance matrix. The eigenvalues represent the variance of the eigenvectors in the data, while the eigenvectors define the direction of the principal components. Through step S21, a feature set with high discrimination power can be obtained to facilitate the establishment of a pre-training model.
[0115] In the S22, a vocabulary mapping in the range of [0, 65535] is used to divide the features extracted by CICFlowMeter into blocks of 16 bytes each, and feature values exceeding 16 bytes, i.e., greater than 65535, are subjected to block processing, separated by a delimiter token[sep] to indicate that the separated blocks belong to the same feature value. After block processing, each feature is converted into two forms: token[val] and [token[sep], token[val_1], ..., token[val_n], token[sep]].
[0116] Specifically, to improve the performance of the pre-trained model in subsequent tasks, the features extracted by CICFlowMeter first need to be normalized. To enable the pre-trained model to better understand network traffic characteristics and their correlations, a vocabulary mapping in the range of [0, 65535] is used. Since BERT-Base is a neural network model derived from natural language processing (NLP), it is necessary to set rules for the feature values extracted by CICFlowMeter, similar to constructing a language, to fully utilize the performance of the pre-trained model.
[0117] refer to Figure 3 By splitting large values into smaller units and encoding them with token[sep] as a separator, it can not only effectively reduce the dimension of the values and avoid the numerical stability problems caused by directly processing large values, but also prevent the sharp increase of neural network parameters, standardize the numerical range of eigenvalues, and reduce the computational complexity of the model. At the same time, the block processing method can also improve the learning efficiency of the pre-trained model. Since the pre-trained model is good at processing discrete text data, converting values into smaller, discrete tags can make full use of the advantages of the pre-trained model in processing discrete output space. In addition, using the separator token[sep] can help the model better understand the boundaries and relative position relationships between values.
[0118] Furthermore, the step S3 includes the following steps:
[0119] S31. Set up a common knowledge representation layer architecture;
[0120] S32. Setting up a specific task presentation layer architecture;
[0121] In S31, each network flow of the extracted data collection is processed by the pre-trained model to generate a feature vector of a specific network flow with a size of 1×768, which is recorded as the first feature vector Input bert , the common knowledge representation layer represents a matrix of size 192×768, and the first eigenvector Input bert Add to the end of the common knowledge representation layer to realize the first feature vector Input bert Fusion with the shared knowledge representation layer to form the first fusion matrix temp net , whose size is 193×768, where the first 192 rows are the rows of the matrix represented by the original common knowledge representation layer, and the last row is the first eigenvector Input bert ;
[0122] Use the attention mechanism to the first fusion matrix temp net Perform attention calculation, that is, generate the first query vector Q, the first key vector K and the first value vector V;
[0123] The first query vector Q is expressed as:
[0124] Q=W q Input bert
[0125] Among them, W q is the weight matrix of the first query vector Q, used to transform the first fusion matrix Input bert Mapping to query space;
[0126] The first key vector K is expressed as:
[0127] K=W k tmp net
[0128] Among them, W k is the weight matrix of the first key vector K, which is used to make the first key vector contain both the representation of shared knowledge and the representation of input features;
[0129] The first value vector V is represented as:
[0130] V=W v tmp net
[0131] Among them, W v is the weight matrix of the first value vector V, used to transform the first fusion matrix Input bert Mapping to value vector space;
[0132] According to the first query vector Q, the first key vector K and the first value vector V, the first attention matrix A is calculated and multiplied by the first value vector V to obtain the output vector Output of the shared knowledge representation layer, whose size is 1×768;
[0133] The first attention matrix A is expressed as:
[0134]
[0135] Where √D is the scaling factor, D represents the dimension of the hidden layer vector, which is used to prevent the gradient from disappearing due to excessive values, and the softmax function represents the normalization process;
[0136] The output vector Output of the shared knowledge representation layer is expressed as:
[0137] Output=W o AV+b o
[0138] Among them, W o is the output linear transformation weight matrix, bo is the bias term.
[0139] Specifically, in the multi-task incremental fine-tuning learning process, the key innovation of the present invention lies in the effective fusion and separation of the common knowledge representation layer and the task-specific knowledge layer. Specifically, the present invention designs a neural network matrix of size 192×768 as the common knowledge representation layer. The role of the common knowledge representation layer in multi-task learning is to extract and share common features for different tasks, ensuring that each task can benefit from the existing knowledge representation while sharing the same knowledge foundation. This design greatly improves the effect of knowledge transfer in multi-task learning.
[0140] refer to Figure 4 In the entire process from the input feature vector to the final output vector of the shared knowledge representation layer, the pre-trained model shares the existing knowledge representation and the combination of input features, enabling it to not only cope with different tasks in multi-task learning, but also improve the performance of tasks in specific fields. The structure of the shared knowledge representation layer ensures that the pre-trained model has stronger semantic understanding and information generalization capabilities during the knowledge transfer process, and can use these enhanced knowledge representations to make more accurate decisions when processing tasks in specific fields;
[0141] The first eigenvector Input bert Contains the essential characteristics of the network flow, which is a deep representation of the input network flow. The first eigenvector Input bert As the input of a specific task, it contains the unique information of the current task. The common knowledge representation layer contains the common knowledge learned from multiple tasks. However, in order to make full use of the common knowledge features, the present invention integrates it with the common knowledge representation layer. In order to achieve this integration, the present invention adopts a clever splicing method. Specifically, the first eigenvector Input bert Rather than adding the matrix represented by the shared knowledge representation layer element by element, a concatenation operation is performed along the first dimension. This concatenation allows the pre-trained model to utilize both the shared knowledge representation layer and the task-specific input features in subsequent calculations, improving the model's understanding of the current task. This not only helps the model excel on new tasks but also ensures that the shared knowledge is utilized to the maximum extent possible. In particular, in the subsequent attention mechanism, this concatenation enables the model to effectively weight general knowledge and task-specific features, further enhancing the performance of the pre-trained model.
[0142] The attention mechanism is generally used to improve the model's weighting of the importance of different information sources. In the present invention, the attention mechanism can dynamically adjust the weights between common knowledge and specific task features according to the needs of the current task. netDifferent weights are assigned to features in different rows in the pre-trained model. The pre-trained model can flexibly decide whether to pay more attention to certain common features in the shared knowledge representation layer or the input features of specific tasks in the current task. The attention mechanism ensures that the model has sufficient adaptability and generalization capabilities in different tasks. By introducing the attention mechanism, the pre-trained model of the present invention can not only realize knowledge sharing between different tasks, but also dynamically adapt to the needs of each task, making the model performance more robust. The effective fusion of the shared knowledge representation layer and the input features of specific tasks greatly improves the performance of multi-task incremental fine-tuning learning. Especially in the field of network traffic analysis, the pre-trained model can better capture the key features in complex network environments.
[0143] The attention mechanism calculates queries, keys, and values, resulting in three basic objects: the query vector (Query), the key vector (Key), and the value vector (Value). These basic objects combine input features with shared knowledge, helping the model better learn the relationship between network traffic features and existing knowledge, thereby achieving better performance in multi-task learning.
[0144] By multiplying the first attention matrix with the first value vector V, the pre-trained model can apply the attention weights to the value vector to generate an output vector of the common knowledge representation layer that fuses the input features and the shared knowledge representation.
[0145] Furthermore, the step S3 includes the following steps:
[0146] S31. Set up a common knowledge representation layer architecture;
[0147] S32. Set up a specific task knowledge layer architecture;
[0148] In S31, each network flow of the extracted data collection is processed by the pre-trained model to generate a feature vector of a specific network flow with a size of 1×768, which is recorded as the first feature vector Input bert , the common knowledge representation layer represents a matrix of size 192×768, and the first eigenvector Input bert Add to the end of the common knowledge representation layer to realize the first feature vector Input bert Fusion with the shared knowledge representation layer to form the first fusion matrix temp net , whose size is 193×768, where the first 192 rows are the rows of the matrix represented by the original common knowledge representation layer, and the last row is the first eigenvector Input bert ;
[0149] Use the attention mechanism to the first fusion matrix temp netPerform attention calculation, that is, generate the first query vector Q, the first key vector K and the first value vector V;
[0150] The first query vector Q is expressed as:
[0151] Q=W q Input bert
[0152] Among them, W q is the weight matrix of the first query vector Q;
[0153] The first key vector K is expressed as:
[0154] K=W k tmp net
[0155] Among them, W k is the weight matrix of the first key vector K;
[0156] The first value vector V is represented as:
[0157] V=W v tmp net
[0158] Among them, W v is the weight matrix of the first value vector V;
[0159] According to the first query vector Q, the first key vector K and the first value vector V, the first attention matrix A is calculated and multiplied by the first value vector V to obtain the output vector Output of the shared knowledge representation layer, whose size is 1×768;
[0160] The first attention matrix A is expressed as:
[0161]
[0162] Where √D is the scaling factor, D represents the dimension of the hidden layer vector, and the softmax function represents the normalization process;
[0163] The output vector Output of the shared knowledge representation layer is expressed as:
[0164] Output=W o AV+b o
[0165] Among them, W o is the output linear transformation weight matrix, b o is the bias term;
[0166] In S32, each network flow of the extracted data collection is processed by the pre-training model and combined with the common knowledge representation layer to generate a feature vector of size 1×768, which is recorded as the second feature vector Output G , the task-specific knowledge layer represents a size of num task ×768 matrix, where num task Indicates the number of current tasks, the second eigenvector Output G Splice with the task-specific knowledge layer on the first dimension to obtain the matrix tmp of spliced specific tasks and shared knowledge process , whose dimension is (n+1)×768. For each task, extract the one-dimensional task-specific knowledge vector corresponding to the task, which is recorded as the specific knowledge vector tmp input , which represents the specific knowledge required for the current task, using the specific knowledge vector tmp input Perform attention calculation with the shared knowledge to generate the second query vector Q', the second key vector K' and the second value vector V';
[0167] The second query vector Q' is expressed as:
[0168] Q'=W q 'tmp input
[0169] Among them, W q ' is the weight matrix of the second query vector Q';
[0170] The second key vector K' is expressed as:
[0171] K'=W k 'tmp process
[0172] Among them, W k ' is the weight matrix of the second key vector K', which is based on the matrix tmp of the spliced specific task and shared knowledge process , which is used to measure the relevance of the query vector to task-specific knowledge and shared knowledge;
[0173] The second value vector V' is represented as:
[0174] V'=W v 'tmp process
[0175] Among them, W v ' is the weight matrix of the second value vector V', which acts on the matrix tmp of the spliced specific task and shared knowledge process , helping the pre-trained model extract useful information from task-related specific knowledge and shared knowledge.
[0176] Specifically, to achieve incremental learning and enhance the model's ability to transfer and retain knowledge from previous tasks when tackling new tasks, the present invention adds a task-specific knowledge layer on top of the common knowledge representation layer. The introduction of the task-specific knowledge layer aims to encode the unique knowledge of each task, ensuring that the model can maintain its understanding and performance of previous tasks when tackling new tasks. At the same time, it extracts useful knowledge from previous tasks for new tasks, thereby enhancing the model's overall learning ability and flexibility.
[0177] The task-specific knowledge layer consists of a matrix of size n×768. Each task has a corresponding 1×768 vector, which is used to capture the specific knowledge representation corresponding to the task. The second eigenvector Output G It incorporates a fusion of deep-level features of network flows and general knowledge, representing the pre-trained model's understanding of the essential characteristics of network flows. Unlike the shared knowledge representation layer, the task-specific knowledge layer is optimized specifically for the uniqueness of each task, providing the model with a refined representation of task knowledge. This allows the pre-trained model to quickly adapt to new tasks while retaining important knowledge representations from previous tasks, preventing forgetting.
[0178] Each row vector in the task-specific representation matrix represents the specific knowledge of each task, and the concatenated matrix of task-specific and shared knowledge is tmp process It contains both general knowledge representation and task-specific knowledge representation, forming a complete feature fusion. This is then fed into the attention mechanism, which dynamically assigns weights to determine whether the current task should focus more on common knowledge or task-specific knowledge. The pre-trained model can flexibly switch its focus between different tasks, ensuring optimal processing of the current task and ensuring that new tasks can benefit from the specific knowledge of old tasks without affecting the model's performance on the old tasks.
[0179] The design goal of the task-specific knowledge layer is to ensure sufficient model plasticity. That is, when learning new tasks, the model will not interfere with the performance of previously learned tasks. Instead, through sharing mechanisms, the new task can benefit from the knowledge of the old tasks, thereby improving the model's performance on complex tasks. This is especially true in multi-task learning or incremental learning scenarios, where the pre-trained model faces continuous input from different tasks and can effectively manage and utilize the knowledge shared between these tasks.
[0180] refer to Figure 5 , layer n is the nth layer of the task-specific knowledge layer, layer iThe attention mechanism of the task-specific knowledge layer is designed to combine task-specific knowledge and shared knowledge so that the model can benefit from both sources of knowledge at the same time, improve the performance of the model in specific tasks, and maintain its good performance on existing tasks. The task-specific knowledge layer realizes the sharing and transfer of new task knowledge and old task knowledge by dynamically adding task representation vectors. New tasks can not only benefit from shared knowledge, but also draw experience from old task-specific knowledge, further enhancing the generalization ability and learning efficiency of the model, while maintaining the robustness of the pre-trained model in processing complex and changeable network tasks.
[0181] Although the present invention has been described with respect to a limited number of embodiments, it will be apparent to those skilled in the art, having benefit of the foregoing description, that other embodiments are contemplated within the scope of the invention thus described. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and didactic purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Consequently, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is intended to be illustrative rather than restrictive of the scope of the invention, which is defined by the appended claims.
Claims
1. A method for designing a network traffic multi-task incremental pre-training model architecture, characterized by: The following steps are involved: S1. Construct a multi-task incremental pre-training model architecture that separates the pre-training model from the common knowledge representation layer and introduces a task-specific representation layer; Specifically, the multi-task incremental pre-training model architecture includes a pre-training model, a common knowledge representation layer, and a task-specific knowledge layer. The pre-training model and the common knowledge representation layer are separated, and the pre-training model is used as an independent module, and the common knowledge representation layer is connected to the task-specific knowledge layer. S2. Train the pre-trained model in the multi-task incremental pre-trained model architecture to obtain a trained multi-task incremental pre-trained model architecture; S3. Fine-tune the trained multi-task incremental pre-training model architecture to obtain the final multi-task incremental pre-training model architecture based on the set common knowledge representation layer architecture and specific task representation layer architecture; The S3 includes the following steps: S31. Set up a common knowledge representation layer architecture; S32. Set up a specific task knowledge layer architecture; In the above S31, each network flow of the extracted data collection is processed by the pre-training model to generate a feature vector of a specific network flow with a size of 1×768, which is recorded as the first feature vector , the common knowledge representation layer represents a matrix of size 192×768, and the first eigenvector Add to the end of the common knowledge representation layer to achieve the first eigenvector Fusion with the common knowledge representation layer to obtain the first fusion matrix , whose size is 193×768, where the first 192 rows are the rows of the matrix represented by the original common knowledge representation layer, and the last row is the first eigenvector ; Use the attention mechanism to the first fusion matrix Perform attention calculation to generate the first query vector , first bond vector and the first value vector ; According to the first query vector , first bond vector and the first value vector , calculate the first attention matrix A, and combine the first attention matrix with the first value vector Multiply them together to get the output vector of the shared knowledge representation layer , its size is 1×768; In the above S32, each network flow of the extracted data collection is processed by the pre-training model and combined with the common knowledge representation layer to generate a feature vector of size 1×768, which is recorded as the second feature vector. , the task-specific knowledge layer represents a ×768 matrix, where Indicates the number of current tasks, and the second eigenvector Splice with the task-specific knowledge layer on the first dimension to obtain the matrix of spliced specific tasks and shared knowledge , whose dimension is (n+1)×768. For each task, a one-dimensional task-specific knowledge vector corresponding to the task is extracted and recorded as a specific knowledge vector , using a specific knowledge vector Perform attention calculation with shared knowledge to generate the second query vector , the second key vector and the second value vector .
2. A network traffic multi-task incremental pre-training model architecture design method according to claim 1, characterized in that: Said S2 comprises the following steps: S21. Calculate and rank the discriminative power of the features extracted by CICFlowMeter using principal component analysis; S22. Block processing is performed on the features of different distinguishing strengths extracted by CICFlowMeter to achieve feature form conversion; S23. According to the features after the form conversion, two pre-training tasks, namely, the masking task and the service prediction task, are set for the pre-training model for training; In the step S21, calculating and sorting the distinguishing strength of the features extracted by CICFlowMeter includes data preprocessing, calculating the covariance matrix, calculating the eigenvalues and eigenvectors to select the principal components, calculating the distinguishing strength of the features, and normalizing the distinguishing strength; In the data preprocessing process, the input data of the pre-trained model, that is, the network dataset, is cleaned, and the network dataset is checked for missing values. The missing values are filled with 0, and the input data is standardized by the standardization formula to obtain the standardized data matrix. ; Normalized data matrix Expressed as: ; Where X is the original data matrix, μ is the mean vector of the feature, and σ is the standard deviation vector of the feature; In the process of calculating the covariance matrix, according to the standardized data matrix , calculate the covariance matrix C; The covariance matrix C is expressed as: ; in, represents matrix transpose, n represents the number of features; Perform eigenvalue decomposition on the covariance matrix C and calculate the eigenvalues and eigenvectors of the covariance matrix C; The decomposition process is expressed as: ; in, is the i-th eigenvalue of the covariance matrix, is the corresponding eigenvector, and the integration results in a set of eigenvalues and the corresponding feature vector set ; Sort the principal components according to the size of the eigenvalue, select the principal component whose eigenvalue is greater than the set contribution value, and obtain the eigenvector corresponding to the principal component And combined into a dimensionality reduction matrix ,in, The number of principal components is the standardized data matrix Project it onto the selected principal component to obtain the dimension-reduced data Z; ; In the process of calculating the distinguishing strength of features, the importance of each feature is measured by its contribution to each principal component, and the contribution rate is calculated. Get the importance of each feature, that is, get the discrimination power of CICFlowMeter features and sort them; Contribution rate Expressed as: ; Among them, Variance explained by the feature represents the variance explained by the feature, and Total variance explained by all features represents all variance explained by all features; In the S22, a vocabulary mapping in the range of [0, 65535] is used to divide the features extracted by CICFlowMeter into blocks of 16 bytes each, and the feature values exceeding 16 bytes, i.e., greater than 65535, are processed in blocks, separated by a delimiter token [sep] to indicate that the separated blocks belong to the same feature value. After the block processing, each feature is converted into and Two forms.
3. A method for designing a network traffic multi-task incremental pre-training model architecture according to claim 2, characterized in that: In S3, the first query vector Expressed as: ; in, is the first query vector The weight matrix of First key vector Expressed as: ; in, is the first key vector The weight matrix of First value vector Expressed as: ; in, is the first value vector The weight matrix of The first attention matrix A is expressed as: ; in, is the scaling factor, D represents the dimension of the hidden layer vector, and the softmax function represents the normalization process; Output vector of the shared knowledge representation layer Expressed as: ; in, is the output linear transformation weight matrix, is the bias term; Second query vector Expressed as: ; in, is the second query vector The weight matrix of Second key vector Expressed as: ; in, is the second key vector The weight matrix of Second value vector Expressed as: ; in, is the second value vector The weight matrix of .
Citation Information
Patent Citations
ET-BERT traffic classification method based on multi-task learning, storage medium and equipment
CN116155821A
Multi-rhetorical text generation method based on BART model
CN117725964A