Model training method and apparatus, device, and medium
By classifying the original data set and expanding the data dimensions, the extended data set is generated, the data imbalance problem is solved, the accuracy of network traffic detection is improved, and the detection effect of the model is enhanced.
Patent Information
- Application Number
- PCT/CN2024/135861
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-03
AI Technical Summary
Using data imbalanced network traffic datasets to establish an attack detection model will reduce the accuracy of subsequent detection results of network traffic detection using the model.
By classifying the original data set, attack characteristics corresponding to attack network traffic with a small amount of data are obtained, and then the attack characteristics are expanded multiple times to generate an extended data set, and the bidirectional time convolution network Bi-TCN model is trained together with the original data set to solve the data imbalance problem.
The accuracy of network traffic detection results is improved, and an extended data set with attack characteristics is generated, which enhances the effectiveness of the model for network traffic detection.
Smart Images

Figure CN2024135861_03072025_PF_FP_ABST
Abstract
Description
Model training method, device, equipment and medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 28, 2023, with application number 202311845143.8 and application name "A model training method, device, equipment and medium", the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of network security technology, and in particular to a model training method, apparatus, device and medium. Background Art
[0004] With the rapid development of information technology, various infrastructures are inseparable from the network and are therefore inevitably subject to threats and attacks from network traffic. The security of network traffic plays a vital role in the smooth operation of infrastructure systems. Currently, network traffic detection methods can be mainly divided into rule-based detection, statistical detection, and machine learning-based detection.
[0005] Rule-based network traffic detection uses prior knowledge of attacks, such as attack characteristics, to create rules and perform threat traffic detection. Statistical detection, on the other hand, detects anomalies by establishing the statistical distribution of intrusion patterns.
[0006] However, rule-based traffic detection methods require manual creation of rules tailored to specific network traffic types, making rule updates difficult and resulting in low detection accuracy and efficiency. Statistical methods also have very high computational costs and limited ability to process large amounts of data. Consequently, a growing number of researchers are applying machine learning-based network traffic detection methods to the security of various infrastructure systems.
[0007] Currently, the primary approach to building attack detection systems for network traffic is to create attack detection models using publicly available network traffic datasets. However, most network traffic datasets suffer from data imbalance. Using unbalanced datasets to build attack detection models reduces the accuracy of subsequent network traffic detection results using the model. Summary of the Invention
[0008] The embodiments of the present application provide a model training method, apparatus, device and medium for solving the problem that using an unbalanced network traffic data set to establish an attack detection model will reduce the accuracy of the detection results of subsequent network traffic detection using the model.
[0009] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0010] Obtaining an original data set and a bidirectional temporal convolutional network (Bi-TCN) model to be trained; the original data set includes multiple network traffic data;
[0011] Performing feature classification on the original data set to obtain multiple attack features in the original data set;
[0012] Based on the multiple target attack features, an iterative operation is performed until the similarity between the multiple extended data of the current iteration and the multiple network flow data in the original data set is greater than a similarity threshold, and the iterative operation is terminated; the iterative operation includes: performing data dimension expansion on each network flow data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain the multiple extended data; if the current iteration is the first iteration, the multiple target attack features are the multiple attack features; if the current iteration is not the first iteration, the multiple target attack features are the attack features corresponding to the multiple extended data obtained in the previous iteration;
[0013] The multiple extended data obtained in each iteration are used as the extended data corresponding to the multiple attack features to obtain an extended data set;
[0014] The Bi-TCN model to be trained is trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0015] In some embodiments, after obtaining the extended data set and before training the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model, the method further includes:
[0016] Performing quality checks on each network flow extension data in the extended data set to obtain network flow extension data that meets preset quality inspection rules;
[0017] The network traffic extension data that meets the preset quality detection rules is used as the network traffic extension data in the extended data set.
[0018] In some embodiments, the network traffic extension data that meets the preset quality detection rules includes:
[0019] Network traffic extension data that does not include illegal values;
[0020] When the maximum mean difference between the extended data set and the original data set is less than a preset difference threshold, the extended data set includes network traffic extended data;
[0021] The network traffic extension data includes a plurality of preset attack features; the plurality of preset attack features are used to perform feature classification on the original data set.
[0022] In some embodiments, the feature classification of the original data set to obtain multiple attack features in the original data set includes:
[0023] For each network traffic data in the original data set:
[0024] Extracting features from the network traffic data to obtain a plurality of feature information corresponding to the network traffic data;
[0025] The feature information that is the same as any one of the preset multiple attack features is used as the attack feature in the original data set.
[0026] In some embodiments, the data dimension expansion of each network traffic data corresponding to the multiple target attack features to obtain multiple reference data includes:
[0027] For each network traffic data, perform the following operations:
[0028] Combining the target attack features corresponding to the network traffic data into a one-dimensional feature vector, and multiplying the one-dimensional feature vector by the identity matrix to obtain a first matrix;
[0029] Determining the row number and column number of each parameter in the two-dimensional matrix according to the number of target attack features corresponding to the network traffic data;
[0030] For each parameter in the two-dimensional matrix: if the row number and the column number of the parameter are the same, the value of the parameter is zero; if the row number and the column number of the parameter are different, determining, based on the first matrix, a column vector corresponding to the row number and a column vector corresponding to the column number, and determining the value of the parameter based on the column vector corresponding to the row number and the column vector corresponding to the column number;
[0031] The two-dimensional matrix is input into a generator of the CL-WGAN model to generate the plurality of reference data.
[0032] In some embodiments, performing traffic classification on the plurality of reference data to obtain the plurality of extended data includes:
[0033] Inputting the multiple reference data into a convolutional bidirectional long short-term memory neural network (CNN-BiLSTM), performing feature extraction on the multiple reference data to obtain feature vectors;
[0034] Performing forward and reverse feature extraction on the feature vector to obtain sequence feature vectors in two directions, and fusing the sequence feature vectors in the two directions to obtain traffic classification results of the multiple reference data;
[0035] The traffic classification result is characterized as reference data of attack network traffic as extended data; the attack network traffic is network traffic having feature information identical to any one of a plurality of preset attack features.
[0036] In some embodiments, training the Bi-TCN model to be trained based on the original dataset and the extended dataset to obtain a trained Bi-TCN model includes:
[0037] The original data set and the extended data set are used as training data sets, input into the Bi-TCN model to be trained, and classification results of each network traffic data in the training data set are obtained;
[0038] Determining a loss function of the Bi-TCN model to be trained based on the classification result and the labeled data in the training data set; the labeled data in the training data set indicates whether each network traffic data in the training data set is attack network traffic;
[0039] Based on the loss function of the Bi-TCN model to be trained, the weight parameters and bias parameters of the Bi-TCN model to be trained are adjusted until the classification result is consistent with the labeled data in the training data set, thereby obtaining a trained Bi-TCN model.
[0040] In some embodiments, the Bi-TCN model includes an input layer, a convolutional layer, a fully connected layer, and a softmax layer, the convolutional layer includes n residual modules, each residual module includes a causal void convolution unit; the expansion coefficient of the causal void convolution unit of each residual module increases exponentially; the method further includes:
[0041] Input the original data set into the convolution layer, perform forward and reverse feature extraction in each residual module according to the dilation coefficient of the causal hole convolution unit, obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in the two directions and input them into the next residual module until the last residual module outputs the feature vector after convolution processing;
[0042] Inputting the convolution-processed feature vector into the fully connected layer, and determining a reference feature vector based on weight parameters and bias parameters in the trained Bi-TCN model;
[0043] The reference feature vector output by the fully connected layer is input into the softmax layer for activation operation to obtain the classification results of the multiple network traffic data in the original data set.
[0044] In a second aspect, an embodiment of the present application provides a model training device, the device comprising:
[0045] A first acquisition module is used to acquire an original data set and a bidirectional temporal convolutional network Bi-TCN model to be trained; the original data set includes multiple network traffic data;
[0046] A classification module, configured to perform feature classification on the original data set to obtain a plurality of attack features in the original data set;
[0047] The expansion module is configured to perform an iterative operation based on multiple target attack features until the similarity between the multiple extended data of this iteration and the multiple network flow data in the original data set is greater than a similarity threshold, thereby terminating the iterative operation; the iterative operation includes: performing data dimension expansion on each network flow data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain the multiple extended data; if this iteration is the first iteration, the multiple target attack features are the multiple attack features; if this iteration is not the first iteration, the multiple target attack features are the attack features corresponding to the multiple extended data obtained in the previous iteration;
[0048] an extended data set determining module, configured to use the plurality of extended data obtained in each iteration as the extended data corresponding to the plurality of attack features to obtain an extended data set;
[0049] A training module is used to train the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0050] In some embodiments, the apparatus further comprises:
[0051] A quality detection module is used to perform quality inspection on each network flow extension data in the extended data set to obtain network flow extension data that meets preset quality detection rules;
[0052] The network traffic extension data that meets the preset quality detection rules is used as the network traffic extension data in the extended data set.
[0053] In some embodiments, the network traffic extension data that meets the preset quality detection rules includes:
[0054] Network traffic extension data that does not include illegal values;
[0055] When the maximum mean difference between the extended data set and the original data set is less than a preset difference threshold, the extended data set includes network traffic extended data;
[0056] The network traffic extension data includes a plurality of preset attack features; the plurality of preset attack features are used to perform feature classification on the original data set.
[0057] In some embodiments, the classification module is specifically configured to:
[0058] For each network traffic data in the original data set:
[0059] Extracting features from the network traffic data to obtain a plurality of feature information corresponding to the network traffic data;
[0060] The feature information that is the same as any one of the preset multiple attack features is used as the attack feature in the original data set.
[0061] In some embodiments, the expansion module is specifically configured to:
[0062] For each network traffic data, perform the following operations:
[0063] Combining the target attack features corresponding to the network traffic data into a one-dimensional feature vector, and multiplying the one-dimensional feature vector by the identity matrix to obtain a first matrix;
[0064] Determining the row number and column number of each parameter in the two-dimensional matrix according to the number of target attack features corresponding to the network traffic data;
[0065] For each parameter in the two-dimensional matrix: if the row number and the column number of the parameter are the same, the value of the parameter is zero; if the row number and the column number of the parameter are different, determining, based on the first matrix, a column vector corresponding to the row number and a column vector corresponding to the column number, and determining the value of the parameter based on the column vector corresponding to the row number and the column vector corresponding to the column number;
[0066] The two-dimensional matrix is input into a generator of the CL-WGAN model to generate the plurality of reference data.
[0067] In some embodiments, the expansion module is specifically configured to:
[0068] Inputting the multiple reference data into a convolutional bidirectional long short-term memory neural network (CNN-BiLSTM), performing feature extraction on the multiple reference data to obtain feature vectors;
[0069] Performing forward and reverse feature extraction on the feature vector to obtain sequence feature vectors in two directions, and fusing the sequence feature vectors in the two directions to obtain traffic classification results of the multiple reference data;
[0070] The traffic classification result is characterized as reference data of attack network traffic as extended data; the attack network traffic is network traffic having feature information identical to any one of a plurality of preset attack features.
[0071] In some embodiments, the training module is specifically used to:
[0072] The original data set and the extended data set are used as training data sets, input into the Bi-TCN model to be trained, and classification results of each network traffic data in the training data set are obtained;
[0073] Determining a loss function of the Bi-TCN model to be trained based on the classification result and the labeled data in the training data set; the labeled data in the training data set indicates whether each network traffic data in the training data set is attack network traffic;
[0074] Based on the loss function of the Bi-TCN model to be trained, the weight parameters and bias parameters of the Bi-TCN model to be trained are adjusted until the classification result is consistent with the labeled data in the training data set, thereby obtaining a trained Bi-TCN model.
[0075] In some embodiments, the Bi-TCN model includes an input layer, a convolutional layer, a fully connected layer, and a softmax layer, the convolutional layer includes n residual modules, each residual module includes a causal void convolution unit; the expansion coefficient of the causal void convolution unit of each residual module increases exponentially;
[0076] The device further includes a detection module, configured to:
[0077] Input the original data set into the convolution layer, perform forward and reverse feature extraction in each residual module according to the dilation coefficient of the causal hole convolution unit, obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in the two directions and input them into the next residual module until the last residual module outputs the feature vector after convolution processing;
[0078] Inputting the convolution-processed feature vector into the fully connected layer, and determining a reference feature vector based on weight parameters and bias parameters in the trained Bi-TCN model;
[0079] The reference feature vector output by the fully connected layer is input into the softmax layer for activation operation to obtain the classification results of the multiple network traffic data in the original data set.
[0080] In a third aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein:
[0081] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the above-mentioned model training method.
[0082] In a fourth aspect, an embodiment of the present application provides a storage medium. When the computer program in the storage medium is executed by a processor of an electronic device, the electronic device can execute the above-mentioned model training method.
[0083] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is executed by an electronic device, the electronic device can implement the above-mentioned model training method provided in the present application.
[0084] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0085] In an embodiment of the present application, an original data set and a bidirectional temporal convolutional network Bi-TCN model to be trained are obtained; the original data set includes multiple network traffic data; feature classification is performed on the original data set to obtain multiple attack features in the original data set; an iterative operation is performed based on multiple target attack features until the similarity between the multiple extended data of this iteration and the multiple network traffic data in the original data set is greater than a similarity threshold, and the iterative operation is terminated; the iterative operation includes: data dimension expansion of each network traffic data corresponding to the multiple target attack features to obtain multiple reference data; traffic classification is performed on the multiple reference data to obtain multiple extended data; if this iteration is the first iteration, the multiple target attack features are multiple attack features; if this iteration is not the first iteration, the multiple target attack features are attack features corresponding to the multiple extended data obtained in the previous iteration; the multiple extended data obtained in each iteration are used as extended data corresponding to the multiple attack features to obtain an extended data set; the Bi-TCN model to be trained is trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0086] Therefore, by classifying the original data set, we obtain the attack features corresponding to the attack network traffic with a smaller amount of data, and then perform multiple data dimension expansion on the attack features to generate an extended data set, thereby improving the attack network traffic with attack features. Then, we use the extended data set and the original data set together to train the Bi-TCN model, which effectively solves the data imbalance problem and further improves the accuracy of the detection results of network traffic detection using the model.
[0087] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0089] FIG1 is a schematic diagram of the system structure of a model training method provided in an embodiment of the present application;
[0090] FIG2 is a flow chart of a model training method provided in an embodiment of the present application;
[0091] FIG3 is a schematic diagram of the structure of a CL-WGAN model provided in an embodiment of the present application;
[0092] FIG4 is a schematic diagram of a process of data dimension expansion provided in an embodiment of the present application;
[0093] FIG5 is a flow chart of a traffic classification process provided by an embodiment of the present application;
[0094] FIG6 is a schematic diagram of a process for training a Bi-TCN model provided in an embodiment of the present application;
[0095] FIG7 is a schematic structural diagram of a Bi-TCN model provided in an embodiment of the present application;
[0096] FIG8 is a flow chart of a method for detecting network traffic data provided by an embodiment of the present application;
[0097] FIG9 is a substructure diagram of a convolutional layer provided in an embodiment of the present application;
[0098] FIG10 is a schematic structural diagram of a model training device provided in an embodiment of the present application;
[0099] FIG11 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0100] To make the purpose, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Among them, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0101] Moreover, in the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0102] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0103] To facilitate understanding of this application, some technical terms involved in this application are introduced below:
[0104] 1. Data imbalance: This refers to a significant disparity in the number of samples across different categories within the same dataset, making it difficult to extract features from a small number of samples. Data imbalance can be categorized into two main types: large data imbalance and small data imbalance. Large data imbalance refers to a small proportion of samples from a particular category within a larger dataset; small data imbalance refers to a smaller number of samples from a particular category within a smaller dataset.
[0105] 2. Temporal Convolutional Network (TCN): It refers to a new model that combines convolutional neural networks with time series models and applies it to time series data classification.
[0106] 3. Generative Adversarial Network (GAN): This consists of a generative network and a discriminative network. Inspired by the zero-sum game in game theory, the data generation problem in GAN can be viewed as an adversarial zero-sum game between the discriminative and generative networks, achieving model optimization in this process. It is currently widely used in the field of deep learning for images.
[0107] With the rapid development of information technology, various infrastructures are inseparable from the network and are therefore inevitably subject to threats and attacks from network traffic. The security of network traffic plays a vital role in the smooth operation of infrastructure systems. Currently, network traffic detection methods can be mainly divided into rule-based detection, statistical detection, and machine learning-based detection.
[0108] Rule-based network traffic detection uses prior knowledge of attacks, such as attack characteristics, to create rules and perform threat traffic detection. Statistical detection, on the other hand, detects anomalies by establishing the statistical distribution of intrusion patterns.
[0109] However, rule-based traffic detection methods require manual creation of rules tailored to specific network traffic types, making rule updates difficult and resulting in low detection accuracy and efficiency. Statistical methods also have very high computational costs and limited ability to process large amounts of data. Consequently, a growing number of researchers are applying machine learning-based network traffic detection methods to the security of various infrastructure systems.
[0110] Currently, the primary approach to building attack detection systems for network traffic is to create attack detection models using publicly available network traffic datasets. However, most network traffic datasets suffer from data imbalance. Using unbalanced datasets to build attack detection models reduces the accuracy of subsequent network traffic detection results using the model.
[0111] In view of this, the embodiments of the present application provide a model training method, apparatus, equipment and medium for solving the problem that using a data-unbalanced network traffic data set to establish an attack detection model will reduce the accuracy of the detection results of subsequent network traffic detection using the model.
[0112] The inventive concept of the embodiment of the present application: In the embodiment of the present application, by classifying the original data set, the attack features corresponding to the attack network traffic with a smaller amount of data are obtained, and then the attack features are expanded in data dimensions multiple times to generate an extended data set, thereby improving the attack network traffic with attack features, and then the extended data set and the original data set are used together to train the Bi-TCN model, which effectively solves the data imbalance problem and further improves the accuracy of the detection results of the network traffic detection using the model.
[0113] In order to build a real-time, efficient, and accurate attack detection model in the embodiments of this application, on the one hand, the data set must have sufficient attack traffic characteristics; on the other hand, the model must be able to learn the attack characteristics of the traffic for classification. Therefore, first of all, in terms of data, in order to address the problem that the imbalance of traffic data in the public data set will reduce the accuracy of abnormal traffic detection, the abnormal traffic data with a small amount of data is supplemented by sample generation, and the pass rate of the generated traffic quality inspection is high, which effectively solves the data imbalance problem and obtains a balanced data set. In terms of the model, the existing time convolutional network is improved to a bidirectional time convolutional network model, which can better capture traffic feature information in a longer time dimension and has low model space complexity.
[0114] As shown in Figure 1, a schematic diagram of the system structure of the model training method in the embodiment of the present application is shown. First, the CL-WGAN model is used to process the data imbalance of the original data set. Then, the balanced data is preprocessed by data integration, data cleaning, data normalization, etc., and then the Bi-TCN deep learning model is trained and tested to detect network traffic and obtain classification results, that is, the attack network traffic and normal network traffic in the original data set are obtained.
[0115] To further illustrate the technical solutions provided by the embodiments of the present application, the following is a detailed description of the technical solutions in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of the present application provide the method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or no creative work. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided in the embodiments of the present application.
[0116] See Figure 2, which is a flow chart of a model training method provided in an embodiment of the present application. The method includes the steps shown in Figure 2:
[0117] In step 201, an original data set and a bidirectional temporal convolutional network Bi-TCN model to be trained are obtained; the original data set includes a plurality of network traffic data.
[0118] After obtaining the original data set, the original data in the original data set is subjected to data preprocessing such as data cleaning and normalization.
[0119] In step 202, feature classification is performed on the original data set to obtain multiple attack features in the original data set.
[0120] In some embodiments, feature classification is performed on the original data set to obtain multiple attack features in the original data set. The following steps may be performed for each network traffic data in the original data set:
[0121] Feature extraction is performed on the network traffic data to obtain multiple feature information corresponding to the network traffic data; feature information that is the same as any one of the preset multiple attack features is used as the attack feature in the original data set.
[0122] Among them, the preset multiple attack features can be set according to actual experience or according to actual needs, and this application does not impose any restrictions on this.
[0123] In practice, the data expansion in this application is generated only for attack signatures. Therefore, the original dataset must first be divided into attack network traffic and normal network traffic. Therefore, multiple attack signatures are pre-set and stored on the server. The features of each network traffic data point in the original dataset are then extracted and compared with the pre-set attack signatures. The features of all data in the original dataset are then classified as attack signatures or non-attack signatures. A feature that matches any of the pre-set attack signatures is considered an attack signature; otherwise, it is considered a non-attack signature.
[0124] In step 203, an iterative operation is performed based on multiple target attack features until the similarity between the multiple extended data of this iteration and the multiple network traffic data in the original data set is greater than a similarity threshold, and the iterative operation is terminated; the iterative operation includes: expanding the data dimension of each network traffic data corresponding to the multiple target attack features to obtain multiple reference data; and performing traffic classification on the multiple reference data to obtain multiple extended data.
[0125] If this iteration is the first iteration, the multiple target attack features are multiple attack features; if this iteration is not the first iteration, the multiple target attack features are attack features corresponding to multiple extended data obtained in the previous iteration.
[0126] In specific implementation, as shown in Figure 3, when the CL-WGAN model is used to process the data imbalance of the original dataset, the original dataset is input into the CL-WGAN model, and passes through the data dimension expansion unit, generator, classification control unit, and discriminator in the model to finally generate an extended dataset.
[0127] In some embodiments, when data dimension expansion is performed on each network traffic data corresponding to multiple target attack features to obtain multiple reference data, the operations shown in FIG4 are performed for each network traffic data:
[0128] In step 401, target attack features corresponding to network traffic data are formed into a one-dimensional feature vector, and the one-dimensional feature vector is multiplied by the identity matrix to obtain a first matrix;
[0129] In step 402, the row number and column number of each parameter in the two-dimensional matrix are determined according to the number of target attack features corresponding to the network traffic data;
[0130] In step 403, for each parameter in the two-dimensional matrix: if the row number and column number of the parameter are the same, the value of the parameter is zero; if the row number and column number of the parameter are different, the column vector corresponding to the row number and the column vector corresponding to the column number are determined based on the first matrix, and the value of the parameter is determined based on the column vector corresponding to the row number and the column vector corresponding to the column number;
[0131] In step 404 , the two-dimensional matrix is input into the generator of the CL-WGAN model to generate a plurality of reference data.
[0132] In specific implementation, the network traffic corresponding to the attack features in the original data set is expanded by data dimension, as follows:
[0133] The target attack features corresponding to the network traffic data are formed into a one-dimensional feature vector. Assume that a network traffic corresponding to the attack feature can be expressed as X = [x1,…,x t ,…,x n ], where x1 represents a certain feature information. For each network flow, the network flow X can be transformed into a two-dimensional matrix X′ by expanding the data dimension.
[0134] Specifically, the one-dimensional eigenvector is multiplied by the identity matrix to obtain the first matrix, as shown in formula (1):
[0135] Then for i∈[n], according to formula (1) and formula (2), the matrix F corresponding to each feature information i is obtained i :
[0136] Next, for the row number j and column number k of each parameter in the two-dimensional matrix, j,k∈[n], the value x of each parameter in the two-dimensional matrix is defined according to formula (1): jk :
[0137] Among them, F j is the matrix F corresponding to the feature information j in formula (2)j , that is, the first matrix XI m The column vector of the jth column in ;
[0138] Finally, after calculating the values of each parameter in the two-dimensional matrix according to formula (3), we can get the two-dimensional matrix X ′ :
[0139] Therefore, the characteristic information x1 in the network traffic X is transformed into X'1=[x 11 ,…,x 1t ,…,x 1n ], and so on, other feature information x i Also transformed into the corresponding X i‘ =[x i1 ,…,x it ,…,x in ], the original one-dimensional vector is transformed into a two-dimensional matrix, and the matrix contains all the characteristic information of the original network traffic X, and the data after the data dimension is expanded is obtained. Then, the two-dimensional matrix after the data dimension is expanded is input into the generator to obtain the generated data, and the generated data is converted into a vector form through the inverse operation of the dimension conversion, that is, through the inverse operation of formulas (1), (2), and (3), multiple parameter flows corresponding to the network traffic are obtained.
[0140] In some embodiments, traffic classification is performed on a plurality of reference data to obtain a plurality of extended data, which may be performed as follows:
[0141] Input multiple reference data into the convolutional bidirectional long short-term memory neural network CNN-BiLSTM, perform feature extraction on the multiple reference data, and obtain feature vectors;
[0142] Perform forward and reverse feature extraction on the feature vector to obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in the two directions to obtain the traffic classification results of multiple reference data;
[0143] The traffic classification result is characterized as reference data of attack network traffic as extended data; the attack network traffic is network traffic having feature information identical to any one of a plurality of preset attack features.
[0144] In specific implementation, as shown in Figure 5, multiple parameter flows corresponding to network traffic X are input into the pre-trained CNN-BiLSTM classification control unit for sample classification inspection to obtain traffic classification results, including traffic that can be correctly classified and traffic that cannot be correctly classified, that is, attack network traffic and normal network traffic. Then, the reference data that passes the inspection, that is, the traffic data that can be correctly classified, is input into the discriminator, and the reference data that fails the inspection, that is, the traffic data that cannot be correctly classified, is directly discarded; the result of the discriminator is used as the extended data of this iteration and fed back to the generator, and the data dimension expansion and traffic classification operations are repeatedly performed until a high-quality extended data set is finally generated.
[0145] The CNN-BiLSTM classification control unit is primarily composed of an input layer, a CNN (Convolutional Neural Network) layer, a BiLSTM (Bi-directional Long Short-Term Memory) layer, a connection layer, and an output layer. The training dataset first enters the input layer, then passes through the CNN layer to extract relationships between features and generate feature vectors. It then enters the BiLSTM layer to learn the patterns between features, and finally, the output layer generates the classification results.
[0146] The description of each layer structure of the CNN-BiLSTM classification control unit is as follows:
[0147] Input layer: reads the pre-processed reference data into the model. For example, with n as the batch length of the reference data, the input vector can be expressed as X = [x1,…,x t ,…,x n ].
[0148] CNN layer: This layer extracts features from the input layer's reference data, identifies relationships between features, constructs important feature vectors, and then feeds these feature vectors to the BiLSTM layer. CNN layers can be further divided into convolutional layers, pooling layers, and fully connected layers. The convolutional and pooling layers primarily extract features from the input vectors, which are then passed through the fully connected layers to output feature vectors.
[0149] The BiLSTM layer performs forward and reverse feature extraction on the feature vectors extracted by the CNN layer to obtain sequence feature vectors in both directions. It then learns the characteristic relationships between the time series features and then fuses the sequence feature vectors in both directions based on these relationships to obtain a new feature vector. This layer uses a forward LSTM network and a reverse LSTM network, with two LSTMs interconnected on the input sequence. Data passes through the recurrent neural network in both the forward and reverse directions, better capturing the contextual features of the traffic data.
[0150] Output layer: The new feature vector obtained by the BiLSTM layer passes through the output layer to obtain the final traffic classification result.
[0151] In step 204, the multiple extended data obtained in each iteration are used as extended data corresponding to the multiple attack features to obtain an extended data set.
[0152] Among them, the condition for terminating the iterative operation in this application is: the similarity between the multiple extended data of this iteration and the multiple network flow data in the original data set is greater than the similarity threshold, and the iterative operation is terminated. The similarity between the multiple extended data of this iteration and the multiple network flow data in the original data set refers to the proportion of the number of the multiple extended data of this iteration that are identical to the multiple network flow data in the original data set in the multiple extended data. For example, this iteration generates 10 extended data, but 9 of them are identical to the multiple network flow data in the original data set, that is, the similarity reaches 90%, which is greater than the similarity threshold of 70%, then the iterative operation is terminated.
[0153] Then, multiple extended data obtained in each previous iteration are all used as extended data corresponding to multiple attack features to generate an extended data set.
[0154] In some embodiments, after obtaining the extended data set, and before training the Bi-TCN model to be trained based on the original data set and the extended data set to obtain the trained Bi-TCN model, the model training method provided in this application can also be performed as follows:
[0155] Perform quality checks on each network flow extension data in the extended data set to obtain network flow extension data that meets preset quality detection rules; and use the network flow extension data that meets the preset quality detection rules as the network flow extension data in the extended data set.
[0156] Among them, network traffic extended data that meets the preset quality detection rules includes:
[0157] Network traffic extension data that does not include illegal values;
[0158] When the maximum mean difference between the extended dataset and the original dataset is less than a preset difference threshold, the extended dataset includes network traffic extended data;
[0159] The network traffic extension data includes multiple preset attack features; the multiple preset attack features are used to perform feature classification on the original data set.
[0160] In specific implementation, after obtaining the extended dataset, each network traffic extended data in the extended dataset is quality checked both globally and locally. Overall, the maximum mean discrepancy (MMD) between the extended dataset and the original dataset is calculated; locally, each network traffic extended data is checked for illegal values and attack signatures.
[0161] Illegal value check: Checks each network traffic extension data in the extended data set for any illegal values. For example, illegal values include at least one of the following: character data that does not exist in the original data set, decimal points, or null values. The system then checks whether character data exists in the extended data set, whether multiple decimal points exist, or whether null values exist. Network traffic extension data that does not contain illegal values is considered to meet the preset quality inspection rules.
[0162] Maximum mean difference check: Check whether the original data set and the expanded data set are derived from the same distribution, that is, assuming D s =(x1,x2,…,x n )~P(x) and D t =(y1,y2,…,y n )~Q(y), calculate D s and D t The MMD value between the two datasets is calculated. The smaller the MMD value, the better the model performance and the more similar the two datasets are. Therefore, when the maximum mean difference between the extended dataset and the original dataset is less than the preset difference threshold, the network traffic extension data included in the extended dataset is regarded as network traffic extension data that meets the preset quality detection rules.
[0163] Among them, the preset difference threshold can be set according to experience or according to actual needs, and this application does not impose any restrictions on this.
[0164] Attack feature check: Check whether the network traffic extension data in the extended data set contains multiple preset attack features. The network traffic extension data including the preset multiple attack features is regarded as the network traffic extension data that meets the preset quality detection rules. The preset multiple attack features are used to perform feature classification on the original data set. The attack feature check is performed based on the attack principle. For example, taking the Reconnaissance attack in the Modbus protocol as an example, its attack features are command_address (command address) and response_address (response address), that is, the attack features of the generated network traffic extension data include command_address and response_address, while other non-attack features should not include these features.
[0165] Finally, the network traffic extension data in the extended dataset that meets the preset quality detection rules is used as the network traffic extension data in the extended dataset, and the network traffic extension data in the extended dataset that does not meet the preset quality detection rules is deleted and not used in the training process of the Bi-TCN model.
[0166] In step 205, the Bi-TCN model to be trained is trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0167] In some embodiments, training the Bi-TCN model to be trained based on the original dataset and the extended dataset to obtain a trained Bi-TCN model can be performed as follows:
[0168] The original dataset and the extended dataset are used as training datasets and input into the Bi-TCN model to be trained to obtain the classification results of each network traffic data in the training dataset;
[0169] Based on the classification results and the labeled data in the training dataset, the loss function of the Bi-TCN model to be trained is determined; the labeled data in the training dataset indicates whether each network traffic data in the training dataset is attack network traffic;
[0170] Based on the loss function of the Bi-TCN model to be trained, the weight parameters and bias parameters of the Bi-TCN model to be trained are adjusted until the classification result is consistent with the labeled data in the training dataset, thereby obtaining a trained Bi-TCN model.
[0171] As shown in Figure 6, the extended dataset is mixed with the original dataset to obtain a new dataset. The new traffic data is then preprocessed, including data integration, data cleaning, data normalization, and other operations. The new dataset is then divided into a training dataset and a test dataset. For example, the new dataset is divided into 10 parts, 9 of which are used as training datasets and 1 as a test dataset, and a 10-fold cross-validation is performed.
[0172] The training dataset is then fed into the Bi-TCN model for training to obtain the classification results. Based on the classification results and the labeled data in the training dataset, the loss function of the Bi-TCN model to be trained is determined. The convergence of the model is judged based on the loss function, and the weight parameters and bias parameters are updated according to the loss function through the optimizer. The model is continuously iterated and updated until the model converges to obtain the trained Bi-TCN model. The test dataset is then fed into the trained Bi-TCN model, and a 10-fold cross-validation is performed on the trained Bi-TCN model, and the test results are output.
[0173] In some embodiments, as shown in Figure 7, the trained Bi-TCN model includes an input layer, a convolutional layer, a fully connected layer and a softmax layer. The convolutional layer includes n residual modules (Residual block), each residual module includes two causal void convolution units (Dilated Causal Conv), two weight normalization (WeightNorm), two activation units (ReLU) and two regularization units (Dropout); the expansion coefficient of the causal void convolution unit of each residual module increases exponentially.
[0174] Therefore, in this application, after the Bi-TCN model is trained, the original data set is input into the trained Bi-TCN model to obtain the classification results of multiple network traffic data in the original data set, that is, the attack network traffic and normal network traffic in the original data set. Specifically, the steps shown in Figure 8 can be performed:
[0175] In step 801, the original data set is input into the convolution layer. In each residual module, forward and reverse feature extraction is performed according to the dilation coefficient of the causal hole convolution unit to obtain sequence feature vectors in two directions. The sequence feature vectors in the two directions are fused and input into the next residual module until the last residual module outputs the feature vector after convolution processing.
[0176] In step 802, the convolutional feature vector is input into the fully connected layer, and a reference feature vector is determined based on the weight parameters and bias parameters in the trained Bi-TCN model.
[0177] In step 803, the reference feature vector output by the fully connected layer is input into the softmax layer for activation operation to obtain the classification results of the multiple network traffic data in the original data set.
[0178] Before the original data set is input into the convolution layer, the original data set is processed according to the dimension information of the original data set, the size of the preset sliding window, and the step size of the preset sliding window to obtain a processed original data set, and then the processed original data set is input into the convolution layer.
[0179] Among them, the size of the preset sliding window and the step size of the preset sliding window can be set according to experience or according to actual needs, and this application does not impose any restrictions on this.
[0180] In specific implementation, the time series input vector of the original data set can be expressed as X = [x1, x2, ..., x t-1 ,x t ], where t is the batch length of the original dataset. The input x at time i is i It can be expressed as x i =[r,r2,…,r k-1 ,r k ], where k is the dimension of the data at that moment. The input vector X is sliced using the sliding window concept. Assuming the sliding window size is m, each slice along the time series produces a two-dimensional matrix of size m × k. When the sliding window step size is s, the final result is a three-dimensional matrix of size p × m × k, where p = ts, the number of sliding window slides. This converts the original two-dimensional matrix into a three-dimensional matrix, capturing all sequence information while preserving temporal characteristics.
[0181] Then, the three-dimensional matrix is input into the convolution layer for training. The convolution layer consists of n time blocks connected in series, each of which is a residual module, which can prevent the model from getting worse during training. As the network depth increases, the expansion coefficient d in the causal hole convolution will increase exponentially (d = 2 n-1 ), the expansion coefficient d increases the expansion field, allowing the model to learn information from a longer time ago and reducing the network depth.
[0182] Figure 9 shows the substructure of the convolutional layer. The input value of the next module is the linear superposition output value obtained by the skip mechanism in the residual connection of the previous module. This process is repeated to complete the convolution process, resulting in a two-dimensional matrix of size m×k as the final output. In each residual module, forward and reverse feature extraction is performed based on the dilation coefficient of the causal dilated convolution unit to obtain sequence feature vectors in two directions. These sequence feature vectors in two directions are then fused and input to the next residual module until the last residual module outputs the convolution-processed feature vector. For example, in a hidden layer with a dilation coefficient of 1, one neuron (the gray neuron) is selected for every two neurons. This neuron then obtains feature information from the corresponding three adjacent neurons in the input layer in both the forward and reverse directions. The sequence feature vectors in the two directions are fused and then input to the neurons in the next hidden layer. Similarly, the corresponding sequence feature vectors are obtained in the other hidden layers and finally output at the output layer.
[0183] The convolutional feature vector, an m×k two-dimensional matrix, is then passed through a fully connected layer for linear dimensionality reduction. This is then multiplied by the weight parameters and added with the bias parameters to produce a one-dimensional reference feature vector. Finally, this one-dimensional reference feature vector is activated through a sofamax layer, which maps the multiple outputs to the (0, 1) interval, yielding the probability distribution for each network traffic data point in the original dataset.
[0184] Among them, the probability distribution represents the probability that the network traffic data is normal network traffic. If the probability is greater than a preset probability threshold, the network traffic data is determined to be normal network traffic; otherwise, it is attack network traffic, and a classification result is obtained.
[0185] Among them, the preset probability threshold can be set according to experience or according to actual needs, and this application does not impose any restrictions on this.
[0186] The model training method provided in the embodiment of the present application can be used to expand attack traffic samples: constructing nearly realistic high-quality attack traffic samples based on real attack traffic data of a small sample, balancing the distribution ratio of normal traffic and attack traffic, and enriching the attack characteristics of attack traffic to help better optimize and establish subsequent attack detection models, or be used by security research experts for research experiments to reduce network security risks.
[0187] The model trained according to the model training method provided in the embodiment of the present application can also be used to detect attack events and block attacks: the Bi-TCN model trained in this application is used to detect attack traffic, which is highly efficient, and the Bi-TCN model has low complexity and consumes little resources on the deployed machine. It can obtain detection results in real time and accurately, and block attacks in real time.
[0188] The model training method provided in the embodiment of the present application can be used to enrich attack target range cases: in the current environment where security incidents occur frequently, the method of generating an extended data set in the model training method can generate similar real attack traffic for building network target ranges in different environments, and provide security personnel of each system with an attack actual combat simulation environment, so that the generated traffic can effectively act on the network target range, with high ease of use and practicality.
[0189] Based on the foregoing description, in an embodiment of the present application, an original data set and a bidirectional temporal convolutional network Bi-TCN model to be trained are obtained; the original data set includes multiple network traffic data; the original data set is feature classified to obtain multiple attack features in the original data set; based on multiple target attack features, an iterative operation is performed until the similarity between the multiple extended data of this iteration and the multiple network traffic data in the original data set is greater than a similarity threshold, and the iterative operation is terminated; the iterative operation includes: performing data dimension expansion on each network traffic data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain multiple extended data; if this iteration is the first iteration, the multiple target attack features are multiple attack features; if this iteration is not the first iteration, the multiple target attack features are attack features corresponding to the multiple extended data obtained in the previous iteration; the multiple extended data obtained in each iteration are used as extended data corresponding to the multiple attack features to obtain an extended data set; the Bi-TCN model to be trained is trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0190] Therefore, by classifying the original data set, we obtain the attack features corresponding to the attack network traffic with a smaller amount of data, and then perform multiple data dimension expansion on the attack features to generate an extended data set, thereby improving the attack network traffic with attack features. Then, we use the extended data set and the original data set together to train the Bi-TCN model, which effectively solves the data imbalance problem and further improves the accuracy of the detection results of network traffic detection using the model.
[0191] Based on the same technical concept, the embodiment of the present application also provides a model training device. The principle of solving the problem by the model training device is similar to that of the above-mentioned model training method. Therefore, the implementation of the model training device can refer to the implementation of the model training method, and the repeated parts will not be repeated.
[0192] FIG10 is a schematic diagram of the structure of a model training device provided in an embodiment of the present application, wherein the device includes a first acquisition module 1001, a classification module 1002, an expansion module 1003, an extended data set determination module 1004, and a training module 1005; wherein:
[0193] The first acquisition module 1001 is used to acquire an original data set and a bidirectional temporal convolutional network Bi-TCN model to be trained; the original data set includes multiple network traffic data;
[0194] A classification module 1002 is configured to perform feature classification on the original data set to obtain a plurality of attack features in the original data set;
[0195] The expansion module 1003 is configured to perform an iterative operation based on multiple target attack features until the similarity between the multiple extended data of this iteration and the multiple network flow data in the original data set is greater than a similarity threshold, thereby terminating the iterative operation. The iterative operation includes: performing data dimension expansion on each network flow data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain the multiple extended data; if this iteration is the first iteration, the multiple target attack features are the multiple attack features; if this iteration is not the first iteration, the multiple target attack features are the attack features corresponding to the multiple extended data obtained in the previous iteration;
[0196] The extended data set determining module 1004 is configured to use the multiple extended data obtained in each iteration as the extended data corresponding to the multiple attack features to obtain an extended data set;
[0197] The training module 1005 is used to train the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
[0198] In some embodiments, the apparatus further comprises:
[0199] A quality detection module is used to perform quality inspection on each network flow extension data in the extended data set to obtain network flow extension data that meets preset quality detection rules;
[0200] The network traffic extension data that meets the preset quality detection rules is used as the network traffic extension data in the extended data set.
[0201] In some embodiments, the network traffic extension data that meets the preset quality detection rules includes:
[0202] Network traffic extension data that does not include illegal values;
[0203] When the maximum mean difference between the extended data set and the original data set is less than a preset difference threshold, the extended data set includes network traffic extended data;
[0204] The network traffic extension data includes a plurality of preset attack features; the plurality of preset attack features are used to perform feature classification on the original data set.
[0205] In some embodiments, the classification module 1002 is specifically configured to:
[0206] For each network traffic data in the original data set:
[0207] Extracting features from the network traffic data to obtain a plurality of feature information corresponding to the network traffic data;
[0208] The feature information that is the same as any one of the preset multiple attack features is used as the attack feature in the original data set.
[0209] In some embodiments, the expansion module 1003 is specifically configured to:
[0210] For each network traffic data, perform the following operations:
[0211] Combining the target attack features corresponding to the network traffic data into a one-dimensional feature vector, and multiplying the one-dimensional feature vector by the identity matrix to obtain a first matrix;
[0212] Determining the row number and column number of each parameter in the two-dimensional matrix according to the number of target attack features corresponding to the network traffic data;
[0213] For each parameter in the two-dimensional matrix: if the row number and the column number of the parameter are the same, the value of the parameter is zero; if the row number and the column number of the parameter are different, determining, based on the first matrix, a column vector corresponding to the row number and a column vector corresponding to the column number, and determining the value of the parameter based on the column vector corresponding to the row number and the column vector corresponding to the column number;
[0214] The two-dimensional matrix is input into a generator of the CL-WGAN model to generate the plurality of reference data.
[0215] In some embodiments, the expansion module 1003 is specifically configured to:
[0216] Inputting the multiple reference data into a convolutional bidirectional long short-term memory neural network (CNN-BiLSTM), performing feature extraction on the multiple reference data to obtain feature vectors;
[0217] Performing forward and reverse feature extraction on the feature vector to obtain sequence feature vectors in two directions, and fusing the sequence feature vectors in the two directions to obtain traffic classification results of the multiple reference data;
[0218] The traffic classification result is characterized as reference data of attack network traffic as extended data; the attack network traffic is network traffic having feature information identical to any one of a plurality of preset attack features.
[0219] In some embodiments, the training module 1005 is specifically configured to:
[0220] The original data set and the extended data set are used as training data sets, input into the Bi-TCN model to be trained, and classification results of each network traffic data in the training data set are obtained;
[0221] Determining a loss function of the Bi-TCN model to be trained based on the classification result and the labeled data in the training data set; the labeled data in the training data set indicates whether each network traffic data in the training data set is attack network traffic;
[0222] Based on the loss function of the Bi-TCN model to be trained, the weight parameters and bias parameters of the Bi-TCN model to be trained are adjusted until the classification result is consistent with the labeled data in the training data set, thereby obtaining a trained Bi-TCN model.
[0223] In some embodiments, the Bi-TCN model includes an input layer, a convolutional layer, a fully connected layer, and a softmax layer, the convolutional layer includes n residual modules, each residual module includes a causal void convolution unit; the expansion coefficient of the causal void convolution unit of each residual module increases exponentially;
[0224] The device further includes a detection module, configured to:
[0225] Input the original data set into the convolution layer, perform forward and reverse feature extraction in each residual module according to the dilation coefficient of the causal hole convolution unit, obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in the two directions and input them into the next residual module until the last residual module outputs the feature vector after convolution processing;
[0226] Inputting the convolution-processed feature vector into the fully connected layer, and determining a reference feature vector based on weight parameters and bias parameters in the trained Bi-TCN model;
[0227] The reference feature vector output by the fully connected layer is input into the softmax layer for activation operation to obtain the classification results of the multiple network traffic data in the original data set.
[0228] The division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the functional modules in the embodiments of the present application may be integrated into one processor, or may exist physically separately, or two or more modules may be integrated into one module. The coupling between the modules can be achieved through some interfaces, which are usually electrical communication interfaces, but it is not ruled out that they may be mechanical interfaces or other forms of interfaces. Therefore, the modules described as separate components may or may not be physically separated, and may be located in one place or distributed to different locations of the same or different devices. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0229] After introducing the model training method and device according to an exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.
[0230] The electronic device 130 implemented according to this embodiment of the present application is described below with reference to Figure 11. The electronic device 130 shown in Figure 11 is only an example and should not limit the functions and scope of use of the embodiment of the present application.
[0231] As shown in Figure 11, electronic device 130 is a general electronic device. Components of electronic device 130 may include, but are not limited to, at least one processor 131, at least one memory 132, and a bus 133 connecting different system components (including memory 132 and processor 131).
[0232] Among them, at least one memory 132 stores a computer program that can be executed by at least one processor 131. When the computer program is executed by at least one processor 131, the at least one processor 131 can execute the steps of any model training method provided in the embodiments of the present application.
[0233] Bus 133 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a processor or local bus using any of a variety of bus architectures.
[0234] The memory 132 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1321 and / or a cache memory 1322 , and may further include a read-only memory (ROM) 1323 .
[0235] The memory 132 may also include a program / utility 1325 having a set (at least one) of program modules 1324, such program modules 1324 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0236] The electronic device 130 may also communicate with one or more external devices 134 (e.g., a keyboard, pointing device, etc.), one or more devices that enable a user to interact with the electronic device 130, and / or any device that enables the electronic device 130 to communicate with one or more other electronic devices (e.g., a router, a modem, etc.). Such communication may occur via an input / output (I / O) interface 135. Furthermore, the electronic device 130 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 136. As shown, the network adapter 136 communicates with other modules of the electronic device 130 via a bus 133. It should be understood that, although not shown, other hardware and / or software modules may be used in conjunction with the electronic device 130, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0237] In an exemplary embodiment, a storage medium is also provided, and when a computer program in the storage medium is executed by a processor of an electronic device, the electronic device can perform any of the above-mentioned model training methods. Optionally, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device.
[0238] In an exemplary embodiment, a computer program product is also provided. When the computer program product is executed by an electronic device, the electronic device can implement the steps of any model training method provided in this application.
[0239] Furthermore, the computer program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0240] In the embodiments of the present application, the program product for device discovery may be a CD-ROM and include program code, and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0241] A readable signal medium may include a data signal transmitted in baseband or as part of a carrier wave, which carries readable program code. Such a transmitted data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0242] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, radio frequency (RF), etc., or any suitable combination of the foregoing.
[0243] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, such as a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0244] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.
[0245] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0246] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0247] The present application is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0248] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0249] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0250] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0251] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application also includes these modifications and variations.
Claims
1. A model training method, characterized in that, The method includes: Obtaining an original data set and a bidirectional temporal convolutional network (Bi-TCN) model to be trained; the original data set includes multiple network traffic data; Performing feature classification on the original data set to obtain multiple attack features in the original data set; Based on multiple target attack features, performing iterative operations until the similarity between the multiple extended data in the current iteration and the multiple network traffic data in the original data set is greater than a similarity threshold, and ending the iterative operations; the iterative operations include: performing data dimension expansion on each network traffic data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain the multiple extended data; if the current iteration is the first iteration, the multiple target attack features are the multiple attack features; if the current iteration is not the first iteration, the multiple target attack features are the attack features corresponding to the multiple extended data obtained in the previous iteration; Taking the multiple extended data obtained in each iteration as the extended data corresponding to the multiple attack features to obtain an extended data set; Training the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
2. The method according to claim 1, characterized in that After obtaining the extended data set and before training the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model, the method further includes: Performing quality inspection on each network traffic extended data in the extended data set to obtain network traffic extended data that conforms to a preset quality detection rule; Taking the network traffic extended data that conforms to the preset quality detection rule as the network traffic extended data in the extended data set.
3. The method according to claim 2, wherein The network traffic extended data that conforms to the preset quality detection rule includes: Network traffic extended data that does not include illegal values; When the maximum mean difference value between the extended data set and the original data set is less than a preset difference threshold, the network traffic extended data included in the extended data set; Network traffic extended data that includes a preset multiple of attack features; the preset multiple of attack features are used for performing feature classification on the original data set.
4. The method according to claim 1, wherein The performing feature classification on the original data set to obtain multiple attack features in the original data set includes: For each network traffic data in the original data set: Performing feature extraction on the network traffic data to obtain multiple feature information corresponding to the network traffic data; Taking the feature information that is the same as any one of the preset multiple of attack features as the attack feature in the original data set.
5. The method according to claim 1, characterized in that The performing data dimension expansion on each network traffic data corresponding to the multiple target attack features to obtain multiple reference data includes: For each network traffic data, respectively performing the following operations: Forming a one-dimensional feature vector from the target attack features corresponding to the network traffic data, and multiplying the one-dimensional feature vector by an identity matrix to obtain a first matrix; Determine the row numbers and column numbers of each parameter in the two-dimensional matrix according to the number of target attack features corresponding to the network traffic data; For each parameter in the two-dimensional matrix: if the row number and column number of the parameter are the same, the value of the parameter is zero; if the row number and column number of the parameter are different, based on the first matrix, determine the column vector corresponding to the row number and the column vector corresponding to the column number, and determine the value of the parameter based on the column vector corresponding to the row number and the column vector corresponding to the column number; Input the two-dimensional matrix into the generator of the CL-WGAN model to generate the multiple reference data.
6. The method according to claim 5, wherein The traffic classification of the multiple reference data to obtain the multiple extended data includes: Input the multiple reference data into the convolutional bidirectional long short-term memory neural network CNN-BiLSTM to perform feature extraction on the multiple reference data to obtain feature vectors; Perform forward and backward feature extraction on the feature vectors to obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in two directions to obtain the traffic classification results of the multiple reference data; Characterize the traffic classification results as reference data of attack network traffic as extended data; the attack network traffic is network traffic having the same feature information as any one of a preset plurality of attack features.
7. The method according to claim 1, characterized in that, Train the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model, including: Use the original data set and the extended data set as a training data set and input it into the Bi-TCN model to be trained to obtain the classification results of each network traffic data in the training data set; Based on the classification results and the labeled data in the training data set, determine the loss function of the Bi-TCN model to be trained; the labeled data in the training data set indicates whether each network traffic data in the training data set is attack network traffic; Based on the loss function of the Bi-TCN model to be trained, adjust the weight parameters and bias parameters of the Bi-TCN model to be trained until the classification results are consistent with the labeled data in the training data set to obtain a trained Bi-TCN model.
8. The method according to any one of claims 1 to 7, characterized in that, The Bi-TCN model includes an input layer, a convolutional layer, a fully connected layer, and a softmax layer. The convolutional layer includes n residual modules, and each residual module includes a causal dilated convolutional unit; the dilation coefficients of the causal dilated convolutional units of each residual module increase exponentially; the method further includes: Input the original data set into the convolutional layer, perform forward and backward feature extraction in each residual module according to the dilation coefficient of the causal dilated convolutional unit to obtain sequence feature vectors in two directions, and fuse the sequence feature vectors in two directions and input them into the next residual module until the last residual module outputs the feature vectors after convolutional processing; Input the feature vector after convolution processing into the fully connected layer, and determine a reference feature vector based on the weight parameters and bias parameters in the trained Bi-TCN model. Input the reference feature vector output by the fully connected layer into the softmax layer for activation operation to obtain the classification results of multiple network traffic data in the original data set.
9. A model training device, characterized in that, The device includes: A first acquisition module, configured to acquire an original data set and a bidirectional temporal convolutional network (Bi-TCN) model to be trained; the original data set includes multiple network traffic data. A classification module, configured to perform feature classification on the original data set to obtain multiple attack features in the original data set. An expansion module, configured to perform iterative operations based on multiple target attack features until the similarity between the multiple extended data in the current iteration and the multiple network traffic data in the original data set is greater than a similarity threshold, and end the iterative operations; the iterative operations include: expanding the data dimensions of each network traffic data corresponding to the multiple target attack features to obtain multiple reference data; performing traffic classification on the multiple reference data to obtain the multiple extended data; if the current iteration is the first iteration, the multiple target attack features are the multiple attack features; if the current iteration is not the first iteration, the multiple target attack features are the attack features corresponding to the multiple extended data obtained in the previous iteration. An extended data set determination module, configured to use the multiple extended data obtained in each iteration as the extended data corresponding to the multiple attack features to obtain an extended data set. A training module, configured to train the Bi-TCN model to be trained based on the original data set and the extended data set to obtain a trained Bi-TCN model.
10. An electronic device, characterized in that, It includes: At least one processor, and a memory communicatively connected to the at least one processor, wherein: The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.
11. A storage medium, characterized in that, When the computer program in the storage medium is executed by the processor of the electronic device, the electronic device can execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Hybrid network traffic attack prediction method based on TMB
CN115694985A
Model training method and device, equipment and medium
CN117811801A
Privacy-sensitive neural network training using data augmentation
US20230351042A1
Conversation intent recognition model training method, apparatus, computer device, and medium
WO2022141864A1