Air-Ground-Space Network Intrusion Detection Method Based on Multimodal Conditional Countermeasure Domain Adaptation
By introducing feature fusion and multilinear mapping in the adversarial domain adaptive method, combined with entropy conditions, the problem of difficult multimodal data distribution processing in the prior art is solved, effective alignment of inter-domain and intra-domain features is achieved, and the performance of intrusion detection is improved.
Patent Information
- Application Number
- CN202211009783.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-08-22
AI Technical Summary
When handling multimodal data distribution, the existing adversarial domain adaptive methods are difficult to capture the multimodal characteristics of the data structure, resulting in inter-domain alignment errors and poor migration effect.
The multimodal conditional adversarial domain adaptation method based on feature fusion is adopted, and the source domain, target domain and domain invariant features are extracted through the three feature extraction networks of Fs, Fm and Ft, and the adversarial network is introduced through multilinear mapping and entropy conditions to achieve inter-domain feature alignment and multimodal structure capture.
While achieving the alignment of edge distribution between domains, the multimodal structure of intra-domain categories is captured, which enhances the intra-domain category resolution capability and improves the performance of intrusion detection tasks.
Smart Images

Figure CN115412324B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed air-ground-space network, and in particular to an air-ground-space network intrusion detection method based on multi-modal condition confrontation field adaptation. Background Art
[0002] The current adversarial domain adaptation methods have good overall effects on tasks such as classification and segmentation, but these adversarial domain adaptation methods may still be constrained by two bottlenecks. First, to extract domain-invariant features using adversarial training, the source domain and the target domain need to be mapped to a high-dimensional space, in which the domain-invariant characteristics of the two domains exist. However, most adversarial domain adaptation methods can only align the source domain and target domain data in terms of overall distribution, while ignoring the multimodal distribution characteristics of the data within the domain. This is because different categories of data in each domain also correspond to different features, and these different features appear as different distributions in the mapped high-dimensional space. Therefore, even data of the same category from different domains may not be consistent in location and form after being mapped to the high-dimensional space, that is, the overall data will present more than one multimodal structure. When the data distribution contains a complex multimodal structure, the adversarial adaptation method may not be able to capture this multimodal structure. It only optimizes the marginal distribution and ignores the structural distribution characteristics of the data itself, which may lead to incorrect alignment, such as Figure 1 As shown, the migration effect becomes worse.
[0003] This drawback comes from the balance problem of adversarial learning, because even if the discriminator can no longer identify fake samples, it does not mean that the data distributions of the two domains are sufficiently similar. This drawback cannot be simply solved by aligning the distributions of features and classes using a separate domain discriminator, because the multimodal structure can only be fully captured by the mutual covariance dependency between features and classes. Secondly, when the prediction information of the classifier is uncertain, using the prediction information as a condition for the domain discriminator is potentially risky. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a multimodal conditional adversarial domain adaptation method based on feature fusion, which provides an optimal detection model suitable for the domain environment for each network domain through domain adaptive training of the model. The purpose of the present invention is to propose a multimodal conditional adversarial domain adaptation algorithm based on feature fusion in view of the existing problems.
[0005] The air-ground-space network intrusion detection method based on multi-modal conditional confrontation domain adaptation includes the following steps:
[0006] Step 1: Construct a multimodal conditional adversarial domain adaptation network model based on feature fusion.
[0007] Model using F s 、Fm 、F t Three feature extraction networks, all of which use the ShuffleNet_Lite_ECA network. s 、F t are used to extract unique features from the source and target domains, respectively, m Used to extract domain-invariant features between domains. s 、F m The extracted features are fused and passed to the classifier C for supervised training; the classifier trained with source domain data learns the information of the target domain, and then F m The extracted features and the classification labels of the source domain data are processed according to the multilinear mapping method and then input into the discriminator D of the adversarial network for domain discrimination.
[0008] Features and labels are merged, that is, f and g represent features and classification labels respectively. According to a multilinear mapping method The cross-covariance information is extracted through multilinear mapping, which is defined as the outer product of multiple random vectors. Assume that the linear mapping φ(x) = x and the one-hot label variable y has C categories, and the mean mapping The means of x and y are calculated independently, while the multilinear mapping Compute the mean of the conditional distribution P(x|y) for each category.
[0009] The entropy condition is introduced to measure the uncertainty of the sample prediction category. A smaller weight is given to samples with larger entropy, and a larger weight is given to samples with smaller entropy, so as to reduce the negative impact of samples with poor prediction effect on adversarial training. The calculation formula of entropy H(·) is:
[0010]
[0011] N c Indicates the number of categories, Indicates that the sample belongs to the nth c The probability of the class. Therefore, the entropy weight ω(H(·)) is obtained according to the entropy value:
[0012]
[0013] The discriminator D determines whether the received data sample is from the source domain or the target domain. This part adopts the conditional domain confrontation idea. Since the feature data received by the discriminator D incorporates label information; the target domain data is also passed through F m 、F tThe feature extraction network extracts shared features and unique features; classifier C is used to classify and predict the target domain data. Classifier C is first trained for one epoch under supervision, and then adversarial training is started. The target domain data has classification labels. After the outer product calculation, the conditional discriminator is input for adversarial training.
[0014] The feature alignment loss penalty term is introduced into the objective function for feature alignment, and the CORAL algorithm is used to calculate the difference between features. The CORAL algorithm aligns the second-order statistics of the feature distribution of the source domain and the target domain, thereby narrowing the difference in feature distribution between the two domains. The CORAL difference calculation formula is:
[0015]
[0016] d is the feature dimension, C S and C T They are the feature covariance matrices of the source domain and the target domain respectively:
[0017]
[0018]
[0019] 1 is a column variable whose elements are all 1, M S N is the number of data contained in each batch S The characteristics of source domain data samples, M T N is the number of data contained in each batch T The characteristics of the target domain data samples.
[0020] This model adopts a three-stage training method, namely pre-training stage, training stage and testing stage. The pre-training stage model network only contains F s and F m Two feature mapping networks and classifier C and discriminator D. When pre-training converges, freeze the feature extraction network F. m The parameters of the discriminator are frozen at the beginning of the formal training, so that the classifier C can be supervised for one epoch, and then the adversarial training is started. After the formal training model converges, the F m 、F t , C, and D parameters are verified and evaluated on the test set.
[0021] Step 2: Construct the objective function.
[0022] Two feature extraction networks, a domain discriminator and a classifier of the target domain are learned to perform intrusion detection tasks on the target domain.
[0023] Defining source domain data represents the i-th sample in the source domain data, represents the label of the i-th sample in the source domain, N s Represents the number of samples in the source domain, and the distribution of the source domain data is recorded as P s Similarly, define the target domain data represents the i-th sample in the target domain data. Unlike the source domain, the target domain has no label. N t Represents the number of samples in the target domain, and the distribution of the target domain data is recorded as P t , and define is a sample set of two domains, and d is defined at the same time i is the domain label of the i-th sample, d i =0 represents the source domain, d i =1 represents the target domain.
[0024] In the F-MCADA architecture, the objective function is divided into three parts.
[0025] 1) Classification loss of supervised training generated by classifier C
[0026] The unique features extracted from the source domain are combined with F m The extracted confusion features are fused and passed to C for supervised learning. The classifier C performs supervised training on the labeled source domain data and uses cross entropy loss for optimization:
[0027]
[0028] Representative feature generation network F s Parameters, Representative feature confusion network F m The parameter θ C represents the parameters of the classifier C, L c represents the cross entropy loss.
[0029] 2) Conditional domain classification loss
[0030] The overall optimization goal of the conditional domain adversarial training loss is:
[0031]
[0032] L adv-s , L adv-t The source domain data and the target domain data are respectively in the feature extraction network F m The adversarial error term on the discriminator D:
[0033]
[0034]
[0035] Indicates that the source domain data passes through F s and F m The predicted label of the fused features on the classifier C is Indicates that the target domain data passes through F t and F m The pseudo prediction label of the fused feature on the classifier C; T(·) is the multilinear mapping conditional fusion strategy between the feature and the prediction label, f m Indicates F m The extracted features, c is the classification label, and d c Respectively represent the vector f m and the dimension of c, is a multilinear mapping such that the conditional domain discriminator captures f m and c’s multimodal information and joint distribution; H(·) is the entropy used to measure the uncertainty of sample prediction category; ω(H(·)) represents the entropy weight calculated according to the entropy value.
[0036] 3) Inter-domain CORAL alignment loss
[0037] According to formula (3), the inter-domain CORAL feature alignment loss is:
[0038]
[0039] d is the feature dimension, C S and C T are the feature covariance matrices of the source domain and the target domain, respectively.
[0040] The final training objective function of F-MCADA is:
[0041]
[0042] α and β are loss weight coefficients.
[0043] The technical effects of the present invention are as follows:
[0044] 1. First, the source domain data and the target domain data are mapped to the domain-invariant space respectively by using the domain adversarial method to align the marginal distributions of the source domain and the target domain.
[0045] 2. Use the labeled data in the source domain to train the source domain classifier.
[0046] 3. In order to solve the problem that the traditional adversarial network only uses one common feature extraction network, which leads to the loss of unique features of the source domain and the target domain, thereby causing blurred class boundaries within the domain, three independent feature extraction networks are used to extract source domain, target domain and domain-invariant features respectively, so as to achieve the purpose of extracting domain-invariant features while retaining unique information within the domain.
[0047] 4. In order to solve the problem that inter-domain adaptation will ignore the multimodal information within the domain, conditional domain adaptation is introduced. The prediction results of the classifier are used to guide the identification work of the discriminator, making the features of each category in the domain more distinguishable and enhancing the ability to distinguish categories within the domain.
[0048] 5. In order to solve the problem that the source domain and the target domain characteristics may be too different due to adversarial training, which may lead to poor alignment between the domains, an alignment penalty term is introduced to achieve feature alignment by minimizing the difference between the source domain characteristics and the target domain characteristics. Ultimately, the marginal distribution between domains is aligned, and the multimodal structure of the categories within the domain can also be correctly aligned.
[0049] 6. Experiments are conducted to prove the excellent performance of the F-MCADA model in intrusion detection tasks, and a t-SNE visualization experiment is designed to intuitively display the feature distribution status before and after F-MCADA domain adaptation training. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of adaptive inter-category error alignment in the prior art;
[0051] Figure 2 This is a flow chart of a multi-modal conditional adversarial domain adaptation method based on feature fusion in the present invention;
[0052] Figure 3 Pre-training network structure for the present invention;
[0053] Figure 4 Formal training of the network structure for the present invention;
[0054] Figure 5 Testing network structure for the model of the present invention;
[0055] Figure 6 This is a loss variation diagram of the field discriminator of the present invention;
[0056] Figure 7 This is a graph showing the accuracy change of the field discriminator of the present invention;
[0057] Figure 8 This is the loss change diagram of the classifier C of the present invention;
[0058] Fig. 9 This is a graph showing the change in the penalty term loss for feature alignment of the present invention;
[0059] Fig.10 is the change in accuracy of the present invention on the test set;
[0060] Fig.11 Comparison of the scores of various evaluation indicators of different models of the present invention;
[0061] FIG12( a ) is a comparison of sample feature distribution before and after field adaptive training of the present invention;
[0062] FIG12( b ) is a comparison of sample feature distribution before and after field adaptive training of the present invention. DETAILED DESCRIPTION
[0063] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific implementation. According to an embodiment of the present invention, a multi-modal conditional adversarial domain adaptation method based on feature fusion is provided. Figure 2 The flowchart of the proposed method for air-ground-space network intrusion detection based on multi-modal conditional adversarial domain adaptation includes:
[0064] Step 1: Construct a multimodal conditional adversarial domain adaptation network model based on feature fusion.
[0065] In order to improve the general performance of the model in the target domain and make full use of the unique domain features of each domain, a multi-modal conditional adversarial domain adaptation network model based on feature fusion (Multi-modal Conditional Adversarial Domain Adaptation Network Based on Feature Fusion, F-MCADA) is proposed. The model can simultaneously extract features with domain-invariant characteristics and individual features of the domain, and fully explore the implicit information carried by the domain data; the model label classifier incorporates the information of the target domain, which enhances its generalization in the target domain data set; the extracted features and the classification labels are calculated as the outer product as the input of the discriminator in the adversarial network, and the features with inter-class discrimination are fully learned, thereby achieving the purpose of multi-modal domain adaptation.
[0066] In order to simultaneously extract shared features between fields and features specific to each field, this model uses F s 、F m 、F t Three feature extraction networks, all of which use the ShuffleNet_Lite_ECA network. s 、F t are used to extract unique features from the source and target domains, respectively, m Used to extract domain-invariant features between domains. s、F m The extracted features are fused and passed to the classifier C for supervised training. By using the supervised information, the classification accuracy of the source domain is improved while minimizing the classification risk of the target domain. m contains the target domain information and domain-invariant features, so the classifier trained with source domain data also learns the target domain information, and then F m The extracted features and the classification labels of the source domain data are processed according to the multilinear mapping method and then input into the discriminator D of the adversarial network for domain discrimination.
[0067] Current adversarial domain adaptation methods have the following problems: (1) They only consider individual alignment features without considering alignment labels, and often fail to make full use of effective information; (2) When the data distribution reflects a complex multimodal structure, methods that simply consider feature alignment cannot capture this multimodal structure. Even if the model training converges and the discriminator is completely confused and unable to distinguish the source of the samples, there is no guarantee that the same categories in the source domain and the target domain are sufficiently similar; (3) General adversarial domain adaptation methods treat samples as equally important. At this time, samples with uncertain label predictions that are difficult to transfer may have a negative effect on the adversarial network.
[0068] Problems (1) and (2) can be solved by aligning the joint distribution of features and categories. The conditional generative adversarial network (CGAN) reveals that different distributions can obtain auxiliary information on relevant information so that the generator and discriminator can be well matched, for example, by associating features with labels. The classifier prediction results carry potential information that can reveal the multimodal characteristics of the data distribution. If this information can be used to guide the training of the generator and discriminator, then multimodal information can be better extracted. However, simply merging features and labels, that is, f and g represent features and classification labels respectively, which makes f and g still independent of each other, resulting in the inability to capture the multiplicative interaction between features and predicted classification results. Therefore, the multimodal information contained in the classifier prediction results cannot be fully extracted to match the multimodal distribution of complex domains. According to a multilinear mapping method This problem can be solved by extracting the cross-covariance information through multilinear mapping, which is defined as the outer product of multiple random vectors. Assuming the linear mapping φ(x) = x and the one-hot label variable y with C categories, a simple mean mapping The means of x and y are calculated independently, while the multilinear mapping The mean of the conditional distribution P(x|y) for each category is calculated. Compare Advantageously, multilinear mapping simulates the multiplicative interactions between different variables, and the biggest advantage of multilinear mapping is that it can fully capture the multimodal structure behind complex data distribution.
[0069] For problem (3), we introduce entropy conditions to measure the uncertainty of sample prediction categories. We give smaller weights to samples with larger entropy and larger weights to samples with smaller entropy.
[0070] Reduce the negative impact of samples with poor prediction effects on adversarial training. The calculation formula of entropy H(·) is:
[0071]
[0072] N c Indicates the number of categories, Indicates that the sample belongs to the nth c The probability of the class. Therefore, the entropy weight ω(H(·)) is obtained according to the entropy value:
[0073]
[0074] The discriminator D determines whether the received data sample is from the source domain or the target domain. This part adopts the idea of conditional domain confrontation. Since the feature data received by the discriminator D incorporates label information, it increases the distinguishability between different categories between domains and solves the problem of mixed and indistinguishable distribution of categories between domains. Similar to the processing flow of source domain data, the target domain data is also processed by F m 、F t The feature extraction network extracts shared features and unique features. Unlike the source domain data, the source domain data has its own real classification labels, while the target domain is unlabeled data. Therefore, this method adopts a semi-supervised pseudo-label generation method and uses classifier C to classify and predict the target domain data. In order to prevent the initial poor classification ability of classifier C for the target domain data, resulting in a large error in the pseudo-label generated by classifier C for the target domain data, classifier C is first trained for one epoch of supervised training before starting adversarial training. Since then, the target domain data also has classification labels, which is similar to After the outer product calculation, the conditional discriminator is input for adversarial training.
[0075] In addition, in order to prevent the unique features of the source domain and the target domain from being extracted, the features extracted by the network s and f m As the training progresses, the source domain tends to be far away from the target domain. Therefore, this paper introduces a feature alignment loss penalty term in the objective function to perform feature alignment and uses the CORAL algorithm to calculate the difference between features. The CORAL algorithm aligns the second-order statistics of the feature distribution of the source domain and the target domain, thereby narrowing the feature distribution difference between the two domains. The CORAL difference calculation formula is:
[0076]
[0077] d is the feature dimension, C S and C T They are the feature covariance matrices of the source domain and the target domain respectively:
[0078]
[0079]
[0080] 1 is a column variable whose elements are all 1, M S N is the number of data contained in each batch S The characteristics of source domain data samples, M T N is the number of data contained in each batch T The characteristics of the target domain data samples.
[0081] In order to make the model converge better and faster, this model adopts a three-stage training method, namely pre-training stage, training stage and testing stage. The model network in the pre-training stage only contains F s and F m Two feature mapping networks and classifier C and discriminator D, such as Figure 3 As shown. When the pre-training converges, freeze the feature extraction network F m The parameters of the formal training are used. The formally trained model network is as follows Figure 4 As shown in Figure 1, the discriminator parameters are frozen at the beginning, so that the classifier C can be supervised for one epoch, and then the adversarial training is started. After the formal training model converges, the F m 、F t The parameters of C and D are verified and evaluated on the test set. The verification process is as follows: Figure 5 shown.
[0082] Step 2: Construct the objective function.
[0083] F-MCADA belongs to an unsupervised conditional domain adaptation network. There are two domains: source domain and target domain. The data in the source domain is labeled, and the data in the target domain is unlabeled. There are differences in the data distribution in the two domains. The goal of the algorithm proposed in this chapter is to learn two feature extraction networks, a domain discriminator, and a classifier in the target domain to perform intrusion detection tasks on the target domain.
[0084] Defining source domain data represents the i-th sample in the source domain data, represents the label of the i-th sample in the source domain, N s Represents the number of samples in the source domain, and the distribution of the source domain data is recorded as Ps Similarly, define the target domain data represents the i-th sample in the target domain data. Unlike the source domain, the target domain has no label. N t Represents the number of samples in the target domain, and the distribution of the target domain data is recorded as P t , and define is a sample set of two domains, and d is defined at the same time i is the domain label of the i-th sample, d i =0 represents the source domain, d i =1 represents the target domain.
[0085] In the F-MCADA architecture, the objective function can be mainly divided into three parts.
[0086] 1) Classification loss of supervised training generated by classifier C
[0087] Because the source domain data is labeled, it is necessary to maximize the source domain supervision information. m The extracted confusion features are fused and passed to C for supervised learning. Because the confusion features contain feature information of the target domain, they can further promote feature alignment between the source domain and the target domain. In addition, the label information of the source domain can also allow the model to learn features that distinguish between classes, which is beneficial for the classifier to distinguish between various samples in the domain. Classifier C performs supervised training on labeled source domain data and uses cross entropy loss for optimization:
[0088]
[0089] Representative feature generation network F s Parameters, Representative feature confusion network F m The parameter θ C represents the parameters of the classifier C, L c represents the cross entropy loss.
[0090] 2) Conditional domain classification loss
[0091] In conditional adversarial training, data from the source domain and the target domain are used to fuse classification information. The overall optimization goal of conditional domain adversarial training loss is:
[0092]
[0093] L adv-s , L adv-t The source domain data and the target domain data are respectively in the feature extraction network F m The adversarial error term on the discriminator D:
[0094]
[0095]
[0096] Indicates that the source domain data passes through F s and F m The predicted label of the fused features on the classifier C is Indicates that the target domain data passes through F t and F m The pseudo prediction label of the fused feature on the classifier C; T(·) is the multilinear mapping conditional fusion strategy between the feature and the prediction label, f m Indicates F m The extracted features, c is the classification label, and d c Respectively represent the vector f m and the dimension of c, is a multilinear mapping such that the conditional domain discriminator captures f m and c’s multimodal information and joint distribution; H(·) is the entropy used to measure the uncertainty of sample prediction category; ω(H(·)) represents the entropy weight calculated according to the entropy value.
[0097] 3) Inter-domain CORAL alignment loss
[0098] According to formula (3), the inter-domain CORAL feature alignment loss is:
[0099]
[0100] d is the feature dimension, C S and C T are the feature covariance matrices of the source domain and the target domain, respectively.
[0101] In summary, the final training objective function of F-MCADA is:
[0102]
[0103] α and β are loss weight coefficients.
[0104] F-MCADA algorithm steps, F-MCADA algorithm pre-training steps:
[0105]
[0106]
[0107] Formal training steps of F-MCADA algorithm:
[0108]
[0109]
[0110] Step 3: Experimentally verify the excellent performance of the model.
[0111] 1) Dataset settings
[0112] The dataset used for training the F-MCADA network model is based on the two public datasets CIC-IDS-2017 and CSE-CIC-IDS-2018, adding location information such as longitude and latitude, and altitude. 76 statistical features common to CIC-IDS-2017 and CSE-CIC-IDS-2018 are selected to simulate the satellite intrusion traffic dataset, named SAT-IDS-1 and SAT-IDS-2. SAT-IDS-1 is used as the source domain dataset and SAT-IDS-2 is used as the target domain dataset. The network topology environment and network attack methods of these two datasets are very different when collecting data, so they are very suitable for the complex characteristics of the air-space-ground network environment. In order to facilitate the experiment, the 76-dimensional feature data is padded to 81 dimensions with 0 and converted into a 9*9 feature matrix to better adapt to the two-dimensional input format of the convolutional network model.
[0113] 2) Specific network settings for each module of F-MCADA
[0114] (1) Feature mapping network F: Three independent feature mapping networks are used in the F-MCADA network, all of which adopt the proposed SuffleNet_Lite_ECA. The specific structure has been introduced in detail and will not be repeated here.
[0115] (2) Classifier network C: The classifier part adopts a two-layer fully connected layer structure, which is 1024→100→2. Batch Normalization is added after the previous layer, ReLU activation is used, and Pytorch's CrossEntropyLoss is used to calculate the classification loss.
[0116] (3) Domain discrimination network D: The domain discrimination network D adopts a three-layer fully connected layer structure, which is 1024→512→100→1. ReLU activation is used after the first two layers, and dropout (0.5) is used to prevent overfitting. Since Pytorch's BCELoss is used to calculate the domain classification loss, the Sigmoid function activation needs to be used before calculating the loss.
[0117] 3) Hardware Configuration
[0118] The hardware environment configuration for model training is GPU: TITAN Xp*1, video memory: 12GB, and memory: 16GB.
[0119] 4) Hyperparameter setting
[0120] The number of model training epochs is set to 20, the batch size batch_size is 512, the learning rate is μ=1e-3, the learning rate decay coefficient is 0.5, and it decays every 4 epochs. The SGD optimizer is used to update the model parameters, and the loss weight coefficients α=1 and β=5.
[0121] Experimental results and analysis:
[0122] In the migration experiment from SAT-IDS-1 to SAT-IDS-1, the data set is first preprocessed, then the model is trained according to the description of the above two algorithms, and finally the test set is tested according to Figure 5 The network structure of the model is experimentally evaluated, and the experimental results are obtained and analyzed based on the experimental results. The evaluation indicators of this experiment are precision, accuracy, recall, AUC and F1-score.
[0123] The overall optimization objective of the multimodal conditional adversarial domain adaptation neural network consists of domain discrimination loss, classification loss, and inter-domain individual feature alignment penalty loss. During the adversarial training process, Figure 6 is the loss change of the domain discriminator, Figure 7 And the accuracy change curve.
[0124] As can be seen from the figure, during the training of the multimodal conditional adversarial domain adaptation network model, the domain discrimination loss initially decreases rapidly and then oscillates and rises. In contrast, the accuracy of the domain discriminator initially increases rapidly and then oscillates and decreases. The decrease in the discriminator loss indicates that the common feature extraction network F m The common features of the two fields have not been extracted yet, and the feature extraction network F m The extracted features contain more differentiated features between domains, so the discriminator has a high accuracy. Then the conditional discriminator loss oscillates and rises at a relatively fast rate and gradually stabilizes, and finally shows a small oscillation state, which remains at around 0.7. This shows that as the model training progresses, the domain-invariant feature extraction network has successfully extracted domain-invariant features of the two domains, which interferes with the judgment of the conditional domain discriminator. Therefore, ACC shows a large oscillation and then stabilizes, and finally oscillates in the range of around 0.5. The accuracy of 0.5 shows that the discriminator is no longer able to identify the domain label and has reached the probability level of random selection. The curve oscillation during the training process reflects the adversarial properties of the network. In the end, the conditional domain discriminator loss and ACC are both stable at a certain level, indicating that the adversarial relationship has reached a relative balance.
[0125] Figure 8is the loss change curve of classifier C. It can be seen from the figure that the loss of source domain classifier C decreases rapidly with the progress of training. Due to the adversarial property of adversarial learning, it finally presents a stable oscillating form, which means that in F m While successfully extracting domain-invariant features between two domains, the feature extraction network F s Feature information from the source domain that is helpful for intra-domain classification is also extracted, which is beneficial to the classification work of the classifier. Fig. 9 F s 、F t The alignment loss curve between the extracted individual features of the two domains shows the difference between the features. As the alignment loss gradually decreases and tends to a stable oscillation state, it means that the features extracted by the source domain feature extraction network and the target domain feature extraction network are very different in the early stage of training. With the constraint of the inter-domain feature alignment loss penalty, the distance between the two features gradually decreases. The domain invariant features and the unique features of the source domain and the target domain are retained respectively, and these features have achieved good classification results on the classifier.
[0126] Fig.10 This is the detection accuracy curve of the target domain test set in the model prediction stage. It can be seen from the figure that the model has achieved a high accuracy in 2 epochs. With the increase in the number of training times, the accuracy gradually increases until it stabilizes at around 0.76.
[0127] In order to demonstrate the superiority of the multimodal conditional adversarial domain adaptation intrusion detection model proposed in this chapter, Table 1 shows the binary classification comparison experimental results with several other classic migration algorithms.
[0128] Table 1 Comparison of several models on various evaluation indicators
[0129]
[0130] For intuitive comparison, convert the table into a histogram, such as Fig.11 shown.
[0131] From the comparison chart, it can be seen that the F-MCADA proposed in this chapter performs very well under various evaluation indicators. Due to the large differences between the experimental data sets, the feature-based transfer learning algorithm TCA cannot mine the deep common features between the two domains, so the scores of each item are very low and the transfer effect is poor; the AUC of the DANN model is only 51.33%, close to 50%, indicating that DANN has almost no ability to distinguish traffic data and is almost randomly selected. The Recall and AUC scores of the NSP-GAN model are higher than those of the F-MCADA model, but the ACC and F1 scores are much lower than those of the F-MCADA model, indicating that although the NSP-GAN model has a strong ability to distinguish, the judgment error rate is extremely high, resulting in a low final accuracy rate.
[0132] As a trade-off indicator between Presion and Recall, F1 can more accurately reflect the comprehensive performance of the model. In terms of the comprehensive evaluation indicator F1, F-MCADA's performance far exceeds that of other methods, and it is 34.2% higher than the second-ranked CDAN, proving the excellent performance of the F-MCADA model in intrusion detection tasks.
[0133] t-SNE visualization analysis:
[0134] In order to more intuitively show the feature distribution status before and after F-MCADA domain adaptation training, this section designs a t-SNE visualization experiment. The data for visualization comes from the test set after F m Feature extraction network and F tAfter a series of operations such as dimensionality reduction, the fusion features extracted by the feature extraction network are visualized using t-SNE. Figures 12(a) and 12(b) show the visualization results, with the perplexity value set to 20. In the figure, No. 1 and No. 2 represent normal traffic samples and abnormal traffic samples in the source domain, respectively, and No. 3 and No. 4 represent normal traffic samples and abnormal traffic samples in the target domain, respectively. According to the results, the visualization results without domain adaptation training show that the data distribution of the source domain and the target domain is chaotic, as shown in Figure 12 (a). There is no obvious distribution pattern for the distribution of normal traffic samples and abnormal traffic samples, and they are intertwined with each other. After the F-MCADA algorithm in this chapter, as shown in Figure 12 (b), it can be found that the results after domain adaptation have been improved. The normal traffic samples and abnormal traffic samples of the source domain and the target domain are clustered together, and there is a relatively obvious boundary between the normal traffic samples and the abnormal traffic samples. This shows that the F-MCADA algorithm has achieved relatively good results. While aligning the data distribution of the two domains, each category has a more obvious inter-class distinction. This also shows that while obtaining the domain invariant characteristics, the distinction between the categories within the domain is also retained, which is helpful for the classification task of the classifier. However, due to the large difference in data distribution between the SAT-IDS-1 and SAT-IDS-2 datasets, from the visualization results, there are still some samples in a confused state after adaptive training and they are not effectively distinguished. But overall, the proposed F-MCADA model has effective domain adaptation capabilities.
Claims
1. An air-ground-space network intrusion detection method based on multimodal conditional confrontation domain adaptation, characterized in that: The following steps are involved: Step 1: construct a multimodal conditional adversarial domain adaptation network model based on feature fusion; Model using F s 、F m 、F t Three feature extraction networks, all of which use ShuffleNet_Lite_ECA network; F s 、F t They are used to extract the unique features of the source domain and the target domain respectively, m Used to extract domain-invariant features between domains; F s 、F m The extracted features are fused and passed to the classifier C for supervised training; the classifier trained with source domain data learns the information of the target domain, and then F m The extracted features and the classification labels of the source domain data are processed according to the multilinear mapping method and then input into the discriminator D of the adversarial network for domain discrimination; In step 1, a feature alignment loss penalty term is introduced into the objective function to perform feature alignment, and the difference between features is calculated using the CORAL algorithm to align the second-order statistics of the feature distribution of the source domain and the target domain, thereby narrowing the feature distribution difference between the two domains. The CORAL difference calculation formula is: d is the feature dimension, C S and C T They are the feature covariance matrices of the source domain and the target domain respectively: 1 is a column variable whose elements are all 1, M S N is the number of data contained in each batch S The characteristics of source domain data samples, M T N is the number of data contained in each batch T The characteristics of the target domain data samples; Step 2, construct the objective function; Learning two feature extraction networks, a domain discriminator and a classifier for the target domain to perform intrusion detection tasks on the target domain; Step 2: Defining source domain data represents the i-th sample in the source domain data, represents the label of the i-th sample in the source domain, N s Represents the number of samples in the source domain, and the distribution of the source domain data is recorded as P s ; Similarly, define the target domain data represents the i-th sample in the target domain data. Unlike the source domain, the target domain has no label. N t Represents the number of samples in the target domain, and the distribution of the target domain data is recorded as P t , and define is a sample set of two domains, and d is defined at the same time i is the domain label of the i-th sample, d i =0 represents the source domain, d i =1 represents the target domain; The objective function in step 2 is divided into three parts; 1) Classification loss of supervised training generated by classifier C The unique features extracted from the source domain are combined with F m The extracted confusion features are fused and passed to C for supervised learning; classifier C performs supervised training on the labeled source domain data and uses cross entropy loss for optimization: Representative feature generation network F s Parameters, Representative feature confusion network F m The parameter θ C represents the parameters of the classifier C, L c represents the cross entropy loss; 2) Conditional domain classification loss The overall optimization goal of the conditional domain adversarial training loss is: L adv-s , L adv-t The source domain data and the target domain data are respectively in the feature extraction network F m The adversarial error term on the discriminator D: Indicates that the source domain data passes through F s and F m The predicted label of the fused features on the classifier C is Indicates that the target domain data passes through F t and F m The pseudo prediction label of the fused feature on the classifier C; T(·) is the multilinear mapping conditional fusion strategy between the feature and the prediction label, f m Indicates F m The extracted features, c is the classification label, and d c Respectively represent the vector f m and the dimension of c, is a multilinear mapping such that the conditional domain discriminator captures f m and c’s multimodal information and joint distribution; H(·) is the entropy used to measure the uncertainty of sample prediction category; ω(H(·)) represents the entropy weight calculated according to the entropy value; 3) Inter-domain CORAL alignment loss According to formula (3), the inter-domain CORAL feature alignment loss is: d is the feature dimension, C S and C T are the feature covariance matrices of the source domain and the target domain respectively; The final training objective function of F-MCADA is: α and β are loss weight coefficients.
2. The air-ground-space network intrusion detection method based on multimodal conditional confrontation domain adaptation according to claim 1 is characterized in that: In step 1, features and labels are also merged, that is, f⊕g, where f and g represent features and classification labels respectively. According to a multilinear mapping method The cross-covariance information is extracted through multilinear mapping, which is defined as the outer product of multiple random vectors; assuming that the linear mapping φ(x) = x and the one-hot label variable y has C categories, the mean mapping E xy [x⊕y]=E x [x]⊕E y [y] will calculate the mean of x and y independently, while the multilinear mapping Compute the mean of the conditional distribution P(x|y) for each category.
3. The air-ground-space network intrusion detection method based on multimodal conditional confrontation domain adaptation according to claim 1 is characterized in that: In step 1, the entropy condition is also introduced to measure the uncertainty of the sample prediction category. A smaller weight is given to samples with larger entropy, and a larger weight is given to samples with smaller entropy, so as to reduce the negative impact of samples with poor prediction effect on adversarial training. The calculation formula of entropy H(·) is: N c Indicates the number of categories, Indicates that the sample belongs to the nth c The probability of the class; thus, the entropy weight ω(H(·)) is obtained according to the entropy value: The discriminator D determines whether the received data sample is from the source domain or the target domain. This part adopts the conditional domain confrontation idea. Since the feature data received by the discriminator D incorporates label information; the target domain data is also passed through F m 、F t The feature extraction network extracts shared features and unique features; classifier C is used to classify and predict the target domain data. Classifier C is first trained for one epoch under supervision, and then adversarial training is started; the target domain data has classification labels, which are similar to f m After the outer product calculation, the conditional discriminator is input for adversarial training.
4. The air-ground-space network intrusion detection method based on multimodal conditional confrontation domain adaptation according to claim 1 is characterized in that: In step 1, the model is trained in three stages, namely pre-training stage, training stage and testing stage. In the pre-training stage, the model network only contains F s and F m Two feature mapping networks and classifier C and discriminator D. When pre-training converges, freeze the feature extraction network F. m Parameters of the discriminator are used for formal training. At the beginning of formal training, the discriminator parameters are frozen, so that the classifier C can be supervised for one epoch, and then adversarial training is started. After the formal training model converges, the F m 、F t , C, and D parameters are verified and evaluated on the test set.
Citation Information
Patent Citations
Cross-user human body behavior recognition method based on antagonism domain adaptation strategy
CN113705339A
Small sample underwater target identification method based on multi-platform auditory perception feature deep transfer learning
CN114202056A