Intrusion detection model training method and intrusion detection method
By adopting a semi-supervised learning model of pseudo-label and interpolation technology in network intrusion detection, combined with consistency regularization algorithm and data enhancement technology, the problem of network intrusion detection under very small amounts of labeled data is solved, efficient intrusion detection and classification performance is achieved, and labeling costs are reduced.
Patent Information
- Application Number
- CN202510414009.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is difficult to effectively utilize very small amounts of labeled data in network intrusion detection, especially in the face of encrypted traffic and new attacks, traditional methods cannot achieve effective network intrusion detection under extremely limited labeled data.
A semi-supervised learning model combining pseudo-label and interpolation technology is adopted to train with a very small amount of labeled data and a large amount of unlabeled data. The robustness and generalization capabilities of the intrusion detection model are improved by using a consistent regularization algorithm and data augmentation technology.
Under the condition of very small amount of labeled data, the performance of the intrusion detection model is improved, so that it can achieve the same detection level as using a large amount of labeled data for model supervision training in very small amounts of labeled data, reducing the time and labor cost required for labeling.
Smart Images

Figure CN119922022A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and network intrusion technology, and in particular to an intrusion detection model training method and an intrusion detection method. Background Art
[0002] Network intrusion detection system (IDS) is a key line of defense for network security, and its performance is crucial to protecting network systems from attacks. With the development of deep learning technology, the application of this technology in IDS has shown great potential. However, the training of deep learning models usually relies on a large amount of labeled data, which faces two major challenges in reality. First, it is difficult to collect labeled network traffic data. On the one hand, network attacks are hidden and diverse, and the scarcity of network security incidents makes it difficult to obtain a large amount of labeled network traffic data; on the other hand, the problem of data security protection is becoming increasingly severe, making it difficult to collect network traffic data containing personal sensitive information. Second, traffic data labeling is arduous and time-consuming, and is highly dependent on expert experience, which greatly limits the application of deep learning algorithms in IDS.
[0003] To overcome these challenges, semi-supervised learning (SSL) methods have become an effective solution that can use unlabeled data to assist model training. However, existing SSL methods usually still require a relatively large number of labeled samples, at least hundreds of them. Especially for new attacks, such as 0-day and n-day vulnerabilities, it is difficult to obtain a sufficient number of labeled training samples. Therefore, how to effectively perform network intrusion detection when labeled data is extremely limited (for example, only 3 or 5 labeled samples per class) has become a problem that needs to be solved urgently.
[0004] Another challenge of IDS is that the use of network traffic data encryption technology increases the difficulty of traffic detection. Although there are many studies focusing on intrusion detection technology for encrypted flows, there are few studies that combine the challenges of encryption and few labels. Therefore, it is also urgent to study network intrusion detection with few labels in encrypted scenarios. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides an intrusion detection model training method and an intrusion detection method to eliminate or improve one or more defects existing in the prior art.
[0006] A first aspect of the present invention provides an intrusion detection model training method, the method comprising the following steps: Training a preset intrusion detection model based on a plurality of data streams with respective intrusion category labels, the intrusion detection model comprising an encoder and a classifier connected in series; Based on each data stream of each intrusion category, each strongly enhanced and weakly enhanced data stream corresponding to each intrusion category is transformed through strong enhancement and weak enhancement to train the preset intrusion detection model, so that the intrusion detection model outputs each intrusion category prediction result corresponding to each strongly enhanced data stream and each pseudo label corresponding to each weakly enhanced data stream; Based on each data stream of each intrusion category and each interpolation sample obtained by interpolating every two data streams in each data stream, a preset feature extraction model is trained so that the feature extraction model outputs each embedding representation corresponding to each data stream and each interpolation embedding representation sample corresponding to each interpolation sample, and interpolates every two embedding representations in each embedding representation to obtain each embedding representation interpolation sample, wherein the feature extraction model includes the encoder and a projector connected to the encoder; A consistency regularization algorithm is used to keep each intrusion category prediction result and the corresponding pseudo-label, and each interpolation embedding representation sample and the corresponding embedding representation interpolation sample consistent, so that the trained intrusion detection model can output the corresponding intrusion category prediction result based on the data stream to be detected.
[0007] In some embodiments of the present invention, during the training process, the trained intrusion detection model is obtained by minimizing the total loss, where the total loss includes supervised classification loss, unsupervised classification loss, and unsupervised contrast loss.
[0008] In some embodiments of the present invention, in the process of minimizing the unsupervised classification loss, a pseudo label whose credibility is not less than a preset threshold is selected to calculate the unsupervised classification loss; By minimizing the unsupervised contrast loss, each positive sample pair is close to each other in each corresponding embedding space, and each negative sample pair is far away from each other in each corresponding embedding space. The positive sample pairs are formed by the embedding representation interpolation samples and interpolation embedding representation samples corresponding to any two data streams in each data stream of each intrusion category, and the negative sample pairs are formed by the two embedding representations corresponding to any two data streams in each data stream of each intrusion category.
[0009] In some embodiments of the present invention, the supervised classification loss adopts cross entropy loss and is calculated by the following formula: in, represents supervised classification loss, b and B represent the corresponding intrusion category labels. Data Flow The sequence and quantity of, H(·) represents the entropy function, Represents the data flow output by the intrusion detection model F(·) The probability distribution of the corresponding intrusion category prediction result y; The unsupervised classification loss adopts cross entropy loss and is calculated by the following formula: in, represents the unsupervised classification loss; k and Represents the data flow of each intrusion category The serial number and quantity, and ; According to the pseudo label Credibility Control The key function of , max(·) represents the maximum value function; Weakly enhanced data stream representing the output of the intrusion detection model The probability distribution of the corresponding intrusion category prediction result y, the y with the largest probability value is taken as the pseudo label, that is , Indicates the preset threshold; Strongly enhanced data stream representing the output of an intrusion detection model The probability distribution of the corresponding intrusion category prediction result y; The unsupervised contrast loss is calculated by the following formula: in, represents the unsupervised contrast loss; Indicates any two data flows among all data flows passing through each intrusion category and The corresponding embedding represents the interpolated sample and interpolated embedding representation samples The positive sample pairs formed, i and N represent the sequence number and number of positive sample pairs respectively. Indicates that the data flow and The corresponding embedding representation and The negative sample pairs formed, and k≠j; represents the indicator function, ; T represents temperature parameter; The total loss is calculated by the following formula: Where L represents the total loss and m represents the trade-off hyperparameter for controlling the unsupervised contrastive loss.
[0010] In some embodiments of the present invention, the data stream used to train the preset intrusion detection model and feature extraction model includes an encrypted data stream, or an encrypted data stream and a non-encrypted data stream; wherein, when the data stream is an encrypted data stream, before each training step, the method further includes: Perform feature extraction on the encrypted data stream to obtain corresponding data packet length sequence features, so as to train a preset intrusion detection model based on the data packet length sequence features, or perform strong enhancement and weak enhancement transformation on the data packet length sequence features, or train a preset feature extraction model based on the data packet length sequence features and interpolate the data packet length sequence features.
[0011] In some embodiments of the present invention, the strong enhancement transformation includes random permutation and random masking based on a strong enhancement factor, and the weak enhancement transformation includes random masking based on a weak enhancement factor.
[0012] A second aspect of the present invention provides an intrusion detection method, the method comprising the following steps: The data stream to be detected is input into a pre-trained intrusion detection model so that the intrusion detection model outputs an intrusion category prediction result corresponding to the data stream to be detected. The intrusion detection model is pre-trained by the intrusion detection model training method as described in the first aspect above.
[0013] In some embodiments of the present invention, the data stream to be detected includes an encrypted data stream and a non-encrypted data stream; wherein, when the data stream to be detected is an encrypted data stream, before the step of inputting the data stream to be detected into the pre-trained intrusion detection model, the method further includes: Feature extraction is performed on the data stream to be detected to obtain data packet length sequence features, so as to input the data packet length sequence features into a pre-trained intrusion detection model.
[0014] The third aspect of the present invention provides an electronic device, which includes: a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, when the computer instructions are executed by the processor, the device implements the steps of the intrusion detection model training method as described in the first aspect above, or implements the steps of the intrusion detection method as described in the second aspect above.
[0015] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the intrusion detection model training method as described in the first aspect above, or implements the steps of the intrusion detection method as described in the second aspect above.
[0016] The fifth aspect of the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the intrusion detection model training method as described in the first aspect above, or implement the steps of the intrusion detection method as described in the second aspect above.
[0017] The intrusion detection model training method and the intrusion detection method of the present invention utilize a very small amount of labeled data and a large amount of unlabeled data, and combine interpolation and pseudo-labeling technology to perform semi-supervised training on the intrusion detection model, which can fully utilize a large amount of unlabeled data and learn effective feature representations therefrom to accurately distinguish, identify and classify network data traffic of different intrusion categories, and can improve the performance of the intrusion detection model and its IDS trained under the condition of a very small amount of labeled data (such as only 3 or 5 labeled samples per category), and can achieve the same detection level of intrusion detection accuracy as that of supervised training of the model using a large amount of labeled data.
[0018] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.
[0019] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.
[0021] Figure 1 A schematic diagram of a flow chart of an intrusion detection model training method according to an embodiment of the present invention; Figure 2 A schematic diagram of the network architecture and specific training process of an intrusion detection model in one embodiment of the present invention; Figure 3 The figure is a flow chart of an intrusion detection method in one embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0023] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0024] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0025] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0026] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0027] Since a large amount of network labeled data of many new network intrusions or attack types is difficult to obtain and collect, however, the traditional intrusion detection method based on semi-supervised learning still needs to use a large amount of labeled data to train the deep learning model to realize network intrusion detection. Therefore, the traditional method cannot realize effective network intrusion detection when the labeled data is extremely limited, and the application is limited. In order to solve this technical problem, an embodiment of the present invention proposes an intrusion detection model training method and an intrusion detection method. The method proposes a semi-supervised learning model combining pseudo-labels and interpolation technology. By combining pseudo-labels, interpolation technology and semi-supervised learning methods, the model is trained simultaneously using a very small amount of labeled data and a large amount of unlabeled data, and the entire training process includes a supervised training process and two unsupervised training processes, which can enable the trained model to learn effective data distribution and feature representation from a large amount of unlabeled data, thereby improving the intrusion detection and classification performance when the labeled data is extremely limited.
[0028] Figure 1 and Figure 2 They are respectively a flow chart of an intrusion detection model training method according to an embodiment of the present invention and a network architecture of an intrusion detection model and a specific training flow chart. Figure 1 As shown, the method comprises the following steps: Step S110: training a preset intrusion detection model based on a plurality of data streams with respective intrusion category labels, wherein the intrusion detection model comprises an encoder and a classifier connected in series.
[0029] This step is as follows Figure 2The supervised training process shown in FIG. Specifically, the data stream with intrusion category labels includes two major categories of data, one of which is abnormal data stream belonging to network intrusion, and the other is normal data stream belonging to non-network intrusion. The intrusion categories include network intrusions and non-network intrusions of subdivided categories such as internal penetration attacks, denial of service attacks, distributed denial of service (DDoS) attacks, cracking, port scanning attacks, botnets, Web attacks, zero-day attacks, and n-day attacks. The training data for each intrusion category used for supervised training of the model above only includes a very small number of data stream samples with corresponding labels, such as 3 or 5, to form a supervised training data set consisting of a small number of labeled training samples. In this supervised training process, the intrusion detection model is trained by optimizing and minimizing the difference between the intrusion category prediction result of the intrusion detection model for the data stream with the real label and the real label of the data stream. The intrusion detection model (encoder and classifier) can adopt deep learning network structures such as LSTM, CNN, MLP, WRN-28-2, etc.
[0030] Step S120, training the preset intrusion detection model based on the strong enhancement and weak enhancement data streams corresponding to the data streams of each intrusion category obtained through strong enhancement and weak enhancement transformation, so that the intrusion detection model outputs the prediction results of each intrusion category corresponding to each strong enhancement data stream and each pseudo label corresponding to each weak enhancement data stream.
[0031] This step is as follows Figure 2The first unsupervised training process shown in Figure 1 is shown in Figure 2. Specifically, since network flows and their features are usually temporal sequences, such as raw byte sequences and packet length sequence features, the commonly used data enhancement methods in the traditional image field, such as rotation, flipping and scaling, are not suitable for network flow data enhancement. In addition, considering that packet loss, disorder and other induced phenomena may occur during network transmission, especially under high-speed network transmission or unstable conditions, this phenomenon is more serious. This method proposes a data enhancement algorithm suitable for network flow data, which has certain practical significance in the semantic feature analysis and learning of flow data. The data enhancement algorithm includes strong enhancement transformation and weak enhancement transformation, and the degree of enhancement transformation is controlled by introducing two parameters, strong enhancement factor and weak enhancement factor. Among them, the strong enhancement transformation includes random permutation (Random Permutation) and random masking (Random Masking) based on the strong enhancement factor. The unlabeled data stream is subjected to random permutation and random masking in turn to simulate disorder and packet loss phenomena, and the strong enhancement transformation of the unlabeled data stream is realized. The specific process may include: first, according to a preset strong enhancement factor Ks, determine Ks-1 segmentation points from the integer range of the sequence length of the unlabeled data, and based on the determined Ks-1 segmentation points, cut and segment the sequence of the unlabeled data stream to obtain a plurality of segmented sequence segments, and shuffle and re-randomly sort the order of the plurality of segmented sequence segments, and then splice the randomly sorted plurality of segmented sequence segments to form a spliced new sequence corresponding to the unlabeled data, and then randomly mask the Ks feature points in the spliced new sequence, and finally obtain the strongly enhanced data stream corresponding to the unlabeled data stream. The weak enhancement transformation includes a random mask based on a weak enhancement factor, by only performing a small degree of random masking on the original data, that is, the unlabeled data stream, so as to retain the main features of the data stream while introducing slight disturbances. The specific process may include: according to a preset weak enhancement factor Kw, randomly mask the Kw feature points in the sequence of the unlabeled data stream to obtain the weakly enhanced data stream corresponding to the unlabeled data stream. The above data augmentation algorithms, on the one hand, not only increase the diversity of training data, but also help the model learn more robust feature representations, so that the model can have a certain tolerance to data noise; on the other hand, they have a certain robustness against adversarial attacks. Due to the need to comply with protocol standards and ensure the effectiveness of attacks, network traffic cannot be arbitrarily tampered with to implement adversarial attacks. Recent studies have shown that available adversarial examples, such as evasion packets, can be automatically mined through selective symbolic execution (S2E) technology. The so-called evasion packets refer to different packet sequences that appear between intermediate devices and terminal hosts due to differences in protocol implementation. In principle, this is very homogeneous with packet loss.Therefore, the data enhancement method that simulates network traffic data packet loss and disorder in this method can resist similar adversarial attacks to a certain extent. When facing a real network transmission environment, even if real network induced phenomena occur, it will not have much impact on the final intrusion detection and classification results.
[0032] In the first unsupervised training process of the intrusion detection model, the intrusion detection model predicts the intrusion category of the weakly enhanced data stream to generate the corresponding pseudo-label of the weakly enhanced data stream; and predicts the intrusion category of the strongly enhanced data stream to generate the intrusion category prediction result corresponding to the strongly enhanced data stream. By optimizing and minimizing the difference between the pseudo-label corresponding to the weakly enhanced data stream corresponding to the unlabeled data stream and the intrusion category prediction result corresponding to the strongly enhanced data stream, the intrusion detection model is trained, and the robustness and generalization ability of the trained intrusion detection model can be enhanced, and it has a certain fault tolerance and robustness even in the face of an unstable or high-speed transmission network environment.
[0033] Step S130, based on each data stream of each intrusion category and each interpolation sample obtained by interpolating every two data streams in each data stream, a preset feature extraction model is trained, so that the feature extraction model outputs each embedded representation corresponding to each data stream and each interpolation embedded representation sample corresponding to each interpolation sample, and interpolates every two embedded representations in each embedded representation to obtain each embedded representation interpolation sample, and the feature extraction model includes the encoder and a projector connected to the encoder.
[0034] This step is as follows Figure 2 The second unsupervised training process shown. The design of this unsupervised training process stems from an observation that if the output of the network model can change linearly, the decision boundary will become larger. In the first and second unsupervised training processes, the unsupervised training dataset includes a large number of unlabeled data streams, and the various intrusion categories involved can be the same as the various intrusion categories involved in the supervised training dataset, which can enable the trained intrusion detection model to achieve accurate network intrusion or non-intrusion detection, and can achieve accurate detection of data streams of specific subdivided intrusion categories. The various intrusion categories involved in the unlabeled data streams in the unsupervised training dataset may not be exactly the same as the various intrusion categories involved in the supervised training dataset, which can also enable the trained intrusion detection model to achieve accurate intrusion or non-intrusion detection, but there may be a problem of poor detection performance for data streams of subdivided intrusion categories not included in the unsupervised training dataset. In the case of limited labeled training data, in order to accurately capture the intrinsic characteristics of unlabeled training data, this method proposes and adopts a strategy of interpolating the unlabeled data stream in the original space and embedded space of the unlabeled data stream used to train the feature extraction network model.
[0035] Specifically, set any two unlabeled data stream samples in the unsupervised training dataset and On the one hand, the two unlabeled data stream samples are firstly transformed through the encoder and projector in the feature extraction model and Mapped to their respective embedding spaces to obtain low-dimensional feature representations and : Among them, G(·) represents the feature extraction model. Then in the embedding space and Perform interpolation operations to generate embedded representation interpolation samples ,in, , It follows a Beta distribution, and On the other hand, in the original data space, for two unlabeled data stream samples and Interpolation is performed to generate interpolated samples, and then the interpolated samples are mapped to the corresponding embedding space using the feature extraction model G(·) to obtain the interpolated embedding representation samples .
[0036] In the second unsupervised training process of the feature extraction model that shares the same encoder with the intrusion detection model, the feature extraction model first extracts and maps the original unlabeled data to obtain the feature representation, and interpolates the feature representation in the embedding space to obtain the interpolated embedding representation result; and extracts and maps the interpolated result obtained by interpolating the unlabeled data in the original data space to obtain the interpolated embedding representation result. By optimizing and minimizing the difference between the interpolated embedding representation results generated by the unlabeled data stream under the two interpolation strategies and the corresponding interpolated embedding representation results, the feature extraction model is trained, and the data feature representation and discrimination capabilities of the encoder shared by the trained feature extraction model and the intrusion detection model can be enhanced.
[0037] Step S140, using a consistency regularization algorithm to keep each intrusion category prediction result and the corresponding pseudo-label, and each interpolation embedding representation sample and the corresponding embedding representation interpolation sample consistent, so that the trained intrusion detection model outputs the corresponding intrusion category prediction result based on the data stream to be detected.
[0038] The consistency regularization algorithm in this step is the regularized consistency constraint. In the first unsupervised training process, the consistency regularization forces the intrusion detection model to predict the intrusion category of the strongly enhanced data stream samples to be consistent with the pseudo-labels of the corresponding weakly enhanced data stream samples, which can improve the accuracy and reliability of the pseudo-labels. In the second unsupervised training process, the regularized consistency constraint is used to expect the intrusion detection model to come from the same original unlabeled data stream. and The corresponding two features represent and As close as possible in the embedding space, the feature representations obtained by the two interpolation strategies in the embedding space can be made consistent. and Combined with the two interpolation strategies to construct a positive sample pair, this positive sample pair construction method based on two interpolation strategies can make the embedding relationship of the network model change linearly. By performing this interpolation consistency training on a large amount of unlabeled data, it is possible to force the embedding representation of the feature extraction network model to introduce linear changes, which is conducive to expanding the margin of the decision boundary, so as to improve the ability of the intrusion detection network model that shares the same encoder with the feature extraction network model to distinguish data streams of different intrusion categories, thereby further improving the intrusion detection model's ability to identify intrusion categories with very little labeled training data.
[0039] During the entire training process of the intrusion detection model, the trained intrusion detection model is obtained by minimizing the total loss, and the total loss includes supervised classification loss, unsupervised classification loss and unsupervised contrast loss. In the process of minimizing the unsupervised classification loss, a pseudo label whose credibility (confidence) is not less than a preset threshold is selected to calculate the unsupervised classification loss. By minimizing the unsupervised contrast loss, each positive sample pair is close to each other in each corresponding embedding space, and each negative sample pair is far away from each other in each corresponding embedding space. The positive sample pair is formed by the embedded representation interpolation sample and the interpolated embedded representation sample corresponding to any two data streams in each data stream of each intrusion category, and the negative sample pair is formed by the two embedded representations corresponding to any two data streams in each data stream of each intrusion category.
[0040] Specifically, the supervised classification loss adopts cross entropy loss and can be calculated by the following formula: in, represents supervised classification loss, b and B represent the corresponding intrusion category labels. Data Flow The sequence and quantity of, H(·) represents the entropy function, Represents the data flow output by the intrusion detection model F(·) The probability distribution of the corresponding intrusion category prediction result y.
[0041] The unsupervised classification loss adopts cross entropy loss and can be calculated by the following formula: in, represents the unsupervised classification loss; k and Represents the data flow of each intrusion category The serial number and quantity, and ; According to the pseudo label Credibility Control The key function of , max(·) represents the maximum value function; Weakly enhanced data stream representing the output of the intrusion detection model The probability distribution of the corresponding intrusion category prediction result y, the y with the largest probability value is taken as the pseudo label, that is , Indicates the preset threshold; Strongly enhanced data stream representing the output of an intrusion detection model The probability distribution of the corresponding intrusion category prediction result y. When the maximum category probability value in the intrusion category prediction result corresponding to the weakly enhanced data stream is higher than the threshold hour, , the unsupervised classification loss is calculated normally; otherwise, the generated pseudo-label is considered inaccurate and does not participate in the calculation of the unsupervised classification loss. , which can ensure the credibility of pseudo labels. In the first unsupervised training process, by selecting pseudo labels with a credibility not less than a preset threshold and ignoring pseudo labels with lower credibility to calculate the unsupervised classification loss, the negative impact of wrong labels on the intrusion detection model training can be avoided, thereby improving the generalization ability and stability of the intrusion detection model. At the same time, it also allows the intrusion detection model to gradually learn and improve the model's ability to generate high-quality pseudo labels during the training process, further improving the performance of the intrusion detection model and the quality of the generated pseudo labels.
[0042] The unsupervised contrast loss can be calculated by the following formula: in, represents the unsupervised contrast loss; Indicates any two data flows among all data flows passing through each intrusion category and The corresponding embedding represents the interpolated sample and interpolated embedding representation samples The positive sample pairs formed, i and N represent the sequence number and number of positive sample pairs respectively. Indicates that the data flow and The corresponding embedding representation and The negative sample pairs formed, and k≠j; represents the indicator function, ; T represents the temperature parameter. In the second unsupervised training process, the contrastive learning method is used to optimize the unsupervised contrast loss to force the positive samples to Positive samples in and are close to each other in the embedding space, while the negative samples Negative samples in and The labels are far away from each other in the embedding space, so that the decision boundary is expanded, thereby improving the ability of the trained intrusion detection model to distinguish different intrusion category data in the case of very few labeled training data.
[0043] Therefore, the total loss for training the intrusion detection model is calculated by the following formula: Among them, L represents the total loss, and m represents the trade-off hyperparameter used to control the unsupervised contrast loss, that is, The weight parameter of . The calculation of , the weight parameter of the unsupervised classification loss can be left unchanged.
[0044] In some embodiments, the data stream used to train the preset intrusion detection model and feature extraction model includes an encrypted data stream, or an encrypted data stream and a non-encrypted data stream.
[0045] Wherein, in the case where the data stream is an encrypted data stream, before each training step, the method further comprises the following steps: Perform feature extraction on the encrypted data stream to obtain corresponding data packet length sequence features, so as to train a preset intrusion detection model based on the data packet length sequence features, or perform strong enhancement and weak enhancement transformation on the data packet length sequence features, or train a preset feature extraction model based on the data packet length sequence features and interpolate the data packet length sequence features.
[0046] The data streams included in the training data sets of the supervised training process and the two unsupervised training processes can be all encrypted data streams, or they can be two types of data streams, including encrypted data streams and non-encrypted data streams. Both of them can enable the trained intrusion detection model to realize network intrusion detection in both encrypted and non-encrypted scenarios. In the case where the training data is non-encrypted data streams, in order to capture more fine-grained information, the original byte sequence of the network data stream is directly used as the model input to train the model, and no feature engineering is required. In addition, the encrypted data stream classification in the intrusion detection method of the traditional encrypted network data stream faces the characterization difficulties caused by factors such as few available features and unreliable time features. Most of them rely on various statistical features obtained by feature engineering, which increases the complexity of feature extraction, and the statistical features of only coarse-grained traffic data cannot fully reflect the network behavior characteristics when the labeled samples are extremely limited. Moreover, in the encrypted scenario, the payload of the data packet is unreadable and the bytes contained have no practical meaning. However, network traffic of different intrusion categories will show different length feature sequences in the time dimension due to their inherent behavior and pattern differences. For example, the network traffic of DDoS (distributed denial of service) attacks tends to increase dramatically in a short period of time, and the packet length sequence will show obvious spikes and fluctuation patterns. In contrast, the packet length sequence of normal traffic that is not a network attack tends to be more stable and uniform. Therefore, when the training data is an encrypted data stream, this method uses the packet length sequence characteristics of the encrypted data stream to reveal the pattern differences in network traffic behavior, and combines context analysis technology to fully explore deep data association features, improve the understanding of the behavior pattern of the encrypted data stream in the time dimension, and thus help the trained intrusion detection model to better classify the encrypted stream.
[0047] Figure 3 FIG. 1 is a flow chart of an intrusion detection method according to an embodiment of the present invention. Figure 3 As shown, the method comprises the following steps: Step S310, inputting the data stream to be detected into a pre-trained intrusion detection model so that the intrusion detection model outputs an intrusion category prediction result corresponding to the data stream to be detected, and the intrusion detection model is pre-trained by the intrusion detection model training method described in any of the aforementioned embodiments.
[0048] In some embodiments, the data stream to be detected includes an encrypted data stream and a non-encrypted data stream. Wherein, in the case where the data stream to be detected is an encrypted data stream, before the step of inputting the data stream to be detected into the pre-trained intrusion detection model, the method further includes the following steps: Feature extraction is performed on the data stream to be detected to obtain data packet length sequence features, so as to input the data packet length sequence features into a pre-trained intrusion detection model.
[0049] Through the above steps, the intrusion detection method can realize network attack or intrusion detection in encrypted scenarios and non-encrypted scenarios, and can further realize network intrusion detection in subdivided categories such as internal penetration attack, denial of service attack, DDoS attack, cracking, port scanning attack, botnet, Web attack, zero-day attack and n-day attack.
[0050] In summary, the intrusion detection model training method and the intrusion detection method in the embodiment of the present invention use a very small amount of labeled data and a large amount of unlabeled data, and combine interpolation and pseudo-labeling technology to perform semi-supervised training on the intrusion detection model, which can make full use of a large amount of unlabeled data and learn effective feature representations from it to accurately distinguish, identify and classify network data traffic of different intrusion categories, and can improve the performance of the intrusion detection model and its IDS trained under the condition of a very small amount of labeled data (such as only 3 or 5 labeled samples per category), and can achieve the same level of intrusion detection accuracy as the supervised training of the model using a large amount of labeled data. This can greatly reduce the time, manpower and other costs required for labeling in practical applications. Among them, in the unsupervised training process corresponding to the pseudo-labeling technology, a data enhancement technology suitable for network flow data is used to reduce pseudo-label noise while enhancing the robustness of the intrusion detection model. In the unsupervised training process corresponding to the interpolation technology, two interpolation strategies are introduced to construct positive sample pairs, so that the embedding representation of the feature extraction model that shares the same encoder with the intrusion detection model changes linearly, which greatly improves the intrusion detection model under the condition of very few labeled data. The ability to detect and distinguish intrusion categories and the ability to distinguish intrusion categories in the intrusion detection model.
[0051] Corresponding to the above method, the present invention also provides an electronic device, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the device implements the steps of the above method.
[0052] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0053] An embodiment of the present invention further provides a computer program product, comprising computer instructions, which implement the steps of the aforementioned method when executed by a processor.
[0054] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0055] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0056] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training an intrusion detection model, characterized in that: The method comprises: Training a preset intrusion detection model based on a plurality of data streams with respective intrusion category labels, the intrusion detection model comprising an encoder and a classifier connected in series; Based on each data stream of each intrusion category, each strongly enhanced and weakly enhanced data stream corresponding to each intrusion category is transformed through strong enhancement and weak enhancement to train the preset intrusion detection model, so that the intrusion detection model outputs each intrusion category prediction result corresponding to each strongly enhanced data stream and each pseudo label corresponding to each weakly enhanced data stream; Based on each data stream of each intrusion category and each interpolation sample obtained by interpolating every two data streams in each data stream, a preset feature extraction model is trained so that the feature extraction model outputs each embedding representation corresponding to each data stream and each interpolation embedding representation sample corresponding to each interpolation sample, and interpolates every two embedding representations in each embedding representation to obtain each embedding representation interpolation sample, wherein the feature extraction model includes the encoder and a projector connected to the encoder; A consistency regularization algorithm is used to keep each intrusion category prediction result and the corresponding pseudo-label, and each interpolation embedding representation sample and the corresponding embedding representation interpolation sample consistent, so that the trained intrusion detection model can output the corresponding intrusion category prediction result based on the data stream to be detected.
2. The method according to claim 1, characterized in that During the training process, the trained intrusion detection model is obtained by minimizing the total loss, where the total loss includes supervised classification loss, unsupervised classification loss and unsupervised contrast loss.
3. The method according to claim 2, characterized in that In the process of minimizing the unsupervised classification loss, selecting a pseudo label whose credibility is not less than a preset threshold to calculate the unsupervised classification loss; By minimizing the unsupervised contrast loss, each positive sample pair is close to each other in each corresponding embedding space, and each negative sample pair is far away from each other in each corresponding embedding space. The positive sample pairs are formed by the embedding representation interpolation samples and interpolation embedding representation samples corresponding to any two data streams in each data stream of each intrusion category, and the negative sample pairs are formed by the two embedding representations corresponding to any two data streams in each data stream of each intrusion category.
4. The method according to claim 3, characterized in that The supervised classification loss adopts cross entropy loss and is calculated by the following formula: in, represents supervised classification loss, b and B represent the corresponding intrusion category labels. Data Flow The sequence and quantity of, H(·) represents the entropy function, Represents the data flow output by the intrusion detection model F(·) The probability distribution of the corresponding intrusion category prediction result y; The unsupervised classification loss adopts cross entropy loss and is calculated by the following formula: in, represents the unsupervised classification loss; k and Represents the data flow of each intrusion category The serial number and quantity, and ; According to the pseudo label Credibility Control The key function of , max(·) represents the maximum value function; Weakly enhanced data stream representing the output of the intrusion detection model The probability distribution of the corresponding intrusion category prediction result y, the y with the largest probability value is taken as the pseudo label, that is , Indicates the preset threshold; Strongly enhanced data stream representing the output of an intrusion detection model The probability distribution of the corresponding intrusion category prediction result y; The unsupervised contrast loss is calculated by the following formula: in, represents the unsupervised contrast loss; Indicates any two data flows among all data flows passing through each intrusion category and The corresponding embedding represents the interpolated sample and interpolated embedding representation samples The positive sample pairs formed, i and N represent the sequence number and number of positive sample pairs respectively. Indicates that data flow and The corresponding embedding representation and The negative sample pairs formed, and k≠j; represents the indicator function, ; T represents temperature parameter; The total loss is calculated by the following formula: Where L represents the total loss and m represents the trade-off hyperparameter for controlling the unsupervised contrastive loss.
5. The method according to claim 1, characterized in that The data stream used to train the preset intrusion detection model and feature extraction model includes an encrypted data stream, or an encrypted data stream and a non-encrypted data stream; wherein, when the data stream is an encrypted data stream, before each training step, the method further includes: Perform feature extraction on the encrypted data stream to obtain corresponding data packet length sequence features, so as to train a preset intrusion detection model based on the data packet length sequence features, or perform strong enhancement and weak enhancement transformation on the data packet length sequence features, or train a preset feature extraction model based on the data packet length sequence features and interpolate the data packet length sequence features.
6. The method according to claim 1, characterized in that The strong enhancement transformation includes a random permutation and a random mask based on a strong enhancement factor, and the weak enhancement transformation includes a random mask based on a weak enhancement factor.
7. An intrusion detection method, characterized in that: The method comprises: The data stream to be detected is input into a pre-trained intrusion detection model so that the intrusion detection model outputs an intrusion category prediction result corresponding to the data stream to be detected, and the intrusion detection model is pre-trained by the method described in any one of claims 1 to 6.
8. The method according to claim 7, characterized in that The data stream to be detected includes an encrypted data stream and a non-encrypted data stream; wherein, when the data stream to be detected is an encrypted data stream, before the step of inputting the data stream to be detected into the pre-trained intrusion detection model, the method further includes: Feature extraction is performed on the data stream to be detected to obtain data packet length sequence features, so as to input the data packet length sequence features into a pre-trained intrusion detection model.
9. An electronic device comprising a processor, a memory and computer instructions stored in the memory, characterized in that: The processor is used to execute the computer instructions, and when the computer instructions are executed, the device implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Interpolation comparison learning method in few-mark semi-supervised learning
CN114372571A
Training method of fine-grained network intrusion detection model and network intrusion detection method
CN116232699A
Prediction model training method and device, equipment, medium and program product
CN117217368A
Semi-supervised video fire detection method based on consistency regularization and distribution alignment
CN117274881A
Encrypted stream threat detection method and system based on context analysis
CN117640252A