A Network Anomaly Traffic Detection Method and Device Based on PU-MIL

By integrating positive unlabeled learning and multiple instance learning with a dual-part loss function, the method addresses the challenges of costly labeling and class imbalance in network anomaly detection, achieving efficient and accurate anomaly identification.

CN119996076BActive Publication Date: 2025-07-15BEIJING CHAITIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510449733.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

In the prior art, abnormal traffic detection methods rely on comprehensive labeled data, and the positive and negative samples are unbalanced, resulting in increased difficulty in model training and weakened abnormal traffic recognition capabilities.

Method used

Using a network abnormal traffic detection method based on PU-MIL, combined with positive sample, unlabeled sample learning and multi-example learning, data flow-level and packet-level classifiers are built by introducing an automatic encoder model of two-component loss function, and deep learning backpropagation training.

Benefits of technology

In the case of a small number of known exception samples and a large number of unlabeled data, efficient and accurate identification of abnormal traffic in the network has improved the detection ability of a few types of abnormal events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996076B_ABST
    Figure CN119996076B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for detecting abnormal network traffic based on PU-MIL, which relates to the field of network security technology. The method includes: dividing a network data stream containing multiple data packets into positive packets and unlabeled packets containing multiple examples; constructing an autoencoder model with a two-component loss function; through deep learning backpropagation training, iteratively solving the minimum value of the two-component loss function to obtain a data stream-level classifier and a data packet-level classifier; using the data stream-level classifier to discriminate abnormal data streams, and using the data packet-level classifier to detect abnormal data packets. By introducing an abnormal traffic detection technology combining positive sample, unlabeled sample learning and multi-instance learning, the present invention can efficiently and accurately identify abnormal traffic in the network when there are only a small number of known abnormal samples and a large amount of unlabeled data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and particularly to a method and device for detecting network abnormal traffic based on PU-MIL. Background Art

[0002] In current network traffic monitoring, traditional abnormal traffic detection methods rely on each data sample having a clear label (normal or abnormal). However, in practical applications, due to the high cost and long time consumption of the annotation process, it is extremely difficult to obtain comprehensive and accurate label information. In addition, traditional abnormal traffic detection methods usually assume that the positive and negative samples are balanced, but in reality, abnormal traffic only accounts for a small part of the total traffic, resulting in a serious imbalance between positive and negative samples in the training set. This imbalance not only increases the difficulty of model training, but also may cause the model to be biased towards the majority class, i.e., normal traffic, thus weakening the ability to identify the minority class, i.e., abnormal traffic.

[0003] Therefore, how to provide a technology that can not only process efficiently but also make full use of limited abnormal data stream information is crucial for improving the accuracy and efficiency of abnormal detection. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a method and device for detecting network abnormal traffic based on PU-MIL, aiming to solve the challenges of the serious dependence on fully annotated data and the imbalance between positive and negative samples in the existing abnormal traffic detection methods. By introducing an abnormal traffic detection technology that combines learning based on positive samples, unlabeled samples, and multi-instance learning (referred to as PU-MIL in the present invention), it is possible to efficiently and accurately identify abnormal traffic in the network with only a small number of known abnormal samples and a large amount of unlabeled data.

[0005] In one aspect of the present invention, a method for detecting network abnormal traffic based on PU-MIL is provided, including: dividing a network data stream containing multiple data packets into positive packets and unlabeled packets containing multiple instances; wherein, each instance corresponds to a data packet, the positive packet is a data stream with at least one abnormal data packet, and the unlabeled packet is a data stream without any class label; constructing an autoencoder model with a two-component loss function; wherein, the two-component loss function includes a first-component loss function and a second-component loss function, the first-component loss function is used to calculate the loss value of all unlabeled packets, and the second-component loss function is used to calculate the loss value of all positive packets;

[0006] Through deep learning backpropagation training, iteratively solving the minimum value of the two-component loss function to obtain a data stream-level classifier and a data packet-level classifier; using the data stream-level classifier to discriminate abnormal data streams, and using the data packet-level classifier to detect abnormal data packets.

[0007] Furthermore, the loss value of all unlabeled bags is calculated by the following formula:

[0008] ;

[0009] , where ;

[0010] Among them, represents the first component loss function, represents the bag-level reconstruction error in multi-instance learning, represents including data packets a data stream, represents the instance-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams.

[0011] Furthermore, the loss value of all positive bags is calculated by the following formula:

[0012] ;

[0013] ;

[0014] ;

[0015] Among them, represents the second component loss function, represents the data stream-level classifier, represents the data packet-level classifier, represents weight, represents including data packets a data stream, represents the instance-level reconstruction error in multi-instance learning, represents the set of abnormal data streams, represents the set of normal data streams, takes a constant, and represent the parameters to be learned by the autoencoder.

[0016] Furthermore, the data packet type, data packet length, data packet payload length, and data packet character statistical features are used as the instance features of each instance.

[0017] Furthermore, the set of abnormal data streams and the set of normal data streams The number of data streams is the same.

[0018] Another aspect of the present invention further provides a network abnormal traffic detection device based on PU-MIL, including: a first module configured to divide a network data stream containing multiple data packets into a positive packet and an unlabeled packet containing multiple examples; wherein each example corresponds to a data packet, the positive packet is a data stream with at least one abnormal data packet, and the unlabeled packet is a data stream without any class label; a second module configured to construct an autoencoder model with a two-component loss function; wherein the two-component loss function includes a first-component loss function and a second-component loss function, the first-component loss function is used to calculate the loss value of all unlabeled packets, and the second-component loss function is used to calculate the loss value of all positive packets; a third module configured to iteratively solve the minimum value of the two-component loss function through deep learning backpropagation training to obtain a data stream-level classifier and a data packet-level classifier; a fourth module configured to use the data stream-level classifier to discriminate abnormal data streams and use the data packet-level classifier to detect abnormal data packets.

[0019] Further, the loss value of all unlabeled packets is calculated by the following formula:

[0020] ;

[0021] , where ;

[0022] where, represents the first-component loss function, represents the packet-level reconstruction error in multi-instance learning, represents including data packets of a data stream, represents the example-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams.

[0023] Further, the loss value of all positive packets is calculated by the following formula:

[0024] ;

[0025] ;

[0026] ;

[0027] where, represents the second-component loss function, represents the data stream-level classifier, Represents a packet-level classifier, Represents the weight of, Represents containing packets a data stream of, Represents the example-level reconstruction error in multi-instance learning, Represents the abnormal data stream set, Represents the normal data stream set, Takes a constant, and Represents the parameters to be learned of the autoencoder.

[0028] Furthermore, the packet type, packet length, packet payload length, and packet character statistical features are used as the example features of each example.

[0029] Furthermore, the abnormal data stream set and the normal data stream set have the same number of data streams.

[0030] The network abnormal traffic detection method and device based on PU-MIL provided by the present invention introduce an abnormal traffic detection technology that combines positive sample, unlabeled sample learning (PU learning) and multi-instance learning (MIL learning), propose a new loss function for the autoencoder model, introduce the Platt scaling technique to reconstruct the error to adapt to the PU-MIL learning algorithm, and combine the example-level reconstruction error with the packet-level label to optimize the performance of the autoencoder model. Thus, in the case of only a small number of known abnormal samples and a large amount of unlabeled data, the abnormal traffic in the network can be efficiently and accurately identified, the detection ability for minority-class abnormal events is effectively improved, and a more practical and effective solution is provided for network security protection. Brief Description of the Drawings

[0031] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more obvious:

[0032] Figure 1 is a flowchart of the network abnormal traffic detection method based on PU-MIL provided by an embodiment of the present application;

[0033] Figure 2 is a schematic structural diagram of the network abnormal traffic detection device based on PU-MIL provided by an embodiment of the present application;

[0034] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include the plural forms unless the context clearly indicates otherwise.

[0037] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe the acquisition modules, these acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.

[0038] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0039] It should be noted that the orientation terms such as "upper", "lower", "left", "right", etc. described in the embodiments of the present invention are described from the angles shown in the accompanying drawings and should not be construed as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that one element is formed "on" or "under" another element, it can not only be directly formed "on" or "under" another element, but also be indirectly formed "on" or "under" another element through an intermediate element.

[0040] An embodiment of the present application provides a method for detecting abnormal network traffic based on PU-MIL (combining positive sample, unlabeled sample learning and multi-instance learning). This method uses an autoencoder as the underlying anomaly detector and introduces a new loss function for PU-MIL learning, which can achieve effective learning in the case of only partial positive packets and unlabeled packets, and solves the problem of the traditional method's dependence on fully labeled data.

[0041] To facilitate the understanding of the technical solutions of the present invention, some technical terms will be explained below:

[0042] Packet - In network communication, a packet is the basic unit transmitted by the TCP / IP protocol. By splitting data into multiple small packets for transmission, network reliability and efficiency are improved. Each packet is transmitted independently and reassembled into the original data at the destination.

[0043] Data stream - A collection of data continuously transmitted from a source to a destination over a period of time. A data stream contains several packets.

[0044] Bag - Bags and packets are two completely different concepts. A bag is a specialized term in multi-instance learning, representing a collection of multiple sample examples. A bag has a class label, while sample examples do not. When the label of a bag is negative, it means that all the samples in this bag are negative; when the label of a bag is positive, it means that at least one of the samples in this bag is positive.

[0045] The following details the network anomaly traffic detection method based on PU-MIL in this embodiment. Refer to Figure 1 , this method includes the following steps:

[0046] Step S101, divide the network data stream containing multiple packets into positive bags and unlabeled bags containing multiple examples; where each example corresponds to a packet, the positive bag is a data stream with at least one abnormal packet, and the unlabeled bag is a data stream without any class label.

[0047] Specifically, network traffic includes several data streams at different times, and each data stream contains multiple packets. Abnormal data streams are usually fewer compared to normal data streams. Therefore, in this embodiment, abnormal data streams are regarded as positive bags in PU multi-instance learning. The unlabeled bags in PU multi-instance learning contain both normal service data streams and unknown data streams that may contain malicious attacks. Packets are regarded as examples of each bag in PU multi-instance learning. If there is an abnormal packet (i.e., a positive example) in a certain data stream, then this data stream is considered an abnormal data stream (i.e., a positive bag). Among them, the packet type, packet length, payload length, character statistical features, etc. extracted from each packet are used as the example features of each example.

[0048] Step S102, construct an autoencoder model with a two-component loss function; where the two-component loss function includes a first-component loss function and a second-component loss function. The first-component loss function is used to calculate the loss value of all unlabeled bags, and the second-component loss function is used to calculate the loss value of all positive bags.

[0049] This step designs a brand-new loss function for the autoencoder model to achieve the effect of being able to distinguish both abnormal data streams and abnormal data packets.

[0050] Specifically, an autoencoder with a two-component loss function is trained, and the two-component loss function can be expressed as: . Among them, the first-component loss function is used to calculate the loss values of all unlabeled packets, that is, unlabeled packets are used to simulate the distribution of the seen examples. In other words, it is to use the unlabeled data stream (the data stream to be judged) to learn the distribution characteristics of each data packet in each data stream. The second-component loss function is used to calculate the loss values of all positive packets, that is, positive packets (data streams labeled as abnormal) are used to learn to distinguish positive packets and negative packets (i.e., abnormal data streams and normal data streams).

[0051] Furthermore, the loss value of all unlabeled packets is calculated by the following formula:

[0052] ;

[0053] , where ;

[0054] Among them, represents the first-component loss function, represents the packet-level reconstruction error in multi-instance learning, represents a data stream containing data packets , represents the example-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams.

[0055] The above calculation process obtains the packet-level reconstruction error through the example-level reconstruction error, and then obtains the loss value of all unlabeled packets through the packet-level reconstruction error.

[0056] Furthermore, the loss value of all positive packets is calculated by the following formula:

[0057] ;

[0058] ;

[0059] ;

[0060] Among them, represents the second-component loss function, represents the data stream-level classifier, represents a data packet - level classifier, represents the weight of represents containing data packets a data stream of represents the example - level reconstruction error in multi - instance learning, represents the set of abnormal data streams, represents the set of normal data streams, takes a constant and is preferably set to 0.01, and represents the parameters to be learned by the auto - encoder.

[0061] In the calculation process of the above loss value, first, by introducing Platt scaling , the example anomaly score range of the auto - encoder is converted from [0, + to probability values to unify the proportions of different abnormal data streams; second, since only learning the information of abnormal data streams will cause the auto - encoder model to overfit and have an imbalance problem, so in this step, the same number of normal data streams as the set of abnormal data streams are collected to form the set of normal data streams , thus transforming the problem into a classification task with balanced classes; finally, the negative value of the log - likelihood of abnormal data streams and normal data streams is used as the loss function, and by introducing the parameter to alleviate overfitting.

[0062] After completing the calculation of the above two parts of the loss function, the two parts are added to obtain the two - component loss function = , which is the network abnormal data stream detection model based on PU - MIL.

[0063] Step S103, through deep - learning backpropagation training, iteratively solve the minimum value of the two - component loss function to obtain the data - stream - level classifier and the data - packet - level classifier.

[0064] Specifically, by using the backpropagation training technique in deep learning to iteratively solve the minimum value of the two - component loss function , the data - stream - level classifier and the data - packet - level classifier can be obtained.

[0065] Step S104, use the data - stream - level classifier to discriminate abnormal data streams and use the data - packet - level classifier to detect abnormal data packets.

[0066] The network abnormal traffic detection method based on PU-MIL provided in this embodiment introduces an abnormal traffic detection technology that combines positive samples, unlabeled sample learning, and multi-instance learning. A new loss function for the autoencoder model is proposed, and the Platt scaling technique is introduced to reconstruct the error to adapt to the PU-MIL learning algorithm. By combining the instance-level reconstruction error and the packet-level label, the performance of the autoencoder model is optimized, so as to efficiently and accurately identify the abnormal traffic in the network with only a small number of known abnormal samples and a large amount of unlabeled data, and effectively improve the detection ability for rare abnormal events.

[0067] See Figure 2 Another embodiment of the present invention further provides a network abnormal traffic detection device 200 based on PU-MIL, including a first module 201, a second module 202, a third module 203, and a fourth module 204. The device 200 can execute the network abnormal traffic detection method in the method embodiment.

[0068] Specifically, the network abnormal traffic detection device 200 based on PU-MIL includes:

[0069] The first module 201 is configured to divide the network data stream containing multiple data packets into positive packets and unlabeled packets containing multiple instances; wherein, each instance corresponds to a data packet, the positive packet is a data stream with at least one abnormal data packet, and the unlabeled packet is a data stream without any class label;

[0070] The second module 202 is configured to construct an autoencoder model with a two-component loss function; wherein, the two-component loss function includes a first-component loss function and a second-component loss function, the first-component loss function is used to calculate the loss value of all unlabeled packets, and the second-component loss function is used to calculate the loss value of all positive packets;

[0071] The third module 203 is configured to iteratively solve the minimum value of the two-component loss function through deep learning backpropagation training to obtain a data stream-level classifier and a data packet-level classifier;

[0072] The fourth module 204 is configured to use the data stream-level classifier to discriminate abnormal data streams and use the data packet-level classifier to detect abnormal data packets.

[0073] Furthermore, the loss value of all unlabeled packets is calculated by the following formula:

[0074] ;

[0075] , where ;

[0076] where, represents the first component loss function, represents the bag-level reconstruction error in multi-instance learning, represents including data packets a data stream, represents the instance-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams.

[0077] Furthermore, the loss value of all positive bags is calculated by the following formula:

[0078] ;

[0079] ;

[0080] ;

[0081] wherein, represents the second component loss function, represents the data stream-level classifier, represents the data packet-level classifier, represents the weight of, represents including data packets a data stream, represents the instance-level reconstruction error in multi-instance learning, represents the set of abnormal data streams, represents the set of normal data streams, takes a constant, and represent the parameters to be learned by the autoencoder.

[0082] Furthermore, the data packet type, data packet length, data packet payload length, and data packet character statistical features are used as the instance features of each instance.

[0083] Furthermore, the number of data streams in the set of abnormal data streams is the same as that in the set of normal data streams.

[0084] It should be noted that the network abnormal traffic detection device 200 provided in this embodiment corresponding to the technical solutions that can be used to execute the method embodiments has the same implementation principle and technical effects as the method, which will not be elaborated here.

[0085] See Figure 3, Another embodiment of the present invention provides a schematic structural diagram of an electronic device 300, which is used to implement the network abnormal traffic detection method based on PU-MIL in the method embodiment. The electronic device 300 in the embodiment of the present invention may include, but is not limited to, a PC, a server, a smart phone, a tablet computer, and a PDA. Figure 3 The illustrated electronic device 300 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0086] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303 to implement the method of the embodiments as described in the present invention. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 305. The input / output (I / O) interface 304 is also connected to the bus 305.

[0087] Generally, the following devices may be connected to the I / O interface 304: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 3 the illustrated electronic device 300 has various devices, it should be understood that it is not required to implement or include all the illustrated devices. Instead, more or fewer devices may be implemented or included.

[0088] The above description is only the preferred embodiment of the present invention. Those skilled in the art should understand that the disclosed scope in the present invention is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.

Claims

1. A method for detecting network abnormal traffic based on PU-MIL, characterized in that, Including the following steps: Dividing a network data stream containing multiple data packets into positive packets and unlabeled packets containing multiple examples; wherein, each example corresponds to a data packet, the positive packet is a data stream with at least one abnormal data packet, and the unlabeled packet is a data stream without any class label; Constructing an autoencoder model with a two-component loss function; wherein, the two-component loss function includes a first-component loss function and a second-component loss function, the first-component loss function is used to calculate the loss value of all unlabeled packets, and the second-component loss function is used to calculate the loss value of all positive packets; The loss value of all unlabeled packets is calculated by the following formula: ; , among which ; Among them, represents the first component loss function, represents the bag-level reconstruction error in multi-instance learning, represents including number of data packets a data stream of represents the instance-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams; The loss value of all positive packets is calculated by the following formula: ; ; ; Among them, represents the second component loss function, represents the data stream level classifier, represents the data packet level classifier, represents the weight of represents containing data packets of a data stream, represents the instance level reconstruction error in multi-instance learning, represents the abnormal data stream set, represents the normal data stream set, takes a constant, and represents the parameters to be learned of the autoencoder; Through deep learning backpropagation training, iteratively solving the minimum value of the two-component loss function to obtain a data stream-level classifier and a data packet-level classifier; Using the data stream-level classifier to discriminate abnormal data streams and using the data packet-level classifier to detect abnormal data packets.

2. The network abnormal traffic detection method based on PU-MIL according to claim 1, characterized in that Taking the data packet type, data packet length, data packet payload length, and data packet character statistical features as the example features of each example.

3. The network abnormal traffic detection method based on PU-MIL according to claim 1, characterized in that, The set of abnormal data streams and the set of normal data streams have the same number of data streams.

4. A network abnormal traffic detection device based on PU-MIL, characterized in that, Including: A first module configured to divide a network data stream containing multiple data packets into positive packets and unlabeled packets containing multiple examples; wherein, each example corresponds to a data packet, the positive packet is a data stream with at least one abnormal data packet, and the unlabeled packet is a data stream without any class label; A second module configured to construct an autoencoder model with a two-component loss function; wherein, the two-component loss function includes a first-component loss function and a second-component loss function, the first-component loss function is used to calculate the loss value of all unlabeled packets, and the second-component loss function is used to calculate the loss value of all positive packets; The loss value of all unlabeled packets is calculated by the following formula: ; , where ; Among them, represents the first component loss function, represents the bag-level reconstruction error in multi-instance learning, represents including number of data packets a data stream of represents the instance-level reconstruction error in multi-instance learning, represents the number of all data streams, represents the set of all data streams; The loss value of all positive packets is calculated by the following formula: ; ; ; Among them, represents the second component loss function, represents the data stream level classifier, represents the data packet level classifier, represents the weight of, represents including data packets a data stream of, represents the example level reconstruction error in multi-instance learning, represents the abnormal data stream set, represents the normal data stream set, takes a constant, and represents the parameters to be learned of the autoencoder; A third module configured to, through deep learning backpropagation training, iteratively solve the minimum value of the two-component loss function to obtain a data stream-level classifier and a data packet-level classifier; A fourth module configured to use the data stream-level classifier to discriminate abnormal data streams and use the data packet-level classifier to detect abnormal data packets.

5. The network abnormal traffic detection device based on PU-MIL according to claim 4, wherein Taking the data packet type, data packet length, data packet payload length, and data packet character statistical features as the example features of each example.

6. The network abnormal traffic detection device based on PU-MIL according to claim 4, wherein, The abnormal data flow set and the normal data flow set have the same number of data flows.

Citation Information

Patent Citations

  • Weak supervision detection method and system for encrypting malicious traffic

    CN114826776A