Network traffic detection method and device, equipment and storage medium
By using teacher-student model architecture and knowledge distillation technology in network traffic detection, the problem that detection methods in the existing technology cannot be flexibly deployed and adapted to complex network environments is solved, and efficient, flexible and adaptive network traffic detection is achieved.
Patent Information
- Application Number
- CN202510332741.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-24
AI Technical Summary
The existing network security detection methods rely on massive data training models and cannot be flexibly deployed in the existing network architecture. The model lacks self-learning ability and has poor ability to adapt to complex and changeable network environments.
By training a pre-built neural network model based on the traffic data sample library, the teacher model is obtained, and its knowledge is transferred to the student model through knowledge migration. Using knowledge distillation technology, soften the probability distribution of the teacher model by adjusting the distillation temperature, and train the student model to improve its detection performance and adaptability.
It significantly improves the adaptability and detection performance of network traffic detection, so that network security detection systems can better deal with complex and changing network environments and growing security threats.
Smart Images

Figure CN120200796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular, to a network traffic detection method, device, equipment and storage medium. Background Art
[0002] In the context of the rapid development of digitalization today, network security threats emerge in an endless stream, various network incidents occur frequently, threats such as hacker attacks, data leaks, and malware continue to evolve, posing significant risks to personal privacy, enterprise operations, and even national security, making security challenges increasingly severe. Network traffic monitoring, as an important part of network security monitoring, is of great significance for defending against network attacks and maintaining the security of the network environment.
[0003] Domestic and foreign research scholars have proposed targeted anomaly detection methods for different network application scenarios, including: multi-level security protection of data centers and cloud computing environments by combining distributed computing frameworks with machine learning, a lightweight detection algorithm based on the cloud-edge-end computing framework to achieve a security protection mechanism for rapid response in the Internet of Things network, and capturing network "jitter" features by integrating deep learning and ensemble learning algorithms to achieve abnormal traffic detection in the mobile Internet. However, most of the above security protections rely on a large amount of data to train models and cannot be flexibly deployed into existing network architectures. In addition, the models themselves do not have the ability of self-learning and their adaptability to the real environment is poor. Therefore, in the face of a complex and changing network environment, there is an urgent need for a network security detection method with strong adaptability, high detection performance, and continuous learning to cope with the growing security threats and challenges. Summary of the Invention
[0004] The embodiments of the present invention provide a network traffic detection method, which can improve the adaptability and detection performance of network traffic detection.
[0005] In a first aspect, the embodiments of the present invention provide a network traffic detection method, including:
[0006] Training a pre-constructed neural network model based on a traffic data sample library to obtain a teacher model, and transferring the knowledge of the teacher model to a student model by using knowledge transfer;
[0007] Inputting the traffic data to be detected in the network environment to be detected into the teacher model and the student model, and respectively outputting the probability distribution predicted by the teacher model and the first prediction result predicted by the student model; wherein, the traffic data to be detected includes traffic data and corresponding original labels;
[0008] Based on knowledge distillation technology, softening the probability distribution by adjusting the distillation temperature to obtain the soft probability distribution of the teacher model;
[0009] Train the student model according to the soft probability distribution, so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and obtain a network traffic detection result according to the second prediction result.
[0010] Further, the method further includes:
[0011] Calculate the total loss during the training process of the student model based on the first prediction result, the original label, the soft probability distribution, and the second prediction result;
[0012] Update the student model and the teacher model according to the total loss.
[0013] Further, the calculating the total loss during the training process of the student model based on the first prediction result, the original label, the soft probability distribution, and the second prediction result includes:
[0014] Calculate a hard loss based on the first prediction result and the original label;
[0015] Calculate a soft loss based on the soft probability distribution and the second prediction result;
[0016] Calculate the total loss during the training process of the student model according to the soft loss and the hard loss.
[0017] Further, the calculating the hard loss based on the first prediction result and the original label includes:
[0018] Adopt the cross-entropy algorithm to calculate the hard loss based on the first prediction result and the original label, and the formula is as follows:
[0019]
[0020] where Loss hard is the hard loss, is the first prediction result, y i is the original label, λ i is the focusing level of the i-th class, and γ is the level regulation parameter used to control the attention to different classes in the sample.
[0021] Further, the calculating the soft loss based on the soft probability distribution and the second prediction result includes:
[0022] Adopt the cross-entropy algorithm to calculate the soft loss based on the soft probability distribution and the second prediction result, and the formula is as follows:
[0023]
[0024] Among them, Loss soft is the soft loss, is the soft probability distribution, that is, the value of the softmax output of the teacher model on the i-th class under the condition that the temperature is equal to T, is the second prediction result, that is, the value of the softmax output of the student model on the i-th class under the condition that the temperature is equal to T.
[0025] Furthermore, calculating the total loss according to the soft loss and the hard loss includes:
[0026] Constructing a total loss function according to the soft loss and the hard loss, and the function is expressed as follows:
[0027] Loss = αLoss soft +(1 - α)Loss hard ;
[0028] Among them, Loss is the total loss, Loss soft is the soft loss, Loss hard is the hard loss, α is the focal loss factor, and its value range is [0, 1], which is used to represent the dependence degree of the student model on the teacher model during the training process, and the dependence degree is positively correlated with the value size;
[0029] Calculating the total loss during the training process of the student model according to the total loss function.
[0030] Furthermore, updating the student model and the teacher model according to the total loss includes:
[0031] Sorting the traffic data to be detected according to the total loss to obtain a traffic data sequence;
[0032] Sequentially selecting a preset number of traffic data from the traffic data sequence and adding them to the traffic data sample library;
[0033] Updating the teacher model using the updated traffic data sample library, and the update formula is as follows:
[0034]
[0035] Among them, is the training parameter of the teacher model in the (t + 1)-th round of training, is the training parameter of the teacher model in the t-th round of training, Δw (t+1) is the degree of parameter change of the teacher model after the (t + 1)-th round of training compared with the t-th round of training, is the learning rate of the teacher model in the (t + 1)-th round of training, e (t+1)is the average value of the error between the original label and the probability distribution in the (t + 1)-th round of training of the teacher model, with a value range of [0, 1], κ is the original learning rate, and ε is a constant with a value range of [0, 1].
[0036] In a second aspect, an embodiment of the present invention provides a network traffic detection device, including:
[0037] A teacher model training module, configured to train a pre-constructed neural network model based on a traffic data sample library to obtain a teacher model, and transfer the knowledge of the teacher model to a student model by using knowledge transfer;
[0038] A student model hard prediction module, configured to input traffic data to be detected in a network environment to be detected into the teacher model and the student model, and respectively output a probability distribution predicted by the teacher model and a first prediction result predicted by the student model; wherein, the traffic data to be detected includes traffic data and corresponding original labels;
[0039] A probability distribution softening module, configured to soften the probability distribution based on knowledge distillation technology by adjusting the distillation temperature to obtain a soft probability distribution of the teacher model;
[0040] A student model training module, configured to train the student model according to the soft probability distribution, so that the student model learns from the teacher model based on the soft probability distribution and outputs a second prediction result, and obtain a network traffic detection result according to the second prediction result.
[0041] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0042] A memory, configured to store a computer program;
[0043] A processor, configured to execute the computer program;
[0044] Wherein, when the processor executes the computer program, the network traffic detection method according to any one of the first aspects is implemented.
[0045] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed, the network traffic detection method according to any one of the first aspects is implemented.
[0046] Compared with the prior art, a network traffic detection method provided by an embodiment of the present invention has the beneficial effects that: a neural network model pre-constructed is trained based on a traffic data sample library to obtain a teacher model, and knowledge transfer is used to transfer the knowledge of the teacher model to a student model; the traffic data to be detected in the network environment to be detected is input into the teacher model and the student model, and the probability distribution predicted by the teacher model and the first prediction result predicted by the student model are respectively output; wherein, the traffic data to be detected includes traffic data and corresponding original labels; based on the knowledge distillation technology, by adjusting the distillation temperature, the probability distribution is softened to obtain the soft probability distribution of the teacher model; the student model is trained according to the soft probability distribution, so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and a network traffic detection result is obtained according to the second prediction result; the adaptability and detection performance of network traffic monitoring can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical features of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the following described drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings according to these drawings without creative efforts.
[0048] Figure 1 is a schematic flowchart of an embodiment of a network traffic detection method provided by the present invention;
[0049] Figure 2 is a schematic structural diagram of an embodiment of a network traffic detection device provided by the present invention;
[0050] Figure 3 is a schematic structural diagram of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] It should be noted that although functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different sequence from that in the flowchart. Terms such as "first" and "second" in the description, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein are only for the purpose of describing the embodiments of this invention and are not intended to limit this invention.
[0054] In a first aspect, an embodiment of the present invention provides a network traffic detection method. Refer to Figure 1 , which is a schematic flowchart of an embodiment of a network traffic detection method provided by the present invention.
[0055] As Figure 1 shown, the method includes the following steps:
[0056] S1: Train a pre-constructed neural network model based on a traffic data sample library to obtain a teacher model, and transfer the knowledge of the teacher model to a student model by using knowledge transfer;
[0057] S2: Input the traffic data to be detected in the network environment to be detected into the teacher model and the student model, and respectively output the probability distribution predicted by the teacher model and the first prediction result predicted by the student model; wherein, the traffic data to be detected includes traffic data and corresponding original labels;
[0058] S3: Based on knowledge distillation technology, soften the probability distribution by adjusting the distillation temperature to obtain the soft probability distribution of the teacher model;
[0059] S4: Train the student model according to the soft probability distribution, so that the student model learns from the teacher model based on the soft probability distribution and outputs a second prediction result, and obtain a network traffic detection result according to the second prediction result.
[0060] It should be noted that knowledge distillation adopts a teacher-student mode, using a complex and large model as the teacher model. The structure of the student model is relatively simple. The teacher model is used to assist the training of the student model. The teacher model has strong learning ability and can transfer the knowledge it has learned to the student model with relatively weak learning ability, so as to enhance the generalization ability of the student model. The complex and bulky but effective teacher model is not put online and is simply a tutor role. The truly deployed and online prediction task is performed by the flexible and lightweight student model.
[0061] Through the end-edge-cloud collaborative framework, the present invention makes full use of computing resources at different levels. The end device is responsible for data collection, the edge node conducts local preliminary processing, and the cloud performs complex model training and updating. The lightweight design of the student model enables it to run on resource-constrained end devices and edge nodes, while the teacher model can be trained and optimized in the cloud using powerful computing resources. In this way, reasonable allocation and optimal utilization of resources can be achieved, improving the efficiency of the entire system.
[0062] In specific implementation, first, a teacher model and a student model are constructed. A traffic data sample library containing a large amount of network traffic data and their corresponding labels is used to train the neural network model. The trained model is used as the teacher model. Using knowledge transfer technology, the knowledge learned by the teacher model is transferred to a smaller and simpler-structured student model. The traffic data to be detected in the network environment to be detected is obtained, including traffic data and the corresponding original labels. The original labels are hard labels, which are respectively input into the teacher model and the student model. The teacher model outputs a probability distribution, representing the predicted probabilities of the input traffic data belonging to each category. The student model outputs a hard prediction result, directly predicting the category to which the input traffic data belongs.
[0063] Furthermore, knowledge distillation technology is used to utilize the predicted probability distribution of the teacher model to guide the training of the student model. By adjusting the distillation temperature, the predicted probability distribution of the teacher model is softened. It can be understood that the probability distribution directly output by the teacher model assigns a probability value to the prediction of each category. The largest probability value (i.e., the most likely category) usually dominates, while other smaller probability values may be greatly suppressed. When training the student model, relying only on the hard label (i.e., the category corresponding to the largest probability) for learning, the student model may not be able to fully capture the subtle differences between categories from the data. Therefore, by increasing the distillation temperature, the distribution of the teacher model becomes smoother, reducing extreme values, which helps the student model better learn the decision logic of the teacher model.
[0064] The probability distribution of the teacher model after softening is a soft probability distribution, also known as a soft label. Using the soft probability distribution as the training target to train the student model, the student model will try to learn the prediction logic of the teacher model while maintaining a small model size and fast inference speed. The trained student model will output a soft prediction result. Based on the soft prediction result of the student model, the category to which the input traffic data belongs can be determined.
[0065] In summary, the present invention trains a pre-constructed neural network model based on a traffic data sample library to obtain a teacher model, and uses knowledge transfer to transfer the knowledge of the teacher model to a student model; inputs the traffic data to be detected in the network environment to be detected into the teacher model and the student model, and respectively outputs the probability distribution predicted by the teacher model and the first prediction result predicted by the student model; wherein, the traffic data to be detected includes traffic data and corresponding original labels; based on the knowledge distillation technology, by adjusting the distillation temperature, softens the probability distribution to obtain the soft probability distribution of the teacher model; trains the student model according to the soft probability distribution, so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and obtains a network traffic detection result according to the second prediction result; the student model obtained by knowledge transfer in the present invention can inherit the generalization ability of the teacher model and has good adaptability to different network environments. By combining the traffic data sample library, the neural network model, knowledge transfer and knowledge distillation technology, an efficient, flexible and highly adaptable network traffic anomaly detection system is constructed, which can significantly improve the detection performance of network traffic monitoring and provide strong technical support for network security protection.
[0066] In an alternative embodiment, the method further includes:
[0067] Calculating the total loss during the training process of the student model based on the first prediction result, the original label, the soft probability distribution and the second prediction result;
[0068] Updating the student model and the teacher model according to the total loss.
[0069] Specifically, based on the first prediction result, the original label, the soft probability distribution and the second prediction result, calculate the total loss during the training process of the student model, and update the student model and the teacher model according to the total loss. When updating the student model, perform non-architectural pruning on the student model in real time to determine which connection channels and neuron weights are relatively small and remove them from the model.
[0070] In an alternative embodiment, the calculating the total loss during the training process of the student model based on the first prediction result, the original label, the soft probability distribution and the second prediction result includes:
[0071] Calculating a hard loss based on the first prediction result and the original label;
[0072] Calculating a soft loss based on the soft probability distribution and the second prediction result;
[0073] Calculate the total loss during the training process of the student model according to the soft loss and the hard loss.
[0074] Specifically, calculate the hard loss based on the first prediction result and the original label. The hard loss measures the difference between the result directly predicted by the student model for the sample and the true label of the sample. Calculate the soft loss based on the soft probability distribution and the second prediction result. The soft loss measures the difference between the prediction result obtained by the student model after learning from the teacher model and the prediction result of the teacher model. These two losses together constitute the total loss during the training process of the student model.
[0075] In an alternative implementation, calculating the hard loss based on the first prediction result and the original label includes:
[0076] Use the cross-entropy algorithm to calculate the hard loss based on the first prediction result and the original label. The formula is as follows:
[0077]
[0078] where Loss hard is the hard loss, is the first prediction result, y i is the original label, λ i is the focusing level of the i-th class, and γ is the level adjustment parameter used to control the attention to different classes in the sample.
[0079] Specifically, use the cross-entropy algorithm to calculate the hard loss. The hard loss measures the difference between the first prediction result of the student model and the original label. The calculation formula is as follows:
[0080]
[0081] where Loss hard is the hard loss, is the first prediction result, y i is the original label, usually one-hot encoded, λ i is the focusing level of the i-th class, and γ is the level adjustment parameter used to control the attention to different classes in the sample.
[0082] It can be understood that λ iIndicates the focusing level for each category, generally between 0 and 4, which can adjust the importance of different categories in loss calculation. For example, if a certain category is very important and the model needs to pay more attention, the focusing level of this category can be appropriately increased, so that the model pays more attention to the learning of this category during training. γ is used to control the modulation level, which can balance the data volume of classification samples, thereby increasing the attention to certain classification samples. For example, if the number of samples in some categories is small but very important, the loss weight of these categories can be increased by increasing the value of γ, so that the model pays more attention to these categories during training.
[0083] In an alternative embodiment, calculating the soft loss based on the soft probability distribution and the second prediction result includes:
[0084] Using the cross-entropy algorithm, based on the soft probability distribution and the second prediction result, calculate the soft loss, and the formula is as follows:
[0085]
[0086] where Loss soft is the soft loss, is the soft probability distribution, that is, the value of the softmax output of the teacher model on the i-th class under the condition that the temperature is equal to T, is the second prediction result, that is, the value of the softmax output of the student model on the i-th class under the condition that the temperature is equal to T.
[0087] Specifically, use the cross-entropy algorithm to calculate the soft loss. The soft loss measures the difference between the second prediction result of the student model and the soft probability distribution of the teacher model. By minimizing this loss, the student model can learn the prediction logic of the teacher model and imitate its output as much as possible. The calculation formula is as follows:
[0088]
[0089] where Loss soft is the soft loss, is the soft probability distribution, that is, the probability value of the output of the teacher model on the i-th class through the softmax function under the condition that the temperature is equal to T, and i represents the i-th category, z i represents the possibility of belonging to the i-th class, is the second prediction result, that is, the probability value of the output of the student model on the i-th class through the softmax function under the condition that the temperature is equal to T, z i' represents the probability belonging to the i-th class after being processed by the student model. The temperature T is a hyperparameter used to control the smoothness of the softmax output. A higher temperature makes the output smoother, while a lower temperature makes the output sharper. C is the total number of classes.
[0090] In an alternative embodiment, calculating the total loss according to the soft loss and the hard loss includes:
[0091] Constructing a total loss function according to the soft loss and the hard loss, and the function is expressed as follows:
[0092] Loss = αLoss soft +(1 - α)Loss hard ;
[0093] Where Loss is the total loss, Loss soft is the soft loss, Loss hard is the hard loss, α is the focal loss factor, and its value range is [0, 1], which is used to represent the degree of dependence of the student model on the teacher model during the training process. The degree of dependence is positively correlated with the value.
[0094] Calculating the total loss during the training process of the student model according to the total loss function.
[0095] Specifically, constructing a total loss function that combines the soft loss and the hard loss, and the function is expressed as follows:
[0096] Loss = αLoss soft +(1 - α)Loss hard ;
[0097] Where Loss is the total loss, Loss soft is the soft loss, Loss hard is the hard loss, α is the focal loss factor, and its value range is [0, 1], which is used to represent the degree of dependence of the student model on the teacher model during the training process. When α is close to 1, it means that the student model depends more on the prediction of the teacher model during the training process. When α is close to 0, it means that the student model depends more on the original label, and an appropriate α value needs to be selected according to the specific task and model performance.
[0098] The present invention provides a flexible way to balance the degree of dependence of the student model on the teacher model's prediction and the original label by introducing the focal loss factor, which helps to achieve better performance during the knowledge distillation process.
[0099] In an alternative embodiment, updating the student model and the teacher model according to the total loss includes:
[0100] Sort the traffic data to be detected according to the total loss to obtain a traffic data sequence;
[0101] Select a preset number of traffic data from the traffic data sequence in turn and add them to the traffic data sample library;
[0102] Use the updated traffic data sample library to update the teacher model, and the update formula is as follows:
[0103]
[0104] Among them, is the training parameter for the (t + 1)-th round of training of the teacher model, is the training parameter for the t-th round of training of the teacher model, Δw (t+1) is the degree of parameter change of the teacher model after the (t + 1)-th round of training compared to the t-th round of training, is the learning rate for the (t + 1)-th round of training of the teacher model, e (t+1) is the average value of the error between the original label and the probability distribution in the (t + 1)-th round of training of the teacher model, with a value range of [0, 1], κ is the original learning rate, and ε is a constant with a value range of [0, 1].
[0105] Specifically, the performance of the teacher model has a great impact on the training effect of the student model. Therefore, it should be ensured that the teacher model is accurate and reliable. After calculating the total loss, sort the traffic data to be detected according to the size of the total loss, select a preset number of traffic data and add them to the traffic data sample library. Exemplarily, samples with the top 10% of the total loss can be selected to expand the traffic data sample library, and the updated sample library is used to retrain or fine-tune the teacher model. The selection criterion can be determined according to actual needs. In the embodiments of the present invention, samples with the highest total loss and challenges are selected to focus on difficult-to-classify data.
[0106] Use the updated traffic data sample library to update the teacher model, and the update formula is as follows:
[0107]
[0108] Among them, is the training parameter for the (t + 1)-th round of training of the teacher model, is the training parameter for the t-th round of training of the teacher model, Δw (t+1) is the degree of parameter change of the teacher model after the (t + 1)-th round of training compared to the t-th round of training, is the learning rate for the (t + 1)-th round of training of the teacher model, which is a dynamic value, e (t+1) is the average value of the error between the original label and the probability distribution in the (t + 1)-th round of training of the teacher model, with a value range of [0, 1]. When e (t+1)The closer it is to 1, the worse the quality of the training data in the current (t + 1) round. Therefore, the learning rate is smaller, and the update of the teacher model parameters is slower. κ is the original learning rate, and ε is a constant with a value in [0, 1], which is used to prevent the situation where e (t+1) becomes zero.
[0109] In a second aspect, an embodiment of the present invention provides a network traffic detection device. Refer to Figure 2 , which is a schematic structural diagram of an embodiment of a network traffic detection device provided by the present invention.
[0110] As Figure 2 shown, the device includes:
[0111] A teacher model training module 21, configured to train a pre-constructed neural network model based on a traffic data sample library to obtain a teacher model, and transfer the knowledge of the teacher model to a student model by using knowledge transfer;
[0112] A student model hard prediction module 22, configured to input the traffic data to be detected in a network environment to be detected into the teacher model and the student model, and respectively output a probability distribution predicted by the teacher model and a first prediction result predicted by the student model; wherein, the traffic data to be detected includes traffic data and corresponding original labels;
[0113] A probability distribution softening module 23, configured to soften the probability distribution based on knowledge distillation technology by adjusting the distillation temperature to obtain a soft probability distribution of the teacher model;
[0114] A student model training module 24, configured to train the student model according to the soft probability distribution, so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and obtain a network traffic detection result according to the second prediction result.
[0115] In an optional implementation manner, the device further includes a model update module, configured to:
[0116] Calculate the total loss in the training process of the student model based on the first prediction result, the original label, the soft probability distribution, and the second prediction result;
[0117] Update the student model and the teacher model according to the total loss.
[0118] In an optional implementation manner, the model update module is further configured to:
[0119] Calculate a hard loss based on the first prediction result and the original label;
[0120] Calculate a soft loss based on the soft probability distribution and the second prediction result;
[0121] Calculate the total loss during the training process of the student model according to the soft loss and the hard loss.
[0122] In an alternative embodiment, the model update module is further configured to:
[0123] Use the cross-entropy algorithm to calculate the hard loss based on the first prediction result and the original label, and the formula is as follows:
[0124]
[0125] where Loss hard is the hard loss, is the first prediction result, y i is the original label, λ i is the focusing level of the i-th class, and γ is the level adjustment parameter used to control the attention to different classes in the sample.
[0126] In an alternative embodiment, the model update module is further configured to:
[0127] Use the cross-entropy algorithm to calculate the soft loss based on the soft probability distribution and the second prediction result, and the formula is as follows:
[0128]
[0129] where Loss soft is the soft loss, is the soft probability distribution, that is, the value of the softmax output of the teacher model on the i-th class under the condition that the temperature is equal to T, is the second prediction result, that is, the value of the softmax output of the student model on the i-th class under the condition that the temperature is equal to T.
[0130] In an alternative embodiment, the model update module is further configured to:
[0131] Construct a total loss function according to the soft loss and the hard loss, and the function is expressed as follows:
[0132] Loss = αLoss soft +(1 - α)Loss hard ;
[0133] where Loss is the total loss, Loss soft is the soft loss, Loss hardis the hard loss, α is the focal loss factor, and its value ranges from [0, 1], which is used to represent the dependence degree of the student model on the teacher model during the training process. The dependence degree is positively correlated with the value;
[0134] The total loss during the training process of the student model is calculated according to the total loss function.
[0135] In an alternative embodiment, the model update module is further configured to:
[0136] Sort the traffic data to be detected according to the total loss to obtain a traffic data sequence;
[0137] Sequentially select a preset number of traffic data from the traffic data sequence and add them to the traffic data sample library;
[0138] Update the teacher model using the updated traffic data sample library, and the update formula is as follows:
[0139]
[0140] where, are the training parameters of the teacher model in the (t + 1)-th round of training, are the training parameters of the teacher model in the t-th round of training, and Δw (t+1) is the degree of parameter change of the teacher model after the (t + 1)-th round of training compared to the t-th round of training, is the learning rate of the teacher model in the (t + 1)-th round of training, e (t+1) is the average value of the error between the original label and the probability distribution in the (t + 1)-th round of training of the teacher model, and its value ranges from [0, 1], κ is the original learning rate, and ε is a constant with a value ranging from [0, 1].
[0141] In a third aspect, an embodiment of the present invention provides an electronic device. Refer to Figure 3 shown, which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0142] As Figure 3 shown, the device includes:
[0143] A memory 31 for storing a computer program;
[0144] A processor 32 for executing the computer program;
[0145] Wherein, when the processor 32 executes the computer program, the network traffic detection method described in any of the above embodiments is implemented.
[0146] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 32 to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0147] The processor 32 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] The memory 31 can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory 31, and calling the data stored in the memory 31, the processor 32 realizes various functions of the electronic device. The memory 31 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.), etc. In addition, the memory 31 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0149] It should be noted that the above-mentioned electronic device includes, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 3 The structural schematic diagram is only an example of the above-mentioned electronic device, and does not constitute a limitation on the electronic device. It may include more components than shown in the figure, or combine some components, or different components.
[0150] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the network traffic detection method described in any of the above embodiments is implemented.
[0151] It should be understood that the present invention can implement all or part of the processes in the above-mentioned network traffic detection method, and can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned network traffic detection method. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0152] The above description is only a preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. It should be pointed out that for those skilled in the art, several equivalent obvious variations and / or equivalent substitutions can be made without departing from the technical principles of the present invention. These obvious variations and / or equivalent substitutions should also be regarded as the protection scope of the present invention.
Claims
1. A network traffic detection method, characterized in that: include: A pre-built neural network model is trained based on a traffic data sample library to obtain a teacher model, and knowledge transfer is used to transfer the knowledge of the teacher model to the student model; Input the traffic data to be detected of the network environment to be detected into the teacher model and the student model, and output the probability distribution predicted by the teacher model and the first prediction result predicted by the student model respectively; wherein the traffic data to be detected includes traffic data and corresponding original labels; Based on the knowledge distillation technology, the probability distribution is softened by adjusting the distillation temperature to obtain the soft probability distribution of the teacher model; The student model is trained according to the soft probability distribution so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and a network traffic detection result is obtained according to the second prediction result.
2. The network traffic detection method according to claim 1, characterized in that: The method further comprises: Calculate the total loss during the student model training process based on the first prediction result, the original label, the soft probability distribution and the second prediction result; The student model and the teacher model are updated according to the total loss.
3. The network traffic detection method according to claim 2, characterized in that: The calculating the total loss in the student model training process based on the first prediction result, the original label, the soft probability distribution and the second prediction result includes: Calculate a hard loss based on the first prediction result and the original label; Calculate a soft loss based on the soft probability distribution and the second prediction result; The total loss during the student model training process is calculated based on the soft loss and the hard loss.
4. The network traffic detection method according to claim 3, characterized in that: The calculating the hard loss based on the first prediction result and the original label includes: The cross entropy algorithm is used to calculate the hard loss based on the first prediction result and the original label. The formula is as follows: Among them, Loss hard For hard loss, is the first prediction result, y i is the original label, λ i is the focus level of the i-th category, and γ is the level control parameter, which is used to control the attention paid to different categories in the sample.
5. The network traffic detection method according to claim 3, characterized in that: The calculating the soft loss based on the soft probability distribution and the second prediction result includes: The cross entropy algorithm is used to calculate the soft loss based on the soft probability distribution and the second prediction result. The formula is as follows: Among them, Loss soft is a soft loss, is the soft probability distribution, that is, the value of the softmax output of the teacher model on the i-th class under the condition that the temperature is equal to T. is the second prediction result, that is, the value of the softmax output of the student model on the i-th category under the condition that the temperature is equal to T.
6. The network traffic detection method according to claim 3, characterized in that: The total loss is calculated according to the soft loss and the hard loss, including: According to the soft loss and the hard loss, a total loss function is constructed, and the function is expressed as follows: Loss=αLoss soft +(1-α)Loss hard ; Among them, Loss is the total loss, Loss soft is the soft loss, Loss hard is the hard loss, α is the focal loss factor, and its value is [0,1]. It is used to indicate the degree of dependence of the student model on the teacher model during the training process. The degree of dependence is positively correlated with the value. The total loss during the student model training process is calculated according to the total loss function.
7. The network traffic detection method according to claim 2, characterized in that: The updating of the student model and the teacher model according to the total loss includes: Sorting the flow data to be detected according to the total loss to obtain a flow data sequence; Selecting a preset number of flow data in sequence from the flow data sequence and adding them to the flow data sample library; Use the updated traffic data sample library to update the teacher model. The update formula is as follows: in, is the training parameter of the teacher model for the t+1th round of training, is the training parameter of the teacher model for the tth round of training, Δw (t+1) is the degree of parameter change of the teacher model after the t+1th round of training compared to the tth round of training, is the learning rate of the teacher model for the t+1th round of training, e (t+1) is the average value of the error between the original label and the probability distribution in the t+1th round of training of the teacher model, and its value is [0,1]. κ is the original learning rate, and ε is a constant with a value of [0,1].
8. A network traffic detection device, characterized in that: include: A teacher model training module is used to train a pre-built neural network model based on a traffic data sample library to obtain a teacher model, and use knowledge transfer to transfer the knowledge of the teacher model to the student model; The student model hard prediction module is used to input the traffic data to be detected of the network environment to be detected into the teacher model and the student model, and output the probability distribution predicted by the teacher model and the first prediction result predicted by the student model respectively; wherein the traffic data to be detected includes traffic data and corresponding original labels; A probability distribution softening module, which is used to soften the probability distribution based on the knowledge distillation technology by adjusting the distillation temperature to obtain a soft probability distribution of the teacher model; The student model training module is used to train the student model according to the soft probability distribution so that the student model learns the teacher model based on the soft probability distribution and outputs a second prediction result, and obtains a network traffic detection result according to the second prediction result.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program; Wherein, when the processor executes the computer program, the network traffic detection method as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the network traffic detection method according to any one of claims 1 to 7 is implemented.