Model training method, network abnormal flow detection method using model, system, device and medium

By using a discrete cosine transform-based autoencoder model in network anomaly traffic detection for feature extraction, and combining meta-learning, transfer learning, generative adversarial network and ant colony optimization algorithm, the problems of insufficient samples, unbalanced categories, feature redundancy and local optimal solutions in network anomaly traffic detection are solved, and higher detection accuracy and generalization capabilities are achieved.

CN120031105APending Publication Date: 2025-05-23CHENGDU ANZHUN NETWORK SECURITY TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510498558.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art faces insufficient sample, category imbalance, feature redundancy and local optimal solution problems in network abnormal traffic detection, resulting in insufficient model accuracy and generalization capabilities.

Method used

The autoencoder model based on discrete cosine transformation is used for feature extraction, and combined with meta-learning and transfer learning, the model's feature extraction ability and parameter adjustment efficiency are improved by generating adversarial networks and ant colony optimization algorithms.

Benefits of technology

It effectively solves the high-dimensionality and complexity problems of network traffic data, improves the detection accuracy and generalization capabilities of the model, and avoids the problems of local optimal solutions and feature redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031105A_ABST
    Figure CN120031105A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, a network anomaly traffic detection method, system, equipment and medium using the model, in a network anomaly traffic detection task, meta learning and transfer learning are fused, the effect of few-sample and cross-domain network anomaly detection is improved, meta knowledge is extracted through a multi-task learning normal form of meta learning, and the network anomaly traffic detection efficiency is improved. And in combination with knowledge migration from a source domain to a target domain in transfer learning, the problems of insufficient network traffic data samples and generalization of unlabeled data in the target domain are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer security technology, and in particular to a model training method, a network abnormal traffic detection method based on a fusion framework using the model, a system device and a medium. Background Art

[0002] With the rapid development of network technology and the complexity of network environment, the scale and types of network traffic are growing exponentially, and the forms of network attacks are becoming increasingly diverse and covert. As the core technology to ensure network security, network abnormal traffic detection faces many technical challenges in the face of complex and changeable network traffic. First, network traffic data usually has high-dimensional, dynamic and complex characteristics, abnormal traffic accounts for a small proportion of the overall traffic, and the category imbalance problem is serious. Secondly, in practical applications, the target domain network traffic data usually lacks sufficient annotations, and it is difficult to achieve ideal detection results directly using traditional supervised learning methods. In addition, due to the differences in network environment and traffic characteristics between the source domain and the target domain, it is difficult for existing technologies to achieve cross-domain knowledge transfer, resulting in insufficient adaptability and generalization of the model in the target domain. In the face of these problems, the existing network anomaly detection methods based on a single algorithm have certain limitations in performance, efficiency and generalization ability.

[0003] The Chinese invention patent with the specific prior art publication number CN117669655A proposes a network intrusion detection deep learning model compression method, which trains the deep learning teacher model through network intrusion traffic samples, calculates the student model Model-S after knowledge distillation according to the probability distribution of each type of traffic of the edge computing node k to be deployed, and then prunes the student model Model-S based on the model pruning method of the node traffic type distribution to obtain the compression model required to deploy the node k to achieve compression for different models. The present invention combines knowledge distillation technology and model pruning technology to perform targeted pruning on the network intrusion detection deep learning algorithm in the edge computing node based on the distribution of network traffic types faced by each node, and can perform personalized compression on the deep learning detection model according to the distribution of network traffic types of the edge computing nodes to be released, thereby effectively reducing the loss of computing and storage resources and improving the accuracy of traffic detection.

[0004] The Chinese invention patent with publication number CN112749028B proposes a network traffic processing method, related equipment and readable storage medium. In a multi-threaded network intrusion detection system, a network traffic receiving queue and a connection tracking table are set for each thread. The network data of the same network connection is received by a network traffic receiving queue. For the same network connection, there will not be multiple threads processing its network data. Moreover, when each thread processes the network data, it only needs to operate the corresponding connection tracking table. The threads will not interfere with each other, so that each thread in the system can process the network data in each network traffic receiving queue in parallel, thereby improving the performance of the multi-threaded network intrusion detection system.

[0005] It can be seen that the existing technology has the following problems that still need to be further solved:

[0006] 1. In the task of abnormal network traffic detection, traditional methods find it difficult to use source domain network traffic data knowledge to achieve effective model training when the target domain network traffic data samples are insufficient or unlabeled, and the model accuracy and generalization ability are insufficient.

[0007] 2. In the task of detecting abnormal network traffic, the existing generative adversarial networks often generate samples with uneven category distribution and feature redundancy when generating abnormal traffic data, which results in the detection model being unable to fully cover the diversity of actual abnormal traffic.

[0008] 3. In the task of detecting abnormal network traffic, the existing feature extraction methods have limited ability to process redundant information in high-dimensional network traffic data, making it difficult to highlight the main features, resulting in a decrease in classification accuracy.

[0009] 4. In the task of abnormal network traffic detection, existing neural network optimization methods are prone to fall into local optimal solutions when processing complex network traffic data, lack dynamic adjustment capabilities, and affect the training efficiency and final performance of the model. Summary of the invention

[0010] In a first aspect, the present invention provides a model training method, wherein the model is an autoencoder model that meets the characteristics of network traffic data;

[0011] The autoencoder model is an autoencoder algorithm model based on discrete cosine transform, and the training method includes:

[0012] Initializing the weights and biases of the autoencoder, wherein the initialization method is random initialization;

[0013] According to the high-dimensional characteristics and smoothness of network traffic data, the feature vector after feature extraction is regarded as the time domain feature before smoothing. The input network traffic data is mapped to the frequency space through discrete cosine transformation;

[0014] In the feature encoding process, linear transformation with dynamic sparse regularization is used to encode network traffic data;

[0015] receiving the encoded network traffic data through a decoder and attempting to reconstruct the original input network traffic data through an inverse transform process;

[0016] Calculate the reconstruction error and sparse regularization loss, update the model weights and biases through the back-propagation algorithm, and optimize the performance of the entire network structure;

[0017] Repeat the iteration until the preset stop iteration condition is met and the autoencoder model training is completed.

[0018] Furthermore, the linear transformation with dynamic sparse regularization is used to encode the network traffic data, and the encoding process is expressed as:

[0019] , ,

[0020] In the formula, To represent the frequency domain after the feature extraction is transformed, is the hidden layer feature, For the hidden layer features, is the encoder weight matrix, is the network traffic data after discrete cosine transform processing, is the bias term, is the ReLU activation function, is the sparsity loss function, For the The sparse regularization parameter for the iteration, is the dimension of the hidden layer, The number of samples entered for the current batch, is the first hyperparameter, For the The second adjustment hyperparameter of the iteration is is the third hyperparameter.

[0021] In a second aspect, the present invention provides a method for detecting abnormal network traffic, wherein the fusion framework includes a transfer learning module, a meta-learning module and a classification decision fusion module;

[0022] Obtain a source domain network traffic dataset and a target domain network traffic dataset;

[0023] Using a generative adversarial network algorithm based on feature sparsity constraints to increase the number of samples in the source domain network traffic data set to form a network traffic data sample set;

[0024] The network traffic data sample set is obtained by a meta-learning module to extract feature data I, and the feature data I is used to pre-train a network abnormal traffic detection model I, wherein the meta-learning module is a neural network model; at the same time, a transfer learning module extracts feature data II, and the feature data II is used to pre-train a network abnormal traffic detection model II, wherein the transfer learning module is an autoencoder model, wherein the transfer learning module is an autoencoder model obtained by the training method described in claim 1 or 2;

[0025] Fix the parameters of the models of the meta-learning module and the transfer learning module;

[0026] The target domain network traffic data set is divided into a query set and a support set, the pre-trained network abnormal traffic detection model I is fine-tuned by the query set and the support set, and at the same time, the pre-trained network abnormal traffic detection model II fine-tunes the support set; after the fine-tuning is completed, the network abnormal traffic detection model is obtained;

[0027] The network abnormal traffic detection model receives a target domain network traffic data set;

[0028] Output network abnormal traffic detection category.

[0029] Furthermore, the neural network model uses an ant colony optimization algorithm based on local adaptive adjustment to optimize parameters, thereby realizing neural network training, and the training steps include:

[0030] Initialize the neural network structure, the neural network includes 5 hidden layers, and use ant colony optimization to search for the optimal solution of the neural network parameters;

[0031] Initialize the individuals and search areas of the ant colony in the ant colony optimization algorithm. Each ant represents a potential solution of a neural network, including weight matrix and structure information.

[0032] Conduct ant colony optimization search and information sharing. In the ant colony optimization algorithm, ants search the weight space by simulating ant colony behavior and use pheromones to transmit the quality of the current solution.

[0033] Adaptive breeding strategy is used to adjust the search direction, dynamically adjust the search step size and balance the global and local search capabilities through adaptive breeding strategy;

[0034] After completing multiple rounds of search by ant colony optimization, the optimal solution is evaluated and selected, and the optimal solution is selected as the final structure of the neural network according to the fitness value;

[0035] Repeat the iteration until the preset stop iteration condition is met, that is, the neural network model training is completed.

[0036] Furthermore, the input and output relationship of the neural network is expressed as:

[0037]

[0038] In the formula, The neural network The output of the layer, The neural network The weight matrix of the layer, The neural network The output of the layer, The neural network The bias term of the layer, is the activation function of the neural network;

[0039] Furthermore, the method for adjusting the search step size is:

[0040]

[0041] In the formula, For the The ants are in Step size changes in round search; is the step size adjustment factor, is the step size variation parameter, is the distance between the current solution and the global optimal solution;

[0042] The step size change is controlled by a dynamic factor, and the calculation method of the step size adjustment factor is expressed as:

[0043]

[0044] In the formula, is the step size adjustment parameter; Control ratio for dynamic factors; It is the parameter for changing speed control.

[0045] Furthermore, the generative adversarial network algorithm training method based on feature sparsity constraints includes:

[0046] Initialize the network structure of the generator and discriminator of the generative adversarial network;

[0047] The generator optimizes the characteristics of generated samples by adopting sparsity constraints;

[0048] Train the discriminator, input the real network traffic data set and the network traffic data sample set generated by the generator, and improve the ability to distinguish between real samples and fake samples;

[0049] Adjust the generated target according to the category distribution offset based on the dynamic expansion mechanism;

[0050] Evaluate the authenticity and diversity of generated network traffic data and verify the quality of generated network traffic data;

[0051] Repeat the iteration until the evaluation function value of the generated network traffic data is less than the preset threshold, and the training of the generative adversarial network algorithm based on feature sparsity constraints is completed.

[0052] Furthermore, the sparsity constraint term is defined as:

[0053]

[0054] In the formula, is the sparsity loss function; To generate the sample Features, is the total number of features for generating samples; is the adaptive noise adjustment item; The weight of the adaptive noise adjustment term.

[0055] In a third aspect, the present invention provides a network abnormal traffic detection system, comprising:

[0056] A data acquisition module, wherein the data acquisition module is used to acquire a source domain network traffic data set and a target domain network traffic data set;

[0057] A data sample set module, wherein the data sample set module increases the number of samples in the source domain network traffic data set by using a generative adversarial network algorithm based on feature sparsity constraints to form a network traffic data sample set;

[0058] A fusion framework module, wherein the fusion framework module includes a meta-learning module and a transfer learning module;

[0059] The meta-learning module is used to extract feature data I from the network traffic data sample set through the meta-learning module, and the feature data I is used to pre-train the network abnormal traffic detection model I, wherein the meta-learning module is a neural network model;

[0060] The transfer learning module is used to extract feature data II from the network traffic data sample set through the transfer learning module, and the feature data II is used to pre-train the network abnormal traffic detection model II, wherein the transfer learning module is an autoencoder model; finally, the parameters of the models of the meta-learning module and the transfer learning module are fixed;

[0061] The network abnormal traffic detection model completion module divides the target domain network traffic data set into a query set and a support set, and the pre-trained network abnormal traffic detection model I is fine-tuned by the query set and the support set. At the same time, the pre-trained network abnormal traffic detection model II fine-tunes the support set; after the fine-tuning is completed, the network abnormal traffic detection model is obtained;

[0062] A data receiving module, wherein the data receiving module is used for the network abnormal traffic detection model to receive a target domain network traffic data set;

[0063] A data output module, wherein the data output module is used to output the network abnormal traffic detection category.

[0064] In a fourth aspect, the present invention provides a computer device, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of a method for detecting abnormal network traffic based on a fusion framework when executing the computer program.

[0065] In a fifth aspect, the present invention provides a storage medium having computer instructions stored thereon, wherein the computer instructions implement a method for detecting abnormal network traffic when executed.

[0066] Compared with the prior art, the present invention is innovative in the following aspects:

[0067] 1. In the task of abnormal network traffic detection, meta-learning and transfer learning are integrated to improve the effect of few-sample and cross-domain network anomaly detection. Meta-knowledge is extracted through the multi-task learning paradigm of meta-learning, and transfer learning is combined with knowledge transfer from the source domain to the target domain to solve the generalization problem of insufficient network traffic data samples and unlabeled data in the target domain.

[0068] 2. In the task of detecting abnormal network traffic, a generative adversarial network with feature sparsity constraints is used to alleviate the problems of imbalanced network traffic data and feature redundancy. The feature distribution of generated network traffic data is optimized through sparsity constraints, making it closer to the actual network traffic characteristics and improving the anomaly detection performance in a small sample environment.

[0069] 3. In the task of detecting abnormal network traffic, the autoencoder model based on discrete cosine transform is used to enhance the ability to extract network traffic data features. The network traffic data is mapped to the frequency space through discrete cosine transform, which enhances the expression ability of the main features and reduces the interference of redundant information on the detection task.

[0070] 4. In the task of abnormal network traffic detection, the ant colony optimization algorithm is used to improve the efficiency of neural network parameter adjustment and feature selection ability. By dynamically adjusting the step size and balancing global and local searches, the network parameters are optimized during the feature extraction of high-dimensional traffic network traffic data, making training more efficient and avoiding the problem of local optimal solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is a fusion framework diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0073] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0074] The network abnormal traffic detection model training framework proposed in the present invention is a meta-learning and transfer learning fusion framework. The meta-learning learns the meta-knowledge of different tasks through a multi-task learning paradigm, and uses this meta-knowledge to quickly learn in new tasks, so that it can achieve better learning results with a small number of network traffic data samples. In addition, the transfer learning relies on the knowledge transfer from the source domain network traffic data to the target domain network traffic data to solve the small sample and cross-domain learning problems, and needs to rely on a large amount of labeled network traffic data to pre-train the network abnormal traffic detection model.

[0075] The fusion framework proposed in this paper uses the advantages of meta-learning and transfer learning to improve the accuracy and generalization performance of the network abnormal traffic detection model. The fusion framework is as follows Figure 1 As shown:

[0076] The fusion framework includes a transfer learning part, a meta-learning part, and a classification decision fusion part. Among them, both the transfer learning part and the meta-learning part have their respective feature extraction modules and feature classification modules, and the two parts independently complete the training process before the final classification decision fusion.

[0077] The main idea of the fusion framework is to first divide the source domain network traffic data by batches and tasks as the basic units of the network traffic training data, and the task is the classification task of the present invention;

[0078] In the way of dividing network traffic data by task, the divided network traffic data is used as the basic input unit of meta-learning. Each task is further divided into a support set and a query set. The support set and the query set contain various categories of network traffic data. Under each category, there are various network traffic data samples, and the division ratio of the support set to the query set is 8:2.

[0079] The pre-training of transfer learning and meta-learning uses the same source domain network traffic data set and is carried out simultaneously;

[0080] Specifically, the collection sources of the source domain network traffic data set cover multiple network environments and network attack types, including:

[0081] Real production environment: Real network traffic data collected from various data centers, cloud platforms, and enterprise network environments;

[0082] Simulated attack environment: Network traffic data generated through simulated attack experiments (such as DDoS, SQL injection, XSS, etc.);

[0083] Public network traffic data set: Open source network traffic data sets such as KDDCup99 and NSL-KDD.

[0084] The network traffic data collection method adopts packet capture technology, and network packets are captured in real time through network monitoring devices (such as network traffic collectors, IDS / IPS systems, etc.). In one embodiment, the attributes of the constructed network traffic data include:

[0085] Ra is the source IP address (IPv4 address, indicating the IP address of the sender), Da is the destination IP address (IPv4 address, indicating the IP address of the receiver), Pa is the source port number (16-bit integer, indicating the source port), Qa is the destination port number (16-bit integer, indicating the destination port), Ta is the transmission protocol (integer, identifying the protocol type, such as TCP=1, UDP=2, etc.), La is the flow duration (floating timestamp, in seconds), Ba is the packet size (in bytes, indicating the total size of the network traffic), Ca is the number of packets (integer, indicating the number of packets contained in the traffic within a specific time), Fa is the flow direction (enumeration type, indicating whether the traffic is inbound or outbound), and Ha is the host feature of the traffic source (character, indicating the operating system or specific features of the host where the traffic comes from).

[0086] It should be noted that this embodiment is only intended to illustrate a format and type of network traffic data of the present invention. In actual applications, the attributes of network traffic data are usually more than 10 attributes, and the number of attributes of network traffic data may reach dozens or even hundreds.

[0087] Furthermore, the collected network traffic data is labeled. The labeling method of the present invention is manual labeling. In one embodiment, the labeling categories include: normal traffic (labeled as "0") and abnormal traffic (labeled as "1").

[0088] It should be noted that the target domain network traffic data has the same attributes as the source domain network traffic data, but the target domain network traffic data has no labeled categories.

[0089] In addition, it is understandable that in some cases, the collection, acquisition, labeling and preprocessing of network traffic training data are time-consuming and labor-intensive, and insufficient training samples can easily lead to poor model generalization ability and affect the accuracy of the model. The present invention adopts a generative adversarial network algorithm based on feature sparsity constraints. The feature sparsity constraints use sparsity constraints in the process of generating adversarial networks to ensure that network traffic data maintains diversity in the feature space and conforms to the actual traffic distribution, and alleviates the sample distribution offset problem caused by category imbalance, thereby improving the authenticity and category coverage of generated samples.

[0090] Specifically, the training process of the generative adversarial network algorithm based on feature sparsity constraints is as follows:

[0091] 1. Initialize the network structure of the generator and discriminator of the generative adversarial network. The generator receives the noise vector and part of the original network traffic data features as input, and outputs the generated traffic sample. Its output mapping function is defined as:

[0092] In the formula, It is the generated fake network traffic data; is the generator function; is the random noise of the generator input, Represents the parameters of the generator.

[0093] Moreover, the goal of the discriminator is to determine whether the input sample is real network traffic data, and its output is:

[0094] In the formula, Represents the discriminator's input The discriminant probability of is the discriminator function; It is the real network traffic data; and are the weight and bias of the discriminator, is the Sigmoid activation function.

[0095] 2. The generator optimizes the features of the generated samples by adopting sparsity constraints to make them more sparse to avoid redundant features and meaningless redundant features, so as to ensure that the generated network traffic data has actual abnormal traffic behavior. The sparsity constraint term is defined as:

[0096] In the formula, is the sparsity loss function; To generate the sample Features, is the total number of features for generating samples; is the adaptive noise adjustment item; The weight of the adaptive noise adjustment term. Preferably, Set to 0.3.

[0097] Furthermore, the loss function of the generator contains adversarial loss and sparsity constraint, which is expressed as:

[0098]

[0099] In the formula, express expectations; Indicates compliance with a specific distribution; Represents the distribution of real network traffic data; is the weight coefficient of the sparsity constraint; represents the parameters of the discriminator; is the distribution of noise; is the characteristic difference term, Control the weight of feature difference terms. Preferably, Set to 0.2.

[0100] Furthermore, the calculation method of the feature difference term is expressed as:

[0101] In the formula, is the feature vector of real network traffic data; Generate feature vectors of network traffic data for the generator; is the L2 norm.

[0102] Furthermore, the adaptive noise adjustment term represents the noise intensity when the generator generates network traffic data, so that the output of the generator is not only subject to feature constraints but also affected by the noise level, thereby enhancing the diversity and uncertainty of the samples generated in the network traffic. The calculation method is expressed as:

[0103]

[0104] In the formula, is the adaptive noise weight coefficient, which indicates the control of the noise intensity of each feature by the generator when generating samples; is the input noise dimensions, To generate the sample Features.

[0105] Furthermore, the adaptive noise weight coefficient is predicted by a preset auxiliary neural network, and the output of the network is mapped to interval, ensuring that the noise coefficient is within a reasonable range. The calculation method is expressed as:

[0106]

[0107] In the formula, is the output feature vector of the preset auxiliary neural network.

[0108] 3. The task of the discriminator is to evaluate the authenticity of each input network traffic data, that is, to determine whether it is real network traffic data. When training the discriminator, the input network traffic data includes real network traffic data and network traffic data samples generated by the generator. The discriminator gradually improves its ability to distinguish between real samples and fake samples by minimizing its discrimination loss. The loss function of the discriminator is:

[0109] In the formula, Generate the distribution of network traffic data respectively, is the first category weight, is the second category weight. Preferably, Set to 0.5, Set to 0.4.

[0110] 4. The dynamic expansion mechanism adjusts the generated target according to the category distribution offset. The role of the category distribution constraint in the discriminator's loss function is to minimize the difference between the generated samples and the real samples. The calculation method is expressed as:

[0111]

[0112] In the formula, is the KL divergence weight; It is The weight of each feature is a learnable parameter; It is a function that measures the difference between the class distributions of real and generated network traffic data; It is the category distribution of real network traffic data; To generate a category distribution of network traffic data. Preferably, Set to 0.2, Set to 0.2.

[0113] Furthermore, the difference function between the real and generated network traffic data category distribution is calculated by KL divergence, which is expressed as:

[0114]

[0115] In the formula, is the category distribution of real network traffic data corresponding to Features The corresponding first Features.

[0116] 5. Evaluate the authenticity and diversity of the generated network traffic data and verify the quality of the generated network traffic data. The evaluation function is:

[0117]

[0118] In the formula, To generate an evaluation function for network traffic data, ensure that the generated network traffic data is close to the real network traffic data and can improve the effectiveness of the actual network traffic data set.

[0119] 5. Repeat the above process until the evaluation function value of the generated network traffic data is less than the preset threshold, then stop the iteration, indicating that the training of the generative adversarial network is completed;

[0120] After the network traffic data expansion model training is completed, the trained network traffic data expansion model is used to increase the number of samples. In one embodiment, assuming that the original collected samples are 800, and the network traffic data expansion model expands and generates 200 samples, the expanded network traffic data set contains 1000 samples.

[0121] Furthermore, the fusion framework training is divided into two stages, the pre-training stage and the fine-tuning stage;

[0122] In the pre-training stage, for the feature extraction modules of transfer learning and meta-learning, network traffic data samples with batches as training units are input into the feature extraction module of transfer learning, and network traffic data samples with tasks as training units are input into the feature extraction module of meta-learning;

[0123] In the fine-tuning stage, the training set of the target domain network traffic data is also divided into batches and tasks as basic units. The training set of the target domain network traffic data is a set of network traffic data with labels in the target domain network traffic data. The pre-trained network anomaly traffic detection models of meta-learning and transfer learning are used to classify the target domain network traffic data, and the scores of the classification results of the two network anomaly traffic detection models are fused as the final classification result of the fusion framework.

[0124] In the fusion framework, meta-learning adopts a model weak correlation meta-learning strategy, the learning goal of which is to learn the optimal parameters that can quickly adapt to new tasks through a specific optimization algorithm. It is specifically divided into two layers of loop modes: inner loop and outer loop.

[0125] The inner loop uses a basic learner to extract features of a specific task, and the basic learner is a neural network model; the outer loop uses a meta-learning strategy to update the initialization parameters of the basic learner through gradient descent.

[0126] In the meta-learning strategy, each category is regarded as a task. The goal of meta-learning is to learn how to quickly distinguish different categories from these category tasks. The network abnormal traffic detection model should not only learn to complete the classification task itself, but also learn to extract knowledge from a small number of examples to more efficiently solve other unseen category tasks.

[0127] Specifically, the meta-learning strategy maps each sample in the support set and each sample in the query set to a high-dimensional feature space as the input of the neural network model, and randomly selects several categories during the training process to train the network abnormal traffic detection model.

[0128] After completing the task learning of all batches, the optimal initialization parameters are obtained. When encountering a new task, the support set of the new task is used to fine-tune the optimal initialization parameters to obtain the parameters that best suit the new task.

[0129] The neural network model uses an ant colony optimization algorithm based on local adaptive adjustment to optimize parameters, thereby realizing neural network training. The local adaptive adjustment improves optimization efficiency and global search performance by dynamically adjusting the step size and balancing the search direction in local search, and solves the local optimality and feature selection complexity problems of traditional feature extraction algorithms in high-dimensional network traffic data processing, thereby realizing accurate and rapid feature adaptive refinement.

[0130] 1. Initialize the neural network structure. The neural network contains 5 hidden layers. Ant colony optimization is used to search for the optimal solution of the neural network parameters. The input and output relationship of the neural network is expressed as:

[0131]

[0132] In the formula, The neural network The output of the layer, The neural network The weight matrix of the layer, The neural network The output of the layer, The neural network The bias term of the layer, is the activation function of the neural network.

[0133] Furthermore, the parameters of the neural network are initialized as variables to be optimized, which are determined by the ant colony optimization algorithm. In order to improve the adaptability of the neural network to network traffic data, the weight update strategy is adjusted through a dynamic adaptability adjustment function, which is expressed as:

[0134]

[0135] In the formula, For the The first iteration of the neural network The weight matrix of the layer; For the The first iteration of the neural network The weight matrix of the layer; The neural network Ant colony feedback adjustment factor of the layer.

[0136] Furthermore, the calculation method of the ant colony feedback adjustment factor is expressed as:

[0137]

[0138] In the formula, Adjust the amplitude for the ant colony, is the ant colony step length control factor, is the fitness difference between the current solution and the optimal solution. Preferably, Set to 0.3, Set to 0.2.

[0139] 2. Initialize the individuals and search areas of the ant colony in the ant colony optimization algorithm. Each ant represents a potential solution of a neural network, including weight matrix and structural information. The fitness of each individual is evaluated by the following method:

[0140]

[0141] In the formula, For the The fitness of an ant, is the loss function of the neural network, is the regularization factor of the neural network; is the regularization term of the neural network, indicating the complexity of the current solution; represents a potential solution of the neural network. Preferably, Set to 0.2.

[0142] Furthermore, the loss function measures the error of the neural network under the current solution, and the calculation method is expressed as:

[0143]

[0144] In the formula, The number of samples input to the neural network for the current batch; is the mean square error calculation function of the sample; For the The true labels of samples; For the neural network The feature vector of each sample after feature extraction is used to calculate the classification result using the preset Softmax function; is the regularization coefficient of the neural network, The neural network The absolute value sum of the weight matrices of the neurons in the layer; is the number of layers of the neural network. Preferably, Set to 0.2.

[0145] Furthermore, the regularization term of the neural network is used to dynamically adjust the weight sparsity, and the calculation method is expressed as:

[0146]

[0147] In the formula, is a dynamic adjustment factor and is a training parameter.

[0148] 3. Perform ant colony optimization search and information sharing. In the ant colony optimization algorithm, ants search the weight space by simulating ant colony behavior and use pheromones to transmit the quality of the current solution. The calculation method of path selection probability is expressed as:

[0149]

[0150] In the formula, For the The ants are in The probability of selection on the path; It is the set of ant paths in the ant colony optimization algorithm; For the The iteration The pheromone concentration on the path, For the Heuristic information on the path, is the first control parameter, is the second control parameter.

[0151] Furthermore, in the process of ant colony search, pheromone update is an important mechanism to guide the search direction. A nonlinear adjustment method is adopted in the exploration and development of the search space. The pheromone update rule is:

[0152]

[0153] In the formula, is the evaporation coefficient; For the The iteration pheromone concentration along the path; is the pheromone increment. Preferably, Set to 0.2.

[0154] Furthermore, the pheromone increment is updated based on the feedback of fitness, and the calculation method is expressed as:

[0155]

[0156] In the formula, is the pheromone intensity, is a smoothing factor. Preferably, Set to 2, Set to 0.01.

[0157] 4. Use adaptive breeding strategy to adjust the search direction. Use adaptive breeding strategy to dynamically adjust the search step size and balance the global and local search capabilities. The step size adjustment method is:

[0158]

[0159] In the formula, For the The ants are in Step size changes in round search; is the step size adjustment factor, is the step size variation parameter, is the distance between the current solution and the global optimal solution. Preferably, Set to 0.3.

[0160] Furthermore, the step size change is controlled by a dynamic factor, and the calculation method of the step size adjustment factor is expressed as:

[0161]

[0162] In the formula, is the step size adjustment parameter; Control ratio for dynamic factors; is a variable speed control parameter. Preferably, Set to 0.2, Set to 0.4, Set to 2.

[0163] 5. After completing multiple rounds of search by ant colony optimization, evaluate and select the optimal solution. The optimal solution is selected as the final structure of the neural network according to the fitness value, which is expressed as:

[0164]

[0165] In the formula, is the optimal solution, i.e. the optimal weight configuration of the neural network; For the The fitness of each ant; Indicates the index corresponding to the maximum value; For the The pheromone concentration of the iteration.

[0166] 6. Repeat the iteration until the preset stop iteration condition is met, indicating that the model training is completed. In one embodiment, the preset stop iteration condition is reaching a preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0167] In the fusion framework, unlike meta-learning based on the weak correlation strategy of the model, transfer learning can mine subtle features under the shallow features of network traffic data. The advantage of transfer learning is that it can make full use of the source domain network traffic data samples to pre-train the network abnormal traffic detection model, and specifically use the autoencoder model for migration.

[0168] The autoencoder model is an autoencoder algorithm based on discrete cosine transform, which includes two parts: an encoder and a decoder. The encoder uses discrete cosine transform to transform the input network traffic data into a new frequency space. The transformation can highlight the main features in the network traffic data while ignoring noise and redundant information. The decoder part is responsible for reconstructing the encoded network traffic data back to the original network traffic data space. The discrete cosine transform extracts the main characteristic frequencies of the network traffic data by using a discrete cosine transform strategy in the encoder, reduces the information loss of the network traffic data in the encoding process, and promotes the sparsity and generalization ability of the model.

[0169] Specifically, the training process of the autoencoder algorithm based on discrete cosine transform is as follows:

[0170] 1. Initialize the weights and biases of the autoencoder. In one embodiment, the initialization method is random initialization.

[0171] 2. Since the collected network traffic data has high-dimensional characteristics and the characteristics of the network traffic data have a certain degree of smoothness, that is, the feature vector after feature extraction can be regarded as the time domain feature before smoothing. The input network traffic data is mapped to the frequency space through discrete cosine transform. The encoder part receives the transformed network traffic data to generate feature coding, which is expressed as:

[0172]

[0173] In the formula, After feature extraction, The frequency domain representation after the dimensional feature transformation, is the nth sample of the characteristic network traffic data after the original feature extraction, is the length of network traffic data, is the normalized coefficient of discrete cosine transform.

[0174] In one embodiment, the calculation method of the normalized coefficient of the discrete cosine transform is expressed as:

[0175]

[0176]

[0177] In the formula, is the discriminant factor.

[0178] 3. In the feature encoding process, linear transformation with dynamic sparse regularization is used to encode network traffic data. The encoding process is expressed as:

[0179]

[0180]

[0181]

[0182] In the formula, To represent the frequency domain after the feature extraction is transformed, is the hidden layer feature, For the hidden layer features, is the encoder weight matrix, is the network traffic data after discrete cosine transform processing, is the bias term, is the ReLU activation function, is the sparsity loss function, For the The sparse regularization parameter for the iteration, is the dimension of the hidden layer, The number of samples entered for the current batch, is the first hyperparameter, For the The second adjustment hyperparameter of the iteration is, is the third control hyperparameter. Preferably, Set to 2.

[0183] In one embodiment, The calculation method of the sparse regularization parameter of the iteration is expressed as:

[0184]

[0185] Furthermore, in this embodiment, The calculation method of the second control hyperparameter of the iteration is expressed as:

[0186]

[0187] In the formula, is a sparse hyperparameter, Indicates Preferably, Set to 2.

[0188] 4. The decoder receives the encoded network traffic data and attempts to reconstruct the original input network traffic data through the inverse transformation process. The calculation method is expressed as:

[0189]

[0190]

[0191] In the formula, is the adjusted feature weight, is the reconstructed network traffic data sample, is the reconstructed frequency domain representation, is the transpose of the decoder weight matrix, is the bias term of the decoder.

[0192] In one embodiment, the adjustment of feature weights depends on information entropy. Specifically, for the reconstructed frequency domain representation, Features , calculate its information entropy The way is expressed as:

[0193]

[0194] In the formula, is the reconstructed frequency domain representation of Features, is the reconstructed frequency domain representation of The information entropy corresponding to the features is It is a feature Distribution probability function in network traffic dataset.

[0195] Furthermore, the feature weight The calculation method is expressed as:

[0196]

[0197] In the formula, is the reconstructed frequency domain representation of The information entropy corresponding to each feature.

[0198] 5. Calculate the reconstruction error and sparse regularization loss, update the weights and biases of the model through the back propagation algorithm, and optimize the performance of the entire network structure. The calculation method is expressed as:

[0199]

[0200] In the formula, the autoencoder based on discrete cosine transform is is the total loss function; in the autoencoder based on discrete cosine transform is a sparsity loss function, wherein the sparsity loss is calculated by the L1 norm of the network parameters; is the regularization strength parameter. Preferably, Set to 0.3.

[0201] 6. Repeat the above steps until the preset stop iteration condition is met, indicating that the model training is completed. In one embodiment, the preset stop iteration condition is reaching a preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0202] After the autoencoder training is completed, the intermediate feature vector output by the encoder is the feature extracted by the autoencoder.

[0203] After the pre-training, the parameters of the two feature extraction networks are fixed and transferred to the target domain network traffic data for fine-tuning. The target domain network traffic data is divided into a support set and a query set according to the meta-learning strategy. The network traffic data of the support set is used to fine-tune the transfer learning network anomaly traffic detection model. The fine-tuning strategy can use the gradient descent method.

[0204] After fine-tuning the network abnormal traffic detection model, the query set of the target domain network traffic data is input into the neural network model and the autoencoder model respectively to obtain their corresponding prediction scores. The prediction scores are then normalized using the Softmax function. Finally, the two prediction scores are fused and output as the final prediction result. The classification category is the category corresponding to the maximum prediction score.

[0205] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A model training method, wherein the model is an autoencoder model that meets the characteristics of network traffic data; The autoencoder model is an autoencoder algorithm model based on discrete cosine transform, characterized in that: The training method comprises: Initializing the weights and biases of the autoencoder, wherein the initialization method is random initialization; According to the high-dimensional characteristics and smoothness of network traffic data, the feature vector after feature extraction is regarded as the time domain feature before smoothing. The input network traffic data is mapped to the frequency space through discrete cosine transformation; In the feature encoding process, linear transformation with dynamic sparse regularization is used to encode network traffic data; receiving the encoded network traffic data through a decoder and attempting to reconstruct the original input network traffic data through an inverse transform process; Calculate the reconstruction error and sparse regularization loss, update the model weights and biases through the back-propagation algorithm, and optimize the performance of the entire network structure; Repeat the iteration until the preset stop iteration condition is met and the autoencoder model training is completed.

2. A model training method according to claim 1, characterized in that: The linear transformation with dynamic sparse regularization is used to encode the network traffic data. The encoding process is expressed as: , , In the formula, To represent the frequency domain after the feature extraction is transformed, is the hidden layer feature, For the hidden layer features, is the encoder weight matrix, is the network traffic data after discrete cosine transform processing, is the bias term, is the ReLU activation function, is the sparsity loss function, For the The sparse regularization parameter for the iteration, is the dimension of the hidden layer, The number of samples entered for the current batch, is the first hyperparameter, For the The second adjustment hyperparameter of the iteration is is the third hyperparameter.

3. A method for detecting abnormal network traffic, wherein the fusion framework includes a transfer learning module, a meta-learning module and a classification decision fusion module; characterized in that: include, Obtain a source domain network traffic dataset and a target domain network traffic dataset; Using a generative adversarial network algorithm based on feature sparsity constraints to increase the number of samples in the source domain network traffic data set to form a network traffic data sample set; The network traffic data sample set is obtained by a meta-learning module to extract feature data I, and the feature data I is used to pre-train a network abnormal traffic detection model I, wherein the meta-learning module is a neural network model; at the same time, a transfer learning module extracts feature data II, and the feature data II is used to pre-train a network abnormal traffic detection model II, wherein the transfer learning module is an autoencoder model, wherein the transfer learning module is an autoencoder model obtained by the training method described in claim 1 or 2; Fix the parameters of the models of the meta-learning module and the transfer learning module; The target domain network traffic data set is divided into a query set and a support set, the pre-trained network abnormal traffic detection model I is fine-tuned by the query set and the support set, and at the same time, the pre-trained network abnormal traffic detection model II fine-tunes the support set; after the fine-tuning is completed, the network abnormal traffic detection model is obtained; The network abnormal traffic detection model receives a target domain network traffic data set; Output network abnormal traffic detection category.

4. A method for detecting abnormal network traffic according to claim 3, characterized in that: The neural network model uses an ant colony optimization algorithm based on local adaptive adjustment to optimize parameters, thereby realizing neural network training. The training steps include: Initialize the neural network structure, the neural network includes 5 hidden layers, and use ant colony optimization to search for the optimal solution of the neural network parameters; Initialize the individuals and search areas of the ant colony in the ant colony optimization algorithm. Each ant represents a potential solution of a neural network, including weight matrix and structure information. Conduct ant colony optimization search and information sharing. In the ant colony optimization algorithm, ants search the weight space by simulating ant colony behavior and use pheromones to transmit the quality of the current solution. Adaptive breeding strategy is used to adjust the search direction, dynamically adjust the search step size and balance the global and local search capabilities through adaptive breeding strategy; After completing multiple rounds of search by ant colony optimization, the optimal solution is evaluated and selected, and the optimal solution is selected as the final structure of the neural network according to the fitness value; Repeat the iteration until the preset stop iteration condition is met, that is, the neural network model training is completed.

5. A method for detecting abnormal network traffic according to claim 4, characterized in that: The method for adjusting the search step size is: In the formula, For the The ants are in Step size changes in round search; is the step size adjustment factor, is the step size variation parameter, is the distance between the current solution and the global optimal solution; The step size change is controlled by a dynamic factor, and the calculation method of the step size adjustment factor is expressed as: In the formula, is the step size adjustment parameter; Control ratio for dynamic factors; It is the speed control parameter of the change.

6. A method for detecting abnormal network traffic according to claim 5, characterized in that: The generative adversarial network algorithm training method based on feature sparsity constraints includes: Initialize the network structure of the generator and discriminator of the generative adversarial network; The generator optimizes the characteristics of generated samples by adopting sparsity constraints; Train the discriminator, input the real network traffic data set and the network traffic data sample set generated by the generator, and improve the ability to distinguish between real samples and fake samples; Adjust the generated target according to the category distribution offset based on the dynamic expansion mechanism; Evaluate the authenticity and diversity of generated network traffic data and verify the quality of generated network traffic data; Repeat the iteration until the evaluation function value of the generated network traffic data is less than the preset threshold, and the training of the generative adversarial network algorithm based on feature sparsity constraints is completed.

7. A method for detecting abnormal network traffic according to claim 6, characterized in that: The sparsity constraint is defined as: In the formula, is the sparsity loss function; To generate the sample Features, is the total number of features for generating samples; is the adaptive noise adjustment item; The weight of the adaptive noise adjustment term.

8. A network abnormal traffic detection system, characterized in that: include: A data acquisition module, wherein the data acquisition module is used to acquire a source domain network traffic data set and a target domain network traffic data set; A data sample set module, wherein the data sample set module increases the number of samples in the source domain network traffic data set by using a generative adversarial network algorithm based on feature sparsity constraints to form a network traffic data sample set; A fusion framework module, wherein the fusion framework module includes a meta-learning module and a transfer learning module; The meta-learning module is used to extract feature data I from the network traffic data sample set through the meta-learning module, and the feature data I is used to pre-train the network abnormal traffic detection model I, wherein the meta-learning module is a neural network model; The transfer learning module is used to extract feature data II from the network traffic data sample set through the transfer learning module, and the feature data II is used to pre-train the network abnormal traffic detection model II, wherein the transfer learning module is an autoencoder model; finally, the parameters of the models of the meta-learning module and the transfer learning module are fixed; The network abnormal traffic detection model completion module divides the target domain network traffic data set into a query set and a support set, and the pre-trained network abnormal traffic detection model I is fine-tuned by the query set and the support set. At the same time, the pre-trained network abnormal traffic detection model II fine-tunes the support set; after the fine-tuning is completed, the network abnormal traffic detection model is obtained; A data receiving module, wherein the data receiving module is used for the network abnormal traffic detection model to receive a target domain network traffic data set; A data output module, wherein the data output module is used to output the network abnormal traffic detection category.

9. A computer device, a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the network abnormal traffic detection method as described in any one of claims 3 to 7 are implemented.

10. A storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, a method for detecting abnormal network traffic as described in any one of claims 3-7 is implemented.

Citation Information

Patent Citations

  • Network traffic processing method, related equipment and readable storage medium

    CN112749028B

  • Deep learning model compression method for network intrusion detection

    CN117669655A

  • Semi-supervised anomaly detection method based on transfer learning

    CN113128613A

  • Fund management method and system based on intelligent supervision

    CN118657610A

  • Fast adaptive model training method for novel attack detection

    CN119227781A