Balanced intrusion detection method and processing device based on deep learning
By using KAN network and time convolution-based generative adversarial network model in the deep learning-based balanced intrusion detection method, the network intrusion detection data set is preprocessed and balanced, which solves the bias problem of the model in the unbalanced data set and improves the detection performance of a few types of attacks.
Patent Information
- Application Number
- CN202510196581.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
When facing an imbalanced data set, the model prefers most class samples, thereby reducing the detection performance of a few class samples.
The intrusion detection model based on KAN network is adopted, and the balanced data set is obtained by pre-processing the original data set of network intrusion detection, including data cleaning, single-hot encoding, normalization processing and dimensionality reduction processing. At the same time, a generative adversarial network model based on time convolution is introduced, and the data set is balanced and pseudo-samples are generated to augment a few types of attack samples.
The performance and training speed of the balanced intrusion detection model based on deep learning are improved, the detection ability of a few types of attacks is enhanced, and the model's bias towards most types of samples is reduced.
Smart Images

Figure CN120050091A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information security intrusion detection, and specifically to a balanced intrusion detection method and processing device based on deep learning. Background Art
[0002] With the rapid evolution of network technology, network security issues have become one of the key factors restricting its development. As an important means of protecting network security, intrusion detection systems are attracting more and more attention from researchers. As an active defense strategy, intrusion detection can significantly enhance network security. Early intrusion detection methods mainly relied on traditional machine learning techniques. However, the scale, complexity, and dimension of data in the current network environment are far greater than before. Traditional machine learning methods are usually difficult to effectively classify such complex high-dimensional data, and their limited feature expression and learning capabilities make it difficult to cope with increasingly complex and changeable network attack methods.
[0003] The introduction and rapid development of deep learning theory has brought significant progress to the field of intrusion detection. As an important branch of machine learning, deep learning has attracted more and more researchers' attention with its excellent automatic feature learning ability and nonlinear expression ability. Some scholars have begun to apply deep learning technology to intrusion detection and have achieved remarkable results. Deep learning algorithms can effectively improve the performance of intrusion detection systems, and have significantly improved detection accuracy, detection speed, and generalization ability for unknown attacks.
[0004] Although deep learning has made good progress in the field of intrusion detection, the class imbalance problem is still an important factor that restricts the performance of deep learning-based balanced intrusion detection processing devices. An imbalanced data set refers to a skewed data distribution, that is, the number of samples in a certain class is much less than that in other classes. In intrusion detection scenarios, attack data is usually significantly less than normal data, which leads to the class imbalance problem. Class imbalance is a common problem in the field of intrusion detection, which causes the model to tend to favor majority class samples, thereby weakening the detection ability of minority class samples, and may even completely ignore certain types of attacks. This poses a major challenge to the practical application of intrusion detection systems. Because in actual network environments, attack events are usually low-probability events, but the harm they bring is very serious. Therefore, how to effectively solve the class imbalance problem and improve the detection performance of minority class attacks is a key issue that needs to be solved in the field of deep learning intrusion detection. Summary of the invention
[0005] To overcome the problem that in the current deep learning-based balanced intrusion detection method when facing an imbalanced dataset, the model tends to be biased towards majority-class samples, thereby reducing the detection performance for minority-class samples, this application provides a deep learning-based balanced intrusion detection method and processing device, adopting the following technical solutions:
[0006] In a first aspect, this application provides a deep learning-based balanced intrusion detection method, including:
[0007] Obtain the original network intrusion detection dataset, and preprocess the original network intrusion detection dataset to obtain a balanced network intrusion detection dataset;
[0008] Input the balanced network intrusion detection dataset into the intrusion detection model based on the KAN network, train the intrusion detection model based on the KAN network, and obtain the trained intrusion detection model based on the KAN network;
[0009] Input the network traffic dataset to be detected into the trained intrusion detection model based on the KAN network, and output the detection result.
[0010] Further, obtaining the original network intrusion detection dataset and preprocessing the original network intrusion detection dataset to obtain a balanced network intrusion detection dataset includes:
[0011] Perform data cleaning on the original network intrusion detection dataset;
[0012] Perform one-hot encoding on the pre-cleaned original network intrusion detection dataset to convert the categorical variables of the original data into digital form;
[0013] Perform normalization processing on the one-hot encoded original network intrusion detection dataset using the min-max method to obtain a standardized network intrusion detection dataset;
[0014] Perform dimensionality reduction processing on the standardized network intrusion detection dataset to remove redundant information in the standardized network intrusion detection dataset, and obtain a dimensionality-reduced network intrusion detection dataset;
[0015] Introduce a generative adversarial network model based on temporal convolution to perform balancing processing on the dimensionality-reduced network intrusion detection dataset, and obtain a balanced network intrusion detection dataset.
[0016] Further, performing dimensionality reduction processing on the standardized network intrusion detection dataset includes:
[0017] Perform dimensionality reduction processing on the standardized network intrusion detection dataset through a stacked denoising autoencoder to remove redundant information in the standardized network intrusion detection dataset, and obtain a dimensionality-reduced network intrusion detection dataset.
[0018] Furthermore, a generative adversarial network model based on temporal convolution is introduced to balance the network intrusion detection dimensionality reduction dataset, resulting in a network intrusion detection balanced dataset, including:
[0019] Partition the network intrusion detection dimensionality reduction dataset;
[0020] Construct a generator and a discriminator of the generative adversarial network model based on temporal convolution;
[0021] Construct a value function of the generative adversarial network model based on temporal convolution based on the generator and the discriminator;
[0022] Train the generative adversarial network model based on temporal convolution through the partitioned network intrusion detection dimensionality reduction dataset;
[0023] Construct a network intrusion detection balanced dataset through the generative adversarial network model based on temporal convolution.
[0024] Furthermore, construct a generator and a discriminator of the generative adversarial network model based on temporal convolution, including:
[0025] The constructed temporal convolutional network generator adopts a temporal convolutional network, introduces dilated convolution to extract features from the preset time steps, and finally the temporal convolutional network adopts residual connection;
[0026] During the training phase, the generator continuously updates the convolution kernel weights f and biases b through repeated forward and backward propagation to learn the latent distribution of the data; after training is completed, by inputting random noise, the generator can perform forward propagation to generate new data, that is, pseudo samples;
[0027] Adopt a convolutional neural network structure to construct the discriminator. The discriminator consists of multiple convolutional layers, and finally outputs a binary classification probability value p ∈ [0, 1] through a sigmoid layer, indicating whether the data is real;
[0028] Construct a value function of the generative adversarial network model based on temporal convolution. The generator learns the distribution of real data, the discriminator distinguishes real data and generated data, and the generator and the discriminator are alternately optimized.
[0029] Furthermore, the value function is as follows:
[0030]
[0031] In the formula, x represents real data, and G(z) represents the pseudo sample data generated by the generator G, where P data is the distribution of real data, and P z (z) is the distribution of fake data implicitly defined by G(z) and z ∼ P(z), where the noise z is sampled from the Gaussian noise distribution P.
[0032] Further, a network intrusion detection balanced dataset is constructed through a generative adversarial network model based on temporal convolution, including:
[0033] Input random Gaussian noise into the generator G of the generative adversarial network model based on temporal convolution, so that the generator G generates minority-class attack samples, thereby expanding the network intrusion detection dimensionality-reduced dataset, and finally obtaining a network intrusion detection balanced dataset.
[0034] In a second aspect, the present application also provides a balanced intrusion detection processing device based on deep learning, including:
[0035] A data acquisition module, configured to acquire a network intrusion detection original dataset, and preprocess the network intrusion detection original dataset to obtain a network intrusion detection balanced dataset;
[0036] A model training module, configured to input the network intrusion detection balanced dataset into an intrusion detection model based on the KAN network, train the intrusion detection model based on the KAN network, and obtain a trained intrusion detection model based on the KAN network;
[0037] A detection result output module, configured to input a network traffic dataset to be detected into the trained intrusion detection model based on the KAN network, and output a detection result.
[0038] In a third aspect, the present application provides an electronic device, including:
[0039] One or more processors; a memory; and one or more computer programs, where the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the device, cause the device to execute the method described in the first aspect.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when it runs on a computer, it causes the computer to execute the method described in the first aspect.
[0041] In a fifth aspect, the present application provides a computer program, which, when executed by a computer, is used to execute the method described in the first aspect.
[0042] In a possible design, the program in the fifth aspect can be stored in whole or in part on a storage medium packaged together with the processor, or can be stored in whole or in part on a memory not packaged together with the processor.
[0043] Compared with the prior art, the present application mainly has the following beneficial effects:
[0044] 1. By processing the missing values and outliers, performing one-hot encoding, and normalizing the original network intrusion detection dataset, this application can improve the performance and training speed of the balanced intrusion detection prediction model based on deep learning.
[0045] 2. This application uses a stacked denoising autoencoder to perform dimensionality reduction on the standardized network intrusion detection dataset. The stacked denoising autoencoder adds noise to the input data through multiple denoising autoencoders and then trains the model to reconstruct the original noise-free data, thereby improving the generalization ability and robustness of the model, and also enhancing the expression ability of data features.
[0046] 3. By introducing a temporal convolutional network as the generator of the generative adversarial network, this application can improve the ability of the generative adversarial network to capture the temporal features of intrusion detection data, thereby enhancing the authenticity of the pseudo-samples.
[0047] 4. By introducing a temporal convolutional network as the generator of the generative adversarial network, this application can improve the ability of the generative adversarial network to capture the temporal features of intrusion detection data, thereby enhancing the authenticity of the pseudo-samples.
[0048] 5. This application proposes a multi-class intrusion detection model based on the Kolmogorov-Arnold Networks theory. The model adopts a deeper and narrower network structure and combines residual connections to improve the expression ability and generalization performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is an exemplary system architecture diagram to which the embodiments of this application can be applied.
[0050] Figure 2 It is a flowchart of the balanced intrusion detection method based on deep learning in this application.
[0051] Figure 3 It is a schematic framework diagram of the balanced intrusion detection method based on deep learning in this application.
[0052] Figure 4 It is a flowchart of the operation of the balanced intrusion detection method based on deep learning in this application.
[0053] Figure 5 It is a flowchart of data dimensionality reduction of the balanced intrusion detection method based on deep learning in this application.
[0054] Figure 6 It is a flowchart of data balancing of the balanced intrusion detection method based on deep learning in this application.
[0055] Figure 7It is the architecture diagram of the data balance generator for the balanced intrusion detection method based on deep learning in this application.
[0056] Figure 8 It is the data balance training flow chart of the balanced intrusion detection method based on deep learning in this application.
[0057] Figure 9 It is the structure diagram of the KAN network prediction model for the balanced intrusion detection method based on deep learning in this application.
[0058] Figure 10 It is the schematic diagram of the device for the balanced intrusion detection method based on deep learning in this application.
[0059] Figure 11 It is the schematic diagram of the computer device of the embodiment of this application. Detailed implementation manners
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0061] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0062] In order to enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0063] Such as Figure 1As shown, the terminal devices represent various user devices connected to the network, such as servers, computers, laptops, and mobile phones. Terminal 1 and Terminal 2 are laptops and desktops, Terminal 3 can be a mobile device, and Terminal 4 and Terminal 5 can be file servers responsible for storing and processing data. Switches 1 and 2 play the role of connecting each terminal in the local area network and are responsible for forwarding data packets to the router. Common devices such as Cisco Catalyst and Juniper EX series switches are widely used in enterprise networks to improve network connection efficiency and bandwidth utilization. The router is responsible for managing the forwarding of network traffic and connecting the internal network to the external Internet. Typical routers include Cisco iSR and TP-Link Archer, which are responsible for communication between devices and ensure that data reaches the destination safely. The firewall, as an important line of defense for network security, blocks unauthorized access according to preset rules. Common firewall brands include Palo Alto Networks, Fortinet, and Cisco ASA, which not only protect the network from external threats but also manage the access rights of internal users.
[0064] The topology of the local area network is composed of terminals connected to the router through switches. The network traffic of all terminals first converges to the router through the switch, and the latter is responsible for data forwarding and communication. When data from the external network flows into the local area network, it first passes through the firewall to block unauthorized access and filter malicious traffic. Subsequently, the intrusion detection system monitors the network traffic in real time to detect abnormal behaviors and potential security threats. If an intrusion or abnormal situation is detected, the intrusion detection system can issue an alarm or take protective measures in a timely manner to ensure the security of the internal network and transmit the network traffic that has passed through the security screening to each device in the local area network. At the same time, the intrusion detection system can be deployed in any internal network composed of switches to protect the security of the internal network.
[0065] It should be noted that the balanced intrusion detection method based on deep learning provided in the embodiments of the present application is generally executed by terminal devices. Correspondingly, the balanced intrusion detection processing device based on deep learning is generally set between the gateway and the Internet or inside the subnet.
[0066] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0067] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to Figure 3 , the flowchart of the balanced intrusion detection method based on deep learning of the present application is shown in the figure. The overall framework of the balanced intrusion detection method based on deep learning is as shown inFigure 4 As shown in Figure 4 , the method includes the following steps:
[0068] Step S1: Obtain the original network intrusion detection dataset, and preprocess the original network intrusion detection dataset to obtain a balanced network intrusion detection dataset, including:
[0069] Step S11: Clean the original network intrusion detection dataset;
[0070] Step S12: Perform one-hot encoding on the cleaned original network intrusion detection dataset to convert the categorical variables of the original data into digital form;
[0071] Step S13: Normalize the one-hot encoded original network intrusion detection dataset using the min-max method to obtain a standardized network intrusion detection dataset;
[0072] Step S14: Perform dimensionality reduction on the standardized network intrusion detection dataset to remove redundant information in the standardized network intrusion detection dataset, and obtain a dimensionality-reduced network intrusion detection dataset;
[0073] Step S15: Introduce a generative adversarial network model based on temporal convolution to balance the dimensionality-reduced network intrusion detection dataset, and obtain a balanced network intrusion detection dataset.
[0074] In the embodiment of the present application, the method for obtaining the network intrusion detection dataset is as follows: Collect the network traffic data of security events from a security information and event management system, and splice and fuse it with an existing public network intrusion detection dataset to obtain the original network intrusion detection dataset.
[0075] In the embodiment of the present application, the formula used in the min-max method is as follows:
[0076]
[0077] In the above expression, x represents the attribute value of a certain feature, x max is the maximum value of this feature, x min is the minimum value of its feature, and x' is the result after normalizing x.
[0078] In the embodiment of the present application, the dimensionality reduction of the standardized network intrusion detection dataset is performed by stacking denoising autoencoders to remove redundant information in the standardized network intrusion detection dataset and obtain a dimensionality-reduced network intrusion detection dataset. The specific implementation method is as follows:
[0079] For the stacked denoising autoencoder composed of L layers of denoising autoencoders, the original data is input layer by layer from the first layer to the L layer, and noise is added to obtain corrupted data;
[0080] Use an encoding function to map the corrupted data to a low-dimensional representation;
[0081] Use a decoding function to map the low-dimensional representation back to the original data space;
[0082] Construct a mean squared error loss function to calculate the reconstruction error;
[0083] Use the stochastic gradient descent method to minimize the loss function and perform backpropagation;
[0084] Until all L-layer denoising autoencoders are trained, obtain the optimal weights, biases, and features of the last-layer denoising autoencoder model to get the stacked denoising autoencoder;
[0085] Reduce the dimension of the network intrusion detection standardized dataset through the stacked denoising autoencoder to obtain the network intrusion detection dimension-reduced dataset.
[0086] The specific data dimension reduction process is as Figure 5 shown, and the specific implementation steps are as follows:
[0087] For the stacked denoising autoencoder composed of L-layer denoising autoencoders, the training steps are as follows from the first layer to the Lth layer:
[0088] Step 51, input the original data and add noise to obtain the corrupted data. Specifically, for the ith-layer denoising autoencoder (i = 1, 2,..., L), if i is 1, the input is the original data x, and noise is added and a random function is used to obtain the corrupted data If i is greater than 1, the input is the hidden layer output h (i-1) , and noise is added and a random function is used to obtain
[0089] Step 52, use the encoding function to map the corrupted data to a low-dimensional representation. Specifically, use the encoding function f to map the input data to a low-dimensional representation. If i = 1, it is expressed as:
[0090]
[0091] If i > 1, it is expressed as:
[0092]
[0093] where h (i) is the hidden layer output of the ith-layer encoder; or h (i-1) is the input of the ith-layer encoder; W (i) represents the weight matrix of the ith-layer encoder; b (i)It represents the bias vector of the i-th layer encoder; f represents the sigmoid activation function.
[0094] Step 53, use the decoding function to map the low-dimensional representation back to the original data space, specifically: use the decoding function g to map the low-dimensional representation back to the original data space, which is expressed as:
[0095]
[0096] In the formula is the reconstructed data of the i-th layer decoder; W (i) is the weight matrix of the i-th layer decoder, usually the transpose of W (i) ; b (i) represents the bias vector of the i-th layer decoder; g represents the sigmoid activation function.
[0097] Step 54, construct the mean square error loss function and calculate the reconstruction error, specifically: construct the mean square error loss function and calculate the reconstruction error, and the loss function is expressed as:
[0098]
[0099] In the formula, N is the number of samples, is the original data of the m-th sample or the output of the previous layer of denoising autoencoder, is the reconstructed data of the m-th sample.
[0100] Step 55, use the stochastic gradient descent method to minimize the loss function and backpropagate; specifically: minimize the loss function and continuously update the weights W (i) 、W (i) and the biases b (i) 、b (i) to find the optimal solution.
[0101] Step 56, repeat steps 52 - 55 until all L layers of denoising autoencoders are trained, and obtain the optimal weights W (L) 、biases b (L) and features h (L) , and finally obtain the stacked denoising autoencoder.
[0102] Step 57, use the obtained stacked denoising autoencoder to perform dimensionality reduction processing on the network intrusion detection standardized dataset to obtain a network intrusion detection dimensionality reduction dataset.
[0103] In the embodiments of the present application, introduce the generative adversarial network model based on temporal convolution, perform balancing processing on the network intrusion detection dimensionality reduction dataset to obtain a network intrusion detection balanced dataset, and its operation process is as Figure 6 shown.
[0104] Step 61: Divide the network intrusion detection dimensionality reduction dataset into a training set and a test set.
[0105] Step 62: Construct a generator and a discriminator of a generative adversarial network model based on temporal convolution.
[0106] In the embodiment of the present application, constructing a generator and a discriminator of a generative adversarial network model based on temporal convolution includes:
[0107] Use a temporal convolutional network to construct a temporal convolutional network generator. A temporal convolutional network is a convolutional neural network architecture specifically designed for sequence prediction, especially suitable for processing time series data. The temporal convolutional network only uses the information of past time steps for convolution operations. In other words, the output value at time t is only convolved with the elements before time t in the previous layer. To capture the dependencies with longer time spans in network intrusion detection data, the constructed temporal convolutional network generator introduces dilated convolutions, enabling the model to effectively capture more distant historical information. Specifically, for an input sequence and a filter The dilated convolution operation acting on the sequence element s can be defined as:
[0108]
[0109] where k represents the filter size, d represents the dilation factor, * d represents the dilated convolution operation, and b is the bias. The dilation factor indicates the depth of capturing historical information in the convolution operation *. d Therefore, when d = 1, its formula is a normal convolution operation, and the larger the dilation factor d, * d the deeper the depth of capturing historical information. When dealing with long sequences, dilated convolutions can effectively extract features from more distant time steps. Finally, to avoid the problem of gradient vanishing or explosion during training, the temporal convolutional network generator adopts a residual connection, and the residual connection expression is:
[0110] O(x) = Φ(x + T(x))
[0111] where T(x) is the output of the convolutional layer, Φ represents the activation function, and x is the input of the convolutional layer. The activation function enhances the flow of gradients by adding the input x to the output T(x) of the convolutional layer.
[0112] During the training phase, the generator continuously updates the convolutional kernel weights f and the bias b through repeated forward and backward propagation, learning the latent distribution of the data. After training is completed, new data, that is, pseudo-samples, can be generated by passing random noise through the generator for forward propagation.
[0113] A discriminator is constructed using a convolutional neural network structure. The discriminator consists of multiple convolutional layers and finally outputs a binary classification probability value p ∈ [0, 1] through a sigmoid layer, indicating whether the data is real. The convolutional neural network can better capture data features, thus showing stronger discrimination ability.
[0114] Both the generator and the discriminator use the cross-entropy loss function for loss evaluation.
[0115] Step 63: Construct a value function of the generative adversarial network model based on temporal convolution using the generator and the discriminator.
[0116] In the embodiment of the present application, when constructing the value function of the generative adversarial network model based on temporal convolution, the main objective of the generator G is to learn the distribution of real data so that the generated data is as "real" as possible. The objective of the discriminator D is to distinguish real data and generated data as much as possible. Among them, the generator and the discriminator are alternately optimized, and the value function is as follows:
[0117]
[0118] In the formula, x represents real data, and G(z) represents the pseudo-sample data generated by the generator G, where P data is the distribution of real data, and P z (z) is the distribution of fake data implicitly defined by G(z) and z ∼ P(z), where the noise z is sampled from the Gaussian noise distribution P.
[0119] The generator G achieves its objective by minimizing the following loss function:
[0120]
[0121] In the formula, D(G(z)) is the judgment probability of the discriminator for the generated sample G(z).
[0122] The objective of the discriminator D is to maximize the judgment probability for real data and minimize the judgment probability for generated data at the same time. The discriminator loss function is:
[0123]
[0124] Finally, the ideal state of the generative adversarial network model based on temporal convolution is to enable the generator G to fit the distribution of the data, the generated pseudo-samples can deceive the discriminator D, and the discriminator D also has strong binary classification ability.
[0125] Step 64: Train the generative adversarial network model based on temporal convolution using the partitioned network intrusion detection dimensionality reduction dataset.
[0126] In the embodiments of the present application, the data balancing training process of the generative adversarial network model based on temporal convolution is as follows Figure 8 As shown, first, initialize the maximum number of iterations and the parameters of the generator and discriminator, aiming to improve the training speed of the generator and discriminator. Subsequently, input noise into the generator to make the generator generate pseudo-sample data. Input the pseudo-sample data and real data into the discriminator for training and update the discriminator loss. Then, use the loss of the discriminator for pseudo-samples to update the generator. Finally, check whether the maximum number of iterations is reached. If the maximum number of iterations is reached, stop training; otherwise, continue training.
[0127] Step 65, construct a network intrusion detection balanced dataset through the generative adversarial network model based on temporal convolution.
[0128] In the embodiments of the present application, constructing a network intrusion detection balanced dataset through the generative adversarial network model based on temporal convolution includes: inputting random Gaussian noise into the generator G of the generative adversarial network model based on temporal convolution, enabling the generator G to generate minority-class attack samples, thereby expanding the network intrusion detection dimensionality-reduced dataset, and finally obtaining the network intrusion detection balanced dataset.
[0129] Step S2, input the network intrusion detection balanced dataset into the intrusion detection model based on the KAN network, train the intrusion detection model based on the KAN network, and obtain the trained intrusion detection model based on the KAN network.
[0130] The structure of the intrusion detection model based on the KAN network in the above step S2 is as follows Figure 9 As shown, training the intrusion detection model based on the KAN network is specifically manifested as:
[0131] If f is a multivariate continuous function on a bounded domain, then f can be expressed as a finite combination of continuous functions, that is, for
[0132]
[0133] where and For the supervised learning task composed of input-output {x i , y i}, it is necessary to find the function f such that y of all data points i ≈ f(x i ). In the above equation, φ q,p and Φ q represent appropriate univariate functions. And the univariate function is composed of a locally learnable coefficient c i and a B-spline curve B i (x), that is
[0134]
[0135] Based on the Kolmogorov - Arnold Networks theory, expanded in terms of depth and breadth, the KAN network layer is determined by the following function matrix:
[0136] Φ = {φ q,p}, p = 1, 2, …, n in , q = 1, 2, …, n out
[0137] where n in is the input data dimension, and n out is the output data dimension. φ q,p represents a suitable unary function and has trainable parameters. The shape of the KAN network is represented by the integer array [n 1 , n 2 , ···, n L , where n i is the number of nodes in the i - th layer. The pre - activation value of the i - th neuron in the l - th layer is denoted as x l,i . Between the l - th layer and the (l + 1) - th layer, there are a total of n l n l+1 activation functions. The activation function connecting the i - th neuron in the l - th layer and the j - th neuron in the (l + 1) - th layer is denoted as φ l,j,i . The activation value obtained after the pre - activation value x l,i passes through the activation function φ l,j,i is expressed as:
[0138]
[0139] where i = 1, ..., n l ; j = 1, ..., n l+1 . Taking the sum of the activation values of the l - th layer neurons as the input of the j - th neuron in the (l + 1) - th layer, the pre - activation value of the j - th neuron in the (l + 1) - th layer is:
[0140]
[0141] Finally, the KAN network structure features result in the following weight - parameter - like matrix composed of activation functions:
[0142]
[0143] At this time, Φ l is the parameter function matrix of the l - th layer. Further, the KAN network can be expressed in the form of function composition, that is:
[0144]
[0145] Step S3: Input the network traffic dataset to be detected into the trained intrusion detection model based on the KAN network, and output the detection result.
[0146] In step S3, the network traffic dataset to be detected is used as the input of the model. This data may come from traffic data in the network, system logs in the server, behavioral data of applications, etc. At this time, the intrusion detection model based on the KAN network has been fully trained and can effectively identify different attack types and output highly accurate detection results.
[0147] Continue to refer to Figure 10 , the balanced intrusion detection processing device based on deep learning described in this embodiment includes:
[0148] A data acquisition module 1001, configured to acquire the original network intrusion detection dataset and preprocess the original network intrusion detection dataset to obtain a balanced network intrusion detection dataset;
[0149] A model training module 1002, configured to input the balanced network intrusion detection dataset into the intrusion detection model based on the KAN network, train the intrusion detection model based on the KAN network, and obtain the trained intrusion detection model based on the KAN network;
[0150] A detection result output module 1003, configured to input the network traffic dataset to be detected into the trained intrusion detection model based on the KAN network and output the detection result.
[0151] To solve the above technical problems, the embodiments of the present application also provide a computer device structure. For details, please refer to Figure 11 , Figure 11 is a schematic diagram of the computer device in this embodiment. The memory refers to a system for storing data for processing. In modern computers, this includes various types of storage. The memory usually interacts with the processor to read and write data. The processor, as the central processing unit of the computer, is responsible for executing instructions and performing calculations, and at the same time interacts with the memory and the input / output (I / O) system to process data and instructions. The interface refers to the input / output interfaces of the system, and these interfaces may include USB ports, network interfaces, and any connection form between the computer and external devices. The integrator or system bus refers to a communication system for transmitting data between internal components of the computer (such as between the processor, memory, and I / O devices). It connects all important parts of the system and constitutes the core pillar of the computer.
[0152] The schematic diagram shows the basic architecture of a typical computer system, where three main hardware components (storage, processor, and interface) are connected through a central integrator or bus system. The storage device is used to save the data processed by the central processor. There are mainly two types of storage in modern computers. In this embodiment, the memory is usually used to store the operating system and various application software installed on the computer device. The CPU extracts data from the memory to perform calculations, which is the main computing engine of the computer. It processes the instructions in the program by reading the program instructions in the memory, performing arithmetic and logical operations, and storing the results back in the memory. At the same time, the CPU also controls the flow of data between the interface and external devices, such as running the program code of the deep learning-based balanced intrusion detection method and processing device. The interface is the connection point between the computer and the external world. The I / O ports enable data to enter and leave the computer, allowing users to interact with the system and connecting it to peripheral devices. Typical examples include USB ports, network interfaces, and display ports.
[0153] The present application also provides another implementation manner, that is, to provide a non-volatile computer-readable storage medium storing a program of the deep learning-based balanced intrusion detection method and processing device, and the deep learning-based balanced intrusion detection method and processing device can be executed by at least one processor, so that the at least one processor executes the steps of the deep learning-based balanced intrusion detection method and processing device as described above.
[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device to execute the methods described in various embodiments of the present application.
[0155] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The accompanying drawings show the preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure made by using the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is equally within the scope of patent protection of this application.
Claims
1. A balanced intrusion detection method based on deep learning, characterized in that: include: Obtaining a network intrusion detection original data set, and preprocessing the network intrusion detection original data set to obtain a network intrusion detection balanced data set; Inputting the network intrusion detection balanced data set into the intrusion detection model based on the KAN network, training the intrusion detection model based on the KAN network, and obtaining the trained intrusion detection model based on the KAN network; The network traffic data set to be detected is input into the trained KAN network-based intrusion detection model, and the detection results are output.
2. The balanced intrusion detection method based on deep learning according to claim 1 is characterized in that: Obtain the original network intrusion detection data set, and preprocess the original network intrusion detection data set to obtain a balanced network intrusion detection data set, including: Performing data cleaning on the network intrusion detection original data set; Performing one-hot encoding on the cleaned original network intrusion detection data set to convert the categorical variables of the original data into digital form; The network intrusion detection original data set encoded by one-hot encoding is normalized by using the minimum-maximum method to obtain a standardized data set for network intrusion detection; Performing dimensionality reduction processing on the network intrusion detection standardized data set to remove redundant information in the network intrusion detection standardized data set to obtain a network intrusion detection reduced dimensionality data set; A generative adversarial network model based on temporal convolution is introduced to balance the network intrusion detection dimensionality reduction data set to obtain a network intrusion detection balanced data set.
3. The balanced intrusion detection method based on deep learning according to claim 2 is characterized in that: Performing dimensionality reduction processing on the network intrusion detection standardized data set includes: The network intrusion detection standardized data set is subjected to dimensionality reduction processing by stacking denoising autoencoders, redundant information in the network intrusion detection standardized data set is removed, and a network intrusion detection dimensionality reduction data set is obtained.
4. The balanced intrusion detection method based on deep learning according to claim 2 is characterized in that: A generative adversarial network model based on temporal convolution is introduced to balance the network intrusion detection dimensionality reduction data set to obtain a network intrusion detection balanced data set, including: Partitioning network intrusion detection dimensionality reduction datasets; Construct a generator and discriminator for a generative adversarial network model based on temporal convolution; Construct a value function of a generative adversarial network model based on temporal convolution based on the generator and discriminator; The temporal convolution-based generative adversarial network model is trained by partitioning the network intrusion detection dimensionality reduction dataset; Constructing a balanced dataset for network intrusion detection through a temporal convolution-based generative adversarial network model.
5. The balanced intrusion detection method based on deep learning according to claim 4 is characterized in that: Construct a generator and discriminator of a generative adversarial network model based on temporal convolution, including: A temporal convolutional network generator is constructed using a temporal convolutional network, and dilated convolution is introduced to extract features from preset time steps. Finally, the temporal convolutional network generator adopts residual connection; During the training phase, the generator continuously updates the convolution kernel weight f and bias b through repeated forward and backward propagation to learn the potential distribution of the data. After the training is completed, new data, i.e., pseudo samples, can be generated by passing random noise to the generator for forward propagation. The discriminator is constructed using a convolutional neural network structure. The discriminator consists of multiple convolutional layers, and finally outputs a binary probability value p∈[0,1] through a sigmoid layer, indicating whether the data is true or not. Construct a value function of a generative adversarial network model based on temporal convolution. The generator learns the distribution of real data, the discriminator distinguishes between real data and generated data, and the generator and discriminator are optimized alternately.
6. The balanced intrusion detection method based on deep learning according to claim 5 is characterized in that: The value function is as follows: Where x represents real data, and G(z) represents pseudo sample data generated by generator G, where P data is the distribution of real data, P z (z) is the distribution of false data implicitly defined by G(z) and z~P(z), where the noise z is sampled from the Gaussian noise distribution P.
7. The balanced intrusion detection method based on deep learning according to claim 4 is characterized in that: A balanced dataset for network intrusion detection is constructed through a generative adversarial network model based on temporal convolution, including: Random Gaussian noise is passed into the generator G of the generative adversarial network model based on temporal convolution, so that the generator G generates minority attack samples, thereby expanding the network intrusion detection dimensionality reduction data set, and finally obtaining a balanced data set for network intrusion detection.
8. A balanced intrusion detection processing device based on deep learning, characterized in that: include: The data acquisition module is used to acquire the original data set of network intrusion detection, and pre-process the original data set of network intrusion detection to acquire the balanced data set of network intrusion detection; The model training module is used to input the network intrusion detection balanced data set into the intrusion detection model based on the KAN network, train the intrusion detection model based on the KAN network, and obtain the trained intrusion detection model based on the KAN network; The detection result output module is used to input the network traffic data set to be detected into the trained KAN network-based intrusion detection model and output the detection result.