Network intrusion detection method and system

By building a deep confidence network and width learning model, combining particle swarm optimization and ridge regression algorithm, the traditional intrusion detection method has solved the shortcomings in accuracy, rate and resource utilization, and achieved efficient and accurate network intrusion detection.

CN120074914APending Publication Date: 2025-05-30HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217218.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the prior art faces massive and diverse network traffic data, the traditional intrusion detection method seems to be unscrupulous in terms of accuracy, rate and ability to identify unknown attacks. In addition, deep learning model training takes a long time, poor interpretability and high degree of dependence on hardware resources.

Method used

A deep confidence network containing multiple constrained Boltzmann machines is adopted to train and optimize weights through the particle swarm method, output dimensionality reduction data of high-order features, and train the width learning model through a pseudo-inverse algorithm of ridge regression for rapid classification detection.

Benefits of technology

It improves the accuracy, real-time and robustness of the intrusion detection system, shortens training time, reduces dependence on hardware resources, and enhances the detection rate of the model under diversified attack samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074914A_ABST
    Figure CN120074914A_ABST
Patent Text Reader

Abstract

The invention provides a network intrusion detection method and system, and belongs to the technical field of intrusion detection in a network, and the method comprises the steps: carrying out the data cleaning of original data, carrying out the standardization and data normalization of the cleaned data, standardizing the data according to a unified standard, and carrying out the normalization of the data; carrying out dimension reduction processing on the processed data set by using a deep belief network, optimizing the weight and bias of the deep belief network in combination with a particle swarm optimization algorithm, finding a global optimal solution by searching a whole solution space, and carrying out learning training on low-dimensional data by using a width learning network to obtain a global optimal solution; and putting a to-be-tested data set into the trained network model to obtain a final classification result. Through the intrusion detection model based on the combination of deep learning and width learning, the real-time performance and reliability of network attack detection are improved, the training time is shortened, and the detection rate of the model under diversified attack samples is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intrusion detection technology in the network, and specifically to a network intrusion detection method and system. Background Art

[0002] In recent years, with the rapid progress of the Internet and information technology, network intrusion behaviors have become increasingly frequent, and traditional intrusion detection methods based on feature matching and statistical analysis are difficult to effectively cope with. Especially when faced with a large amount of diverse network traffic data, traditional detection methods are unable to handle well in terms of accuracy, speed, and the ability to identify unknown attacks. Therefore, how to improve the accuracy, real-time performance, and robustness of intrusion detection systems has become a key research topic in the field of information security. Against this background, deep learning and wide learning technologies, with their powerful data feature extraction capabilities, parallel computing efficiency, and generalization performance, provide new solutions for network intrusion detection.

[0003] Deep learning has achieved remarkable results in many fields such as image recognition, natural language processing, and autonomous driving. Its core idea is to automatically extract hierarchical features of data through a multi-layer neural network architecture and can effectively capture complex patterns in the data. Deep learning models can perform deep feature extraction on raw data, capture temporal and spatial dependencies in the data, thereby improving the detection accuracy. However, deep learning models generally require a large amount of labeled data, and the time consumed for model training is relatively long, with a high degree of dependence on hardware resources, which will be restricted to a certain extent in the actual application process.

[0004] On the other hand, as an emerging shallow learning model, wide learning achieves fast learning through feature expansion and enhanced node parallelization structure, and can obtain good feature representation ability without multiple iterations. Compared with deep learning, wide learning has obvious advantages in training speed and is suitable for the rapid processing of large-scale data sets. In addition, the wide learning system shows good adaptability and scalability when dealing with high-dimensional data, and can efficiently add new features or new categories without changing the data structure. Therefore, by integrating wide learning and deep learning, leveraging the feature extraction ability of deep learning and the fast training advantages of wide learning, the efficiency and accuracy of intrusion detection systems can be effectively improved. Summary of the Invention

[0005] In order to overcome the problems of long training time, poor interpretability, and high dependence on hardware resources of current deep learning models, this application provides a network intrusion detection method and system, and adopts the following technical solutions:

[0006] In a first aspect, this application provides a network intrusion detection method, including:

[0007] Construct a deep belief network containing multiple restricted Boltzmann machines, and layer by layer learn and extract high-order features of the preprocessed historical network traffic data;

[0008] Use the particle swarm method to train the deep belief network and update and optimize the weights, and through iterative search for the global optimal weights, output the dimensionality-reduced data of the high-order features;

[0009] Train a wide learning model through the dimensionality-reduced data, solve the wide learning network weights through the pseudo-inverse algorithm of ridge regression, and obtain the trained wide learning model;

[0010] Input the network traffic data to be tested into the trained wide learning model for classification detection, and output the classification detection results.

[0011] Furthermore, layer by layer learn and extract high-order features of the preprocessed historical network traffic data, including:

[0012] The first-layer restricted Boltzmann machine learns low-level features from the input data;

[0013] Take the output of the hidden layer of the first-layer restricted Boltzmann machine as the input of the next-layer restricted Boltzmann machine, and push forward layer by layer in this way. Each layer gradually learns higher-order information of the features extracted by the previous layer;

[0014] As the number of layers increases, the deep belief network layer by layer extracts more complex features of the input data, realizes the high-order representation of the input data, and finally forms a feature expression that captures the deep patterns of the input data.

[0015] Furthermore, use the particle swarm method to train the deep belief network and update and optimize the weights, including:

[0016] The deep belief network is realized by training each restricted Boltzmann machine layer by layer. The CD algorithm is an effective approximation method for solving the maximum likelihood function of the restricted Boltzmann machine. The weights of each layer of the restricted Boltzmann machine are updated based on a preset formula, and the particle swarm optimization algorithm is introduced to find the initial global optimal weights.

[0017] Furthermore, introduce the particle swarm optimization algorithm to find the initial global optimal weights, including:

[0018] Initialize each particle as a weight combination, that is, the weights and bias parameters of the restricted Boltzmann machine;

[0019] Generate the reconstruction error of each particle through contrastive divergence, and take the reconstruction error as the fitness;

[0020] Update the velocity and position of each particle, so that the particle approaches its own historical best position and the global best position;

[0021] If the current fitness of the particle is better than the historical best fitness, update the historical best position of the particle; if the current fitness of the particle is better than the global best fitness, update the global best position of the particle;

[0022] Repeat the above steps until the maximum number of iterations is reached or the global optimal solution converges.

[0023] Furthermore, when the particle swarm converges to a global optimal solution, the obtained weight combination is used as the initial weight, and the fine-tuning CD algorithm performs local optimization near the initial solution obtained in the particle swarm;

[0024] In the fine-tuning stage, several iterations are performed through the CD algorithm to continue minimizing the reconstruction error, and the formula is as follows:

[0025] θ best ←θ best +η( <vh> data - <vh> recon )

[0026] where θ best is the optimal parameter of the current model, η represents the learning rate, v represents the activation value of the nodes in the visible layer, and h represents the activation value of the nodes in the hidden layer, <vh> data is the data expected value, <vh> recon It is to reconstruct the data expected value.

[0027] Furthermore, the width learning model is trained with the dimensionality-reduced data, and the weights of the width learning network are solved by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model, including:

[0028] Taking the dimensionality-reduced data as the input of the width learning model, the dimensionality-reduced data passes through the feature nodes and enhancement nodes in sequence to construct the width features required by the width learning model, including:

[0029] Width learning first maps the input data into feature nodes through a mapping function;

[0030] Expand and enhance the feature nodes, and map the feature nodes into enhancement nodes through a mapping function;

[0031] Merge the feature mapping nodes and the enhancement nodes to obtain the width learning input matrix of the width learning model.

[0032] Furthermore, the weights of the width learning network are solved by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model, including:

[0033] The size of the width learning network weights is restricted by the regularization of the pseudo-inverse algorithm of ridge regression, and the solution model is as follows:

[0034]

[0035] where u and v represent a kind of norm regularization, and σ 1 , σ 2 represent the powers of the loss function and the regularization term respectively, Y is the target value matrix, and λ is the regularization coefficient. When σ 1 = σ 2 = u = v = 2, the solution model is the ridge regression model.

[0036] On the second aspect, the present application also provides a network intrusion detection system, including:

[0037] A high-order feature extraction module, which is used to construct a deep belief network containing multiple restricted Boltzmann machines, and layer by layer learn and extract the high-order features of the preprocessed historical network traffic data;

[0038] A dimensionality-reduced data output module, which is used to train the deep belief network by using the particle swarm method and update and optimize the weights, and output the dimensionality-reduced data of the high-order features by iteratively searching for the global optimal weights;

[0039] A width learning model training module, which is used to train the width learning model with the dimensionality-reduced data, and solve the weights of the width learning network by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model;

[0040] A classification detection result output module, which is used to input the network traffic data to be tested into the trained width learning model for classification detection and output the classification detection result.

[0041] In a third aspect, the present application provides an electronic device, including:

[0042] One or several processors, a memory, and one or several computer programs, where the computer programs are stored in the memory and include instructions. When these instructions are executed by the device, the device can implement the method described in the first aspect.

[0043] In a fourth aspect, the present application presents a computer-readable storage medium, in which a computer program is stored. When this program runs on a computer, it can cause the computer to execute the method described in the first aspect.

[0044] In a fifth aspect, the present application provides a computer program, which is used to execute the method described in the first aspect when the computer program is executed by a computer.

[0045] In a possible design, the program in the fifth aspect can be stored completely or partially on a storage medium packaged together with the processor, or can be stored completely or partially in a memory separate from the processor.

[0046] Compared with the prior art, the present application mainly has the following beneficial effects:

[0047] 1. By adopting data cleaning technology, the present application performs operations such as missing value processing, outlier processing, and deduplication on historical network traffic data, and clears incomplete, duplicate, and incorrect data. Avoiding redundant information from affecting the efficiency and accuracy of model training. By the data cleaning step, the present application greatly improves the data quality and lays a solid foundation for the subsequent feature extraction and detection accuracy of the model.

[0048] 2. By standardizing and normalizing the data, the present application normalizes the data. Standardization converts the feature values into a normal distribution with a mean of 0 and a variance of 1, and normalization scales the data to a fixed range of [0,1] or [-1,1]. By normalizing the data, the present application can reduce the negative impact of feature value differences on model training, prevent some features from dominating the model weights due to excessive numerical values, make the model more stable and the training more efficient, thereby improving the accuracy and robustness of intrusion detection.

[0049] 3. This application constructs a deep belief network, which includes multiple restricted Boltzmann machines, and uses the particle swarm method to train and optimize its weight parameters. This method can effectively search for the global optimal weights, avoid local optima, thereby improving the stability and accuracy of the model, ensuring that the data after dimensionality reduction has high information density, and is beneficial to subsequent detection or classification tasks.

[0050] 4. This application obtains a trained wide learning model by constructing a wide learning model and introducing the pseudoinverse algorithm of ridge regression to solve the network weights. It ensures the stability and overfitting resistance of the model, and can achieve efficient and accurate detection while maintaining a low computational cost.

[0051] 5. This application improves the real-time performance and reliability of network attack detection through an intrusion detection model that combines deep learning and wide learning, shortens the training time, and effectively improves the detection rate of the model under diverse attack samples. In the intrusion detection model that combines deep learning and wide learning, it not only utilizes the advantages of deep learning models in feature extraction and pattern recognition, but also gives play to the advantages of wide learning in fast classification and model update, and can achieve a good balance between detection accuracy and speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is an exemplary system architecture diagram to which the embodiments of this application can be applied.

[0053] Figure 2 It is a flowchart of a network intrusion detection method of this application.

[0054] Figure 3 It is an overall framework diagram of a network intrusion detection method of this application.

[0055] Figure 4 It is a flowchart of the operation of a network intrusion detection method of this application.

[0056] Figure 5 It is a schematic diagram of the structure of the deep belief network model of a network intrusion detection method of this application.

[0057] Figure 6 It is a flowchart of the operation of the deep belief network model of a network intrusion detection method of this application.

[0058] Figure 7 It is a schematic diagram of the structure of the wide learning model of a network intrusion detection method of this application.

[0059] Figure 8 It is a flowchart of the operation of the wide learning model of a network intrusion detection method of this application.

[0060] Figure 9 It is a module relationship diagram of a network intrusion detection system of this application.

[0061] Figure 10 It is a schematic diagram of a computer device according to an embodiment of this application. Detailed implementation manners

[0062] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms mentioned herein are mainly used for the purpose of describing specific embodiments and do not limit their broader applicability. This application is not intended to limit its scope. The terms "including" and "having" and any variants thereof used in the description, claims, and related drawings of the application are intended to represent a non-exclusive inclusion relationship. In addition, terms such as "first" and "second" that appear in the description and claims are mainly used to distinguish different objects and are not used to indicate a specific order. The selection of these terms is to ensure a broad understanding and application of the content of the application.

[0063] As used herein, "embodiment" refers to an instance that combines specific features, structures, or characteristics, and these features may be included in at least one embodiment of this application. This phrase may not refer to the same embodiment at different positions in the description, nor is it necessarily mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0064] In order to enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0065] As Figure 1 shown, the figure shows a typical network structure based on a network intrusion detection system, which includes multiple network devices and terminals. First, the intrusion detection system is located at the entrance of the entire network, connected to the firewall, and connected to the switch through the router to monitor the internal and external traffic of the network. The external network is connected to the Internet through the firewall, and the firewall acts as a protection barrier for the network boundary at this position to filter the traffic entering the network. The intrusion detection system monitors the traffic from the external Internet and the internal network and real-time detects possible malicious activities or abnormal behaviors. The combined use of the firewall and the intrusion detection system can block known threats while real-time discovering potential intrusion behaviors, providing more complete network security protection for the administrator.

[0066] In the internal structure of the network, the intrusion detection system is connected to two switches (Switch 1 and Switch 2) through a router, and then connected to multiple terminal devices (Terminal 1 to Terminal 4). The switch is used to distribute data packets to ensure smooth communication and data transmission among internal devices. The terminal devices include user terminals such as mobile phones and computers, which are distributed on different ports of the switch. This distribution structure allows the intrusion detection system to monitor the communication of all terminal devices with the outside world to identify potential threats. Data flows from the terminal to the switch and finally passes through the router and firewall to be transmitted to the Internet. The intrusion detection system can intercept all incoming and outgoing network traffic through this topology and analyze whether there is any abnormal activity. This structure also facilitates quickly locating the specific terminal device when an intrusion behavior is detected. Through this layered network architecture and the mutual cooperation among devices, the entire system can efficiently monitor, analyze, and alarm network traffic in real time, improving the security and protection capabilities of the entire network.

[0067] It should be clear that the intrusion detection method based on the combination of deep learning and wide learning given in the embodiments of the present application is usually executed by terminal devices. Correspondingly, a network intrusion detection system is generally installed between the gateway and the Internet. It should be understood that Figure 1 the number of terminal devices, networks, and network devices here is only for illustrative purposes. According to the actual implementation requirements, any number of terminal devices, networks, and network devices can be available.

[0068] Continue to refer to Figure 2 , the figure shows a flowchart of a network intrusion detection method of the present application. The overall framework of a network intrusion detection method is as Figure 3 shown, and the operation flowchart is as Figure 4 shown. The method includes the following steps:

[0069] Step S1, construct a deep belief network containing multiple restricted Boltzmann machines, and layer by layer learn and extract the high-order features of the preprocessed historical network traffic data.

[0070] In the embodiments of the present application, the preprocessing of historical network traffic data includes: using data cleaning techniques to perform operations such as missing value processing, outlier processing, and deduplication on historical network traffic data, clearing incomplete, duplicate, and incorrect data to ensure data quality;

[0071] Using data cleaning techniques to perform operations such as missing value processing, outlier processing, and deduplication on historical network traffic data, clearing incomplete, duplicate, and incorrect data, which is specifically manifested as:

[0072] First, handle missing values, usually using interpolation, mean filling, or deletion to ensure data integrity. Second, for outliers, identify and filter out abnormal data by setting thresholds or based on statistical methods to reduce noise interference.

[0073] Normalize the historical network traffic data through standardization and normalization, scaling the feature values to a unified range, specifically as follows:

[0074] Network traffic data may include various features such as packet size, transmission frequency, protocol type, etc. The value ranges of these features may vary greatly. If standardization and normalization are not performed, during the training process of the model, more attention will be given to features with large value ranges, while ignoring features with small value ranges. First, calculate the mean and mean absolute error of each feature data, and the formula is as follows:

[0075]

[0076] where μ is the mean of this feature, σ is the standard deviation of this feature, and x i represents the attribute of the i-th record. Then standardize and measure each data record, and the formula is as follows:

[0077]

[0078] where x is the original feature value. After standardization like this, the mean of the feature is 0 and the standard deviation is 1, which can effectively convert the data into a form that conforms to the standard normal distribution, enabling the model to treat each feature more fairly when processing data. Then perform normalization on the standardized data, and the formula is as follows:

[0079]

[0080] where x min and x max are the minimum and maximum values of the feature respectively. In this way, the data is scaled to the specified range and the proportional relationship between the data is retained.

[0081] In the embodiment of the present application, construct a deep belief network containing multiple restricted Boltzmann machines, and layer by layer learn and extract high-order features of the preprocessed historical network traffic data as Figure 5 shown, specifically as follows:

[0082] The deep belief network is stacked by multiple restricted Boltzmann machines. Each restricted Boltzmann machine contains a visible layer and a hidden layer, and learns features from the input data through an unsupervised layer-by-layer training method.

[0083] During training, the first layer of restricted Boltzmann machines directly learns low-level features from the input data, such as the basic structure or local patterns of the data. Subsequently, the output of the hidden layer of the first layer of restricted Boltzmann machines is used as the input for the next layer of restricted Boltzmann machines, and so on layer by layer. Each layer gradually learns higher-order information about the features extracted by the previous layer. As the number of layers increases, the deep belief network can extract increasingly complex features layer by layer, thereby achieving a high-order representation of the data and ultimately forming a feature representation that can capture the deep patterns in the data. In this application, layer-by-layer training not only helps to avoid the vanishing gradient problem in deep neural networks but also enables more reasonable initialization of the network weights through unsupervised pre-training, improving the learning efficiency and convergence speed of the network.

[0084] Step S2, use the particle swarm method to train the deep belief network and update and optimize the weights, search for the global optimal weights through iteration, and output the dimensionality-reduced data of high-order features;

[0085] In the embodiment of this application, using the particle swarm method to train the deep belief network and update and optimize the weights is specifically manifested as:

[0086] Combining the particle swarm optimization algorithm and the deep belief network for dimensionality reduction of high-dimensional data is an effective strategy. In the deep belief network, the goal of training is to optimize the weights and biases by maximizing the log-likelihood function of the data. However, directly solving the gradient of the log-likelihood function is very complex, especially when the data dimension is high. Therefore, the deep belief network is usually implemented by training each restricted Boltzmann machine layer by layer, and the CD algorithm is an effective approximate method for solving the maximum likelihood function of the restricted Boltzmann machine. The weights of each layer of the restricted Boltzmann machine are updated by the following formula:

[0087] ΔW = η(<ν i h j > data -<ν i h j > recon )

[0088] where η is the learning rate, <v i h j > data represents the expectation under the data distribution, and <v i h j > recon represents the expectation under the reconstructed distribution.

[0089] In the embodiment of this application, the weight update formula of the restricted Boltzmann machine is actually an approximation of the gradient of the log-likelihood function with respect to the weights. By minimizing the reconstruction error (i.e., <v i h j > data -<v i h j > recon ) The CD algorithm indirectly maximizes the log-likelihood function. The fitness function of the particle swarm optimization algorithm is based on the reconstruction error of the CD algorithm. The reduction of the reconstruction error means that the model can better fit the data distribution, thus indirectly maximizing the log-likelihood function. The CD algorithm is essentially a local search method, suitable for fine-tuning weights, but it is not easy to find the global optimal solution. Therefore, it is of practical significance to introduce the particle swarm optimization algorithm to find the initial global optimal weights.

[0090] In the embodiment of the present application, the particle swarm optimization algorithm is introduced to find the initial global optimal weights, and the specific content is as follows:

[0091] First, initialize each particle as a weight combination θ p =(W p , b p , c p ), that is, the weights and bias parameters of the restricted Boltzmann machine. Among them, W p represents the weight matrix corresponding to the p-th particle, b p represents the bias vector of the visible layer corresponding to the p-th particle, and c p represents the bias vector of the hidden layer corresponding to the p-th particle. In this way, the position of each particle in the search space represents a possible initial weight.

[0092] Then, set the fitness function as the reconstruction error based on the CD algorithm to evaluate the weight quality of each particle. The CD algorithm generates the reconstruction error of each particle through contrastive divergence, and this reconstruction error is used as the fitness value. The formula is as follows:

[0093]

[0094] where i represents the index of the sample, N represents the total number of samples, v i is the original input data of the i-th sample, v i ′ is the output of the visible layer reconstructed through the current weight θ p of particle p, and ||v i -v i ′|| 2 is the reconstruction error of each sample.

[0095] Then, each particle updates its velocity and position to make the particle approach its own historical best position and the global best position. The formula is as follows:

[0096]

[0097] where t is the current iteration number, w is the inertia weight, c 1 is the individual learning factor that controls the movement intensity towards the individual's historical best position, c 2 is the social learning factor that controls the movement intensity towards the global best position, r 1 , r 2 is a random number is the velocity of particle p is the position of the particle (i.e., the weight θ p ), p best,p and g best are the historical best position and the global best position of this particle respectively.

[0098] Then, fitness evaluation and update are carried out. For each particle, the contrast divergence algorithm (i.e., CD algorithm) is used to calculate its reconstruction error, and it is used as the fitness value. If the current fitness of the particle is better than the historical best fitness, then update the historical best position of the particle; if the current fitness of the particle is better than the global best fitness, then update the global best position of the particle. Repeat the above steps until the maximum number of iterations is reached or the global best solution converges. When the maximum number of iterations is reached or the global best solution converges, the particle swarm converges to a global best solution.

[0099] Finally, fine-tune the CD algorithm for local optimization. When the particle swarm converges to a global best solution, the obtained weight combination is used as the initial weight. Next, use the CD algorithm to further fine-tune this weight to make it fit the data distribution more precisely. In the fine-tuning stage, several iterations are carried out through the CD algorithm to continue minimizing the reconstruction error. The formula is as follows:

[0100] θ best ←θ best +η( <vh> data - <vh> recon )

[0101] where θ best is the optimal parameter of the current model, η represents the learning rate, v represents the activation value of the nodes in the visible layer, and h represents the activation value of the nodes in the hidden layer. <vh> data is the data expected value, <vh> recon It is to reconstruct the data expected value.

[0102] In the fine-tuning stage, when the CD algorithm meets the preset number of iterations, the data output by the deep belief network is the dimensionality-reduced data of the high-order features.

[0103] In this way, the CD algorithm can effectively perform local optimization near the initial solution obtained in the particle swarm to further improve the model performance.

[0104] In the embodiment of the present application, the global optimization and the local optimization are optimization processes in two different stages. First, the global optimization method (particle swarm optimization) is used to find a suitable initial solution. At this time, the global optimization method searches in the entire solution space to avoid falling into the local optimum prematurely. Next, the local optimization method (i.e., the CD algorithm) is used to fine-tune this initial solution to further improve the accuracy and fitting effect. The CD algorithm no longer adjusts the weights on a large scale, but focuses on optimizing the local area near the current solution, so that the solution can better fit the data.

[0105] Step S3, construct a width learning model, and use the dimensionality-reduced data as the input to train the model, and solve the network weights through the pseudo-inverse algorithm of ridge regression as Figure 7 shown to obtain the trained width learning model;

[0106] In the embodiment of the present application, constructing a width learning model, using the dimensionality-reduced data as the input to train the model, and solving the network weights through the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model is specifically manifested as:

[0107] Use the dimensionality-reduced data as the model input, where m represents the number of samples and n represents the feature dimension. The input data will pass through the feature nodes and the enhancement nodes in sequence to construct the width features required by the model.

[0108] Width learning first maps the input data into feature nodes through a mapping function, and the formula is as follows:

[0109] Z f =σ f (XW f +b f ), f = 1, 2,..., p

[0110] where Z f is the output of the feature node, X is the input feature matrix, W f is the weight matrix of the feature node, b f is the bias, and σ f represents the activation function. Then, the feature nodes are expanded and enhanced, and the feature nodes are mapped into enhancement nodes through a mapping function, and the formula is as follows:

[0111] H e = ξ e (Z p W e + b e ), e = 1, 2, …, q

[0112] where H e is the output of the enhancement node, Z p is the mapped feature node matrix, W e is the weight matrix of the enhancement node, b e is the bias of the enhancement node, and ξ e is the activation function. Then, the feature mapping nodes and the enhancement nodes are combined to obtain the total input matrix of the width learning model, denoted as A = [Z f , H e .

[0113] Pseudo-inverse is the method used by width learning to solve the network weight W. The pseudo-inverse algorithm of ridge regression is a linear regression method with a regularization term, which is used to solve the problems of data overfitting and feature collinearity. Its goal is to limit the weight size through regularization to prevent the model from being too sensitive to noise. The solution model is as follows:

[0114]

[0115] where u and v represent a kind of norm regularization, σ 1 , σ 2 respectively represent the power of the loss function and the regularization term, Y is the target value matrix, and λ is the regularization coefficient. When σ 1 = σ 2 = u = v = 2, the above formula is actually a ridge regression model. Thus, we can obtain the formula for the weight matrix as follows:

[0116] W = (λI + AA T ) -1 A T Y

[0117] where I is the identity matrix and A is the width learning input matrix.

[0118] Step S4: Input the dataset to be tested into the trained width learning model for classification detection, and output the corresponding classification detection results;

[0119] Inputting the dataset to be tested into the trained width learning model for classification detection and outputting the corresponding classification detection results in step S4 are specifically manifested as follows:

[0120] First, the test data needs to go through the same preprocessing and feature expansion steps as the training data to ensure that the features of the input data are consistent with the structure of the model, that is, the above steps S1 and S2 preprocess the data and obtain the dimensionality-reduced data through a deep belief network. Then, the dimensionality-reduced data obtained after being processed by the above steps S1 and S2 is input into the feature nodes and enhancement nodes of the wide learning model. The wide learning model quickly classifies each input sample according to the trained weights and feature patterns. The wide learning model uses linear or non-linear activation functions to extract the feature patterns of the data, and then completes the classification prediction through the weight matrix obtained by the pseudo-inverse algorithm of ridge regression, and outputs the detection labels of each sample, such as "normal traffic" or "abnormal attack type", etc. These classification detection results help security personnel quickly identify potential intrusion behaviors and take corresponding measures, effectively improving the efficiency and accuracy of network intrusion detection.

[0121] Continue to refer to Figure 9 , a network intrusion detection system described in this embodiment includes:

[0122] A high-order feature extraction module 901, which is used to construct a deep belief network containing multiple restricted Boltzmann machines, and layer by layer learn and extract the high-order features of the preprocessed historical network traffic data;

[0123] A dimensionality-reduced data output module 902, which is used to train the deep belief network by using the particle swarm method and update and optimize the weights, and output the dimensionality-reduced data of the high-order features by iteratively searching for the global optimal weights;

[0124] A wide learning model training module 903, which is used to train the wide learning model with the dimensionality-reduced data, solve the wide learning network weights by the pseudo-inverse algorithm of ridge regression, and obtain the trained wide learning model;

[0125] A classification detection result output module 904, which is used to input the network traffic data to be tested into the trained wide learning model for classification detection and output the classification detection results.

[0126] To solve the above technical problems, the embodiments of the present application also provide a computer device structure. For details, please refer to Figure 10 , Figure 10 Schematic diagram of the computer device in this embodiment. In the field of network intrusion detection, the structure of a computer device usually consists of four parts: a network interface, an integrator, a processor, and a memory. They work together to capture, analyze, and store network data in order to detect potential intrusion behaviors. First of all, the network interface is the bridge between the device and the external network, responsible for capturing network data packets in transit and delivering them to the internal system. The importance of the network interface lies in its ability to collect traffic data in real time, enabling the intrusion detection system to obtain the latest network activity, including normal communication and potential threat traffic. The intrusion detection system relies on the efficient data acquisition ability of the network interface to ensure the real-time and accuracy of analysis, so as to detect abnormal traffic or malicious behaviors in a timely manner. The network interface sends the captured data to the integrator, which plays the role of a "dispatching center" for the internal data flow of the system, reasonably allocating the data to the memory and the processor for subsequent analysis and recording. The efficient operation of the integrator can ensure that the system can still maintain a smooth processing flow when facing a large amount of network traffic, avoiding data loss or detection lag caused by a sudden increase in traffic.

[0127] After the data flows through the integrator, it enters the processor and the memory for further processing. The processor is the analysis core of the entire system, undertaking the tasks of parsing and analyzing data, mainly using two detection methods: rule matching and behavior analysis. The rule matching method compares the characteristics of the traffic according to known attack signatures, which can quickly identify common attack behaviors, while the behavior analysis method captures unknown threats by detecting abnormal patterns in the traffic. After detecting suspicious traffic, the processor triggers an alarm mechanism to report potential intrusion behaviors in a timely manner, ensuring that the administrator can respond quickly. At the same time, the processor works in cooperation with the memory. The memory is used to store logs, historical traffic data, and rule libraries, helping the system to refer to historical data or attack feature libraries when conducting analysis. The memory can not only provide real-time data support, but also be used for subsequent event tracing and trend analysis. In the entire architecture, each module cooperates closely to achieve real-time monitoring and efficient processing of network traffic, ensuring that the intrusion detection system has fast and stable analysis capabilities.

[0128] This application also provides an implementation approach, that is, to provide a non-volatile computer-readable storage medium that stores a program for a network intrusion detection method. This network intrusion detection method can be executed by at least one processor, so that the at least one processor executes the steps of the network intrusion detection method as described above.

[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the embodiments can be implemented by software in combination with necessary general hardware platforms. Although hardware implementation is possible in some cases, generally speaking, the former is a more appropriate choice. Based on such understanding, the core essence of the technical solution of this application or the part contributing to the prior art can be presented in the form of a software product. This software product is stored in a storage medium and contains several instructions, enabling a terminal device to execute the methods of the various embodiments described in this application.

[0130] Obviously, the above embodiments are only part of the examples of this application, not all. The preferred embodiments of this application are shown in the drawings, but they should not be regarded as limiting the scope of the patent of this application. This application can be implemented in various forms. The purpose of providing these embodiments is to help readers understand the content of this application more clearly and comprehensively. Although this application has been described in detail in combination with the previous embodiments, those skilled in the art can still modify the technical solutions therein or perform equivalent replacements for some technical features. Any equivalent structure formed based on the content of the specification and drawings of this application and directly or indirectly applied to other related technical fields should be included within the scope of patent protection of this application.< / vh> < / vh> < / vh> < / vh> < / vh> < / vh> < / vh> < / vh>

Claims

1. A network intrusion detection method, characterized in that: include: Construct a deep belief network consisting of multiple restricted Boltzmann machines to learn and extract high-order features of preprocessed historical network traffic data layer by layer; The particle swarm method is used to train the deep belief network and update the optimization weights. The global optimal weights are searched iteratively to output the reduced-dimensional data of high-order features. The width learning model is trained by reducing the dimensionality of the data, and the weight of the width learning network is solved by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model; The network traffic data to be tested is input into the trained width learning model for classification detection, and the classification detection results are output.

2. The network intrusion detection method according to claim 1, characterized in that: Learn and extract high-level features of preprocessed historical network traffic data layer by layer, including: The first layer of restricted Boltzmann machines learns low-level features from the input data; The hidden layer output of the first layer of restricted Boltzmann machine is used as the input of the next layer of restricted Boltzmann machine, and each layer gradually learns the higher-order information of the features extracted by the previous layer. As the number of layers increases, the deep belief network extracts more complex features of the input data layer by layer, realizes high-order representation of the input data, and finally forms a feature expression that captures the deep pattern of the input data.

3. The network intrusion detection method according to claim 1, characterized in that: The particle swarm method is used to train the deep belief network and update the optimization weights, including: The deep belief network is implemented by training each restricted Boltzmann machine in layers. The CD algorithm is used to solve the effective approximation method of the maximum likelihood function of the restricted Boltzmann machine. The weights of each layer of the restricted Boltzmann machine are updated based on a preset formula, and the particle swarm optimization algorithm is introduced to find the initial global optimal weights.

4. The network intrusion detection method according to claim 3, characterized in that: The particle swarm optimization algorithm is introduced to find the initial global optimal weights, including: Initialize each particle to a weight combination, namely the weight and bias parameters of the restricted Boltzmann machine; Generate the reconstruction error of each particle by contrasting the divergence, and use the reconstruction error as the fitness; Update the speed and position of each particle to make the particle approach its own historical best position and the global best position; If the current fitness of the particle is better than the historical optimal fitness, the historical optimal position of the particle is updated; if the current fitness of the particle is better than the global optimal fitness, the global optimal position of the particle is updated; Repeat the above steps until the maximum number of iterations is reached or the global optimal solution converges.

5. The network intrusion detection method according to claim 4, characterized in that: When the particle swarm converges to a global optimal solution, the obtained weight combination is used as the initial weight, and the CD algorithm is fine-tuned to perform local optimization near the initial solution obtained in the particle swarm; In the fine-tuning stage, the CD algorithm is used for several iterations to continue minimizing the reconstruction error. The formula is as follows: i best ←θ best +n( <vh> data - <vh> recon )< / vh> < / vh> Among them, θ best is the optimal parameter of the current model, η represents the learning rate, v represents the activation value of the node in the visible layer, and h represents the activation value of the node in the hidden layer. <vh> data is the expected value of the data, <vh> recon is the expected value of the reconstructed data.< / vh> < / vh> 6. The network intrusion detection method according to claim 1, characterized in that: The width learning model is trained by reducing the dimensionality of the data, and the weight of the width learning network is solved by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model, including: The dimension-reduced data is used as the input of the width learning model. The dimension-reduced data passes through the feature nodes and enhancement nodes in turn to construct the width features required by the width learning model, including: Width learning first maps the input data into feature nodes through a mapping function; Expand and enhance the feature nodes, and map the feature nodes into enhanced nodes through mapping functions; The feature mapping node is merged with the enhancement node to obtain the width learning input matrix of the width learning model.

7. The network intrusion detection method according to claim 6, characterized in that: The width learning network weights are solved by the pseudo-inverse algorithm of ridge regression to obtain the trained width learning model, including: The network weights are learned by regularizing the width of the pseudo-inverse algorithm of ridge regression, and the solution model is as follows: Among them, u and v represent a norm regularization, σ1 and σ2 represent the powers of the loss function and the regularization term respectively, Y is the target value matrix, and λ is the regularization coefficient. When σ1=σ2=u=v=2, the solved model is the ridge regression model.

8. A network intrusion detection system, characterized in that: The system comprises: The high-order feature extraction module is used to build a deep belief network containing multiple restricted Boltzmann machines, and learn and extract high-order features of preprocessed historical network traffic data layer by layer; The dimension reduction data output module is used to train the deep belief network and update the optimization weights using the particle swarm method, and output the dimension reduction data of high-order features by iteratively searching for the global optimal weights; The width learning model training module is used to train the width learning model through dimension reduction data, solve the width learning network weights through the pseudo-inverse algorithm of ridge regression, and obtain the trained width learning model; The classification detection result output module is used to input the network traffic data to be tested into the trained width learning model for classification detection and output the classification detection results.