Balanced intrusion detection method and system based on width focus learning

By combining a width-based focus learning approach with generative adversarial networks and focus-focused loss mechanisms, the performance degradation of intrusion detection systems on imbalanced datasets is addressed, enabling effective identification of minority class attack samples and improving the detection rate.

CN119696853BActive Publication Date: 2025-12-09HENAN UNIVERSITY OF TECHNOLOGY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411781851.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-12-09
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

Existing intrusion detection systems suffer from severe performance degradation when faced with imbalanced datasets, failing to effectively identify different attack samples, especially minority attack samples.

Method used

We employ a width-focused learning approach, using a generative adversarial network to learn and train on the dataset, generating more diverse and realistic minority class samples. We also introduce a focus-focused loss mechanism to optimize the width-focused learning weight representation, constructing a width-focused learning model to address the class imbalance problem.

Benefits of technology

It improved the intrusion detection model's recognition rate for minority attack samples, enhanced the model's resistance to imbalance, and improved the detection rate and recognition ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119696853B_ABST
    Figure CN119696853B_ABST
Patent Text Reader

Abstract

The application provides a balanced intrusion detection method and system based on width focus learning, and belongs to the field of information security intrusion detection. The method comprises the following steps: obtaining a standardized data set, screening and filtering information features of the standardized data set through information variance and multicollinearity, eliminating invalid features and redundant features, and optimizing data quality; adopting a generative adversarial network to construct a balanced data set, introducing a deep belief network to optimize generated data, increasing data diversity and authenticity, and using a sigmoid model as a discriminator of the adversarial network. Finally, a width learning network is used to learn and train the balanced data set, a focal loss mechanism is introduced to strengthen the attention degree of the width learning model to different types of attack samples, so that the detection and recognition of different attack samples are achieved. The balanced intrusion detection model based on width focus learning improves the precision and reliability of intrusion detection attack detection, shortens the training time, and improves the detection rate of different attack samples.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information security intrusion detection, in particular to a balanced intrusion detection method and system based on width focus learning. BACKGROUND

[0002] The network intrusion detection system as an important barrier of computer security provides a powerful help for resisting network attacks. The intrusion detection system monitors and filters network behavior by analyzing host audit data and network traffic characteristics, identifies abnormal access and notifies the network administrator, so as to protect network information security. In recent years, due to the vigorous development of artificial intelligence, machine learning algorithm has better adaptability and generalization ability than traditional intrusion detection system, and can better learn automatically, so more and more researchers introduce machine learning technology into intrusion detection system, and have achieved certain results. However, the traditional machine learning model performs poorly when dealing with complex and high-dimensional data, and it is difficult to capture the deep patterns in the data.

[0003] With the rise of deep learning, researchers have found that it can automatically extract and fully learn complex features in data, achieve high accuracy and robustness, so deep learning technology is widely used in intrusion detection system. However, the deep learning model training process is computationally intensive, time-consuming and energy-consuming. In addition, the deep learning model usually lacks interpretability and is difficult to understand and explain, which has caused some problems to the performance of the intrusion detection system, and new technology development is urgently needed to solve it.

[0004] Recently, researchers have proposed a new kind of random neural network named width learning system, which is based on shallow network, unlike extending to deep level, it does not need to use gradient guided update weight, but based on pseudo-inverse operation to solve weight, so the network structure size is small and the calculation speed is faster, researchers introduce width learning into intrusion detection system and achieve good results. However, the data set of intrusion detection is often seriously unbalanced, and the traditional width learning model cannot handle the unbalanced data set well, resulting in poor detection performance of minority class attacks, which is often fatal to the original intention of intrusion detection system, and appropriate methods must be found to alleviate this situation. SUMMARY

[0005] In order to overcome the problem that the current intrusion detection method has serious detection performance decline when facing unbalanced data set and cannot effectively identify different attack samples, the present application provides a balanced intrusion detection method and system based on width focus learning, which adopts the following technical scheme:

[0006] In the first aspect, the present application provides a balanced intrusion detection method based on width focus learning, comprising:

[0007] obtaining a standardized dataset;

[0008] obtaining a dimension-reduced dataset based on the standardized dataset;

[0009] introducing a deep belief network as a generator of the generative adversarial network to generate generated data more fitting the real data, and selecting a sigmoid discriminator to distinguish the authenticity of the data, to jointly construct the generative adversarial network;

[0010] learning and training the dimension-reduced dataset based on the generative adversarial network, performing data enhancement on the minority class samples to generate more minority class samples with diversity and authenticity, so that the class distribution in the dimension-reduced dataset tends to be balanced, and obtaining a balanced dataset;

[0011] introducing a focal loss mechanism into the width learning model to optimize the expression of the width learning weight, and constructing a width focal learning model;

[0012] taking the balanced dataset as the input of the width focal learning model, training the width focal learning model, and solving the optimal weight through pseudo-inverse calculation and adjusting the focal factor, to obtain the trained width focal learning model;

[0013] inputting the dataset samples to be detected into the trained width focal learning model for classification and detection, and the width focal learning model effectively identifies samples of different attack types through the learned feature representation and the focal loss mechanism, and outputs corresponding detection results.

[0014] Further, the obtaining of the dimension-reduced dataset based on the standardized dataset comprises: filtering and screening the standardized dataset through information variance and multicollinearity, eliminating low-variance features with small contribution, and removing redundant features with large interference by calculating a correlation matrix, to obtain the dimension-reduced dataset.

[0015] Further, the filtering and screening of the standardized dataset through information variance and multicollinearity, and the elimination of low-variance features with small contribution, and the removal of redundant features with large interference by calculating a correlation matrix to obtain the dimension-reduced dataset specifically comprise:

[0016] In the entire dimension-reduced data processing process, the standardized dataset is input in the form of a two-dimensional matrix containing multiple samples and features, and the original standardized dataset is denoted as X, which contains n samples and p features;

[0017] calculating the standard deviation of each feature, screening and removing features with a standard deviation lower than a preset threshold, and obtaining a screened dataset;

[0018] The dataset is input as a matrix. The Pearson correlation coefficient between each feature in the filtered dataset is calculated, generating a feature-to-feature correlation matrix R, where each value in the correlation matrix R is R0. ij This indicates the correlation between feature i and feature j;

[0019] The upper triangular mask is used to cover the upper half of the correlation matrix. The correlation data of the lower triangular part is filtered out by the mask operation, and features with correlation coefficients exceeding the set threshold are removed to obtain data X′.

[0020] The data stream is transformed from the original standardized dataset X into a simplified dataset X′, and the simplified dataset X′ is used as a dimensionality-reduced dataset.

[0021] Furthermore, the generative adversarial network uses BinaryCrossEntropyLoss for loss assessment.

[0022] Furthermore, a deep belief network is introduced as the generator of the generative adversarial network to produce generated data that better fits the real data. Specifically, this is manifested as follows:

[0023] During the training phase, the deep belief network is trained layer by layer in unsupervised manner using a multi-layer restricted Boltzmann machine, and the weights of each layer are updated according to the following formula:

[0024] ΔW=η( <v i h j > data - <v i h j > model )

[0025] Where η is the learning rate. <v i h j > data and <v i h j > model Let represent the joint probability expectation under the data sample and the model-generated sample, respectively. By maximizing the difference between the two, the Restricted Boltzmann Machine can learn the data features layer by layer.

[0026] After training is complete, the deep belief network generates new data through the forward propagation process, and the hidden units h of each layer... j From the visible unit v of the previous layer i And the activation of the weight matrix W, the specific formula is as follows:

[0027]

[0028] Where σ(·) is the activation function, W ij Let b be the weight matrix. jis the bias term; when generating new data, the DBN generates new visible unit data by backpropagation from the activation of the top hidden layer, layer by layer, through back sampling;

[0029] visible unit v i The reconstructed value of the visible unit v j is calculated by backpropagation with the weight of the hidden unit h

[0030]

[0031] where c i is the bias term of the visible unit; through repeated forward and backpropagation, the DBN can generate new samples conforming to the distribution of the training data; each time the sampling updates the visible unit v i and the hidden unit h j , until a sample similar to the real data distribution is generated; finally, the new data v' output by the DBN is obtained through multiple samplings, through these steps, the DBN learns the latent distribution of the data and can generate new data with similar distribution.

[0032] Further, the focal loss mechanism is introduced into the width learning model, which is specifically as follows:

[0033] The width learning first maps the input data into a feature node matrix through a mapping function, and the process of generating the mapping node from the input data is as follows:

[0034]

[0035] where: is an activation function, X is an input feature matrix, is the i-th group of feature mapping weight matrix, is the i-th group of feature mapping bias matrix, N is the total number of samples, and q is the corresponding number of feature nodes; all mapping matrices are integrated into a total mapping feature node matrix Z, and the process of generating enhanced nodes is as follows:

[0036]

[0037] where: ξ j is an activation function, Z is a mapping feature node matrix, is the j-th group of feature enhancement weight matrix, is the j-th group of feature enhancement bias matrix, N is the total number of samples, and r is the number of enhanced nodes corresponding to each group of enhanced transformation;

[0038] Then all enhanced feature node matrices are integrated; finally, the mapping nodes and the enhanced nodes are combined to obtain the total input matrix of the width learning model, denoted as A[Z|H];

[0039] The core idea of the focal loss mechanism is that, on the basis of the standard cross-entropy loss, a modulation factor and a weight term are introduced to enable the loss function to dynamically adjust the weight of the sample, which is as follows:

[0040] L focal-loss =-a(1-p t ) y log(p t )

[0041] Where p t is the predicted probability of the model for the real class t; a is the weighting coefficient for different class samples, used to balance the class imbalance; y is the modulation factor parameter, used to adjust the weight of difficult and easy samples;

[0042] In order to minimize the error between the predicted value and the real value of the width focal learning model, and as much as possible to make the probability of positive samples larger, the optimization objective of the width focal learning model can be obtained as:

[0043] min:L(W)=λ f (1-AW) Y log(AW)+(AW-Y) 2 +λW 2

[0044] Where λ f is the focal penalty factor, A is the width learning input matrix, W is the weight of the width focal system, Y is the real label, and λ is the width learning penalty factor.

[0045] In a second aspect, the application also provides a balanced intrusion detection system based on width focal learning, comprising:

[0046] A standardized data set acquisition module is used to acquire a standardized data set;

[0047] A dimensionality reduction data set acquisition module is used to acquire a dimensionality reduction data set based on the standardized data set;

[0048] A generative adversarial network module is used to introduce a deep belief network as a generator of the generative adversarial network to generate realistic data samples, and select a sigmoid discriminator to distinguish the authenticity of the data, and jointly construct the generative adversarial network;

[0049] A balanced data set acquisition module is used to learn and train the dimensionality reduction data set based on the generative adversarial network, perform data augmentation on the minority class samples, generate more minority class samples with diversity and authenticity, make the class distribution in the dimensionality reduction data set tend to be balanced, and obtain a balanced data set;

[0050] The width focal learning model construction module is configured to introduce a focal loss mechanism into the width learning model, optimize the width learning weight expression, and construct a width focal learning model.

[0051] The width focal learning model training module is configured to take the balanced data set as an input of the width focal learning model, train the width focal learning model, and solve the optimal weight by pseudo-inverse calculation and focal factor adjustment to obtain the trained width focal learning model.

[0052] The detection result output module is configured to input the data set sample to be detected into the trained width focal learning model for classification and detection, and the width focal learning model effectively identifies samples of different attack types by using the learned feature representation and the focal loss mechanism, and outputs corresponding detection results.

[0053] In a third aspect, the present application provides an electronic device, comprising:

[0054] One or more processors, memories and computer programs stored therein, containing instructions, enabling the device to perform the method of the first aspect.

[0055] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program that, when executed, causes a computer to perform the method of the first aspect.

[0056] In a fifth aspect, the present application provides a computer program for performing the method of the first aspect when the program is executed by a computer.

[0057] In one possible design, the program of the fifth aspect can be stored entirely or partially on a storage medium packaged with the processor, or in a memory independent of the processor.

[0058] Compared with existing algorithms and technologies, the present application has the following advantages:

[0059] 1. The present application filters and filters the standardized data set by information variance and multicollinearity, eliminates low-variance features that contribute less to the model, and removes redundant features that interfere more with the model by calculating the correlation matrix, thereby obtaining a simplified and efficient reduced dimension data set, thereby improving the detection accuracy and reliability of intrusion detection.

[0060] 2. The present application introduces a deep belief network that efficiently represents complex multi-layer features to generate more realistic minority class samples, and selects a targeted sigmoid discriminator to jointly construct a generative adversarial network structure, which can improve the authenticity and anti-interference of the generated balanced data set and enhance the anti-unbalance ability of the intrusion detection model.

[0061] 3. The application introduces the focal point loss mechanism into the width learning model, optimizes the width learning weight expression, improves the weight and attention of the minority class, thereby enhancing the recognition rate of the intrusion detection model for minority class attack samples and improving the detection rate.

[0062] 4. The application solves the optimal weight by pseudo-inverse calculation and adjustment of the focal factor, dynamically adjusts the sensitivity and adaptability of the intrusion detection model to different data, enhances the anti-interference ability of the model to different sample types, and reduces the training time of the model.

[0063] 5. In the balanced intrusion detection model of width focal learning, the end-to-end structure of balanced data, input samples, feature nodes, enhanced nodes and model prediction can be easily realized, so that the model converges to the global optimal solution, and the detection performance and recognition rate of the model to unbalanced data set are maximized. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 An exemplary system architecture diagram to which embodiments of the application can be applied.

[0065] Figure 2 is a flowchart of the balanced intrusion detection method based on width focal learning of the application.

[0066] Figure 3 is a frame diagram of the balanced intrusion detection method based on width focal learning of the application.

[0067] Figure 4 is a flowchart of the balanced intrusion detection method based on width focal learning of the application.

[0068] Figure 5 is a deep belief generation adversarial network diagram of the balanced intrusion detection method based on width focal learning of the application.

[0069] Figure 6 is a discriminator training flowchart of the balanced intrusion detection method based on width focal learning of the application.

[0070] Figure 7 is a generator training flowchart of the balanced intrusion detection method based on width focal learning of the application.

[0071] Figure 8 is a width focal learning structure diagram of the balanced intrusion detection method based on width focal learning of the application.

[0072] Figure 9 is a width focal learning training flowchart of the balanced intrusion detection method based on width focal learning of the application.

[0073] Figure 10is a module relationship diagram of a balanced intrusion detection method system based on width focus learning of the present application.

[0074] Figure 11 is a schematic diagram of a computer device of an embodiment of the present application. DETAILED DESCRIPTION

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art; the terminology used in the specification herein is for describing specific embodiments only and is not intended to be limiting of the present application; "comprising," "having," and variations thereof herein are intended to mean a non-exclusive inclusion; for example, a process, method, article, or apparatus that comprises a variety of components or steps does not include only those components or steps that are recited but also other components or steps that are not expressly recited. Furthermore, the use of "first," "second," and the like does not imply a particular order but is used for naming purposes only.

[0076] To help those skilled in the art better understand the schemes of the present application, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings.

[0077] As shown in Figure 1 , terminal devices represent various user devices connected to the network, such as servers, computers, laptops, and mobile phones. Terminal 1 and terminal 2 can be file servers responsible for storing and processing data, terminal 3 can be a mobile device, and terminal 4 and terminal 5 can be office desktops or laptops. Switch 1 and switch 2 are responsible for connecting various terminals in the local area network and forwarding data packets to the router. Common examples include Cisco Catalyst or Juniper EX Series switches, which are widely used in enterprise networks to improve network connection efficiency and bandwidth utilization. The router is responsible for managing the forwarding of network traffic and connecting the internal network to the external Internet. Common routers include Cisco ISR or TP-Link Archer, which manage communication between devices and ensure that data can safely reach its destination. The firewall is an important security line of defense for the network, responsible for blocking unauthorized access according to preset rules. Palo Alto Networks, Fortinet, or Cisco ASA are commonly used to protect the network from external threats while managing access permissions for internal users.

[0078] These terminals are connected to the router through the switch, forming a local area network topology. The network traffic of all terminals is first aggregated to the router, which manages data forwarding and communication. After data flows from the external network to the local area network, it first passes through the firewall to block unauthorized access and filter some malicious traffic,

[0079] Subsequently, the intrusion detection system monitors network traffic in real time, detects abnormal behavior and potential security threats. If an intrusion or anomaly is detected, the IDS can issue a warning or take measures to protect, ensure the security of the internal network, and transmit secure network traffic to each device in the local area network.

[0080] It should be noted that the balanced intrusion detection method based on width focus learning provided in the embodiments of the present application is generally executed by a terminal device, and accordingly, the balanced intrusion detection system based on width focus learning is generally arranged between a gateway and the Internet. Figure 1 The number of terminal devices, networks and network devices in the above-mentioned network environment is only illustrative, and any number can be set according to requirements.

[0081] With reference to Figure 2 , a flowchart of the balanced intrusion detection method based on width focus learning of the present application is shown, the overall framework of the balanced intrusion detection method based on width focus learning is shown in Figure 3 , and the running flowchart is shown in Figure 4 The method comprises the following steps:

[0082] Step S1, obtaining a standardized data set.

[0083] In the present application, obtaining a standardized data set comprises: using data cleaning techniques to eliminate abnormal values, bad values and default values in the original data set, and converting the data through standardization techniques to convert the features to a unified scale, eliminate the differences in numerical ranges of different features, and obtain a standardized data set;

[0084] Explanation: Data cleaning techniques are used to preprocess the original data set to ensure the quality and reliability of the data. The core goal of data cleaning techniques is to eliminate abnormal values, bad values and default values in the data set to improve the integrity and consistency of the data set. Abnormal values refer to extreme data points that deviate from the normal range, which may be introduced due to input errors, sensor failures or other uncontrollable factors, and may mislead the model training. In order to eliminate these abnormal values, common methods include statistical detection or algorithm-based anomaly detection.

[0085] Explanation: Bad values usually refer to incorrect or invalid data in the data, which may be caused by transmission errors, system crashes or other unexpected situations. For example, in a user age field, if there are negative numbers or unreasonable large values, these are considered as bad values. Eliminating these data is a key step to ensure the accuracy of the data set. Default values are blank fields left due to missing information during data collection, and the commonly used methods include deleting samples containing a large number of default values or using other statistical methods, so that the data set is more complete.

[0086] Explanation: After data cleaning, the data is processed using standardization techniques. Because the numerical values of different features in the data set have different meanings, this difference can cause the model to be biased towards features with larger values and ignore features with smaller values during training. The purpose of standardization is to convert all features to a unified scale, usually achieved through statistical methods such as scaling. Through standardization, the differences in numerical ranges between features are eliminated, ensuring that different features are treated equally. After data cleaning and standardization, a more uniform and reliable standardized data set is obtained.

[0087] Step S2, based on the standardized data set, obtain the dimensionality reduction data set.

[0088] In this application, based on the standardized data set, the dimensionality reduction data set is obtained, including: through information variance and multicollinearity, screening and filtering the standardized data set, eliminating low-variance features that contribute less to the model, and removing redundant features that interfere with the model more by calculating the correlation matrix, obtaining the dimensionality reduction data set, which is specifically represented as:

[0089] In the entire dimensionality reduction data processing process, the standardized data set is input in the form of a two-dimensional matrix containing multiple samples and features, denoted as X, which contains n samples and p features;

[0090] Calculate the standard deviation of each feature, filter and remove features with a standard deviation below a preset threshold, and obtain a filtered data set;

[0091] The filtered data set is input in the form of a matrix, and the Pearson correlation coefficient between each feature in the filtered data set is calculated to generate a correlation matrix R between features, where each value R ij in the correlation matrix R represents the correlation between feature i and feature j;

[0092] Use the upper triangular mask to mask the upper half of the correlation matrix, filter out the correlation data in the lower triangular part through the mask operation, remove features with correlation coefficients exceeding the set threshold, and obtain data X';

[0093] The data flow is converted from the original standardized data set X to the simplified data X', and the simplified data X' is used as the dimensionality reduction data set.

[0094] Explanation: Information variance is a measure of how much each feature in the dataset varies across the entire sample set. A feature with a small variance indicates that it changes little between samples and remains essentially constant. Such features often provide little useful information to the model and may even introduce noise, reducing the model's predictive performance. Therefore, removing these low-variance features is an important step in feature selection. If a feature in the dataset takes the same value for almost all samples, it is likely that this feature does not contribute significantly to the model. By setting a threshold for variance, features with a variance below this threshold can be removed. Removing these low-variance features not only simplifies the dimensionality of the dataset but also improves the model's ability to focus on the main features.

[0095] Explanation: Multicollinearity refers to the high linear correlation between multiple features, causing their explanatory power for the target variable to overlap. This can make it difficult for the model to distinguish between features that contribute to the prediction result, affecting the stability and accuracy of the model. Therefore, detecting and removing redundant features is a key step in improving model performance. To detect multicollinearity, the correlation matrix of the dataset is calculated to show the correlation between each feature and other features. By analyzing these correlation coefficients, we can find features that are highly correlated with each other. When the correlation coefficient between two features exceeds a certain threshold (such as 0.98), they are considered highly correlated, and one of them can be considered redundant and removed from the dataset. Removing these redundant features can reduce the complexity of the model while avoiding multicollinearity problems.

[0096] Explanation: By removing low-variance features with little contribution through information variance selection, and then removing redundant features through multicollinearity detection, we obtain a simplified and efficient reduced dimension dataset. This reduced dataset retains the most valuable features for model prediction while eliminating features that contribute less to the model and redundant features, ensuring the simplification and performance optimization of the model. This reduced dataset, although with fewer feature dimensions, still represents the main information in the dataset, greatly improving the computational efficiency and generalization ability of the model.

[0097] Step S3: Introduce a deep belief network that efficiently represents complex multi-layer features as a generator of the generative adversarial network to generate generated data that better fits the real data, and select a sigmoid discriminator to distinguish the authenticity of the data, together constructing a generative adversarial network structure.

[0098] The generative adversarial network structure of step S3 is as shown in Figure 5 The process of generating new data using a deep belief network includes several key steps.

[0099] In the training phase, the deep belief network is trained layer by layer unsupervisedly through multiple layers of restricted Boltzmann machines. The weights of each layer are updated according to the following formula:

[0100] ΔW = η(<v i h j > data -v i h j > model )

[0101] where η is the learning rate, <v i h j > data and <v i h j > model represent the joint probability expectation under the data sample and the model generated sample respectively; by maximizing the difference between the data sample and the model generated sample, the restricted Boltzmann machine can learn the data features layer by layer.

[0102] When the training is completed, the deep belief network generates new data through the forward propagation process. The hidden units h j of each layer are activated by the visible units v i of the previous layer and the weight matrix W, and the specific formula is:

[0103]

[0104] where σ(·) is the activation function, W ij is the weight matrix, and b j is the bias term. When generating new data, the deep belief network generates new visible unit data by backward sampling, starting from the activation value of the top hidden layer, and propagating backward layer by layer. The reconstructed value of the visible unit v i is calculated by backward propagating the weight of the hidden unit h j , and the formula is:

[0105]

[0106] where c i is the bias term of the visible unit. The visible unit v i and the hidden unit h j are updated each time the sampling is performed until samples similar to the real data distribution are generated. Finally, the new data v' output by the deep belief network is obtained through multiple samplings. The deep belief network learns the latent distribution of the data and can generate new data with similar distribution.

[0107] Step S4, learning and training the reduced dimension data set based on the generative adversarial network, performing data enhancement on the minority class samples to generate more minority class samples with diversity and authenticity, so that the class distribution in the reduced dimension data set tends to be balanced, and an balanced data set is obtained;

[0108] The step S4 of learning and training the reduced dimension data set based on the generative adversarial network includes discriminator training of the generative adversarial network, and a flowchart is as shown in Figure 6 The step S4 of learning and training the reduced dimension data set based on the generative adversarial network includes discriminator training of the generative adversarial network, and a flowchart is as shown in

[0109] The real data is input into the discriminator to distinguish whether the data is real or fake;

[0110] The discriminator learns to distinguish the real data and the fake data generated by the generator, and makes a prediction on the input data, classifies based on the current discriminator parameters, and outputs the prediction result;

[0111] According to the prediction result of the discriminator, information for training the discriminator is fed back, which usually includes the prediction error of the discriminator, and the prediction error is used to update the weight of the discriminator;

[0112] The training state of the current discriminator is judged to check whether the expected training effect is achieved, if the training effect has not yet reached the expectation, the discriminator needs to be further optimized; if the discrimination effect reaches the expectation, the training is completed, and the next step is entered.

[0113] When the training effect is not ideal, the discriminator is optimized, and the optimization process includes adjusting the parameters of the discriminator, using the back propagation algorithm and the gradient descent method to update the weight of the discriminator, so as to improve the discrimination ability of the discriminator, and the optimized discriminator will be used again in the subsequent training cycle.

[0114] The step S4 of learning and training the reduced dimension data set based on the generative adversarial network includes discriminator training of the generative adversarial network, and a flowchart is as shown in Figure 7 The step S4 of learning and training the reduced dimension data set based on the generative adversarial network includes discriminator training of the generative adversarial network, and a flowchart is as shown in

[0115] The generator training is to input random noise into the generator, and the random noise is a vector composed of random numbers;

[0116] The generator generates data samples using the noise vector, and compares them with the real data set to learn to generate samples similar to the real data distribution, and generates new data samples according to the random noise;

[0117] The new data sample is input into the discriminator as fake data and compared with the real data;

[0118] The discriminator makes a judgment according to the input data, and outputs a probability indicating whether the input data is "real" or "fake"; the probability output by the discriminator will be used as feedback to guide the adjustment of the generated data;

[0119] According to the probability feedback output by the discriminator, the generator obtains an evaluation of the "good or bad" of the generated data. If the discriminator can better distinguish that the data generated by the generator is fake, the generator needs to further improve its generation mechanism, otherwise it indicates that the generation ability of the generator is enhanced; if the generator does not reach the expected target, the generator is optimized; the optimization method usually includes adjusting the model structure of the generator or updating the weight of the generator through back propagation and gradient descent method, so that the data generated by the generator is closer to the real data.

[0120] When the generator reaches the expected target and can generate data samples close enough to the real data, the training is completed, the parameters of the generator model are no longer updated, and the training enters the end stage, and the generator is also trained.

[0121] In this application, the generator and the discriminator are trained to enable the generator to generate enough realistic minority class samples, and to perform data enhancement on the minority class samples, while enabling the discriminator to accurately distinguish between true and false data, and the two cooperate with each other, so that the class distribution of the reduced dimension dataset tends to be balanced, and a balanced dataset is obtained.

[0122] Step S5, introduce focal loss mechanism into the width learning model, optimize the width learning weight expression, make the model pay more attention to the minority class samples, improve the classification performance on the unbalanced dataset, and thus construct a width focal learning model;

[0123] As described in step S5, the width focal learning structure in the balanced intrusion detection method based on width focal learning is as described in the following Figure 8 .

[0124] The width learning first maps the input data into a feature node matrix, wherein the process of generating the mapping node from the input data is as follows:

[0125]

[0126] Wherein: is an activation function, X is an input feature matrix, is the i-th group of feature mapping weight matrix, is the i-th group of feature mapping bias matrix, N is the total number of samples, and q is the number of feature nodes corresponding to each feature mapping. Then all the mapping feature node matrices are integrated to obtain the total mapping feature node matrix Z,

[0127] The process of generating the enhanced node from the mapping node is as follows:

[0128]

[0129] Wherein: ξ j is an activation function, Z is a mapping feature node matrix, is the jth group of feature enhancement weight matrix, is the jth group of feature enhancement bias matrix, N is the total number of samples, and r is the number of enhancement nodes corresponding to each group of enhancement transformation. Then, all the enhanced feature node matrices are integrated to obtain the total enhanced feature node matrix H. Then, the feature mapping nodes and the enhanced nodes are merged to obtain the total input matrix of the width learning model, denoted as A[Z|H].

[0130] The core idea of the focus concentration loss mechanism is to introduce a modulation factor and a weight term on the basis of the standard cross-entropy loss, so that the loss function can dynamically adjust the weight of the sample. The form is as follows:

[0131] L focal-loss =-a(1-p t ) y log(p t )

[0132] Where p t is the predicted probability of the model for the real class t; a is the weighting coefficient for different class samples, used to balance the class imbalance; y is the modulation factor parameter, used to adjust the weight of difficult and easy samples.

[0133] In order to minimize the error between the predicted value and the true value of the width focus learning model, and as far as possible to make the probability of positive samples larger, the optimization objective of the width focus learning model can be obtained as:

[0134] min:L(W)=λ f (1-AW) Y log(AW)+(AW-Y) 2 +λW 2

[0135] Where λ f is the focus penalty factor, A is the width learning input matrix, W is the weight of the width focus system, Y is the real label, and λ is the width learning penalty factor.

[0136] Step S6, taking the balanced data set as the input of the width focus learning model, training the width focus learning model, and solving the optimal weight through pseudo-inverse calculation and adjusting the focus factor to obtain the trained width focus learning model;

[0137] As described in step S6, the width focus learning training process of the balanced intrusion detection method based on width focus learning is as described in Figure 9 .

[0138] The width focus learning model maps the balanced data set into a feature matrix Z through the mapping node, enhances the feature representation ability of the model, and the feature matrix Z contains the effective features extracted from the input data.

[0139] The width focus learning model maps the feature matrix Z to an enhanced matrix H through an enhancement node. The role of the enhancement node is to expand the model's representation ability for input data by introducing a nonlinear transformation, thereby improving the model's generalization ability. The enhanced matrix H can further enhance the representation of features, enabling the model to capture complex patterns and characteristics in the data.

[0140] The feature matrix Z and the enhanced matrix H are merged horizontally to form an expanded input matrix A = [Z | H]. The matrix A serves as the final input for the width focus learning model, containing two parts of information: the original feature matrix and the enhanced feature representation. This ensures that the model can utilize both the original features and the complex features extracted through the enhancement node during the learning process. The width focus learning model introduces a focus concentration loss mechanism and a width learning penalty term. The focus concentration loss mechanism aims to address the class imbalance problem by adjusting the model's attention to different classes, especially focusing on minority class samples, thereby improving the detection ability of difficult-to-classify samples. During model training, the constraint term is optimized to minimize it, ensuring that the model's learning process not only accurately fits the data but also maintains good generalization ability. By introducing the focus concentration loss mechanism and optimizing the penalty term, the model exhibits better learning ability and detection performance when dealing with class-imbalanced datasets. Through these optimizations, the width focus learning model can more accurately identify samples of different classes, especially significantly improving the detection ability of minority class samples.

[0141] On the basis of the merged input matrix, the focus concentration loss mechanism and the width learning penalty term are introduced. The width focus learning optimization goal is set by combining the focus concentration loss mechanism and the width learning penalty term.

[0142] Based on the width focus learning optimization goal, the optimal weight is output by pseudo-inverse calculation and adjustment of the focus factor, and the trained width focus learning model is obtained.

[0143] Explanation: In the training process of the width focus learning model, the weight needs to be solved. The solution of the weight usually involves matrix operations. When an irreversible matrix appears, a suitable approximate solution is found through pseudo-inverse calculation. For example, during the forward propagation of the model, the input data may form a linear equation system y = Ax (here A is the coefficient matrix, x is the model parameter, and y is the output) after a series of transformations. When A is irreversible, the pseudo-inverse of A + , A + , can be calculated to obtain the approximate parameter solution y = A + x. Pseudo-inverse calculation helps to find a weight value that meets certain conditions during optimization. In the optimization process aimed at minimizing the loss function, it helps to adjust the weight so that the model's output is closer to the true value.

[0144] Explanation: When dealing with imbalanced datasets, the focus factor can guide the model to pay more attention to minority class samples. Minority class samples are usually more difficult to classify correctly. By adjusting the focus factor, more weight is given to minority class samples during training, allowing the model to better learn the characteristics of minority class samples. When the optimization goal is to improve the model's detection ability for minority class samples, the value of the focus factor can be appropriately increased to make the model focus more on learning minority class samples during training. At the same time, the adjustment of the focus factor will also affect the calculation of the model's loss function, as it changes the model's attention to different classes of samples, thereby affecting the model's weight update process. By continuously adjusting the focus factor and adjusting the weights based on pseudo-inverse calculations, the model eventually reaches the optimization goal and outputs the optimal weights.

[0145] Step S7, input the dataset samples to be detected into the trained width-focused learning model for classification and detection. The model can effectively identify samples of different attack types and output corresponding detection results through learned feature representations and focus-centralized loss mechanisms.

[0146] As described in step S7, the dataset to be detected is first input into the model. These data samples may come from network traffic data, system logs, or application behavior data, etc. At this time, the width-focused learning model has been fully trained in the previous stage and can extract feature representations from the input data.

[0147] The width-focused learning model uses the feature representations learned in the previous training process to map the input data to be detected into the feature space. Since in the training stage, the width-focused learning model has formed the ability to extract deep features from input data through the combination of feature matrices and enhancement matrices, the current input samples will also undergo similar feature mapping. These features can reveal the potential patterns and differences in the input data, especially for fine-grained characterization of possible attack behaviors.

[0148] The width-focused learning model strengthens the focus on minority class samples and difficult-to-classify samples by introducing a focus-centralized loss mechanism. In actual network attack detection scenarios, minority class attack types are often difficult for the model to identify, such as rare attacks or new attacks. These minority class attack samples may be ignored or misclassified during the model's learning process. The focus-centralized loss mechanism reduces the weight of easy-to-classify samples and enhances the influence of difficult-to-classify samples, thereby improving the detection accuracy of minority class attack samples.

[0149] The width focal learning model classifies and detects the input sample to be detected based on the learned weights and feature mapping inside it. By analyzing the features of each sample, the width focal learning model can determine whether the sample belongs to a certain known attack type or normal behavior. The output of the width focal learning model is a probability distribution representing the probability value of the sample belonging to each class. Different attacks have different feature patterns, and the width focal learning model identifies them through these feature patterns. The width focal learning model inputs the sample to be detected into the trained width focal learning model, and the model uses the learned deep features and focal concentrated loss mechanism to effectively identify and classify different attack types, and finally outputs high-precision detection results.

[0150] With reference to the foregoing Figure 10 The balanced intrusion detection system based on width focal learning includes:

[0151] The standardized dataset acquisition module 101 is configured to acquire a standardized dataset.

[0152] The dimensionality reduction dataset acquisition module 102 is configured to acquire a dimensionality reduction dataset based on the standardized dataset.

[0153] The generative adversarial network module 103 is configured to introduce a deep belief network as a generator of a generative adversarial network to generate realistic data samples, and select a sigmoid discriminator to distinguish the authenticity of the data, thereby jointly constructing the generative adversarial network.

[0154] The balanced dataset acquisition module 104 is configured to learn and train the dimensionality reduction dataset based on the generative adversarial network, perform data augmentation on the minority class samples to generate more minority class samples with diversity and authenticity, make the class distribution in the dimensionality reduction dataset tend to be balanced, and obtain a balanced dataset.

[0155] The width focal learning model construction module 105 is configured to introduce a focal concentrated loss mechanism into the width learning model to optimize the width learning weight expression and construct a width focal learning model.

[0156] The width focal learning model training module 106 is configured to input the balanced dataset into the width focal learning model as an input, train the width focal learning model, and solve the optimal weights by pseudo-inverse calculation and adjustment of the focal factor, thereby obtaining the trained width focal learning model.

[0157] The detection result output module 107 is configured to input the dataset sample to be detected into the trained width focal learning model for classification and detection. The width focal learning model effectively identifies samples of different attack types through learned feature representation and a focal concentrated loss mechanism, and outputs corresponding detection results.

[0158] To solve the above technical problems, the embodiment of the present application further provides a computer device structure. Specifically, refer to Figure 11 The figure is a schematic diagram of the computer device of the embodiment. The memory represents a storage system for processing data, which can be of various types in modern computer systems. This unit usually interacts with the processor to read / write data. The processor, as the central processor, is the "brain" of the computer, responsible for executing instructions and calculations, and interacting with the memory and the input / output system to process data and instructions. Further, the interface refers to the input / output interface of the system, which may include USB ports, network interfaces or any form of communication between the computer and external devices. Further, the integrator may refer to the system bus or data bus, which is a communication system for transmitting data between internal components of the computer (for example, between the processor, memory and I / O devices), and is the backbone connecting all important parts of the system.

[0159] The schematic diagram shows the basic architecture of how a typical computing system operates, in which the three main hardware components (storage, processor and interface) are connected by a central integrator or bus system. The storage device retains the data processed by the central processor. In the embodiment, the memory is used to store the operating system and various application software on the computer device, such as the program code of the balanced intrusion detection method and system based on width focus learning. The CPU extracts data from this memory to perform calculations, which is the main computing engine of the computer. It processes the instructions in the program by reading program instructions from memory, performing arithmetic and logical operations and storing the results back to memory, and it also controls the data flow to and from the interface and external devices. For example, running the program code of the balanced intrusion detection method and system based on width focus learning. The interface is the connection point between the computer and the outside world. The I / O port enables data to be moved into and out of the computer, allowing users to interact with the system and connecting it to peripheral devices. Typical examples include USB ports, network interfaces and display ports.

[0160] The present application also provides another embodiment, i.e. a non-volatile computer readable storage medium storing the program of the balanced intrusion detection method and system based on width focus learning, which can be executed by at least one processor to enable the at least one processor to perform the steps of the balanced intrusion detection method and system based on width focus learning as described above.

[0161] Through the description of the foregoing embodiments, those skilled in the art can clearly understand that the method can be implemented by software and a necessary universal hardware platform. Although hardware implementation is also feasible in some cases, the former is generally more optimal. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, stored in a storage medium, and contain a number of instructions, enabling an end device to perform the methods of various embodiments described in the present application.

[0162] Obviously, the above embodiments are only a part of examples of the present application, not all. The preferred embodiments of the present application are shown in the drawings, but should not be considered as limiting the scope of the patent of the present application. The present application can be implemented in various forms, and the purpose of providing these embodiments is to help the reader more clearly and comprehensively understand the content. Although the foregoing embodiments have been described in detail, those skilled in the art can still modify or equivalently replace the solutions. Any equivalent structure based on the content of the specification and drawings, and directly or indirectly applied to other related technical fields, should be included in the patent protection scope of the present application.

Claims

1. A balanced intrusion detection method based on width focus learning, characterized in that, include: Obtain a standardized dataset; Obtain a dimensionality-reduced dataset based on a standardized dataset; A deep belief network is introduced as the generator of the generative adversarial network to produce generated data that better fits the real data, and a sigmoid discriminator is selected to distinguish the authenticity of the data, thus jointly constructing the generative adversarial network; Generative adversarial networks are used to learn and train on dimensionality reduction datasets, and data augmentation is performed on minority class samples to generate more diverse and realistic minority class samples, so that the class distribution in the dimensionality reduction dataset tends to be balanced, resulting in a balanced dataset. The focal loss mechanism is introduced into the width learning model to optimize the weight expression of the width learning model and construct a width focal learning model. The balanced dataset is used as input to the width focus learning model to train the width focus learning model. The optimal weights are solved by pseudo-inverse calculation and adjustment of focus factor to obtain the trained width focus learning model. The dataset samples to be detected are input into the trained width-focus learning model for classification and detection. The width-focus learning model effectively identifies samples of different attack types and outputs corresponding detection results through the learned feature representations and focus concentration loss mechanism.

2. The equalization intrusion detection method based on width focus learning according to claim 1, characterized in that, The process of obtaining a dimensionality-reduced dataset based on a standardized dataset includes: screening and filtering the standardized dataset through information variance and multicollinearity to eliminate low-variance features with small contributions, and removing redundant features with large interference by calculating the correlation matrix to obtain the dimensionality-reduced dataset.

3. The equalization intrusion detection method based on width focus learning according to claim 2, characterized in that, The process involves screening and filtering the standardized dataset using information variance and multicollinearity to eliminate low-variance features with small contributions, and removing redundant features with significant interference by calculating the correlation matrix to obtain a dimensionality-reduced dataset. Specifically, this is manifested as follows: Throughout the dimensionality reduction data processing, the normalized dataset is input as a two-dimensional matrix containing multiple samples and features. Let the original normalized dataset be X, which contains n samples and p features. Calculate the standard deviation of each feature, filter and remove features whose standard deviation is below a preset threshold, and obtain the filtered dataset; The dataset is input as a matrix. The Pearson correlation coefficient between each feature in the filtered dataset is calculated, generating a feature-to-feature correlation matrix R, where each value in the correlation matrix R is R0. ij This indicates the correlation between feature i and feature j; The upper triangular mask is used to cover the upper half of the correlation matrix. The correlation data of the lower triangular part is filtered out by the mask operation, and features with correlation coefficients exceeding the set threshold are removed to obtain data X′. The data stream is transformed from the original standardized dataset X into a simplified dataset X′, and the simplified dataset X′ is used as a dimensionality-reduced dataset.

4. The equalization intrusion detection method based on width focus learning according to claim 1, characterized in that, Generative adversarial networks use Binary Cross Entropy Loss for loss assessment.

5. The equalization intrusion detection method based on width focus learning according to claim 1, characterized in that, By introducing a deep belief network as the generator of a generative adversarial network, generated data that better fits the real data is produced. Specifically, this manifests as follows: During the training phase, the deep belief network is trained layer by layer in unsupervised manner using a multi-layer restricted Boltzmann machine, and the weights of each layer are updated according to the following formula: ΔW=η( <v i h j > data - <v i h j > model ) Where η is the learning rate. <v i h j > data and <v i h j > model Let represent the joint probability expectation under the data sample and the model-generated sample, respectively. By maximizing the difference between the two, the Restricted Boltzmann Machine can learn the data features layer by layer. After training is complete, the deep belief network generates new data through the forward propagation process, and the hidden units h of each layer... j From the visible unit v of the previous layer i And the weight matrix W is activated, and the hidden units include h i with h j It can be seen that the set of units v includes v i v j The specific formula is as follows: Where σ(·) is the activation function, W ij Let b be the weight matrix. i As a bias term; when generating new data, the deep belief network generates new visible unit data by backsampling, starting from the activation value of the top hidden layer and propagating backward layer by layer. Visible unit v j The reconstructed value is obtained by comparing it with the hidden unit h. i The weights are calculated using backpropagation, and the formula is: Among them, c i As the bias term for the visible units, the deep belief network can generate new samples that conform to the training data distribution through repeated forward and backward propagation; each sampling updates the visible unit v. i and hidden unit h j The process continues until samples similar to the real data distribution are generated. Ultimately, the new data v′ output by the deep belief network is obtained through multiple samplings. Through these steps, the deep belief network learns the latent distribution of the data and can generate new data with a similar distribution.

6. The equalization intrusion detection method based on width focus learning according to claim 1, characterized in that, The focal loss mechanism is introduced into the width learning model, specifically manifested as follows: Width learning first maps the input data into a feature node matrix using a mapping function. The process of generating the mapped nodes from the input data is as follows: in: Let X be the activation function, and X be the input feature matrix. It is the weight matrix of the i-th feature map. Let q be the bias matrix of the i-th feature mapping group, N be the total number of samples, and q be the corresponding number of feature nodes. The process of integrating all mapping matrices into the total mapped feature node matrix Z and generating enhanced nodes is as follows: Where: ξ j Let Z be the activation function, and Z be the mapping feature node matrix, including Z0... i ; It is the feature enhancement weight matrix of the j-th group. is the feature enhancement bias matrix of the j-th group, N is the total number of samples, and r is the number of enhancement nodes corresponding to each enhancement transformation; Then, all the enhanced feature node matrices are integrated; finally, the mapping nodes and enhanced nodes are merged to obtain the total input matrix of the width learning model, denoted as A[Z|H]. The core idea of ​​the focus-focus loss mechanism is to introduce a modulation factor and a weight term into the standard cross-entropy loss, so that the loss function can dynamically adjust the weights of the samples, as follows: L focal-loss =-a(1-p t ) y log(p t ) Where p t is the model's predicted probability of the true class t; a is the weighting coefficient for samples of different classes, used to balance class imbalance; y is the modulation factor parameter, used to adjust the weights of easy and difficult samples; To minimize the error between the predicted and actual values ​​of the width-focused learning model and increase the probability of positive samples, the optimization objective of the width-focused learning model can be obtained as follows: min:L(W)=λ f (1-AW) Y log(AW)+(AW-Y) 2 +λW 2 Where λ f λ is the focus penalty factor, A is the width learning input matrix, W is the weight of the width focus system, Y is the true label, and λ is the width learning penalty factor.

7. A balanced intrusion detection system based on width focus learning, used to implement the balanced intrusion detection method based on width focus learning according to claims 1-6, characterized in that, The system includes: The standardized dataset acquisition module is used to acquire standardized datasets. The dimensionality reduction dataset acquisition module is used to obtain a dimensionality reduction dataset based on a standardized dataset. The Generative Adversarial Network (GAN) module is used to introduce a deep belief network as the generator of the GAN, generate realistic data samples, and select a sigmoid discriminator to distinguish the authenticity of the data, thus jointly constructing the GAN. The balanced dataset acquisition module is used to learn and train the dimensionality reduction dataset based on the generative adversarial network, perform data augmentation on the minority class samples, generate more diverse and realistic minority class samples, so that the class distribution in the dimensionality reduction dataset tends to be balanced, and a balanced dataset is obtained. The width-focused learning model building module is used to introduce the focal loss mechanism into the width-learning model, optimize the width-learning weight expression, and build the width-focused learning model. The width-focus learning model training module is used to train the width-focus learning model by taking the balanced dataset as input. It solves for the optimal weights and obtains the trained width-focus learning model by pseudo-inverse calculation and adjusting the focus factor. The detection result output module is used to input the dataset samples to be detected into the trained width-focus learning model for classification and detection. The width-focus learning model effectively identifies samples of different attack types and outputs the corresponding detection results through the learned feature representation and focus concentration loss mechanism.

8. An electronic device, characterized in that, include: One or more processors, a memory, and a computer program stored therein, comprising instructions that cause the device to perform any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium contains a computer program that, when run, causes the computer to execute any of the methods of claims 1 to 6.

Citation Information

Patent Citations

  • Network intrusion detection method and system based on ensemble learning

    CN113922985A

  • Network security early warning method and system based on deep learning

    CN118353667A