Intrusion detection method and system

Through the methods of feature selection, oversampling and improving network structure, the problem of insufficient detection capabilities of existing intrusion detection systems for diversified abnormal behavior is solved, and more efficient and accurate intrusion detection is achieved.

CN120034381APending Publication Date: 2025-05-23GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187567.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing deep learning-based intrusion detection systems rely on a large number of fixed training data and are difficult to apply to diverse anomalies.

Method used

By acquiring traffic data, selecting features based on representativeness, oversampling using adversarial networks, combining Actor-Critic algorithm and Mamba-optimized U-Net to improve network structure, model training is performed based on Markov decision-making process.

Benefits of technology

The intrusion detection system's detection ability of diversified anomalies is improved, and the training efficiency and detection accuracy of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034381A_ABST
    Figure CN120034381A_ABST
Patent Text Reader

Abstract

The invention discloses an intrusion detection method and system, and the method comprises the steps: obtaining traffic data, considering the representativeness of a data set, and carrying out the feature selection of the traffic data; constructing an adversarial network, and performing oversampling on the data; an initial Actor-Critic algorithm model is improved, and an intrusion detection model is obtained; and based on a Markov decision process, training an intrusion detection model by using the oversampled data. The system comprises a data preprocessing module, an oversampling module, a model building module and a model training module. According to the invention, the accuracy of intrusion detection can be improved. The method can be widely applied to the field of network security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and in particular to an intrusion detection method and system. Background Art

[0002] With the rapid development of information and communication technology, people have gained great convenience in all aspects of life, and the Internet has become an indispensable part of people's lives. However, due to the special information transmission mechanism of cyberspace, exposing key information in cyberspace will bring immeasurable risks.

[0003] In order to maintain the security of the network environment, many security technologies for combating network attacks have emerged, including firewall technology, network information encryption technology, vulnerability scanning technology, and intrusion detection technology. Intrusion detection technology is a proactive security defense technology that mainly detects attack behaviors by collecting and analyzing relevant network information and formulating reasonable defense strategies.

[0004] Existing deep learning-based intrusion detection systems (IDS) usually rely on a large amount of training data to capture the behavior patterns of normal traffic. However, abnormal traffic is often diverse and unpredictable, which causes the model to perform poorly when facing unknown abnormal behavior. Summary of the invention

[0005] In view of this, in order to solve the technical problem that most of the existing intrusion detection methods rely on a large amount of fixed training data, which leads to the inability to be applicable to diverse abnormal behaviors, the present invention proposes an intrusion detection method, which includes the following steps:

[0006] Acquiring flow data and performing feature selection on the flow data based on representativeness;

[0007] Oversampling the selected data based on adversarial networks;

[0008] Based on the Actor-Critic algorithm model, the network structure is improved by using the U-Net optimized by Mamba.

[0009] Based on the Markov decision process, the model is trained using the preprocessed network data.

[0010] In some embodiments, the step of obtaining traffic data and performing feature selection on the traffic data based on representativeness is aimed at extracting representative features from the original network traffic data to reduce the data dimension and improve the training efficiency of the model, which specifically includes:

[0011] Capture traffic data from network environments;

[0012] The features in the original traffic data are divided into string features and numerical features, and the string features are one-hot encoded and converted into numerical form;

[0013] Perform numerical normalization on the digitized feature set to prevent certain features from dominating the model training due to excessively large values;

[0014] The chimpanzee-chicken algorithm is used to select the most representative features from the normalized data, reducing the data dimension and improving the model efficiency.

[0015] In some embodiments, the step of oversampling the selected data based on the adversarial network to obtain the oversampled data aims to solve the data imbalance problem and generate more minority class samples, which specifically includes:

[0016] A stacked WGAN-GP framework is used to construct a generative adversarial network, including a generator and a discriminator, to generate minority class samples;

[0017] Through pre-training, the model can generate attack samples similar to real malicious traffic samples;

[0018] The selected data is input into the trained generative model to oversample the minority class samples to solve the data imbalance problem.

[0019] In some embodiments, the Actor-Critic algorithm model is used as the basic model, and the network structure is improved by using the U-Net optimized by Mamba. The purpose is to combine reinforcement learning and deep learning technology to build an efficient intrusion detection model, which specifically includes:

[0020] Model the environment to form an agent that can simulate the real network environment for training and testing intrusion detection models;

[0021] Through adversarial learning, the classification agent (intrusion detection model) and the environmental agent are jointly optimized;

[0022] Improve the traditional SAC framework to make it more suitable for intrusion detection tasks.

[0023] In some embodiments, based on the Markov decision process, the step of training the intrusion detection model using the preprocessed network data aims to train the intrusion detection model through a reinforcement learning framework so that it can dynamically adapt to changes in the network environment, specifically including:

[0024] Define state space, action space and reward value;

[0025] The classification agent obtains the data flow at the current moment from the environment agent;

[0026] Select real data labels according to the maximum entropy strategy to guide the training of classification agents;

[0027] Generate real-time reward values ​​by comparing the true labels with the classification agent’s predictions;

[0028] The environment agent selects the next data flow based on the current state and reward value;

[0029] Store the current state, action, reward, and next state into the experience replay pool for subsequent training.

[0030] The present invention also proposes an intrusion detection system, the system comprising:

[0031] A data preprocessing module, which obtains flow data and performs feature selection on the flow data based on representativeness;

[0032] Oversampling module, oversampling the selected data based on adversarial network;

[0033] The model building module uses the Actor-Critic algorithm model as the basic model and improves the network structure through the U-Net optimized by Mamba;

[0034] The model training module, based on the Markov decision process, uses the preprocessed network data to train the model.

[0035] Based on the above scheme, the present invention provides an intrusion detection method and system, which enhances the exploratory ability of the intelligent agent through the SoftActor-Critic algorithm, improves the sample imbalance problem based on the attack sample generation model of the improved WGAN-GP, and effectively extracts the characteristics of traffic data by using the enhanced deep learning strategy network and action network composed of Mamba-UNet to facilitate strategy formation, thereby improving the accuracy of intrusion detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flow chart of steps of an intrusion detection method of the present invention;

[0037] Figure 2 It is the algorithm framework of the present invention;

[0038] Figure 3 This is the structure diagram of Mamba-UNet of the present invention;

[0039] Figure 4 This is a diagram of the Mamba block structure of the present invention. DETAILED DESCRIPTION

[0040] In addition to the problems raised in the background technology, existing algorithms in intrusion detection systems (IDS) still have some limitations, which mainly focus on detection uncertainty, data imbalance and lack of historical data.

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0042] It should be noted that, for the convenience of description, only the parts related to the relevant invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0043] It should be understood that the "system", "device", "unit" and / or "module" used in this application is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.

[0044] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "a kind" and / or "the" do not refer to the singular, but also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. The elements defined by the sentence "includes a..." do not exclude the existence of other identical elements in the process, method, commodity or device that includes the elements.

[0045] In the description of the embodiments of the present application, "plurality" means two or more than two. The following terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.

[0046] In addition, flow charts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, the various steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0047] Reference Figure 1, is a flow chart of an optional example of the intrusion detection method proposed in the present invention. The method can be applied to computer devices. The detection method proposed in this embodiment may include but is not limited to the following steps:

[0048] Step S1, a preprocessing step, obtaining flow data and performing feature selection on the flow data based on representativeness to obtain selected data;

[0049] Step S2, an oversampling step, oversampling the selected data based on an adversarial network to obtain oversampled data;

[0050] Step S3, model construction step, taking the Actor-Critic algorithm model as the basic model, improving the network structure of the initial Actor-Critic algorithm model through the U-Net optimized by Mamba, and obtaining the intrusion detection model;

[0051] Step S4, model training step, based on the Markov decision process, uses the oversampled data to train the intrusion detection model to obtain a trained intrusion detection model.

[0052] In order to classify normal traffic and abnormal traffic of corresponding attack types, an intrusion detection model consisting of the interaction between the classification agent and the detection environment is constructed, such as Figure 2 As shown. First, the chimpanzee-chicken optimization algorithm is used to select features of network traffic data. In order to alleviate the problem of traffic data sample imbalance, an attack sample generation model of the improved WGAN-GP framework is constructed. WGAN-GP networks are created for samples of different categories to form a stacked WGAN-GP, which effectively prevents noise interference of different categories. Adding gradient penalty terms can effectively improve the training stability of GAN and the quality of generated data. The cosine similarity is used to improve the objective function to enhance the learning ability of sample distribution. High-quality attack samples are generated through the improved WGAN-GP model to alleviate network data imbalance.

[0053] The intrusion detection environment is a reinforcement learning environment based on the improved SoftActor-Critic framework. The environment is changed to be suitable for intrusion detection tasks by improving the action space, state space and reward function, and the environment is modeled as an environmental agent. Specifically, the environmental agent randomly selects a state, inputs the current state into the action network formed by Mamba-UNet to select the action distribution, and selects the traffic that the classification agent should currently process based on the probability of the action distribution. The above is used as the simulation environment to select actions. The classification agent is a model responsible for extracting traffic features and classifying them. Specifically, the classification agent extracts features from the input traffic data through Mamba-UNet, converts it into an exact state space, and uses the traffic label as its corresponding action to form an action space.

[0054] After extracting features from the traffic data, the agent classifies the currently input traffic data, i.e., it determines the possibility of normal traffic based on the features, or abnormal traffic of the corresponding attack type.

[0055] In some feasible embodiments, step S1, the preprocessing step specifically includes:

[0056] S1.1. Obtain raw traffic data through network traffic capture tools (such as Wireshark, tcpdump, etc.).

[0057] Among them, the captured data usually includes basic information of the network packet (such as source IP, destination IP, port, protocol type, packet size, timestamp, etc.).

[0058] S1.2. The features in the original traffic data are divided into string features and numerical features, where string features include protocol type (such as TCP, UDP), service type (such as HTTP, FTP), flags (such as SYN, ACK), etc.; numerical features include packet size, transmission rate, time interval, etc. Then, the string features are one-hot encoded and converted into binary vector form; the digitized feature set includes numerical features and one-hot encoded string features.

[0059] S1.3. Perform numerical normalization on the digitized feature set to normalize the numerical features to the same scale. The normalization methods include Min-Max normalization and Z-Score normalization. Output the normalized data so that all feature values ​​are at the same scale.

[0060] S1.4. Perform feature selection on the normalized data, where the feature selection methods include filtering (selecting features based on statistical indicators), wrapping (selecting features based on model performance) and embedding (selecting features during model training); Output: selected data, including the most representative features.

[0061] In some feasible embodiments, step S1.4 specifically includes:

[0062] The chimpanzee-chicken algorithm is used to perform feature selection on the normalized data.

[0063] The network traffic data set contains many low-value features. The chimpanzee-chicken swarm algorithm is used to select the most representative feature set for network traffic. By modeling the chimpanzee group's behaviors such as capturing prey and social communication, the search process for the optimal feature subset in the feature set space is simulated. At the same time, the hierarchical structure of the chicken swarm algorithm is used to optimize the chimpanzees' hunting and search behaviors. The lower the fitness, the larger the search range, which accelerates the chimpanzees' convergence ability and prevents the chimpanzees from falling into the optimal solution of currency aggregation.

[0064] S141, Initialization: In the initialization stage, each solution represents the position of a chimpanzee, that is, the feature subset L n ,The normalized dataset ensures that all features are on the same scale, which helps the chimpanzee-chicken algorithm to more accurately evaluate the importance of features and then adaptively search for the optimal subset.

[0065] L={L 1 ,L 2 ,...,L n}; 1≤t≤n

[0066] S142, fitness calculation: calculate the fitness value of each solution and select the optimal solution as the "prey".

[0067] The fitness function is calculated to measure the dependency between features and target variables, thereby evaluating the fitness of each feature subset. The feature set with a lower fitness value is the best feature set for network intrusion detection traffic. The fitness value is estimated using the following expression:

[0068]

[0069] Where e specifies the total number of samples, ρ represents the fitness metric, and J e is the output of the network, Indicates the target output.

[0070] S143. During the hunting and chasing phases, chimpanzees chase their “prey” (i.e., the optimal feature subset) based on their distance from it. The random search behavior of the chicks in the swarm algorithm is combined with the directional search behavior of the chimpanzees to improve the convergence of the algorithm, simulating the search process for the optimal feature subset in the feature space. The randomness of the chicks helps to explore more feature combinations, while the directional nature of the chimpanzees helps to quickly approach the optimal feature set.

[0071]

[0072] Where a represents the number of iterations, x and k specify the coefficient values, and L chimp Indicates the location of the chimpanzee, L prey represents the current position of the prey, Randu(0,ω 2) represents a Gaussian distribution with a standard deviation of ω and a mean of 0, and n represents a coefficient vector.

[0073] S144, Attack process: In the attack phase, the prey is attacked intensively. The attack process is divided into two stages: exploring the location of the prey and surrounding the prey for attack. The attacking chimpanzee usually performs the attack process. In addition, other chimpanzees in the group, such as barrier chimpanzees, driving chimpanzees, and chasing chimpanzees, also contribute to the attack process, and the combination and interaction of different feature subsets are used to improve the performance of the model. The attack strategy of each character is based on the analysis of the current location and historical behavior of the prey (optimal feature subset).

[0074] Therefore, the four optimal solutions are specified as:

[0075]

[0076] And, each optimal solution is specified as:

[0077] L 1 =L attacker -n 1 (N attacker )

[0078] L 2 =L barrier -n 2 (N barrier )

[0079] L 3 =L chaser -n 3 (N chaser )

[0080] L 4 =L driver -n 4 (N driver )

[0081] L 1 , L 2 , L 3 , L 4 Represents attack, obstacle, pursuit and driving, N attacker 、N barrier 、N chaser 、N driver They represent attacker prey, barrier prey, pursuer prey, and driver prey, respectively.

[0082] S145, Social Incentives: Obtaining social satisfaction and appropriate motivation in the last segment will cause the chimpanzees to abandon the hunting task and restart the chimpanzee's position between the normal update model or random values ​​depending on the pursuit and drive process of the prey.

[0083]

[0084] Where k is either 0 or 1. Here, the random value contains a sequence of evolving variables that repeats itself completely, so that regular intervals starting from any point in the sequence.

[0085] S146, Fitness Re-evaluation: Re-evaluate the quality of the solution based on the fitness function.

[0086] S147, termination condition: repeat the above process until the best sub-feature set is reached.

[0087] The chimpanzee-chicken algorithm adaptively searches the feature space to find the optimal feature subset, which is the feature subset that is most representative of the data set and has the relatively least number of selected features.

[0088] In some feasible embodiments, step S2 specifically includes:

[0089] S2.1, using the stacked WGAN-GP framework to build a generative adversarial network for generating attack samples;

[0090] The model builds independent generators, discriminators, and classifiers for samples of each category, and defines the generator as G N , the discriminator is D N , the classifier is C N . G N , D N and C N All are feed-forward neural networks. N It includes input layer, output layer and hidden layer. The activation function of the output layer is Linear and the activation function of the hidden layer is ReL. u Among them, the number of neurons in the input layer is equal to the noise dimension, and the number of neurons in the output layer is equal to the real sample dimension. N and C N It includes input layer, output layer and hidden layer, where the number of neurons in the input layer is equal to the real sample dimension, D N and C N Shared hidden layer, C N The output layer outputs the probability of each sample, D N Then determine whether the sample is a generated sample. The hidden layer activation function is ReL u The objective function expression of the improved WGAN-GP is as follows:

[0091]

[0092] Among them, P r is the real data distribution; P g To generate data distribution; z~P ris a Gaussian noise distribution; the function of G(z) is to map the noise vector to the data space of the generated samples, where D(x) represents the probability that a given sample x is a true sample. C(x k ) represents the classification loss used to update the discriminator D, represents the classification loss used to update the generator G. k Represents the generated samples The result of calculating the cosine similarity with other samples. In order to distinguish the generated samples from the real samples, the discriminator will try to increase the value of D(x) and When the objective function reaches the global optimal solution, P r =P g .

[0093] The corresponding original sample is x k After feature extraction by the classifier network C, the high-dimensional features of the generated sample and the original sample are F(G k (z)) and F(x k ), then by calculating the cosine similarity of high-dimensional features, we can get:

[0094]

[0095] in, are high-dimensional features of generated samples of different categories, represents the cosine similarity between the high-dimensional features of the generated sample and the high-dimensional features of the original sample, Represents the cosine similarity between the high-dimensional features of generated samples of different categories. We expect the distribution distance between the generated samples and the original samples to be as close as possible, that is, the smaller the better, and we expect the distribution distance between generated samples of different categories to be as far away as possible to prevent noise interference between different categories of data.

[0096] S2.2. Pre-train the adversarial sample generation model to obtain a trained attack sample generation model;

[0097] First, initialize the parameters of the two networks, the generator and the discriminator, and define a noise distribution that follows a Gaussian distribution. Second, prepare the real data and the noise data. The real data is obtained from a few attack classifications in the dataset. Take a subset of one of the attack traffic categories and at the same time take the same amount of noise as the attack data from the noise distribution. Third, fix the generator and train the discriminator. The noise generates the same number of generated samples through the generator, and the discriminator is trained using the attack traffic subset and the generated samples. The discriminator judges the source of the data through a neural network to distinguish between the attack traffic subset and the generated samples. Fourth, fix the discriminator and train the generator. After training the discriminator obtained by training different categories of data 100 rounds according to the third step to train the generator, make the discriminator as confused as possible about the source of the data so that it cannot distinguish whether the data is the real number of the original traffic dataset or from the generated samples. After 1000 update iterations according to the third and fourth steps, the final attack sample generator is obtained. Fifth, use each attack subset in turn to train the generative adversarial network according to the above steps, and finally obtain different attack sample generators.

[0098] S2.3. Input the preprocessed data into the trained attack sample generation model for oversampling, and use the output data as the new traffic data.

[0099] In some feasible embodiments, step S3 specifically includes:

[0100] S3.1. Model the environment to form an environment proxy, and the classification agent and the environment proxy perform adversarial learning to construct an improved SoftActor-Critic network framework;

[0101] A model for traffic data classification is constructed using the SoftActor-Critic architecture, as Figure 2 shown. The model includes that the Critic is the evaluation network. When the input is the environmental state, it can evaluate the value of the current state. When the input is the environmental state and the action taken, it can evaluate the value of taking this action in the current state. The Actor is the action network, which takes the current state as the input, and the output is the probability distribution of the action or the continuous action value, and then the Critic network evaluates the quality of this action to adjust the strategy, including an Actor network, two V Critic networks (evaluation network and target network), and two Q Critic networks (Q 0 and Q 1 networks).

[0102] By modeling the environment to form an environment proxy, the environment proxy and the classification agent have the same network structure, which also includes an Actor network, two V Critic networks (evaluation network and target network), and two Q Critic networks (Q0 and Q 1 network), thus forming an adversarial SoftActor-Critic model, in which the classification agent and the environment agent are trained adversarially to improve the accuracy of the classification agent.

[0103] S3.2. Use Mamba-UNet to improve the network in the SoftActor-Critic framework.

[0104] Mamba-UNet Figure 3 It means that the whole structure includes multiple Mamba blocks, and the adjacent Mamba blocks are connected through a fully connected network, so that the model can efficiently extract the similarity information between samples. At the same time, the entire U-Net network is divided into two parts, the encoding network and the decoding network. The encoding network is composed of multiple Mamba blocks, which is used to learn the features of the data while maintaining the feature dimension; the decoding network is also composed of multiple Mamba blocks to reconstruct the features of the network data. The encoder and decoder at each level use jump connections to mix multi-scale features with the amplified output, and enhance spatial details by merging shallow and deep layers.

[0105] In the Mamba block, the input features first encounter a linear embedding layer and then fork into a dual path. One branch goes through the convolutional layer and activation function layer, enters the SSM module and layer normalization, and merges with the alternate stream after activation. The Mamba module chooses a streamlined structure without an MLP stage, thereby achieving a denser block stack within the same depth budget. Figure 4 shown.

[0106] A discrete state space model (SSM) defines a function from an input signal (a function of time t) x(t)∈R M To the output signal y(t)∈R M Through a latent state h(t)∈R M The linear mapping of As a parameter, as the projection parameter of the state size M, and the skip connection By giving a time scale parameter Δ∈R D , the SMM is expressed using a linear ordinary differential equation:

[0107] y(t)=Ch(k)+Dx(k)

[0108] h(k)=Ah(k-2)+Bx(k)

[0109] A= e ΔA

[0110] B=(e ΔA -I)A -1 B

[0111]

[0112] In some feasible embodiments, step S4 specifically includes:

[0113] The intrusion detection problem is abstracted into a Markov decision process. The Markov decision process includes four tuples, namely:

[0114] State space S, S = {s t} represents the traffic feature dataset currently being processed;

[0115] Action space A, A={a t} represents the real label set of the current processed traffic data;

[0116] {s t+1} represents the traffic feature data set to be processed in the next step;

[0117] Reward value R, R = {r t} represents a set of reward values, where r t Represents the feedback reward value generated by comparing the classification result with the true label.

[0118] The classification agent obtains data flow s selected by the environment agent through the action network at time t t , according to the maximum entropy strategy Select the real data label a t ∈A, generates a real-time reward value r based on the data label and the classification agent prediction results t , and then the environment agent selects the data flow as the flow data s for the next step of processing through the action network t+1 , store the quadruple into the experience replay pool and enter the next iteration step.

[0119] Among them, the classification agent does not receive any reward if the classification is wrong, and receives a reward value if the classification is correct. The formula is:

[0120]

[0121] The environment agent does not receive any reward if the classification is correct, but receives a reward value if the classification is wrong. The formula is:

[0122]

[0123] Maximum entropy learning strategy, including:

[0124] When the experience replay pool reaches the preset number, batch network traffic data is sampled from the experience replay pool, and the maximum entropy learning strategy is performed using approximate reasoning through the value function;

[0125] The purpose of maximum entropy learning is to maximize the sum of the cumulative reward value and the policy entropy value, which is expressed as:

[0126]

[0127] Among them, α is the temperature parameter, which determines the importance of entropy value relative to reward, r(a t ,s t ) means in s t Moment a t The instant reward value obtained, π represents the strategy, that is, the probability distribution mapping from state to action, and the optimization goal is to find the strategy π that maximizes the objective function * , represents the expected reward value under the maximum entropy strategy, and H(π(·|s t ) means in s t The policy entropy under the state measures the randomness of the action distribution and is calculated as:

[0128] H(π(·|s t ))=-logπ(·|s t )

[0129] The SoftActor-Critic algorithm is defined as:

[0130]

[0131] Q(s t )=r(a t ,s t )+γV(s t+1 )

[0132] Among them, Q(a t ,s t ) is the value function of the state-action pair, π(a t |s t ) is in state s t Next select action a t The probability of α is the temperature parameter used to balance exploration and utilization, r(a t ,s t ) is in state s t Take action a t The immediate reward obtained, where γ∈[0,1] is the discount factor used to weigh the immediate reward and future rewards.

[0133] The discrete action space uses the maximum entropy strategy, which is beneficial to enhancing the exploration rate of the agent and ensuring that the agent does not fall into a local optimal solution.

[0134] An intrusion detection system, comprising:

[0135] A preprocessing module, configured to obtain traffic data and perform feature selection on the traffic data based on representativeness to obtain the selected data;

[0136] An oversampling module, configured to perform oversampling on the selected data based on an adversarial network to obtain the oversampled data;

[0137] A model construction module, based on the Actor-Critic algorithm model as the basic model, improves the network structure of the initial Actor-Critic algorithm model through the U-Net optimized by Mamba to obtain an intrusion detection model;

[0138] A model training module, based on the Markov decision process, uses the oversampled data to train the intrusion detection model to obtain a trained intrusion detection model.

[0139] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0140] An intrusion detection device:

[0141] At least one processor;

[0142] At least one memory, configured to store at least one program;

[0143] When the at least one program is executed by the at least one processor, the at least one processor implements an intrusion detection method as described above.

[0144] The content in the above method embodiments is applicable to the present device embodiment. The functions specifically implemented by the present device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0145] A storage medium, in which instructions executable by a processor are stored, and the instructions executable by the processor are used to implement an intrusion detection method as described above when executed by the processor.

[0146] The content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0147] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. An intrusion detection method, characterized in that: The following steps are involved: Acquire flow data and perform feature selection on the flow data based on representativeness to obtain selected data; Oversampling the selected data based on an adversarial network to obtain oversampled data; Taking the Actor-Critic algorithm model as the basic model, the network structure of the initial Actor-Critic algorithm model is improved by using the U-Net optimized by Mamba to obtain the intrusion detection model. Based on the Markov decision process, the intrusion detection model is trained using the oversampled data to obtain a trained intrusion detection model.

2. The intrusion detection method according to claim 1, characterized in that: The step of acquiring the flow data and performing feature selection on the flow data based on representativeness to obtain selected data specifically includes: Get traffic data; The features in the traffic data are divided into string features and numerical features, and the string features are one-hot encoded to obtain a digitized feature set; Performing numerical normalization processing on the digitized feature set to obtain normalized data; Based on the representativeness of the data set, feature selection is performed on the normalized data to obtain selected data.

3. The intrusion detection method according to claim 2, characterized in that: The chimpanzee-chicken algorithm is used to perform feature selection on the normalized data.

4. The intrusion detection method according to claim 1, characterized in that: The step of oversampling the selected data based on the adversarial network to obtain oversampled data specifically includes: The stacked WGAN-GP framework is used to build an attack sample generation model; Pre-training the adversarial sample generation model to obtain a trained attack sample generation model; The selected data is input into the trained attack sample generation model for oversampling processing to obtain oversampled data.

5. The intrusion detection method according to claim 1, characterized in that: The step of taking the Actor-Critic algorithm model as the basic model and improving the network structure of the initial Actor-Critic algorithm model through the U-Net optimized by Mamba to obtain the intrusion detection model specifically includes: Model the environment to form an environmental agent, classify the agent and the environmental agent for adversarial learning, and build an improved SoftActor-Critic network framework; Use Mamba-UNet to improve the network in the SoftActor-Critic framework.

6. An intrusion detection method according to claim 1, characterized in that: The step of training the intrusion detection model based on the Markov decision process using the oversampled data to obtain a trained intrusion detection model specifically includes: Define state space, action space and reward value; Based on the state space, the action space and the reward value, the intrusion detection model is trained in combination with a maximum entropy strategy to obtain a trained intrusion detection model.

7. An intrusion detection method according to claim 6, characterized in that: The step of training the intrusion detection model based on the state space, the action space and the reward value in combination with the maximum entropy strategy specifically includes: The classification agent obtains data flow selected by the environment agent through the action space at a certain moment; Select the real data labels based on the maximum entropy strategy; A real-time reward value is generated based on the real data label and the classification agent's prediction results. The environment agent then selects the data flow as the flow data for the next step of processing through the action space, stores the quadruple in the experience replay pool, and enters the next iteration step.

8. An intrusion detection method according to claim 7, characterized in that: The step of generating a real-time reward value based on the real data label and the classification agent prediction result specifically includes: The environment agent and the classification agent obtain comparative results based on the classification results and the true labels of the actual traffic; Optimizing the classification agent and the environment agent based on the comparison result; For the classification agent, no reward is obtained for incorrect classification, and a reward value is obtained for correct classification; For the environmental agent, there is no reward for correct classification, and a reward value is obtained for incorrect classification.

9. An intrusion detection method according to claim 6, characterized in that: The expression of the maximum entropy strategy is as follows: Among them, α represents the temperature parameter, r(a t ,s t ) means in s t Moment a t The immediate reward value obtained, H(π(·|s t ) means in s t The policy entropy under the state measures the randomness of the action distribution, π represents the policy, Represents the expected reward value under the maximum entropy strategy.

10. An intrusion detection system, characterized in that: include: A data preprocessing module, used for acquiring flow data and performing feature selection on the flow data based on representativeness to obtain selected data; An oversampling module, oversampling the selected data based on an adversarial network to obtain oversampled data; The model building module is used to improve the network structure of the initial Actor-Critic algorithm model through the U-Net optimized by Mamba based on the Actor-Critic algorithm model to obtain the intrusion detection model; The model training module trains the intrusion detection model based on the Markov decision process using the preprocessed network data to obtain a trained intrusion detection model.