Reinforcement learning intrusion detection method and system based on dynamic network feature screening
By adopting the feature screening method of adaptive reward reinforcement learning network based on Transformer framework and crayfish optimization algorithm in the intrusion detection system, the adaptability and accuracy of intrusion detection in dynamic network environments are solved, and more efficient attack recognition and detection performance is achieved.
Patent Information
- Application Number
- CN202510206124.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing intrusion detection technologies are difficult to effectively identify abnormal behaviors in dynamic network environments, especially in the face of sample imbalance and complex attack behaviors, which lack sufficient adaptability and accuracy.
Adaptive reward reinforcement learning network enhanced based on Transformer framework is adopted, and network traffic feature screening is combined with crayfish optimization algorithm. Through dynamic network feature screening and adaptive reward mechanism, the adaptability and detection performance of the model are improved.
It significantly improves the robustness, adaptability and overall performance of the intrusion detection system, can more accurately identify complex attack behaviors in a dynamically changing network environment, and improves the detection ability of a few types of attacks.
Smart Images

Figure CN119696934B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and in particular relates to a reinforcement learning intrusion detection method and system based on dynamic network feature screening. Background Art
[0002] In the field of network security, intrusion detection system (IDS) is an important defense measure designed to identify potential malicious activities or abnormal behaviors in the network. With the development of network technology and the continuous evolution of attack methods, the research on intrusion detection system faces more and more challenges, especially in high-dimensional, dynamic, and nonlinear network traffic data. How to effectively identify abnormal behaviors has become a problem that needs to be solved urgently.
[0003] Machine learning-based intrusion detection methods mainly identify potential attack behaviors by learning normal network behavior patterns, because attack behaviors usually deviate from normal patterns, and this deviation can be used for intrusion detection. For example, Ahmed Abdullah Alqarni et al. proposed a model combining support vector machine (SVM) and ant colony optimization (ACO) algorithm for intrusion detection, and Joseph Bamidele Awotunde et al. developed a multi-level random forest algorithm based on fuzzy inference system, combining the advantages of filter and wrapper methods, and using multi-level feature selection technology and fuzzy logic for intrusion classification. These methods rely on manual feature selection and supervised learning, and can identify potential attacks by modeling and analyzing network traffic. However, existing machine learning-based intrusion detection methods face some challenges. First, the sample imbalance problem is still a serious problem. Attack behaviors usually account for a small proportion of the data, which leads to the model's more accurate prediction of normal traffic and ignores the detection of attack traffic. Second, the feature selection process relies heavily on manual design. Although some methods try to improve detection accuracy through automated feature selection, their effectiveness is still constrained by feature quality. Finally, due to the constant changes in network environment and attack behaviors, machine learning models based on static training data lack sufficient adaptability and are unable to maintain high detection performance when facing new types of attacks.
[0004] With the development of deep neural networks, intrusion detection methods based on deep learning have been widely used in the field of intrusion detection. These methods can effectively identify complex attack behaviors by automatically extracting features from a large amount of historical data. For example, Ayesha S. Dina et al. proposed a deep learning method using the Focal Loss loss function, which can dynamically adjust the model's weights for difficult-to-classify negative examples; the BLoCNet model proposed by Brandon Bowen et al. combines convolutional neural networks (CNNs) and bidirectional long short-term memory networks (BLSTMs), enabling the intrusion detection system to identify characteristic patterns in network data in a shorter time; Hongpo Zhang et al. proposed a two-stage intrusion detection model that combines LightGBM and convolutional neural networks to solve the problem of class imbalance and improve the detection accuracy of abnormal traffic. Although deep learning methods perform well in feature extraction and attack identification, there are still some problems. First, deep learning models usually require a large amount of labeled data for training, and the labeling of network traffic data is often expensive and difficult. The lack of labeled data will affect the training effect of the model. Secondly, deep learning models are prone to overfitting, especially when there is insufficient training data, the generalization ability of the model is poor. Furthermore, deep learning algorithms usually require a lot of computing resources and time, which makes it difficult to perform real-time detection under large-scale network traffic data. Finally, deep learning methods are still plagued by data class imbalance, especially when attack traffic samples are scarce, the model may not be able to identify minority attacks.
[0005] Intrusion detection methods based on reinforcement learning have also begun to be applied in intrusion detection due to their adaptive learning capabilities, especially in dynamic environments. For example, the intrusion detection model based on dual deep Q network (CBL_DDQN) proposed by Sagar Dhanraj Pande et al. aims to solve the problem of imbalanced data between regular traffic and attack traffic and improve the detection rate of attack traffic; Antonio Maci et al. proposed a classifier based on dual deep Q network (DDQN) using imbalanced classification Markov decision process (ICMDP) to solve the data imbalance problem in phishing detection; Haonan Tan et al. constructed a dual experience replay mechanism with adaptive sample distribution through deep reinforcement learning to solve the problem of imbalanced distribution of traffic samples. Although reinforcement learning can respond flexibly in dynamic environments, its application also faces some challenges. First, the training efficiency of reinforcement learning in intrusion detection is low, especially in complex network environments. Training an efficient reinforcement learning model requires a lot of computing resources and time. Secondly, when faced with unbalanced data, reinforcement learning models are still prone to classify regular traffic, resulting in unsatisfactory detection of attack traffic. In addition, although reinforcement learning has certain adaptive capabilities, how to ensure that the model can quickly adapt and make accurate judgments when faced with rapid changes in network environment and attack behavior is still a problem that needs to be solved urgently.
[0006] In summary, existing intrusion detection technologies have significant deficiencies in practical applications, mainly in terms of their ability to process complex network traffic data and their adaptability in dynamic and nonlinear environments. First, most models based on traditional machine learning rely on artificially designed features and static data, which makes it difficult for them to cope with the rapid changes in network environments and attack behaviors. This static training method leads to insufficient recognition of new attacks by the model, and often exhibits poor detection performance when faced with data imbalance. Secondly, although deep learning-based methods have certain advantages in feature extraction, due to their reliance on large-scale labeled data and their high computing resource requirements, deep learning models face high training costs and difficulties in real-time deployment in practical applications. Moreover, the generalization ability of deep learning methods is often affected by sample imbalance, especially when there are fewer attack samples, the performance of the model is often difficult to maintain stability. Finally, although existing reinforcement learning methods have adaptive capabilities and can be optimized in dynamic environments, they still face the challenges of low training efficiency and difficulty in coping with sample imbalance. In addition, reinforcement learning models have poor adaptability to rapid changes in the environment, especially in complex network environments. How to quickly adjust strategies to cope with changing attack patterns is still an urgent problem to be solved. In summary, existing technologies have certain limitations when dealing with the diversity, dynamics and complexity of network traffic, and more flexible and intelligent solutions are urgently needed. Summary of the invention
[0007] In response to the above problems, the present invention provides a reinforcement learning intrusion detection method and system based on dynamic network feature screening, aiming to effectively solve the limitations of traditional intrusion detection technology, especially the adaptability and accuracy problems in the face of sample imbalance and complex attack behaviors in a dynamically changing network environment.
[0008] According to a first aspect of an embodiment of the present disclosure, a reinforcement learning intrusion detection method based on dynamic network feature screening is provided, the method comprising constructing an adaptive reward reinforcement learning network enhanced based on a Transformer framework, comprising the following steps:
[0009] Build a reinforcement learning network based on the training set data and adaptively determine the reinforcement learning reward mechanism;
[0010] Position encoding is performed on the input network traffic data features to enhance the temporal feature representation;
[0011] The network traffic data processed by position encoding is input into the intelligent agent based on the Transformer framework. The attack recognition results of the input network traffic data are generated through the feature extraction and decision network enhanced by the Transformer framework. The attack recognition results are compared with the actual labels, and the recognition accuracy and classification error rate are calculated. In combination with the adaptive reward mechanism, the intelligent agent is given corresponding rewards and punishments, and the rewards and punishments are fed back to the intelligent agent.
[0012] New samples are cyclically extracted and input into the agent to continuously update the state space and strengthen the learning of the experience pool until the agent model converges, completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework.
[0013] In some embodiments, the method further includes performing network traffic feature screening in combination with the crayfish optimization algorithm, including the following steps:
[0014] Initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and individual crayfish;
[0015] The temperature decay formula is used to guide the search to converge to the local optimum;
[0016] The flip probability is used to dynamically adjust the selection strategy of the feature subspace to optimize the exploration ability and fitness value of key features;
[0017] In the case that there are lobster individuals that do not produce better fitness values during the iteration process, a global random search is triggered to break the local convergence situation;
[0018] Based on the optimal feature set output after the iteration is completed, combined with the training set and the validation set, an adaptive reward reinforcement learning network enhanced based on the Transformer framework is constructed and trained for deployment applications.
[0019] In some embodiments, the reward mechanism for reinforcement learning includes:
[0020] The reward function expression is:
[0021] ,
[0022] ,
[0023] ,
[0024] Among them, X is a one-dimensional array, which represents the number of samples in each category. Formula Z is used to scale the number of samples to the same dimension. Formula Continue to scale based on formula Z, scaling different sample ratios to a preset maximum value U max and minimum valueU min In the range, the formula The sample ratio is further reversed and amplified by 10 times to obtain an integer as the final reinforcement learning reward function.
[0025] In some embodiments, the input network traffic data features are positionally encoded using a position encoding method based on sine and cosine functions to generate unique position information for each dimension of feature embedding, explicitly representing the time series relationship of the network traffic features.
[0026] In some embodiments, the initialization of the characteristic subspace, temperature, number of iterations and fitness value of the crayfish population and individual crayfish representation specifically includes:
[0027] The feature set is represented as individuals in a crayfish population, where each crayfish population has n crayfish O=[O1, O2, O3, …, O n ], characteristic state O i,j ∈{0,1} indicates whether the jth feature of the i-th lobster in the population is selected;
[0028] Set the initial temperature T and the maximum number of iterations max_iterations, use the reinforcement learning network to verify the feature subspace represented by each crayfish individual on the validation set, calculate its classification accuracy in the invasion detection task, and use the accuracy as the fitness value;
[0029] By comparing the fitness values, record the current best fitness value fitness best and the corresponding global optimal individual.
[0030] In some embodiments, the use of the temperature attenuation formula to guide the search to converge to the local optimum specifically includes:
[0031] The temperature attenuation formula is expressed as:
[0032] ,
[0033] In each iteration, the temperature T is updated by the temperature decay formula to achieve gradual cooling, where iteration represents the current number of iterations and max_iterations represents the maximum number of iterations. is the maximum temperature, and the temperature decay formula is used to ensure that the temperature gradually decreases with the increase of the number of iterations, thereby guiding the search process to converge from the initial extensive exploration to a more accurate local optimization.
[0034] In some embodiments, the method of dynamically adjusting the feature subspace selection strategy using the flip probability to optimize the exploration capability and fitness value of key features specifically includes:
[0035] The specific expression of the flip probability of the jth feature of the i-th lobster individual is:
[0036] ,
[0037] ,
[0038] Among them, the characteristic state O i,j ∈{0,1} indicates whether the jth feature of the i-th lobster in the population is selected. Yes i,j , k is the factor that controls the flip sensitivity, β is the scaling factor, and f i It indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set. rand() is a random number in the range of 0 to 1.
[0039] According to the flip probability, when a feature meets the conditions, its state will change from 0 to 1 or from 1 to 0, thereby dynamically adjusting the selection strategy of the feature subspace, improving the ability to explore key features and optimizing the detection performance on the validation set;
[0040] After completing the feature exploration, recalculate the fitness value of each lobster individual, that is, evaluate the accuracy of its invasion detection on the validation set. If the fitness value of the individual after exploration is better than the fitness value before exploration, update the individual and its corresponding fitness value;
[0041] Determine whether the individual fitness value after exploration exceeds the current global optimal fitness value; if it is better than the global optimal value, update the global optimal individual and its corresponding optimal fitness value at the same time to ensure that the algorithm gradually approaches a better solution.
[0042] In some embodiments, the global random search specifically includes:
[0043] If a lobster individual does not produce a better fitness value during N consecutive iterations, a global random search is triggered to break the local convergence situation. The expression of the global random search is:
[0044] ,
[0045] ,
[0046] in, is the characteristic state O i,j Update, rand j is a random number in the range [0,1], used to introduce randomness to the jth feature, α is the expansion factor, is the maximum temperature, f iis the fitness value, which indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set; if after random search, the fitness value f of the i-th lobster individual i If it is better than the original fitness value, the individual and its corresponding fitness value are updated, thereby improving the algorithm's ability to search for the global optimal solution.
[0047] According to a second aspect of an embodiment of the present disclosure, a reinforcement learning intrusion detection system based on dynamic network feature screening is provided, wherein the system includes constructing an adaptive reward reinforcement learning network unit enhanced based on a Transformer framework, specifically including:
[0048] A network building module is used to build a reinforcement learning network based on the training set data and adaptively determine the reward mechanism of reinforcement learning;
[0049] A data position encoding module is used to position encode the input network traffic data features to enhance the temporal characteristics representation;
[0050] The feature extraction and recognition module is used to input the network traffic data processed by position encoding into the intelligent agent based on the Transformer framework, generate the attack recognition results of the input network traffic data through the feature extraction and decision network enhanced by the Transformer framework, compare the attack recognition results with the actual labels, calculate the recognition accuracy and classification error rate, and combine the adaptive reward mechanism to give the intelligent agent corresponding rewards and punishments, and feedback to the intelligent agent;
[0051] The network training module is used to cyclically extract new samples and input them into the intelligent agent to continuously update the state space and strengthen the learning of the experience pool until the intelligent agent model converges, completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework.
[0052] In some embodiments, the system further includes a network traffic feature screening unit in combination with the crayfish optimization algorithm, specifically including:
[0053] Parameter initialization module, used to initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and individual crayfish;
[0054] Temperature decay module, used to guide the search to converge to the local optimum using the temperature decay formula;
[0055] Individual exploration module, which is used to dynamically adjust the selection strategy of feature subspace using flip probability, and optimize the exploration ability and fitness value of key features;
[0056] The global exploration module is used to trigger global random search to break the local convergence situation when there are lobster individuals that do not produce better fitness values during the iteration process;
[0057] Based on the new data set, the network and training module are built to build an adaptive reward reinforcement learning network enhanced based on the Transformer framework based on the optimal feature set output after the iteration, combined with the training set and the validation set, and then deployed for application after training.
[0058] The embodiments of the present disclosure provide a reinforcement learning intrusion detection method and system based on dynamic network feature screening, and the beneficial effects are as follows:
[0059] Based on the existing network intrusion detection technology, a new solution is proposed for feature selection, sample imbalance problems, and model adaptability and decision-making ability in dynamic network environments. Through the network lobster optimization and dimensionality reduction method based on lobster foraging behavior, and the adaptive reward reinforcement learning architecture based on Transformer enhancement, this invention has achieved breakthroughs at multiple levels and significantly improved the detection performance, stability and adaptability of the model.
[0060] Innovative feature selection and dimensionality reduction methods: Existing intrusion detection technologies often rely on manual design in feature selection and dimensionality reduction, or use traditional dimensionality reduction algorithms such as PCA, which usually ignore the complexity and time-varying characteristics of network traffic data. In contrast, the present invention innovatively designs a feature dimensionality reduction method based on the crayfish optimization algorithm, which can efficiently extract the most representative feature subsets from complex and changeable network traffic data. This method simulates the intelligent foraging behavior of lobsters and automatically screens out key features from massive network traffic, effectively avoiding the limitations of traditional methods, while greatly improving the level of automation of feature selection and reducing the need for human intervention.
[0061] Improved adaptability of reinforcement learning framework: In the prior art, many intrusion detection models based on reinforcement learning often face the problem of being unable to fully cope with the dynamic changes of network traffic and sample imbalance. Although traditional reinforcement learning methods such as Q-learning and DQN can handle decision-making problems, they are difficult to adapt quickly in a dynamic network environment. In order to overcome this problem, the present invention proposes an adaptive reinforcement learning framework based on Transformer enhancement. This method combines the Transformer mechanism with the DQN reinforcement learning method to significantly enhance the feature extraction and decision-making capabilities of the intelligent agent. Under the action of the Transformer's multi-head self-attention layer, this method can capture the global dependencies between network traffic features and improve the perception accuracy of the model in the multi-dimensional feature space. Compared with existing methods, this method enables the model to show stronger adaptability and decision-making accuracy in a dynamically changing network environment, thereby improving the accuracy and robustness of intrusion detection.
[0062] Effective solution to the sample imbalance problem: The sample imbalance problem has always been a difficulty in the field of intrusion detection, especially when there are fewer attack samples, traditional models often find it difficult to effectively identify minority attack events. In order to solve this problem, the present invention innovatively proposes an adaptive reinforcement learning reward optimization mechanism, which can adaptively determine the reward value according to the sample sparsity, helping the intelligent agent to achieve an effective balance between exploration and utilization. Specifically, the adjustment of the reward value is achieved through strategies such as transformation, translation, and scaling. This adaptive mechanism can help the model better learn minority attacks and significantly improve the detection capability under sample imbalance conditions.
[0063] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention;
[0065] Figure 1 It is a schematic diagram of the process of the reinforcement learning intrusion detection method based on dynamic network feature screening in an embodiment of the present invention;
[0066] Figure 2 Schematic diagram of the structure of a reinforcement learning intrusion detection system based on dynamic network feature screening in an embodiment of the present invention;
[0067] Figure 3 It is a schematic diagram of the structure of an adaptive reward reinforcement learning network unit based on the Transformer framework enhancement in an embodiment of the present invention;
[0068] Figure 4 It is a schematic diagram of the structure of a network traffic feature screening unit in combination with a crayfish optimization algorithm in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0070] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0071] The present invention specifically targets intrusion detection tasks in dynamic network environments, and aims to solve the technical difficulties of existing technologies in terms of sample imbalance, inefficient feature selection, and insufficient adaptability to complex attack patterns. Traditional intrusion detection methods are difficult to accurately extract key features from dynamically changing network traffic, and in the case of uneven sample distribution, the model tends to ignore sparse but important attack samples, resulting in decreased detection performance. In addition, when facing a dynamic network environment, existing reinforcement learning methods have problems with unstable decision-making processes and limited generalization capabilities. The present invention innovatively introduces a network lobster optimization and dimensionality reduction method based on lobster foraging behavior, combined with an adaptive reward reinforcement learning architecture based on Transformer enhancement, to provide an efficient and intelligent intrusion detection solution, which effectively improves the robustness, adaptability and overall performance of the detection system.
[0072] The purpose of the present invention is to propose an innovative method to effectively solve the limitations of traditional intrusion detection technology, especially in the dynamically changing network environment, the adaptability and accuracy problems when facing sample imbalance and complex attack behaviors. To this end, the present invention designs a network lobster optimization and dimensionality reduction method based on lobster foraging behavior, through which the most representative feature subset can be efficiently extracted from complex and changeable network traffic data, thereby reducing the dependence on manual feature selection and improving the automation and accuracy of the feature selection process. In addition, the present invention proposes an adaptive reward reinforcement learning architecture based on Transformer enhancement, which significantly enhances the adaptive learning and decision optimization capabilities in a dynamic network environment by combining the Transformer mechanism with reinforcement learning, and significantly improves the perception accuracy and decision-making ability of the model in a multidimensional feature space, especially in the face of complex attack patterns. It shows stronger robustness and generalization ability. In order to further solve the problem of sample imbalance, the present invention also proposes an adaptive reward mechanism to help the model better cope with sample sparsity in dynamically changing network traffic, balance the relationship between exploration and utilization, thereby improving decision-making efficiency and improving the overall performance of the intrusion detection system. Through these innovations, the present invention can effectively overcome the deficiencies of the prior art in terms of sample imbalance, feature selection, dynamic adaptability, etc., and provide a more intelligent and efficient intrusion detection solution.
[0073] The present invention provides the following embodiments for a reinforcement learning intrusion detection method and system based on dynamic network feature screening:
[0074] A reinforcement learning intrusion detection method based on dynamic network feature screening. The embodiment combines the crayfish optimization algorithm to extract features from high-dimensional network traffic to capture the most valuable feature set for intrusion detection, and proposes a reinforcement learning network architecture based on adaptive rewards enhanced by the Transformer framework, which significantly improves the perception accuracy and decision-making ability of the model in multi-dimensional feature space.
[0075] The specific implementation steps are as follows: Figure 1 As shown, the method includes constructing an adaptive reward reinforcement learning network based on the Transformer framework enhancement, including the following steps:
[0076] S101: Build a reinforcement learning network based on training set data and adaptively determine the reward mechanism for reinforcement learning;
[0077] Specifically, firstly, the reward function of reinforcement learning is adaptively determined according to the number of samples in each category of the data set (including normal samples and attack samples of different categories). The calculation method is as follows:
[0078] ,
[0079] ,
[0080] ,
[0081] Among them, X is a one-dimensional array that records the number of samples in each category. Due to the extremely unbalanced samples in the intrusion detection scenario, some highly concealed attacks are hidden in a large number of benign samples (the number of benign samples may be hundreds of thousands of times the number of highly concealed attack samples), so it is necessary to use the formula Scale the sample size to the same dimension. Formula In the formula Continue to scale based on the specified U max and U min Within the range, U max and U min The values of are set to 1 and 0.1 respectively. The sample ratio is further processed in reverse order and amplified by 10 times to take the integer as the final reinforcement reward function. The purpose is to make the reinforcement learning framework pay more attention to minority samples, that is, the higher the concealment, the higher the attack reward value.
[0082] S102: Position encoding the input network traffic data features to enhance the temporal feature representation;
[0083] Specifically, a batch of input network traffic data features are positionally encoded to enhance the representation of temporal characteristics. A position encoding method based on sine and cosine functions is used to generate unique position information for each dimension of feature embedding, explicitly representing the temporal relationship of network traffic features. Through this step, the position encoding is combined with the original features to provide more accurate temporal information support for subsequent feature extraction and decision optimization.
[0084] S103: Input the network traffic data processed by position encoding into the intelligent agent based on the Transformer framework for attack detection. Use the multi-head self-attention mechanism to capture the global dependencies between network traffic features and further enhance the distinguishing ability of features. Combine residual connections, normalization and feedforward neural network layers to improve the robustness and expressiveness of feature extraction.
[0085] S104: In step S103, the attack recognition results of the input network traffic data are generated through the feature extraction and decision network enhanced by the Transformer framework. In this step, the attack recognition results output by the agent are compared with the actual labels, and the recognition accuracy and classification error rate are calculated. Combined with the adaptive reward mechanism designed in step S101, the agent is given corresponding rewards and punishments according to the correctness of the recognition results, and feedback is given to the agent. In this way, the agent can achieve a balance between exploring unknown features and optimizing existing strategies, gradually improve the attack recognition ability of complex network traffic, and effectively deal with the problem of uneven sample distribution.
[0086] S105: Re-extract a batch of samples from the data set as the new state space input to the agent. The agent performs actions by interacting with the environment, obtains feedback according to the reward mechanism, and stores information such as the current state, action, reward value, and next state in the experience pool. The experience pool supports the policy update of the model by recording the complete state transition process, and prioritizes the learning of key samples. The above process is repeated until the model converges, completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework.
[0087] like Figure 1 As shown, the method also includes combining the crayfish optimization algorithm to perform network traffic feature screening, including the following steps:
[0088] S106: Initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and individual crayfish;
[0089] Specifically, the feature set is represented as individuals in a crayfish population, where each crayfish population has n crayfish O=[O1, O2, O3, …, O n ], characteristic state O i,j ∈{0,1} indicates whether the jth feature of the i-th lobster in the population is selected (1 for selection, 0 for elimination). The initial temperature T is set to 100, and the maximum number of iterations max_iterations is set to 500. The reinforcement learning network constructed using the above process is verified on the validation set for the feature subspace represented by each crayfish individual, and its classification accuracy in the invasion detection task is calculated, and the accuracy is used as the fitness value. By comparing the fitness values, the current best fitness value fitness is recorded best and the corresponding global optimal individual.
[0090] S107: Use the temperature decay formula to guide the search to converge to the local optimum. Specifically, in each iteration, the temperature T is updated by the following formula to achieve gradual cooling. Among them, iterations represents the current number of iterations, and max_iterations is the maximum number of iterations. This formula ensures that the temperature gradually decreases as the number of iterations increases, thereby guiding the search process from the initial extensive exploration to a more accurate local optimization convergence;
[0091] ,
[0092] S108: Use the flip probability to dynamically adjust the selection strategy of the feature subspace to optimize the exploration ability and fitness value of key features. Specifically, in each iteration, when exploring the feature subspace, the flip probability of the jth feature of the i-th lobster individual is dynamically calculated by the following formula, where k is a factor that controls the flip sensitivity (set to 1), β is a scaling factor (set to 0.03), and f i It indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set, and rand() is a random number ranging from 0 to 1. According to the flip probability, when a feature meets the conditions, its state will change from 0 to 1 or from 1 to 0, thereby dynamically adjusting the selection strategy of the feature subspace, further improving the ability to explore key features and optimizing the detection performance on the validation set.
[0093] The specific expression of the flip probability of the jth feature of the i-th lobster individual is:
[0094] ,
[0095] ,
[0096] Among them, the characteristic state O i,j ∈{0,1} indicates whether the jth feature of the i-th lobster in the population is selected. Yes i,j , k is the factor that controls the flip sensitivity, β is the scaling factor, and f i It represents the intrusion detection accuracy of the current lobster individual on the validation set based on the selected feature set. rand() is a random number in the range of 0 to 1.
[0097] After completing the feature exploration, it is necessary to recalculate the fitness value of each lobster individual, that is, to evaluate the accuracy of its intrusion detection on the validation set. If the fitness value of the individual after exploration is better than the fitness value before exploration, the individual and its corresponding fitness value are updated. In addition, it is also necessary to determine whether the fitness value of the individual after exploration exceeds the current global optimal fitness value; if it is better than the global optimal value, the global optimal individual and its corresponding optimal fitness value are updated at the same time, so as to ensure that the algorithm gradually approaches a better solution.
[0098] S109: In the case that a lobster individual does not produce a better fitness value during the iteration process, a global random search is triggered to break the local convergence situation. Specifically, in order to balance feature exploration and population competition and avoid the algorithm from falling into a local optimal solution, a global random search mechanism is set up: if a lobster individual does not produce a better fitness value during 20 consecutive iterations, a global random search is triggered to break the local convergence situation.
[0099] If a lobster individual does not produce a better fitness value during N consecutive iterations, a global random search is triggered to break the local convergence situation. The expression of the global random search is:
[0100] ,
[0101] ,
[0102] in, is the characteristic state O i,j Update, rand j is a random number in the range [0,1], used to introduce randomness to the jth feature, α is the expansion factor, is the maximum temperature, f i is the fitness value, which indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set; if after random search, the fitness value f of the i-th lobster individual iIf it is better than the original fitness value, the individual and its corresponding fitness value are updated, thereby improving the algorithm's ability to search for the global optimal solution.
[0103] S110: Repeat steps S106 to S109 until one of the following termination conditions is met: the global optimal fitness value does not change after 50 consecutive iterations, or the preset maximum number of iterations is reached. After the algorithm ends, the current optimal lobster individual, the selected optimal feature set, and the classification accuracy of the feature set on the validation set are output.
[0104] S111: Combine the training set and the validation set to form a new data set, and re-execute steps S101 to S105 based on the optimal feature set selected in step S110 to train a new reinforcement learning model. Finally, the optimized model is deployed to actual application scenarios for real-time intrusion detection and network protection.
[0105] Another embodiment is used to illustrate a reinforcement learning intrusion detection system based on dynamic network feature screening, such as Figure 2 As shown, the system 300 includes a Transformer framework-based enhanced adaptive reward reinforcement learning network unit 310 and a network traffic feature screening unit 320 combined with a crayfish optimization algorithm, wherein the Transformer framework-based enhanced adaptive reward reinforcement learning network unit 310 is constructed, as shown in FIG. Figure 3 As shown, specifically including:
[0106] A network construction module 311 is used to construct a reinforcement learning network based on training set data and adaptively determine a reward mechanism for reinforcement learning;
[0107] A data position encoding module 312 is used to position encode the input network traffic data features to enhance the temporal characteristics representation;
[0108] The feature extraction and recognition module 313 is used to input the network traffic data processed by the position encoding into the intelligent agent based on the Transformer framework, generate the attack recognition result of the input network traffic data through the feature extraction and decision network enhanced by the Transformer framework, compare the attack recognition result with the actual label, calculate the recognition accuracy and classification error rate, and combine the adaptive reward mechanism to give the intelligent agent corresponding rewards and punishments, and feedback to the intelligent agent;
[0109] The network training module 314 is used to cyclically extract new samples and input them into the intelligent agent to continuously update the state space and strengthen the learning of the experience pool until the intelligent agent model converges, thus completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework.
[0110] like Figure 4 As shown, the network traffic feature screening unit 320 is combined with the crayfish optimization algorithm, specifically including:
[0111] Parameter initialization module 321, used to initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and crayfish individual representation;
[0112] A temperature decay module 322, for guiding the search to converge to a local optimum using a temperature decay formula;
[0113] Individual exploration module 323, used to dynamically adjust the selection strategy of feature subspace using flip probability, and optimize the exploration ability and fitness value of key features;
[0114] The global exploration module 324 is used to trigger a global random search to break the local convergence situation when there are lobster individuals that do not produce better fitness values during the iteration process;
[0115] A network and training module 325 is constructed based on the new data set, which is used to construct an adaptive reward reinforcement learning network enhanced based on the Transformer framework based on the optimal feature set output after the iteration is completed, combined with the training set and the verification set, and then used for deployment application after training.
[0116] In addition to the upper module, the system 300 may also include other components. However, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.
[0117] The other specific working processes of the system 300 refer to the description of the above-mentioned basic network traffic detection method embodiment and will not be repeated here.
[0118] Based on the technical solutions provided by the above embodiments, a reinforcement learning intrusion detection method and system based on dynamic network feature screening, the present invention is based on the optimization and dimensionality reduction method of lobster foraging behavior, effectively extracts the most representative features in network traffic, and improves the accuracy and stability of the intrusion detection model; current network intrusion detection technologies mostly rely on traditional feature selection and dimensionality reduction methods, such as principal component analysis (PCA), linear discriminant analysis (LDA), etc. These methods usually face the risk of information loss in high-dimensional data, and fail to dynamically adapt to complex network traffic data features. The dimensionality reduction method in the prior art is usually based on static analysis and lacks the ability to dynamically perceive data changes, resulting in poor feature selection and dimensionality reduction effects of the model when dealing with a changeable network environment, affecting the accuracy and stability of the subsequent detection model. The present invention innovatively proposes a feature dimensionality reduction method based on a crayfish optimization algorithm. By simulating the dynamic behavior of lobster foraging, the method can adaptively extract the most representative feature subset from complex and changeable network traffic data. Different from the traditional method, the proposed Transformer-enhanced adaptive reward reinforcement learning Q network is embedded into the network lobster optimization algorithm in feature selection to perform feature screening on the training set and validation set, and the best network feature combination obtained by screening is applied to the test set and the actual scenario after deployment. This method can effectively reduce redundant features and improve the performance of network traffic data in subsequent intrusion detection models, especially in dynamically changing network environments, showing stronger adaptability and stability.
[0119] The present invention is based on a Transformer-enhanced reinforcement learning architecture to improve the adaptability and decision-making accuracy in a dynamic network environment; Implementation scheme of the prior art: Traditional reinforcement learning methods, such as Q-learning and DQN, usually have certain limitations when dealing with dynamically changing network environments. Since these methods fail to effectively process the temporal information and global dependencies of network traffic features, their adaptability in complex network environments is weak. In addition, traditional reinforcement learning methods often face instability caused by frequent updates in the decision-making process, and cannot fully deal with the problem of sample imbalance, which limits the application effect of the prior art in intrusion detection. The present invention proposes an adaptive reward reinforcement learning architecture based on Transformer enhancement, which enhances the temporal dependency and global relationship modeling ability of the agent for network traffic features by introducing the Transformer mechanism and combining it with the agent in reinforcement learning. Specifically, the method explicitly captures and transmits the sequence information of network traffic features through an absolute position encoding method based on sine and cosine functions, and deeply mines the global dependencies between features through a multi-head self-attention mechanism, which greatly enhances the feature extraction and decision-making ability of the agent, and significantly enhances the adaptability and decision-making accuracy of the model in a dynamic network environment. Compared with the existing technology, the present invention has made significant breakthroughs in the time series modeling and feature extraction of network traffic data, improved the adaptability and decision-making accuracy of the model in a dynamic network environment, and especially can effectively capture the changes in attack patterns in the multi-dimensional feature space, thereby improving the accuracy and robustness of the intrusion detection system.
[0120] The present invention sets an adaptive reinforcement learning reward optimization mechanism to optimize the sample imbalance problem and improve the detection performance; most current reinforcement learning methods usually use a static reward function to optimize the training process when dealing with the sample imbalance problem, but this method often cannot fully cope with the scarcity of minority class samples, resulting in poor results of the model in detecting minority class attacks. Although some methods attempt to generate more minority class samples to balance the sample imbalance problem through oversampling technology, if the minority class samples contain noise data, the oversampling method may generate more noise-based samples, thereby amplifying the impact of the noise and reducing the performance of the model. Implementation scheme of the present invention: In order to effectively deal with the sample imbalance problem, the present invention proposes an adaptive reinforcement learning reward optimization mechanism that can adaptively determine the reward value of reinforcement learning according to the sample sparsity during the training process (the reward mechanism will allocate a higher reward value to the minority class samples to increase the model's attention to the minority class samples), and use strategies such as transformation, translation and scaling to help the intelligent agent achieve an effective balance between exploration and utilization. Compared with the prior art, the adaptive reward mechanism of the present invention can flexibly adjust the reward signal during the training process, further optimize the impact of sample imbalance on model performance, and significantly improve the detection accuracy and robustness of the model in diversified attack scenarios. Through this innovative combination of technologies, the present invention demonstrates better processing capabilities and detection effects when facing the problem of sample imbalance.
[0121] In this document, the terms "comprises," "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such step or method.
[0122] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A reinforcement learning intrusion detection method based on dynamic network feature screening, characterized in that: The method includes constructing an adaptive reward reinforcement learning network based on a Transformer framework enhancement, including the following steps: Build a reinforcement learning network based on the training set data and adaptively determine the reinforcement learning reward mechanism; Position encoding is performed on the input network traffic data features to enhance the temporal feature representation; The network traffic data processed by position encoding is input into the intelligent agent based on the Transformer framework for attack detection. The attack identification results of the input network traffic data are generated through the feature extraction and decision network enhanced by the Transformer framework. The attack identification results are compared with the actual labels, and the identification accuracy and classification error rate are calculated. In combination with the adaptive reward mechanism, the intelligent agent is given corresponding rewards and punishments, and the rewards and punishments are fed back to the intelligent agent. New samples are cyclically extracted and input into the agent to continuously update the state space and strengthen the learning of the experience pool until the agent model converges, completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework; The method also includes combining the crayfish optimization algorithm to perform network traffic feature screening, including the following steps: Initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and individual crayfish; The temperature decay formula is used to guide the search to converge to the local optimum; The flip probability is used to dynamically adjust the selection strategy of the feature subspace to optimize the exploration ability and fitness value of key features; In the case that there are lobster individuals that do not produce better fitness values during the iteration process, a global random search is triggered to break the local convergence situation; Based on the optimal feature set output after the iteration, combined with the training set and the validation set, an adaptive reward reinforcement learning network based on the Transformer framework is constructed and trained for deployment applications. Global random search specifically includes: If a lobster individual does not produce a better fitness value during N consecutive iterations, a global random search is triggered to break the local convergence situation. The expression of the global random search is: , , in, It is the characteristic state Update, rand j is a random number in the range [0,1], used to introduce randomness to the jth feature, α is the expansion factor, is the maximum temperature, f i is the fitness value, which indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set; if after random search, the fitness value f of the i-th lobster individual i If it is better than the original fitness value, the individual and its corresponding fitness value are updated, thereby improving the algorithm's ability to search for the global optimal solution.
2. The reinforcement learning intrusion detection method based on dynamic network feature screening according to claim 1 is characterized in that: The reward mechanism of the reinforcement learning includes: The reward function expression is: , , , Among them, X is a one-dimensional array, which represents the number of samples in each category. Formula Z is used to scale the number of samples to the same dimension. Formula Continue to scale based on formula Z, scaling different sample ratios to a preset maximum value U max and minimum value U min In the range, the formula The sample ratio is further reversed and amplified by 10 times to obtain an integer as the final reinforcement learning reward function.
3. The reinforcement learning intrusion detection method based on dynamic network feature screening according to claim 1 is characterized in that: The position encoding of the input network traffic data features adopts a position encoding method based on sine and cosine functions to generate unique position information for each dimension of feature embedding, and explicitly represents the time series relationship of the network traffic features.
4. The reinforcement learning intrusion detection method based on dynamic network feature screening according to claim 1 is characterized in that: The initialization of the characteristic subspace, temperature, number of iterations and fitness value of the crayfish population and individual crayfish representation specifically includes: Represent the feature set as individuals in a crayfish population, where each crayfish population has n crayfish , feature state Indicates whether the jth feature of the i-th lobster in the population is selected; Set the initial temperature T and the maximum number of iterations max_iterations, use the reinforcement learning network to verify the feature subspace represented by each crayfish individual on the validation set, calculate its classification accuracy in the invasion detection task, and use the accuracy as the fitness value; By comparing the fitness values, record the current best fitness value fitness best and the corresponding global optimal individual.
5. The reinforcement learning intrusion detection method based on dynamic network feature screening according to claim 1 is characterized in that: The method of using the temperature attenuation formula to guide the search to converge to the local optimum specifically includes: The temperature attenuation formula is expressed as: , During each iteration, the temperature is updated by the temperature decay formula To achieve gradual cooling, where iteration represents the current number of iterations and max_iterations represents the maximum number of iterations. is the maximum temperature, and the temperature decay formula is used to ensure that the temperature gradually decreases with the increase of the number of iterations, thereby guiding the search process to converge from the initial extensive exploration to a more accurate local optimization.
6. The reinforcement learning intrusion detection method based on dynamic network feature screening according to claim 1 is characterized in that: The method of dynamically adjusting the feature subspace selection strategy by using the flip probability to optimize the exploration capability and fitness value of key features specifically includes: The specific expression of the flip probability of the jth feature of the i-th lobster individual is: , , Among them, the characteristic state Indicates whether the jth feature of the i-th lobster in the population is selected, Yes , k is the factor that controls the flip sensitivity, β is the scaling factor, and f i It indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set, and rand() is a random number in the range of 0 to 1; According to the flip probability, when a feature meets the conditions, its state will change from 0 to 1 or from 1 to 0, thereby dynamically adjusting the selection strategy of the feature subspace, improving the ability to explore key features and optimizing the detection performance on the validation set; After completing the feature exploration, recalculate the fitness value of each lobster individual, that is, evaluate the accuracy of its invasion detection on the validation set. If the fitness value of the individual after exploration is better than the fitness value before exploration, update the individual and its corresponding fitness value; Determine whether the individual fitness value after exploration exceeds the current global optimal fitness value; if it is better than the global optimal value, update the global optimal individual and its corresponding optimal fitness value at the same time to ensure that the algorithm gradually approaches a better solution.
7. A reinforcement learning intrusion detection system based on dynamic network feature screening, characterized in that: The system includes building an adaptive reward reinforcement learning network unit based on the Transformer framework enhancement, specifically including: A network building module is used to build a reinforcement learning network based on the training set data and adaptively determine the reward mechanism of reinforcement learning; A data position encoding module is used to position encode the input network traffic data features to enhance the temporal characteristics representation; The feature extraction and recognition module is used to input the network traffic data processed by position encoding into the intelligent agent based on the Transformer framework for attack detection. The attack recognition results of the input network traffic data are generated through the feature extraction and decision network enhanced by the Transformer framework. The attack recognition results are compared with the actual labels, and the recognition accuracy and classification error rate are calculated. In combination with the adaptive reward mechanism, the intelligent agent is given corresponding rewards and punishments, and the rewards and punishments are fed back to the intelligent agent. The network training module is used to cyclically extract new samples and input them into the agent to continuously update the state space and strengthen the learning of the experience pool until the agent model converges, completing the training of the adaptive reward reinforcement learning network enhanced based on the Transformer framework; The system also includes a network traffic feature screening unit in combination with the crayfish optimization algorithm, specifically including: Parameter initialization module, used to initialize the characteristic subspace, temperature, number of iterations and fitness value of crayfish population and individual crayfish; Temperature decay module, used to guide the search to converge to the local optimum using the temperature decay formula; Individual exploration module, which is used to dynamically adjust the selection strategy of feature subspace using flip probability, and optimize the exploration ability and fitness value of key features; The global exploration module is used to trigger global random search to break the local convergence situation when there are lobster individuals that do not produce better fitness values during the iteration process; Build a network and training module based on the new data set, which is used to build an adaptive reward reinforcement learning network based on the Transformer framework based on the optimal feature set output after the iteration, combined with the training set and the validation set, and then deploy the network after training; Global random search specifically includes: If a lobster individual does not produce a better fitness value during N consecutive iterations, a global random search is triggered to break the local convergence situation. The expression of the global random search is: , , in, It is the characteristic state Update, rand j is a random number in the range [0,1], used to introduce randomness to the jth feature, α is the expansion factor, is the maximum temperature, f i is the fitness value, which indicates the intrusion detection accuracy of the current lobster individual based on the selected feature set on the validation set; if after random search, the fitness value f of the i-th lobster individual i If it is better than the original fitness value, the individual and its corresponding fitness value are updated, thereby improving the algorithm's ability to search for the global optimal solution.
Citation Information
Patent Citations
Garbage detection method of neural network model based on deep learning
CN116258905A
Intrusion detection method based on deep reinforcement learning and structured data Transform
CN118740475A