Intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning

By employing the SEA-DDQN adaptive reinforcement learning method, combined with attention and dynamic weight reward mechanisms, the network security situation prediction model is optimized, solving the problem of insufficient adaptability of traditional methods in dynamic network environments and achieving more efficient intrusion situation prediction.

CN120979838AActive Publication Date: 2025-11-18CHENGDU UNIV OF INFORMATION TECH

Patent Information

Application Number
CN202511493849.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Traditional machine learning methods are not adaptable enough to predict network security situations, struggle to handle dynamic network environments and new types of attacks, and consume high computational resources.

Method used

We employ an adaptive reinforcement learning approach based on SEA-DDQN, using a dual-value deep attention Q-network (SEA-DDQN) model that combines attention and dynamic weight reward mechanisms to optimize the model's adaptability and decision-making capabilities for predicting network security intrusion scenarios.

Benefits of technology

It improves the model's predictive and anti-interference capabilities in complex threat scenarios, enhances its adaptability to dynamic environments, reduces computational resource consumption, and improves prediction accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979838A_ABST
    Figure CN120979838A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security situation awareness, and discloses an intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning, and the method comprises the following steps: obtaining a network flow data set; preprocessing the obtained network flow data set to obtain a standardized data set; and constructing a deep attention Q network SEA-DDQN model based on double values, and carrying out network security intrusion situation prediction ISP to realize an ISP process. According to the intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning, through continuous interaction with a traffic environment, the prediction capability is continuously improved, so that the model can flexibly cope with complex and continuously changing threat scenes; the SEA-DDQN uses the advantages of reinforcement learning to optimize the adaptivity and decision making process of the model, thereby enhancing the anti-interference capability of the model in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security situation awareness, and in particular to an intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning. BACKGROUND

[0002] With the increasing complexity and diversity of network threats, the importance of cyber security situation awareness (CSA) is increasingly prominent. As a key component of CSA, situation prediction bears the key responsibility of identifying and responding to potential security threats in the future. Traditional machine learning (ML) methods usually rely on supervised learning on network security data to predict and identify security incidents. However, these methods often face challenges such as insufficient adaptability and limited generalization ability when dealing with dynamic network environments and their changing attack patterns, which limits their effectiveness in real-time monitoring and emergency response.

[0003] Therefore, in order to solve the problems of dependence on labeled data, insufficient prediction ability for new attacks, poor adaptability in dynamic environments, and increased computational resource consumption and reduced training efficiency caused by "state explosion" when dealing with large-scale data sets due to the high dimensionality of attack behavior characteristics in existing supervised learning and unsupervised learning methods for intrusion situation prediction (ISP), the present application proposes a deep attention Q network (SEA-DDQN) adaptive reinforcement learning framework based on double value to predict network security intrusion situation. SUMMARY

[0004] The purpose of the present application is to provide an intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning, which continuously improves its prediction ability through continuous interaction with the traffic environment, enabling the model to flexibly cope with complex and changing threat scenarios; SEA-DDQN utilizes the advantages of reinforcement learning to optimize the adaptability and decision-making process of the model, thereby enhancing the model's anti-interference ability in dynamic environments.

[0005] To achieve the above purpose, the present application provides an intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning, comprising the following steps: Step S1, obtaining NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020 network traffic dataset; Step S2, preprocessing the obtained network traffic dataset to obtain a normalized dataset; Step S3, constructing a deep attention Q network SEA-DDQN model based on double value to predict network security intrusion situation ISP; Step S4, based on the SEA-DDQN model, the ISP process is realized.

[0006] Preferably, in step S2, the obtained network traffic dataset is preprocessed to obtain a normalized dataset, and the specific process is as follows: Step S21, the network traffic dataset is defined as , wherein is the number of traffic samples, and the feature space is , as follows: ; , wherein represents a feature, represents an index of the number of features in the dataset; represents a feature subset; ; is a subset of traffic features; is a subset of content features; is a subset of time statistical features; is a subset of host behavior features; represents a union symbol, indicating that the four feature subsets are combined into a feature space ; Step S22, for discrete features , hot one coding processing is performed; First, the discrete feature subset ; wherein represents a discrete feature; represents a discrete feature index; the corresponding value range of each discrete feature is respectively: ; ; ; wherein , and respectively represent the value range of the discrete feature, ; for any discrete feature, define the indicator function , as follows: ; , wherein ; represents the value of the sample in the discrete feature ; Then, through hot one coding transformation , i.e. , the feature space is reconstructed, as follows: ; , wherein d is the total number of features after hot one coding processing; represents a hot one coding processing function; denotes a vector concatenation operation; is the index of all non-discrete features in the original features; is the subset of discrete features; Finally, the original training set and test set are respectively subjected to hot one coding to generate their respective extended feature matrices , and are normalized as follows: ; wherein, denotes the normalized features; and denote the maximum and minimum values of the features , respectively; is an indicator function, taking the value 1 when , and 0 otherwise; the attack classes in the data set are classified and mapped into a set of numerical classes .

[0007] Preferably, based on the NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020 training sets and test sets after data preprocessing, a network ISP reinforcement learning environment is constructed and defined; the ISP process is abstracted as a Markov decision process MDP, and is formally defined as a four-tuple , as follows: ; wherein, the state is a set of feature vector spaces of each network flow; the action is a set of flow type decisions defined in NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020, at time The state and action of the agent are respectively denoted as and ; the reward is the reward or punishment obtained by the agent for predicting the flow correctly or not; the calculation of the reward vector is based on the comparison between the prediction result of the agent and the actual label to evaluate the accuracy of the prediction of the agent; the discount factor is used to balance the discount coefficient of the immediate reward and the long-term return.

[0008] Preferably, a SEA-DDQN model agent is constructed, including an action selection network and a target Q value network; the SEA-DDQN model agent performs prediction and target Q value calculation on the captured network flow; the action selection network and the target Q value network The Q value functions of the SEA-DDQN model are respectively composed of four-layer fully connected feedforward neural networks, and attention mechanisms and ReLu activation functions are used between the fully connected layers, wherein two hidden layers each contain 128 neurons; The state of the SEA-DDQN model agent at each time step is represented by a feature vector of network traffic , and the state space of the agent is defined as Therefore, the state vector of each piece of network traffic obtained by the SEA-DDQN model agent is , as shown below: ; wherein each dimension represents a certain feature of the network traffic.

[0009] Preferably, an attention mechanism is introduced to optimize the attention allocation of the SEA-DDQN model agent to the network traffic features, enhance discriminative features and suppress irrelevant features, and the specific process is as follows: For the network traffic feature state vector input at each time step, after processing by the fully connected layer, a 128-dimensional feature vector is obtained, wherein represents the network traffic feature vector processed by the fully connected layer, represents the dimension index; then it is sent to the attention layer to enhance the attention to important features, and the attention weight is as follows: ; wherein is a dimension reduction projection matrix; is a dimension increase reconstruction matrix; represents the dimension reduction ratio, which is a hyperparameter; represents a Sigmoid gating function, which compresses the input value to the interval , and the generated attention weight is used to measure the importance of each feature; the generated attention weight is fused with the original feature by modulation, as shown below: ; wherein represents the fused feature representation; is a Hadamard product, i.e., an element-wise multiplication operation; this operation dynamically adjusts each feature value by the corresponding weight .

[0010] Preferably, the SEA-DDQN model agent generates an action list in the form of an action vector according to the input features of the deep neural network , and the final Q value is used to evaluate whether the attack behavior is successfully predicted; First, the class labels in the NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020 datasets are respectively mapped into a set of digital categories as follows: ; wherein 0 represents Normal; 1 represents DoS; 2 represents Probe; 3 represents U2R; and 4 represents R2L; ; wherein 0 represents normal traffic; and 1 represents abnormal traffic; ; wherein 0 represents Bots; 1 represents Brute Force; 2 represents DDos; 3 represents Dos; 4 represents Normal; 5 represents Port Scanning; and 6 represents Web Attacks; ; wherein 0 represents Normal; 1 represents Bruteforce; 2 represents Scan_A; 3 represents Scan_sU; and 4 represents Sparta; Then, the action space of the SEA-DDQN model agent is defined as: ; ; ; ; wherein each action corresponds to a predicted class decision in the action list; represents the class index.

[0011] Preferably, a dynamic weight reward mechanism is designed to balance the class bias by constructing a class-sensitive reward, and an inverse frequency weighting strategy is used to determine the class weight to enhance the sensitivity of the SEA-DDQN model agent to different attack classes, balance the influence of class bias in the dataset on the learning process, and the specific process is as follows: First, the specific form of the reward function is defined as follows: ; wherein represents the weight of the attack class; represents correct prediction; represents incorrect prediction; Then, the inverse frequency weighting strategy is used to assign weights by considering the frequency of occurrence of each attack category, as follows: ; ; where, represents a set of attack category weights of represents the weight of attack represents the weight of no attack represents the weight of attack represents the weight of attack represents the weight of attack represents a set of attack category weights of represents the weight of attack represents the weight of attack represents the weight of attack represents the weight of attack represents the weight of attack The weight rule is updated as follows: ; where, is the total number of samples, is the number of attack types, represents the total number of attack samples; For the two prediction tasks, a positive reward is given when the SEA-DDQN model agent correctly predicts the sample, and a negative penalty is given otherwise, as follows: ; where, represents the reward setting rule of each flow in the dataset; In order to simulate the delay feedback scene in the real network, the n-step reward accumulation mechanism is adopted, and the reward actually stored in the experience replay pool is the n-step cumulative reward , rather than the single-step immediate reward ; the agent stores the transition sequence of the last n steps in the process of interacting with the environment, and calculates the discounted cumulative reward after reaching n steps , as follows: .

[0012] Preferably, a priority sampling strategy PER is adopted to optimize the sampling process of empirical network traffic. Due to the importance difference of attack categories, TD-Error is introduced as a key indicator to measure the importance of samples. By defining a priority update formula, the dynamic adjustment of sample sampling probability is realized, as follows: ; wherein, denotes the sampling probability of the i-th empirical network traffic; is a constant to ensure that all empirical network traffic has a sampling probability, avoiding that some samples have a sampling probability of zero due to too small TD-Error; is the number of network traffic in the current experience replay pool; is a priority parameter to control the influence degree of TD-Error on the sampling probability; is the TD-Error corresponding to the i-th empirical network traffic, as follows: ; ; wherein, denotes the target Q value network; denotes the i-th state; denotes the action of the i-th state; denotes the Q value estimate of the target network; denotes the maximum value of the Q value estimate of the target network; denotes the i-th state; denotes the action of the i-th state; denotes the Q value estimate of the current network. Preferably, in step S4, the ISP process is realized based on the SEA-DDQN model, including the following steps: Step S41, based on the SEA-DDQN model, the state feature vector space of each network traffic sample is mapped to the input of the action selection network as follows: Step S42, in the input layer, the d-dimensional feature vector in the state feature vector space of the network traffic sample is mapped to a 128-dimensional hidden space, as follows:

[0013] Step S43, the action selection network is used to select the action of the network traffic sample, as follows: Step S44, the state feature vector space of the network traffic sample is mapped to the input of the value network as follows: Step S45, the value network is used to estimate the value of the network traffic sample, as follows: Step S46, the value of the network traffic sample is calculated as follows: Step S47, the reward of the network traffic sample is calculated as follows: Step S48, the target network is used to calculate the target value of the network traffic sample, as follows: Step S49, the target value of the network traffic sample is calculated as follows: Step S410, the target value of the network traffic sample is calculated as follows: Step S411, the target value of the network traffic sample is calculated as follows: Step S412, the target value of the network traffic sample is calculated as follows: ​W1 b1 Step S43, pass through the attention layer and use The activation function realizes a nonlinear transformation, so that the neural network learns a more complex traffic feature representation, as follows: ; ; wherein, is an element-wise multiplication, used to combine the weighted results of the attention mechanism with the original features; represents the result after ReLU activation function processing, representing the nonlinearly transformed feature representation learned by the network; Step S44, the second layer fully connected layer and the subsequent layers and the above layers are the same, as follows: ; ; ; ; wherein, represents the pre-activation vector of the second layer neural network; represents the weight of the second layer neural network; represents the bias of the second layer neural network; represents the result after ReLU activation function processing of the second layer; Step S45, in the last output layer, the corresponding Q value of each action in the agent action space is generated, and the output value represents the expected reward accumulation of the corresponding action in the current state , as follows: ; wherein, is the weight of the fourth layer neural network; is the result after ReLU activation function processing of the third layer, representing the nonlinearly transformed feature representation learned by the network; is the bias of the fourth layer neural network; is the Q value of the first action in the action space element vector; represents the Q value estimation of the network; is the expected cumulative reward corresponding to the element in the action space ; For the Q value vector obtained by the action selection network for the current state , the agent adopts The policy to make the best prediction action selection as follows: ; wherein, denotes the size of the prediction action space; denotes the optimal prediction action under the current state ; denotes the exploration rate of the current time step; denotes the probability of selecting action under state ; The update rule of the target Q-value network is as follows: ; wherein, is the initial exploration rate, is the minimum exploration rate threshold, is the decay coefficient, is the current learning time step.

[0014] The calculation process of the target Q-value network is the same as that of the action selection network , and the input is , and the target Q-value is calculated as follows: ; wherein, is the reward obtained by taking action under the current state ; is the target network Q-value; Based on PER, the sampling probability is assigned according to the importance of experience, the sampling bias is introduced, and the mean square error loss function with importance sampling weight is constructed to compensate for the bias introduced by non-uniform sampling as follows: ; wherein, denotes the loss function; denotes the bth traffic in the experience replay pool; denotes the state; denotes the action; denotes the predicted Q-value of the action selection network for the state and action of the bth traffic; denotes the target Q-value of the bth traffic; denotes the probability of sampling the bth traffic; denotes the importance sampling coefficient; is the number of network traffics drawn from the experience replay pool; In the initial stage of learning, as learning proceeds, The value gradually increases and finally approaches 1, as follows: ; Wherein, Indicates the beta value at time step t; Indicates the beta value at time step t-1; Indicates the initial value of the beta value; Indicates the final target value of the beta value; the beta value is a parameter for controlling the degree of forgetting or retaining past experience; In the action selection network The parameters are updated in real time by back propagation, and the update formula of the parameters is as follows: ; Wherein, Indicates the parameters of the neural network; Indicates the learning rate; Indicates the gradient of the parameters ; The parameters of the target Q value network are updated by combining soft update and hard update of Polyak average, that is, every 100 time steps, the parameters in the action selection network are updated by weighted average with the parameters of the target Q value network ; every 1000 time steps, the parameters in the action selection network are directly copied to the parameters of the target Q value network , as follows: ; Wherein, Indicates the parameters of the target Q value network; Is The average soft update coefficient.

[0015] Therefore, the present application adopts the above-mentioned one kind based on SEA-DDQN adaptive reinforcement learning's invasion situation prediction method, and the beneficial effects are as follows: (1) The present application adopts the architecture combining DDQN and attention mechanism, and separates the agent to predict the captured network traffic and calculate the target Q value, thereby reducing the overestimation bias of the agent when predicting the network traffic. At the same time, the attention mechanism is introduced to optimize the attention distribution of the agent to the network traffic features, enhance the discriminative features and suppress irrelevant features.

[0016] (2) The application introduces an n-step reward accumulation mechanism and a lag reward distribution strategy to simulate the inherent delay of threat confirmation in a real network environment. In this way, the agent must learn to predict the cumulative return based on the historical state sequence to predict the future and make autonomous decisions about the potential intrusion situation, rather than relying on immediate supervision signals.

[0017] (3) The application combines the priority experience replay (PER) strategy mechanism and introduces the time difference error (TD-Error) as an important index of experience network flow samples, and preferentially uses TD error large experience network flow samples for learning, so as to preferentially use those experience network flow samples that have the greatest impact on the current strategy for updating, and improve the ISP efficiency.

[0018] (4) The application uses benchmark NSL-KDD, UNSW-NB15, CICIDS-2017 and NSL-KDD, UNSW-NB15 data sets to evaluate SEA-DDQN, and compares the results with other mainstream research methods. The experimental results show that the prediction accuracy of SEA-DDQN is higher. This shows that SEA-DDQN has higher accuracy and robustness when dealing with complex network flow data. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is the ISP model implementation flowchart of the application; Figure 2 is the DRL agent environment interaction flowchart of the application; Figure 3 is the flowchart of the feature processing process of the attention mechanism of the application; Figure 4 is the flowchart of the priority experience sampling process of the application; Figure 5 is the flowchart of the ISP attack category prediction decision process of the application; Figure 6 is the accuracy and exploration rate change graph under different discount factors in the embodiment of the application; Figure 7 is the average loss comparison graph under different discount factors in the embodiment of the application; Figure 8 is the accuracy distribution graph of the last 50 rounds under different discount factors in the embodiment of the application; the orange part represents gamma=0.5; the blue part represents gamma=0.005; and the green part represents gamma=0.9; Figure 9are confusion matrix figures under different discount factors in embodiments of the present application; wherein, (a) is a confusion matrix when the discount factor is 0.9; (b) is a confusion matrix when the discount factor is 0.5; (c) is a confusion matrix when the discount factor is 0.005; Figure 10 are accuracy comparison figures under different parameter values of Batch-size & β_final in embodiments of the present application; Figure 11 are ablation comparison figures in embodiments of the present application; wherein, (a) is a comparison of F1-Score of each category; (b) is a comparison of overall performance; Figure 12 are minority category comparison figures in embodiments of the present application; wherein, (a) is a performance comparison of R2L category; (b) is a performance comparison of R2L category (reduce key category); Figure 13 are confidence curve figures of different attack types in embodiments of the present application; Figure 14 are performance box plots in embodiments of the present application; wherein, different colors represent different attack categories; cyan green represents DoS; apricot orange represents Normal; mist blue represents Probe; rose pink represents R2L; tender yellow green represents U2R; Figure 15 are training loss comparison figures under different random seeds in embodiments of the present application; Figure 16 are 95% confidence interval figures under different random seeds in embodiments of the present application; Figure 17 are NSL-KDD exploration figures in embodiments of the present application; wherein, (a) is the accuracy rate and exploration rate of NSL-KDD data set with the change of training rounds; (b) is the average loss of NSL-KDD with the change of training rounds; Figure 18 are UNSW-NB15 exploration figures in embodiments of the present application; wherein, (a) is the accuracy rate and exploration rate of UNSW-NB15 with the change of training rounds; (b) is the average loss of UNSW-NB15 of each round; Figure 19 are CICIDS2017 exploration figures in embodiments of the present application; wherein, (a) is the accuracy rate and exploration rate of CICIDS2017 with the change of training rounds; (b) is the average loss of CICIDS2017 of each round; Figure 20 are MQTT-IoT-IDS2020 exploration figures in embodiments of the present application; wherein, (a) is the accuracy rate and exploration rate of MQTT-IoT-IDS2020 with the change of training rounds; (b) is the average loss of MQTT-IoT-IDS2020 of each round. DETAILED DESCRIPTION

[0020] The technical solutions of the present application are further illustrated by the accompanying drawings and examples.

[0021] Example 1 As Figure 1 shown, a method for intrusion situation prediction based on SEA-DDQN adaptive reinforcement learning includes the following steps: Step S1, obtaining NSL-KDD, UNSW-NB15, CICIDS-2017 and NSL-KDD, UNSW-NB15 network flow data sets.

[0022] Obtaining NSL-KDD, UNSW-NB15, CICIDS-2017 and NSL-KDD, UNSW-NB15 network flow data sets. Among them, the NSL-KDD data set and the UNSW-NB15 data set contain 41 features and 1 category label, the CICIDS-2017 data set contains 52 features and 1 category label, and the MQTT-IoT-IDS2020 data set contains 31 features and one category label, as shown in Tables 1, 2, 3 and 4 respectively. The features cover basic network connection attributes, content features and flow-based statistical information, which can fully reflect the characteristics of network flow.

[0023] The NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020-Biflow data sets used in the present application contain 148517, 257673, 524288 and 173371 network flow records respectively. Among them, the CICIDS-2017 data set originally contains 1048576 records, which is reduced by half through stratified sampling method, ensuring that the sample proportion of each category can be highly consistent with the original data set.

[0024] Table 1 NSL-KDD data set features ;

[0025] Table 2 UNSW-NB15 data set features ;

[0026] Table 3 CICIDS-2017 data set features ;

[0027] Table 4 MQTT-IoT-IDS2020 data set features ;

[0028] Step S2, pre-process the obtained network traffic dataset to obtain a normalized dataset.

[0029] The pre-processing methods of NSL-KDD, UNSW-NB15, CICIDS-2017 and NSL-KDD, UNSW-NB15 datasets are consistent. This embodiment takes the NSL-KDD dataset as an example, and the specific processing process is as follows: Step S21, define the network traffic dataset as , wherein is the number of traffic samples, then the feature space is , as follows: ; , wherein represents a feature subset, ; represents a feature, ; represents a union symbol, and represents that four feature subsets are combined into a feature space ;

[0030] is a traffic feature subset, which is based on the most basic network connection level, provides a direct description of the packet transmission duration and the communication protocol adopted, and is a basic element for understanding network behavior patterns, including , and other features.

[0031] is a content feature subset, which reflects the data interaction scale and directionality of the communication parties by quantifying the actual transmission content in the data packet, and includes , and other features.

[0032] is a time statistical feature subset, which summarizes and analyzes the network traffic from the time dimension, reveals the frequency of network connection and the activity status of the server within a certain time window, and includes , and other features.

[0033] is a host behavior feature subset, which focuses on the network behavior pattern of a specific host, provides connection frequency, data transmission volume and other indicators of the destination host, and includes and other features.

[0034] Step S22, for discrete features , , , perform hot one-hot encoding processing.

[0035] First, the discrete feature subset ; wherein, represents a discrete feature; represents a discrete feature index. The corresponding value range of each discrete feature is respectively: ; ; ; wherein, , and respectively represent the value range of the discrete feature, . For any discrete feature, define the indicator function as follows: ; wherein, ; represents a sample The value of the discrete feature .

[0036] Then, by the hot one-hot encoding transformation that is , the feature space is reconstructed as follows: ; wherein d is the total number of features after hot one-hot encoding processing; represents a hot one-hot encoding processing function; represents a vector splicing operation; is the index of all non-discrete features in the original 41-dimensional feature; is the discrete feature subset.

[0037] Finally, the original training set and test set are respectively hot one-hot encoded to generate their respective extended feature matrices , and normalized as follows: ; wherein, represents a normalized feature; and respectively represent the maximum value and the minimum value of the feature ; is an indicator function, which takes the value 1 when , otherwise 0.

[0038] The NSL-KDD dataset contains 59 attack categories, which are divided into 5 main categories and mapped into a set of digital categories , which is convenient for the model to process, as shown in Table 5.

[0039] Table 5 NSL-KDD dataset record categories ;

[0040] In Table 5, the normal traffic has no attack behavior, which is classified as Normal. Attacks such as Back, Worm, etc. occupy system or network resources by a large number of legal requests, making normal users unable to obtain services, which are classified as Dos. Attacks such as Nmap, Ipsweep, etc. frequently perform port scanning and IP scanning, which are classified as Probe. Attacks such as Snmpguess, tp_write, etc. obtain access to the local system by remotely sending data packets, which are classified as R2L. Attacks such as Perl, Rootkit, etc. are local users who elevate their own permissions to root users through system vulnerabilities, which are classified as U2R.

[0041] Similarly, in the CICIDS-2017 and MQTT-IoT-IDS2020-Biflow data sets, they are divided into 7 main categories and 5 main categories and mapped into a digital category set, the original record class of CICIDS-2017 and the record class used in this study are shown in Tables 6 and 7, respectively, and MQTT-IoT-IDS2020-Biflow is shown in Table 8. In the UNSW-NB15 data set, it is two-predicted, and its record class is shown in Table 9.

[0042] Table 6 Record class of CICIDS-2017 original data set ;

[0043] Table 7 Record class of CICIDS-2017 data set of the present application ;

[0044] Table 8 Record class of MQTT-IoT-IDS2020-Biflow data set ;

[0045] Table 9 Record class of UNSW-NB15 data set ;

[0046] Step S3, a deep attention Q network (SEA-DDQN) model based on double value is constructed, and network security intrusion situation prediction (ISP) is performed.

[0047] Step S31, based on the NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020 training set and test set after data preprocessing, the network ISP reinforcement learning environment is constructed and defined.

[0048] The ISP process is abstracted as a Markov decision process MDP, and it is formally defined as a four-tuple As shown below: ; wherein the state is a set of feature vector spaces for each network flow; the action is a set of flow type decisions defined in NSL-KDD, UNSW-NB15, CICIDS-2017 and MQTT-IoT-IDS2020, at a specific time The state and action of the agent are represented as and respectively; the reward is the reward or punishment obtained by the agent for predicting the flow correctly or not, which is a scalar signal; in the present application, the calculation of the reward vector is based on the comparison between the prediction result of the agent and the actual label to evaluate the accuracy of the agent prediction; the discount factor is used to balance the immediate reward and the discounted coefficient of long-term return.

[0049] Since each flow has independence, the MDP is degraded to a conditional reward maximization problem. The basic process of the agent interacting with the environment in the deep reinforcement learning used in the present application is shown in Figure 2 .

[0050] Step S32, constructing a SEA-DDQN model agent, including an action selection network and a target Q value network.

[0051] The SEA-DDQN model agent includes two key neural networks, namely the action selection network and the target Q value network. The SEA-DDQN model agent performs prediction (i.e. action selection) and target Q value calculation on the captured network flow.

[0052] The Q value functions of the action selection network and the target Q value network are respectively composed of four-layer fully connected feedforward neural networks, and the attention mechanism and ReLu activation function are used between the fully connected layers, wherein each of the two hidden layers contains 128 neurons, as shown below: ; wherein, represents the action selection network or the target network; is the fourth fully connected layer; is the third fully connected layer; is the second fully connected layer; is the first fully connected layer; represents inputting the state S into the first fully connected layer; is the Relu layer; is the attention layer.

[0053] In this application, the features other than the category label in the four data sets are used as the state representation of the SEA-DDQN model agent; the category label is used to calculate the reward vector based on the model prediction result. The calculation of the reward vector is based on the comparison between the prediction result of the agent and the actual label, and the agent only obtains the action for calculating the reward vector, without performing any actual action on the environment. The action selection of the agent is only used to evaluate the accuracy of its prediction, and the feedback is given to the SEA-DDQN model agent through the reward function to guide its learning process.

[0054] The state of the SEA-DDQN model agent at each time step is represented by the feature vector of the network traffic, and the state space of the agent is defined as Therefore, the state vector of the SEA-DDQN model agent for each piece of network traffic is as follows: ; where each dimension represents a certain feature of the network traffic.

[0055] Step S33, introduce attention mechanism to optimize the attention allocation of the SEA-DDQN model agent to the network traffic features, enhance discriminative features and suppress irrelevant features.

[0056] Step S331, for the input network traffic feature state vector at each time step, after processing by the fully connected layer, a 128-dimensional feature vector is obtained, where represents the network traffic feature vector processed by the fully connected layer, represents the dimension index; then it is sent to the attention layer to enhance the attention to important features, and the attention weight is as follows: ; where is a dimension reduction projection matrix, which projects the original feature vector to a low-dimensional space to reduce the computational complexity and extract the main feature trend; is a dimension lifting reconstruction matrix, which is used to map the dimension-reduced features back to the original feature space to generate attention weights matching the original feature dimensions; represents the dimension reduction ratio, which is a hyperparameter; represents the Sigmoid gating function, as follows: ; wherein, represents a linear combination calculated by the previous layer of the model; is a Sigmoid function.

[0057] Here, the Sigmoid gating function compresses the input value to the interval , and the generated attention weight can be regarded as the original "attention coefficient" of the feature, which is used to measure the importance of each feature.

[0058] Step S332, the generated attention weight is fused with the original feature by modulation, as follows: ; wherein, represents the fused feature representation; is the Hadamard product, i.e., the element-wise multiplication operation.

[0059] This operation makes each feature value be dynamically adjusted by the corresponding weight . Since , the features with larger weights will be enhanced, while the features with smaller weights will be weakened. The entire attention mechanism processing feature process is as shown in Figure 3 .

[0060] Step S34, the SEA-DDQN model agent generates an action list according to the input features of the deep neural network, which is presented in the form of an action vector; the final Q value is used to evaluate whether the attack behavior is successfully predicted.

[0061] Step S341, first, respectively map the class labels in the NSL-KDD, UNSW-NB15, CICIDS-2017, and MQTT-IoT-IDS2020 data sets into a set of digital classes, as follows: ; wherein 0 represents Normal; 1 represents DoS; 2 represents Probe; 3 represents U2R; and 4 represents R2L.

[0062] ; wherein 0 represents normal traffic; and 1 represents abnormal traffic.

[0063] ; wherein 0 represents Bots; 1 represents Brute Force; 2 represents DDos; 3 represents Dos; 4 represents Normal; 5 represents Port Scanning; and 6 represents Web Attacks.

[0064] ; wherein 0 represents Normal; 1 represents Bruteforce; 2 represents Scan_A; 3 represents Scan_sU; and 4 represents Sparta.

[0065] Step S342, the action space of the SEA-DDQN model agent is defined as: ; ; ; ; wherein each action corresponds to a predicted category decision in the action list; represents a category index.

[0066] Step S35, a dynamic weight reward mechanism is designed, and a category-sensitive reward is constructed to effectively cope with the imbalance of different category attack data.

[0067] In reinforcement learning, the design of the reward function is crucial for guiding the agent to learn the optimal strategy. The present application proposes a dynamic reward function design method, which balances the category bias by constructing a category-sensitive reward, and uses a reverse frequency weighting strategy to determine the category weight, so as to enhance the sensitivity of the SEA-DDQN model agent to different attack categories and balance the influence of category bias in the data set on the learning process.

[0068] Step S351, the specific form of the reward function is as follows: ; wherein, represents the weight of the attack category; represents correct prediction; represents incorrect prediction.

[0069] Step S352, the reverse frequency weighting strategy, by considering the frequency of occurrence of each attack category, the category with lower frequency obtains higher weight, so as to give more attention in the reward function to balance the category bias, as follows: ; ; wherein, represents the attack category weight set of ; represents the weight of attack; represents the weight of no attack; represents the weight of attack; represents the weight of attack; represents the weight of attack; represents the attack category weight set of represents the weight of attack; represents the weight of attack; represents the weight of attack; represents the weight of attack; represents the weight of attack.

[0070] The weight rule is updated as follows: ; wherein, is the total number of samples, is the number of attack types, represents the total number of attack samples.

[0071] Through this design, the reward function can accurately feedback the predicted behavior of the agent, thereby improving the prediction accuracy of different attack category samples.

[0072] Since the UNSW-NB15 performs a two-prediction task in this study, a positive reward is given when the SEA-DDQN model agent correctly predicts the sample, and a negative penalty is given otherwise, as follows: ; wherein, represents reward setting rule of each flow in the data set.

[0073] In order to simulate the delay feedback scene in the real network, the n-step reward accumulation mechanism is adopted, and the reward actually stored in the experience replay pool is the n-step cumulative reward , rather than the single-step immediate reward ; the agent stores the transition sequence of the last n steps in the process of interacting with the environment, and calculates the discount cumulative reward after reaching n steps , as follows: .

[0074] Step S36: Based on the Priority Experience Playback (PER) strategy mechanism, the Time Differential Error (TD-Error) is introduced as an indicator of the importance of experience network traffic samples. Experience network traffic samples with large TD errors are used for learning first, thereby prioritizing the use of experience network traffic samples that have the greatest influence on the current strategy for updating, thus improving ISP efficiency.

[0075] During the interaction between the agent and the environment, experiences (states, actions, rewards, and the next state) are stored in an experience replay pool and periodically replayed for learning. This approach balances the effects of rewards and penalties by weighting samples of different categories, enabling the agent to learn the characteristics of different attack categories more evenly.

[0076] A Priority Sampling (PER) strategy is adopted to optimize the sampling process of empirical network traffic. Taking into full account the differences in importance of different attack categories in NSL-KDD and CICIDS-2017, TD-Error is introduced as a key indicator to measure sample importance. A priority update formula is defined to dynamically adjust the sample sampling probability, as shown below: ; in, Indicates the first The probability that a given network traffic sample will be taken; It is a small constant used to ensure that all empirical network traffic has a certain probability of being sampled, and to avoid some samples having a zero sampling probability due to the TD-Error being too small; This represents the current network traffic count in the experience replay pool. This is a priority parameter used to control the degree of influence of TD-Error on the sampling probability; For the first The TD-Error corresponding to each empirical network traffic is shown below: ; in, Represents the target Q-value network; Indicates the first One state; Indicates the first Actions in each state; This represents the Q-value estimate of the target network; This represents the maximum value of the Q-value estimate for the target network; Indicates the first One state; Indicates the first Actions in each state; This represents the Q-value estimate of the current network.

[0077] Priority experience sampling process, such as Figure 4 As shown.

[0078] Step S4: Implement the ISP process based on the SEA-DDQN model.

[0079] Step S41: Based on the SEA-DDQN model, convert the state feature vector space of each network traffic sample into a single vector. As an action selection network The input is processed through a hierarchical nonlinear transformation to capture the features of each traffic sample, as shown below: ; in, Indicates the action selection network; This represents the input of the first fully connected layer.

[0080] Step S42: In the input layer, convert the network traffic sample state feature vector space... The d-dimensional feature vectors in the model are mapped to a 128-dimensional hidden space as follows: ; in, This represents the preactivation vector of the first layer of the neural network; This represents the weights of the first layer of the neural network; This indicates the bias of the second layer of the neural network.

[0081] Step S43: Through the attention layer and using The activation function implements a non-linear transformation, enabling the neural network to learn more complex representations of traffic features, as shown below: ; ; in, This is an element-wise multiplication method used to combine the weighted result of the attention mechanism with the original features; This represents the result after processing with the ReLU activation function, signifying the feature representation learned by the network after nonlinear transformation.

[0082] Step S44, the second fully connected layer, and subsequent layers are similar to the previous layers, as shown below: ; ; ; ; in, This represents the preactivation vector of the second layer of the neural network; weights of the second layer neural network; bias of the second layer neural network; result of the second layer after ReLU activation function processing.

[0083] Step S45, in the last output layer, the Q value of each action in the agent action space is generated , that is, the output values of the 5 neurons of the output layer represent the expected reward accumulation of the 5 possible actions in the current state , as follows: ; wherein, is the weight of the fourth layer neural network; is the result of the third layer after ReLU activation function processing, representing the nonlinearly transformed feature representation learned by the network; is the bias of the fourth layer neural network; is the Q value of the first action in the action space element vector; represents the Q value estimation of the network; is the expected cumulative reward corresponding to the action space element.

[0084] For the current state , the Q value vector obtained by the action selection network , the agent uses policy to select the best predicted action, as follows: ; wherein, represents the size of the predicted action space; represents the optimal predicted action under the current state ; represents the exploration rate of the current time step; represents the probability of selecting action under state ; the entire predicted attack category decision process is as shown in Figure 5 .

[0085] In the reinforcement learning process, as the agent accumulates experience, it is hoped that it will gradually reduce the frequency of exploration and increase the proportion of utilization. Let the exploration rate decay over time, then the update rule is as follows: ; wherein, is the initial exploration rate, is the minimum exploration rate threshold, is the decay coefficient, This is the current learning time step.

[0086] In the early stages of learning, intelligence is experienced... The probability of randomly predicting an action is selected, and as learning progresses, It decays exponentially. This ensures that the agent mainly relies on the learned strategy for prediction in the later stages of training, without completely losing its exploration ability, thus maintaining a stable balance between exploration and utilization in the long run.

[0087] Step S46, Target Q-value network Its purpose is to provide a stable target value for training the action selection network. It calculates the next state. The Q value is used to provide the target Q value, thereby achieving an approximation of the Q-value function.

[0088] Target Q-value network The computation process and action selection network The calculation process is similar, and its input is... The target Q value is calculated as follows: ; in, The current state Take action below The rewards received; Let Q be the target network Q-value.

[0089] Step S47: Since PER assigns sampling probabilities based on empirical importance, this leads to a difference between the sampling distribution and the original empirical distribution, introducing sampling bias. Therefore, the loss function of this invention is the mean squared error loss with importance sampling weights to compensate for the bias introduced by non-uniform sampling, as shown below: ; in, Represents the loss function; This represents the b-th flow in the experience replay pool; Indicates state; Indicates an action; This indicates the state of the action selection network for the b-th traffic. and actions The predicted Q value; This represents the target q value of the b-th flow; This indicates the probability that the b-th traffic item will be sampled. Indicates the importance sampling coefficient; This represents the number of network traffic points drawn from the experience replay pool; at this point, In for .

[0090] In the initial stage of learning, in order to reduce the weight influence of high-priority network traffic in the experience pool, The value is relatively low. As learning progresses, The value gradually increases and eventually approaches 1, as follows: ; Wherein, βt represents the value of β at time step t; βt-1 represents the value of β at time step t-1 (i.e. the previous time step); β0 represents the initial value of β; β1 represents the final target value of β; β is a parameter for controlling the degree of "forgetting" or "retaining" of past experiences (experience pool).

[0091] Step S48, the network for calculating target Q value in the present application and have the same structure, but their parameter update methods are not the same. Polyak average soft update strategy and hard update strategy are adopted to update the target network parameters smoothly, reducing the instability caused by gradient fluctuations in the learning process.

[0092] In the action selection network Parameter is updated in real time by backpropagation, and the update formula of parameter is as follows: ; Wherein, represents the parameter of the neural network; represents the learning rate; represents the gradient of parameter .

[0093] In order to enable the target Q value network to effectively follow the learning progress of the action selection network while maintaining stability, the parameters of the target Q value network are updated by combining Polyak average soft update and hard update, that is, every 100 time steps, the parameters in the action selection network and the parameters of the target Q value network are weighted and averaged to update the parameters of the target Q value network; every 1000 time steps, the parameters in the action selection network are directly copied to the parameters of the target Q value network For example, as shown below: ; wherein, denote the parameters of the target q-value network; is an average soft update coefficient.

[0094] Example 2 This embodiment carries out multi-prediction tasks on NSL-KDD, CICIDS-2017 datasets, and two-prediction tasks on UNSW-NNB15. In the context of multi-prediction and two-prediction, TP (true positive) refers to the number of positive samples correctly predicted by the agent, FP (false positive) refers to the number of negative samples incorrectly predicted as positive, TN (true negative) refers to the number of samples correctly predicted as negative, and FN (false negative) refers to the number of negative samples incorrectly predicted as positive.

[0095] With the help of TP, FP, TN and FN, more interpretable performance indicators such as accuracy, precision, recall and F1 score can be calculated to comprehensively and quantitatively evaluate the performance of the agent.

[0096] In the performance evaluation system of the prediction task, accuracy As the most commonly used measure, its essence is to reflect the proportion of correct prediction by the agent in the overall sample. Specifically, it is the ratio of the number of samples correctly predicted by the agent to the total number of samples, as shown below: ; Precision is defined as the proportion of samples that actually belong to a certain positive class in the sample set predicted by the agent to belong to that positive class, as shown below: ; Recall is defined as the proportion of samples correctly predicted by the agent as positive in all actual positive samples, as shown below: ; F1 score As a comprehensive performance indicator, it is defined as the harmonic mean of precision and recall. It comprehensively measures the accuracy and integrity of the agent in predicting positive samples. By integrating the information of precision and recall, F1 score can comprehensively reflect the balanced performance of the agent in predicting positive and negative samples, as shown below: ; To provide reference and lessons for subsequent research, in this study, the parameters of the proposed SEA-DDQN model are shown in Table 10. These parameters are determined after systematic research and optimization, and they play a crucial role in the training and testing process of the model. Through experimental verification, the combination of these parameters performs excellently on the specific task and data set targeted by this study, effectively balancing the exploration and utilization capabilities of the model, and improving the accuracy and stability of its decision-making.

[0097] Table 10 SEA-DDQN model parameter table ;

[0098] I. Sensitivity analysis of discount factor γ and Batch-size & β_final

[0099] The value of the discount factor γ has a decisive influence on the time horizon and decision-making strategy of the reinforcement learning agent. To demonstrate the selection of the γ parameter in the proposed model, in-depth sensitivity experiments were conducted on the NSL-KDD dataset, comparing three different γ values: γ = 0.9 (long-term planning), γ = 0.5 (medium-term planning, main setting), and γ = 0.005 (extremely short-sighted). The final accuracy on the test set for γ = 0.5 was 97.50%, for γ = 0.005 was 97.14%, and for γ = 0.9 was 96.75%. The experimental results consistently show that γ = 0.5 achieves the best balance in model performance and learning stability. Given that this parameter outperforms the other two groups on NSL-KDD, it is directly migrated to UNSW-NB15 and CICIDS-2017, further verifying the cross-dataset robustness of γ = 0.5.

[0100] From Figure 6 , Figure 7 and Figure 8 , it is clear that γ = 0.5 is the undisputed optimal choice. It not only achieves the highest final performance, but more importantly, this achievement is realized through the most stable and efficient learning process. In the learning process of γ = 0.5, its average loss decreases the fastest and smoothest, indicating an efficient and stable learning process. From Figure 8It can be seen that in the last 50 episodes, the exploration rate is close to 0.1, and the median line in the box is obviously higher than the other two groups, indicating that in more than 50% of the tests, its accuracy is the highest among the three. The box is very short, which means that 50% of the test results are concentrated in a very narrow high-precision interval (97.10%-97.50%). This is a sign of high stability and reliability of performance. The upper and lower lines are short and have no outliers, further proving that its output is very consistent and there is no performance anomaly or surge. This shows that γ = 0.5 enables it to quickly learn the reward and plan for long-term benefits.

[0101] The "short-sighted" strategy of γ = 0.005, although the performance is acceptable, the learning process is unstable, and its performance upper limit cannot surpass γ = 0.5. In the learning process of γ = 0.005, its average loss curve fluctuates significantly in the later period, indicating that its learning strategy is unstable. From Figure 8 It can be seen that in the last 50 episodes, the exploration rate is close to 0.1, and the median accuracy line is lower than γ = 0.5 but higher than γ = 0.9. The box is significantly longer than the box of γ = 0.5. This means that its 50% test results are distributed in a wider range (96.80%-97.15%). This indicates that its performance is volatile and not stable. The model learned by the "short-sighted" strategy has poor robustness, and its performance is more susceptible to changes in test samples, although it can sometimes achieve good results, but it is unreliable.

[0102] γ = 0.9 is significantly behind in terms of final performance due to low learning efficiency. In the learning process of γ = 0.9, its average loss value decreases the slowest and is always the highest, indicating that its learning process is inefficient. From Figure 8 It can be seen that in the last 50 episodes, the exploration rate is close to 0.1, and the median line in the box is obviously higher than the other two groups, indicating that in more than 50% of the tests, its accuracy is the highest among the three. The box is very short, which means that 50% of the test results are concentrated in a very narrow high-precision interval (97.10%-97.50%). This is a sign of high stability and reliability of performance. The upper and lower lines are short and have no outliers, further proving that its output is very consistent and there is no performance anomaly or surge. This shows that γ = 0.5 enables it to quickly learn the reward and plan for long-term benefits.

[0103] In contrast, choosing γ = 0.5 exhibits the best overall performance. It not only achieves the highest median accuracy, but more importantly, its box is short and compact, with symmetrical upper and lower lines and no outliers. This indicates that the performance of the model is highly concentrated and stable, with excellent robustness. A stable model is crucial in real network security applications, as its output must be reliable and predictable.

[0104] Figure 9The confusion matrix of Table 11 and the detailed results reveal the profound impact of the discount factor γ on the model behavior, especially on the detection ability of minority attack categories. The setting of γ = 0.9, due to low learning efficiency, performs poorly on all categories, especially failing to detect U2R attacks completely (Recall: 16.42%), which proves that the overly "far-sighted" strategy is ineffective in this task. The setting of γ = 0.005 exhibits "short-sighted" behavior. Although its overall accuracy (97.14%) is quite misleading, in-depth analysis finds that its detection strategy for U2R attacks is conservative and fails, trading a high false negative rate (Recall only 47.76%) for a higher prediction accuracy. In practical security applications, this is an unacceptable strategy. Its ability to detect R2L attacks deteriorates significantly, with a significant drop in recall to 84.00%, proving that it cannot learn complex patterns that require multi-step correlation. This confirms that a very low γ value will simplify the agent into a short-sighted classifier, which is fundamentally at odds with the goal of simulating attack evolution. In contrast, the choice of γ = 0.5 achieves excellent balanced performance, significantly improving the detection ability of the most critical high-level threats (U2R), nearly doubling the recall rate, and achieving the only available F1-Score. It achieves the best performance in R2L attack detection, achieving a perfect balance between high recall and high precision. At the same time, it maintains top-notch performance on most categories. This proves that the "mid-term vision" provided by γ = 0.5 is the key for the agent to learn to detect various attack patterns from simple to complex and make the best trade-off.

[0105] Table 11 Detailed indicators under different discount factors ;

[0106] To verify the influence of Batch-size and β_final on the model accuracy in Table 11, under the premise of keeping the rest of the hyperparameters unchanged, set them to 10w, 20w and 24w for training. From Table 12, it can be seen that the three curves are almost completely coincident, and the average accuracy of 10w, 20w and 24w configurations is 97.47%, 97.46% and 97.49% respectively, with a standard deviation of only 0.013%, and the maximum difference is not more than 0.03%. The above results show that within the range of values considered, the variation of Batch-size and β_final has little effect on the final accuracy of the model, verifying the robustness of the method proposed in the present invention to the setting of this hyperparameter. Figure 10

[0107] II. Ablation experiment.

[0108] ​To evaluate the contribution of each key component in the proposed network intrusion detection model, an ablation experiment on the NSL-KDD dataset was designed. The attack class inverse frequency weight mechanism and the PER module were removed respectively, and compared with the complete model. The experiment was evaluated from two dimensions of overall performance and category performance, focusing on the detection ability of the model for minority class attacks (U2R, R2L).

[0109] As shown in Table 12 and Figure 11 It can be seen that the complete model achieves the optimal performance in accuracy, precision, recall and F1-Score. Removing the class weight or PER will cause the overall performance of the model to decline, and removing the PER mechanism has the most significant impact on the overall performance (accuracy decreases by 0.87%), which indicates that PER plays a key role in improving the overall learning efficiency and stability of the model. For the majority or more common categories such as DoS, Normal, Probe, the three model variants all maintain a high performance level (F1-Score is higher than 0.96), and the difference between them is small.

[0110] Table 12 Comparison of overall performance indicators of the complete model, removing the class weight and removing the PER mechanism ;

[0111] Although the overall performance gap seems small, the performance difference in each category, especially in the minority class attacks, reveals the importance of each component. U2R and R2L are the rarest attack types in the data and the most difficult to detect categories, and their F1-Score comparison shows that the complete model has a significant advantage over the other two model variants. Figure 12The complete model achieved the highest F1-Score (0.5825) for the U2R category, with a precision of 0.8333 significantly higher than the recall of 0.4478, indicating that the model's prediction results for this category are very reliable, but there is still room for improvement in detection capability. After removing the category weight, the model's precision decreased to 0.7143, resulting in a decrease in F1-Score to 0.5505. This indicates that the category weight module effectively ensures the model's learning of minority class features, preventing them from being overwhelmed by the majority class. After removing PER, the model maintained a high precision (0.8113), but the F1-Score was similar to that of the complete model (0.5767). This indicates that PER has a positive effect on improving the detection stability of rare attacks such as U2R, but is not the most critical factor. For the R2L category, the complete model achieved the highest F1-Score (0.9417) on the R2L category, achieving a balance between high precision (0.9182) and high recall (0.9664). After removing the category weight, the model's recall decreased, and the F1-Score decreased to 0.9243. This indicates that without the category weight, the model's detection capability for R2L attacks will be weakened. After removing PER, the model's performance decreased relatively small (F1-Score: 0.9383), indicating that for the R2L category, the role of the category weight is more critical than that of PER.

[0112] Through the ablation experiment, it can be known that the category weight module is the core component of improving the model's detection performance for minority class attacks (such as U2R, R2L). Its role mainly lies in preventing the model training process from being dominated by the majority class, ensuring that the loss function can effectively reflect the classification error of minority class samples, thereby significantly improving the recall rate and F1-Score for minority classes. The PER mechanism is the key to improving the overall performance and stability of the model. By focusing on "difficult to learn" samples (usually misclassified or uncertain samples, often belonging to minority classes), the PER mechanism effectively speeds up the model's convergence speed and improves overall performance indicators. Its optimization of minority class performance is not as direct as the category weight, but provides an important supplement. The complete model combines the category weight and PER mechanism, achieving the best balance in overall performance and performance for each category, especially for the detection of key minority classes, verifying the effectiveness and necessity of the model design.

[0113] III. Random seed sensitivity analysis and model stability evaluation.

[0114] To verify the potential impact of random seeds on model reproducibility, this study systematically evaluates the influence of different random initialization seeds on model performance on the NSL-KDD dataset, demonstrating the reliability and stability of the results. Experiments were repeated on the standard test set using five different random seeds (42, 2023, 2024, 3407, 12345), and statistical analysis was performed on overall and category-specific performance metrics.

[0115] The model of this invention exhibits high repeatability and stability in performance across the entire system and most categories. Figure 13 , Figure 14 As shown in Table 13, the overall classification accuracy of the model exhibits extremely high stability under different random seeds. Its mean is as high as 97.60%, with a very low standard deviation (0.15%) and a very narrow 95% confidence interval ([97.46%, 97.73%]). This indicates that the overall performance of the model is not sensitive to the choice of random seed and has excellent reproducibility. For the four attack categories with large sample sizes—DoS, Normal, Probe, and R2L—the standard deviation of the accuracy is less than 1%, the confidence intervals are concentrated, and the box plot distribution is compact. This proves that the model's ability to identify these major categories is reliable and consistent. Due to the extremely small number of samples in the U2R category in the training data (severe class imbalance), its performance is most sensitive to the randomness of model initialization, with a standard deviation of accuracy as high as 11.06%. However, despite the fluctuations in absolute values, its performance is statistically significantly better than random guessing, and multiple experiments have provided an expected range of its performance (mean 46.74%, 95% CI [37.04%, 56.43%]), which provides a reliable benchmark for subsequent research.

[0116] Table 13 Statistical analysis of model performance under different random seeds ;

[0117] To rigorously determine the convergence of model training, this invention proposes a "double threshold" criterion: the exploration rate ε must first drop below 0.1, and the average difference in consecutive losses over the last 50 episodes must be less than 0.03. This criterion's design takes into account both the "exploration-exploitation" game dynamics of reinforcement learning and the local fluctuations in the loss function. First, ε < 0.1 ensures that the policy has entered the "exploitation"-dominated phase. At this point, the agent's traversal of the state space almost stops, and subsequent weight updates mainly rely on the first-order information of the policy gradient. If training continues, the incremental gain comes only from small perturbations, and their marginal contribution to the final performance can be ignored. Second, requiring the average loss fluctuation over 50 consecutive episodes to be below 0.03 is to statistically eliminate Gaussian noise introduced by stochastic gradients: such as... Figure 15 , Figure 16As shown, the average loss range of all random seeds within 95% confidence interval does not exceed 0.0256 (maximum appears in seed 2023), indicating that the loss sequence has entered a stationary phase, and there is no trend decline or periodic oscillation; the threshold of 0.03 is about 1.2 times the maximum observed range, which can accommodate the difference between seeds and environmental randomness.

[0118] Four, CICIDS-2017, NSL-KDD, MQTT-IoT-IDS2020 multi-ISP and UNNSW-NB15 two-ISP.

[0119] For the multi-ISP analysis of the CICIDS2017 dataset, Table 14 shows that the method proposed by the present application maintains strong and balanced detection performance in the seven types of traffic. Specifically, the F1-score of DDoS and DoS attacks is as high as 99.46% and 98.47%, respectively, and the recall rate is close to or exceeds 99%, indicating almost no false negatives for large-scale blocking attacks; Brute Force also achieves an F1 of 97.63%, with an accuracy of 98.44% and a recall of 96.82%; Port Scanning maintains an accuracy of 90.37% with a high sensitivity of 99.75% recall, with an F1 of 94.83%, effectively balancing the risk of false positives; Web Attacks and Normal traffic are stable in the F1 interval of 90% and 98%, reflecting the robust characterization of complex application layer anomalies and normal behavior; Only Bots, due to the scarcity of samples, reduces all evaluations to 67.78%, becoming the only shortcoming, but the remaining six categories are better than the same period results reported in the field, verifying the wide applicability and reliability of the model in multi-attack scenarios.

[0120] Table 14 CICIDS2017 multi-prediction ;

[0121] Table 15 shows the performance of representative methods for multi-classification tasks on the CICIDS2017 dataset. The accuracy of mainstream research has generally exceeded 97%. Among them, Zihan W et al. lead with the highest accuracy of 99.35% and F1-score of 99.17%; Yaser Alhasawi et al. and Jieling L et al. also follow closely behind with F1-scores of 98.90% and 98.54%, respectively, showing balanced detection capabilities for each type of attack. In contrast, W. Elmasry et al. achieved an accuracy of 98.95%, but Precision and Recall both dropped to about 95.8%, indicating that the recall for the minority class is still insufficient. The method of the present application maintains a high accuracy of 98.14% while stabilizing Precision, Recall, and F1-score at around 98.2%, with a difference of less than 1 percentage point from the optimal result, verifying its effectiveness and competitiveness in overall detection performance and class balance.

[0122] Table 15 CICIDS2017 overall multi-prediction comparison ;

[0123] The multi-ISP experimental results for the NSL-KDD dataset presented in Table 16 show that SEA-DDQN can achieve efficient and accurate prediction in complex network attack scenarios. Specifically, the model performs particularly well in DoS attack detection, with an accuracy of 98.71%, precision of 99.19%, recall of 99.52%, and F1-score of 99.35%, indicating that it can identify DoS attacks with extremely high confidence, providing reliable technical support for preventing such attacks. In the detection of Probe attacks, the model's accuracy is 94.54%, with precision, recall, and F1-score of 95.10%, 99.38%, and 97.19%, respectively, showing strong identification capabilities for Probe attacks and effectively detecting potential network reconnaissance behaviors in the early detection stage. For the more complex R2L attacks, the model's accuracy is 88.98%, with precision, recall, and F1-score of 91.82%, 96.64%, and 94.17%, respectively, although the accuracy is relatively low, the recall is high, indicating that the model can capture most of the real attack samples when identifying R2L attacks, which helps to take timely protective measures. In the prediction of normal traffic, the model's accuracy, precision, recall, and F1-score are 94.78%, 98.56%, 96.12%, and 97.32%, respectively, which can also accurately distinguish between normal network behavior and various attack behaviors.

[0124] Table 16 NSL-KDD multi-prediction ;

[0125] Table 17 shows that compared with existing research, the schemes of C.P.R. Kanna et al. and Chadia E L A et al. achieve 98.67% and 98.92% in accuracy, respectively, but the former sacrifices part of precision at 100% recall, and the latter falls to 95.44% in recall, resulting in no absolute advantage in F1-score; the works of Zhendong W et al. and Wei et al. also achieve 98.60% and 92.95% in accuracy, but their precision or recall fluctuates significantly, and the comprehensive F1-score is lower than 98%. In contrast, the method in this paper maintains a highly consistent balanced performance in 97.50% accuracy, 97.54% precision, 97.50% recall, and 97.48% F1-score, neither over-biased to a certain index nor robust in overall performance.

[0126] Table 17 NSL-KDD overall multi-prediction comparison ;

[0127] Table 18 compares the two ISP results of SEA-DDQN on the UNSW-NB15 dataset with existing research. From the key indicators of accuracy, precision, recall, and F1-score, most of the literature presents a clear trade-off between recall and precision: Yousefnezhad et al. push the F1-score to 93.34% with an extremely high recall of 99.72%, but the precision drops to 87.37%; Zhendong W et al., Li Jieling et al., and Marwa K et al. also sacrifice precision to achieve more than 90% recall, resulting in F1-score hovering between 84% and 89%. In contrast, the method in this paper achieves a high consistency of four indicators in 95.32% accuracy, 95.47% precision, 95.32% recall, and 95.30% F1-score, indicating that SEA-DDQN can maintain extremely low false positive rate while still fully capturing attack samples, achieving balanced and robust two-class detection performance.

[0128] Table 18 UNSW-NB15 two-prediction comparison ;

[0129] The multi-ISP experimental results of the MQTT-IoT-IDS2020 dataset presented in Table 19 show that the proposed model achieves excellent and balanced detection performance for all attack types: Sparta and Scan_A almost achieve full marks in the four indicators, with F1-score of 100% and 99.99%, respectively; Normal traffic is also stably identified, with F1-score of 99.66%; Scan_sU and Bruteforce are relatively slightly lower, with F1-score of 99.48% and 98.33%, respectively, fully verifying the strong generalization ability and robustness of the model in complex MQTT Internet of Things environment.

[0130] Table 19 MQTT-IoT-IDS2020 multi-prediction ;

[0131] Table 20 shows that compared with existing research, the existing methods generally have a significant gap in the four indicators: Khan et al. and Ullah-DT et al. have an overall accuracy of more than 98%, but the recall rate drops to 86.71% and 82.06%, respectively, causing the F1-score to stay around 90%; Pandey et al.'s precision and recall simultaneously drop to about 87%, and F1 further drops to 85.64%; Lucia et al.'s recall rate is as high as 99.17%, but the overall accuracy is missing and the precision is only 92.14%, with F1-score of 95.53%. In contrast, Otokwala et al., Shirodkar and Akbar et al. have pushed a single indicator to 100% or more than 99.5%, but still have a slight gap in recall or precision. The method in this paper locks the accuracy, precision, recall and F1-score at 99.59% at the same time, while maintaining high detection sensitivity, it completely eliminates the trade-off defect of false positives and false negatives, and is better than the best existing results.

[0132] Table 20 MQTT-IoT-IDS2020 overall multi-prediction comparison ;

[0133] V. Agent Exploration Analysis

[0134] During the experiment, the agent continuously adjusts its strategy through interaction with network traffic to achieve the best balance between exploration and exploitation. This study records the accuracy and exploration rate changes of the model at each training stage on the test set in detail. Figures 17-20In the learning process of the four data sets NSL-KDD, UNSW-NB15, CICIDS2017 and MQTT-IoT-IDS2020 respectively, the exploration rate of the agent on all data sets shows a monotonic decreasing trend, and the accuracy of traffic prediction continues to improve with the training, and finally tends to be convergent, which gradually changes from extensive exploration in the early stage to stable utilization in the later stage, reflecting that the agent effectively identifies the high return action in the strategy space and forms a stable decision strategy. In addition, the loss function curve shows a downward trend on all data sets, and there is slight shock in the early stage of loss, mainly due to the instability of the strategy in the exploration stage, and as the exploration rate decreases, the loss curve gradually becomes smooth, reflecting the controllability and stability of the strategy convergence process. The agent on the four types of ISP data sets shows a reasonable exploration-exploitation trade-off mechanism, which can effectively avoid local optimization in the early stage of training and stably improve the detection performance in the later stage.

[0135] In summary, the present application continuously enhances the adaptability and optimization of the model through continuous interaction with the traffic environment, improves the feature discrimination ability through the combination of attention mechanism, and improves the learning efficiency through focusing on samples with large information amount through priority experience replay, and has higher accuracy and robustness in processing complex network traffic data.

[0136] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application but not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that: it can still modify or equivalently replace the technical solutions of the present application, and these modifications or equivalent replacements also cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. An intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning, characterized in that, Includes the following steps: Step S1: Obtain network traffic datasets for NSL-KDD, UNSW-NB15, CICIDS-2017, and MQTT-IoT-IDS2020; Step S2: Preprocess the acquired network traffic dataset to obtain a normalized dataset; Step S3: Construct a deep attention Q-network SEA-DDQN model based on dual value to predict network security intrusion situations for ISPs; Step S4: Implement the ISP process based on the SEA-DDQN model.

2. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 1, characterized in that, In step S2, the acquired network traffic dataset is preprocessed to obtain a normalized dataset. The specific process is as follows: Step S21: Define the network traffic dataset as... ,in Let be the number of traffic samples, then the feature space As shown below: ; in, Indicates features, An index representing the number of features in the dataset; Represents a subset of features; ; It is a subset of traffic features; It is a subset of content features; It is a subset of time statistical features; It is a subset of host behavior characteristics; The union symbol represents the combination of four feature subsets into a feature space. ; Step S22: For discrete features Perform hot-unique encoding processing; First, discrete feature subsets ;in, Represents discrete characteristics; This represents the discrete feature index; the corresponding value ranges for each discrete feature are as follows: ; ; ;in, , and These represent the ranges of discrete features. Define an indicator function for any discrete feature. As shown below: ; in, ; Indicates sample In discrete features The possible values ​​of ; Then, through hot-unique encoding transformation Right now The feature space is reconstructed as follows: ; Where d is the total number of features after hot unique coding; This represents the hot-unique encoding processing function; This represents a vector concatenation operation; This is the index of all non-discrete features in the original features; It is a discrete feature subset; Finally, hot-coded versions of the original training and test sets are performed separately to generate their respective extended feature matrices. Then normalize it, as shown below: ; in, Indicates normalized features; and Representing features respectively The maximum and minimum values; For indicator functions, when The value is 1 if the condition is met, and 0 otherwise; the attack categories in the dataset are classified and mapped to a set of numerical categories. .

3. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 2, characterized in that, Based on the preprocessed training and test sets of NSL-KDD, UNSW-NB15, CICIDS-2017, and MQTT-IoT-IDS2020, a network ISP reinforcement learning environment is constructed and defined. The ISP process is abstracted into a Markov decision process (MDP) and formally defined as a quadruple. As shown below: ; Among them, state The set of feature vectors for each network traffic item; Action For the traffic type decision set defined in NSL-KDD, UNSW-NB15, CICIDS-2017, and MQTT-IoT-IDS2020, in time The state and action of the agent are represented as follows: and ;award The reward or penalty given to the agent for correctly predicting traffic; the calculation of the reward vector is based on the comparison between the agent's prediction results and the actual labels to evaluate the accuracy of the agent's predictions; discount factor. This is a discount factor used to balance immediate rewards with long-term returns.

4. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 3, characterized in that, Construct a SEA-DDQN model agent, which includes an action selection network and a target Q-value network; the SEA-DDQN model agent predicts captured network traffic and calculates the target Q-value. Action Selection Network and target Q-value network The Q-value function is composed of four fully connected feedforward neural networks, with attention mechanism and ReLU activation function used between the fully connected layers, and each of the two hidden layers contains 128 neurons; The SEA-DDQN model agent at each time step The state is determined by the feature vector of network traffic. The state space of an intelligent agent is defined as follows: Therefore, the SEA-DDQN model agent obtains the state vector for each network traffic. As shown below: ; Each dimension It represents a certain characteristic of network traffic.

5. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 4, characterized in that, An attention mechanism is introduced to optimize the attention allocation of the SEA-DDQN model agent to network traffic features, enhancing discriminative features and suppressing irrelevant features. The specific process is as follows: For each time step, the input network traffic feature state vector After processing by the fully connected layer, a 128-dimensional feature vector is obtained. ,in This represents the network traffic feature vector after processing by the fully connected layer. This represents the dimension index; it is then fed into an attention layer to enhance the focus on important features, with its attention weights... As shown below: ; in, The dimension-reduced projection matrix; To reconstruct the matrix for higher dimensions; The dimensionality reduction ratio is a hyperparameter. This represents the Sigmoid gate function, which compresses the input value to... Intervals, generated attention weights This is used to measure the importance of each feature; the generated attention weights With original features Fusion is achieved through modulation, as shown below: ; in, This represents the feature representation after fusion; The Hadamard product is an element-wise multiplication operation; this operation makes each eigenvalue... All are corresponding weights Dynamic adjustment.

6. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 5, characterized in that, The SEA-DDQN model agent relies on the input features of a deep neural network. The system generates a list of actions, which is presented as action vectors; the final Q-value is used to evaluate whether the attack behavior was successfully predicted. First, the category labels in the NSL-KDD, UNSW-NB15, CICIDS-2017, and MQTT-IoT-IDS2020 datasets are mapped to sets of digit categories, as shown below: Where 0 represents Normal; 1 represents DoS; 2 represents Probe; 3 represents U2R; and 4 represents R2L. Where 0 represents normal traffic and 1 represents abnormal traffic; Where 0 represents Bots; 1 represents Brute Force; 2 represents DDoS; 3 represents DoS; 4 represents Normal; 5 represents Port Scanning; and 6 represents Web Attacks. Where 0 represents Normal; 1 represents Bruteforce; 2 represents Scan_A; 3 represents Scan_sU; and 4 represents Sparta. Then, the action space of the SEA-DDQN model agent is defined as: ; ; ; ; Each action These correspond to the prediction category decisions in the action list; Indicates a category index; This indicates a definition symbol, where the left side of the symbol is defined by the right side.

7. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 6, characterized in that, A dynamic weighted reward mechanism is designed to balance class bias by constructing class-sensitive rewards and using an inverse frequency-weighted strategy to determine class weights. This enhances the sensitivity of the SEA-DDQN model agent to different attack categories and balances the impact of class bias in the dataset on the learning process. The specific process is as follows: First, define the reward function. The specific form is as follows: ; in, Indicates the weight of the attack category; This indicates a correct prediction; Indicates an incorrect prediction; Then, using an inverse frequency-weighted strategy, weights are assigned by considering the frequency of occurrence of each attack category, as shown below: ; ; in, express The set of attack category weights; express The weight of the attack; Indicates a weight that is not under attack; express The weight of the attack; express The weight of the attack; express The weight of the attack; express The set of attack category weights; express The weight of the attack; express The weight of the attack; express The weight of the attack; express The weight of the attack; express The weight of the attack; The weighting rules have been updated as follows: ; in, The total sample size is 1. Number of attack types This indicates the total number of attack samples; For the two-prediction task, a positive reward is given when the SEA-DDQN model agent correctly predicts the sample, and a negative penalty is given otherwise, as shown below: ; in, express The reward setting rules for each traffic item in the dataset; To simulate latency feedback scenarios in real-world networks, an n-step reward accumulation mechanism is adopted, and the actual reward stored in the experience replay pool is the n-step accumulated reward. Instead of one-step instant rewards The agent stores the transition sequence of the most recent n steps during its interaction with the environment, and calculates the discounted cumulative reward after reaching n steps. As shown below: 。 8. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 7, characterized in that, A priority sampling strategy (PER) is adopted to optimize the sampling process of empirical network traffic. Due to the difference in the importance of attack categories, TD-Error is introduced as a key indicator to measure the importance of samples. By defining a priority update formula, the sampling probability of samples is dynamically adjusted, as shown below: ; in, Indicates the first The probability that a given network traffic sample will be taken; It is a constant used to ensure that all empirical network traffic has a probability of being sampled, and to avoid some samples having a zero sampling probability due to the TD-Error being too small; This represents the current network traffic count in the experience replay pool. This is a priority parameter used to control the degree of influence of TD-Error on the sampling probability; For the first The TD-Error corresponding to each empirical network traffic is shown below: ; in, Represents the target Q-value network; Indicates the first One state; Indicates the first Actions in each state; This represents the Q-value estimate of the target network; This represents the maximum value of the Q-value estimate for the target network; Indicates the first One state; Indicates the first Actions in each state; This represents the Q-value estimate of the current network.

9. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 7, characterized in that, In step S4, the ISP process is implemented based on the SEA-DDQN model, including the following steps: Step S41: Based on the SEA-DDQN model, convert the state feature vector space of each network traffic sample into a single vector. As an action selection network The input is used to capture the characteristics of each traffic sample through hierarchical nonlinear transformation; Step S42: In the input layer, convert the network traffic sample state feature vector space... The d-dimensional feature vectors in the model are mapped to a 128-dimensional hidden space as follows: ; in, This represents the preactivation vector of the first layer of the neural network; This represents the weights of the first layer of the neural network; This indicates the bias of the second layer of the neural network; Step S43: Through the attention layer and using The activation function implements a non-linear transformation, enabling the neural network to learn more complex representations of traffic features, as shown below: ; ; in, This is an element-wise multiplication method used to combine the weighted result of the attention mechanism with the original features; This represents the result after processing with the ReLU activation function, signifying the feature representation learned by the network after nonlinear transformation. Step S44, the second fully connected layer, and subsequent layers are similar to the previous layers, as shown below: ; ; ; ; in, This represents the preactivation vector of the second layer of the neural network; This represents the weights of the second layer of the neural network; This indicates the bias of the second layer of the neural network; This represents the result of the second layer after processing with the ReLU activation function; Step S45: In the final output layer, the action space of each agent is generated. The corresponding Q value, its output value represents the current state. The expected reward accumulation for the corresponding action is as follows: ; in, These are the weights of the fourth layer of the neural network; The result of the third layer after processing with the ReLU activation function represents the feature representation learned by the network after nonlinear transformation. This is the bias for the fourth layer of the neural network; The Q-value of the first action in the action space element vector; This represents the Q-value estimate of the network; For the corresponding action space The expected cumulative reward of the middle element; For the current state After action selection network The agent uses the obtained Q-value vector. The strategy for selecting the best predictive action is as follows: ; in, Indicates the size of the predicted action space; Indicates the current state The optimal predictive action is given below; This indicates the exploration rate at the current time step; Indicates the state Select action The probability of; The update rules are as follows: ; in, It is the initial exploration rate. It is the minimum exploration rate threshold. It is the attenuation coefficient. This is the current learning time step.

10. The intrusion situation prediction method based on SEA-DDQN adaptive reinforcement learning according to claim 9, characterized in that, Target Q-value network The computation process and action selection network The calculation process is similar, and its input is... The target Q value is calculated as follows: ; in, The current state Take action below The rewards received; The target network's Q-value; Based on PER, sampling probabilities are assigned according to empirical importance, and sampling bias is introduced. A mean squared error loss function with importance sampling weights is constructed to compensate for the bias introduced by non-uniform sampling, as shown below: ; in, Represents the loss function; This represents the b-th flow in the experience replay pool; Indicates state; Indicates an action; This indicates the state of the action selection network for the b-th traffic. and actions The predicted Q value; This represents the target q value of the b-th flow; This indicates the probability that the b-th traffic item will be sampled; Indicates the importance sampling coefficient; The number of network traffic samples drawn from the experience replay pool; In the initial stage of learning, as learning progresses, The value gradually increases and eventually approaches 1, as shown below: ; in, This represents the value of β at time step t; This represents the value of β at time step t-1; This represents the initial value of β; This represents the final target value for the β value; the β value is a parameter used to control the degree of forgetting or retention of past experiences. In action selection network Medium parameters Parameters are updated in real time through backpropagation. The update formula is as follows: ; in, Represents the parameters of the neural network; Indicates the learning rate; Indicates parameters The gradient; Target Q-value network parameters Updates are achieved by combining Polyak average soft updates and hard updates, specifically by performing an action selection network every 100 time steps. Medium parameters Network with target Q value parameters The parameters of the target Q-value network are updated by performing a weighted average; the action selection network is updated every 1000 time steps. Medium parameters Directly copy to the network with the target Q value parameters In the middle, as shown below: ; in, The parameters of the target q-value network are represented. for Average soft update coefficient.

Citation Information

Patent Citations

  • Network intrusion detection method for reinforcement learning near-end strategy optimization

    CN117579343A

  • Network intrusion detection model construction method based on deep reinforcement learning

    CN119051891A

  • Reinforcement learning intrusion detection method and system based on dynamic network feature screening

    CN119696934A

  • Network intrusion prevention method and device, equipment and storage medium

    CN119814459A

Cited By

  • Data detection method based on layered three-stage deep reinforcement learning

    CN121509111A

  • A data detection method based on hierarchical three-stage deep reinforcement learning

    CN121509111B