Security policy matching method based on reinforcement learning, computer program product, storage medium and terminal
Through the security policy matching method based on reinforcement learning, network security policies are dynamically adjusted to deal with new threats, solving the problem that existing technologies are difficult to deal with new attacks, and achieving efficient response and policy updates to complex network environments.
Patent Information
- Application Number
- CN202510172710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively deal with new or variant cyber attacks, and the rules are maintained and updated in a complex manner, and statistical methods have limited capabilities when identifying unknown threats.
The security policy matching method based on reinforcement learning is adopted to optimize the interaction between the model and the network environment through the policy optimization method, and the network state changes are sensed in real time, and the strategies are dynamically adjusted to deal with new threats. The method includes initializing the network environment and policy optimization model, interactive data storage, calculating probability ratios and dominant functions, updating the policy network, and performing a blocking operation if necessary.
It realizes rapid optimization and update of new network threats, improves the ability to respond to complex and dynamic network environments, simplifies the policy update process, and improves the efficiency of security policy generation and response speed.
Smart Images

Figure CN120017375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a security policy matching method based on reinforcement learning, a computer program product, a storage medium and a terminal. Background Art
[0002] Security policy generation models play an important role in the field of network security. The rule-based security policy generation model is one of the earliest security policy generation models. This model formulates a series of rules to detect and defend against security threats based on known threat characteristics and attack patterns. However, the limitation of this model is that it can only deal with known threats and often cannot output effective security policies for new or variant attack methods. In addition, as threats continue to increase and change, the maintenance and update of rules are becoming more and more complex and difficult.
[0003] With the development of statistics, security policy generation models based on statistics have gradually emerged. This model analyzes data such as network traffic and user behavior, uses statistical methods to identify abnormal behaviors, and then generates corresponding security policies. However, statistical methods can often only identify abnormal behaviors that are different from known patterns, and have limited ability to identify unknown threats. Summary of the invention
[0004] The purpose of the present invention is to overcome the problems of the prior art and provide a security policy matching method based on reinforcement learning, a computer program product, a storage medium and a terminal.
[0005] The purpose of the present invention is to achieve the following technical solution: a security strategy matching method based on reinforcement learning, the method comprising the following steps:
[0006] Initialize the network environment, policy optimization model and experience buffer, the network environment includes the network state, the network state includes the network traffic, and the policy optimization model includes the policy network and the value network;
[0007] Given a state s t , the policy network is distributed according to the probability Select an action t And execute, the network environment is in the following state s t+1 and immediate reward r t Respond and get the given status s t The state value V(s) output by the next value network t ), the interaction data s t ,a t ,r t ,s t+1 , V(s t ) is stored in the experience buffer;
[0008] During iterative updates, the current policy is calculated to be in a given state s t Take action a t The probability ρ θ (a t ∣s t ), the previous strategy is in a given state s t Take action a t Probability And calculate the probability ρ θ (a t ∣s t ) and probability The ratio k t ; The advantage function is calculated using the generalized advantage estimation method Advantage function Indicates state s t The next line is a t Deviation from the mean;
[0009] According to the interaction data, the ratio k t , advantage function Calculate the objective function, including the objective function based on KL penalty optimization and the objective function of limiting the ratio of new and old strategies for shearing operations, and update the strategy network according to the objective function;
[0010] Repeat the above steps until the termination condition is reached, the policy network gradually converges to the optimal scheduling strategy, and the optimal scheduling strategy is used to deal with network threats.
[0011] In one example, before the network environment, the policy optimization model and the experience buffer are initialized, the process further includes:
[0012] Monitor the network in real time, collect traffic information and store it in the network traffic database, and determine network attack information based on the traffic information, including source IP address, destination IP address, traffic rate, and threat type;
[0013] The network attack information is matched in the policy database. If the match is successful, the corresponding matching policy in the policy database is executed. If the match fails, the network environment, policy optimization model and experience buffer are initialized.
[0014] In one example, after the optimal scheduling strategy is used to handle the network threat, the method further includes:
[0015] Storing the optimal scheduling strategy in a strategy database;
[0016] Calculate the ratio r1 of events in the policy database that have the same source IP address and target IP address as network attack events to the total number of all processed network attack events; calculate the ratio r2 of network traffic in the network traffic database that has the same source IP address and target IP address as network attack events to the total network traffic. If r1>r2*threshold, it is determined to be a high-risk path and a blocking operation is performed.
[0017] In one example, the objective function L optimized based on the KL penalty term is KL The calculation expression of (θ) is:
[0018]
[0019] in, Represents the expected value of the estimate; μ represents the penalty coefficient of KL divergence; Represents a given state s t Next, the old strategy The probability distribution of all possible actions; ρ θ (·|st) represents a given state s t Under the current strategy ρ θ Probability distribution over all possible actions; represents the KL divergence between the new and old strategies.
[0020] In one example, after calculating the objective function based on the KL penalty term optimization, the method further includes calculating the expected value of the KL divergence between the new and old strategies, and adjusting the penalty coefficient according to the relationship between the expected value and the target expected threshold.
[0021] In one example, the calculation expression of the expected value d of the KL divergence between the new and old strategies is:
[0022]
[0023] in, represents the expected value of the estimate; represents the KL divergence between the new and old strategies.
[0024] In one example, the objective function L of limiting the ratio of the new and old strategies to perform the shear operation is CLIP The calculation expression of (θ) is:
[0025]
[0026] in, represents the expected value estimated at time t; σ represents the truncation hyperparameter; clip() represents the truncation function, limiting the proportion k t Greater than or equal to 1-σ and less than or equal to 1+σ.
[0027] It should be further explained that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.
[0028] The present invention also includes a computer program product, including a computer program, which, when executed by a processor, implements the steps of the security policy matching method based on reinforcement learning formed by any one of the above examples or a combination of multiple examples.
[0029] The present invention also includes a storage medium on which computer instructions are stored. When the computer instructions are executed, the steps of the security policy matching method based on reinforcement learning formed by any one of the above examples or multiple examples are executed.
[0030] The present invention also includes a terminal, including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor executes the computer instructions, the steps of the security policy matching method based on reinforcement learning formed by any one or more of the above examples are executed.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] 1. In one example, reinforcement learning perceives changes in the state of the network environment in real time through continuous interaction between the policy optimization model and the network environment, and can promptly capture dynamic changes in the network environment, such as the emergence of new attack traffic; based on the received reward signal, the model can evaluate the effectiveness of the current strategy and dynamically adjust the strategy accordingly; when faced with new network threats, if the old strategy cannot effectively respond and the reward is reduced, the model will automatically explore new strategies to find better response plans, thereby achieving rapid optimization and updating of the strategy, enabling the model to better cope with the complexity and dynamics of the network environment, and when the network environment or attack methods change, it can quickly adjust the strategy to meet new challenges.
[0033] 2. In one example, the KL divergence constraint on the amplitude of parameter changes is added to the objective function, and a new objective function clipping advantage function is constructed, which effectively limits the step size of the policy update and avoids excessive changes in the policy parameters in each iteration, thereby ensuring the smoothness and stability of the policy update. At the same time, it simplifies the problem-solving method, improves the efficiency of security policy generation, and achieves efficient response to network threats.
[0034] 3. In one example, by introducing a policy database, when a network attack occurs, the information of the network attack is matched in the policy database first, and the corresponding matching historical policy in the policy database can be quickly selected, further improving the response rate to network threats.
[0035] 4. In one example, by comparing the data in the policy database and the network traffic database, it is possible to more accurately identify high-risk paths, promptly discover and respond to potential threats, immediately perform blocking operations, and reduce the impact of attacks on the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings. The accompanying drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0037] Figure 1 A method flow chart provided for an example of the present invention;
[0038] Figure 2 A method structure diagram provided for an example of the present invention;
[0039] Figure 3 The objective function L provided for an example of the present invention is CLIP In the advantage function Graph of functions greater than 0;
[0040] Figure 4 The objective function L provided for an example of the present invention is CLIP In the advantage function Graph of functions less than 0;
[0041] Figure 5 A simulation result diagram of the security strategy generated based on the PPO algorithm and other algorithms in terms of defense cycle provided by an example of the present invention;
[0042] Figure 6 A simulation result diagram of a security policy generated based on the PPO algorithm and other algorithms in terms of intercepted data packets provided as an example of the present invention;
[0043] Figure 7 This is a simulation result diagram of throughput of security policies generated based on the PPO algorithm and other algorithms provided as an example of the present invention. DETAILED DESCRIPTION
[0044] The technical solution of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In the description of the present invention, it should be noted that the directions or positional relationships indicated by “center”, “up”, “down”, “left”, “right”, “vertical”, “horizontal”, “inside” and “outside”, etc., are directions or positional relationships based on the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0046] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, "connected" and "connection" should be understood in a broad sense, for example, it can be directly connected or indirectly connected through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0047] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0048] In one example, if Figure 1-Figure 2 As shown, a security policy matching method based on reinforcement learning includes the following steps:
[0049] S1: Initialize the network environment, strategy optimization model and experience buffer.
[0050] Among them, the network environment includes network status, action space and reward function. The network status includes network traffic. The network traffic includes normal traffic and potential attack traffic, such as source IP address, destination address, traffic size, traffic rate, network type, threat type, etc. Network types include local area network, wide area network, metropolitan area network, etc. Threat types include network scanning, malware attack, phishing, denial of service attack, etc. The action space includes security policies and, optionally, defense policies. The reward function is used to give positive rewards when an attack is successfully blocked and negative rewards when a false alarm occurs. The policy optimization model in this example is a proximal policy optimization (PPO) model, which includes a policy network and a value network. The policy network is in a given state s tThe probability distribution of the action is selected, and the objective function of the PPO algorithm is used to update the parameters of the policy network. The value network is used to estimate the state value or action value. Every time the PPO algorithm updates the policy, it ensures that the difference between the new policy and the old policy is not too large. This restriction helps maintain the stability of the training and prevents the policy from changing drastically during the update process, which can lead to training failure or slow convergence. Of course, the policy optimization model can also use the Trust Region Policy Optimization (TRPO) model, the Deep Deterministic Policy Gradient (DPG) algorithm model, the SAC algorithm model, the Policy Gradient (PG) algorithm model, etc. The experience buffer is used to store the interaction information between the policy optimization model and the network environment, including state, action, reward and other information, for subsequent policy updates.
[0051] S2: Collect interaction data.
[0052] In each iteration, the current policy network is used to interact with the network environment. Given a state s t , the policy network is distributed according to the probability Select an action t And execute, the network environment is in the following state s t+1 and immediate reward r t Respond and get the given status s t The state value V(s) output by the next value network t ), the interaction data s t ,a t ,r t ,s t+1 , V(s t ) is stored in the experience buffer for subsequent strategy updates. t Including: fast flow control, slow flow control, threat processing strategy, etc. t It is the current status information about the network traffic control environment, including source IP address, destination IP address, traffic bytes per second, network type and threat type.
[0053] S3: Calculate the new and old strategies to take a t The probability ratio of the behavior k t and advantage functions.
[0054] During iterative updates, the current policy is calculated at time t when the policy network is in a given state s t Take action a t The probability ρ θ(a t ∣s t ) and the previous strategy (old strategy) takes action a in the same state t Probability And calculate the probability ρ θ (a t ∣s t ) and probability The ratio k t , to control the update amplitude of the new strategy, the ratio k t The calculation expression is:
[0055]
[0056] Among them, k t The closer the ratio is to 1, the more it proves that the current strategy is in state s at time t. t Next take a t The probability of the behavior ρ θ (a t ∣s t ), taking action a in the same state as the previous strategy t Probability The closer the values are, the smaller the difference between the new and old strategies is. If the difference between the new and old strategies is obvious (exceeding the set difference threshold) and the advantage function is large (when the advantage function is large, it indicates that the state s t Take action a t If an action has a higher value relative to other actions, which means that the action not only obtains a higher immediate reward in the current state, but is also better than other possible actions), the update amplitude is appropriately increased.
[0057] Furthermore, the generalized advantage estimation method is used to calculate the advantage function Advantage function Indicates state s t The next line is a t Deviation from the mean. This example uses the generalized advantage estimation (GAE) method to calculate the advantage function. The calculation expression is:
[0058]
[0059] Right now:
[0060]
[0061] Among them, l represents the time step index; γ is a hyperparameter for GAE to balance the bias and variance. When γ is close to 1, GAE tends to multi-step returns, thereby reducing the variance; when γ is close to 0, GAE tends to one-step returns, thereby reducing the bias. λ∈[0,1] is the introduced hyperparameter, represents the timing difference error, The calculation formula is:
[0062]
[0063] Among them, V represents a learned state-value function; ω represents the parameters of the state-value function.
[0064] S4: Update the policy network.
[0065] Specifically, according to the interaction data, the ratio k t , advantage function Calculate the objective function, including the objective function optimized based on the KL penalty term and the objective function of limiting the ratio of the new and old strategies for shearing operations, and update the strategy network according to the objective function.
[0066] The objective function usually consists of two parts: one is the goal of policy gradient optimization (i.e., maximizing the expected cumulative reward, in this example, the objective function is optimized based on the KL penalty term), and the other is the clipping function (or clipping term) used to limit the difference between the new and old policies. When the difference between the new and old policies is obvious, the clipping function takes effect to limit the amplitude of the update. At the same time, when the advantage function is large, it tends to increase the probability of the action in the current policy. Specifically, the difference between the new and old policies is measured by the KL divergence, a threshold of the KL divergence is set, and an attempt is made to keep the KL divergence between the new and old policies within this threshold during the optimization process. This can be achieved by adding a penalty term of the KL divergence to the objective function, or by adaptively adjusting the penalty coefficient to maintain a stable update of the policy. For the update amplitude, a clipping term is introduced to limit the range of the probability ratio of the new and old policies, that is, when calculating the objective function, a clipping function is introduced to limit the range of the probability ratio of the new and old policies. If the probability of the new policy exceeds the probability of the old policy by a certain range (for example, between 0.8 and 1.2), it is clipped to ensure that the amplitude of the policy update is not too large. This approach helps avoid overly aggressive policy updates and thus maintains training stability.
[0067] S5: Clear the experience buffer, enter the next iteration, and repeat steps S1-S4 until the termination condition is reached, that is, the iteration is stopped when the predefined number of iterations is reached or the convergence standard is met. It is considered that the policy network has converged to the optimal scheduling strategy, and the optimal scheduling strategy is used to deal with network threats, that is, actions are selected according to the optimal strategy (such as fast flow control, slow flow control, banned flow control, etc.) to maximize the cumulative reward, thereby effectively responding to network threats.
[0068] Preferably, if Figure 2 As shown, after calculating the advantage function, the value network is also updated. The input of the value network includes the current state and the action probability distribution generated by the policy network; the output of the value network includes the value estimate of the state-action pair, which is used to calculate the policy gradient and thus to calculate the advantage function. The update of the value network and the update of the policy network are performed alternately. Before each update of the policy network, the value network is updated first to ensure its accuracy. This alternating update method contributes to the stability and efficiency of the algorithm. Specifically, the value network update includes the following steps:
[0069] a) Calculate the target value: The target value is the true value that the value network attempts to approximate. This example uses the generalized advantage estimation (GAE) method to calculate it.
[0070] b) Calculate the loss function: This is achieved by the error between the predicted value and the target value. In the PPO algorithm, this error is usually measured using the mean squared error (MSE).
[0071] c) Update the value network parameters: Use the stochastic gradient descent optimization algorithm to minimize the loss function to update the value network parameters and improve the accuracy of the value network's estimation of state value.
[0072] In one example, the objective function L optimized based on the KL penalty term KL The calculation expression of (θ) is:
[0073]
[0074] in, represents the expected value of the estimate; Represents a given state s t Next, the old strategy The probability distribution of all possible actions; ρ θ (·∣s t ) represents a given state s t Under the current strategy ρ θ Probability distribution over all possible actions; represents the KL divergence between the new and old strategies; μ represents the penalty coefficient of KL divergence. The selection of the initial value of the penalty term μ has almost no effect on the PPO algorithm because it can adapt to the KL divergence of the new and old strategies at each iteration. First, the KL divergence threshold is set, and then the expected value d of the KL divergence between the new and old strategies is calculated, and then the expected value is calculated based on the expected value and the target expected threshold [d tar g / 1.5,d tar g ×1.5] to adjust the penalty coefficient. The calculation expression of the expected value d is:
[0075]
[0076] in, represents the expected value of the estimate; represents the KL divergence between the new and old strategies. tar g / 1.5, which proves that the divergence is small and needs to weaken the penalty, and μ is adjusted to μ / 2; if d>d tar g ×1.5, which proves that the divergence is large and the penalty needs to be increased, so μ is adjusted to 2μ.
[0077] In this example, the objective function L KL Contains a penalty term to penalize the difference between the new policy and the old policy. Specifically, the objective function L KL The penalty term calculates the difference between the new policy and the old policy, multiplies it by a penalty coefficient (usually a positive number), and then adds it to the objective function. When the difference between the new policy and the old policy is large, the penalty term gives a large negative reward, thereby encouraging the algorithm to keep the difference small when updating the policy.
[0078] In one example, the objective function L that limits the ratio of the new and old strategies to perform the shear operation is CLIP The calculation expression of (θ) is:
[0079]
[0080] in, represents the expected value estimated at time t; σ represents the truncation hyperparameter; clip() represents the truncation function, limiting the proportion k t is greater than or equal to 1-σ and less than or equal to 1+σ to ensure convergence; finally L CLIP According to the min function, the smaller value between the untruncated and truncated targets is selected to form the target lower limit to constrain the range of change. CLIP In the advantage function Greater than 0 and The function graphs when it is less than 0 are as follows Figure 3-Figure 4 As shown, according to the advantage function Is it greater than 0, L CLIP It can be divided into two cases: if the advantage function is positive, it is necessary to increase the ratio of the new and old strategies k t , however, when k t >1+σ, no additional incentive will be provided; if the advantage function is negative, the ratio of the new and old strategies k needs to be reduced t , but in k t When <1+σ, no additional incentives are provided, so that the difference between the new and old strategies is limited to a reasonable range.
[0081] In this example, the objective function L CLIP Contains a clipping term to limit the difference between the new policy and the old policy. Specifically, the clipping term calculates the probability ratio of the new policy and the old policy in a given state and clips this ratio to a predetermined range (usually [1-σ, 1+σ], where σ is a small positive number). When the difference between the new policy and the old policy exceeds this range, the clipping term pulls it back to this range, thereby ensuring that the magnitude of the policy update is not too large.
[0082] In order to further illustrate the beneficial effect of the present invention in generating security policies based on the PPO model, corresponding experiments are now carried out to verify the actual effect of the present invention. The specific contents of the experiments are as follows.
[0083] In device-to-device (D2D) communication, distributed denial of service (DDoS) attacks are very harmful because they can cause network structure damage. To this end, the present invention optimizes security policies for DDos attacks in a D2D communication environment. The datasets used to train the PPO-based security policy optimization model include a simulation dataset and a CICDDoS2019 dataset, as follows:
[0084] 1. Simulate the Slowloris attack in the D2D communication network and generate a Slowloris dataset specific to the D2D network. The Slowloris attack is a denial of service (DoS) attack that exploits vulnerabilities or design flaws in a Web server, aiming to make it unavailable by occupying all available connections on the server. The simulation process is as follows: establish an HTTP connection with the target server; send a request containing only part of the HTTP header information and keep the connection open. The request sent by the attacker is legitimate, but lacks complete HTTP header information, causing the server to wait for the request to complete; maintain the connection by sending a keep-connected request and periodically send bytes to keep the connection from being disconnected; repeat the above steps with multiple such connections to occupy the server's connection resources; when the number of concurrent connections on the server reaches the upper limit, it will not be able to accept new connection requests, resulting in a denial of service.
[0085] 2. The CICDDoS2019 dataset is a dataset for distributed denial of service (DDoS) attack detection released by the Cyber Intelligence and Cybersecurity Center (CIC) at the University of New Brunswick, Canada. The dataset covers a variety of DDoS attack types, such as HTTP flood, UDP flood, etc., as well as normal traffic data, which helps to comprehensively analyze the characteristics and patterns of DDoS attacks. The dataset records real DDoS attack events and has high real-time and practicality.
[0086] In the process of PPO model training, using real data sets and simulated data sets each has significant advantages. The real data set can accurately reflect the complexity and diversity of the real world, which helps to improve the generalization ability of the model and reduce the risk of overfitting because it contains noise and anomalies in the actual data. The simulated data set is highly controllable and repeatable, and can generate data as needed, reduce data collection costs, and avoid privacy and security issues. Therefore, the real data set and simulated data set are used together to train the model to make full use of the advantages of both and further improve the performance and stability of the model.
[0087] The present invention is compared with TRPO, DPG and SAC methods in the data set, and the results are as follows Figure 5-Figure 7 As shown, Figure 5-Figure 7 The horizontal axis in the middle represents the defense cycle. Figure 5 The middle vertical axis represents delay, Figure 6 The vertical axis represents the intercepted data packets. Figure 7 The vertical axis represents the throughput. It can be seen that the present invention can achieve relatively good results in evaluation indicators such as security, network quality, and time efficiency. The present invention can effectively improve security, ensure network quality, improve time efficiency, and ensure positive optimization when generating corresponding security policies in response to network threats.
[0088] In one example, before initializing the network environment, the policy optimization model, and the experience buffer, the process further includes:
[0089] S01: Monitor the network in real time, collect traffic information and store it in the network traffic database, and determine network attack information through fine-grained monitoring of traffic, including source IP address, destination IP address, traffic rate (such as the number of traffic bytes per second), threat type, and network type;
[0090] S02: Match the network attack information in the policy database. If the match is successful, execute the corresponding matching policy in the policy database. If the match fails, go to step S1, that is, output the optimal scheduling policy through the policy network of the policy optimization model to handle the network threat.
[0091] Furthermore, in step S01, a flow monitoring module is used to obtain flow information. The flow monitoring module monitors and analyzes network flow based on SFlow, SNMP and NetFlow protocols. Real-time capture of network device status is achieved by integrating multiple protocols (SNMP, NetFlow, sFlow). These protocols can provide device status information at different levels and angles, ensuring that the system can fully understand the working conditions of network devices and respond quickly to changes. SNMP, NetFlow and sFlow are now explained:
[0092] (1) SNMP
[0093] SNMP is a network management protocol used to collect and organize management information in a computer network in order to monitor and control network devices and applications. It uses a client-server model, in which a management system (usually a network management system) communicates with managed devices through the SNMP protocol. SNMP is based on an agent-manager architecture, in which there are two key roles: manager and agent. The manager is part of the network management system and is responsible for monitoring and controlling devices and applications in the network. The manager sends SNMP requests to the agent to obtain the status and information of the device; the agent is a software module installed on the managed device, which is responsible for collecting information about the device and responding to SNMP requests from the manager. The agent can be a hardware module on a network device (such as a router or switch) or a software program on a computer system. Once the SNMP protocol is started in the network, the network management system NMS, as the network management center of the entire network, will manage the device. Each managed device contains an agent residing on the device, multiple managed objects, and a management information base MIB. The NMS interacts with the agent running on the managed device, and the agent completes the NMS's instructions by operating the MIB on the device side. The working principle of SNMP is to send protocol data units (also called SNMP GET requests) to network devices that respond to SNMP. Users can track all communication processes and obtain data from SNMP through network monitoring tools.
[0094] The workflow of SNMP is as follows:
[0095] a) The manager sends an SNMP request to the agent to obtain device load and connection status information, such as the device's CPU utilization and memory usage.
[0096] b) After receiving the request, the agent performs corresponding operations according to the request type, such as obtaining device status, configuration information, etc.
[0097] c) The agent packages the results of the request into an SNMP response and sends it back to the manager.
[0098] d) After receiving the response, the manager parses the response data and processes it accordingly, such as displaying status information, generating alarms, etc.
[0099] Traffic monitoring based on SNMP protocol collects variables related to specific devices and traffic information through network device MIB. Including: number of input bytes, number of input non-broadcast packets, number of input broadcast packets, number of input packet discards, number of input packet errors, number of input unknown protocol packets, number of output packets, number of output non-broadcast packets, number of output broadcast packets, number of output packet discards, number of output packet errors, number of output team leaders, etc. Similar methods also include RMON. This method is implemented using software methods, does not require network modification or additional components, has simple configuration and low cost. In this module, SNMP is used to collect core operating indicators of the device, such as CPU utilization, memory usage, network interface status, etc. By periodically polling to obtain these indicators, it is possible to judge the load of the device and trigger an alarm or adjust the policy when the set threshold is exceeded.
[0100] (1) NetFlow protocol
[0101] NetFlow goes a step further. The information that SNMP targets is generally centered around network element devices, such as Interface throughput, the number of bad frames received, CPU / RAM utilization, etc. NetFlow focuses on the characteristic information of the traffic transmitted on the network link, and this information can more directly reflect the current distribution of access behaviors on the network and the actual service quality level obtained by the contract customers at this time. Netflow is a network packet switching technology that is used to accelerate data exchange and count IP data flows passing through network devices. Netflow is built on the session level, and each data flow corresponds to a session information. Netflow consists of two parts, one is the collection and caching of data flows; the other is the data export mechanism through UDP. In fact, a Netflow contains data from the same TCP session, but a TCP session may contain multiple Netflow flows.
[0102] The working principle of NetFlow is as follows: NetFlow uses the standard exchange mode to process the first IP packet data of the data flow, generates the NetFlow cache, and then transmits the same data in the same data flow based on the cache information, and no longer matches the relevant access control and other policies. The NetFlow cache also contains the statistical information of the subsequent data flow. NetFlow has two core components: NetFlow cache, which stores IP flow information; NetFlow's data export or transmission mechanism, which NetFlow uses to send data to the network management collector.
[0103] The main information and functions of NetFlow flow records include: who: source IP address; when: start time, end time; where: source IP address, source port number, destination IP address, destination port number (access path); what: protocol type, destination IP address, destination port (what application); why: baseline, threshold, characteristics (whether normal); how: traffic size, number of data packets (access status).
[0104] Therefore, in this example, the NetFlow protocol is used to monitor detailed data of network traffic, such as packet rate, protocol usage, etc.
[0105] (3) sFlow protocol
[0106] NetFlow's cache and NDE mechanism mean that there is an impact on the device CPU and memory, but sFlow uses a dedicated chip built into the hardware to eliminate the impact on the CPU and memory, and sFlow can monitor each port, while NetFlow obviously cannot, one because it can only monitor interfaces with enabled functions, and the other is that it uses the port mirroring function. Based on the above two main reasons, the present invention introduces the sFlow protocol for traffic monitoring. sFlow (SampledFlow) is a network traffic monitoring technology based on message sampling, which is mainly used for statistical analysis of network traffic. sFlow provides interface-based traffic analysis, which can monitor traffic conditions in real time, promptly discover abnormal traffic and the source of attack traffic, and provide great convenience for daily inspection and maintenance.
[0107] Specifically, the sFlow system includes an sFlowAgent (sFlow agent) embedded in the device and a remote sFlow Collector (sFlow collector). Among them, the sFlowAgent obtains interface statistics and data information through sFlow sampling, and encapsulates the information into sFlow messages. When the sFlow message buffer is full or the sFlow message cache time (cache time is 1 second) times out, the sFlowAgent will send the sFlow message to the specified sFlow Collector. The sFlowCollector analyzes the sFlow message and displays the analysis results. The sFlowAgent provides two sampling methods for users to analyze network traffic conditions from different angles, namely Flow sampling and Counter sampling. Among them, Flow sampling is that the sFlowAgent device samples and analyzes the message according to a specific sampling direction and sampling ratio on the specified interface to obtain relevant information about the message data content. This sampling method mainly focuses on the details of the traffic, so that the flow behavior on the network can be monitored and analyzed. Counter sampling is that the sFlowAgent device periodically obtains traffic statistics on the interface. Compared with Flow sampling, Counter sampling only focuses on the amount of traffic on the interface, not the detailed information of the traffic. Therefore, this example monitors detailed data of network traffic based on the sFlow protocol and collects traffic statistics and performance data.
[0108] Through SNMP, NetFlow, and sFlow protocols, we can obtain information such as source IP, destination IP address, protocol type, packet size, and flow rate, so as to conduct in-depth analysis of the device's flow situation. Through fine-grained monitoring of flow, the system can detect potential anomalies or attack patterns in advance and adjust defense strategies at the appropriate time.
[0109] In one example, after using the optimal scheduling strategy to handle network threats, the following steps are also included:
[0110] Store the optimal scheduling strategy and processed network attack information in the strategy database,
[0111] Calculate the ratio r1 of events in the policy database that have the same source IP address and target IP address as network attack events to the total number of all processed network attack events; calculate the ratio r2 of network traffic in the network traffic database that has the same source IP address and target IP address as network attack events to the total network traffic. If r1>r2*threshold, it is determined to be a high-risk path and a blocking operation is performed.
[0112] Specifically, the calculation expression of the proportion r1 is:
[0113]
[0114] where s same Represents the number of network attacks that occur with the same source IP address and target ID address in the policy database, s total Represents the total number of network attacks recorded in the policy database.
[0115] The calculation expression of proportion r2 is:
[0116]
[0117] where t same Represents the network traffic between the same source IP address and destination ID address in the traffic database, s total Represents all network traffic on the network.
[0118] In this example, the threshold is 1.5. If r1>r2*1.5, it is determined to be a high-risk path and is directly blocked.
[0119] Combining the above examples, a preferred example of the present invention is obtained, which includes the following steps:
[0120] S10: Monitor the network in real time, collect traffic information and store it in the network traffic database, and determine network attack information through fine-grained monitoring of traffic;
[0121] S20: Match the network attack information in the policy database. If the match is successful, execute the corresponding matching policy in the policy database. If the match fails, go to step S30;
[0122] S30: Initializing the network environment, the strategy optimization model and the experience buffer;
[0123] S40: given state s t , the policy network is distributed according to the probability Select an action t And execute, the network environment is in the following state s t+1 and immediate reward r t Respond and get the given status s t The state value V(s) output by the next value network t ), the interaction data s t ,a t ,r t ,s t+1 , V(s t ) is stored in the experience buffer
[0124] S50: During iterative update, calculate the current strategy in a given state s t Take action a tThe probability ρ θ (a t ∣s t ), the previous strategy is in a given state s t Take action a t Probability And calculate the probability ρ θ (a t ∣s t ) and probability The ratio k t ; The advantage function is calculated using the generalized advantage estimation method Advantage function Indicates state s t The next line is a t Deviation from the mean;
[0125] S60: Based on the interaction data, ratio k t , advantage function Calculate the objective function, including the objective function based on KL penalty optimization and the objective function of cutting operation by limiting the ratio of new and old strategies, update the strategy network according to the objective function, and update the value network parameters alternately;
[0126] S70: Repeat steps S3-S60 until a predefined number of iterations is reached or convergence criteria are met, the policy network gradually converges to the optimal scheduling policy, the policy is used to handle network threats, and the policy is stored in the policy database;
[0127] S80: Calculate the ratio r1 of events in the policy database that have the same source IP address and target IP address as the network attack events to the total number of all processed network attack events; calculate the ratio r2 of network traffic in the network traffic database that has the same source IP address and target IP address as the network attack events to the total network traffic. If r1>r2*threshold, determine it as a high-risk path and execute a blocking operation.
[0128] The present invention introduces the mechanism of reinforcement learning and uses the PPO algorithm to construct a security policy generation model. In order to maintain a positive optimization policy matching model, the model automatically learns and finds the optimal security configuration and response strategy according to the current network environment to adapt to the ever-changing network environment and attack methods, and respond to potential security threats in a timely manner, thereby more effectively responding to new network threats, improving the efficiency and accuracy of security protection, and reducing the complexity of security policy configuration and operating costs.
[0129] Furthermore, by adding the KL divergence constraint on the amplitude of parameter changes to the objective function and constructing a new objective function clipping advantage function, the step size of the policy update is effectively limited, avoiding excessive changes in the policy parameters in each iteration, thereby ensuring the smoothness and stability of the policy update and simplifying the problem-solving method.
[0130] At the same time, by introducing a policy database, when a network attack occurs, the information of the network attack is matched in the policy database first. If the match is successful, it is executed directly; if the match fails, the PPO algorithm is used to generate, store and execute the policy. Under the experimental background of this invention, the model effectively ensures the network steady state when generating the policy, and can improve the analysis efficiency of network security incidents.
[0131] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the security policy matching method based on reinforcement learning formed by any one of the above examples or a combination of multiple examples. The processor may be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.
[0132] The present invention also provides a storage medium, which has the same inventive concept as the security policy matching method based on reinforcement learning formed by any one of the above examples or a combination of multiple examples, and stores computer instructions on the storage medium, which execute the steps of the security policy matching method based on reinforcement learning formed by any one of the above examples or a combination of multiple examples when the computer instructions are executed.
[0133] Based on this understanding, the technical solution of this embodiment, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0134] The present invention also provides a terminal, which has the same inventive concept as any example or a combination of multiple examples corresponding to the above-mentioned security policy matching method based on reinforcement learning, including a memory and a processor, wherein the memory stores computer instructions that can be run on the processor, and the processor executes the steps of the above-mentioned security policy matching method based on reinforcement learning when running the computer instructions. The processor can be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.
[0135] In one example, the terminal, i.e., the electronic device, is presented in the form of a general-purpose computing device, and the components of the electronic device may include but are not limited to: at least one processing unit (processor) mentioned above, at least one storage unit mentioned above, and a bus connecting different system components (including storage units and processing units).
[0136] The storage unit stores a program code, and the program code can be executed by the processing unit, so that the processing unit performs the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit can perform the above-mentioned security policy matching method based on reinforcement learning.
[0137] The storage unit may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 3201 and / or a cache memory unit, and may further include a read-only memory unit (ROM).
[0138] The storage unit may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0139] The bus may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0140] The electronic device may also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface. Furthermore, the electronic device may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter. The network adapter communicates with other modules of the electronic device through a bus. It should be understood that other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0141] Through the above description, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to this exemplary embodiment can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method of the exemplary embodiment of the present application.
[0142] The above specific implementation methods are detailed descriptions of the present invention. It cannot be determined that the specific implementation methods of the present invention are limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions and substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A security policy matching method based on reinforcement learning, characterized in that: The following steps are involved: Initialize the network environment, policy optimization model and experience buffer, the network environment includes the network state, the network state includes the network traffic, and the policy optimization model includes the policy network and the value network; Given a state s t , the policy network is distributed according to the probability Select an action t And execute, the network environment is in the following state s t+1 and immediate reward r t Respond and get the given status s t The state value V(s) output by the value network t ), the interaction data s t ,a t ,r t ,s t+1 , V(s t ) is stored in the experience buffer; During iterative updates, the current policy is calculated to be in a given state s t Take action a t The probability ρ θ (a t ∣s t ), the previous strategy is in a given state s t Take action a t Probability And calculate the probability ρ θ (a t ∣s t ) and probability The ratio k t ; The advantage function is calculated using the generalized advantage estimation method Advantage function Indicates state s t The next line is a t Deviation from the mean; According to the interaction data, the ratio k t , advantage function Calculate the objective function, including the objective function based on KL penalty optimization and the objective function of limiting the ratio of new and old strategies for shearing operations, and update the strategy network according to the objective function; Repeat the above steps until the termination condition is reached, the policy network gradually converges to the optimal scheduling strategy, and the optimal scheduling strategy is used to deal with network threats.
2. The security policy matching method based on reinforcement learning according to claim 1 is characterized in that: Before the network environment, the policy optimization model and the experience buffer are initialized, the following steps are also included: Monitor the network in real time, collect traffic information and store it in the network traffic database, and determine network attack information based on the traffic information, including source IP address, destination IP address, traffic rate, and threat type; The network attack information is matched in the policy database. If the match is successful, the corresponding matching policy in the policy database is executed. If the match fails, the network environment, policy optimization model and experience buffer are initialized.
3. The security policy matching method based on reinforcement learning according to claim 2 is characterized in that: After the optimal scheduling strategy is adopted to deal with the network threat, the method further includes: Storing the optimal scheduling strategy in a strategy database; Calculate the ratio r1 of events in the policy database that have the same source IP address and target IP address as network attack events to the total number of all processed network attack events; calculate the ratio r2 of network traffic in the network traffic database that has the same source IP address and target IP address as network attack events to the total network traffic. If r1>r2*threshold, it is determined to be a high-risk path and a blocking operation is performed.
4. The security policy matching method based on reinforcement learning according to claim 1 is characterized in that: The objective function L optimized based on the KL penalty term KL The calculation expression of (θ) is: in, Represents the expected value of the estimate; μ represents the penalty coefficient of KL divergence; Represents a given state s t Next, the old strategy The probability distribution of all possible actions; ρ θ (·∣s t ) represents a given state s t Under the current strategy ρ θ Probability distribution over all possible actions; represents the KL divergence between the new and old strategies.
5. The security policy matching method based on reinforcement learning according to claim 1 is characterized in that: After calculating the objective function based on the KL penalty term optimization, the method further includes calculating the expected value of the KL divergence between the new and old strategies, and adjusting the penalty coefficient according to the relationship between the expected value and the target expected threshold.
6. The security policy matching method based on reinforcement learning according to claim 5 is characterized in that: The calculation expression of the expected value d of the KL divergence between the new and old strategies is: in, represents the expected value of the estimate; represents the KL divergence between the new and old strategies.
7. The security policy matching method based on reinforcement learning according to claim 1 is characterized in that: The objective function L of limiting the ratio of new and old strategies to perform shearing operations CLIP The calculation expression of (θ) is: in, represents the expected value estimated at time t; σ represents the truncation hyperparameter; clip() represents the truncation function, limiting the proportion k t Greater than or equal to 1-σ and less than or equal to 1+σ.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the security policy matching method based on reinforcement learning described in any one of claims 1 to 7 are implemented.
9. A storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, the steps of the security policy matching method based on reinforcement learning described in any one of claims 1 to 7 are executed.
10. A terminal comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, characterized in that: When the processor runs the computer instructions, the processor performs the steps of the security policy matching method based on reinforcement learning described in any one of claims 1 to 7.
Citation Information
Patent Citations
CPPS optimal defense strategy game method for uncertain attacks
CN117439794A
Cited By
Network attack active defense strategy optimization method based on deep reinforcement learning
CN120934876A
Mobile robot navigation control method and system based on reinforcement learning and storage medium
CN121277190A