Intelligent network security risk monitoring and disposal system

The integration of reinforcement learning with real-time data and adaptive strategy updates in the network security system addresses the limitations of static rules, enhancing threat detection and response efficiency.

CN120320985APending Publication Date: 2025-07-15ANHUI SANSHI SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510402146.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional network security protection technologies are difficult to effectively deal with unknown or zero-day attacks, and are prone to false alarms or missed reports, affecting system reliability and network equipment operation.

Method used

Combining reinforcement learning technology with real-time data collection, reward evaluation and policy updates, an intelligent network security risk monitoring and disposal system is designed, and protection strategies are optimized through reinforcement learning models to reduce false alarms and missed reports.

Benefits of technology

It has improved the automation and intelligence level of network security protection, can effectively respond to new network security threats, reduce false alarms and missed reports, and ensure the normal operation of network equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120320985A_ABST
    Figure CN120320985A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent network security risk monitoring and handling system. The system comprises a data acquisition module, an action decision module, a network security monitoring and response module, a reward evaluation module, a reinforcement learning model updating module and an evaluation and tuning module. The data acquisition module collects network traffic in real time and constructs a network state vector; the action decision module makes a protection decision based on the reinforcement learning model Q; the reward evaluation module evaluates the protection effect and optimizes the strategy; the model updating module adjusts the Q value according to the reward and continuously optimizes the decision; and the evaluation and optimization module periodically updates the model, the reward function and the state characteristics to cope with novel threats. The method has the advantages that novel attacks can be automatically identified and dealt with, false alarms and missing alarms are reduced, and compared with a traditional rule-based protection system, the intelligent and automatic level of network security protection is improved by reinforcing the adaptive capacity of the learning model and continuously optimizing a protection strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and particularly to an intelligent network security risk monitoring and disposal system. Background Art

[0002] With the rapid development of the Internet and information technology, network security issues have become increasingly serious. Especially with the continuous change and evolution of network attack methods and means, traditional network security protection technologies and strategies have become difficult to effectively cope with complex and changing network security threats. Therefore, network security protection systems urgently need to be updated and iterated to address more complex and diverse security challenges.

[0003] Existing network security protection technologies usually rely on static rules and preset security policies, such as firewalls, intrusion detection systems (IDS), intrusion prevention systems (IPS), etc. These systems identify and block potential attacks through predefined rules, but this static protection strategy has some limitations:

[0004] First, traditional security protection systems often rely on signature matching or rule bases and are unable to effectively cope with unknown or zero-day attacks. Due to the continuous evolution and change of network attack means, static rules and strategies are difficult to cover all potential threats.

[0005] Second, traditional network security protection systems are prone to false positives or false negatives when detecting unknown attacks. For example, normal network traffic may be misjudged as an attack, while new attacks may be missed. This not only affects the reliability of the system but may also have a negative impact on network devices and business operations.

[0006] To solve these problems, in recent years, researchers have begun to explore network security protection methods based on artificial intelligence (AI) and machine learning. Among them, reinforcement learning (RL), as a self-learning intelligent algorithm, has become a research hotspot for intelligent network security protection systems. Through interaction with the network environment, the reinforcement learning model can automatically optimize the protection strategy to effectively cope with new network security threats. Summary of the Invention

[0007] The purpose of the present invention is to design an intelligent network security risk monitoring and disposal system by combining reinforcement learning technology with real-time data collection, reward evaluation, and policy update, which can effectively cope with new network security threats, optimize the protection strategy, reduce false positives and false negatives, and improve the automation and intelligence level of the protection system.

[0008] For this reason, the technical solution adopted by the present invention is as follows:

[0009] An intelligent network security risk monitoring and disposal system includes the following modules,

[0010] M1, a data acquisition module, which is used to collect network device security information in real time for constructing a network state vector;

[0011] M2, an action decision-making module, which selects an action a based on the current state vector s t through a reinforcement learning model Q, t that is, makes corresponding protection actions according to the network security state at time step t; the action is selected from a preset action space A, and the action space A contains all protection actions, and the reinforcement learning model Q is a neural network;

[0012] M3, a network security monitoring and response module, where the network device executes corresponding protection decisions according to the selected action a t and feeds back a new network state vector s t+1 to the reward evaluation module and the reinforcement learning model;

[0013] M4, a reward evaluation module, which calculates an immediate reward r through a reward function according to the current state vector s t , the corresponding selected action a t in combination with the network state vector s fed back by the network device t+1 , and the reward function is used to evaluate the effect of the current action; t

[0014] M5, a reinforcement learning model update module, which updates the reinforcement learning model according to the actions taken by the network device and the received rewards, so as to learn the optimal network security protection decision;

[0015] M6, an evaluation and tuning module, which is used to regularly evaluate the performance of the reinforcement learning model, update the model parameters, the reward function and the feature selection of the state vector to ensure that the system can effectively respond to new network security attacks.

[0016] Furthermore, the state vector in the data acquisition module is represented as s t , which contains the features of the network device security information at time step t;

[0017] The state vector where represents the i-th type of network device security information at time step t, and n represents the number of types of network device security information.

[0018] Furthermore, the action selection process in the action decision-making module is defined in the following form:

[0019]

[0020] ​Among them, Q represents the reinforcement learning model, and θ represents the current neural network parameters of Q. At the current state vector s t Select the action a t The Q value of.

[0021] Furthermore, the reward function is expressed as follows:

[0022] Including positive reward +P1, negative reward -P2, and no reward 0. When the network device feedbacks that the network security attack is effectively blocked, set r t = +P1; when the network device feedbacks that the network security attack is not effectively blocked, or the network device does not suffer from a network security attack, but the action a t Affects the normal operation of the network device, set r t = -P2; when the network device does not suffer from a network security attack, and the action a t Does not affect the normal operation of the network device, set r t = 0.

[0023] Furthermore, the update process of the reinforcement learning model is as follows:

[0024] When the reward evaluation module calculates the reward r t After that, based on the network state vector s at the current time step t t , reward r t And the network device security information s at time step t+1 t+1 , to update the Q value of the reinforcement learning model.

[0025] Furthermore, the Q value update process is expressed as follows:

[0026]

[0027] Among them, α is the learning rate, γ is the discount factor, which determines the influence degree of future rewards. Is the maximum Q value of the action selected by the state vector at the next time step t+1, θ - Represents the target neural network parameters, θ - Is updated after a preset integer multiple of time steps, that is, the update operation is θ - = θ, the target neural network parameters θ - Are equal to θ when the reinforcement learning model Q is initialized.

[0028] Furthermore, the update steps of the reinforcement learning model of the evaluation and tuning module are as follows:

[0029] First, calculate the loss function L(θ), let The loss function formula is

[0030] L(θ) = [y t -Q(s t , a t ; θ)] 2

[0031] Then calculate the gradient of the loss function and expand it using the chain rule. The formula is

[0032]

[0033] Finally, update the model parameters. The formula is

[0034]

[0035] where β is the learning rate.

[0036] Furthermore, the update of the reward function refers to adjusting the magnitudes of the positive reward +P1 and the negative reward -P2.

[0037] Compared with the prior art, the advantages of the present invention are as follows:

[0038] By combining a reinforcement learning model with real-time data collection and analysis, the present invention enables network security protection to automatically make optimal decisions according to the real-time network status. The reinforcement learning model can continuously adjust its protection strategy through interaction with the environment to cope with the ever-changing network security threats. Compared with traditional protection systems based on rules or static policies, the present invention can make decisions more efficiently and accurately in the face of new attacks and complex network security environments, thereby greatly improving the automation and intelligence levels of network security protection.

[0039] By continuously updating the Q-value and the reward function of the reinforcement learning model, the present invention can continuously optimize the protection strategy and reduce the probabilities of false positives and false negatives. When facing network attacks, the reinforcement learning model can be dynamically adjusted through the reward mechanism, avoiding misjudgment of normal business traffic and thus reducing interference with the performance of network devices. In this way, the system can achieve more accurate threat identification and response, improving the protection effect.

[0040] By combining real-time data collection, reward evaluation, and policy update, the present invention enhances the adaptability of the network security protection system to new network security threats. As the means of network attacks evolve, the system can continuously learn new attack patterns through the reinforcement learning model and automatically adjust the protection strategy. This adaptability enables the system to timely identify and effectively respond to unknown attack types that have not been covered by the preset rules. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 It is the system working flow chart of the present invention;

[0043] Figure 2 It is the schematic diagram of Q - value selection of the present invention;

[0044] Figure 3 It is the schematic diagram of Q - value update of the present invention. Detailed implementation manners

[0045] To achieve the above - mentioned objectives, the present invention is realized through the following technical solutions. The present invention provides an intelligent network security risk monitoring and disposal system. Combining with the attached Figure 1 , this system includes:

[0046] M1, Data acquisition module

[0047] This module is used to collect network device security information in real - time for constructing a network state vector.

[0048] The network device security information includes: network traffic data, system log data, security event data, network device status data, etc.

[0049] The network state vector is denoted as s t , which contains the features of the network device security information at time step t; the state vector where, represents the i - th network device security information at time step t, and n represents the types of network device security information. For example: represents the traffic rate, represents the CPU load, represents the number of accessed IPs.

[0050] M2, Action decision - making module

[0051] Based on the current state vector s t select an action a through the reinforcement learning model Q t , that is, make corresponding protection decisions according to the network security state at time step t; the action is selected from the action space A preset by the system, and the action space A contains all protection decisions, such as: continuously monitoring a specific type of network traffic, blocking the source IP, updating firewall rules, etc.; the reinforcement learning model Q is a neural network;

[0052] As Figure 2 shown, the action selection process is defined in the following form:

[0053]

[0054] where Q represents the reinforcement learning model, and θ represents the current neural network parameters of Q. When selecting action a t at the current state vector s t of the Q value.

[0055] M3, Network Security Monitoring and Response Module

[0056] The network device executes the corresponding protection decision according to the selected action a t and feeds back the new network state vector s t+1 to the reward evaluation module and the reinforcement learning model.

[0057] M4, Reward Evaluation Module

[0058] According to the current state vector s t and the corresponding selected action a t combining the network state vector s t +1 fed back by the network device, calculate the immediate reward r t through the reward function. The reward function is used to evaluate the effect of the current action. The design of the reward takes into account the network security protection effect. The reward function is expressed as follows:

[0059] There are positive rewards +P1, negative rewards -P2, and no rewards 0. When the network device feeds back that a network security attack has been effectively blocked, set r t = +P1; when the network device feeds back a false positive or missed report of a network security attack, set r t = -P2; when the network device has no feedback, set r t = 0.

[0060] M5, Reinforcement Learning Model Update Module

[0061] This module updates the reinforcement learning model according to the actions taken by the network device and the received rewards, so as to learn the optimal network security protection decision, as Figure 3 shown.

[0062] After the reward evaluation module calculates the reward r t , based on the network state vector s t at the current time step t, the reward r t and the network device security information s t+1 at time step t+1, to update the Q - value of the reinforcement learning model. The Q - value update process is expressed as follows:

[0063]

[0064] where α is the learning rate and γ is the discount factor, which determines the impact degree of future rewards. is the maximum Q - value of the action selected by the state vector at the next time step t + 1, and θ - represents the target neural network parameters, and θ - is updated after a preset integer - multiple number of time steps, that is, the update operation is θ - = θ.

[0065] The target neural network parameters θ - are equal to θ when the reinforcement learning model Q is initialized.

[0066] M6. Evaluation and Tuning Module

[0067] Regularly evaluate the performance of the reinforcement learning model, and update the model parameters, reward function, and feature selection of the state vector to ensure that the system can effectively respond to new cybersecurity threats.

[0068] The update steps of the reinforcement learning model are as follows.

[0069] First, calculate the loss function L(θ), let The formula of the loss function is

[0070] L(θ)=[y t -Q(s t , a t ; θ)] 2

[0071] Then, calculate the gradient of the loss function, expand it using the chain rule, and the formula is

[0072]

[0073] Finally, perform the model parameter update, and the formula is

[0074]

[0075] where β is the learning rate.

[0076] The update of the reward function means adjusting the magnitudes of the positive reward +P1 and the negative reward -P2 by combining past actions and the feedback of network devices.

[0077] The working processes of the system in three different scenarios are given below.

[0078] Example 1: A DDoS (Distributed Denial of Service) attack is a type of attack that occupies the target network resources through a large number of malicious requests, and its goal is to make the network service unable to operate normally.

[0079] Deploy the intelligent network security detection and disposal system proposed by the present invention in a network device that needs to defend against DDoS attacks. The defense steps of this system are as follows:

[0080] First, collect network traffic data, including: source IP, traffic rate (such as number of requests per second), traffic packet size, abnormal traffic, etc. And convert the collected data into the network state vector, which may include traffic rate, number of source IPs, CPU load, etc.

[0081] Secondly, the reinforcement learning model Q selects the best protection action according to the current network state vector. The action space includes: blocking IP, traffic limiting, enabling a traffic filter to discard abnormal requests, etc.

[0082] Thirdly, the network device executes corresponding protection measures according to the selected protection action. For example, limit the traffic from the attack source through a firewall or a load balancer; enable a firewall or a DDoS protection system to filter illegal requests; dynamically adjust the traffic limit to prevent network overload.

[0083] After executing the protection action, the network device returns a new network state vector.

[0084] Then, the reward function calculates the reward according to the effect of the protection action. For example, if the protection action is effective, the traffic rate, the number of source IPs, and the CPU load will decrease to a certain extent;

[0085] For example, the reward function can be designed as follows,

[0086] Including, positive reward: if the DDoS attack is effectively blocked and the network device returns to normal, the reward is positive; negative reward: if normal traffic is wrongly blocked, the reward is negative; zero reward: if the network device has no positive or negative impact, the reward is 0. For example, the value of the reward function can be: r t = 10 (positive reward), r t = -5 (negative reward), r t = 0 (zero reward).

[0087] Next, the reinforcement learning model calculates and updates the Q value according to the new network state vector, the current network state vector s t and the corresponding action a t and the reward r t Calculate and update the Q value.

[0088] When the reward is a negative reward, it means that the selected action at Aggravate or fail to resolve the consequences brought about by the current DDoS attack suffered by the network device, or the current network device is not suffering from a DDoS attack, and the selected action a t Affects the normal operation of the network device, then Q(s t ,a t ; θ) value should be appropriately lowered. When the reward is a positive reward, it indicates that the selected action a t Can resolve the DDoS attack suffered by the current network device, then Q(s t ,a t ; θ) value should be appropriately raised; when the reward is a zero reward, it indicates that the current network device is not suffering from a DDoS attack, and the selected action will not affect the normal operation of the network device, then Q(s t ,a t ; θ) value remains unchanged.

[0089] By updating the Q value, the reinforcement learning model can select more suitable actions.

[0090] Finally, regularly evaluate the performance of the reinforcement learning model during the evaluation period, and adjust the reward function and reinforcement learning model parameters according to the evaluation results to optimize the system response strategy.

[0091] Through the above steps, the reinforcement learning model gradually learns to recognize the traffic characteristics of DDoS attacks and automatically selects the optimal protection strategy. And it can detect attacks in real time, take corresponding protection measures, such as blocking the source IP, enabling traffic filtering, etc., and adjust the protection strategy according to the feedback to reduce false positives and false negatives.

[0092] Example 2: A brute force attack is a common type of attack in the network. The attacker attempts to log in to the target network device through a large number of password attempts. Deploying this system on a network device that needs to defend against brute force attacks, the steps for the system to monitor and handle brute force attacks are as follows:

[0093] First, collect login-related information in the network device system log to construct a state vector. For example, the state vector includes: the number of failed login attempts. For example, an IP has 50 failed login attempts within 10 minutes; the login attempt success rate. For example, the success rate <5%; the login time distribution. For example, the number of login attempts increases abnormally late at night; the geographical location distribution of the login source IP, etc.

[0094] Second, the reinforcement learning model selects protection actions based on the state vector to prevent the attack from expanding; for example, the protection actions can include: no action, continue to monitor; block the suspicious IP; extend the login attempt interval, etc.

[0095] Again, the protected network device performs the protection actions selected by the model. Meanwhile, the system monitors the new state vector of the network device in real time.

[0096] Then, the reward function calculates the reward according to the effect of the protection action. For example, if the protection action is effective, the number of failed login attempts will decrease, the success rate of login attempts will increase, the login time distribution will conform to the statistical law, the geographical location of the source IP of the login will conform to the statistical law, and so on.

[0097] For example, the reward function can be designed as follows

[0098] Including: positive reward: the attack stops and no legitimate user is wrongly blocked, the reward is positive; negative reward: a normal user is wrongly blocked or the attacker bypasses the protection, the reward is negative; zero reward: there is no positive or negative impact on the network device, the reward is 0. For example, the value of the reward function can be: r t = 10 (positive reward), r t = -8 (negative reward), r t = 0 (zero reward).

[0099] Next, the reinforcement learning model calculates and updates the Q value according to the new network state vector, the current network state vector s t and the corresponding action a t and the reward r t calculates and updates the Q value.

[0100] When the reward is a negative reward, it means that the selected action a t aggravates or cannot solve the consequences brought by the brute - force cracking attack suffered by the current network device, or the current network device is not suffering from a brute - force cracking attack, and the selected action a t affects the normal operation of the network device, then the value of Q(s t , a t ; θ) should be appropriately lowered. When the reward is a positive reward, it means that the selected action a t can solve the brute - force cracking attack suffered by the current network device, then the value of Q(s t , a t ; θ) should be appropriately raised; when the reward is a zero reward, it means that the current network device is not suffering from a brute - force cracking attack, and the selected action does not affect the normal operation of the network device, then the value of Q(s t , a t ; θ) remains unchanged.

[0101] Finally, the performance of the reinforcement learning model is evaluated regularly, and the reward function and the parameters of the reinforcement learning model are adjusted according to the evaluation results to optimize the system response strategy.

[0102] Through the above steps, the reinforcement learning model gradually learns to recognize the characteristics of brute-force attacks and automatically selects the optimal protection strategy.

[0103] Example 3: In an internal network, this system can be used to detect and respond to malicious intranet behaviors such as internal employees abusing their permissions, data leakage, or unauthorized access. The goal is to make the optimal protection decision based on the real-time collected network state vector through the reinforcement learning model to ensure effective control of the threats to the intranet.

[0104] Deploy this system on network devices that need to defend against malicious intranet behavior attacks. The steps for this system to monitor and handle malicious intranet behaviors are as follows:

[0105] First, collect the security information of the network device to construct a state vector. For example, the state vector includes: the number of failed logins, the location of the device, the upload rate, the download rate, the frequency of accessing sensitive files, the time of accessing sensitive files, and access beyond permissions, etc.

[0106] Second, select a protection action through the reinforcement learning model to avoid the leakage of important resources. For example, the protection actions can include: taking no measures; restricting the network traffic of the user and reducing its access permissions; blocking all network connections of the device; forcing the user to re-authenticate and forcing a password change, etc.

[0107] Third, the protected network device executes the protection action selected by the model. At the same time, this system monitors the new state vector of the network device in real time.

[0108] Then, the reward function calculates the reward according to the effect of the protection action. For example, if the protection action is effective, the number of failed logins will decrease, the location of the device meets the requirements, the upload rate and the download rate are within a certain threshold, the frequency and time of accessing sensitive files decrease respectively, and access beyond permissions is blocked, etc.

[0109] For example, the reward function can be designed as follows,

[0110] Including, positive reward: the malicious behavior stops and no legal users are affected, the reward is positive; negative reward: a normal user is wrongly blocked or the malicious behavior is not prevented, the reward is negative; zero reward: there is no positive or negative impact on the network device, the reward is 0. For example, the value of the reward function can be: r t = 10 (positive reward), r t = -5 (negative reward), r t = 0 (zero reward).

[0111] Next, the reinforcement learning model is based on the new network state vector, the current network state vector s t and the corresponding action a t and the reward rt Calculate and update the Q value.

[0112] When the reward is a negative reward, it indicates that the selected action a t aggravates or fails to solve the consequences brought about by the internal network malicious behavior suffered by the current network device, or the current network device has not suffered a malicious behavior attack, and the selected action a t affects the normal operation of the network device, then the Q(s t , a t ; θ) value should be appropriately lowered. When the reward is a positive reward, it indicates that the selected action a t can solve the internal network malicious behavior suffered by the current network device, then Q(s t , a t ; θ) value should be appropriately raised; when the reward is a zero reward, it indicates that the current network device has not suffered a malicious behavior attack, and the selected action will not affect the normal operation of the network device, then Q(s t , a t ; θ) value remains unchanged.

[0113] Finally, periodically evaluate the performance of the reinforcement learning model, and adjust the reward function and the parameters of the reinforcement learning model according to the evaluation results to optimize the system response strategy.

[0114] Through the combination of the reinforcement learning model and real-time data collection, the present invention can automatically identify and respond to new and complex network security threats, including DDoS attacks, brute force cracking attacks, internal network malicious behaviors, etc. The reinforcement learning model continuously adjusts the protection strategy, enabling the system to make optimal decisions based on the real-time network status, effectively improving the intelligence and automation level of network security protection.

[0115] The present invention dynamically adjusts the protection strategy through the reward evaluation mechanism, reducing the occurrence of false positives and false negatives. The system can intelligently determine and select the most appropriate protection measures to ensure the normal operation and security of the network device. After the system is deployed, the present invention will continuously learn new attack patterns, optimize the protection strategy according to real-time feedback, and through the periodic evaluation and optimization module, the system can timely update the model parameters, reward function, and state feature selection to cope with new attack means and changing network environments, improving the adaptability and protection effect of the system.

[0116] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.

Claims

1. An intelligent network security risk monitoring and disposal system, characterized in that, It includes the following modules: M1, a data collection module, which is used to collect network device security information in real time for constructing a network state vector; M2. Action decision-making module, which selects an action a based on the current state vector s t through the reinforcement learning model Q t , that is, makes corresponding protection actions according to the network security state at time step t; the action is selected from the preset action space A, the action space A contains all protection actions, and the reinforcement learning model Q is a neural network; M3, Network Security Monitoring and Response Module, the network device performs corresponding protection decisions according to the selected action a t and feeds back the new network state vector s t+1 to the reward evaluation module and the reinforcement learning model; M4, Reward Evaluation Module, which calculates the immediate reward r based on the current state vector s t , the corresponding selected action a t , combined with the network state vector s fed back by the network device t+1 , and calculates the immediate reward r through the reward function t . The reward function is used to evaluate the effect of the current action; M5, a reinforcement learning model update module, which updates the reinforcement learning model according to the actions taken by the network device and the received rewards, so as to learn the optimal network security protection decision; M6, an evaluation and optimization module, which is used to regularly evaluate the performance of the reinforcement learning model, update the model parameters, the reward function and the feature selection of the state vector to ensure that the system can effectively respond to new network security attacks.

2. The intelligent network security risk monitoring and handling system according to claim 1, characterized in that The state vector in the data acquisition module is represented as s t , which contains the features of the network device security information at time step t; The state vector wherein represents the security information of the i-th network device at the t-th time step, and n represents the number of types of network device security information.

3. The intelligent network security risk monitoring and disposal system according to claim 1, wherein, The action selection process in the action decision module is defined in the following form: Among them, Q represents the reinforcement learning model, and θ represents the current neural network parameters of Q. At the current state vector s t select the action a t of the Q value.

4. The intelligent network security risk monitoring and disposal system according to claim 1, characterized in that, The reward function is expressed as follows: Including positive reward +P1, negative reward -P2, and no reward 0. When the network device feedbacks that a network security attack has been effectively blocked, set r t = +P1; when the network device feedbacks that a network security attack has not been effectively blocked, or the network device has not suffered a network security attack, but action a t affects the normal operation of the network device, set r t = -P2; when the network device has not suffered a network security attack, and action a t does not affect the normal operation of the network device, set r t = 0.

5. The intelligent network security risk monitoring and handling system according to claim 1, characterized in that The reinforcement learning model update process is as follows: When the reward evaluation module calculates the reward r t After that, based on the network state vector s at the current time step t t , Reward t and the network device security information s at time step t+1 t+1 , to update the Q value of the reinforcement learning model.

6. The intelligent network security risk monitoring and handling system according to claim 5, characterized in that, The Q-value update process is expressed as follows: where α is the learning rate, γ is the discount factor that determines the impact of future rewards, is the maximum Q value of the action selected by the state vector at the next time step t+1, θ - represents the target neural network parameters, θ - is updated after an integer multiple of the preset delay time steps, that is, the update operation is θ - = θ, and the target neural network parameters θ - are equal to θ when the reinforcement learning model Q is initialized.

7. The intelligent network security risk monitoring and handling system according to claim 1, characterized in that The reinforcement learning model update steps of the evaluation and optimization module are: First, calculate the loss function \(L(\theta)\) and let The formula for the loss function is L(θ) = [y t -Q(s t , a t ; θ)] 2 Then calculate the loss function gradient and expand it using the chain rule. The formula is: Finally, perform model parameter update. The formula is: where β is the learning rate.

8. The intelligent network security risk monitoring and disposal system according to claim 7, characterized in that, The reward function update means adjusting the magnitudes of the positive reward +P1 and the negative reward -P2.

Citation Information

Patent Citations

  • Network defense method and device, equipment and medium

    CN119628877A

  • Intelligent medical data security guarantee system and method

    CN119646873A

  • Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning

    US20240080270A1

Cited By

  • Intelligent internet asset network security risk detection system

    CN121418182A

  • An intelligent internet asset network security risk detection system

    CN121418182B