Power network terminal cooperative defense method and system based on hierarchical reinforcement learning
Through the collaborative defense method of power network terminals based on hierarchical reinforcement learning, real-time analysis and dynamic adjustment of trust assessment and isolation strategies are carried out, which solves the problems of insufficient dynamic adaptability and limited attack detection capabilities of power network terminal defense technology, and achieves rapid response and efficient protection to complex attack scenarios.
Patent Information
- Application Number
- CN202510654896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
Existing power network terminal defense technologies have problems such as insufficient dynamic adaptability, limited attack detection capabilities, rigid trust assessment, and low isolation decision-making efficiency. They are unable to cope with complex and changing attack scenarios, especially new or unknown threats, with slow response speeds and high false alarm and missed alarm rates.
A collaborative defense method for power network terminals based on hierarchical reinforcement learning is adopted. Traffic data of multiple protocols is captured and analyzed in real time through high-performance network traffic collection equipment, and classified using machine learning algorithms. The dynamic trust evaluation and isolation strategies of upper-layer and lower-layer intelligent agents are combined to achieve real-time adaptation and dynamic adjustment.
It improves the real-time adaptability to complex attack scenarios, reduces false alarm and missed alarm rates, optimizes the utilization efficiency of network resources, and ensures the security and reliability of the power network.
Smart Images

Figure CN120675738A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power network terminal defense technology, and specifically to a power network terminal collaborative defense method and system based on hierarchical reinforcement learning. Background Art
[0002] Currently, the defense of power network terminals mainly relies on security protection systems based on rules and behavioral analysis. These systems are generally composed of four core modules: data acquisition module, attack identification module, trust assessment module, and isolation decision module.
[0003] The data collection module collects behavioral data, network traffic, system logs, and security event information from terminal devices through network traffic monitoring devices, log analysis tools, and intrusion detection systems (IDS). Its purpose is to identify any abnormal activities or potential threats.
[0004] The attack identification module analyzes collected data using predefined rules or traditional machine learning algorithms to identify specific types of network attack behaviors, such as port scans and denial of service (DoS) attacks. However, this approach relies heavily on known patterns and features and lacks the ability to effectively respond to unknown threats or zero-day attacks.
[0005] The trust assessment module uses static feature analysis to determine the trust level of each device based on its historical behavior and current state. While this method is effective in some cases, its static nature makes it difficult to adapt to changing behavior patterns or new attack methods, resulting in high false positive and false negative rates.
[0006] Based on the trust assessment results and identified attacks, the isolation decision module takes appropriate isolation measures, such as restricting access through firewall settings or VLAN technology. However, existing isolation strategies are often reactive. Once a device is isolated, restoring access typically requires manual intervention, which not only increases management burdens but also potentially impacts business continuity.
[0007] Although the above methods have enhanced network protection capabilities to a certain extent, they still have obvious limitations: due to their reliance on fixed rules and models, existing systems have difficulty coping with complex and changing attack scenarios, especially when facing new or unknown threats. When the network environment changes or new types of attack behaviors emerge, the system cannot automatically adjust its policies and usually requires manual rule updates, which has a slow response speed. When dealing with edge cases, high false alarm rates and missed alarm rates are prone to occur, reducing the actual protection effect. Existing isolation measures are mostly passive and lack automated and intelligent dynamic adjustment capabilities. This can easily lead to mis-isolation or a complex recovery process after isolation. This is particularly evident in environments such as power networks where there are many types of terminal devices and complex behavior patterns. Summary of the Invention
[0008] In view of the above-mentioned problems, the present invention is proposed.
[0009] Therefore, the technical problems solved by the present invention are: the existing power network terminal defense technology has the problems of insufficient dynamic adaptability, limited attack detection capability, rigid trust assessment and low isolation decision-making efficiency, and how to achieve real-time adaptation and dynamic adjustment to complex attack scenarios by introducing hierarchical reinforcement learning technology, how to enable the system to autonomously adjust the trust rating according to changes in terminal behavior characteristics to reduce false alarm and missed alarm rates, and how to construct a collaborative isolation decision-making mechanism to deeply combine trust rating and isolation strategy to achieve fast and accurate threat isolation and recovery.
[0010] To solve the above technical problems, the present invention provides the following technical solutions: a power network terminal collaborative defense method based on hierarchical reinforcement learning, comprising: a high-performance network traffic collection device uses the port mirroring function of a switch to capture and analyze traffic data packets covering multiple protocols in the power network in real time;
[0011] Machine learning algorithms classify this traffic data, accurately identifying and distinguishing normal traffic from unsafe traffic, and storing the processed multi-dimensional feature data in a structured format;
[0012] Its upper-level intelligent agent receives behavioral data and historical security events in structured storage for dynamic trust assessment, and autonomously adjusts the trust rating based on changes in terminal behavioral characteristics; the lower-level intelligent agent receives the trust rating results and real-time traffic data from the upper-level intelligent agent and executes targeted isolation strategies.
[0013] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the collection of network traffic data covers multiple communication protocols, including but not limited to Modbus, IEC 61850 and DNP3, and is stored in a central data server.
[0014] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the machine learning algorithm is used to distinguish normal traffic from unsafe traffic and is trained based on expert experience and historical data.
[0015] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the state space of the upper-level intelligent agent is represented by the terminal behavior feature vector, and the action space is three rating adjustment actions: improving, maintaining or reducing the trust rating.
[0016] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the action space of the lower-level intelligent agent includes allowing normal terminal access, partially restricting access rights, completely isolating terminal traffic and restoring the isolation state.
[0017] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the reward function design is adjusted according to the terminal behavior compliance and the overall protection effect of the system to optimize the trust rating and isolation decision.
[0018] As a preferred solution of the power network terminal collaborative defense method based on hierarchical reinforcement learning described in the present invention, the method also includes a feedback mechanism, which realizes self-learning and optimization by continuously monitoring changes in terminal behavior and network status, thereby quickly responding to new attacks, reducing the probability of false alarms and missed alarms, and optimizing the allocation efficiency of network resources.
[0019] Another object of the present invention is to provide a computing platform load balancing system based on a particle swarm genetic algorithm, which can introduce hierarchical reinforcement learning technology through one of the schemes to dynamically evaluate terminal trust ratings and implement precise isolation strategies, thereby solving the problems of current power network terminal defense technologies, such as insufficient dynamic adaptability, limited attack detection capabilities, rigid trust assessment, and inefficient isolation decision-making.
[0020] As a preferred embodiment of the particle swarm genetic algorithm-based computing platform load balancing system described in the present invention, a high-performance network traffic collection device acquires real-time network traffic data from key nodes in the power network (such as the control center, edge devices, and terminal device communication paths). Multiple communication protocols are supported to ensure coverage of all important data exchange locations. The captured data is stored in real-time on a central data server in pcap format and replicated to a high-reliability storage device through a scheduled backup mechanism to ensure data security and integrity.
[0021] The machine learning module uses machine learning algorithms to automatically classify the extracted traffic data, distinguish normal traffic from unsafe traffic, reduce false alarm and missed alarm rates, and use the Wireshark tool to deeply analyze the raw traffic data, extract basic network characteristics, behavioral characteristics, and statistical characteristics, and provide detailed information for subsequent analysis.
[0022] The upper-level agent dynamically adjusts trust ratings based on endpoint behavioral data and historical security incidents, categorizing them into high, medium, and low levels corresponding to varying access rights. A deep Q-network is used to model state-action values. Through experience collection and storage, Q-network updates, and policy refinement, it dynamically adapts to network state changes, improving the accuracy of rating adjustments. Trust rating rules are adjusted based on newly detected attack patterns, ensuring the system's rapid adaptability to unknown threats.
[0023] Based on the trust ratings of the upper-layer agents and real-time network traffic data, the lower-layer agents select and execute targeted isolation policies, such as allowing access, partially restricting access, or completely isolating terminal traffic. Illegal actions are dynamically disabled to ensure appropriate protective measures are taken based on different trust ratings, reducing false isolations and the complexity of recovery. Reinforcement learning algorithms such as Double DQN and PPO are used to continuously learn the optimal isolation decision-making strategy, optimizing the accuracy and efficiency of isolation policies.
[0024] The system's overall functionality enables real-time monitoring of all terminal behavior and network status within the power network, enabling timely identification of potential security threats. Through continuous feedback mechanisms, it enables self-learning and optimization, enabling rapid response to new attacks, reducing the probability of false positives and missed alerts, and optimizing the allocation of network resources.
[0025] A computer device includes a high-performance network traffic collection device, the high-performance network traffic collection device has a computer program, and the network traffic collection device executes the computer program to implement a step of a power network terminal collaborative defense method based on hierarchical reinforcement learning.
[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a power network terminal collaborative defense method based on hierarchical reinforcement learning.
[0027] Beneficial Effects of the Invention: By incorporating high-performance network traffic collection, deep traffic analysis, and a dual-layer reinforcement learning intelligent protection mechanism, this invention achieves significant benefits. First, the system enables real-time and comprehensive monitoring of power network traffic data, ensuring efficient collection and analysis of network communications at key nodes. This provides the system with detailed network status information, enabling timely detection of potential security threats. Dynamic updates ensure the timeliness and accuracy of data sets.
[0028] Secondly, through traffic classification using machine learning algorithms, the system can accurately distinguish between normal and unsafe traffic, reducing false positives and missed alerts. This process, combined with the unique communication protocols and behavioral characteristics of power systems, can accurately identify unauthorized access, abnormal communications, and potential attacks, thereby improving the accuracy and reliability of network security protection.
[0029] Crucially, the system employs a two-layer reinforcement learning architecture. The upper-layer agent evaluates trust ratings based on terminal behavior and historical security incidents, dynamically adjusting terminal access permissions. The lower-layer agent implements isolation policies based on real-time traffic data, accurately identifying and isolating potential threat terminals. This layered protection mechanism enables the system to dynamically optimize protection measures based on changes in network status and attack patterns, enhancing security defense capabilities while maximizing the utilization of network resources.
[0030] Through a continuous feedback mechanism, the system can self-learn and optimize, adapting to the ever-changing network environment and new attack methods, ensuring the security of the power network while improving overall operational efficiency. This flexible adaptive management and precise resource allocation provide a solid guarantee for the security, reliability, and efficiency of the power network. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 This is an overall flow chart of the power network terminal collaborative defense method based on hierarchical reinforcement learning provided in the first embodiment of the present invention.
[0033] Figure 2 This is a diagram of the network security defense system architecture of the power network terminal collaborative defense method based on hierarchical reinforcement learning provided in the second embodiment of the present invention.
[0034] Figure 3 A dynamic security strategy diagram of a power network terminal collaborative defense method based on hierarchical reinforcement learning is provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0035] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0036] Example 1, reference Figure 1 , which is an embodiment of the present invention, provides a power network terminal collaborative defense method based on hierarchical reinforcement learning, including:
[0037] S101: High-performance network traffic collection equipment uses the port mirroring function of the switch to capture and analyze traffic data packets covering multiple protocols in the power network in real time.
[0038] Furthermore, data collection equipment utilizes network traffic mirroring technology, using the switch's port mirroring (SPAN) function to capture all data packets flowing on the network. To ensure the integrity of collected traffic data, collection rules cover multiple communication protocols, including Modbus, IEC 61850, and DNP3, as well as different types of network traffic, such as normal business traffic and unsafe traffic. The captured network traffic data is stored in real time on a central data server in pcap format. The server uses a scheduled backup mechanism to copy the traffic data to high-reliability storage devices to ensure data security and integrity, while preventing accidental loss.
[0039] It should be noted that the Wireshark traffic parsing tool was used to perform in-depth analysis and feature extraction on the collected raw traffic data. This feature extraction includes basic network features (such as source IP address, destination IP address, source port, destination port, and protocol type), behavioral features (such as connection duration, number of packets, and transmission rate), and statistical features (such as average traffic size and peak traffic). For power system-specific protocols, protocol field features were also extracted, such as the function code of a Modbus request or the specific data field content of an IEC 61850 message.
[0040] S102: The machine learning algorithm classifies the traffic data, accurately identifies and distinguishes normal traffic from unsafe traffic, and stores the processed multi-dimensional feature data in a structured format.
[0041] Furthermore, machine learning algorithms are used to automatically classify the extracted traffic data. The classification process divides the traffic into two categories: normal traffic and unsafe traffic. Unsafe traffic characteristics include abnormally high-frequency communications, illegal instruction payloads, and unauthorized access behaviors. The training of the classification model combines expert experience and historical data to ensure the accuracy of the classification results. After the classification is completed, the system generates a dataset containing multi-dimensional features and annotation labels. The dataset is stored in a structured format (such as CSV or JSON), and the content includes the basic feature values of network traffic (protocol type, packet size, etc.), time series features (reflecting the pattern of traffic changes over time), and traffic category labels (normal or unsafe).
[0042] It should be noted that data collection and processing are ongoing, and the system is able to update the dataset in real time. Through this dynamic update mechanism, the system incorporates the latest characteristics of the power network communication environment and potential threat behaviors into the dataset, ensuring that the generated dataset can support subsequent network security analysis and optimization of defense decisions.
[0043] S103: Its upper-level intelligent agent receives behavioral data and historical security events in structured storage for dynamic trust assessment, and autonomously adjusts the trust rating based on changes in terminal behavioral characteristics.
[0044] Furthermore, the upper-layer agent performs upper-layer tasks, dynamically analyzing the terminal's behavioral data and historical security events. The upper-layer agent's state space is represented by a feature vector of the terminal's behavior, including network communication frequency, access patterns, abnormal behavior records, and historical security event statistics. The action space consists of three rating adjustment actions: increasing the trust rating, maintaining the trust rating, or decreasing the trust rating. The reward function is designed based on the compliance of the terminal's behavior and the overall protection effectiveness of the system. Behavior that meets security standards receives positive rewards, while abnormal behavior results in negative rewards.
[0045] It should be noted that the trust rating output by the upper-level task is divided into three levels: high, medium, and low, corresponding to different terminal access permissions. Trust ratings are optimized through a reinforcement learning strategy. During long-term training, the system can dynamically adapt to changes in network status, improving the accuracy of rating adjustments. The trust rating results serve as the environmental state input for the lower-level tasks, guiding firewall isolation decisions.
[0046] The specific steps are as follows:
[0047] The upper-level intelligent agent performs upper-level tasks and dynamically analyzes the terminal's behavioral data and historical security events.
[0048] (1) The state space of the upper-level intelligent agent is represented by the characteristic vector of the terminal behavior, including network communication frequency, abnormal behavior records, historical security event statistics and entropy characteristics. State vector Include:
[0049] Terminal communication frequency: f t =[average number of requests / minute, peak number of requests / minute]
[0050] Abnormal behavior statistics: a t =[abnormal instruction count, illegal protocol ratio]
[0051] Historical security incidents: t =[Number of alarms triggered in the past T hours]
[0052] (2) The action space consists of three rating adjustment actions: trust rating improvement, trust rating unchanged, or trust rating reduction. The reward function is designed based on the compliance of terminal behavior and the overall protection effect of the system. Behaviors that meet security standards receive positive rewards, while abnormal behaviors result in negative rewards.
[0053] 1) Discrete action set A t={+1,0,-1} corresponds to trust rating increase / maintain / decrease respectively
[0054] (3) The trust rating output by the upper-level task is divided into three levels: high, medium, and low, corresponding to different terminal access rights. The trust rating is modeled by the state-action value through the deep Q network:
[0055]
[0056] Where θ is the upper layer Q network parameter.
[0057] The training process is:
[0058] Experience collection, execution of upper-level decision-making, and the action strategy selection formula are:
[0059]
[0060] The initial value of ∈ is set to 0.5 and updated according to the linear decay strategy:
[0061]
[0062] 1) Experience storage
[0063] The tuple (s t ,a t ,r t ,s t+1 ) is stored in the playback buffer D
[0064] 2) Q Network Update
[0065] Sampling batches from D {(s i ,a i ,r i ,s i ')}
[0066] Calculate the target Q value:
[0067]
[0068] where θ m- is the target network parameter, synchronized every 100 steps
[0069] Minimize the loss function:
[0070]
[0071] 3) Strategy Improvement
[0072] Generate the lower-level state input through the trust rating mapping function:
[0073]
[0074] During long-term training, the system can dynamically adapt to changes in network status, improving the accuracy of rating adjustments. The trust rating results serve as the environmental state input for lower-level tasks to guide firewall isolation decisions.
[0075] S104: The lower-level intelligent agent receives the trust rating results and real-time traffic data from the upper-level intelligent agent and executes targeted isolation strategies.
[0076] Furthermore, lower-level agents perform lower-level tasks, analyzing trust ratings and real-time network traffic data to implement targeted isolation strategies. The state space includes the trust ratings output by the upper layer, real-time traffic characteristics (such as protocol type, packet size, and flow direction), and security risk status. The action space includes options such as allowing normal terminal access, partially restricting terminal access rights (such as blocking access to highly sensitive resources), completely isolating terminal traffic, and restoring the isolation state.
[0077] When the trust rating is high, the lower-level task maintains the normal access rights of the terminal through default actions; when the trust rating is medium, the lower-level agent selects a strategy to partially restrict the terminal's access rights based on traffic characteristics and network risk assessment, such as restricting access to critical resources or specific ports; when the trust rating is low, the lower-level agent selects an isolation strategy through reinforcement learning decision-making, including completely blocking the terminal's network connection or isolating the terminal to a specified VLAN.
[0078] The upper-level tasks and lower-level tasks collaborate through trust rating results. The upper-level agent dynamically adjusts the trust rating and provides clear security assessment results; the lower-level agent executes specific isolation strategies based on the rating results and real-time environmental data.
[0079] The lower-level intelligent agents perform lower-level tasks, analyze trust rating results and real-time network traffic data, and implement targeted isolation strategies.
[0080] (1) The state space includes the trust rating output by the upper layer and the characteristic values of real-time traffic (such as protocol type, packet size, and traffic direction). The state vector S t ∈R k Include:
[0081] Real-time traffic characteristics: p t =[protocol type code, packet size standard deviation, TCP flag entropy]
[0082] Trust rating: r∈{0,1,2}(low / medium / high)
[0083] (2) The action space includes allowing normal terminal access, partially restricting terminal access rights (such as blocking access to highly sensitive resources), completely isolating terminal traffic, and restoring the isolation state.
[0084] Discrete action set A w ={0,1,2,3} corresponds to allow / partially restrict / completely isolate / restore
[0085] Action mask mechanism: Dynamically disable illegal actions based on trust rating (e.g., disabling isolation actions when the rating is high)
[0086] When the trust rating is high, the lower-level task maintains the normal access rights of the terminal through default actions; when the trust rating is medium, the lower-level agent selects a strategy to partially restrict the terminal's access rights based on traffic characteristics and network risk assessment, such as restricting access to critical resources or specific ports; when the trust rating is low, the lower-level agent selects an isolation strategy through reinforcement learning decision-making, including completely blocking the terminal's network connection or isolating the terminal to a specified VLAN.
[0087] (3) Reward function design
[0088] R w (s,a)=δ×S(s,a)+ε×E(s,a)-ζ×P(s,a)
[0089] Where S(s,a) is the security benefit (the number of successfully blocked attacks), E(s,a) is the resource efficiency (effective traffic throughput retention rate), P(s,a) is the policy execution cost (CPU / memory consumption increase), δ=0.5, ε=0.3, and ζ=0.2 are dynamic adjustment coefficients.
[0090] The upper-level tasks and lower-level tasks collaborate through trust rating results. The upper-level agent dynamically adjusts the trust rating and provides clear security assessment results; the lower-level agent executes specific isolation strategies based on the rating results and real-time environmental data.
[0091] It should be noted that the training and execution of hierarchical reinforcement learning consists of the following steps:
[0092] Initialize the reinforcement learning parameters of the upper-layer agent (such as the state-value function or policy network) and define the reward function, using the normality of the terminal behavior as the positive reward indicator and the accuracy of detecting abnormal behavior as the negative reward indicator;
[0093] The upper-layer agent outputs the trust rating results through continuous learning and passes the results to the lower-layer agent as part of the input state;
[0094] The lower-level agent receives trust ratings and real-time traffic data, and learns the optimal isolation decision strategy based on reinforcement learning algorithms (such as Double DQN or PPO);
[0095] The lower-level intelligent agent selects and executes isolation actions, and provides real-time feedback on protection effects and changes in network status to the upper-level intelligent agent to further optimize the overall defense strategy.
[0096] S105: The hierarchical reinforcement learning agent dynamically adjusts the ratings of power network terminals and continuously manages the terminals.
[0097] Furthermore, the system continuously monitors the behavior and network status of all endpoints in the network in real time through intrusion detection and traffic analysis modules. The intrusion detection module uses a deep learning-based detection algorithm to analyze network traffic characteristics and identify abnormal behavior patterns, including but not limited to abnormal traffic spikes, unauthorized access attempts, and abnormal communication frequencies. The traffic analysis module extracts key traffic characteristics, such as packet size, transmission rate, and protocol distribution, providing input data for subsequent reinforcement learning.
[0098] Hierarchical reinforcement learning dynamically optimizes trust ratings and isolation policies based on real-time feedback data. This real-time feedback data includes detected anomalies, statistical characteristics of terminal behavior, and network resource usage. The upper-layer agent adjusts trust rating rules based on newly detected attack patterns, ensuring the system's rapid adaptability to unknown threats. The reward function is dynamically updated based on changes in terminal behavior and the effectiveness of defense decisions. Rewards increase when defenses are successful or the false alarm rate decreases; otherwise, they decrease. Lower-layer agents utilize real-time network environment changes and upper-layer trust ratings to optimize isolation policies. Based on reinforcement learning decisions, lower-layer agents adjust firewall rules or VLAN configurations to implement precise isolation policies for low-trust terminals, including restricting access to sensitive resources, reducing bandwidth allocations, or completely disconnecting from the network. When a terminal resumes normal behavior and is reassessed as safe based on trust ratings, the agent automatically lifts isolation and restores its network access.
[0099] It should be noted that the system achieves dynamic defense strategy optimization through continuous training and real-time adjustment of a reinforcement learning model. The results of each defensive action are recorded as feedback to the reinforcement learning model, continuously optimizing the agent's strategy network to ensure that the system maintains optimal defense status as attack methods and network environments change. Through this collaborative mechanism, the system achieves dynamic adaptive management of terminal behavior, enabling rapid response to new attacks, reducing the probability of false positives and missed negatives, and optimizing the efficient allocation of network resources. The reinforcement learning agent's real-time adjustment capabilities ensure that the system maximizes the efficient and safe use of network resources while ensuring security.
[0100] Example 2, reference Figure 1-Figure 3 , which is the second embodiment of the present invention, is different from the previous embodiment in that it specifically describes the application of the power network terminal collaborative defense method based on hierarchical reinforcement learning in the smart grid.
[0101] The system is deployed at key nodes in the smart grid, such as substations, distribution network control centers, and power plants. Using high-performance network traffic collection equipment, it captures real-time traffic data across multiple communication protocols, including but not limited to industrial control protocols Modbus, IEC 61850, and DNP3. All collected traffic data is stored on a central data server and simultaneously transmitted to a machine learning module for analysis.
[0102] Machine learning algorithms (such as deep learning-based classification models) analyze traffic data in real time, extracting terminal behavior characteristics (such as communication frequency, protocol field anomalies, and command patterns) to distinguish between normal and unsafe traffic. For example, if a terminal sends high-frequency Modbus write commands that deviate significantly from historical operating patterns, it is identified as abnormal traffic.
[0103] The upper-layer agent receives the terminal behavior feature vector as state space input and dynamically adjusts the trust rating of the terminal through reinforcement learning strategy. Its action space contains three rating adjustment actions:
[0104] Triggered when endpoint behavior meets security benchmarks and remains compliant;
[0105] When the behavior is not significantly abnormal but requires continued observation, the current rating is maintained;
[0106] Executed when a potential attack is detected (such as illegal instruction payload or abnormal communication pattern).
[0107] Based on the rating results of the upper-level intelligent agent, the lower-level intelligent agent selects a response strategy from the predefined action space: open full permissions to high-trust rated terminals; restrict specific functions for medium-risk terminals (such as prohibiting remote control of key equipment); implement network isolation for low-trust rated or high-risk terminals to block their communication with the core system; after the terminal behavior returns to compliance, lift the isolation based on the real-time evaluation results.
[0108] The system continuously optimizes defense strategies through a reward function, which is designed based on the following core metrics: rewards for correctly classified traffic and timely blocking of attacks; and penalties for business interruptions or missed attack reports caused by incorrect isolation.
[0109] Example 3, reference Figure 2-Figure 3 , which is the third embodiment of the present invention, provides a system for a collaborative defense method for power network terminals based on hierarchical reinforcement learning, including a high-performance network traffic collection device, a machine learning module, an upper-layer intelligent agent and a lower-layer intelligent agent.
[0110] The system is deployed on core network nodes in multi-tenant data centers and supports virtualized environments (such as virtual private clouds (VPCs) or network slicing technologies). Its core components include:
[0111] High-performance network traffic collection devices are deployed in data center core switches and edge access points to capture cross-tenant network traffic in real time, covering protocols including Modbus (industrial control protocol), HTTP / HTTPS (web services), and tenant-defined protocols;
[0112] The machine learning module is based on a deep learning traffic classification model and supports parallel processing and feature extraction of multi-tenant traffic;
[0113] The upper-layer agent cluster assigns an independent upper-layer agent to each tenant, which is responsible for trust rating management across tenant terminals;
[0114] The lower-level agent cluster deploys a lower-level agent for each physical or virtual network node to implement fine-grained access control and resource allocation strategies.
[0115] The machine learning module performs the following analysis on multi-tenant traffic: identifying traffic types (such as Modbus commands, HTTP requests) and protocol field anomalies (such as illegal Modbus function codes); establishing a baseline behavior model for each tenant, including normal communication frequency, resource usage patterns, and terminal access rights; and detecting the cross-tenant propagation of abnormal traffic (such as a malicious terminal launching an attack from Tenant A to Tenant B).
[0116] Each tenant's upper-level agent receives the following status inputs: including the anomaly score of a single terminal and the tenant's overall behavior deviation; such as CPU / bandwidth usage and the number of virtual machine instances.
[0117] Based on the reinforcement learning strategy, the upper-level intelligent agent performs the following actions on the terminal: triggered when the terminal behavior complies with the tenant's security policy and resource usage is compliant; maintains the current rating for terminals that have no significant abnormalities in behavior but require continuous monitoring; and executes when suspicious behavior is detected (such as tenant B's terminal continuously attempting to access tenant A's private API) or resource abuse (such as a precursor to a DDoS attack).
[0118] Based on the rating results of the upper-level intelligent agent, the lower-level intelligent agent implements the following strategies: open full permissions to high-trust terminals (such as tenant C's database access request); restrict specific resources for medium-risk terminals (such as isolating tenant D's virtual machine to a dedicated network slice); implement network isolation for low-trust or high-risk terminals (such as blocking tenant E's abnormal Modbus write instructions).
[0119] Dynamically allocate bandwidth based on terminal trust ratings (e.g., prioritizing real-time business traffic of high-trust tenants); migrate virtual machines of low-trust tenants to isolated resource pools to avoid affecting high-priority tenants; and automatically allocate additional resources when a surge in tenant traffic (e.g., a legitimate business peak) is detected to avoid misjudging it as an attack.
[0120] The system optimizes defense strategies through the following reward functions: rewards for successfully blocking attacks (such as isolating malicious terminals) and penalizes business interruptions caused by missed reports or incorrect isolation; rewards for optimized resource allocation (such as efficient bandwidth utilization) and penalizes resource waste caused by excessive isolation; and ensures balanced quality of service (QoS) among multiple tenants by penalizing resource contention among tenants.
[0121] The feedback mechanism continuously optimizes the system through the following methods: real-time updates to machine learning models to adapt to changes in tenant traffic patterns (such as legitimate traffic characteristics after the launch of new services); when new attacks are detected (such as protocol fuzzing attacks against multiple tenants), the upper-level intelligent agent cluster collaboratively adjusts the global defense strategy; regularly analyzes false positive / missed negative cases and optimizes reinforcement learning parameters to improve decision-making accuracy.
[0122] In a multi-tenant data center scenario, this embodiment achieves the following innovative advantages: through the collaborative decision-making of hierarchical intelligent agents, the abnormal behavior of low-trust tenants (such as isolating virtual machines that initiate DDoS attacks) can be quickly blocked without affecting the business of high-trust tenants; the dynamic resource allocation mechanism improves resource utilization by about 30%, while reducing business interruption events caused by incorrect isolation. The system can identify and block cross-tenant attack chains (such as attackers infiltrating tenant B's core system through the vulnerability of tenant A), reducing the overall risk of the data center; by sharing threat intelligence (such as new Modbus attack characteristics), it accelerates the update of the entire network's defense strategy. The feedback mechanism enables the system to quickly adapt to changes in tenant business models (such as the access of new industrial Internet of Things terminals); supports horizontal expansion (such as automatically deploying independent intelligent agents when adding new tenants) to meet the elastic needs of cloud data centers.
[0123] The system also features a built-in feedback loop mechanism that dynamically updates machine learning model parameters and reinforcement learning strategies by continuously monitoring changes in terminal behavior and network status (such as evolving attack patterns or updates to legitimate traffic patterns). For example, when a new attack (such as a protocol fuzzing attack targeting the IEC 61850 protocol) emerges, the system can quickly adjust feature extraction rules and retrain through reinforcement learning to adapt to the new threat.
[0124] In the smart grid scenario, the system achieves the following advantages: through the collaborative decision-making of hierarchical intelligent agents, abnormal traffic can be blocked within milliseconds, such as isolating malicious terminals that attempt to tamper with the circuit breaker status; isolation strategies are dynamically adjusted according to the terminal trust rating to avoid business interruptions caused by excessive blocking (such as only restricting the permissions of non-critical devices rather than completely disconnecting the network); the feedback mechanism significantly reduces the false alarm rate (such as reducing the false isolation caused by fluctuations in legitimate traffic) and the missed alarm rate (such as identifying new attack variants).
[0125] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0126] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0127] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0128] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. It should be noted that the above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications should be encompassed by the claims of the present invention.
[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A collaborative defense method for power network terminals based on hierarchical reinforcement learning, characterized in that: include: High-performance network traffic collection equipment uses the port mirroring function of the switch to capture and analyze traffic data packets covering multiple protocols in the power network in real time; Machine learning algorithms classify this traffic data, accurately identifying and distinguishing normal traffic from unsafe traffic, and storing the processed multi-dimensional feature data in a structured format; Its upper-layer intelligent agent receives behavioral data and historical security events in structured storage to conduct dynamic trust assessment and autonomously adjust the trust rating based on changes in terminal behavioral characteristics; The lower-level agent receives the trust rating results and real-time traffic data from the upper-level agent and executes targeted isolation strategies.
2. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 1 is characterized in that: The collection of network traffic data covers multiple communication protocols.
3. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 2 is characterized in that: The machine learning algorithm is used to distinguish normal traffic from unsafe traffic and is trained based on historical data.
4. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 3 is characterized in that: The state space of the upper-layer agent is represented by the terminal behavior feature vector, and the action space is three rating adjustment actions: improving, keeping unchanged or reducing the trust rating.
5. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 4 is characterized in that: The action space of the lower-layer intelligent agent includes allowing normal terminal access, partially restricting access rights, completely isolating terminal traffic, and restoring the isolation state.
6. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 5, characterized in that: The reward function design is adjusted according to the terminal behavior compliance and the overall protection effect of the system to optimize the trust rating and isolation decision.
7. The power network terminal collaborative defense method based on hierarchical reinforcement learning according to claim 6, characterized in that: The method also includes a feedback mechanism that enables self-learning and optimization by continuously monitoring changes in terminal behavior and network status, thereby quickly responding to new attacks, reducing the probability of false positives and missed positives, and optimizing the allocation efficiency of network resources.
8. A system using the power network terminal collaborative defense method based on hierarchical reinforcement learning as described in any one of claims 1 to 7, characterized in that: It includes high-performance network traffic collection equipment, machine learning modules, upper-layer intelligent agents and lower-layer intelligent agents.
9. A computer device comprising a high-performance network traffic collection device, wherein the high-performance network traffic collection device has a computer program, characterized in that: The network traffic collection device executing the computer program is a step of implementing the power network terminal collaborative defense method based on hierarchical reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the power network terminal collaborative defense method based on hierarchical reinforcement learning according to any one of claims 1 to 7 are implemented.