Digital power grid network security attack and defense game situation awareness and efficiency evaluation system, method, equipment and medium

The situational awareness and performance evaluation system, which utilizes multi-dimensional data processing and multi-level threshold judgment, solves the problems of detection accuracy and inaccurate decision-making in digital power grid network security protection systems under complex environments, achieving efficient attack identification and reduced false alarm rates.

CN121585441APending Publication Date: 2026-02-27INFORMATION CENT OF YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511787309.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing digital power grid network security protection systems suffer from low detection accuracy in complex environments, are unable to flexibly respond to new types of attacks, and have inaccurate decision-making and a high false alarm rate.

Method used

The system employs a field snapshot generation module, an emergency level determination module, a candidate action generation module, a similar segment retrieval module, a rule-based playback module, a red line gatekeeping module, a selection logic module, an execution rhythm control module, and a consistent evaluation module. Combined with multi-dimensional data processing and multi-level threshold judgment, it performs dynamic situational awareness and effectiveness evaluation.

Benefits of technology

It improves the accuracy of attack identification, reduces the false alarm rate, enhances the system's adaptability and decision reliability, and supports cross-scenario applicability and continuous optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585441A_ABST
    Figure CN121585441A_ABST
Patent Text Reader

Abstract

The invention discloses a digital power grid network security attack and defense game situation awareness and efficiency evaluation method, system and device and a medium, and belongs to the field of electrical engineering. Comprising a field snapshot generation module, an emergency level judgment module, a candidate action generation module, a similar fragment retrieval module, a regularization playback module, a red line goalkeeping module, an accepting and rejecting logic module, an execution rhythm control module, a same-caliber evaluation module and an evidence retention and calibration module. According to the system provided by the invention, through historical data playback verification, the motion effect prediction accuracy is improved, and the false alarm rate is reduced. According to the method, the candidate action effect is simulated on the historical fragment, the distribution expressions of the income side and the cost side are output, the false difference is eliminated by adopting the completely consistent statistical caliber and time grid, and the evaluation is quantified by the median and upper bound quantiles, so that the accuracy of the effect evaluation is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electrical engineering, in particular to a digital power grid network security attack and defense game situation awareness and performance evaluation system, method, device and medium. BACKGROUND

[0002] Currently, the digital power grid network security protection system mainly relies on the following two schemes: static defense based on rule matching: many power grid security systems use pre-set security rules for threat detection. The common practice is to identify abnormal traffic based on threshold values and rules set by expert experience. These methods usually rely on fixed thresholds set by humans, which are difficult to adapt to the operating characteristics of different sites and different time periods, have low detection accuracy, and cannot flexibly respond to new attacks. Dynamic detection based on machine learning: In some smart grid security systems, machine learning algorithms are used to analyze network traffic patterns. By training historical data models, the system can identify abnormal behavior. However, this approach is usually limited to single-dimensional detection and cannot comprehensively consider multi-dimensional features such as operating state, channel quality and job state, making it difficult to make precise disposal decisions. However, traditional power grid security technologies rely mainly on simple rule matching or single model detection, which cannot maintain high precision identification in complex power grid environments. For example, special working conditions during night maintenance periods, concurrent anomalies at multiple sites, etc. may cause the false positive rate to rise, affecting the reliability of the system. Therefore, how to achieve accurate identification, effective disposal and continuous optimization of digital power grid network security threats is a technical problem that needs to be solved by those skilled in the art.

[0003] In view of the above problems, the present application provides a solution. SUMMARY

[0004] In view of the above problems, the present application is proposed.

[0005] Therefore, the present application aims to solve the problems of low detection accuracy, inability to flexibly respond to new attacks, inaccurate disposal decisions and high false positive rate of existing digital power grid network security protection systems in complex power grid environments due to reliance on rule matching or single model detection.

[0006] To solve the above technical problems, the present application provides the following technical scheme: a digital power grid network security attack and defense game situation awareness and performance evaluation system, comprising: a field snapshot generation module, an emergency level determination module, a candidate action generation module, a similar segment retrieval module, a rule-based playback module, a red line gatekeeping module, a selection logic module, an execution rhythm control module, a same caliber evaluation module and an evidence retention and calibration module; The field snapshot generation module is responsible for real-time acquisition and processing of multi-dimensional data to generate field snapshots in a unified format; An emergency level determination module determines an emergency level of the current network security situation according to the running intensity and the channel pressure information in the field snapshot by using a multi-level threshold judgment strategy; A candidate action generation module generates a candidate action set corresponding to the gear according to the emergency level output by the emergency level determination module; A similar segment retrieval module retrieves a historical segment similar to the current field snapshot from historical data; A regularized playback module simulates the execution effect of the candidate action on the historical segment and outputs a quantitative benefit and cost distribution; A red line guard module performs a one-vote veto comparison on the upper bound of the cost distribution output by the regularized playback module with three hard constraints, cuts the strategy space, and makes all potential disposal actions not touch the safety red line; A selection logic module selects an action in the feasible set filtered out by the red line guard module by using a double-account book dominance relationship; An execution rhythm control module performs gradual execution and dynamic adjustment of the disposal action by coupling rhythm control and field perception in two stages; A same caliber evaluation module evaluates the measured effect online during the execution process with the same caliber, and uses the same set of measurement standards among the statistical caliber, the execution caliber and the evaluation caliber; An evidence preservation and calibration module solidifies the shortest causal chain evidence and drives parameter small-step calibration.

[0007] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and performance evaluation system, the field snapshot generation module comprises a data acquisition unit, a baseline calculation unit, a relative quantity calculation unit and a snapshot packaging unit; The data acquisition unit acquires data from the substation safety equipment, network equipment and business system; The baseline calculation unit calculates baseline values of abnormal connection intensity, link delay and data loss rate; The relative quantity calculation unit calculates abnormal surge ratios and channel deviation quantities according to the baseline values and packages them into snapshots by the snapshot packaging unit.

[0008] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and performance evaluation system, the emergency level determination module comprises a threshold learning unit and a determination execution unit; The threshold learning unit adjusts the abnormal surge ratio threshold and the channel deviation quantity tolerance interval according to the relationship between the abnormal surge ratio, the channel deviation quantity and the artificially labeled emergency level in the historical attack and defense data; The determination execution unit uses a decision tree or a rule engine to calculate the abnormal surge ratio and the channel deviation quantity in parallel after receiving the field snapshot and outputs the corresponding emergency level label. The emergency level determination module predefines an abnormal burst ratio threshold and a channel deviation tolerance interval, compares the abnormal burst ratio corresponding to the abnormal connection strength with the threshold, and compares the channel deviation corresponding to the link delay and the data loss rate with the tolerance interval. The high emergency level, the medium emergency level and the low emergency level are divided according to a preset combination relationship, and the emergency level is taken as an input mark of an open action file of the candidate action generation module.

[0009] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and efficiency evaluation system, the candidate action generation module includes four types of action templates, i.e., a speed limit observation, a precise plugging, a strong source sealing and a temporary segment sealing. Each action template includes parameter fields of a source address, a target address, a port, a protocol type and a validity duration. The candidate action generation module selects a corresponding action template set according to the emergency level output by the emergency level determination module, and instantiates a specific candidate action item in combination with the source identifier and the entry identifier in the field snapshot.

[0010] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and efficiency evaluation system, the similar segment retrieval module includes a discrete space constructed by a running intensity file, a channel pressure file and a work state file. The historical snapshots are stored by using a multi-dimensional index structure, the running intensity file, the channel pressure file and the work state file are indexed, and the index buckets are accessed according to the discrete coordinates during retrieval. The discrete coordinates are determined according to the abnormal burst ratio, the channel deviation and the maintenance state in the field snapshot, and the segments consistent with the discrete coordinates are retrieved in the historical snapshots. When the matching segments obtained by relaxing the file are lower than a preset lower limit, the remaining dimensions are relaxed, but the work state file is not relaxed when the maintenance is true, and a segment set with a time distance from the field snapshot not exceeding a preset threshold is selected from the candidate segments as a reference set.

[0011] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and efficiency evaluation system, the regularized playback module includes a regularized playback module that replays the historical traffic on the reference set obtained by the similar segment retrieval module, injects parameters into each candidate action, and records the abnormal connection strength, the link delay, the data loss rate and the maintenance session state. The median value and the high quantile value of the candidate action benefits and costs are calculated according to the statistical caliber, the high quantile value of the cost is filtered according to a preset red line constraint, and a set of actionable actions is output.

[0012] As a preferred scheme of the digital power grid network security attack and defense game situation awareness and performance evaluation system, wherein: the execution rhythm control module includes, the execution rhythm control module selects the target action when the target action is selected by the selection logic module, and a snapshot is generated by calling the live snapshot generation module in the observation window. The deviation between the measured value of the index calculated by the same caliber evaluation module and the predicted value of the rule-based playback module determines whether to upgrade to the target action or a stronger action. The evidence retention and calibration module encapsulates the process into an evidence block with a hash chain and adjusts the strategy parameters accordingly.

[0013] Another object of the present application is to provide a digital power grid network security attack and defense game situation awareness and performance evaluation method.

[0014] To solve the above technical problems, the present application provides the following technical scheme: a digital power grid network security attack and defense game situation awareness and performance evaluation method, comprising: Multi-dimensional data including operation, channel and job information are collected from security devices, network devices and business systems, and the data is preprocessed and historical baseline calculated to form a live snapshot including abnormal connection strength and channel delay, and relative amount of packet loss; Based on the relative amount of abnormal connection and channel deviation in the live snapshot, the hierarchical emergency level is obtained by comparing with the threshold value, and under the constraint of the emergency level, multi-grade security disposal actions are enabled, and execution parameter sets are generated for all actions; In the multi-dimensional space based on operation strength, channel pressure and job state, the historical fragments matching the live snapshot are retrieved, the historical fragments are simulated and played back, and the index changes are counted to generate the benefit and cost distribution of all actions; The upper bound of the cost distribution output by the rule-based playback module is compared by three hard constraints in a one-vote veto manner, the strategy space is cut, and all potential disposal actions do not touch the safety red line; According to the benefit and cost distribution and the safety hard constraint, the candidate actions are excluded, the target disposal action is selected according to the benefit and cost double account book rule, and the light action of the target disposal action is executed first and then switched to the target action; During the execution process, the measured effect is evaluated online with the same caliber, the same set of measurement standards is used among the statistical caliber, the execution caliber and the evaluation caliber, the shortest causal chain evidence is solidified, and the parameter small step calibration is driven.

[0015] The present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, characterized in that the processor executes the computer program to realize the steps of the digital power grid network security attack and defense game situation awareness and performance evaluation system.

[0016] The present invention provides a computer-readable storage medium storing a computer program thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the aforementioned digital power grid network security attack and defense game situation awareness and performance evaluation system.

[0017] The beneficial effects of this invention are as follows: This invention improves the accuracy of action effect prediction and reduces the false alarm rate by verifying historical data playback. This invention simulates candidate action effects on historical segments, outputs distribution representations of the benefit and cost sides, and eliminates spurious differences using completely consistent statistical methods and time grids. It quantifies the evaluation using median and upper bound quantile values, significantly improving the accuracy of effect evaluation.

[0018] This invention improves decision confidence and risk controllability by basing decisions on statistical distribution rather than single-point estimation. First, it requires the median return of the indicator to its normal range. Then, among the actions to achieve the target, it prioritizes those with lower costs. This objective evaluation method, based on distributional statistics rather than weighting factors, provides a verifiable standardized framework for multi-objective decision-making in complex systems, enhancing the reliability of decisions.

[0019] This invention supports data comparison across multiple sites and time periods, improving cross-scenario applicability. By deriving dimensionless or relative dimensional descriptions, cross-site and cross-time period comparability is built into the data structure. Furthermore, through incremental learning, rules gradually converge in the data, continuously optimizing system parameters and thus significantly improving the system's adaptability. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a module structure diagram of a digital power grid network security attack and defense game situation awareness and performance evaluation system provided in one embodiment of the present invention.

[0022] The modules include: 101. On-site snapshot generation module; 102. Emergency level determination module; 103. Candidate action generation module; 104. Similar segment retrieval module; 105. Rule-based playback module; 106. Red line gatekeeping module; 107. Selection logic module; 108. Execution rhythm control module; 109. Same-caliber evaluation module; and 110. Evidence retention and calibration module. Detailed Implementation

[0023] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0024] Example 1, referring to Figure 1 This is one embodiment of the present invention, which provides a digital power grid network security attack and defense game situation awareness and performance evaluation system, including: The module includes: a snapshot generation module 101, an emergency level determination module 102, a candidate action generation module 103, a similar fragment retrieval module 104, a rule-based playback module 105, a red line gatekeeping module 106, a selection logic module 107, an execution rhythm control module 108, a consistent evaluation module 109, and an evidence retention and calibration module 110. The on-site snapshot generation module 101 is responsible for real-time collection and processing of multi-dimensional data to generate on-site snapshots in a unified format. The emergency level determination module 102 uses a multi-level threshold judgment strategy to classify the current network security situation into emergency levels based on the operational intensity and channel pressure information in the on-site snapshot. The candidate action generation module 103 opens up the candidate action set for the corresponding level based on the emergency level output by the emergency level determination module. The similar fragment retrieval module 104 retrieves historical fragments similar to the current scene snapshot from historical data; The rule-based replay module 105 simulates the execution effect of candidate actions on historical segments and outputs quantified benefit and cost distributions. The red line gatekeeper module 106 uses three hard constraints to perform a veto-like comparison of the upper bound of the cost distribution output by the rule-based replay module, trimming the strategy space so that all potential actions do not touch the safety red line. The decision logic module 107 uses a dual-ledger control relationship to select actions from the feasible set filtered by the red-line gatekeeper module. The execution rhythm control module 108 embodies the coupling of rhythm control and on-site perception in a two-stage manner, enabling the gradual execution and dynamic adjustment of handling actions; The same-caliber evaluation module 109 evaluates the measured results online using the same caliber during the execution process, and uses the same set of measurement standards across statistical caliber, execution caliber, and evaluation caliber. The evidence retention and calibration module 110 solidifies the shortest causal chain evidence and drives parameter step-by-step calibration.

[0025] Example 2, refer toFigure 1 This is one embodiment of the present invention, which provides a digital power grid network security attack and defense game situation awareness and performance evaluation system, including: In this embodiment of the invention, the on-site snapshot generation module 101 is responsible for real-time acquisition and processing of multi-dimensional data to generate on-site snapshots in a unified format, including the following steps: Specifically, the on-site snapshot generation module 101 includes a data acquisition unit, which is used to collect raw data of the operating surface, the passage surface, and the work surface in real time from the safety equipment, network equipment, and business systems of each substation.

[0026] An alternative example is the security equipment in each substation, such as firewalls and intrusion detection systems; network equipment, such as routers and switches; and business systems, such as SCADA systems and EMS systems.

[0027] Operational plane data includes network connection logs, traffic statistics, and abnormal behavior alarms; channel plane data includes link latency, packet loss rate, and bandwidth utilization; and work plane data includes maintenance work orders, maintenance session status, and equipment maintenance records.

[0028] The data preprocessing unit is used to clean, denoise, format, and standardize the massive amounts of raw data collected to eliminate data redundancy, errors, and inconsistencies, thereby ensuring data quality.

[0029] For example, standardize the log format generated by different devices and synchronize and calibrate the timestamps.

[0030] The baseline calculation unit calculates baseline values ​​for the same time band based on historical data. For abnormal connection strength λ, it calculates its historical median value; for monitoring link tail latency r and monitoring data loss rate d, it calculates their historical high-quantity statistics, such as the 90th or 95th percentile, to reflect the fluctuation range under normal conditions. These baseline values ​​serve as a reference for subsequent anomaly judgment and relative quantity calculation.

[0031] The relative quantity calculation unit is used to derive dimensionless or relative dimensionless descriptions based on the currently collected real-time data and the calculated historical baseline. Abnormal surge ratio. : Defined as the ratio of the current abnormal connection strength λ to the historical median of the same time band.

[0032] in, The strength of abnormal connections within the current time window. This represents the historical median anomalous connectivity strength within the same time zone.

[0033] Channel deviation Δr: represents the offset of the current monitored link tail delay r relative to the historical median of the same time band. in, This refers to the tail latency of the monitoring link within the current time window. This represents the median latency at the end of the historical monitoring link within the same time band. Channel deviation Δd: represents the offset of the current monitoring data loss rate d relative to the historical median within the same time band.

[0034] in, This represents the data loss rate within the current time window. This represents the historical median data loss rate for the same time period. Calculating these relative values ​​ensures data comparability across sites and time periods, avoiding misjudgments caused by differences in absolute values.

[0035] The snapshot encapsulation unit is used to encapsulate all processed data, including original data values, relative values, timestamps, site identifiers, etc., into a unified format of on-site snapshot data structure. This snapshot serves as the input for all subsequent decision-making modules.

[0036] Furthermore, the data acquisition unit in the field snapshot generation module 101 is deployed in the agent programs of the safety equipment, network equipment and business systems of each substation, and collects data in real time through multiple protocols such as SNMP, NetFlow, Syslog and API; it collects connection logs and traffic statistics from firewalls, NetFlow data from routers, and business operation status from SCADA systems.

[0037] The data preprocessing unit uses streaming frameworks (such as Apache Flink or Kafka Streams) to clean, deduplicat, and convert the format of real-time data. For example, it unifies alarm logs generated by devices from different vendors into JSON format and aligns the timestamps.

[0038] The baseline calculation unit utilizes a sliding time window technique to perform statistical analysis on data from the past 24 hours, 7 days, and 30 days, calculating the historical median and high quantile values ​​of λ, r, and d at different time granularities (such as hours and minutes). The baseline model is updated periodically to adapt to seasonal and periodic changes in power grid operation.

[0039] When calculating ρ, Δr, and Δd, the relative quantity calculation unit introduces a smoothing factor to avoid excessive influence of instantaneous fluctuations on the calculation of relative quantities.

[0040] In an alternative example, the present invention may use an exponentially weighted moving average to smooth the baseline.

[0041] A snapshot encapsulation unit, which is used to generate on-site snapshots. It adopts serialization formats such as Protobuf or Avro to ensure efficient transmission and storage. The snapshot contains a unique snapshot ID, generation time, site ID, as well as detailed data of the operation surface, channel surface, and job surface, and the calculated relative quantities.

[0042] In the embodiment of the present invention, the emergency level determination module 102 uses a multi-level threshold judgment strategy to divide the current network security situation into emergency levels according to the operation intensity and channel pressure-bearing information in the on-site snapshot, including the following steps: According to the operation intensity (ρ) and channel pressure-bearing (Δr and Δd) information in the on-site snapshot, a multi-level threshold judgment strategy is adopted to divide the current network security situation into emergency levels. Specifically, it includes: Threshold setting: Set the sudden increase threshold values of ρ, including the significant sudden increase threshold T1 and the obvious sudden increase threshold T2, where T1>T2; For example, T1 = 5, indicating that the abnormal connection intensity is more than 5 times the historical median, and T2 = 2, indicating that the abnormal connection intensity is more than 2 times the historical median; At the same time, set the tolerance threshold values [L1, L2] of Δr and Δd, where L1<L2; For example, L1 = 10ms, L2 = 50ms; L1 = 0.5%, L2 = 2%. These threshold values and thresholds are not fixed constants, but can be dynamically adjusted according to site policies and historical learning results to adapt to different circadian rhythms and site working conditions.

[0043] Set the emergency level, high emergency level: If ρ≥T1 and (Δr≥L2 or Δd≥L2), that is, the abnormal connection intensity has increased significantly, and the channel delay or packet loss rate has approached or exceeded the high tolerance upper limit, it is determined as the high emergency level; Medium emergency level: If T2≤ρ<T1 and L1≤Δr<L2 and L1≤Δd<L2, that is, the abnormal connection intensity has increased significantly but not reached a significant level, and the channel pressure-bearing is within the medium tolerance range, it is determined as the medium emergency level; Low emergency level: In other cases, that is, both the abnormal connection intensity and the channel pressure-bearing are at a low level, it belongs to the low emergency level.

[0044] Furthermore, in the emergency level determination module 102, the threshold values T1, T2, L1, L2, etc. can be manually adjusted through the system configuration interface, or can be learned from historical attack and defense game data through machine learning algorithms, such as reinforcement learning or adaptive threshold algorithms; The decision-making logic can use a decision tree or a rule engine to implement the judgment of the emergency level; Upon receiving a snapshot of the scene, the system calculates ρ, Δr, and Δd in parallel, and performs conditional judgments based on preset threshold values ​​to quickly output the emergency level.

[0045] In this embodiment of the invention, the candidate action generation module 103, based on the emergency level output by the emergency level determination module, opens a set of candidate actions corresponding to the emergency level, including the following steps: Based on the emergency level output by the emergency level determination module 102, a set of candidate actions for the corresponding level is opened. These actions are semantically mutually exclusive and have clear boundaries to avoid ambiguity. The actions are defined as follows: Rate limiting observation: This only applies to rate shaping from this source to this ingress point, without changing the timing of other sources or ingress points; for example, limiting the traffic rate from a specific source IP to a specific target port below a certain threshold and continuously observing it. Precise blocking: Deny access to the entry point only from that source; for example, configure firewall rules to block specific source IPs from accessing specific target IPs and ports; Forced source blocking: Briefly isolate all access from this source within the site; for example, briefly block a specific source IP address across the entire substation network; Temporary blockade: Briefly disconnecting the remote channel of the target entry point; for example, temporarily isolating a specific service entry point of a substation through the network.

[0046] Action Opening Strategy: High Emergency Level: Activate all four action levels: speed-limited observation, precise blockade, strong source closure, and temporary section closure, to provide the most comprehensive response measures.

[0047] Medium emergency level: The first three levels of action are: speed limit observation, precise blockade, and strong sealing of the source, ensuring a certain level of response while avoiding excessive intervention.

[0048] Low emergency level: Only the first two actions are enabled: speed-limited observation and precise blockade, to conduct observation and initial handling with minimal impact.

[0049] Each action has clear semantic boundaries and execution parameters. For example, rate limiting observation requires setting the rate limit and observation duration, precise blocking requires specifying the source and entry address, strong source blocking requires specifying the source address and blocking duration, and temporary segment blocking requires specifying the entry address and network disconnection duration.

[0050] Furthermore, in the candidate action generation module 103, each action is defined with a detailed parameter template: the parameters for the rate limiting observation action include: source IP, target IP, target port, rate limiting bandwidth, and observation duration. The parameters for precise blocking actions include: source IP, target IP, target port, protocol type, and duration of effect. The set of actions that can be enabled at different urgency levels is configured by policy files, such as YAML or JSON format, and these policies can be customized by security experts according to the power grid security policy.

[0051] In this embodiment of the invention, the similar fragment retrieval module 104 retrieves historical fragments similar to the current scene snapshot from historical data, including the following steps: Used to retrieve historical fragments similar to the current scene snapshot from historical data, providing a reference for subsequent rule-based playback; Three-dimensional discrete space construction: The operating intensity level (ρ), channel pressure level (Δr and Δd), and operation status level (maintenance or not) are discretized and divided; ρ can be divided into normal, slight increase, obvious increase, and significant increase levels; Δr and Δd can be divided into normal, slight deviation, moderate deviation, and severe deviation levels according to their degree of deviation from the baseline; the working status is binary (maintenance / non-maintenance).

[0052] The hierarchical retrieval strategy includes: the first layer is to attempt a three-dimensional full match retrieval, which requires that the operating intensity level, channel pressure level, and work status level be completely consistent with the current site snapshot.

[0053] The second layer sets a lower limit if the number of historical fragments retrieved through a full match is insufficient. , =30, then relax one dimension to search for adjacent levels; if the operating intensity level is significantly increased, relax to search for obvious increase; the relaxation order can be preset, with priority given to relaxing the operating intensity level, followed by the channel pressure level.

[0054] The third layer is if it is still insufficient. If the operation status is true, other dimensions can be relaxed, but the operation status must not be relaxed when it is under maintenance. This is because the maintenance status has a huge impact on the operation of the power grid, and cross-domain mixing should not occur on key constraints to avoid misjudgment on the most critical constraints. The retrieved historical segments also need to meet the requirement of temporal proximity, and segments that are closer to the current snapshot should be selected first to ensure the timeliness of historical experience. The search results include segment index, matching degree and relevant statistics.

[0055] Furthermore, in the similar fragment retrieval module 104, multi-dimensional index structures such as Kd-trees or R-trees are used to index historical snapshots, thereby accelerating the retrieval process; The discretization of the operating intensity level, channel pressure level, and work status level can be achieved by equal-frequency binning or dynamic binning based on clustering algorithms, such as K-Means.

[0056] In this embodiment of the invention, the rule-based replay module 105 simulates the execution effect of candidate actions on historical segments and outputs a quantified distribution of benefits and costs, including the following steps: Simulated replay: For each candidate action, a mapping is applied to the trajectory that its semantics should reach, and simulated replay is performed on the historical fragment set provided by the similar fragment retrieval module 104; the rate limiting action only shapes the concurrency and rate trajectory from the source to the entry point, without changing the timing of other sources or entry points; the precise blocking only changes the acceptance decision when the source accesses the entry point, without affecting other combinations; the strong blocking and the blocking segment correspondingly expand the scope but do not exceed their respective boundaries.

[0057] Consistency of statistical methods: The playback time grid, statistical methods, and synchronization benchmark are completely consistent with the field snapshot to eliminate spurious differences caused by statistical method drift and ensure a high degree of comparability between simulation results and actual conditions; for example, if the field snapshot is a 5-minute average, the playback is also based on the 5-minute simulation results for statistical analysis.

[0058] Distribution Representation: For each action, two distributions are output in the reference set. On the payoff side, the distribution is represented by the median and upper bound quantile (e.g., the 90th quantile) of the decrease in anomalous connection strength λ. The specific formula can be: ;in, This is to simulate the decrease in abnormal connection strength λ after the action is performed.

[0059] The cost is expressed on the monitoring link tail delay *r* and monitoring data loss rate *d* as the median and upper bound quantile values, as well as the maintenance downtime probability as the median and upper bound quantile values. The specific formula can be: The cost is: in, To simulate the offset of r after the action is performed, To simulate the offset of d after the action is performed, This is to simulate the probability of disconnection during maintenance after the action is performed.

[0060] In this embodiment of the invention, the red line gatekeeper module 106 uses three hard constraints to perform a veto-review comparison of the upper bound of the cost distribution output by the rule-based replay module, pruning the strategy space to ensure that all potential actions do not touch the safety red line, including the following steps: Business continuity constraint: The upper bound of the offset of r must not exceed That is, after actual execution, the increase in link latency cannot exceed the preset maximum allowable latency threshold. For example, 50ms.

[0061] Data integrity constraint: The upper bound of the offset of d must not exceed [a certain value]. In other words, the increase in data loss rate cannot exceed the preset maximum allowable packet loss rate threshold. For example, 2%.

[0062] Maintenance safety constraints: The upper bound of the probability of maintenance outage must be zero, that is, no action can cause the ongoing maintenance session to be interrupted. This is the highest priority safety requirement for power grid operation. Any action that may cross the upper limit is eliminated in this module to ensure that subsequent selection processes do not bear risks that should not be assumed; these constraint values ​​can be set through the configuration interface and support site-specific configuration.

[0063] Furthermore, in the red line gatekeeper module 106, the hard constraint values ​​Rmax and Dmax are usually specified by national standards, industry specifications, or power grid dispatching departments. After receiving the cost distribution output by the rule-based replay module 105, the system immediately checks whether the upper bound of the cost distribution of each action exceeds Rmax, Dmax, or whether the maintenance disconnection probability is greater than 0. If any condition is not met, the action is removed from the feasible set.

[0064] In this embodiment of the invention, the selection logic module 107 uses the dual ledger control relationship to select actions from the feasible set filtered by the red line gatekeeper module, including the following steps: Comparison of revenue accounts: The decision-making logic module 107 first checks the revenue accounts, that is, the median revenue of each candidate action, which is the median value of the decrease in λ. If the median revenue of a certain action has brought λ back to the normal upper boundary defined based on the historical baseline, it is considered to have met the standard. Those that do not meet the standard will not participate in the subsequent comparison and will be directly eliminated.

[0065] Cost-benefit comparison: Within the set of actions that achieve the benefit target, a cost-benefit comparison is performed, following a clear priority order: Prioritize comparing the median r offset: Select the action with the smaller median r offset; If they are tied, compare the median value of the d-offset: select the action with the smaller median value of the d-offset; If they are still tied, choose the action with the lighter force: according to the above candidate actions, namely speed limit observation, precise blocking, strong blocking of the source, and temporary blockade, the force should be increased in turn, and the action with the lighter force should be selected first.

[0066] This process relies solely on comparable distribution statistics and clear priorities, ensuring the objectivity and transparency of decision-making.

[0067] Furthermore, in the decision-making logic module 107, when determining whether the profit target has been met, for each action, the median profit value is calculated to determine whether it is less than or equal to the normal upper boundary defined by the historical baseline. If the normal upper boundary is Then judge When performing cost sorting, for the set of actions that achieve the target benefit, a multi-level sorter is constructed; First, sort by the median value of the offset r in ascending order; if the median values ​​of the offset r are the same, then sort by the median value of the offset d in ascending order; if they are still the same, then sort by the intensity of the action in ascending order. Finally, select the action that appears first in the sorted list.

[0068] In this embodiment of the invention, the execution rhythm control module 108 embodies the coupling of rhythm control and on-site perception in a two-stage manner, performing progressive execution and dynamic adjustment of the handling action, including the following steps: Phase 1: First, execute the lighter version of the target action. If the target action selected by the decision logic module 107 is precise blocking, then the first phase will execute speed limit observation.

[0069] Observation Window: Enter the observation window and set the observation window duration. For example, for 5 minutes, continuously monitor the network security situation and the effectiveness of actions within the window.

[0070] Trigger condition judgment: When the window expires, determine whether to upgrade to the target action or to the next level based on the pre-agreed trigger conditions; The trigger conditions are defined around the failure to meet the profit target or the occurrence of antagonistic changes in the behavior pattern within the window.

[0071] An alternative example is when the expected return is not met, such as when λ does not return to its normal range, or when the behavior pattern exhibits adversarial changes within the window, such as when the attack source switches ports or protocols.

[0072] Red line guard before upgrade: Before any upgrade, the system will automatically repeat the judgment of the red line guard module 106 to prevent the rhythm switch from triggering the cross line and to ensure that the actions after the upgrade still comply with the safety constraints.

[0073] Furthermore, in the execution rhythm control module 108, when a lighter action is executed, the parameters of the lighter action are sent to the corresponding device for execution. In the observation window Inside, the system continuously acquires real-time data through the data acquisition unit, and the on-site snapshot generation module 101 generates new on-site snapshots to determine whether the revenue targets are met and whether the behavior patterns have changed. The triggering conditions for the judgment include: Returns not met: After the event, the actual decrease in λ did not reach the expected level, or λ remained above the normal upper boundary. Adversarial changes in behavioral patterns: For example, the attack source IP in Frequent changes within the attack network, or a shift in attack traffic from TCP to UDP, indicate that the attacker is circumventing the rules. Before deciding on an upgrade action, the red line gatekeeper module 106 is invoked again to make a judgment using the latest on-site snapshot and the predicted cost distribution of the target action, ensuring the safety of the upgrade action.

[0074] In this embodiment of the invention, the same-caliber evaluation module 109 evaluates the measured effect online using the same caliber during execution, using the same set of measurement standards across statistical caliber, execution caliber, and evaluation caliber, including the following steps: Actual data acquisition: After the observation window ends, actual data is collected using statistical methods that are completely consistent with those of the regularized playback module 105, including the actual decrease of λ, the actual offset of r and d, and the status of the maintenance session.

[0075] Deviation comparison: The measured values ​​are compared item by item with the median and upper bound of each action predicted by the regularization playback module 105, and the deviation value is calculated.

[0076] Deviation threshold setting: Set the profit deviation threshold. and cost deviation threshold .

[0077] Decision adjustment: When the measured deviation exceeds the corresponding threshold, the system will upgrade according to the established schedule; If the expected return is not achieved, escalate to the target action or a more powerful action; otherwise, if the effect is good and stable, conclude the current action.

[0078] Furthermore, in the same-caliber evaluation module 109, after the observation window ends, the same interface and statistical methods as the rule-based playback module 105 are used to collect indicators such as λ, r, and d after the actual execution of the action. Deviation threshold , These thresholds can be set based on historical experience and business tolerance. For example, the revenue deviation threshold. =10%. If the actual return is more than 10% lower than the predicted return, the deviation is considered too large, and the cost deviation threshold is set at 10%. =5%, if the actual cost is more than 5% higher than the predicted cost, then the deviation is considered too large; If the measured deviation exceeds the threshold, the system will trigger an alarm and make adjustments according to the preset strategy, such as automatic upgrade, rollback, or notification for manual intervention.

[0079] In this embodiment of the invention, the evidence retention and calibration module 110 solidifies the shortest causal chain evidence and drives parameter micro-calibration, including the following steps: Evidence solidification: Blockchain technology is used to solidify the chain of evidence, ensuring the immutability and traceability of the evidence; Each decision and execution process generates an evidence block, including on-site snapshots, urgency levels, candidate sets, similar fragment indexes, median and upper bounds of the benefits and costs of each action, red line determination results, the basis for the order of dominant choices, two-stage execution plans, and verification conclusions of online comparison with the same caliber. These evidence blocks are connected by a hash chain to form a complete decision chain, with unified numbering and retrieval capabilities.

[0080] Parameter calibration: The calibration process adopts an incremental learning method, and the system periodically statistically analyzes the distribution of measured versus estimated deviations. If the deviation continues to exceed the allowable range for multiple consecutive periods, the system will automatically adjust the relevant strategy parameters in small increments, such as the gate boundary for similar segment retrieval, red line tolerance and compliance line, deviation threshold, etc. The adjusted parameters will be released in versions, allowing the rules to gradually converge in the data in an incremental manner, thus achieving continuous adaptive optimization of the system.

[0081] Furthermore, in the evidence retention and calibration module 110, key information for each decision and execution step, such as inputs, outputs, intermediate results, timestamps, operator IDs, etc., is packaged into a data block, its hash value is calculated, and it is linked with the hash value of the previous data block to form an immutable blockchain. This evidence is stored in a distributed ledger, ensuring transparency and auditability. Furthermore, the system regularly runs offline calibration tasks. This task will analyze measured-predictive bias data in the historical evidence chain; If the deviation distribution is found to be continuously deviating, and the deviation between the predicted value and the measured value exceeds a certain threshold for N consecutive times, the system will start a parameter adjustment algorithm, such as Bayesian optimization or genetic algorithm, to make small iterative adjustments to the strategy parameters such as the gate boundary of similar segment retrieval, the tolerance value of the red line constraint, and the deviation threshold. The adjusted new parameter version will be signed and published, and recorded in the blockchain to achieve continuous optimization and adaptation of the strategy.

[0082] The scheduled time can be daily or weekly, depending on the actual situation.

[0083] Example 3 is an embodiment of the present invention. The above is an illustrative scheme of a method for situational awareness and performance evaluation in a digital power grid network security attack and defense game. It should be noted that the technical solution of the method for situational awareness and performance evaluation in a digital power grid network security attack and defense game is based on the same concept as the technical solution of the system for situational awareness and performance evaluation in a digital power grid network security attack and defense game described above. Details not described in detail in the technical solution of the method for situational awareness and performance evaluation in a digital power grid network security attack and defense game described above can be found in the description of the technical solution of the system for situational awareness and performance evaluation in a digital power grid network security attack and defense game described above.

[0084] This embodiment provides a method for situational awareness and performance evaluation in digital power grid network security attack and defense game, including: Multidimensional data containing operational, channel, and task information is collected from security equipment, network equipment, and business systems. The data is preprocessed and historical baselines are calculated to form a field snapshot that includes abnormal connection strength, channel latency, and relative packet loss. Based on relative quantities such as abnormal connections and channel deviations in the on-site snapshot, the emergency level is obtained by comparing them with the threshold value. Under the constraint of the emergency level, multiple safety handling actions are activated, and a set of execution parameters is generated for all actions. Retrieve historical segments that match on-site snapshots within a multi-dimensional space based on operational intensity, channel pressure, and operational status; simulate and replay the historical segments and statistically analyze the changes in indicators to generate the distribution of benefits and costs for all actions. The upper bound of the cost distribution output by the rule-based replay module is compared with three hard constraints in a veto-style manner to trim the strategy space so that all potential actions do not touch the safety red line. Candidate actions are eliminated based on the distribution of benefits and costs and hard safety constraints. Target actions are selected according to the dual ledger rules of benefits and costs. The light-level action of the target action is executed first, and then the target action is switched. During the execution process, the measured results are evaluated online using the same criteria. The same set of metrics is used across statistical, execution, and evaluation criteria to solidify the evidence of the shortest causal chain and drive parameter calibration in small steps.

[0085] This embodiment also provides an electronic device applicable to a digital power grid network security attack and defense game situational awareness and performance evaluation method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the digital power grid network security attack and defense game situational awareness and performance evaluation method proposed in the above embodiment.

[0086] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements a digital power grid network security attack and defense game situational awareness and performance evaluation method as proposed in the above embodiment.

[0087] The storage medium proposed in this embodiment and the method for realizing a digital power grid network security attack and defense game situation awareness and performance evaluation proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0088] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A digital power grid network security attack and defense game situational awareness and performance evaluation system, characterized in that: It includes a scene snapshot generation module, an emergency level determination module, a candidate action generation module, a similar segment retrieval module, a rule-based playback module, a red line gatekeeping module, a selection logic module, an execution rhythm control module, a consistent evaluation module, and an evidence retention and calibration module; The on-site snapshot generation module is responsible for real-time collection and processing of multi-dimensional data to generate on-site snapshots in a unified format. The emergency level determination module uses a multi-level threshold judgment strategy to classify the current network security situation into emergency levels based on the operational intensity and channel pressure information in the on-site snapshot; The candidate action generation module opens up a set of candidate actions for the corresponding emergency level based on the emergency level output by the emergency level determination module. The similar fragment retrieval module retrieves historical fragments that are similar to the current scene snapshot from historical data; The rule-based replay module simulates the execution effect of candidate actions on historical segments and outputs quantified benefit and cost distributions. The red line gatekeeper module uses three hard constraints to perform a veto-the-upper bound comparison on the cost distribution output by the rule-based replay module, thus pruning the strategy space and ensuring that all potential actions do not touch the safety red line. The decision-making logic module uses a dual-ledger control relationship to select actions from the feasible set filtered by the red-line gatekeeper module. The execution rhythm control module embodies the coupling of rhythm control and on-site perception in a two-stage manner, enabling the gradual execution and dynamic adjustment of handling actions; The same-caliber evaluation module evaluates the measured results online using the same caliber during the execution process, and uses the same set of measurement standards across statistical caliber, execution caliber, and evaluation caliber. The evidence retention and calibration module solidifies the shortest causal chain evidence and drives small-step parameter calibration.

2. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 1, characterized in that: The on-site snapshot generation module includes a data acquisition unit, a baseline calculation unit, a relative quantity calculation unit, and a snapshot encapsulation unit; The data acquisition unit obtains data from substation safety equipment, network equipment, and business systems; The baseline calculation unit calculates baseline values ​​for abnormal connection strength, link latency, and data loss rate; The relative quantity calculation unit calculates the abnormal surge ratio and channel deviation according to the baseline value and encapsulates them into a snapshot by the snapshot encapsulation unit.

3. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 2, characterized in that: The emergency level determination module includes a threshold learning unit and a determination execution unit; The threshold learning unit adjusts the threshold for abnormal surge ratio and the tolerance range for channel deviation based on the relationship between the abnormal surge ratio, channel deviation, and manually labeled emergency level in historical attack and defense data. The decision execution unit uses a decision tree or rule engine to calculate the abnormal surge ratio and channel deviation in parallel after receiving the on-site snapshot, and outputs the corresponding emergency level label. Among them, the emergency level determination module presets the abnormal surge ratio threshold and the channel deviation tolerance range, compares the abnormal surge ratio corresponding to the abnormal connection strength with the threshold, and compares the channel deviation corresponding to the link latency and data loss rate with the tolerance range. The emergency level is divided into high emergency level, medium emergency level and low emergency level according to the preset combination relationship, and the emergency level is used as the input mark for the candidate action generation module to open the action level.

4. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 3, characterized in that: The candidate action generation module includes four types of action templates predefined by the candidate action generation module: speed limit observation, precise blockade, forced source closure, and temporary section closure. Each action template includes parameter fields such as source address, destination address, port, protocol type, and duration of effect; The candidate action generation module selects the corresponding action template set based on the emergency level output by the emergency level determination module and instantiates specific candidate action items by combining the source identifier and entry identifier in the on-site snapshot.

5. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 4, characterized in that: The similar segment retrieval module includes a discrete space constructed by the similar segment retrieval module based on the running intensity level, channel pressure level, and work status level; A multidimensional index structure is used to store historical snapshots, and indexes are created for operation intensity files, channel pressure files, and operation status files. During retrieval, the index bucket is accessed according to discrete coordinates. Discrete coordinates are determined based on the abnormal surge ratio, channel deviation, and maintenance status in the on-site snapshots, and segments that match the discrete coordinates are retrieved from historical snapshots. When the matched segment obtained by relaxing the range is lower than the preset lower limit, the remaining dimensions are relaxed. However, when the inspection is true, the operation status range is not relaxed. The set of segments whose time distance from the on-site snapshot does not exceed the preset threshold is selected as the reference set from the candidate segments.

6. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 5, characterized in that: The rule-based replay module includes replaying historical traffic on the reference set obtained by the similar fragment retrieval module, injecting parameters for each candidate action and recording abnormal connection strength, link latency, data loss rate and maintenance session status. The median and high percentile values ​​of the benefits and costs of candidate actions are calculated according to statistical standards. The high percentile values ​​of costs are then filtered based on preset red line constraints, and a set of feasible actions is output.

7. The digital power grid network security attack and defense game situational awareness and performance evaluation system as described in claim 6, characterized in that: The execution rhythm control module includes the following: when the selection logic module selects the target action, the execution rhythm control module first executes a lighter action and calls the on-site snapshot generation module to generate a snapshot in the observation window; The deviation between the measured value of the indicator calculated by the same-caliber evaluation module and the predicted value by the rule-based replay module determines whether to upgrade to the target action or a stronger action. The evidence retention and calibration module encapsulates the process into evidence blocks with hash chains and adjusts the strategy parameters accordingly.

8. A method for situational awareness and performance evaluation in digital power grid network security attack and defense game, employing the digital power grid network security attack and defense game situational awareness and performance evaluation system as described in any one of claims 1 to 7, characterized in that, include: Multidimensional data containing operational, channel, and task information is collected from security equipment, network equipment, and business systems. The data is preprocessed and historical baselines are calculated to form a field snapshot that includes abnormal connection strength, channel latency, and relative packet loss. Based on relative quantities such as abnormal connections and channel deviations in the on-site snapshot, the emergency level is obtained by comparing them with the threshold value. Under the constraint of the emergency level, multiple safety handling actions are activated, and a set of execution parameters is generated for all actions. Retrieve historical segments that match on-site snapshots within a multi-dimensional space based on operational intensity, channel pressure, and operational status; simulate and replay the historical segments and statistically analyze the changes in indicators to generate the distribution of benefits and costs for all actions. The upper bound of the cost distribution output by the rule-based replay module is compared with three hard constraints in a veto-style manner to trim the strategy space so that all potential actions do not touch the safety red line. Candidate actions are eliminated based on the distribution of benefits and costs and hard safety constraints. Target actions are selected according to the dual ledger rules of benefits and costs. The light-level action of the target action is executed first, and then the target action is switched. During the execution process, the measured results are evaluated online using the same criteria. The same set of metrics is used across statistical, execution, and evaluation criteria to solidify the evidence of the shortest causal chain and drive parameter calibration in small steps.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the digital power grid network security attack and defense game situation awareness and performance evaluation system according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the digital power grid network security attack and defense game situation awareness and performance evaluation system as described in any one of claims 1 to 7.