Power grid safety protection system and method based on large language model

By adopting a large language model enhancement analysis module in the power grid security protection system, unstructured data is analyzed and multi-stage attack chains are generated, the problem of insufficient detection capabilities of new attack modes in the existing technology is solved, and the reliability of power grid security protection and operation and maintenance decision-making efficiency are improved.

CN120200802APending Publication Date: 2025-06-24TAIAN POWER SUPPLY CO OF STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510345382.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively parse unstructured text data in power grid logs and operation and maintenance reports, lacks adaptive detection capabilities to new attack modes, and lacks understandable semantic outputs in security detection results, which affects operation and maintenance decision efficiency.

Method used

The power grid security protection system based on the large language model is adopted to collect structured and unstructured data in real time through the multi-source data acquisition module, and the large language model enhancement analysis module is used to analyze unstructured data, identify security events and generate multi-stage attack chains, and generate dynamic defense response strategies based on the risk level.

Benefits of technology

It improves the reliability of power grid security protection, can effectively analyze unstructured data, adaptively detect new attack modes, and provides understandable semantic output, improving operation and maintenance decision-making efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200802A_ABST
    Figure CN120200802A_ABST
Patent Text Reader

Abstract

The invention provides a power grid security protection system and method based on a large language model. The system comprises a multi-source data acquisition module, a large language model enhancement analysis module, a security evaluation module and an interaction and output module, wherein the large language model enhancement analysis module generates a multi-stage attack chain; determining a corresponding risk level according to the importance index of the power grid equipment related to the generated attack chain and the real-time alarm data, and generating a dynamic defense response strategy according to the risk level; the security evaluation module calculates a power grid equipment security risk index based on the dynamic security scoring model according to the power field big language model, and determines a vulnerability repair strategy corresponding to the power grid equipment according to the power grid equipment security risk index; the interaction and output module generates a visual security report according to the multi-stage attack chain; and the dynamic defense response strategy and the vulnerability repair strategy are pushed to the dispatching center in real time, so that the reliability of power grid security protection is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power grid security, and in particular to a power grid security protection system and method based on a large language model. Background Art

[0002] With the rapid development of smart grids, the complexity and interconnectivity of power grid systems have increased significantly. Traditional network security detection methods are difficult to cope with the real-time analysis and threat identification of massive heterogeneous data (such as device status logs, power load data, communication protocol traffic).

[0003] The existing technologies mainly rely on rule engines or traditional machine learning models, and there are the following problems:

[0004] Unstructured text data such as power grid logs and operation and maintenance reports are difficult to be effectively parsed; there is a lack of adaptive detection capabilities for new attack patterns (such as APT attacks and protocol camouflage); the security detection results lack understandable semantic outputs, affecting the efficiency of operation and maintenance decision-making. The existing technologies mainly rely on static rule libraries or shallow machine learning models and cannot make full use of the semantic relevance of multi-source heterogeneous data (such as device status logs and communication protocol traffic); this makes the reliability of power grid security protection not high.

[0005] In view of this problem, the present invention provides a power grid security protection system and method based on a large language model to solve the above problems. Summary of the Invention

[0006] In order to solve the problems existing in the prior art, the present invention innovatively proposes a power grid security protection system and method based on a large language model, effectively solving the problem of low reliability of different power grid security protections caused by the prior art, and effectively improving the reliability of power grid security protection.

[0007] In the first aspect of the present invention, a power grid security protection system based on a large language model is provided, including:

[0008] A multi-source data acquisition module for real-time collecting structured data and unstructured data of power grid devices, where the structured data includes power grid device operation data and power grid device sensor data, and the unstructured data includes power grid device operation and maintenance logs and power grid work order texts; the power grid devices include SCADA systems, remote terminal units, programmable logic controllers, phasor measurement units, smart meters, human-machine interfaces, firewalls, intrusion detection / defense systems;

[0009] A large language model enhanced analysis module is used to parse unstructured data of grid terminals through a large language model in the power field, identify security events for grid equipment; generate multi-stage attack chains based on a pre-built power attack knowledge base, structured data, and unstructured data of grid equipment; determine corresponding risk levels according to the importance indicators of grid equipment involved in the generated attack chains and real-time alarm data, and generate dynamic defense response strategies according to the risk levels;

[0010] A security assessment module is used to parse structured data and unstructured data of grid equipment through a large language model in the power field, calculate the security risk index of grid equipment based on a dynamic security scoring model, and determine the corresponding vulnerability repair strategy for grid equipment according to the security risk index of grid equipment; the dynamic defense response strategy includes the corresponding vulnerability repair strategy for grid equipment;

[0011] An interaction and output module is used to generate a visual security report according to the multi-stage attack chain; push the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time.

[0012] Optionally, the large language model enhanced analysis module includes a semantic understanding sub-module, a threat reasoning sub-module, and a dynamic defense response strategy generation sub-module,

[0013] The semantic understanding sub-module is used to parse unstructured text data through a large language model in the power field and extract security events for grid equipment, where the security events include abnormal tripping of grid equipment, the event that the temperature of the insulator of grid equipment exceeds the threshold, and the communication delay alarm between grid equipment;

[0014] The threat reasoning sub-module is used to build a power attack knowledge base according to security events, structured data, and unstructured data of grid equipment, generate multi-stage attack chain hypotheses based on the pre-built power attack knowledge base, structured data, and unstructured data of grid equipment, and determine the corresponding risk levels according to the importance indicators of grid equipment involved in the generated attack chains and real-time alarm data; among them, the power attack knowledge base at least includes an attack pattern library, a defense strategy library, and real-time parameter features, and dynamically corrects the attack feasibility path through distribution network topology constraints and power business time series constraints; the attack pattern library includes the association relationship between each historical attack pattern and the real-time parameter features of each grid equipment, and each historical attack pattern corresponds to a multi-stage attack chain, so that the large language model in the power field can identify the attack traces in the real-time parameter features of each grid equipment through frequency domain clustering analysis and match the similar real-time parameter features in the historical attack patterns;

[0015] The dynamic policy generation sub-module is used to generate a dynamic defense response policy according to the risk level. The dynamic defense response policy includes device isolation instructions for grid devices involved in the attack chain, load balancing adjustment instructions, and the first-priority policy for vulnerability repair priorities. The first-priority policy for vulnerability repair is that the higher the risk level of the attack chain, the higher the repair priority of the involved grid devices.

[0016] Further, the first-priority policy for vulnerability repair priorities is specifically:

[0017] Quantify the risk level of the attack chain as the product of the attack success rate and the real-time impact coefficient; where the attack success rate is the CVSS score of the vulnerability in the attack chain, and the real-time impact coefficient is the proportion of the number of affected devices.

[0018] According to the path weight W of the power grid topology g , calculate the attack chain propagation risk, where the attack chain propagation risk is the product of the quantified risk level of the attack chain and the path weight W of the power grid topology g .

[0019] Construct a reward function according to the attack chain influence range, and sort according to the reward function value. The higher the reward function value, the higher the priority.

[0020] Optionally, the calculation method of the reward function is specifically:

[0021] Reward = -(αP loss + βT recover - γR p ), where α, β, and γ are the weight coefficients of the load caused by the attack chain leading to a power grid outage, the recovery time of the power grid outage caused by the attack chain, and the attack chain propagation risk respectively. The negative sign indicates minimizing the loss; P loss is the load caused by the attack chain leading to a power grid outage; T recover is the recovery time of the power grid outage caused by the attack chain; R p is the attack chain propagation risk, and the weight coefficients of the load caused by the attack chain leading to a power grid outage, the recovery time of the power grid outage caused by the attack chain, and the attack chain propagation risk support update through Q-learning.

[0022] Optionally, the security assessment module includes a protocol anomaly detection sub-module and a security assessment sub-module. The protocol anomaly detection sub-module is used to parse the structured data and unstructured data of grid devices according to the large language model and real-time rule engine in the power field, and identify abnormal fields in the power communication protocol; when an anomaly is detected, associate the anomaly event with the power security control platform to generate a violation work order.

[0023] The security assessment sub-module is used to calculate the power grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the power grid security risk index; the vulnerability repair strategy is the second-priority vulnerability repair strategy, and the second-priority vulnerability repair strategy is that according to the level of the power grid security risk index, the vulnerability repair priority decreases in turn, and the execution order of the second priority is less than that of the first priority.

[0024] Optionally, the formula of the dynamic security scoring model is:

[0025]

[0026] Among them, S index is the power grid security index, α is the positive security capability weight coefficient, β is the vulnerability risk weight coefficient, S is the network security resources and expenditures, P is the reliability of attack prevention measures, C is the attack control capability, I is the event handling efficiency, LS is the number of vulnerabilities, LY is the vulnerability severity, LG is the vulnerability vulnerability, LF is the vulnerability impact range, and M is the confidence score of the large language model in the power field.

[0027] Furthermore, the network security resources and expenditures include power grid equipment assets and security operation and maintenance investments: among them, the power grid network equipment assets include equipment procurement and upgrade costs; the security operation and maintenance investments include equipment security maintenance investments and operation and maintenance personnel management investments;

[0028] The formula for calculating the reliability of attack prevention measures is p = γ×C NIST +δ×E rate , where p is the reliability of attack prevention measures, C NIST is the matching degree of the firewall ACL rule and the industrial control system standard, E rate is the proportion of the protocol coverage rate between power grid devices E rate , γ is the weight of the matching degree of the firewall ACL rule and the industrial control system standard, and δ is the weight of the proportion of the protocol coverage rate between power grid devices;

[0029] The attack control capability is calculated by weighted calculation of the device-level intrusion interception rate and redundancy, and the expression is:

[0030] Among them, C is the attack control capability, IPS i is the interception times of power grid device i, n is the total number of power grid devices i, N attack,i is the total number of attacks on power grid device i, is the intrusion interception rate of power grid device i, R j represents the redundancy of redundant power grid device j, w IPS,i is the weight of power grid device i, w R,j is the weight of redundant power grid device j;

[0031] The method for obtaining the event handling efficiency is the product of the average response time of work orders and the automation disposal ratio;

[0032] The severity of the vulnerability is the CVSS score of the vulnerability; the vulnerability susceptibility is the number of exposed devices; the scope of vulnerability impact is the number of affected business systems; M is the confidence score of the LLM; when an anomaly is detected, the reliability of the attack prevention measures is reduced and the attack control requirements are increased.

[0033] Optionally, after receiving the pushed dynamic defense response policy and vulnerability repair policy, the operation and maintenance platform isolates the power grid devices involved in the attack chain, adjusts the load distribution of the power grid devices involved in the attack chain to the power grid devices not involved, obtains the risk level of the attack chain and the power grid security risk index, and repairs the vulnerabilities of the power grid devices involved in the attack chain in turn according to the execution order of the first priority policy for vulnerability repair and the second priority policy for vulnerability repair.

[0034] The second aspect of the present invention provides a power grid security protection method based on a large language model, which is implemented on the basis of a power grid security protection system based on the large language model described in the first aspect of the present invention, and includes:

[0035] The multi-source data acquisition module collects the structured data and unstructured data of power grid devices in real time. The structured data includes power grid device operation data and power grid device sensor data, and the unstructured data includes power grid device operation and maintenance logs and power grid work order texts; the power grid devices include SCADA systems, remote terminal units, programmable logic controllers, phasor measurement units, smart meters, human-machine interfaces, firewalls, intrusion detection / defense systems;

[0036] The large language model enhanced analysis module analyzes the unstructured data of the power grid terminal through the large language model of the power field to identify security events for power grid devices; generates a multi-stage attack chain based on the pre-built power attack knowledge base, the structured data and unstructured data of power grid devices; determines the corresponding risk level according to the importance index of the power grid devices involved in the generated attack chain and the real-time alarm data, and generates a dynamic defense response policy according to the risk level;

[0037] The security assessment module analyzes the structured data and unstructured data of power grid devices through the large language model of the power field, calculates the power grid device security risk index based on the dynamic security scoring model, and determines the corresponding vulnerability repair policy for the power grid device according to the power grid device security risk index; the dynamic defense response policy includes the corresponding vulnerability repair policy for power grid devices;

[0038] The interaction and output module generates a visual security report according to the multi-stage attack chain; and pushes the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time.

[0039] Optionally, it further includes:

[0040] After receiving the pushed dynamic defense response strategy and the vulnerability repair strategy, the operation and maintenance platform isolates the power grid equipment involved in the attack chain, adjusts the load distribution of the power grid equipment involved in the attack chain to the power grid equipment not involved, obtains the risk level of the attack chain and the power grid security risk index, and repairs the vulnerabilities of the power grid equipment involved in the attack chain in sequence according to the execution order of the vulnerability repair first priority strategy and the vulnerability repair second priority strategy.

[0041] The technical solution adopted by the present invention includes the following technical effects:

[0042] 1. The large language model enhanced analysis module of the present invention generates a multi-stage attack chain based on a pre-built power attack knowledge base, structured data and unstructured data of power grid equipment; determines the corresponding risk level according to the importance index of the power grid equipment involved in the generated attack chain and real-time alarm data, and generates a dynamic defense response strategy according to the risk level; the security assessment module analyzes the structured data and unstructured data of the power grid equipment according to the large language model in the power field, calculates the power grid equipment security risk index based on the dynamic security scoring model, and determines the corresponding vulnerability repair strategy for the power grid equipment according to the power grid equipment security risk index; the interaction and output module generates a visual security report according to the multi-stage attack chain; and pushes the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time, effectively solving the problem of low reliability of different power grid security protection caused by the prior art, and effectively improving the reliability of power grid security protection.

[0043] 2. In the large language model enhanced analysis module of the technical solution of the present invention, the threat inference sub-module generates a multi-stage attack chain hypothesis based on a pre-built power attack knowledge base, structured data and unstructured data of power grid equipment, determines the corresponding risk level according to the importance index of the power grid equipment involved in the generated attack chain and real-time alarm data; generates a dynamic defense response strategy according to the risk level, and the dynamic defense response strategy includes an equipment isolation instruction, a load balancing adjustment instruction and a vulnerability repair first priority strategy for the power grid equipment involved in the attack chain, and the vulnerability repair first priority strategy is that the higher the risk level of the attack chain, the higher the repair priority of the involved power grid equipment, which improves the adaptability of power grid security risk protection.

[0044] 3. The specific strategy of the first priority for vulnerability repair in the technical solution of the present invention is as follows: Quantify the risk level of the attack chain as the product of the attack success rate and the real-time impact coefficient; Calculate the propagation risk of the attack chain according to the path weight of the power grid topology; Construct a reward function based on the influence range of the attack chain, and sort according to the value of the reward function. The higher the value of the reward function, the higher the priority, so that the priority of vulnerability repair further corresponds to the risk level of the attack chain, and further improves the reliability of power grid security risk protection.

[0045] 4. In the technical solution of the present invention, the protocol anomaly detection sub-module is used to parse the structured data and unstructured data of power grid devices according to the large language model in the power field and the real-time rule engine, and identify the abnormal fields in the power communication protocol; When an anomaly is detected, associate the anomaly event with the power security management and control platform to generate a violation work order; The security scoring sub-module is used to calculate the power grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the power grid security risk index; The vulnerability repair strategy is the second priority strategy for vulnerability repair. The second priority strategy for vulnerability repair is that according to the level of the power grid security risk index, the priority of vulnerability repair decreases in turn, and the execution order of the second priority is less than that of the first priority, which further improves the adaptability of the vulnerability repair strategy.

[0046] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic structural diagram of the system in Embodiment 1 of the present invention;

[0049] Figure 2 It is a schematic flow diagram (1) of the method in Embodiment 2 of the present invention;

[0050] Figure 3 It is a schematic flow diagram (2) of the method in Embodiment 2 of the present invention;

[0051] Figure 4 It is a schematic flow diagram (3) of the method in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present invention omits the description of well-known components and processing technologies and processes to avoid unnecessarily limiting the present invention.

[0053] Embodiment 1

[0054] As Figure 1 shown, the present invention provides a power grid security protection system based on a large language model, including:

[0055] A multi-source data acquisition module 1, configured to collect structured data and unstructured data of power grid devices in real time. The structured data includes power grid device operation data and power grid device sensor data, and the unstructured data includes power grid device operation and maintenance logs and power grid work order texts. The power grid devices include SCADA systems, remote terminal units, programmable logic controllers, phasor measurement units, smart meters, human-machine interfaces, firewalls, intrusion detection / defense systems;

[0056] A large language model enhanced analysis module 2, configured to parse the unstructured data of power grid terminals through a large language model in the power field to identify security events for power grid devices; generate a multi-stage attack chain based on a pre-built power attack knowledge base, the structured data, and the unstructured data of power grid devices; determine the corresponding risk level according to the importance indicators of the power grid devices involved in the generated attack chain and real-time alarm data, and generate a dynamic defense response strategy according to the risk level;

[0057] A security assessment module 3, configured to parse the structured data and unstructured data of power grid devices through a large language model in the power field, calculate the power grid device security risk index based on a dynamic security scoring model, and determine the corresponding vulnerability repair strategy for the power grid devices according to the power grid device security risk index; the dynamic defense response strategy includes the corresponding vulnerability repair strategy for power grid devices;

[0058] An interaction and output module 4, configured to generate a visual security report according to the multi-stage attack chain; and push the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time.

[0059] Among them, the network device log collection sub-module 11 in the multi-source data collection module 1 is used to collect the log data of firewalls, routers, and switches; the sensor data collection sub-module 12 is used to collect the real-time monitoring data of temperature and humidity sensors; the communication protocol traffic collection sub-module 13 is used to parse the encrypted traffic data of TCP, UDP, and ICMP protocols; the edge computing sub-module 14 is deployed on the substation side and supports local data preprocessing and real-time transmission.

[0060] The structured data of power grid equipment includes operation data, sensor data, traffic data, and status data. The unstructured data of power grid equipment includes operation and maintenance logs, work order texts, and user behavior data. Power grid equipment includes SCADA systems (Supervisory Control And Data Acquisition), remote terminal units, programmable logic controllers (PLCs), phasor measurement units (PMUs), smart meters, human-machine interfaces (HMIs), firewalls, intrusion detection / defense systems (IDS / IPS);

[0061] The multi-source data collection module 1 supports the parsing and standardization of power-specific protocols, and the power-specific protocols can include IEC61850, DNP3, and Modbus; deploy an edge gateway to collect the data of substation power grid equipment in real time, and integrate the operation and maintenance management system logs through the API interface; data classification: network device logs, sensor data (temperature, humidity), communication protocol traffic (TCP / UDP / ICMP), user behavior data, etc. Deploy an edge gateway to collect the data of substation power grid equipment and collect logs in real time through the Flume tool; data preprocessing: clean noise data (such as invalid timestamps, duplicate records); standardize (unify the time format, protocol field encoding); encrypt sensitive data (using the AES-256 algorithm).

[0062] The large language model enhanced analysis module 2 includes a semantic understanding sub-module 21, a threat inference sub-module 22, and a dynamic defense response strategy generation sub-module 23.

[0063] The semantic understanding sub-module 21 is used to parse unstructured text data through a large language model in the power domain and extract security events for power grid equipment. The security events include abnormal tripping of power grid equipment, over-threshold temperature events of insulators of power grid equipment, and communication delay alarms between power grid equipment.

[0064] The large language model in the power field is a large language model obtained by fine-tuning LLM (such as GPT-4, BERT) to adapt to power field terms (the training dataset includes power equipment manuals, historical work orders and fault cases, and the accuracy of parsing professional terms of the fine-tuned large language model is ≥95%); extract key security events (such as "abnormal circuit breaker tripping", "insulator temperature exceeding the threshold"); use the attention mechanism to optimize text semantic parsing, and the accuracy is increased to 92%; and obtain the real-time parameter characteristics of the device corresponding to the security event of each device, so as to establish a power attack knowledge base according to the real-time parameter characteristics of the device corresponding to the security event of each device in the later stage.

[0065] The threat inference sub-module 22 is used to construct a power attack knowledge base based on security events, structured data of power grid equipment and unstructured data, generate multi-stage attack chain hypotheses based on the pre-built power attack knowledge base, structured data of power grid equipment and unstructured data, and determine the corresponding risk level according to the importance index of the power grid equipment involved in the generated attack chain and real-time alarm data; wherein, the power attack knowledge base includes the association relationship between each historical attack mode and the real-time parameter characteristics of each power grid equipment, and each historical attack mode corresponds to a multi-stage attack chain, so that the large language model in the power field can identify the attack traces in the real-time parameter characteristics of each power grid equipment through frequency domain clustering analysis and match the similar real-time parameter characteristics in the historical attack mode; the importance index can include device type, the location of the device in the power grid, and the number of connections between devices.

[0066] The power attack knowledge base associates historical attack modes with device real-time parameter data, generates multi-stage attack chain hypotheses, and simulates attack paths. The method for constructing the power attack knowledge base includes data collection and integration, mainly including historical attack modes, device real-time parameter data and external knowledge frameworks; the association relationship between each historical attack mode and the real-time parameter characteristics of each power grid equipment, and each historical attack mode corresponds to a multi-stage attack chain, so that the large language model in the power field can identify the attack traces in the real-time parameter characteristics of each power grid equipment through frequency domain clustering analysis and match the similar real-time parameter characteristics in the historical attack mode.

[0067] The power attack knowledge base can also be structured through knowledge structuring and graph construction, mainly including entity relationship modeling and dynamic knowledge updating; the graph nodes represent entities or concepts in the attack chain, including attack themes, attack tools, attack targets, defense measures, etc.; the node attribute types are differentiated according to different node categories. For example, the attributes of the attack tool node include tool name, attack stage, etc., the attributes of the attack target node include device type, IP address, physical location, etc., and the attributes of the defense measure node include defense rule ID, defense coverage, effective time, etc.; the edges represent the dynamic relationships or behaviors between nodes, used to describe the evolution logic of the attack chain or the response path of the defense strategy; the connection of the edges directly reflects the coupling of the power grid physical-information system, and the physical devices and logical systems are associated through the edges. In the simulation of attack path generation, the reachable paths between nodes in the knowledge graph are used to generate possible attack sequences.

[0068] The association analysis and reasoning mechanism mainly includes attack pattern matching and dynamic risk assessment. The specific content of the power attack knowledge base can include attack patterns, defense strategy libraries, and real-time data metrics. The association mechanism is to identify attack traces in real-time measurement parameter data through frequency domain clustering analysis, and match similar features in historical attack patterns. The multi-stage attack chain hypothesis generation method is that the large model matches historical attack cases with abnormal real-time parameter data, matches similar features in historical attack patterns, and makes attack hypotheses based on parameter features (such as "a sudden increase in traffic at a certain node may be a DDoS attack"); the core of attack hypothesis generation is to match key indicators in historical attack patterns through feature attributes, and combine the functional characteristics and security context of power grid devices to infer the attack intention and the evolution direction of the attack chain. Feature attributes include time series patterns (such as periodic anomalies), device behavior deviations (such as a sudden increase in device CPU / memory utilization), network topology association anomalies (such as unusual connections between devices), etc. The corresponding hypotheses for the above attributes are false data injection attacks that affect power grid state perception by tampering with PMU (phasor measurement unit) measurement data, malware residency that directly affects device availability, and lateral movement penetration that breaks through network partitions and threatens core devices. In the power grid environment, the core goal of attack path simulation (such as "traffic surge → protocol spoofing → device overload") is to identify how attackers use information layer vulnerabilities (such as software defects) and physical layer weaknesses (such as device exposure surfaces) to achieve the purpose of disrupting power supply. Some nodes of the attack path are power grid devices, and the constraint categories are four types: physical topology constraints (restricting the possibility of attackers physically contacting devices and affecting the attack propagation path), communication protocol constraints (restricting attackers from parsing protocols), network architecture constraints (restricting lateral movement paths), and device function constraints (preventing malicious firmware implantation).

[0069] According to the importance indicators of the power grid equipment involved in the generated attack chain and the real-time alarm data, determine the corresponding risk level. One implementation method for determining the risk level can be: through multi-source data fusion analysis (network traffic, equipment status, protocol communication, and physical measurement) of the power grid equipment involved in the attack chain, first perform standardization processing and key feature extraction on heterogeneous data (such as protocol anomaly frequency, CPU load, measurement contradiction value), determine the normal range corresponding to each real-time parameter feature, and divide the numerical values of the parameter features that exceed the normal range into 4 intervals on average according to the numerical size. The interval closest to the normal range corresponds to a low risk level, the interval farthest from the normal range corresponds to a severe risk level, and the risk levels corresponding to the middle two intervals are medium and high in sequence according to the interval numerical size. Take the highest risk level corresponding to a certain real-time parameter feature among all power grid equipment as the overall risk level of the power grid equipment involved in the attack chain. Another implementation method can also be to perform parameter clustering on the real-time parameter features of each power grid equipment through a large language model in the power field, train according to the manual annotation results, obtain the real-time parameter ranges of all power grid equipment corresponding to each risk level, and determine the risk level corresponding to the real-time parameter features of the current power grid equipment according to the real-time parameter ranges of all power grid equipment. There may also be other risk level classification methods, as long as it can achieve the corresponding risk level division according to the real-time parameters of different power grid equipment. The embodiments of the present invention do not limit this here.

[0070] Finally, according to the importance indicators of the power grid equipment involved in the generated attack chain, adjust the risk levels corresponding to the real-time parameter characteristics of all current power grid equipment; one implementation method can be: when the equipment type of the power grid equipment involved in the attack chain belongs to the equipment type preset by the importance indicators, or the equipment location belongs to the power grid location preset by the importance indicators, or the number of connections between equipment exceeds the preset quantity threshold, the corresponding risk level is increased by 1 until the highest level. The importance indicators of power grid equipment include equipment type, the power grid location where the equipment is located, and the number of connections between equipment. If the equipment type is SCADA system (Supervisory Control And Data Acquisition system), PLC (Programmable Logic Controller), PMU (Phasor Measurement Unit), firewall, IDS / IPS (Intrusion Detection / Prevention System) equipment (important equipment, that is, the equipment type preset by the importance indicators), the risk level of the power grid equipment involved in the attack chain is increased by 1; if it is other equipment (non-important equipment), the risk level of the power grid equipment involved in the attack chain remains unchanged; if the power grid location where the equipment is located is a key substation (such as a hub substation, or a substation above a certain voltage level, that is, the power grid location preset by the importance indicators), the risk level of the power grid equipment involved in the attack chain is increased by 1; if the power grid location where the equipment is located is a non-key substation, the risk level of the power grid equipment involved in the attack chain remains unchanged; if a certain power grid equipment is connected to 3 or more other power grid equipment (only for illustrative purposes, the specific quantity can be flexibly adjusted according to the actual situation), the risk level of the power grid equipment involved in the attack chain is increased by 1; if a certain power grid equipment is connected to 2 or 1 other power grid equipment, the risk level of the power grid equipment involved in the attack chain remains unchanged.

[0071] Another implementation method can be as follows: when the proportion of the number of power grid devices involved in the attack chain (whose device type belongs to the device type preset by the importance index) in the total number of all power grid devices involved in the attack chain is greater than the first preset percentage threshold, or when the proportion of the number of power grid devices involved in the attack chain (whose device location belongs to the power grid location preset by the importance index) in the total number of all power grid devices involved in the attack chain is greater than the second preset percentage threshold; or when the proportion of the number of power grid devices involved in the attack chain (whose number of connections between devices exceeds the preset quantity threshold) in the total number of all power grid devices involved in the attack chain is greater than the third preset percentage threshold, the corresponding risk level is increased by 1 until the highest level; the first preset percentage threshold, the second preset percentage threshold, and the third preset percentage threshold can be the same or different, and can be flexibly adjusted according to the actual situation. When the proportion of the number of power grid devices involved in the attack chain (whose device type belongs to the device type preset by the importance index) in the total number of all power grid devices involved in the attack chain is not greater than the first preset percentage threshold, or when the proportion of the number of power grid devices involved in the attack chain (whose device location belongs to the power grid location preset by the importance index) in the total number of all power grid devices involved in the attack chain is not greater than the second preset percentage threshold; or when the proportion of the number of power grid devices involved in the attack chain (whose number of connections between devices exceeds the preset quantity threshold) in the total number of all power grid devices involved in the attack chain is not greater than the third preset percentage threshold, the corresponding risk level remains unchanged.

[0072] The dynamic policy generation sub-module 23 is used to generate a dynamic defense response policy according to the risk level. The dynamic defense response policy includes a device isolation instruction, a load balancing adjustment instruction, and a first-priority policy for vulnerability repair for the power grid devices involved in the attack chain. The first-priority policy for vulnerability repair is that the higher the risk level of the attack chain, the higher the repair priority of the involved power grid devices.

[0073] Load balancing adjustment mainly involves the large language model in the power field combining a large number of historical attack patterns, analyzing their characteristics and countermeasures, and dynamically adjusting for real-time data anomalies. By analyzing the characteristics of historical attack patterns (such as protocol abuse, vulnerability exploitation paths, physical-information layer coordinated attacks), the dynamic policy generation module uses the large language model (LLM) in the power field to combine real-time data anomalies (such as traffic mutations, device overloads, measurement contradictions) to generate response instructions, including isolating high-risk devices (such as blocking the lateral communication of infected SCADA), dynamically adjusting the load (such as switching the PMU redundant channel to resist FDI attacks), and preferentially repairing exposed vulnerabilities (such as urgently patching the weak password vulnerability of RTU). At the same time, the policy is dynamically optimized based on the characteristics of power grid devices (such as SCADA redundant control, RTU firmware rollback mechanism) and security constraints (such as network partitioning, real-time requirements).

[0074] Among them, the dynamic policy generation sub-module 23 integrates the reinforcement learning algorithm, dynamically optimizes the priority and execution order of the response policy according to the real-time threat feedback. The dynamic optimization method is as follows: first, quantify the attack features into threat levels, secondly, calculate the policy priority through the reward function according to the power grid topology and the attack influence range, then give priority to executing high-threat actions and delay low-impact operations, and finally update the model parameters based on the policy execution effect to adapt to new attack patterns;

[0075] Specifically, the first-priority policy for vulnerability repair is specifically as follows:

[0076] Quantify the risk level of the attack chain as the product of the attack success rate and the real-time impact coefficient; among them, the attack success rate is the CVSS score of the vulnerability in the attack chain, and the real-time impact coefficient is the proportion of the number of affected devices;

[0077] Quantify the threat level L t as the product of the attack success rate (such as the CVSS score of the vulnerability) and the real-time impact coefficient (such as the proportion of the number of affected devices)

[0078] According to the path weight W of the power grid topology g , calculate the attack chain propagation risk, where the attack chain propagation risk is the product of the quantified product of the risk level of the attack chain and the path weight W of the power grid topology g of the product;

[0079] Combined with the critical path weight W of the power grid topology g (such as the control link weight from SCADA to RTU is 0.8, and the ordinary link is 0.3), calculate the attack propagation risk R p =L t ×W g .

[0080] Construct a reward function according to the attack chain influence range, and sort according to the reward function value. The higher the reward function value, the higher the priority. Among them, the calculation method of the reward function is specifically as follows:

[0081] Reward=-(αP loss +βT recover -γR p ), where α, β, γ are the weight coefficients of the load amount of the power grid outage caused by the attack chain, the recovery time of the power grid outage caused by the attack chain, and the attack chain propagation risk respectively. The negative sign indicates minimizing the loss; P loss is the load amount of the power grid outage caused by the attack chain; T recover is the recovery time of the power grid outage caused by the attack chain; R pFor the propagation risk of the attack chain, and the load amount of power grid outage caused by the attack chain, the recovery time of power grid outage caused by the attack chain, and the weight coefficients of the propagation risk of the attack chain support updating through Q-learning.

[0082] Finally, the policy priorities are sorted by Reward, and the actions of grid equipment corresponding to high negative values (i.e., high P loss , long T recover or high R p ) are preferentially executed (such as isolating the core SCADA to block the attack chain), while low-risk operations (such as non-critical PLC log analysis) are delayed, and the weight coefficients α, β, γ are updated through Q-learning (reinforcement learning algorithm) to adapt to new attack patterns.

[0083] The security assessment module 3 includes a protocol anomaly detection sub-module 31 and a security assessment sub-module 32. The protocol anomaly detection sub-module 31 is used to parse the structured and unstructured data of grid equipment according to the large language model and real-time rule engine in the power field, and identify the abnormal fields in the power communication protocol; when an anomaly is detected, the anomaly event is associated with the power security control platform to generate a violation work order;

[0084] Combined with the LLM and the rule engine, identify the abnormal fields in the power communication protocol (such as IEC 61850 message tampering); support real-time traffic parsing, and the response delay ≤ 500ms.

[0085] The security assessment sub-module 32 is used to calculate the power grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the power grid security risk index; the vulnerability repair strategy is the second priority strategy for vulnerability repair, and the second priority strategy for vulnerability repair is that according to the level of the power grid security risk index, the vulnerability repair priority decreases in turn, and the execution order of the second priority is less than that of the first priority.

[0086] Among them, the formula of the dynamic security scoring model is:

[0087]

[0088] Among them, S index is the power grid security index, α is the positive security capability weight coefficient, β is the vulnerability risk weight coefficient, S is the network security resources and expenditures, P is the reliability of attack prevention measures, C is the attack control ability, I is the event handling efficiency, LS is the number of vulnerabilities, LY is the severity of vulnerabilities, LG is the vulnerability susceptibility, LF is the scope of vulnerability impact, and M is the confidence score of the large language model in the power field.

[0089] Specifically, network security resources and expenditures include power grid equipment assets, security operation and maintenance investments, and network resource occupancy costs of power grid security equipment: Among them, power grid network equipment assets include equipment procurement and upgrade costs (such as deploying industrial control firewalls to protect SCADA and replacing RTUs that support encrypted communication); security operation and maintenance investments include equipment security maintenance investments (such as fixing PLC firmware vulnerabilities and optimizing PMU data encryption policies) and operation and maintenance personnel management investments (such as SCADA real-time monitoring teams and IDS rule library update teams); the network resource occupancy cost of security equipment is associated with traffic processing overheads (such as the CPU consumption of the firewall analyzing DNP3 protocol traffic and the bandwidth occupancy cost when the IPS detects lateral movement), and resources need to be dynamically allocated according to the real-time load of the equipment. Finally, a security cost model centered on equipment and covering the entire life cycle is formed to ensure that high-value equipment (such as substation SCADA) obtains priority resource allocation.

[0090] The reliability calculation formula for attack prevention measures is p = γ × C NIST + δ × E rate , where p is the reliability of attack prevention measures, and C NIST is the matching degree between the firewall ACL rules and the industrial control system standard (NIST SP 800-82 standard), and E rate is the coverage ratio of the protocols between power grid devices (such as the TLS protocol between SCADA and RTU) to E rate , γ is the weight of the matching degree between the firewall ACL rules and the industrial control system standard, and δ is the weight of the coverage ratio of the protocols between power grid devices; among them, the weights γ and δ are assigned according to the criticality of the equipment (such as the weight of the core firewall is 0.6 and that of the ordinary PLC is 0.2).

[0091] The attack control ability is calculated by weighting the device-level intrusion interception rate and redundancy, and the expression is:

[0092] Among them, C is the attack control ability, and IPS i is the number of interceptions of power grid device i, n is the total number of power grid devices i, and N attack,i is the total number of attacks on power grid device i, is the intrusion interception rate of power grid device i, and R j represents the redundancy of redundant power grid device j (the number of redundant devices / the total number of devices), and w IPS,i is the weight of power grid device i, and w R,j is the weight of redundant power grid device j; the weights w IPS,i and w R,j are assigned according to the criticality of the equipment.

[0093] The method for obtaining the event handling efficiency is the product of the average response time of work orders and the automated handling ratio; that is, I is the event handling efficiency, which is determined by the statistical average response time of work orders (MTTR) and the automated handling ratio.

[0094] The vulnerability severity is the CVSS score of the vulnerability (≥7.0 is defined as high risk); the vulnerability susceptibility is the number of exposed devices; the vulnerability impact scope is the number of affected business systems; according to the analysis results of the power domain LLM on the real-time attack mode, the vulnerability weights LS, LY, LG, and LF can be dynamically adjusted. The dynamic adjustment method is that the power domain LLM parses the attack logs and threat intelligence in real time, identifies the frequently exploited vulnerabilities and attack paths, and dynamically increases the weights of LS (activity) and LY (exploitation difficulty) of the relevant vulnerabilities; adjusts LG (impact scope) according to the proportion of affected devices, and updates LF (repair cost) in combination with the repair timeliness. Federated learning synchronizes multi-region attack data to optimize the weight coefficients; aggregates multi-node data through the federated learning framework to optimize the model generalization ability.

[0095] M is the confidence score of the LLM, and the calculation method can be:

[0096] M = 0.6×semantic analysis accuracy + 0.3×logical consistency score + 0.1×historical detection credibility,

[0097] The semantic analysis accuracy, logical consistency score, and historical detection result credibility are all reference standards in the existing confidence calculation process of large models, and are not limited in the embodiments of the present invention.

[0098] When the protocol anomaly detection sub-module 31 detects an anomaly, the security scoring sub-module 32 reduces the reliability P of the attack prevention measure and increases the attack control requirement C.

[0099] Preferably, the security assessment module 3 may further include a device status prediction sub-module 33, which uses the time series analysis ability of the large model LLM in the power field and the LSTM neural network to predict the failure risk of key devices; the LLM component extracts device status features (insulation aging) by analyzing the structured and unstructured features of power grid devices; the LSTM neural network component synchronously analyzes the time series correlation of device parameters (such as temperature) to construct a multi-dimensional failure feature vector; the device status features extracted by the LLM component are used as input data and input into the LSTM neural network component to form a hybrid prediction model to predict the failure risk of power grid devices. That is, using the time series analysis ability of LLM to predict the failure risk of key devices (such as the overload probability of transformers); adopting the LSTM network to enhance time series modeling, the prediction accuracy is increased by 25%. LSTM focuses on numerical time series analysis (such as transformer temperature, current curve) to capture the changing rules of the physical state of devices and predict obvious failures such as overload. LLM analyzes operation and maintenance texts (logs, work orders) to extract hidden risks (such as descriptions of "insulation aging") and correlates with external threat intelligence (such as vulnerability exploitation codes). The output of LSTM is spliced with the semantic features of LLM for joint training of the prediction model. LLM identifies attack-related risks, and LSTM quantifies the impact of attacks to dynamically adjust the device redundancy strategy and enhance the defense resilience against unknown threats. Device status prediction predicts the number LG of vulnerable devices and the impact range LF through LSTM, thereby increasing the vulnerability severity LY; the two act on the security index formula comprehensively.

[0100] The interaction and output module 4 is used to generate a visual security report, support natural language interaction, and push response policies to the operation and maintenance terminal (operation and maintenance platform) in real time through a message queue. The visual report generated by the interaction and output module 4 includes: a three-dimensional threat map, marking the attack path and power grid devices, and the attack path is based on the power grid topology and vulnerability dependencies, and simulates the attacker's penetration path through a knowledge graph (such as from the external network VPN → SCADA server → breaker control); a device health status heat map, which is dynamically updated based on the prediction results; an interactive repair guide, which supports adjusting response policies through natural language instructions. The steps to generate a visual report include a threat map, a device health status heat map, and step-by-step guidance on repair suggestions; support natural language interaction (such as inputting "How to mitigate the current risk?" and the system generates step-by-step operation instructions); push response policies to the operation and maintenance terminal in real time through a message queue (such as Kafka); the steps to generate the repair suggestions are as follows: the system first analyzes the key risk nodes (such as exploited vulnerabilities, abnormal communication ports) and threat types (such as ransomware, DDoS) in the attack path through the threat map, and combines the real-time status indicators (such as CPU / memory overload, firmware version expiration, abnormal process activity) in the device health heat map to locate the vulnerable points of physical devices or logical services; then, through cross-matching with the knowledge base, dynamically generate a repair strategy with a priority ranking.

[0101] Optimize the LLM based on power domain texts (equipment manuals, historical work orders) to increase the accuracy of professional term parsing by 30%; utilize the generative ability of the LLM to simulate multi-stage attack chains to achieve hypothetical detection of unknown threats; generate a visual analysis report of security events through the LLM (such as "insulator temperature exceeds the threshold → heat dissipation system failure → tripping risk"), with an interpretability of over 90%.

[0102] After the operation and maintenance platform receives the pushed dynamic defense response strategy and vulnerability repair strategy, isolate the power grid equipment involved in the attack chain, adjust the load distribution of the power grid equipment involved in the attack chain to the power grid equipment not involved, obtain the risk level of the attack chain and the power grid security risk index, and sequentially perform vulnerability repair on the power grid equipment involved in the attack chain according to the execution order of the first priority strategy for vulnerability repair and the second priority strategy for vulnerability repair.

[0103] When adjusting the load distribution, first conduct real-time monitoring and analysis of the power grid topology structure to identify power grid equipment and important lines. At the same time, combine the operating status data of the power grid, such as voltage, current, power and other information, to evaluate the load capacity of each power grid equipment and the transmission capacity of the lines; on this basis, according to the safe operation requirements of the power grid and the principle of power supply continuity, give priority to isolating the power grid equipment involved in the attack chains identified as having a high risk level and a high power grid security risk index, and redistribute the load by adjusting the switch status, transformer tap position, etc. in the power grid to ensure the stable operation of the power grid and the reliability of power supply; in addition, this strategy will also consider the security status of the power grid equipment, and for power grid equipment with potential safety hazards, corresponding measures will be taken for isolation or repair to prevent the spread of security risks.

[0104] The large language model enhancement analysis module of the present invention generates multi-stage attack chains based on a pre-built power attack knowledge base, structured data and unstructured data of power grid equipment; determines the corresponding risk level according to the importance index of the power grid equipment involved in the generated attack chain and real-time alarm data, and generates a dynamic defense response strategy according to the risk level; the security assessment module analyzes the structured data and unstructured data of the power grid equipment according to the large language model in the power domain, calculates the power grid equipment security risk index based on the dynamic security scoring model, and determines the corresponding vulnerability repair strategy for the power grid equipment according to the power grid equipment security risk index; the interaction and output module generates a visual security report based on the multi-stage attack chain; and pushes the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time, effectively solving the problem of low reliability of different power grid security protections caused by the existing technology, and effectively improving the reliability of power grid security protection.

[0105] In the threat inference sub-module of the large language model enhanced analysis module in the technical solution of the present invention, multi-stage attack chain hypotheses are generated based on a pre-built power attack knowledge base, grid equipment structured data, and unstructured data. According to the importance indicators of the grid equipment involved in the generated attack chain and real-time alarm data, the corresponding risk levels are determined; dynamic defense response strategies are generated according to the risk levels. The dynamic defense response strategies include equipment isolation instructions, load balancing adjustment instructions, and vulnerability repair priority first priority strategies for the grid equipment involved in the attack chain. The vulnerability repair first priority strategy is that the higher the risk level of the attack chain, the higher the repair priority of the involved grid equipment, improving the adaptability of grid security risk protection.

[0106] The vulnerability repair priority first priority strategy in the technical solution of the present invention is specifically: quantifying the risk level of the attack chain as the product of the attack success rate and the real-time impact coefficient; calculating the attack chain propagation risk according to the path weight of the power grid topology; constructing a reward function according to the attack chain influence range, and sorting according to the reward function value. The higher the reward function value, the higher the priority, making the priority of vulnerability repair further correspond to the risk level of the attack chain, and further improving the reliability of grid security risk protection.

[0107] In the technical solution of the present invention, the protocol anomaly detection sub-module is used to parse the structured data and unstructured data of grid equipment according to the large language model and real-time rule engine in the power field, and identify the abnormal fields in the power communication protocol; when an anomaly is detected, the anomaly event is associated with the power security control platform to generate a violation work order; the security scoring sub-module is used to calculate the grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the grid security risk index; the vulnerability repair strategy is the vulnerability repair second priority strategy. The vulnerability repair second priority strategy is that according to the level of the grid security risk index, the vulnerability repair priority decreases in turn, and the execution order of the second priority is lower than that of the first priority, further improving the adaptability of the vulnerability repair strategy.

[0108] Embodiment 2

[0109] As Figures 2 - 3 shown, the technical solution of the present invention also provides a power grid security protection method based on a large language model, which is implemented on the basis of a power grid security protection system based on a large language model in Embodiment 1, and includes:

[0110] S1. The multi-source data acquisition module collects structured data and unstructured data of power grid equipment in real time. The structured data includes power grid equipment operation data and power grid equipment sensor data, and the unstructured data includes power grid equipment operation and maintenance logs and power grid work order texts. The power grid equipment includes SCADA systems, remote terminal units, programmable logic controllers, phasor measurement units, smart meters, human-machine interfaces, firewalls, intrusion detection / defense systems.

[0111] The multi-source data acquisition module mainly performs data acquisition and preprocessing: deploying edge gateways to collect substation equipment data and using the Flume tool to collect logs in real time; cleaning noise data (such as invalid timestamps and duplicate records); performing standardization processing (unifying time formats and protocol field encodings); encrypting sensitive data (using the AES-256 algorithm).

[0112] The data types mainly include: Network device log data: Collecting log files generated by network devices, including logs of devices such as firewalls, routers, and switches.

[0113] Sensor data: Collecting sensor data in the network, such as temperature sensors and humidity sensors.

[0114] Operation and maintenance text data: Collecting text data recorded by operation and maintenance personnel, such as fault reports and operation logs.

[0115] Communication protocol traffic data: Collecting traffic data of network communication protocols, such as traffic of protocols such as TCP, UDP, and ICMP.

[0116] System status data: Collecting system operation status data, such as CPU usage and memory usage.

[0117] User behavior data: Collecting user behavior data in the network, such as login behaviors and access records.

[0118] The data preprocessing steps are as follows: Data cleaning: Removing noise data and invalid data; Data annotation: Annotating the data for use in model training; Data standardization: Converting the data into a unified format and standard; Data deduplication: Removing duplicate data to improve data quality; Data fusion: Fusing different types of data to form a complete data set; Data encryption: Encrypting sensitive data to ensure data security.

[0119] S2, The large language model enhanced analysis module parses the unstructured data of the power grid terminals through the large language model in the power field, identifies security events targeting power grid equipment; generates multi-stage attack chains based on the pre-built power attack knowledge base, structured data, and unstructured data of power grid equipment; determines the corresponding risk levels according to the importance indicators of the power grid equipment involved in the generated attack chains and real-time alarm data, and generates dynamic defense response strategies according to the risk levels;

[0120] Training and optimization of the large language model in the power field;

[0121] Fine-tune the pre-trained model (such as BERT), training dataset: 100,000 pieces of text in the power field (equipment manuals, fault cases); fine-tune the LLM model: fine-tune the large language model based on the collected text data to make it adapt to the terms and protocol specifications in the power field. Optimize the model parameters: adjust the model parameters to optimize the model performance (hyperparameter settings: learning rate = 1e-5, batch size = 32, number of training rounds = 50; evaluation metrics: accuracy ≥ 90%, F1 value ≥ 0.85. The LLM parses key events in the logs (such as "communication delay alarm"), collect text in the power field: collect text data related to the power system, such as equipment manuals, operation and maintenance reports, fault cases, etc.). Construct the training dataset: construct a dataset suitable for model training, including labeled data and unlabeled data. Model verification: verify the trained model and evaluate its performance. Model storage: store the trained model for subsequent use.

[0122] S3, The security assessment module parses the structured data and unstructured data of power grid equipment through the large language model in the power field, calculates the power grid equipment security risk index based on the dynamic security scoring model, and determines the corresponding vulnerability repair strategy for the power grid equipment according to the power grid equipment security risk index; the dynamic defense response strategy includes the corresponding vulnerability repair strategy for the power grid equipment;

[0123] S4, The interaction and output module generates a visual security report based on the multi-stage attack chain; real-time push the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform.

[0124] Use Tableau to generate a threat map, mark high-risk nodes and attack paths; integrate a natural language interaction interface (such as ChatGPT plugin) to support operation and maintenance personnel to query and issue instructions in real time.

[0125] As Figure 4 shown, the technical solution of the present invention also provides a power grid security protection method based on a large language model, further including:

[0126] After the operation and maintenance platform receives the pushed dynamic defense response strategy and vulnerability repair strategy, it isolates the power grid devices involved in the attack chain, adjusts the load distribution of the power grid devices involved in the attack chain to the power grid devices not involved, obtains the risk level of the attack chain and the power grid security risk index, and repairs the vulnerabilities of the power grid devices involved in the attack chain in turn according to the execution order of the vulnerability repair first priority strategy and the vulnerability repair second priority strategy.

[0127] The large language model enhanced analysis module of the present invention generates a multi-stage attack chain based on a pre-built power attack knowledge base, structured data and unstructured data of power grid devices; determines the corresponding risk level according to the importance index of the power grid devices involved in the generated attack chain and real-time alarm data, and generates a dynamic defense response strategy according to the risk level; the security assessment module analyzes the structured data and unstructured data of the power grid devices according to the large language model in the power field, calculates the power grid device security risk index based on the dynamic security scoring model, and determines the corresponding vulnerability repair strategy for the power grid devices according to the power grid device security risk index; the interaction and output module generates a visual security report according to the multi-stage attack chain; and pushes the dynamic defense response strategy and the vulnerability repair strategy to the operation and maintenance platform in real time, effectively solving the problem of low reliability of different power grid security protections caused by the existing technology, and effectively improving the reliability of power grid security protection.

[0128] In the threat inference sub-module of the large language model enhanced analysis module in the technical solution of the present invention, a multi-stage attack chain hypothesis is generated based on a pre-built power attack knowledge base, structured data and unstructured data of power grid device organizations, and the corresponding risk level is determined according to the importance index of the power grid devices involved in the generated attack chain and real-time alarm data; a dynamic defense response strategy is generated according to the risk level, and the dynamic defense response strategy includes a device isolation instruction, a load balancing adjustment instruction and a vulnerability repair first priority strategy for the power grid devices involved in the attack chain. The vulnerability repair first priority strategy is that the higher the risk level of the attack chain, the higher the repair priority of the involved power grid devices, which improves the adaptability of power grid security risk protection.

[0129] The vulnerability repair first priority strategy in the technical solution of the present invention is specifically: quantifying the risk level of the attack chain as the product of the attack success rate and the real-time impact coefficient; calculating the attack chain propagation risk according to the path weight of the power grid topology; constructing a reward function according to the attack chain influence range, and sorting according to the reward function value. The higher the reward function value, the higher the priority, so that the priority of vulnerability repair further corresponds to the risk level of the attack chain, and further improves the reliability of power grid security risk protection.

[0130] In the technical solution of the present invention, the protocol anomaly detection sub-module is used to parse the structured data and unstructured data of grid equipment according to the large language model and real-time rule engine in the power field, and identify the abnormal fields in the power communication protocol; when an anomaly is detected, the anomaly event is associated with the power safety control platform to generate a violation work order; the safety scoring sub-module is used to calculate the grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the grid security risk index; the vulnerability repair strategy is the second priority vulnerability repair strategy, and the second priority vulnerability repair strategy is that according to the level of the grid security risk index, the vulnerability repair priority decreases in turn, and the execution order of the second priority is less than that of the first priority, further improving the adaptability of the vulnerability repair strategy.

[0131] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A power grid security protection system based on a large language model, characterized in that: include: A multi-source data acquisition module is used to collect structured data and unstructured data of power grid equipment in real time. The structured data includes power grid equipment operation data and power grid equipment sensor data, and the unstructured data includes power grid equipment operation and maintenance logs and power grid work order texts. The power grid equipment includes SCADA systems, remote terminal units, programmable logic controllers, phasor measurement units, smart meters, human-machine interfaces, firewalls, and intrusion detection / prevention systems. The large language model enhanced analysis module is used to parse the unstructured data of power grid terminals through the large language model in the power field and identify security incidents targeting power grid equipment; Generate a multi-stage attack chain based on the pre-built power attack knowledge base, structured data of power grid equipment, and unstructured data; determine the corresponding risk level based on the importance indicators of the power grid equipment involved in the generated attack chain and the real-time alarm data, and generate a dynamic defense response strategy based on the risk level; The security assessment module is used to parse the structured and unstructured data of power grid equipment according to the large language model in the power field, calculate the security risk index of power grid equipment based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy of power grid equipment according to the security risk index of power grid equipment; The dynamic defense response strategy includes a corresponding vulnerability repair strategy for power grid equipment; Interaction and output module, used to generate visual security reports based on multi-stage attack chains; Push dynamic defense response strategies and vulnerability repair strategies to the operation and maintenance platform in real time.

2. The power grid security protection system based on a large language model according to claim 1 is characterized in that: The large language model enhanced analysis module includes a semantic understanding submodule, a threat reasoning submodule, and a dynamic defense response strategy generation submodule. The semantic understanding submodule is used to parse unstructured text data through a large language model in the power field and extract security events for power grid equipment, including abnormal tripping of power grid equipment, temperature exceeding threshold of power grid equipment insulators, and communication delay alarms between power grid equipment; The threat reasoning submodule is used to build a power attack knowledge base based on security events, structured data of power grid equipment and unstructured data, generate a multi-stage attack chain hypothesis based on the pre-built power attack knowledge base, structured data of power grid equipment and unstructured data, and determine the corresponding risk level according to the importance index of the power grid equipment involved in the generated attack chain and the real-time alarm data; wherein the power attack knowledge base includes the association relationship between each historical attack mode and the real-time parameter characteristics of each power grid equipment, and each historical attack mode corresponds to a multi-stage attack chain, so that the large language model in the power field can identify the attack traces in the real-time parameter characteristics of each power grid equipment through frequency domain clustering analysis, and match similar real-time parameter characteristics in the historical attack mode; the importance index includes the type of equipment, the location of the equipment in the power grid, and the number of connections between equipment; The dynamic strategy generation submodule is used to generate a dynamic defense response strategy according to the risk level. The dynamic defense response strategy includes device isolation instructions, load balancing adjustment instructions and a first priority strategy for vulnerability repair for the power grid equipment involved in the attack chain. The first priority strategy for vulnerability repair is that the higher the risk level of the attack chain, the higher the repair priority of the power grid equipment involved.

3. The power grid security protection system based on a large language model according to claim 2 is characterized in that: The priority strategy for vulnerability repair is as follows: The risk level of the attack chain is quantified as the product of the attack success rate and the real-time impact coefficient. The attack success rate is the CVSS score of the vulnerability in the attack chain, and the real-time impact coefficient is the proportion of the number of affected devices. According to the path weight W of the power grid topology g , calculate the attack chain propagation risk, where the attack chain propagation risk is the product of the quantized risk level of the attack chain and the path weight W of the power grid topology g The product of Construct a reward function based on the impact range of the attack chain and sort by reward function value. The higher the reward function value, the higher the priority.

4. The power grid security protection system based on a large language model according to claim 2 is characterized in that: The reward function is calculated as follows: Reward=-(αP loss +βT recover -γR p ), where α, β, and γ are the load of the power outage caused by the attack chain, the recovery time of the power outage caused by the attack chain, and the weight coefficient of the risk of the attack chain propagation, respectively. The negative sign indicates minimizing the loss; P loss is the load of the attack chain that causes power outage in the power grid; T recover is the recovery time of the power grid outage caused by the attack chain; R p It is the risk of attack chain propagation, and the load of power outage caused by the attack chain, the recovery time of power outage caused by the attack chain, and the weight coefficient of attack chain propagation risk can be updated through Q-learning.

5. The power grid security protection system based on a large language model according to claim 2 is characterized in that: The security assessment module includes a protocol anomaly detection submodule and a security scoring submodule. The protocol anomaly detection submodule is used to parse the structured data and unstructured data of the power grid equipment according to the large language model of the power field and the real-time rule engine, and identify abnormal fields in the power communication protocol; When an abnormality is detected, the abnormal event is associated with the power safety management and control platform to generate a violation work order; The security scoring submodule is used to calculate the power grid security risk index based on the dynamic security scoring model, and determine the corresponding vulnerability repair strategy according to the power grid security risk index; the vulnerability repair strategy is a vulnerability repair second priority strategy, and the vulnerability repair priority of the second priority strategy decreases in sequence according to the power grid security risk index, and the execution order of the second priority is less than the execution order of the first priority.

6. The power grid security protection system based on a large language model according to claim 2 is characterized in that: The dynamic safety scoring model formula is: Among them, S index is the power grid security index, α is the positive security capability weight coefficient, β is the vulnerability risk weight coefficient, S is the network security resources and expenditure, P is the reliability of attack prevention measures, C is the attack control capability, I is the event handling efficiency, LS is the number of vulnerabilities, LY is the severity of the vulnerability, LG is the vulnerability to attacks, LF is the impact range of the vulnerability, and M is the confidence score of the large language model in the power field.

7. The power grid security protection system based on a large language model according to claim 6 is characterized in that: The network security resources and expenditures include power grid equipment assets and security operation and maintenance investment: among which, power grid network equipment assets include equipment procurement and upgrade costs; security operation and maintenance investment includes equipment security maintenance investment and operation and maintenance personnel management investment; The reliability calculation formula of attack prevention measures is p = γ × C NIST +δ×E rate , where p is the reliability of the attack prevention measure, C NIST is the matching degree between the firewall ACL rules and the industrial control system standards, E rate E is the coverage ratio of the protocol between power grid equipment rate , γ is the weight of the matching degree between the firewall ACL rule and the industrial control system standard, and δ is the weight of the coverage ratio of the protocol between power grid devices; The attack control capability is calculated by weighting the device-level intrusion interception rate and redundancy, and the expression is: Among them, C is the attack control capability, IPS i is the number of interceptions of power grid device i, n is the total number of power grid devices i, N attack,i is the total number of attacks on power grid device i, is the intrusion interception rate of power grid equipment i, R j represents the redundancy of redundant power grid equipment j, w IPS,i is the weight of grid equipment i, w R,j is the weight of redundant power grid equipment j; The event processing efficiency is obtained by multiplying the average response time of the work order by the proportion of automated processing; The severity of the vulnerability is the vulnerability CVSS score; the vulnerability susceptibility is the number of exposed devices; the vulnerability impact scope is the number of affected business systems; M is the LLM confidence score; when an anomaly is detected, the reliability of attack prevention measures is reduced and the need for attack control is increased.

8. The power grid security protection system based on a large language model according to claim 5 is characterized in that: After receiving the pushed dynamic defense response strategy and vulnerability repair strategy, the operation and maintenance platform isolates the power grid equipment involved in the attack chain, adjusts the load distribution of the power grid equipment involved in the attack chain to the power grid equipment not involved, obtains the risk level of the attack chain and the power grid security risk index, and repairs the vulnerabilities of the power grid equipment involved in the attack chain in turn according to the execution order of the vulnerability repair first priority strategy and the vulnerability repair second priority strategy.

9. A power grid security protection method based on a large language model, characterized in that: The invention is implemented on the basis of a power grid security protection system based on a large language model according to any one of claims 1 to 8, comprising: The multi-source data acquisition module collects structured data and unstructured data of power grid equipment in real time. The structured data includes power grid equipment operation data and power grid equipment sensor data, and the unstructured data includes power grid equipment operation and maintenance logs and power grid work order texts. The power grid equipment includes SCADA system, remote terminal unit, programmable logic controller, phasor measurement unit, smart meter, human-machine interface, firewall, intrusion detection / prevention system; The large language model enhanced analysis module uses the large language model in the power field to parse the unstructured data of the power grid terminal and identify security incidents against the power grid equipment. It generates a multi-stage attack chain based on the pre-built power attack knowledge base, the structured data of the power grid equipment, and the unstructured data. It determines the corresponding risk level based on the importance indicators of the power grid equipment involved in the generated attack chain and the real-time alarm data, and generates a dynamic defense response strategy based on the risk level. The security assessment module analyzes the structured data and unstructured data of the power grid equipment according to the large language model in the power field, calculates the security risk index of the power grid equipment based on the dynamic security scoring model, and determines the corresponding vulnerability repair strategy of the power grid equipment according to the security risk index of the power grid equipment; the dynamic defense response strategy includes the corresponding vulnerability repair strategy of the power grid equipment; The interaction and output module generates a visual security report based on the multi-stage attack chain; and pushes dynamic defense response strategies and vulnerability repair strategies to the operation and maintenance platform in real time.

10. The power grid security protection method based on a large language model according to claim 9, characterized in that: include: After receiving the pushed dynamic defense response strategy and vulnerability repair strategy, the operation and maintenance platform isolates the power grid equipment involved in the attack chain, adjusts the load distribution of the power grid equipment involved in the attack chain to the power grid equipment not involved, obtains the risk level of the attack chain and the power grid security risk index, and repairs the vulnerabilities of the power grid equipment involved in the attack chain in turn according to the execution order of the vulnerability repair first priority strategy and the vulnerability repair second priority strategy.

Citation Information

Cited By

  • Network security multi-mode intelligent detection system and method

    CN120498904A

  • Multi-terminal cooperative communication monitoring alarm system for 5G new call

    CN120957172A

  • Method and device for evaluating safety protection capability of power optical transmission system

    CN121012679A

  • AI fault analysis system for box transformer measurement and control device

    CN121461313A

  • Human-computer interaction method and system based on AI large model

    CN121742373A