Substation hidden risk identification method, system and device and storage medium

By employing deep semantic vectorization and heterogeneous graph construction techniques, the problems of fuzzy semantic parsing and multidimensional topology modeling in substation risk identification were solved, enabling accurate identification and systematic assessment of hidden risks and improving the safety and stability of substations.

CN121301760APending Publication Date: 2026-01-09GUANGXI POWER GRID CO LTD FANGCHENGGANG POWER SUPPLY BUREAU

Patent Information

Application Number
CN202511219878.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing substation risk identification methods cannot deeply analyze the ambiguous semantics in unstructured operation and maintenance texts, and it is difficult to model the multidimensional topological relationships between devices. This makes it impossible to systematically identify hidden risks and hazards formed by the aggregation of multiple low-risk events through complex network transmission.

Method used

By processing operation and maintenance text data through deep semantic vectorization, a heterogeneous diagram of substations is constructed. An aggregation algorithm with time decay and a graph attention mechanism are used to iteratively aggregate information, generate equipment risk indices, and produce interpretable risk tracing reports.

Benefits of technology

It enables accurate identification and systematic assessment of potential risks in substations, timely detection of signs of equipment deterioration, generation of interpretable risk tracing reports, and improvement of the substation's ability to ensure safe and stable operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301760A_ABST
    Figure CN121301760A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of substation operation and maintenance. The invention provides a transformer substation hidden risk identification method, system and device and a storage medium. The method comprises the following steps: performing deep semantic vectorization processing on an operation and maintenance text, and constructing an equipment time sequence semantic snapshot; generating an equipment state change trend vector; defining nodes and edges based on the equipment state, the regional topology and the risk type, constructing a transformer substation heterogeneous graph, and designing a meta-path; aggregating heterogeneous neighbor information by using a graph attention mechanism, configuring an independent parameter matrix for different relationships, and expanding a sensing range through multi-round iteration until a node state and a risk entropy are converged; and finally, quantizing and outputting equipment or regional risk indexes, and generating an interpretable traceability report by combining attention weight backtracking key paths and abnormal states. The problems that in the prior art, unstructured text fuzzy semantics cannot be deeply analyzed, the multi-dimensional topological relation between devices is difficult to model, and therefore multiple low-risk events cannot be recognized and hidden dangers are formed through network conduction aggregation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of substation operation and maintenance, in particular to a substation implicit risk identification method, system, device and storage medium. BACKGROUND

[0002] In the field of power system operation and maintenance, ensuring the safety and stability of substations is a core task. Currently, risk identification in substations mainly relies on two types of technology: online real-time monitoring based on SCADA and other systems and offline operation and maintenance data analysis based on expert experience or keyword statistics. Online monitoring can effectively capture explicit faults of single devices that reach the threshold (such as over-temperature and over-current), but it is powerless for the large number of fuzzy and precursory descriptions in operation and maintenance texts such as "slight oil seepage" and "abnormal noise". Offline analysis has significant limitations: it is highly subjective and inefficient due to reliance on human experience; simple keyword statistics cannot deeply understand the true semantics of the text; more importantly, existing methods cannot effectively model the complex electrical connections, mechanical interlocks, and protection logic between the text information and the multi-dimensional topology of the devices in the substation. This results in the inability to identify systemic implicit risks that are formed by the combination, conduction, and aggregation of multiple seemingly independent "low-risk" events (such as minor abnormalities in the main transformer and delayed operation of the associated circuit breaker) through complex device networks, making risk awareness remain at the "point" level and making it difficult to achieve "surface" systemic evaluation, with safety operation and maintenance in a passive response state.

[0003] Further analysis of existing technologies shows that although some studies involve graph theory in power grid applications (such as CN117390003A focusing on topology data verification and visualization) or combine multi-modal data (such as CN119885095A mining misoperation patterns and CN120198106A conducting intelligent operation and maintenance), they have not effectively addressed the core challenges of substation implicit risk identification. These technologies mainly rely on structured historical data or pre-set modalities and lack the ability to deeply analyze the fuzzy semantics in unstructured operation and maintenance texts (work tickets, inspection reports), making it difficult to accurately quantify the potential risks described in "slight oil seepage" and other descriptions; their graph models or correlation analysis are usually simple (such as isomorphic graphs or single relationship types), ignoring the essential characteristics of the diverse types of devices (transformers, circuit breakers, protection devices, etc.) and the heterogeneous connection relationships (electrical connections, mechanical interlocks, protection logic) within the substation, making it difficult to accurately model the complex multi-dimensional interactions between devices. SUMMARY

[0004] The present application aims to provide a substation implicit risk identification method, system, device and storage medium, aiming to solve the problem that the existing substation risk identification method cannot deeply analyze the fuzzy semantics in unstructured text, and it is difficult to model the multi-dimensional topological relationship between devices, resulting in the inability to systematically identify the implicit risk hazards formed by the aggregation of multiple low-risk events through complex networks.

[0005] The present application is realized by the following technical solutions:

[0006] A substation implicit risk identification method comprises the following steps:

[0007] The operation and maintenance text data of the substation is subjected to deep semantic vectorization processing, and the fuzzy description and precursor description are converted into semantic vectors to form a semantic snapshot of the target power device at the corresponding time point;

[0008] Through an aggregation algorithm with time decay, the semantic snapshot of the target power device at the corresponding time point is subjected to time sequence aggregation to generate a comprehensive state vector reflecting the state change trend;

[0009] Based on the comprehensive state vector, the node type and edge type are defined, and the edge sequence connecting different types of nodes is taken as a meta-path to construct a substation heterogeneous graph;

[0010] Based on the substation heterogeneous graph, the node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta-path, and the information transmission weight is calculated in combination with the graph attention;

[0011] Through several rounds of information aggregation iteration, the node perception range is expanded until the L2 norm change rate of the node state vector and the global risk entropy value fluctuation meet the convergence condition;

[0012] The risk quantization output is performed on the iterated node state vector to obtain the risk index of the device or region, and the attention mechanism is used to backtrack the key path and abnormal state to generate an interpretable risk traceability report.

[0013] Based on the same inventive concept, the present application also provides a substation implicit risk identification system for realizing the substation implicit risk identification method, comprising:

[0014] A semantic processing module is used to collect the operation and maintenance text data of the substation, convert the fuzzy description and precursor description into semantic vectors through a pre-trained deep language model, and generate a semantic snapshot of the target power device at the corresponding time point;

[0015] A state aggregation module is connected to the semantic processing module and performs time sequence weighted fusion on the semantic snapshot through an aggregation algorithm with time decay to generate a comprehensive state vector reflecting the state change trend;

[0016] An isomorphic graph construction module connected to the state aggregation module is configured to define six types of node and four types of edge, construct a substation isomorphic graph containing multi-dimensional features of equipment and dynamic edge weights based on a comprehensive state vector, and verify topological consistency through meta-path combination and analytic hierarchy process;

[0017] A graph computing engine module connected to the isomorphic graph construction module is configured to filter neighbor nodes based on a meta-path, configure independent parameter matrices for the four types of edges, calculate dynamic weights of neighbor information transmission combining a graph attention mechanism, perform multiple rounds of information aggregation iteration, and determine convergence through L2 norm change rate and global risk entropy fluctuation.

[0018] A risk quantification and tracing module connected to the graph computing engine module is configured to input the state vector of the node after iteration into a fully connected neural network, output a device risk probability value, aggregate regional device risk probability values to generate a regional risk index, trace key meta-paths and original operation and maintenance texts based on attention weights, and construct a three-level diagnosis report.

[0019] Based on the same inventive concept, the application further provides an electronic device comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program to enable the electronic device to perform the substation implicit risk identification method.

[0020] Based on the same inventive concept, the application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is configured to be executed by a processor to implement the substation implicit risk identification method.

[0021] The technical solution of the application has at least the following advantages and beneficial effects:

[0022] Through deep semantic vectorization processing of the operation and maintenance text data, the fuzzy and precursory descriptions such as "slight oil seepage" and "abnormal sound" are converted into semantic vectors, thereby breaking through the limitations of existing online monitoring technologies that are powerless to fuzzy information and offline analysis that is difficult to deeply understand the real semantics of the text, and enabling the potential risk signals contained in the text to be accurately captured, thereby providing a richer and more accurate information source for risk identification.

[0023] The aggregation algorithm with time decay is used to perform time sequence aggregation on the semantic snapshot to generate a comprehensive state vector, which can effectively reflect the change trend of the state of the target power equipment over time, overcome the shortcomings of the prior art in static analysis of the state of the equipment, and make the risk identification more dynamic and forward-looking, thereby facilitating the timely discovery of signs of deterioration of the state of the equipment.

[0024] The substation heterogeneous graph is constructed based on the comprehensive state vector, node types and edge types are defined, and the edge sequence is taken as a meta path, the diversity of the types of devices in the substation and the heterogeneity of the connection relationship (electrical connection, mechanical locking, protection logic, etc.) are fully considered, the complex multi-dimensional interaction between devices can be more truly and comprehensively modeled, and key support is provided for identifying systematic hidden risks.

[0025] On the basis of the heterogeneous graph, the node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta path, the information transmission weight is calculated in combination with the graph attention, the node sensing range is expanded through multiple rounds of information aggregation iteration until the convergence condition is met, the influence weight between devices can be more accurately calculated, the systematic hidden risk hazards formed by the combination, conduction and aggregation of multiple seemingly independent "low-risk" events through the complex device network can be effectively identified, the limitation that the risk cognition of the prior art stays at the "point" level is broken through, and "area" systematic evaluation is achieved.

[0026] The risk quantization output is performed on the iterated node state vector, the risk index of the device or the area is obtained, the key path and the abnormal state are traced back based on the attention mechanism, and an interpretable risk traceability report is generated, which can not only clearly indicate the specific position and degree of the risk, but also clearly show the formation path and the key influencing factors of the risk, is helpful for workers to quickly locate problems and develop targeted operation and maintenance strategies, changes safety operation and maintenance from passive response to active prevention, and significantly improves the guarantee capability of safe and stable operation of the substation. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 It is a flowchart of a substation hidden risk identification method of an embodiment of the application.

[0028] Figure 2 It is a structural diagram of a substation hidden risk identification system of an embodiment of the application. DETAILED DESCRIPTION

[0029] The following is a specific embodiment combined with the drawings.

[0030] REFERENCE Figure 1 A substation hidden risk identification method, comprising the following steps S101-S106:

[0031] S101, the operation and maintenance text data of the substation is subjected to deep semantic vectorization processing, the vague description and the precursor description are converted into semantic vectors, and a semantic snapshot of the target power device at the corresponding time point is formed.

[0032] In some embodiments, the specific process of the deep semantic vectorization processing of the operation and maintenance text data of the substation, the conversion of the vague description and the precursor description into semantic vectors, and the formation of the semantic snapshot of the target power device at the corresponding time point is as follows:

[0033] collecting operation and maintenance text data of the transformer substation, the operation and maintenance text data including work tickets, equipment inspection records, defect reports, and operation logs;

[0034] encoding semantics of the operation and maintenance text by using a pre-trained deep language model, and converting fuzzy descriptions in the operation and maintenance text into high-dimensional vectors;

[0035] forming a semantic snapshot of the target power equipment at a corresponding time point by using the floating-point number vectors generated after encoding the operation and maintenance text records, the semantic snapshot quantitatively representing differences in risk levels of the text descriptions in a vector space.

[0036] It can be understood that the operation and maintenance text data of the transformer substation is a collection of texts recording various information in the operation, maintenance, and repair processes of the transformer substation, covering work tickets, equipment inspection records, defect reports, and operation logs. The work ticket is a written order for approving work on power equipment, which clearly defines the work content, work location, work time, and work responsible person, and the like. The equipment inspection record is the equipment state information recorded by the operation and maintenance personnel after regularly checking the equipment of the transformer substation, including the appearance of the equipment, the operating parameters, and the like. The defect report is used to record the defect conditions of the equipment, such as defect type, discovery time, severity, and the like. The operation log details the operation process and results of the operation personnel on the equipment.

[0037] Here, the fuzzy description and the precursor description are common but difficult to directly quantify and analyze in the operation and maintenance text data. The fuzzy description refers to those descriptions that are not accurate enough and have uncertainties, such as “slight abnormal noise of the equipment”. Such a description does not clearly define the degree, frequency, and other key information of the abnormal noise. The precursor description is a record of some signs appearing before the equipment may have problems, such as “insulation aging signs”, which implies the risk of future failure of the equipment.

[0038] Specifically, the deep semantic vectorization process is a process of converting text data into numerical vectors using deep learning technology, which can include: first, collecting operation and maintenance text data of the transformer substation, which has a wide range of sources and various formats; then, encoding semantics of the operation and maintenance text by using a pre-trained deep language model (such as BERT or similar models), the pre-trained deep language model being a neural network model pre-trained based on a large amount of text data and capable of understanding semantic information of the text. Under the action of this model, the fuzzy descriptions in the operation and maintenance text are converted into high-dimensional vectors, the high-dimensional vector being a group of numerical values composed of multiple floating-point numbers, each floating-point number representing a feature of the text in a certain semantic dimension.

[0039] Further, by the floating-point number vector generated after the operation and maintenance text record coding, a semantic snapshot of the target power equipment at the corresponding time point can be formed. The semantic snapshot is a quantitative representation of the state of the target power equipment (the target power equipment is various electrical equipment in the substation that needs to be monitored and maintained, such as transformers, circuit breakers, disconnectors, etc.) at a specific time point, which exists in the form of a high-dimensional floating-point number vector. In the vector space, different semantic snapshots correspond to different positions, and the distance between these positions reflects the difference in the risk level of the text description. For example, the distance between the semantic snapshot describing the existence of a serious fault hidden danger of the equipment and the semantic snapshot describing the normal operation state of the equipment in the vector space is far, while the distance between the semantic snapshot describing the existence of a slight abnormality of the equipment and the semantic snapshot describing the normal operation state is relatively close. The dimension of the semantic snapshot depends on the output of the deep language model used, and different models may generate vectors of different dimensions.

[0040] Exemplarily, the semantic snapshot plays a core role in the operation and maintenance management of the power equipment, and it is the basic unit of the equipment state. Based on these semantic snapshots, subsequent time series aggregation and graph construction can be performed. Time series aggregation is to integrate and analyze the semantic snapshots of the same equipment at different time points to observe the trend of the change of the equipment state over time; graph construction is to take the semantic snapshots of different equipment as nodes and construct a graph structure by analyzing the association relationship between them, so as to identify potential risks. For example, by analyzing the association between multiple equipment semantic snapshots, it can be found that the problem of a certain equipment may affect other equipment, and then measures can be taken in advance to prevent risks.

[0041] S102, performing time series aggregation on the semantic snapshot of the target power equipment at the corresponding time point by using an aggregation algorithm with time decay, to generate a comprehensive state vector reflecting the trend of state change.

[0042] In some embodiments, the specific process of performing time series aggregation on the semantic snapshot of the target power equipment at the corresponding time point by using an aggregation algorithm with time decay to generate a comprehensive state vector reflecting the trend of state change is as follows:

[0043] An exponential decay function is used to dynamically calculate the weight of the historical semantic snapshot of the target power equipment; wherein the weight of the semantic snapshot generated at the current time point is the highest, and the weight of the historical semantic snapshot decreases exponentially with the increase of the time interval; by using a weighted aggregation algorithm, all semantic snapshots of the target power equipment are fused into a single vector according to the decay weight, to generate a comprehensive state vector of the target power equipment; when new operation and maintenance text is added, the comprehensive state vector is iteratively updated based on the semantic snapshot of the new operation and maintenance text and the current comprehensive state vector with a time decay weight update mechanism.

[0044] It can be understood that the time-decay aggregation algorithm is a data processing method that comprehensively considers the influence of time factors on data. Among them, time decay embodies the characteristic that the influence of historical data on current state evaluation gradually weakens over time. Specifically, the algorithm can dynamically calculate the weight of the historical semantic snapshot of the target power equipment by using an exponential decay function, which is a mathematical function whose function value decreases exponentially with the increase of the independent variable. In this algorithm, the independent variable is the time interval, that is, the time difference between the historical semantic snapshot generation time and the current time. The semantic snapshot generated at the current time point has the highest weight because it best reflects the current actual state of the device; and the weight of the historical semantic snapshot decreases exponentially with the increase of the time interval, which means that the semantic snapshot that is more distant from the current time contributes less to the comprehensive state vector.

[0045] Further, after obtaining the weight of each historical semantic snapshot of the target power equipment, all semantic snapshots of the target power equipment are fused according to the decay weight by using a weighted aggregation algorithm. Here, the weighted aggregation algorithm is a method of fusing data based on the calculated weight, that is, multiplying each semantic snapshot by its corresponding weight and then adding them together to finally generate a single vector, which is the comprehensive state vector of the target power equipment. The comprehensive state vector integrates the state information of the device at different time points and takes into account the influence of time factors on the state information, and can more accurately reflect the changing trend of the device state.

[0046] Here, when new device state information is generated (i.e., when new operation and maintenance text is added), a new semantic snapshot can be generated based on the new operation and maintenance text, and at this time, the comprehensive state vector needs to be iteratively updated with a time-decay weight update mechanism. The time-decay weight update mechanism will recalculate the weights of all semantic snapshots (including historical semantic snapshots and new semantic snapshots) according to the time information of the new semantic snapshot, and then generate an updated comprehensive state vector using the weighted aggregation algorithm again. This iterative updating method ensures that the comprehensive state vector can accurately reflect the latest state and changing trend of the device in real time.

[0047] In some embodiments, the weight of the historical semantic snapshot of the target power equipment is dynamically calculated by using an exponential decay function as follows:

[0048] Let the semantic snapshot of the target power equipment at time point t i be the vector s i ∈R d , d is the vector dimension, and the current timestamp is t current . Define the decay factor λ>0, which is used to control the decay rate. The larger λ is, the faster the historical snapshot weight decays. The weight of the semantic snapshot is defined as shown in equation (1):

[0049]

[0050] Among them, w i The weight of the i-th semantic snapshot is used to characterize the importance of that snapshot in temporal aggregation; t current This represents the current timestamp, i.e., the "current time" when performing time-series aggregation; t i This represents the timestamp corresponding to the i-th semantic snapshot, i.e., the recording time when the snapshot was generated.

[0051] By using a weighted aggregation algorithm, all semantic snapshots of the target power equipment are fused into a single vector according to the attenuation weights, generating the comprehensive state vector of the target power equipment as follows:

[0052] Suppose that the target power equipment has a total of n semantic snapshots, then the comprehensive state vector is as shown in equation (2):

[0053]

[0054] Where c represents the comprehensive state vector of the target power equipment, which is a vector reflecting the trend of equipment state changes after fusing all historical semantic snapshots.

[0055] When a new operation and maintenance text is added, the comprehensive state vector is iteratively updated based on the semantic snapshot of the new operation and maintenance text and the current comprehensive state vector, using a time decay weight update mechanism as follows:

[0056] Let the current integrated state vector be c. current The current total weight is W. current The corresponding timestamp t current Added semantic snapshots new The timestamp is t new , and t new >t current The updated integrated state vector is shown in equation (3) below:

[0057]

[0058] Among them, c new c represents the updated comprehensive state vector after the addition of a new semantic snapshot; current W represents the current summation state vector, i.e., the current existing comprehensive state vector; decayed The sum of historical weights represents the decayed weights, i.e., the sum of weights before the update after decay over time; Δt represents the time interval; W new This represents the updated total weight.

[0059] By adding semantic snapshots new The comprehensive state vector c new Total weights W new and timestamp t newupdating a current comprehensive state vector c current a weight sum W current and a timestamp t current .

[0060] S103, based on the comprehensive state vector, defining node types and edge types, and combining edge sequences connecting different types of nodes as meta-paths, constructing a substation heterogeneous graph.

[0061] In some embodiments, the specific process of constructing a substation heterogeneous graph based on a comprehensive state vector, defining node types and edge types, and combining edge sequences connecting different types of nodes as meta-paths is as follows:

[0062] The substation equipment is divided into nine node types, including transformers, circuit breakers, disconnectors, buses, mutual inductors, arresters, capacitors, high-voltage cabinets, and relay protection devices.

[0063] A multi-dimensional feature is added to each node, including device physical parameters, real-time monitoring data, and a comprehensive state vector.

[0064] Four types of edge types are defined to represent the relationship between devices, with differentiated weight calculation rules configured and edge weights dynamically adjusted through a time decay mechanism; the four types of edge types include:

[0065] The first type of edge type is used to represent the electrical connection relationship, and the conduction coefficient is dynamically calculated based on the impedance parameters between devices.

[0066] The second type of edge type is used to represent the mechanical locking relationship, and the influence of operation loss is quantified in combination with the mechanical life curve of the device.

[0067] The third type of edge type is used to represent the protection and protected relationship, and a three-level weight matrix is generated through the protection setting value coordination degree.

[0068] The fourth type of edge type is used to represent the physical subordinate relationship, and is directly bound to the device ontology and components.

[0069] Edge sequences connecting different node types are combined into meta-paths, and the consistency of the heterogeneous graph with the actual substation primary wiring diagram is verified through the analytic hierarchy process.

[0070] The node types, edge types, and meta-paths are integrated to form a substation heterogeneous graph.

[0071] Specifically, the transformer is a device that changes alternating voltage using the principle of electromagnetic induction, and plays a key role in voltage transformation and power transmission in power systems; the circuit breaker can close, carry and open the current under normal loop conditions, and close, carry and open the current under abnormal loop conditions within a specified time, and is an important control and protection device in power systems; the disconnecting switch is mainly used for isolating power sources, switching operations, pulling and closing no-current or small-current circuits, etc.; the bus is a conductor that collects and distributes electric energy in a substation; the mutual inductor is divided into voltage mutual inductor and current mutual inductor, which is used to transform high voltage or large current into standard low voltage or standard small current for measurement, protection and control devices; the arrester is an electrical device used to protect electrical equipment from high transient overvoltage and limit the current flow time, usually connected between the conductor and the ground, when overvoltage occurs, the arrester quickly acts to introduce overvoltage into the ground, thereby protecting the safety of electrical equipment; the capacitor is a component that can store electric charge, in the substation, the capacitor is mainly used for reactive power compensation, by adjusting the input and cut-off of the capacitor, the power factor of the power grid can be improved, the power quality can be improved, and the line loss can be reduced; the high-voltage cabinet is a complete power distribution device assembled by primary equipment and secondary equipment according to a certain scheme, which undertakes the important tasks of receiving and distributing electric energy, controlling the on-off of the circuit, etc. in the substation, which usually contains circuit breakers, disconnecting switches, mutual inductors and other devices, with compact structure and high functional integration; the relay protection device is an automatic device that selectively removes faulty components when the power system fails or operates abnormally, to ensure the continued operation of non-faulty parts.

[0072] Here, the multi-dimensional features include device physical parameters, real-time monitoring data and comprehensive state vectors, and adding multi-dimensional features to each node can be used to enhance the richness of node information. Device physical parameters such as the rated capacity of the transformer and the rated current of the circuit breaker are basic attributes of the device itself; real-time monitoring data is the device operating parameters collected in real time by sensors and other devices, such as the oil temperature of the transformer and the opening and closing time of the circuit breaker; the comprehensive state vector reflects the overall state change trend of the device.

[0073] Here, the definition of edge type is used to characterize the relationship between devices, mainly including four types of edge types. Specifically, the first type of edge type is used to represent the electrical connection relationship, and the conduction coefficient is dynamically calculated based on the impedance parameter between devices; the electrical connection relationship is the most basic connection mode between substation devices, the impedance parameter reflects the degree of hindering of current transmission between devices, and the conduction coefficient quantifies the strength of the electrical connection, and will be dynamically adjusted with time. The second type of edge type is used to represent the mechanical interlocking relationship, and the operation loss influence is quantified by combining the device mechanical life curve; the mechanical interlocking relationship ensures the safety and sequence of the device in the operation process, and the device mechanical life curve describes the process of gradual wear of the mechanical parts of the device with the increase of use time and operation times, and the strength of the mechanical interlocking relationship can be more accurately evaluated by quantifying the operation loss influence. The third type of edge type is used to represent the protection and protected relationship, and a three-level weight matrix is generated by the protection setting value coordination degree; the protection and protected relationship is an important relationship between the relay protection device and the protected device, and the protection setting value coordination degree reflects the accuracy and reliability of the protection device action, and the three-level weight matrix quantifies this relationship according to different coordination degrees. The fourth type of edge type is used to represent the physical subordinate relationship, and the device ontology and components are directly bound; the physical subordinate relationship clearly defines the ownership relationship between the device ontology and its attached components, and this relationship is relatively stable, but it will also change with the use and maintenance of the device.

[0074] It can be understood that the meta-path is a sequence of edges connecting different types of nodes in the heterogeneous graph, which describes a specific relationship pattern between nodes. The consistency of the heterogeneous graph with the actual substation primary wiring diagram can be verified by the analytic hierarchy process, which is a method of decomposing complex problems into multiple levels and making decisions by comparing the importance of elements at each level. Through this method, it can be ensured that the constructed heterogeneous graph can accurately reflect the connection relationship and operating state between devices in the actual substation.

[0075] Finally, by integrating node types, edge types and meta-paths, a substation heterogeneous graph can be formed, which is a graph structure that can represent multiple types of nodes and multiple types of edges at the same time, and it can more comprehensively and intuitively display the complex relationship between substation devices.

[0076] In S104, based on the substation heterogeneous graph, the node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta-path, and the information transmission weight is calculated in combination with the graph attention.

[0077] In some embodiments, the specific process of aggregating node neighbor information based on the substation heterogeneous graph, configuring an independent parameter matrix for each relationship on the meta-path, and calculating the information transmission weight in combination with the graph attention is as follows:

[0078] According to the meta-path defined in the transformer substation heterogeneous graph, neighbor nodes having direct or indirect association with the target node are screened out;

[0079] A corresponding parameter matrix is configured for each of the four types of edges. For each type of neighbor node of the target node, a multi-dimensional feature of the neighbor node is multiplied by the parameter matrix of the corresponding edge type through linear transformation to generate a neighbor feature vector.

[0080] The similarity between the target node feature and each neighbor feature vector is calculated, a learnable attention function is used to generate an unnormalized weight, and a Softmax function is used to normalize the neighbor set to obtain a dynamic weight coefficient.

[0081] The normalized weight coefficient and the corresponding neighbor feature vector are weighted and summed to output a new state vector of the target node, completing a single round of information transmission.

[0082] It can be understood that direct association refers to nodes directly connected by an edge, and indirect association refers to nodes indirectly connected through multiple edges and intermediate nodes. For example, if the target node is a transformer, through the meta-path of electrical connection relationship, a circuit breaker node directly connected to the transformer can be found, and a bus node indirectly connected through the circuit breaker node can also be found as a neighbor node.

[0083] The parameter matrix is a tool for linear transformation of node features. Different types of edge relationships have different semantics and feature interaction modes, so independent parameter matrices are needed to capture these differences. For each type of neighbor node of the target node, a multi-dimensional feature of the neighbor node can be multiplied by the parameter matrix of the corresponding edge type through linear transformation to generate a neighbor feature vector. Taking the edge type of electrical connection relationship as an example, the corresponding parameter matrix will perform specific linear combination on the multi-dimensional features of the neighbor nodes (such as circuit breaker nodes) according to the characteristics of electrical connection, highlighting the feature information related to electrical connection, thereby generating a more targeted neighbor feature vector.

[0084] It can be understood that the similarity reflects the closeness of the target node and the neighbor node in the feature space. The higher the similarity, the more similar the two nodes are in state or attribute, and the more important the information transmission is. The learnable attention function is a function that can automatically adjust parameters according to data features. It can dynamically calculate the weight value corresponding to the similarity according to the features of the target node and the neighbor node. The Softmax function can map the unnormalized weight value to the interval [0, 1] and ensure that the sum of the normalized weight coefficients of all neighbor nodes is 1, thereby obtaining a dynamic weight coefficient. These dynamic weight coefficients reflect the contribution degree of different neighbor nodes to the information update of the target node, and will be dynamically adjusted with the changes of node features and edge relationships.

[0085] Further, the normalized weight coefficients are weighted and summed with the corresponding neighbor feature vectors to output the new state vector of the target node, i.e. a single round of information transmission is completed. Here, the weighted sum process is actually a process of fusing neighbor node information. By weighting the feature vector of each neighbor node according to its corresponding weight coefficient, then adding the results, the new state vector of the target node obtained synthesizes the important information of the target node itself and the neighbor nodes. This new state vector not only contains the original state information of the target node, but also incorporates the influence of the surrounding neighbor nodes, and can more comprehensively and accurately reflect the state of the target node in the current graph structure, providing more abundant information for subsequent operation and decision analysis.

[0086] In this way, by continuously iterating the above process, the state vectors of all nodes in the graph can be gradually updated and optimized, thereby realizing dynamic monitoring and evaluation of the state of the substation equipment.

[0087] S105, through several rounds of information aggregation iteration, expand the node perception range until the L2 norm change rate of the node state vector and the global risk entropy value fluctuation meet the convergence condition.

[0088] In some embodiments, the specific process of expanding the node perception range through several rounds of information aggregation iteration until the L2 norm change rate of the node state vector and the global risk entropy value fluctuation meet the convergence condition is as follows:

[0089] Initialize the iteration round, and in each iteration round, based on the current state vector of the node in the substation heterogeneous graph, aggregate the information of the neighbor nodes with a distance greater than that of the last iteration, and update the node state vector;

[0090] After each iteration round, calculate the ratio of the L2 norm difference between the current round node state vector and the last round node state vector to the L2 norm of the last round; at the same time, calculate the information entropy based on the distribution of all node state vectors, and determine the corresponding fluctuation amplitude;

[0091] When the L2 norm change rate of the node state vector is less than the first threshold value for 3 consecutive iteration rounds, and the global risk entropy value fluctuation amplitude does not exceed the second threshold value of the average fluctuation amplitude of the previous 5 rounds, it is determined that the iteration converges.

[0092] It can be understood that, after the initialization of the iteration round, at the beginning of each iteration round, information aggregation is performed based on the current state vector of the node in the substation heterogeneous graph. In this process, the information of the neighbor nodes with a distance greater than that of the last iteration round is aggregated. Here, the "distance" does not refer to the physical space distance, but is based on the path length or specific relationship metric between nodes in the graph structure. For example, under the definition of the meta-path based on the electrical connection relationship, the distance may reflect the complexity of the current transmission path. By aggregating the information of the more distant neighbor nodes, the node can obtain the influence and association information of the devices in a wider range, thereby gradually expanding its own perception range. Subsequently, the node state vector is updated according to the aggregated information, so that the node state can reflect the more comprehensive device operation condition.

[0093] Here, after each iteration round, two key indicators are calculated: first, the ratio of the L2 norm difference between the node state vector of the current round and the node state vector of the last round to the L2 norm of the last round is calculated, which is the L2 norm change rate of the node state vector. Here, the L2 norm is the square root of the sum of squares of the elements of the vector, which can measure the length of the vector in space; the L2 norm change rate reflects the change degree of the node state vector in adjacent two iteration rounds. If the change rate is large, it means that the node state has changed significantly in the iteration process, which may mean that the newly aggregated information has a greater impact on the node state; on the contrary, if the change rate is small, it means that the node state tends to be stable.

[0094] At the same time, the information entropy based on the distribution of all node state vectors needs to be calculated, and the corresponding fluctuation amplitude needs to be determined. Among them, the information entropy is an index for measuring the uncertainty of the system, and the distribution of all node state vectors in the substation heterogeneous graph reflects the complexity and uncertainty of the entire substation device state. The greater the information entropy, the higher the uncertainty of the device state, and there may be more potential risk factors; the smaller the information entropy, the more stable and certain the device state. By calculating the fluctuation amplitude of the information entropy, the change of the uncertainty of the entire system state can be observed.

[0095] Specifically, when a certain convergence condition is met, it is determined that the iteration converges. In the present disclosure, the convergence condition is set as that the rate of change of the L2 norm of the node state vector is less than a first threshold value for three consecutive iterations, and the fluctuation range of the global risk entropy value does not exceed a second threshold value of the average fluctuation range of the previous five rounds. The first threshold value can be set to 0.0001 or 0.0002, which is a very small number, meaning that when the change of the node state vector is very small, it can be considered that the node state basically no longer changes significantly. The second threshold value can be set to 10% or 15%, which limits the fluctuation of the global risk entropy value. The average fluctuation range of the previous five rounds reflects the uncertainty change trend of the system in the early iteration process. When the fluctuation range of the global risk entropy value of the current round does not exceed 10% of the average fluctuation range, it means that the uncertainty change of the system tends to be stable, and the state of the entire substation equipment reaches a relatively stable state.

[0096] In this way, by means of the several rounds of information aggregation iteration and the determination of convergence according to the L2 norm change rate and the global risk entropy value fluctuation, the state of the substation equipment can be more accurately evaluated on the basis of fully considering the complex relationship between the equipment and the extensive information.

[0097] S106, quantifying the risk of the node state vector after iteration to obtain the risk index of the equipment or region, and based on the attention mechanism, tracing the key path and abnormal state to generate an interpretable risk traceability report.

[0098] In some embodiments, the specific process of quantifying the risk of the node state vector after iteration to obtain the risk index of the equipment or region, and based on the attention mechanism, tracing the key path and abnormal state to generate an interpretable risk traceability report is as follows:

[0099] The node state vector after iteration converges is input into a fully connected neural network layer, and the risk probability value of each device node is output by a Softmax classification function;

[0100] The risk probability values of all device nodes in the specified region are weighted and averaged, and the weight is determined by the criticality of the equipment in the power grid topology to generate a regional risk index;

[0101] Extract the attention weight matrix recorded in each layer of the heterogeneous graph of the substation to locate the key neighbor nodes and associated meta-path whose risk contribution value to the target node exceeds a preset threshold;

[0102] Trace back to the source device node along the key meta-path, and associate the original operation and maintenance text records corresponding to the historical semantic snapshot of the corresponding node;

[0103] Based on the backtracking results, a risk transmission path topology graph is constructed, key node contribution degrees and corresponding abnormal description texts are marked, and a three-level diagnosis report including risk causes, transmission paths, and impact ranges is generated.

[0104] Specifically, the fully connected neural network layer is a deep learning structure in which each neuron is connected to all neurons of the previous layer, and can perform complex nonlinear transformation and feature extraction on the input node state vector. The node state vector as input contains comprehensive state information integrated by the device after multiple rounds of information aggregation iteration, which includes physical parameters, running data, and association features of surrounding devices, etc. Through the processing of the fully connected neural network layer, potential patterns and features related to risks in these information can be further mined, and then the risk probability value of each device node is output using the Softmax classification function. Here, the Softmax classification function is a commonly used activation function in multi-classification problems, which can convert the output of the neural network layer into a probability distribution, so that each device node corresponds to a risk probability value in the interval [0, 1], and the sum of the risk probability values of all device nodes is 1. These risk probability values intuitively reflect the possibility of each device in the current state to fail or have a risk event.

[0105] Then, in the process of weighted averaging the risk probability values of all device nodes in the specified area to generate the regional risk index, the weight is determined by the criticality of the device in the power grid topology. The power grid topology is a structure that describes the connection relationship and layout of devices in the power grid, and the criticality of the device in the power grid topology reflects the importance of the device to the stable operation of the entire power grid. For example, a transformer or circuit breaker located at a power grid hub may cause a larger range of power outages if it fails, so it can have a higher criticality. By assigning weights according to the criticality of the device, the overall risk level in the region can be more accurately measured, so that the regional risk index can take into account the risk probability of each device and their important position in the power grid.

[0106] Further, the attention weight matrix of the substation heterogeneous graph is calculated by the attention mechanism in the previous information aggregation iteration process. The attention mechanism is a mechanism that can automatically learn and assign different information importance, and in the substation heterogeneous graph, it can assign an attention weight to each neighbor node according to the association relationship and feature similarity between nodes, indicating the contribution degree of the neighbor node to the information update of the target node. The attention weight matrix of each layer records the attention allocation under different levels and different relationships. By locating the key neighbor nodes and associated meta-paths whose risk contribution values to the target node exceed a preset threshold, other devices that have an important impact on the device risk and their specific relationship patterns can be found.

[0107] It can be understood that the historical semantic snapshot is the result of semantic encapsulation and storage of state information of the device at different time points, which contains rich semantic information such as running parameters, fault records, operation logs and the like of the device. The original operation and maintenance text record is the actual operation and observation recorded by the operation and maintenance personnel during the operation of the device, which can provide more detailed and more intuitive device operation background information.

[0108] Specifically, the risk transmission path topology graph shows the process of risk transmission from the source device node to the target device node along the key path in a graphical manner. In the topology graph, the key node contribution degree and the corresponding abnormal description text are marked. The key node contribution degree reflects the influence degree of each key node on the formation of the final risk in the risk transmission process, which can be quantitatively calculated by attention weight, risk transmission coefficient and the like. The corresponding abnormal description text is a detailed description of the abnormal state and event of the key node in the risk transmission process, which is derived from the historical semantic snapshot and the original operation and maintenance text record. In this way, a three-level diagnosis report containing risk causes, transmission path and influence range can be generated. Among them, the risk causes part analyzes the root causes of the risk in detail, including device faults, external environmental interference and the like; the transmission path part clearly shows the propagation process and key nodes of the risk in the power grid; and the influence range part evaluates the influence degree of the risk on the surrounding devices and the entire power grid, providing comprehensive, accurate and interpretable basis for the operation and maintenance personnel to develop targeted risk response measures.

[0109] The technical scheme of the present application has at least the following advantages and beneficial effects:

[0110] Through deep semantic vectorization processing of the operation and maintenance text data, the fuzzy and precursor descriptions such as "slight oil seepage" and "abnormal noise" are converted into semantic vectors, breaking through the limitations of existing online monitoring technology on fuzzy information and offline analysis on deep understanding of real text semantics, and accurately capturing potential risk signals contained in the text, providing a more rich and accurate information source for risk identification.

[0111] The time-decaying aggregation algorithm is adopted to perform time sequence aggregation on the semantic snapshot to generate a comprehensive state vector, which can effectively reflect the change trend of the state of the target power device with time, overcome the shortcomings of the prior art on static analysis of the device state, and make the risk identification more dynamic and forward-looking, facilitating timely discovery of the signs of device state deterioration.

[0112] The substation heterogeneous graph is constructed based on the comprehensive state vector, node types and edge types are defined, and the edge sequence is taken as a meta path, the diversity of the types of devices in the substation and the heterogeneity of the connection relationship (electrical connection, mechanical locking, protection logic, etc.) are fully considered, the complex multi-dimensional interaction between devices can be more truly and comprehensively modeled, and key support is provided for identifying systematic hidden risks.

[0113] On the basis of the heterogeneous graph, the node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta path, the information transmission weight is calculated in combination with the graph attention, the node perception range is expanded through multiple rounds of information aggregation iteration until the convergence condition is met, the influence weight between devices can be more accurately calculated, the systematic hidden risk hazards formed by the combination, conduction and aggregation of multiple seemingly independent "low-risk" events through the complex device network can be effectively identified, the limitation that the risk cognition of the prior art stays at the "point" level is broken through, and "surface" systematic evaluation is achieved.

[0114] The risk quantization output is performed on the iterated node state vector, the risk index of the device or the area is obtained, the key path and the abnormal state are traced back based on the attention mechanism, and an interpretable risk traceability report is generated, which can not only clearly indicate the specific position and degree of the risk, but also clearly show the formation path and the key influencing factors of the risk, is helpful for workers to quickly locate problems and develop targeted operation and maintenance strategies, changes the safety operation and maintenance from passive response to active prevention, and significantly improves the guarantee capability of the safe and stable operation of the substation.

[0115] Based on the same inventive concept, corresponding to any of the above embodiments, with reference to Figure 2 The application provides a substation hidden risk identification system for realizing the substation hidden risk identification method.

[0116] The semantic processing module is used for collecting operation and maintenance text data of the substation, converting fuzzy description and precursor description into semantic vectors through a pre-trained deep language model, and generating semantic snapshots of target power devices at corresponding time points.

[0117] The state aggregation module is connected with the semantic processing module, performs time sequence weighted fusion on the semantic snapshots through an aggregation algorithm with time attenuation, and generates a comprehensive state vector reflecting the state change trend.

[0118] The heterogeneous graph construction module is connected with the state aggregation module, is used for defining six types of nodes and four types of edges, constructing a substation heterogeneous graph containing multi-dimensional features of devices and dynamic edge weights based on the comprehensive state vector, and verifying the topological consistency through meta path combination and analytic hierarchy process.

[0119] The graph computing engine module is connected to the heterogeneous graph construction module, and is configured to filter neighbor nodes based on a meta-path, configure independent parameter matrices for four types of edges, calculate dynamic weights of neighbor information transmission by combining a graph attention mechanism, perform multiple rounds of information aggregation iteration, and determine convergence by L2 norm change rate and global risk entropy value fluctuation;

[0120] The risk quantification and tracing module is connected to the graph computing engine module, and is configured to input the node state vector after iteration into a fully connected neural network, output a device risk probability value, aggregate regional device risk probability values to generate a regional risk index, trace key meta-paths and original operation and maintenance texts based on attention weights, and construct a three-level diagnosis report.

[0121] Optionally, the semantic processing module is specifically configured to:

[0122] The operation and maintenance text data of the substation includes work tickets, device inspection records, defect reports, and operation logs.

[0123] The semantic coding of the operation and maintenance text is performed by using a pre-trained deep language model, and the ambiguity description in the operation and maintenance text is converted into a high-dimensional vector.

[0124] The floating-point number vector generated by the operation and maintenance text record coding constitutes a semantic snapshot of the target power device at a corresponding time point, and the semantic snapshot quantitatively represents the risk degree difference of the text description in the vector space.

[0125] Optionally, the state aggregation module is specifically configured to:

[0126] An exponential decay function is used to dynamically calculate the weights of the historical semantic snapshots of the target power device; wherein the weight of the semantic snapshot generated at the current time point is the highest, and the weights of the historical semantic snapshots exponentially decay with the increase of the time interval.

[0127] All semantic snapshots of the target power device are fused into a single vector by a weighted aggregation algorithm according to the decay weights, and a comprehensive state vector of the target power device is generated.

[0128] When new operation and maintenance texts are added, the comprehensive state vector is iteratively updated based on the semantic snapshot of the new operation and maintenance texts and the current comprehensive state vector in a time decay weight update mechanism.

[0129] Optionally, the heterogeneous graph construction module is specifically configured to:

[0130] The substation devices are divided into nine node types, including transformers, circuit breakers, disconnectors, busbars, mutual inductors, arresters, capacitors, high-voltage cabinets, and relay protection devices.

[0131] adding a multi-dimensional feature to each node, the multi-dimensional feature including a device physical parameter, real-time monitoring data, and a comprehensive state vector;

[0132] defining four types of edge types representing relationships between devices, configuring differentiated weight calculation rules, and dynamically adjusting edge weights through a time decay mechanism; wherein the four types of edge types include:

[0133] The first type of edge is used to represent the electrical connection relationship, and the conduction coefficient is dynamically calculated based on the impedance parameter between devices;

[0134] The second type of edge is used to represent the mechanical locking relationship, and the influence of operation loss is quantified by combining the mechanical life curve of the device;

[0135] The third type of edge is used to represent the protection and protected relationship, and a three-level weight matrix is generated through the protection setting value matching degree;

[0136] The fourth type of edge is used to represent the physical subordinate relationship, and is directly bound to the device ontology and components;

[0137] The edge sequence connecting different node types is combined into a meta-path, and the consistency of the heterogeneous graph with the actual substation primary wiring diagram is verified by the analytic hierarchy process;

[0138] The node type, the edge type and the meta-path are integrated to form a substation heterogeneous graph.

[0139] Optionally, the graph computing engine module is specifically configured to:

[0140] According to the meta-path defined in the substation heterogeneous graph, neighbor nodes having a direct or indirect association with a target node are screened;

[0141] For each type of neighbor node of the target node, the multi-dimensional features of the neighbor node and the parameter matrix of the corresponding edge type are multiplied through linear transformation to generate a neighbor feature vector;

[0142] The similarity between the target node feature and each neighbor feature vector is calculated, an unnormalized weight is generated through a learnable attention function, and a Softmax function is used to normalize on the neighbor set to obtain a dynamic weight coefficient;

[0143] The normalized weight coefficient and the corresponding neighbor feature vector are weighted and summed to output a new state vector of the target node, and a single round of information transmission is completed.

[0144] Optionally, the graph computing engine module is specifically configured to:

[0145] Initialize the iteration round, in each iteration round, based on the current state vector of the nodes in the transformer station heterogeneous graph, aggregate the information of the neighbor nodes with a distance greater than the last iteration, and update the node state vector;

[0146] After each iteration round, calculate the ratio of the L2 norm difference between the current round node state vector and the last round node state vector to the L2 norm of the last round; and calculate the information entropy based on the distribution of all node state vectors, and determine the corresponding fluctuation amplitude;

[0147] When the L2 norm change rate of the node state vector is less than the first threshold value for three consecutive iteration rounds, and the global risk entropy value fluctuation amplitude does not exceed the second threshold value of the average fluctuation amplitude of the previous five rounds, it is determined that the iteration converges.

[0148] Optionally, the risk quantification and tracing module is specifically configured to:

[0149] Input the node state vector after iteration convergence into a fully connected neural network layer, and output the risk probability value of each device node through a Softmax classification function;

[0150] Weighted average the risk probability values of all device nodes in the specified area, and the weight is determined by the criticality of the device in the power grid topology, to generate a regional risk index;

[0151] Extract each layer of attention weight matrix recorded in the transformer station heterogeneous graph, locate the key neighbor nodes and associated meta-path whose risk contribution value to the target node exceeds a preset threshold;

[0152] Backtrack to the source device node along the key meta-path, and associate the original operation and maintenance text record corresponding to the historical semantic snapshot of the corresponding node;

[0153] Based on the backtracking result, construct a risk transmission path topology graph, label the key node contribution and the corresponding abnormal description text, and generate a three-level diagnosis report containing risk causes, transmission path and impact range.

[0154] Based on the same inventive concept, the present application provides an electronic device, comprising a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to make the electronic device execute the transformer station implicit risk identification method of the embodiment.

[0155] Optionally, the above-mentioned electronic device can be a server.

[0156] In addition, the present embodiment also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the transformer station implicit risk identification method of the embodiment.

[0157] It is appreciated that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0158] The method steps in the embodiments of the present application can be realized by hardware or by the processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically-erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0159] In the embodiments described above, all or some of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded into and executed by a computer, all or some of the procedures or functions according to the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in or transmitted from a storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

Claims

1. A substation implicit risk identification method, characterized in that, The method comprises the following steps: Deep semantic vectorization processing is performed on the operation and maintenance text data of the substation, and fuzzy description and precursor description are converted into semantic vectors to form semantic snapshots of the target power equipment at corresponding time points; Through an aggregation algorithm with time decay, the semantic snapshots of the target power equipment at the corresponding time points are time-series aggregated to generate a comprehensive state vector reflecting the state change trend; Based on the comprehensive state vector, node types and edge types are defined, and edge sequences connecting different types of nodes are taken as meta-paths to construct a substation heterogeneous graph; Based on the substation heterogeneous graph, node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta-path, and information transmission weights are calculated in combination with graph attention; Through several rounds of information aggregation iterations, the node perception range is expanded until the L2 norm change rate of the node state vector and the global risk entropy value fluctuation meet the convergence condition; The risk of the device or the region is quantified by outputting the iterated node state vector to obtain a risk index, and based on the attention mechanism, a key path and an abnormal state are traced back to generate an interpretable risk traceability report.

2. The substation implicit risk identification method of claim 1, wherein, The specific process of the deep semantic vectorization processing on the operation and maintenance text data of the substation, and the conversion of the fuzzy description and the precursor description into semantic vectors to form semantic snapshots of the target power equipment at corresponding time points is as follows: Collecting operation and maintenance text data of the substation, the operation and maintenance text data including work tickets, equipment inspection records, defect reports and operation logs; Performing semantic coding on the operation and maintenance text through a pre-trained deep language model to convert the fuzzy description in the operation and maintenance text into a high-dimensional vector; The floating-point number vector generated by the coding of the operation and maintenance text records constitutes the semantic snapshots of the target power equipment at the corresponding time points, and the semantic snapshots quantitatively represent the risk degree difference of the text description in the vector space.

3. The substation implicit risk identification method of claim 1, wherein, The specific process of the aggregation algorithm with time decay for time-series aggregation of the semantic snapshots of the target power equipment at the corresponding time points to generate a comprehensive state vector reflecting the state change trend is as follows: An exponential decay function is used to dynamically calculate the weights of the historical semantic snapshots of the target power equipment; wherein the weight of the semantic snapshot generated at the current time point is the highest, and the weight of the historical semantic snapshot decreases exponentially with the increase of the time interval; Through a weighted aggregation algorithm, all semantic snapshots of the target power equipment are fused into a single vector according to the decay weights to generate a comprehensive state vector of the target power equipment; When new operation and maintenance text is added, the semantic snapshot based on the new operation and maintenance text and the current comprehensive state vector are iteratively updated with a time decay weight update mechanism to update the comprehensive state vector.

4. The substation implicit risk identification method of claim 1, wherein, The specific process of constructing a substation heterogeneous graph based on the comprehensive state vector, defining node types and edge types, and taking edge sequences connecting different types of nodes as meta-paths is as follows: The substation equipment is divided into nine node types, including transformers, circuit breakers, disconnectors, busbars, mutual inductors, arresters, capacitors, high-voltage cabinets and relay protection devices; Multi-dimensional features are added to each node, including device physical parameters, real-time monitoring data and comprehensive state vectors; Four types of edge types are defined to represent the relationship between devices, and the differentiated weight calculation rules are configured, and the edge weight is dynamically adjusted through the time decay mechanism. Among them, the four types of edges include: The first type of edge is used to represent the electrical connection relationship, and the conduction coefficient is dynamically calculated based on the impedance parameters between devices; The second type of edge is used to represent the mechanical locking relationship, and the operation loss influence is quantified by combining the mechanical life curve of the device; The third type of edge is used to represent the protection and protected relationship, and the three-level weight matrix is generated through the protection setting value matching degree; The fourth type of edge is used to represent the physical subordinate relationship, and is directly bound to the device ontology and component; The edge sequence connecting different node types is combined into a meta-path, and the consistency of the heterogeneous graph with the actual substation primary wiring diagram is verified by the analytic hierarchy process; The node type, edge type and meta-path are integrated to form a substation heterogeneous graph.

5. The substation implicit risk identification method of claim 4, wherein, Based on the substation heterogeneous graph, the node neighbor information is aggregated, an independent parameter matrix is configured for each relationship on the meta-path, and the specific process of combining graph attention to calculate information transmission weight is as follows: According to the meta-path defined in the substation heterogeneous graph, the neighbor nodes having direct or indirect association with the target node are screened; For the four types of edges, the corresponding parameter matrix is configured, and for each type of neighbor node of the target node, the multi-dimensional features of the neighbor node are multiplied with the parameter matrix of the corresponding edge type through linear transformation to generate a neighbor feature vector; The similarity between the target node features and each neighbor feature vector is calculated, the unnormalized weight is generated through a learnable attention function, and the normalized weight is obtained by normalizing the neighbor set using the Softmax function to obtain the dynamic weight coefficient; The normalized weight coefficient and the corresponding neighbor feature vector are weighted and summed to output the new state vector of the target node, and a single round of information transmission is completed.

6. The substation implicit risk identification method of claim 1, wherein, The specific process of expanding the node perception range through several rounds of information aggregation iteration until the L2 norm change rate of the node state vector and the global risk entropy value fluctuation meet the convergence condition is as follows: Initialize the iteration round, in each iteration, based on the current state vector of the node in the substation heterogeneous graph, aggregate the information of the neighbor nodes with a distance greater than the neighbor nodes in the last iteration, and update the node state vector; After each iteration, the ratio of the L2 norm difference value between the current round node state vector and the last round node state vector to the L2 norm of the last round is calculated; at the same time, the information entropy based on the distribution of all node state vectors is calculated, and the corresponding fluctuation amplitude is determined; When the L2 norm change rate of the node state vector is less than the first threshold value for 3 consecutive iterations, and the global risk entropy value fluctuation amplitude does not exceed the second threshold value of the average fluctuation amplitude of the previous 5 rounds, it is determined that the iteration converges.

7. The substation implicit risk identification method of claim 1, wherein, The specific process of risk quantization output of the node state vector after iteration, obtaining the risk index of the device or region, and generating an interpretable risk traceability report based on the attention mechanism backtracking key path and abnormal state is as follows: The node state vector after iteration convergence is input into the fully connected neural network layer, and the risk probability value of each device node is output through the Softmax classification function; The risk probability values of all device nodes in the specified area are weighted and averaged, and the weights are determined by the criticality of the device in the power grid topology, to generate a regional risk index; Extract the attention weight matrix recorded in the heterogeneous graph of the substation, locate the key neighbor nodes and associated meta-path whose risk contribution value to the target node exceeds the preset threshold; Backtrack along the key meta-path to the source device node, associate the original operation and maintenance text records corresponding to the historical semantic snapshot of the corresponding node; Based on the backtracking result, a risk transmission path topology graph is constructed, the contribution of key nodes and the corresponding abnormal description text are labeled, and a three-level diagnosis report containing risk causes, transmission path and impact range is generated.

8. A substation implicit risk identification system for implementing the substation implicit risk identification method of any one of claims 1-7, characterized in that, It comprises: A semantic processing module is used to collect operation and maintenance text data of a substation, convert fuzzy description and precursor description into semantic vectors through a pre-trained deep language model, and generate semantic snapshots of target power equipment at corresponding time points; A state aggregation module is connected to the semantic processing module and performs time-weighted fusion on the semantic snapshots through an aggregation algorithm with time decay to generate a comprehensive state vector reflecting the state change trend; A heterogeneous graph construction module is connected to the state aggregation module and is used to define six types of nodes and four types of edges, construct a substation heterogeneous graph containing device multi-dimensional features and dynamic edge weights based on the comprehensive state vector, and verify the topology consistency through meta-path combination and analytic hierarchy process; A graph computing engine module is connected to the heterogeneous graph construction module and is used to filter neighbor nodes based on meta-path, configure independent parameter matrices for four types of edges, calculate the dynamic weight of neighbor information transmission combined with graph attention mechanism, perform multiple rounds of information aggregation iteration, and determine convergence through L2 norm change rate and global risk entropy value fluctuation; A risk quantification and tracing module is connected to the graph computing engine module and is used to input the node state vector after iteration into a fully connected neural network, output the device risk probability value, weight and aggregate the regional device risk probability value to generate a regional risk index, backtrack the key meta-path and original operation and maintenance text based on the attention weight, and construct a three-level diagnosis report.

9. An electronic device, comprising: An electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to make the electronic device execute the substation implicit risk identification method in any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the substation implicit risk identification method in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for verifying and evaluating topological data of 220kV transformer substation

    CN117390003A

  • Transformer substation misoperation mode mining method for large-scale historical data

    CN119885095A

  • Intelligent operation and maintenance method fusing multi-modal data and active learning

    CN120198106A

Cited By

  • Early warning methods for cascading failure risks

    CN122413192A