Intelligent substation state evaluation method and system based on digital twinning
By constructing a heterogeneous knowledge graph and a relation-aware graph neural network, the problem of equipment and system separation in the status assessment of smart substations is solved, realizing panoramic status perception and risk warning, and improving the predictability and interpretability of the assessment.
Patent Information
- Application Number
- CN202511856166.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for assessing the condition of smart substations suffer from a disconnect between the health status of equipment and the operating status of the power grid. They also struggle to integrate heterogeneous data from multiple sources and lack a detailed characterization of the complex relationships between electrical connections and functional couplings between equipment. Consequently, minor equipment defects are difficult to detect and trace in a timely manner under power grid disturbances.
The intelligent substation status assessment method based on digital twins constructs a heterogeneous knowledge graph, combines ontology definition, asset data, topology data and historical slow-changing data, performs time-series feature encoding of graph nodes and message passing and node embedding of relation-aware graph neural networks, realizes equipment risk prediction and comprehensive substation assessment, and performs risk path tracing.
It enables unified assessment of equipment health and system operating status, possesses interpretable panoramic status perception and risk warning, and improves the predictability and interpretability of the assessment.
Smart Images

Figure CN121599486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of condition assessment technology, and more specifically, to a method and system for condition assessment of smart substations based on digital twins. Background Technology
[0002] As the construction of smart grids continues to deepen, the accurate assessment and prediction of the operating status of smart substations, as key nodes in the power system, is of great significance for improving grid reliability and achieving proactive operation and maintenance. Under the new power system context, the high proportion of renewable energy integration and diverse load fluctuations have made the substation operating environment increasingly complex. Traditional periodic maintenance and single-equipment monitoring models are no longer sufficient to meet the needs of real-time perception, dynamic assessment, and risk tracing. Digital twin technology, by constructing a virtual model highly synchronized with the physical substation, provides a new path for achieving state mapping, behavior inference, and intelligent decision-making, and is gradually becoming an important development direction for smart substation state assessment.
[0003] In existing technologies, condition assessment research based on digital twins mostly focuses on independent analysis at the equipment or system level, and has not yet formed an integrated assessment system that integrates multi-source heterogeneous data and takes into account spatiotemporal characteristics and structural correlations. At the data level, existing technologies often rely on single types of monitoring data, failing to effectively integrate equipment asset information, topological connectivity, historical operation records, and real-time data streams, resulting in limited assessment dimensions and insufficient information utilization. At the model level, most assessment methods still use traditional machine learning or shallow neural networks, which are insufficient in characterizing complex relationships such as electrical connections and functional couplings between equipment, and lack the ability to structurally model heterogeneous knowledge. Of particular note is the common problem of disconnect between equipment health status and grid operation status in existing assessment systems: on the one hand, equipment health assessments often focus on inherent attributes such as insulation aging and mechanical wear, ignoring their functional impact on the dynamic operation of the grid; on the other hand, grid operation status assessments often focus on system-level indicators such as voltage stability and power flow distribution, making it difficult to trace back to the health degradation process of specific equipment. This separation of "health" and "stability" assessments makes it difficult to detect and trace the cascading risks that minor equipment defects may trigger under specific grid disturbances in a timely manner, thus limiting the predictability and interpretability of condition assessments.
[0004] Therefore, there is an urgent need for an optimized method and system for assessing the condition of smart substations based on digital twins. Summary of the Invention
[0005] This application is made in order to solve the above-mentioned technical problems.
[0006] According to one aspect of this application, a method for assessing the status of a smart substation based on digital twins is provided, which includes: constructing a heterogeneous knowledge graph based on ontology definition, asset data, topology data and historical slowly changing data; Based on heterogeneous knowledge graphs and time window lengths, graph node temporal feature encoding is performed on real-time rapidly changing data streams to obtain an initial node feature matrix. Based on heterogeneous knowledge graphs, message passing and node embedding based on relation-aware graph neural networks are performed on the initial node feature matrix to obtain the final node embedding matrix. Equipment risk prediction is performed on the final node embedding matrix to obtain the equipment risk score vector; A comprehensive evaluation of the substation is performed on the final node embedding matrix to obtain a comprehensive substation score. Risk path tracing is performed on equipment risk scoring vectors and heterogeneous knowledge graphs to obtain an explanatory subgraph.
[0007] According to another aspect of this application, a digital twin-based intelligent substation status assessment system is provided, which includes: a heterogeneous knowledge graph construction module for constructing a heterogeneous knowledge graph based on ontology definition, asset data, topology data and historical slowly changing data; The graph node temporal feature encoding module is used to encode the graph node temporal features of real-time fast-changing data streams based on heterogeneous knowledge graphs and time window lengths to obtain an initial node feature matrix. The message passing and node embedding module is used to perform message passing and node embedding based on a relation-aware graph neural network on the initial node feature matrix to obtain the final node embedding matrix based on the heterogeneous knowledge graph. The equipment risk prediction module is used to predict equipment risks from the final node embedding matrix to obtain an equipment risk score vector. The substation comprehensive evaluation module is used to perform a comprehensive evaluation of the substation based on the final node embedded matrix to obtain a comprehensive substation score. The risk path tracing module is used to trace the risk path of equipment risk scoring vectors and heterogeneous knowledge graphs to obtain an interpretable subgraph.
[0008] Compared with existing technologies, this application provides a digital twin-based intelligent substation status assessment method and system. First, it constructs a heterogeneous knowledge graph that accurately maps the logical relationships between physical entities based on the substation's ontology definition, asset information, and topology data. Then, it introduces massive real-time, rapidly changing data streams into this graph framework and uses temporal feature encoding technology to dynamically map node states. Next, a relationship-aware graph neural network performs complex message passing and aggregation mechanisms within the graph structure, generating node embedding representations that possess both local individuality and global commonality. Finally, based on this high-dimensional representation, it simultaneously completes micro-level equipment risk prediction and macro-level comprehensive station assessment, and performs path backtracking and causal analysis for high-risk states based on graph correlations. This solves the problem of the disconnect between equipment health and system operating status, achieving interpretable panoramic status perception and risk early warning for intelligent substations. Attached Figure Description
[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 This is a flowchart of a digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0011] Figure 2 This is a data flow diagram of a digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0012] Figure 3 This is a flowchart of sub-step S1 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0013] Figure 4 This is a flowchart of sub-step S2 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0014] Figure 5 This is a flowchart of sub-step S23 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0015] Figure 6 This is a flowchart of sub-step S3 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0016] Figure 7This is a flowchart of sub-step S4 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0017] Figure 8 This is a flowchart of sub-step S5 of the digital twin-based smart substation condition assessment method according to an embodiment of this application.
[0018] Figure 9 This is a block diagram of a digital twin-based intelligent substation condition assessment system according to an embodiment of this application. Detailed Implementation
[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0020] To address the problems mentioned above in the background technology, this application proposes a smart substation condition assessment method based on digital twins. Figure 1 This is a flowchart of a digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 2 This is a data flow diagram of a digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 1 and Figure 2 As shown, the intelligent substation status assessment method based on digital twins includes the following steps: S1, constructing a heterogeneous knowledge graph based on ontology definition, asset data, topology data, and historical slow-changing data; S2, encoding the temporal features of graph nodes in the real-time fast-changing data stream based on the heterogeneous knowledge graph and the time window length to obtain an initial node feature matrix; S3, performing message passing and node embedding based on a relation-aware graph neural network on the initial node feature matrix based on the heterogeneous knowledge graph to obtain a final node embedding matrix; S4, performing equipment risk prediction on the final node embedding matrix to obtain an equipment risk scoring vector; S5, performing a comprehensive substation assessment on the final node embedding matrix to obtain a comprehensive substation score; S6, performing risk path tracing on the equipment risk scoring vector and the heterogeneous knowledge graph to obtain an interpretable subgraph.
[0021] In the aforementioned digital twin-based intelligent substation status assessment method, step S1 involves constructing a heterogeneous knowledge graph based on ontology definitions, asset data, topology data, and historical slowly varying data. It should be understood that data such as equipment ledgers, topology connections, and historical maintenance records of intelligent substations are scattered across different systems, with heterogeneous data formats and implicit relationships. Existing assessment methods struggle to integrate this data for deep correlation analysis, resulting in limited assessment dimensions. Therefore, this application integrates ontology definitions, asset data, topology data, and historical slowly varying data, constructing a heterogeneous knowledge graph through structured modeling to break down data silos and make complex relationships between data explicit. This transforms multi-source data into a unified graph structure, clearly presenting the relationships between equipment, measurement points, topology, and historical status, providing a structured data foundation for subsequent feature encoding and message passing, and ensuring the accuracy of correlation reasoning and the effectiveness of the assessment process.
[0022] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S1 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 3 As shown, step S1 includes: S11, performing graph ontology modeling and graph initialization based on the ontology definition to obtain an initialized graph pattern; S12, extracting static and historical data from asset data, topology data, and historical slowly changing data, and filling them into the initialized graph pattern to obtain a filled static graph; S13, linking real-time temporal data stream pointers between the time-series database metadata and the filled static graph to obtain a heterogeneous knowledge graph.
[0023] Specifically, step S11 involves graph ontology modeling and graph initialization based on the ontology definition to obtain an initialized graph pattern. It should be understood that, due to the lack of a unified structured standard, directly integrating multi-source heterogeneous data can lead to a chaotic graph structure, poor data compatibility, and an inability to support subsequent association reasoning and data processing. Therefore, this application further performs graph ontology modeling and initialization operations based on a preset ontology definition to clarify the core components and constraint rules of the graph. This establishes a unified graph structure standard, clarifies the definitions and constraints of node labels, edge types, and various attributes, provides a standardized framework for the orderly filling of subsequent multi-source data, ensures the consistency, standardization, and interoperability of graph data, and avoids data conflicts or association failures caused by inconsistent structures.
[0024] Specifically, in one possible embodiment, step S11 is implemented as follows: First, an automated parsing program is started to read the structured descriptions of node labels, edge types, and attribute keys from the ontology definition file. Second, for each node label, its core attributes and data types are defined, and unique constraints are created for key attributes (such as device ID and measurement point ID). Next, the node label combinations corresponding to each type of edge are defined, limiting the connection range of the edges. Then, the above definitions are converted into native operation instructions for the graph database and executed in the target graph database. Finally, the database index creation and constraint activation are completed, forming an initial graph pattern that only contains structural specifications and has no actual data.
[0025] Specifically, step S12 involves extracting static and historical data from asset data, topology data, and historical slowly varying data, and then filling these data into the initialized graph pattern to obtain a filled static graph. It should be understood that since the initialized graph pattern only has a structural framework and lacks actual data support, it cannot reflect the equipment configuration, topology relationships, and historical operating status of the smart substation, making it difficult to meet the data content requirements of status assessment. Therefore, this application further extracts and processes asset data, topology data, and historical slowly varying data, filling them into the initialized graph pattern according to the ontology definition specifications. This enriches the graph's content dimensions and constructs a complete data carrier containing static configuration and historical status. In this way, scattered static and historical data can be transformed into nodes, edges, and attributes in the graph, enabling the graph to describe basic equipment information, relationships, and historical status, providing static data support for subsequent dynamic data association and in-depth assessment.
[0026] Specifically, step S13 involves linking real-time temporal data stream pointers between the time-series database metadata and the populated static graph to obtain a heterogeneous knowledge graph. It should be understood that the populated static graph only contains static and historical data, which cannot perceive the dynamic operating status of the substation in real time. Real-time rapid data streams (such as SCADA data) are massive and high-frequency; directly storing them in the graph would lead to a surge in storage pressure, decreased query efficiency, and affect the timeliness of the assessment. Therefore, this application further utilizes the time-series database metadata to establish real-time data stream access pointers for the measurement point nodes in the static graph, thereby achieving a lightweight association between the static graph and the real-time data stream. In this way, without increasing the graph's storage burden, the graph can acquire dynamic data in real time, integrating static and dynamic data to provide data support for dynamic assessments based on real-time status, ensuring the timeliness and accuracy of the assessment results.
[0027] Specifically, in one possible embodiment, step S13 is implemented as follows: First, all nodes labeled as measurement points in the populated static graph are traversed, and the unique identifier of each measurement point node is extracted. Second, the time-series database metadata is read to obtain the database access address, query interface, and data format specifications corresponding to each measurement point. Then, for each measurement point node, its unique identifier is matched with the time-series database metadata to generate a unique access pointer containing the access address and query parameters. Finally, the access pointer is stored in the graph as a new attribute of the measurement point node, completing the pointer link between all measurement point nodes and the real-time data stream, forming a heterogeneous knowledge graph with the ability to access static data, historical data, and dynamic data.
[0028] In the aforementioned digital twin-based intelligent substation status assessment method, step S2 involves encoding the temporal features of graph nodes in the real-time fast-changing data stream based on a heterogeneous knowledge graph and the time window length to obtain an initial node feature matrix. It should be understood that heterogeneous knowledge graph nodes only contain static and historical information, while the real-time fast-changing data stream is in a high-frequency temporal sequence. These two cannot be directly adapted to subsequent graph neural network calculations, resulting in a lack of timeliness and uniformity in the model input. Therefore, this application further combines the heterogeneous knowledge graph node structure and the time window length to encode the temporal features of graph nodes in the real-time fast-changing data stream, thereby transforming dynamic data into node features of a unified dimension. This allows nodes to possess both static associations and real-time dynamic features, forming a standardized initial node feature matrix, providing compliant input for subsequent calculations, ensuring the model's perception of the real-time state, and avoiding assessment biases caused by data mismatch.
[0029] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S2 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 4 As shown, step S2 includes: S21, extracting the original time-series segments of measurement point nodes from the real-time fast-changing data stream based on the heterogeneous knowledge graph and the time window length; S22, performing time-series feature encoding on the original time-series segments of the measurement point nodes to obtain a measurement point feature mapping table; S23, generating a device feature mapping table based on the heterogeneous knowledge graph and the measurement point feature mapping table; S24, performing category feature encoding on other types of nodes in the heterogeneous knowledge graph to obtain other node feature mapping tables; S25, integrating the measurement point feature mapping table, the device feature mapping table, and other node feature mapping tables into a full-graph node initial feature matrix to obtain an initial node feature matrix.
[0030] Specifically, step S21 involves extracting the original time-series segments of the measurement point nodes from the real-time rapidly changing data stream based on the heterogeneous knowledge graph and the time window length. It should be understood that since the measurement point nodes in the heterogeneous knowledge graph only store data stream access pointers and do not contain specific time-series data, and subsequent time-series feature encoding relies on continuous data within a fixed time range, using the full amount of data would result in redundancy and inefficiency, and may also lead to encoding errors. Therefore, this application further combines the measurement point structure of the heterogeneous knowledge graph with the set time window length to extract the original time-series segments corresponding to the measurement point nodes from the real-time rapidly changing data stream, thereby obtaining effective time-series data that meets the encoding requirements. This allows for precise filtering of dynamic data within a specific time range, ensuring the timeliness and continuity of the data, providing high-quality input for subsequent time-series feature encoding, effectively avoiding encoding errors caused by data issues, and guaranteeing the accuracy of the dynamic features of the measurement point nodes.
[0031] Specifically, step S22 involves encoding the original time-series segments of the measurement point nodes using time-series features to obtain a measurement point feature mapping table. It should be understood that the original time-series segments of the measurement points are variable-length numerical sequences with inconsistent dimensions and lack of extracted key dynamic information, making them unsuitable for direct participation in subsequent graph neural network calculations and hindering the utilization of dynamic features. Therefore, this application further encodes the original time-series segments of the measurement points using time-series features to transform the variable-length data into fixed-dimensional feature vectors and associate them with measurement point identifiers. This allows for the extraction of key time-series information, the formation of standardized feature vectors, and the clarification of the correspondence between measurement points and features through the mapping table. This provides unified dynamic data for equipment feature aggregation, ensuring the real-time status of subordinate measurement points integrated by the equipment.
[0032] Specifically, in one possible embodiment, step S22 is implemented as follows: First, the original time-series segments of each measurement point node are preprocessed, and Z-score normalization is used to eliminate the influence of dimensions. Second, the normalized time-series segments are input into a pre-trained gated recurrent unit model. This model filters key time-series information through internal update and reset gates, and outputs the hidden state of the last time step of the model. Then, the hidden state is used as the time-series feature vector of the measurement point node, and its dimension is fixed according to the model's preset parameters. Next, a measurement point feature mapping table is created, with the measurement point node ID as the key and the corresponding time-series feature vector as the value, and the encoding results of all measurement points are entered one by one. Finally, the dimension of the feature vectors in the mapping table is verified to ensure that all vectors have consistent dimensions, forming the final measurement point feature mapping table.
[0033] Specifically, in step S23, a device feature mapping table is generated based on the heterogeneous knowledge graph and the measurement point feature mapping table. It should be understood that device nodes only contain static attributes, while the dynamic features of measurement points are scattered throughout the mapping table. The lack of fusion between the two results in incomplete device features, affecting the accuracy of the evaluation. Therefore, this application further combines the device-measurement point association from the heterogeneous knowledge graph with the measurement point feature mapping table to generate a device feature mapping table, thereby integrating the static attributes of the device and the dynamic features of the measurement points. This allows device features to possess both inherent attributes and real-time status, forming a comprehensive device feature vector, providing complete device status data for global feature integration and model evaluation.
[0034] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S23 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 5 As shown, step S23 includes: S231, extracting a first device node from a heterogeneous knowledge graph; S232, performing static attribute encoding on the first device node to obtain a static feature part; S233, extracting a set of associated measurement point nodes related to the first device node from a measurement point feature mapping table; S234, performing feature vector aggregation on the set of associated measurement point nodes to obtain a dynamic feature part; S235, performing feature fusion on the static feature part and the dynamic feature part to obtain the features of the first device node.
[0035] More specifically, step S231 involves extracting a first device node from the heterogeneous knowledge graph. It should be understood that batch processing of all device nodes can easily lead to confusion in relationships or errors in attribute matching, resulting in inaccurate device feature generation and an inability to ensure a precise correspondence between the static attributes and associated measurement points of each device. Therefore, this application further extracts a first device node from the heterogeneous knowledge graph to determine the processing unit for generating individual device features. This allows for feature construction centered on a single device, clarifying its static attributes and the range of associated measurement points, avoiding data contamination issues from batch processing, ensuring the accuracy and independence of feature generation for each device, and laying the foundation for subsequent device-by-device coding.
[0036] More specifically, step S232 involves statically encoding the first device node to obtain a static feature portion. It should be understood that the static attribute formats of the first device node are heterogeneous; for example, the commissioning date is in date format and the model number is in text format. These are not converted into machine-recognizable numerical features and cannot be directly fused with dynamic feature vectors, resulting in a lack of standardized representation of device features. Therefore, this application further encodes the static attributes of the first device node to transform the heterogeneous static attributes into a feature vector of a unified dimension. This eliminates attribute format differences, preserves the inherent attribute information of the device, provides a basis for the fusion of static features with dynamic features, offers standardized input for subsequent feature fusion, and ensures the integrity of device features.
[0037] Specifically, in one possible embodiment, step S232 is implemented as follows: First, static attributes are extracted from the basic dataset of the first device node, distinguishing between numerical and categorical attributes. Second, the operating years are calculated for numerical attributes (such as commissioning date), and Min-Max normalization is used to scale them to the [0,1] interval; for categorical attributes (such as model number), a pre-built model number embedding dictionary is called to map the text into a fixed-dimensional dense vector. Then, for other attributes such as manufacturer and rated voltage, one-hot encoding or embedding processing is used respectively to convert them into vectors of the same dimension. Next, all processed vectors are concatenated in a preset order to form the static feature part. Finally, the vector dimension is verified to ensure that it meets the preset standard, and it is stored after confirmation.
[0038] More specifically, step S233 involves extracting a set of associated measurement point nodes related to the first device node from the measurement point feature mapping table. It should be understood that since the dynamic state of the first device node needs to be reflected through its subordinate measurement point data, randomly extracting measurement point features would lead to a mismatch between the dynamic features and the device, failing to accurately represent the device's operating state and thus affecting the accuracy of the device features. Therefore, this application further extracts a set of measurement point features associated with the first device node from the measurement point feature mapping table to obtain the device's unique dynamic data. This allows for precise matching of the association between the device and the measurement points, ensuring that the extracted dynamic features all originate from the device's monitoring data, avoiding interference from irrelevant data, accurately reflecting the device's real-time operating state, providing accurate input for subsequent dynamic feature aggregation, and ensuring the reliability of the overall device feature construction.
[0039] More specifically, step S234 involves aggregating feature vectors of the associated measurement point node feature set to obtain the dynamic feature portion. It should be understood that the associated measurement point node feature set contains independent temporal feature vectors from multiple measurement points. Directly using these scattered vectors cannot reflect the overall dynamic operating state of the device and increases the complexity of subsequent feature fusion, resulting in a one-sided dynamic feature of the device. Therefore, this application further aggregates the vectors of the associated measurement point node feature set to integrate the dynamic features of multiple measurement points into a single dynamic feature vector for the device. This allows for the extraction of common dynamic trends from multiple measurement points, forming a vector representing the overall dynamic state of the device, simplifying the fusion process, avoiding misjudgments of the dynamic state due to one-sided data from a single measurement point, and ensuring the effectiveness of the dynamic features.
[0040] Specifically, in one possible embodiment, step S234 is implemented as follows: First, all temporal feature vectors in the associated measurement point node feature set are read to confirm that the number of vectors is consistent with the number of associated measurement point IDs. Second, an average pooling algorithm is used to average all vectors element-wise according to their corresponding dimensions. If there are abnormal vectors with values exceeding the normal range, they are removed before further calculation. Then, the aggregated vectors are subjected to numerical range verification to ensure they conform to the preset range of feature vectors. Next, if the dimension of the aggregated vector is inconsistent with the static feature part, it is adjusted to the target dimension through a linear transformation layer. Finally, the adjusted vector is used as the dynamic feature part of the first device node and stored together with the static feature part in a temporary cache.
[0041] More specifically, step S235 involves fusing the static and dynamic features to obtain the features of the first device node. It should be understood that since the static features only reflect the inherent attributes of the device, and the dynamic features only reflect its real-time operating state, neither can fully describe the current comprehensive state of the device when existing alone. This makes it difficult for subsequent evaluation models to obtain comprehensive information about the device, thus affecting the evaluation accuracy. Therefore, this application further fuses the static and dynamic features of the first device node to form a feature vector containing complete state information of the device. This organically combines the inherent attributes of the device with its real-time operating state, generating a unified-dimensional device feature vector. This provides comprehensive and accurate device feature input for subsequent equipment risk prediction and substation-wide comprehensive evaluation, effectively ensuring the reliability of the evaluation results.
[0042] Specifically, step S24 involves encoding the category features of other types of nodes in the heterogeneous knowledge graph to obtain a feature mapping table for other nodes. It should be understood that nodes such as alarm events and substation areas in the heterogeneous knowledge graph only possess text or category attributes, lacking structured feature vectors. These cannot participate in the calculation in unison with equipment and measurement point nodes, resulting in incomplete feature coverage of the graph. Therefore, this application further encodes the category features of other types of nodes to transform unstructured category information into fixed-dimensional feature vectors. This ensures that all nodes have standardized features, guarantees that the initial node feature matrix covers all nodes in the graph, avoids calculation errors, and ensures the completeness and accuracy of the graph neural network calculation.
[0043] Specifically, in one possible embodiment, step S24 is implemented as follows: First, other types of nodes besides equipment and measuring points, such as alarm event nodes and substation area nodes, are selected from the heterogeneous knowledge graph. Second, for alarm event nodes, category attributes such as alarm type (e.g., overload, abnormal oil level) are extracted to construct an alarm type dictionary, and each alarm type is mapped to a fixed-dimensional embedding vector through an embedding layer. For substation area nodes, attributes such as area number and functional partition are extracted, and feature vectors are generated using one-hot encoding combined with dimensionality compression. Then, the feature vector of each other type of node is associated with its node ID. Finally, a feature mapping table for other nodes is created, and the feature vectors of all other types of nodes are entered in order of node ID to complete the generation of the feature mapping table for other nodes.
[0044] Specifically, step S25 involves integrating the measurement point feature mapping table, equipment feature mapping table, and other node feature mapping tables into a full-graph node initial feature matrix to obtain an initial node feature matrix. It should be understood that the three types of node features are stored in independent mapping tables with inconsistent node order, making them unsuitable as direct input to the graph neural network and resulting in a lack of standardized global input for the model. Therefore, this application further integrates the three types of mapping tables into a full-graph node initial feature matrix to organize the dispersed features into a matrix in a unified order. This allows the node features to be arranged in a fixed order, forming a standardized initial node feature matrix, ensuring accurate correspondence between features and nodes in subsequent calculations and guaranteeing orderly and accurate model computation.
[0045] Specifically, in one possible embodiment, step S25 is implemented as follows: First, a unique identifier list of all nodes in the heterogeneous knowledge graph is exported, and a fixed sorting rule is determined according to the node creation time. Second, an initial node feature matrix with all zeros is created based on the total number of nodes and the dimension of the feature vector of a single node. Then, the node identifier list is traversed, and for each node identifier, its type is determined: if it is a measurement point node, the vector is retrieved from the measurement point feature mapping table; if it is a device node, the vector is retrieved from the device feature mapping table; if it is another type of node, the vector is retrieved from the other node feature mapping table. Next, the retrieved vector is filled into the corresponding row of the matrix. Finally, the matrix is checked for integrity to ensure that there are no empty rows or abnormal dimensions. After the check passes, it is stored to form the initial node feature matrix.
[0046] In the aforementioned digital twin-based intelligent substation status assessment method, step S3 involves performing message passing and node embedding on the initial node feature matrix using a relation-aware graph neural network based on a heterogeneous knowledge graph to obtain the final node embedding matrix. It should be understood that since the initial node feature matrix only contains the static and dynamic information of the nodes themselves, without incorporating neighboring nodes and inter-node relationships, the node features lack graph structural correlation and cannot support the graph neural network's assessment of the global relational state of the intelligent substation. Therefore, this application further relies on the relational structure of the heterogeneous knowledge graph to perform message passing and node embedding on the initial node feature matrix using a relation-aware graph neural network, thereby integrating node neighborhood information and relational features. This allows the node embedding vector to retain its own features while incorporating multi-hop neighborhood relational states, forming a node representation with global structural awareness. This provides feature input rich in relational information for subsequent assessments, avoiding the one-sidedness of assessments due to a lack of structural correlation.
[0047] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S3 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 6 As shown, step S3 includes: S31, extracting the relation adjacency matrix and relation weight matrix from the heterogeneous knowledge graph; S32, based on the relation adjacency matrix and relation weight matrix, performing neighborhood message generation and relation-based aggregation on the initial node feature matrix to obtain a set of relation-aggregated message matrices; S33, performing cross-relation message fusion on the set of relation-aggregated message matrices to obtain a pre-activation embedding matrix; S34, performing nonlinear activation and inter-layer iteration on the pre-activation embedding matrix to obtain the final node embedding matrix.
[0048] Specifically, step S31 involves extracting a relation adjacency matrix and a relation weight matrix from the heterogeneous knowledge graph. It should be understood that since message transmission in a relation-aware graph neural network relies on the relationship types and connections between nodes, and the impact of different relationships on messages needs to be reflected through specific weights, the lack of a relation adjacency matrix and relation weight matrix would prevent the network from distinguishing message transmission rules for different relationships, leading to message transmission chaos. Therefore, this application further extracts a relation adjacency matrix and a relation weight matrix from the heterogeneous knowledge graph to clarify the connection relationships between nodes and relation-specific parameters. This allows for the construction of an independent adjacency matrix for each relationship to identify the node connection state, initialization of the weight matrix to quantify the degree of influence of the relationship on the message, accurate capture of relation specificity, avoidance of message transmission deviations, and provision of basic data and parameter support for subsequent relation-based aggregation.
[0049] Specifically, in step S32, based on the relation adjacency matrix and relation weight matrix, neighborhood message generation and intra-relational aggregation are performed on the initial node feature matrix to obtain a set of relation aggregated message matrices. It should be understood that the initial node feature matrix does not incorporate neighborhood information, and neighborhood messages under different relations need to be processed separately to retain relation specificity. If neighborhood messages from all relations are directly mixed, the differentiated impact of relations on node states will be lost. Therefore, this application further combines the relation adjacency matrix and weight matrix to perform neighborhood message generation and intra-relational aggregation on the initial node feature matrix, thereby extracting neighborhood information according to relation type. In this way, node neighborhood features can be accurately extracted for each relation, and exclusive aggregated messages can be formed through weight transformation and aggregation, ensuring that messages from different relations are not confused, providing clear input for cross-relationship fusion, and guaranteeing the relation specificity of message transmission.
[0050] Specifically, in one possible embodiment, step S32 is implemented as follows: First, a relation type and its corresponding adjacency matrix and weight matrix are selected from the matrix set. Second, the initial node feature matrix is multiplied by the relation weight matrix to complete the relation-specific transformation of the node features, generating a message feature matrix. Then, based on the relation adjacency matrix, neighborhood aggregation is performed on the message feature matrix: for each node, the transformed features of all its neighboring nodes under the relation are extracted and summed. Next, a normalization constant (the in-degree of the node under the relation) is introduced to normalize the aggregation result to avoid numerical overflow. Finally, the aggregation result of the relation is encapsulated into a relation aggregation message matrix, and the above steps are repeated until all relation types have been processed, forming a set of relation aggregation message matrices.
[0051] Specifically, step S33 involves performing cross-relationship message fusion on the set of relationship aggregation message matrices to obtain a pre-activation embedding matrix. It should be understood that each matrix in the set of relationship aggregation message matrices contains only neighborhood messages of a single relationship, while node states are influenced by multiple relationships. Without fusing multi-relationship messages, the node embedding cannot comprehensively reflect the global association state, resulting in one-sided node features. Therefore, this application further performs cross-relationship message fusion on the set of relationship aggregation message matrices to integrate the neighborhood influence of all relationships on the nodes. This allows for the element-wise fusion of aggregation messages from different relationships along the node dimension, while also incorporating the transformation information of the node's own features, forming a pre-activation embedding matrix that includes the comprehensive influence of multiple relationships. This provides comprehensive input for subsequent nonlinear activation and avoids misjudgments of node states caused by single-relationship messages.
[0052] Specifically, in one possible embodiment, step S33 is implemented as follows: First, all matrices in the relational aggregation message matrix set are read to confirm that the number of nodes in each matrix is consistent with the feature dimension. Second, all matrices are summed element-wise according to the node dimension to obtain the cross-relational aggregation message matrix. Then, the self-loop weight matrix is initialized by multiplying the initial node feature matrix with the self-loop weight matrix to obtain the self-loop message matrix. Next, the cross-relational aggregation message matrix and the self-loop message matrix are added element-wise to complete message fusion. Finally, the fusion matrix is dimension-verified to ensure that it is consistent with the preset hidden layer dimension, forming the pre-activation embedding matrix.
[0053] Specifically, step S34 involves performing nonlinear activation and inter-layer iteration on the pre-activation embedding matrix to obtain the final node embedding matrix. It should be understood that the pre-activation embedding matrix is a linear combination of multi-relational messages and self-loop messages, lacking nonlinear expressive power and unable to capture complex nonlinear relationships between node features. Furthermore, single-layer message passing can only fuse single-hop neighborhood information, failing to reflect deep structural relationships. Therefore, this application further performs nonlinear activation and inter-layer iteration on the pre-activation embedding matrix to enhance feature expressive power and fuse multi-layer neighborhood information. In this way, nonlinear relationships are introduced through a nonlinear activation function, and multi-hop neighborhood information is gradually fused through multi-layer iteration, enabling the final node embedding matrix to possess deep structural perception and complex feature expression capabilities. This provides high-precision feature support for subsequent evaluation, avoiding insufficient evaluation accuracy caused by linear expression or shallow relationships.
[0054] Specifically, in one possible embodiment, step S34 is implemented as follows: First, ReLU is selected as the nonlinear activation function. The pre-activation embedding matrix is input to the function, and the post-activation embedding matrix is calculated element-wise, removing negative values to enhance sparse representation. Second, the post-activation embedding matrix is used as the input for the next layer, repeating the process of intra-relation aggregation, cross-relation fusion, and nonlinear activation. Then, the iteration layer number is recorded, and the gradient of each relation weight matrix is updated after each layer is completed, and the parameters are optimized through gradient descent. Next, when the iteration layer reaches a preset 3 layers, message passing is stopped. Finally, the post-activation embedding matrix of the last layer is used as the final node embedding matrix.
[0055] In particular, instead of constructing a static and binary (0 or 1) set of adjacency matrices as the fixed structure input of the graph neural network, the weighted adjacency matrix is dynamically calculated at each layer of the graph neural network (or at the beginning) through the attention mechanism, using the initial features encoded for each node that contain rich temporal information. This achieves a dynamically weighted adjacency matrix based on the attention mechanism. This means that the structure of the graph (i.e. the connection strength between nodes) is no longer fixed in advance, but is determined by the current state (features) of the nodes and changes dynamically over time.
[0056] Specifically, firstly, static adjacency matrices cannot reflect the intensity of dynamic influences. In a real substation, the intensity of mutual influence between equipment changes dynamically. For example, the connection between a transformer and its connected lines is physically constant, so in a static adjacency matrix, the corresponding element is always 1. However, when the line is under heavy load, its impact on the transformer's thermal state and lifespan loss is far greater than when it is under light load. Static edges with a value of 1 cannot capture the weight of this dynamic change. Secondly, static adjacency matrices ignore implicit functional relationships. In addition to physical connections and explicit attribution relationships, there are also many implicit functional relationships in the system. For example, two feeders that are not directly connected may have highly correlated voltage and power curves at specific times because they serve the same type of highly resilient load. This relationship of shared disturbances cannot be expressed by static topology and needs to be discovered and utilized from the data. Furthermore, after encoding real-time fast-changing data streams and historical slow-changing data into high-quality node features through a temporal encoder, if the static graph structure is unrelated to the node features, it means that only dynamic information is used as node attributes, without being used to optimize and guide the message passing path itself. Therefore, because the temporal feature information is not fully utilized for structural modeling, the model's ability to learn deep structural knowledge from the data is limited. Based on this, in another preferred embodiment, step S3 includes: constructing a dynamically weighted adjacency matrix based on a graph attention mechanism to replace the static adjacency matrix as the structural input of the graph neural network; calculating attention coefficients for the target node and its neighboring nodes in the graph based on their respective node features in the current layer and the relation-specific learnable weight matrix, whereby the attention coefficients characterize the importance of neighboring node features to the target node in the current state; normalizing the attention coefficients using the Softmax function to obtain dynamic attention weights, and using these dynamic attention weights to weighted aggregate the messages of neighboring nodes to update the embedding vector of the target node.
[0057] In other words, instead of directly modifying the adjacency matrix, attention weights are integrated into the message passing computation process, dynamically assigning weights to existing edges, thereby improving the graph neural network into a relation-aware graph attention network. Specifically, in each layer of the graph neural network, for the target node and its neighboring nodes, attention coefficients are calculated based on the node features of the target node and its neighboring nodes in the current layer, as well as the relation-specific learnable weight matrix. The attention coefficient is used to characterize the importance of neighboring nodes to the target node in the current running state. This importance is based on their respective characteristics. and The attention coefficients are calculated using a learnable neural network. Then, the attention coefficients are normalized using the Softmax function to obtain dynamic attention weights. The messages of neighboring nodes are then weighted and aggregated based on the dynamic attention weights to update the embedding vector of the target node.
[0058] That is, the normalization term in the R-GCN update formula is changed. Replace static, graph-based constants with dynamically computed attention weights. Here, for each edge from j to i in the l-th layer and under relation r, the attention coefficient is calculated, which is the feature of node i and node j after a relation-specific linear transformation of relation r. and Joint decision:
[0059] in, and These are the embedding vectors of nodes i and j at layer l, respectively. It is a learnable weight matrix specific to relation r, representing the vector concatenation operation. It is a learnable attention vector for the relation type identifier r. This is the transpose of the vector. As a nonlinear activation function Let these be the attention coefficients. Then, the Softmax function is used to normalize the attention coefficients on all neighbors of node i to ensure that the sum of the weights is 1, thus obtaining the attention weights. .
[0060] Therefore, the new graph neural network layer update formula becomes:
[0061] in, To normalize attention weights, Let l be the self-loop weight matrix of the l-th layer. Let r be the set of neighbors of the target node i. For a set of relation types, It is a non-linear activation function. Let be the embedding vector of the target node i at the (l+1)th layer.
[0062] That is, It dynamically reflects the current state (by...) and This reflects the importance of neighbor j to i. For example, when line j is heavily loaded, its node characteristics change, causing the model to calculate a larger value. This gives its messages a higher weight during aggregation, thus passing on this overload effect more significantly to transformer i.
[0063] Furthermore, considering physically non-adjacent devices, if their operating states are highly similar, they are likely influenced by common upstream factors or have similar operating patterns. Therefore, based on the dynamic attention weights, the similarity between node features, i.e., the features of nodes i and j, can be further fused. and Cosine similarity between them, and attention weights Add them together.
[0064] In this way, attention weights can also handle semantic associations based on data-driven discovery, such as if two non-adjacent transformers exhibit similar harmonic distortion characteristics (reflected in their nodal features) due to sharing similar nonlinear loads. and In the case of a fault warning caused by harmonics, the model will establish attention reinforcement between them through cosine similarity. When one of them has a fault warning caused by harmonics, the model can transmit the risk warning signal to the other through semantic edge weights, even if they are physically far apart, thus realizing the identification of systemic and non-local risks.
[0065] Finally, after repeating the message passing and embedding update process multiple times, a final node embedding matrix containing the embedding vectors of all nodes is obtained. This allows the final node embedding matrix to retain the core features of each node while fully integrating multi-dimensional neighborhood association information, possessing deep structure perception and complex feature characterization capabilities. This provides high-precision feature support for subsequent equipment risk prediction and substation comprehensive assessment, effectively avoiding assessment biases caused by linear expressions or shallow associations, and ensuring the accuracy and comprehensiveness of the assessment results.
[0066] Thus, the model can learn which connections between devices should be strengthened and which should be weakened under different conditions (such as high load and voltage fluctuations), making the assessment results more sensitive to critical and dynamically changing risk transmission paths. Moreover, by dynamically constructing functional connections, the model can discover cooperative anomalies between non-topologically adjacent devices, thereby revealing more complex hidden fault modes, such as accelerated aging of multi-point devices caused by network-wide harmonic problems.
[0067] In the aforementioned digital twin-based intelligent substation status assessment method, step S4 involves performing equipment risk prediction on the final node embedding matrix to obtain an equipment risk score vector. It should be understood that the final node embedding matrix contains deep correlation features of equipment nodes, but exists in the form of a high-dimensional vector, which is not transformed into an intuitive risk result and cannot directly provide risk level references for operation and maintenance personnel, leading to a disconnect between features and decision-making needs. Therefore, this application further performs equipment risk prediction on the final node embedding matrix to map high-dimensional features into a quantitative risk score. This allows for the generation of a 0-1 range score for each device, forming an equipment risk score vector, enabling operation and maintenance personnel to clearly understand the risk level of each device, providing a basis for predictive maintenance, avoiding maintenance delays, and ensuring the safe operation of equipment.
[0068] In particular, in one specific embodiment, Figure 7 This is a flowchart of sub-step S4 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 7 As shown, step S4 includes: S41, extracting the embedding vector of each device node from the final node embedding matrix; S42, inputting the embedding vector of each device node into a shared prediction head to obtain the device risk score vector.
[0069] Specifically, step S41 involves extracting the embedding vectors of each device node from the final node embedding matrix. It should be understood that since the final node embedding matrix contains embedding vectors for all types of nodes in the graph, directly using it for device risk prediction would introduce irrelevant features such as measurement points and alarm events, leading to misjudgments in the prediction model and an inability to accurately focus on the core requirements of device risk assessment. Therefore, this application further extracts the embedding vectors corresponding to each device node from the final node embedding matrix to filter the effective features required for device risk prediction. This eliminates interfering features from non-device nodes, retaining only the deep-level correlation features of the devices, ensuring that the prediction model calculates risk scores based on device-specific features, avoiding assessment bias caused by irrelevant features, and guaranteeing the accuracy of device risk prediction.
[0070] Specifically, in step S42, the embedding vectors of each device node are input into a shared prediction head to obtain a device risk score vector. It should be understood that the embedding vector of each device node is a high-dimensional abstract feature, requiring a unified mapping rule to transform it into a risk score. Constructing a separate prediction head for each device would lead to parameter redundancy, inconsistent scoring standards, and an inability to horizontally compare the risk levels of different devices. Therefore, this application further inputs the embedding vectors of each device into a shared prediction head to achieve risk score transformation through a unified model. This ensures consistent and comparable scoring standards for all devices by relying on the unified parameters and logic of the shared prediction head, reducing model parameters, avoiding scoring confusion, and guaranteeing the consistency and reliability of the vectors.
[0071] Specifically, in one possible embodiment, step S42 is implemented as follows: First, a shared prediction head is constructed: the input layer dimension is consistent with the device embedding vector, the hidden layer has 64-dimensional neurons activated by ReLU, and the output layer has 1-dimensional neurons activated by Sigmoid. Second, pre-trained parameters are loaded, which are obtained by training historical device fault data and corresponding embedding vectors. Then, the device embedding vector dataset is traversed, and each data point is input into the prediction head, processed, and output as a single device score. Next, the scores are recorded in order of device ID to form an initial sequence. Finally, the integrity of the sequence is verified to ensure that there are no missing scores, forming a device risk score vector.
[0072] In the aforementioned digital twin-based intelligent substation status assessment method, step S5 involves performing a comprehensive substation assessment on the final node embedding matrix to obtain a comprehensive substation score. It should be understood that the final node embedding matrix only provides deep features of individual nodes, and the equipment risk score vector only reflects the independent risks of each device. However, as a whole system, the substation's operating status is influenced by equipment relationships and topology; dispersed features cannot reflect the overall health level of the entire substation. Therefore, this application conducts a comprehensive substation assessment on the final node embedding matrix to integrate the node features across the entire graph and generate a score reflecting the overall status of the entire substation. This transforms node-level features into substation-level indicators, quantifies the overall health and risk of the substation, provides the dispatch center with a basis for substation-wide decision-making, and avoids overlooking system-level risks.
[0073] In particular, in one specific embodiment, Figure 8 This is a flowchart of sub-step S5 of the digital twin-based smart substation condition assessment method according to an embodiment of this application. Figure 8 As shown, step S5 includes: S51, performing global graph aggregation on the final node embedding matrix to obtain a global representation vector; S52, inputting the global representation vector into the MLP prediction head to obtain the substation comprehensive score.
[0074] Specifically, step S51 involves performing global graph aggregation on the final node embedding matrix to obtain a global representation vector. It should be understood that each row vector in the final node embedding matrix corresponds to the feature of only a single node; the features of each node exist independently and cannot reflect the global correlation between nodes or the overall characteristics of the system. If directly used for comprehensive evaluation of the entire site, the evaluation results would only reflect the individual state of nodes, ignoring system-level interactions. Therefore, this application further performs global graph aggregation on the final node embedding matrix to integrate the node features of the entire graph into a single vector. This allows for the extraction of common information and global correlation patterns among the node features of the entire graph, condensing the scattered node-level features into a low-dimensional, unified global representation vector, providing standardized full-site feature input for subsequent MLP prediction heads.
[0075] Specifically, in one possible embodiment, step S51 is implemented as follows: First, the dimensions of the final node embedding matrix are determined, confirming that the number of rows is the total number of nodes in the entire graph and the number of columns is the node feature dimension. Second, average pooling is selected as the aggregation method, and the program is started to sum all elements in each column of the matrix, then divided by the total number of nodes to obtain the aggregated value for that dimension. Then, all feature dimensions are calculated sequentially, and the aggregated values are arranged in dimensional order to form a global representation vector. Next, the values of each dimension of the vector are checked to see if they are within the normal range; abnormal dimensions are recalculated. Finally, the vectors that pass the verification are stored in a dedicated cache, recording the vector dimension and aggregation time.
[0076] Specifically, in step S52, the global representation vector is input into the MLP prediction head to obtain the substation comprehensive score. It should be understood that the global representation vector is a condensation of the node features of the entire graph, but it is a high-dimensional abstract numerical vector that has not been transformed into an intuitively interpretable overall substation status indicator. Therefore, it cannot directly provide a clear reference for maintenance personnel and lacks quantitative standards, making it difficult to compare the substation status across different time periods. Therefore, this application further inputs the global representation vector into the MLP prediction head to map the abstract vector into a quantitative comprehensive score through a multi-layer neural network. This allows the non-linear fitting capability of the MLP to be utilized to uncover the overall substation status patterns in the global vector, transforming it into an intuitive score of 0-100 points, facilitating historical data comparison and providing support for the overall substation maintenance plan.
[0077] Specifically, in one possible embodiment, step S52 is implemented as follows: First, pre-trained MLP prediction head parameters are loaded. These parameters are obtained through training with global vectors representing normal operation, pre-fault states, and manually labeled scores. Second, the global representation vector is input into the MLP input layer, passed through the input layer to the hidden layer, and deep global features are extracted through linear transformation and ReLU activation. Then, the hidden layer output is passed to the output layer, and after linear mapping and normalization, it is transformed into a comprehensive score of 0-100. Next, if the score is outside the range, it is corrected to 0-100 using the clip function. Finally, the score and corresponding health level description are output, and the correspondence between the score and the global vector is stored.
[0078] In the aforementioned digital twin-based intelligent substation status assessment method, step S6 involves tracing the risk path of equipment risk scoring vectors and heterogeneous knowledge graphs to obtain an interpretable subgraph. It should be understood that the equipment risk scoring vectors only output the risk values of each device, leaving maintenance personnel unable to identify the source of risk for high-risk equipment. While the heterogeneous knowledge graph stores node relationships, it is not integrated with the risk scoring and therefore cannot pinpoint the risk propagation path, resulting in low fault diagnosis efficiency. Therefore, this application further combines equipment risk scoring vectors and heterogeneous knowledge graphs to conduct risk path tracing, thereby uncovering the risk transmission links of high-risk equipment and generating an interpretable subgraph. This transforms abstract scoring into visualized correlation paths, clarifies the root causes and transmission relationships of risks, helps maintenance personnel quickly locate problems, and shortens troubleshooting time.
[0079] Specifically, in one possible embodiment, step S6 is implemented as follows: First, high-risk equipment nodes with scores higher than 0.8 are selected from the equipment risk scoring vector to determine the source tracing target. Second, based on the topology of the heterogeneous knowledge graph and the equipment risk scoring vector, an attention-weighted path search algorithm is used to trace back along the graph relationship edges from the high-risk equipment nodes, calculating the contribution of each associated node to the risk of the target equipment. Then, key risk transmission paths are selected according to the contribution threshold, and relevant nodes and relationships are extracted. Next, a subgraph containing the risk source equipment, transmission paths, and target high-risk equipment is constructed. Finally, the equipment type and abnormal characteristics of the risk source are labeled, the electrical connection type and influence weight of the transmission relationship are labeled, an annotated explanatory subgraph is generated, and a visualization file is output.
[0080] In summary, the digital twin-based intelligent substation status assessment method based on the embodiments of this application is explained. First, it constructs a heterogeneous knowledge graph that accurately maps the logical relationships between physical entities based on the substation's ontology definition, asset information, and topology data. Then, massive real-time, rapidly changing data streams are introduced into this graph framework, and temporal feature encoding technology is used to dynamically map node states. Next, a complex message passing and aggregation mechanism is implemented within the graph structure using a relation-aware graph neural network to generate node embedding representations that possess both local individuality and global commonality. Finally, based on this high-dimensional representation, micro-level equipment risk prediction and macro-level comprehensive station assessment are simultaneously completed, and path backtracking and causal analysis are performed on high-risk states based on graph correlations. This solves the problem of the disconnect between equipment health and system operating status, achieving interpretable panoramic status perception and risk warning for intelligent substations.
[0081] Figure 9 This is a block diagram of a digital twin-based intelligent substation condition assessment system according to an embodiment of this application. Figure 9As shown, the digital twin-based intelligent substation status assessment system 100 according to an embodiment of this application includes: a heterogeneous knowledge graph construction module 110, used to construct a heterogeneous knowledge graph based on ontology definitions, asset data, topology data, and historical slow-changing data; a graph node temporal feature encoding module 120, used to encode the graph node temporal features of the real-time fast-changing data stream based on the heterogeneous knowledge graph and the time window length to obtain an initial node feature matrix; a message passing and node embedding module 130, used to perform message passing and node embedding based on a relation-aware graph neural network on the initial node feature matrix based on the heterogeneous knowledge graph to obtain a final node embedding matrix; an equipment risk prediction module 140, used to perform equipment risk prediction on the final node embedding matrix to obtain an equipment risk score vector; a substation comprehensive assessment module 150, used to perform a substation comprehensive assessment on the final node embedding matrix to obtain a substation comprehensive score; and a risk path tracing module 160, used to perform risk path tracing on the equipment risk score vector and the heterogeneous knowledge graph to obtain an interpretable subgraph.
[0082] Here, those skilled in the art will understand that the specific operations of each step in the above-mentioned digital twin-based intelligent substation condition assessment system have been referenced above. Figures 1 to 8 The description of the digital twin-based smart substation condition assessment method is detailed here, and therefore, its repeated description will be omitted.
Claims
1. A method for assessing the condition of a smart substation based on digital twins, characterized in that, include: A heterogeneous knowledge graph is constructed based on ontology definition, asset data, topology data, and historical slowly changing data. Based on heterogeneous knowledge graphs and time window lengths, graph node temporal feature encoding is performed on real-time rapidly changing data streams to obtain an initial node feature matrix. Based on heterogeneous knowledge graphs, message passing and node embedding based on relation-aware graph neural networks are performed on the initial node feature matrix to obtain the final node embedding matrix. Equipment risk prediction is performed on the final node embedding matrix to obtain the equipment risk score vector; A comprehensive evaluation of the substation is performed on the final node embedding matrix to obtain a comprehensive substation score. Risk path tracing is performed on equipment risk scoring vectors and heterogeneous knowledge graphs to obtain an explanatory subgraph.
2. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, Based on ontology definitions, asset data, topological data, and historical slowly changing data, a heterogeneous knowledge graph is constructed, including: Based on the ontology definition, graph ontology modeling and graph initialization are performed to obtain the initialized graph pattern; Static and historical data extraction is performed on asset data, topology data, and historical slowly changing data, and these are then filled into the initialized graph pattern to obtain a filled static graph. Real-time temporal data stream pointer links are used to link the metadata of the time-series database and the populated static graph to obtain a heterogeneous knowledge graph.
3. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, Based on heterogeneous knowledge graphs and time window lengths, graph node temporal feature encoding is performed on real-time rapidly changing data streams to obtain an initial node feature matrix, including: Based on heterogeneous knowledge graphs and time window lengths, the original time-series segments of measurement point nodes are extracted from real-time rapidly changing data streams; The original time-series segments of the measurement point nodes are encoded with temporal features to obtain a measurement point feature mapping table; Based on heterogeneous knowledge graphs and measurement point feature mapping tables, generate equipment feature mapping tables; Perform category feature encoding on other types of nodes in the heterogeneous knowledge graph to obtain a feature mapping table for other nodes; The initial node feature matrix is obtained by integrating the measurement point feature mapping table, equipment feature mapping table and other node feature mapping tables into the overall map node initial feature matrix.
4. The method for assessing the condition of a smart substation based on digital twins according to claim 3, characterized in that, Based on heterogeneous knowledge graphs and measurement point feature mapping tables, a device feature mapping table is generated, including: Extracting the first device node from a heterogeneous knowledge graph; Static attribute encoding is performed on the first device node to obtain the static feature part; Extract the set of associated measurement point node features related to the first device node from the measurement point feature mapping table; Feature vector aggregation is performed on the feature set of associated measurement point nodes to obtain the dynamic feature part; The static and dynamic feature components are fused to obtain the first device node feature.
5. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, Based on heterogeneous knowledge graphs, the initial node feature matrix is processed using message passing and node embedding based on a relation-aware graph neural network to obtain the final node embedding matrix, including: Extract the relation adjacency matrix and relation weight matrix from heterogeneous knowledge graphs; Based on the relation adjacency matrix and relation weight matrix, neighborhood message generation and relation-internal aggregation are performed on the initial node feature matrix to obtain a set of relation aggregated message matrices; Perform cross-relation message fusion on the set of relation aggregated message matrices to obtain the pre-activation embedding matrix; The pre-activation embedding matrix is subjected to nonlinear activation and inter-layer iteration to obtain the final node embedding matrix.
6. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, Based on heterogeneous knowledge graphs, the initial node feature matrix is processed using message passing and node embedding based on a relation-aware graph neural network to obtain the final node embedding matrix, including: In each layer of the graph neural network, for the target node and its neighboring nodes in the graph, an attention coefficient is calculated based on the node features of the target node and its neighboring nodes in the current layer and the learnable weight matrix specific to the relationship. The attention coefficient is used to characterize the importance of the neighboring nodes to the target node in the current running state. The attention coefficients are normalized using the Softmax function to obtain dynamic attention weights, and the messages of neighboring nodes are weighted and aggregated based on the dynamic attention weights to update the embedding vector of the target node. After repeating the above message passing and embedding update process multiple times, the final node embedding matrix containing the embedding vectors of all nodes is obtained.
7. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, Device risk prediction is performed on the final node embedding matrix to obtain the device risk score vector, including: Extract the embedding vector of each device node from the final node embedding matrix; The embedding vectors of each device node are input into a shared prediction head to obtain the device risk score vector.
8. The method for assessing the condition of a smart substation based on digital twins according to claim 1, characterized in that, A comprehensive substation evaluation is performed on the final node embedding matrix to obtain a comprehensive substation score, including: Global graph aggregation is performed on the final node embedding matrix to obtain a global representation vector; The global representation vector is input into the MLP prediction head to obtain the substation comprehensive score.
9. A digital twin-based intelligent substation condition assessment system, characterized in that, include: The heterogeneous knowledge graph construction module is used to construct heterogeneous knowledge graphs based on ontology definitions, asset data, topology data, and historical slowly changing data. The graph node temporal feature encoding module is used to encode the graph node temporal features of real-time fast-changing data streams based on heterogeneous knowledge graphs and time window lengths to obtain an initial node feature matrix. The message passing and node embedding module is used to perform message passing and node embedding based on a relation-aware graph neural network on the initial node feature matrix to obtain the final node embedding matrix based on the heterogeneous knowledge graph. The equipment risk prediction module is used to predict equipment risks from the final node embedding matrix to obtain an equipment risk score vector. The substation comprehensive evaluation module is used to perform a comprehensive evaluation of the substation based on the final node embedded matrix to obtain a comprehensive substation score. The risk path tracing module is used to trace the risk path of equipment risk scoring vectors and heterogeneous knowledge graphs to obtain an interpretable subgraph.
Citation Information
Cited By
Digital twinning-based substation computing power health state assessment and self-healing migration method and system
CN121859748A