Power situation awareness system network asset information checking method based on knowledge graph

By building a knowledge graph of power secondary equipment in the power situation awareness system and combining deep learning technology, the problem of difficulty in data verification of power secondary equipment is solved, data accuracy and quality are improved, and comprehensive verification and dynamic update of power secondary equipment information is achieved.

CN120218204APending Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510246696.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to perform effective data field information verification on power secondary equipment, and cannot accurately analyze the communication characteristics of power secondary equipment, thereby correcting the errors in manual filling of information in the ledger, and facing the problems of difficulty in fusion of multi-source data, difficulty in processing massive communication data, and difficulty in checking equipment information.

Method used

By building a secondary equipment knowledge graph for power professionals, the correlation analysis ability of the knowledge graph and the inference and induction ability of deep learning models can be used to verify the secondary equipment manufacturer information and equipment type information in the situational awareness system. Specific steps include data acquisition and preprocessing, building a knowledge graph for power secondary equipment, building a verification rule base, verifying and completing the knowledge graph based on deep learning, and updating the knowledge graph for feedback results.

Benefits of technology

It solves the problem of complex asset data and inability to effectively verify data, further improves the information accuracy of data, improves the quality of asset data, and realizes comprehensive verification and dynamic update of power secondary equipment information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218204A_ABST
    Figure CN120218204A_ABST
Patent Text Reader

Abstract

The invention discloses a method for checking network asset information of a power situation awareness system based on a knowledge graph. The method comprises the following steps: collecting power secondary equipment data from multiple sources and carrying out data preprocessing; based on a plurality of preset field statistics key indexes, communication messages are selected from the communication data, a communication group is constructed, and typical samples are extracted; by taking the equipment information as a center, constructing nodes and relationships of the power secondary equipment knowledge graph on the ontology layer, and importing the power secondary equipment data after data preprocessing into an instance layer of the knowledge graph according to the structure of the ontology layer; checking nodes and relationships in the knowledge graph according to rules, and checking and complementing nodes, edges and attributes in the knowledge graph based on deep learning; and feeding back results of the rule verification and the deep learning verification to the knowledge graph, and updating the knowledge graph. According to the method, data checking can be effectively carried out, the information accuracy of the data is improved, and the asset data quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information verification, and particularly relates to a method for verifying network asset information of a power situation awareness system based on a knowledge graph. Background Art

[0002] Due to the lack of unified management specifications and standardization requirements in the initial stage of power system construction, there are situations of missing and incorrect ledger information in the current power network security situation awareness system, and the integrity, relevance, and accuracy of the current system information can no longer meet the requirements of modern power network security situation awareness systems.

[0003] Existing solutions based on knowledge graph technology mainly include the following aspects: Standardized management of device information: By constructing device naming specifications and data standardization models, device information in different functional modules is uniformly represented;

[0004] Efficient integration of multi-source data: Using the semantic association ability of the knowledge graph, power device information, asset ledgers, operation status data, and other multi-source information are integrated into the same graph.

[0005] Intelligent completion and verification of device information: Using deep learning and natural language processing technologies, missing device information (such as serial numbers, device models, etc.) is automatically completed, and device information is dynamically verified and updated.

[0006] Construction of a data standardization governance system: Constructing a unified device data governance framework, including a full-process management system for data collection, storage, verification, and sharing, to solve the problem of "data islands".

[0007] Support for advanced query and dynamic analysis: Based on the query language and path analysis function of the knowledge graph, dynamic query and analysis of complex device information are supported.

[0008] Existing technologies lack in-depth data mining, cannot effectively verify data field information of secondary power devices, and cannot accurately analyze according to the communication characteristics of secondary power devices, thereby correcting the manually filled incorrect information in the ledger. The main deficiencies are as follows:

[0009] 1. Difficulty in fusing multi-source data: Other technical solutions can use the semantic association ability of the knowledge graph to integrate power asset ledger information, asset scan data, and communication data into the same graph, but this method only associates the data and does not perform in-depth analysis and mining on the data, fails to overcome the problem of the complexity and diversity of multi-source data, fails to establish a graph structure reflecting the physical communication relationship of devices, and is difficult to provide effective data information for the next step of data verification.

[0010] 2. Difficult to process massive communication data: Each substation has about 200 secondary power equipment, which can generate hundreds of thousands of communication data every day. And the entire power grid system has information of tens of thousands of substations, so the communication volume in a day is extremely huge. Ordinary graph processing is difficult to handle such a large amount of data relationships, and the slow graph update speed is also a major bottleneck in the existing technology;

[0011] 3. Difficult to check and verify device information: The existing technical solutions use deep learning or natural language processing technologies, and use models to complete missing or correct device information. This method depends on the data quality in the preprocessing, and almost no existing technical solutions analyze the structural information of device connections together with other feature information, lacking the unified fusion analysis of various structural data. Moreover, the existing technical solutions cannot perform the update and iteration of the model, and the trained model will not be updated or incrementally trained, lacking the process of model improvement.

[0012] In summary, the solutions of the existing technologies can all achieve information integration in the power system, but lack in-depth data mining, cannot effectively check and verify the data field information of the devices detected by the power network security situation awareness system, cannot accurately analyze according to the communication characteristics of secondary power equipment, and thus cannot correct the wrong manual filling information in the account. They face problems such as difficulty in fusing multi-source data for in-depth comprehensive analysis, difficulty in processing massive communication data, and difficulty in checking and verifying the quality of device data. Summary of the Invention

[0013] In order to overcome the defects and deficiencies of the existing technology, the present invention provides a method for checking and verifying network asset information of a power situation awareness system based on a knowledge graph. By constructing a knowledge graph of secondary power equipment in the power specialty, using the association analysis ability of the knowledge graph and the reasoning and induction ability of the deep learning model, it realizes the checking and verification of the manufacturer information and device type information of secondary equipment in the situation awareness system, solves the problem that asset data is complicated and cannot be effectively checked and verified, further improves the information accuracy of the data, and improves the quality of asset data.

[0014] In order to achieve the above object, the present invention adopts the following technical solutions:

[0015] The present invention provides a method for checking and verifying network asset information of a power situation awareness system based on a knowledge graph, including the following steps:

[0016] Collect power secondary equipment data from multiple sources;

[0017] Perform data preprocessing on the power secondary equipment data;

[0018] Based on a plurality of preset field statistical key indicators, select communication messages from the communication data, construct communication groups and extract typical samples;

[0019] Construct a knowledge graph of secondary power equipment. Centered on device information, construct the nodes and relationships of the knowledge graph of secondary power equipment at the ontology layer, and import the pre-processed secondary power equipment data into the instance layer of the knowledge graph according to the structure of the ontology layer;

[0020] Construct a verification rule library, verify the nodes and relationships in the knowledge graph according to the rules, and verify and complete the nodes, edges and attributes in the knowledge graph based on deep learning;

[0021] Feed back the results of rule verification and deep learning verification to the knowledge graph to update the knowledge graph.

[0022] As a preferred technical solution, collect secondary power equipment data from multiple sources, specifically including:

[0023] Call the API interface to obtain real-time or historical power system device data;

[0024] Extract the basic device information from the device asset management database;

[0025] Obtain communication behavior logs from the ES cluster of the power situation awareness system;

[0026] Obtain device vulnerability scanning and port scanning data through a scanning tool;

[0027] Parse the standardized file to extract the device configuration data and communication protocol information.

[0028] As a preferred technical solution, perform data pre-processing on the secondary power equipment data, specifically including:

[0029] Remove duplicate content in the collected data, clear missing invalid data, perform unified format conversion on data from different sources, and associate the device information in the device asset ledger with communication message data, SCD standard file data, port scanning data, and vulnerability scanning data.

[0030] As a preferred technical solution, statistically calculate key indicators based on a preset number of fields. The fields include date, source substation, destination substation, source IP, destination IP, source port, destination port, and communication protocol. The key indicators include the number of communications and the length of communication messages.

[0031] As a preferred technical solution, construct the nodes and relationships of the knowledge graph of secondary power equipment at the ontology layer. The nodes include: device nodes, sub-device nodes, substation nodes, manufacturer nodes, device type nodes, regional nodes, MAC nodes, security partition nodes, IP nodes, port nodes, and message nodes;

[0032] The said relationship is the connection relationship between devices, including the communication protocols and port information between devices.

[0033] As a preferred technical solution, the rule base includes device naming rules, communication relationship rules, and logical relationship rules between devices.

[0034] As a preferred technical solution, the nodes and relationships in the knowledge graph are verified according to the rules, specifically including:

[0035] Verify the device naming rules, and identify abnormal device names that do not conform to the naming rules through regular expressions and graph database queries;

[0036] Verify the missing attributes, and identify and mark the nodes with missing attributes;

[0037] Verify the logical conflicts, and identify the nodes with logical conflicts.

[0038] As a preferred technical solution, the nodes, edges, and attributes in the knowledge graph are verified and completed based on deep learning, specifically including:

[0039] Infer and complete the missing data based on the graph neural network, complete the node attributes based on the text encoding model, obtain the multi-classification results of the node type judgment through the node classifier based on the multi-layer perceptron, and optimize the encoder parameters through the cross-entropy loss function.

[0040] As a preferred technical solution, infer and complete the missing data based on the graph neural network, and the specific steps include:

[0041] Obtain the initial node features based on the attribute information of the secondary power equipment;

[0042] The graph neural network transmits information through the adjacency matrix and node features, and the feature update formula of the node is:

[0043]

[0044] Among them, N(i) represents the set of neighbor nodes of node v i ; A ij represents the adjacency relationship between node v i and node v j ; C ij represents the normalization factor, which is used to balance the influence of each neighbor node on the update result, b i represents the bias term of node v i ; and σ represents the activation function;

[0045] After T rounds of propagation, the final representation of the node is:

[0046]

[0047] Among them, GNN represents a graph neural network;

[0048] The final node representation is input into a classifier for inference of missing attributes, expressed as:

[0049]

[0050] Among them, represents the predicted device type, and Classifier represents the classifier;

[0051] Infer the potential relationship between node v i and node v j , expressed as:

[0052]

[0053] Among them, represents the predicted communication protocol type, are respectively the final representations of node v i and node v j .

[0054] As a preferred technical solution, node attribute completion is performed based on a text encoding model, specifically including:

[0055] Extract text information from device nodes, encode the text information, and input the text of each device into a pre-trained text encoding model to generate the context representation of the device text, expressed as:

[0056]

[0057] Among them, represents the encoded representation of text information T i obtained in the BERT model, represents the semantic vector of device node v i , Pool represents a pooling operation;

[0058] Input the semantic vector of the device node into a multi-layer perceptron for classification, and output the device attribute classification, expressed as:

[0059]

[0060] Among them, represents the attribute prediction value of device node v i , softmax represents the activation function, W represents the weight matrix of the fully connected layer in the MLP, and b represents the bias term;

[0061] Complete the missing device attributes based on the attribute prediction values and feedback them to the device nodes to update the knowledge graph.

[0062] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0063] By constructing a knowledge graph of secondary equipment in the power specialty and utilizing the correlation analysis ability of the knowledge graph and the reasoning and induction ability of the deep learning model, the present invention realizes the verification of the information of secondary equipment manufacturers and equipment types in the situation awareness system, solves the problem that the asset data is complex and cannot be effectively verified, further improves the information accuracy of the data, and improves the quality of asset data. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic flow chart of the method for verifying network asset information of the power situation awareness system based on the knowledge graph of the present invention;

[0065] Figure 2 It is a schematic diagram of the ontology layer structure of the power secondary equipment knowledge graph of the present invention;

[0066] Figure 3 It is a schematic diagram of the instance layer structure of the power secondary equipment knowledge graph of the present invention;

[0067] Figure 4 It is a schematic flow chart of the rule verification of the present invention;

[0068] Figure 5 It is a schematic flow chart of the process of complementing the knowledge graph data based on deep learning reasoning of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0069] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0070] As Figure 1 shown, a method for verifying network asset information of a power situation awareness system based on a knowledge graph includes the following steps:

[0071] S1: Data collection and preprocessing;

[0072] In this embodiment, relevant data of power secondary equipment is collected through various methods, involving multiple internal and external data sources, specifically including:

[0073] (1) API interface call: By calling the API interfaces provided by internal and external systems, real-time or historical power system equipment data is obtained, including equipment status, fault information, etc.;

[0074] (2) Acquisition of asset database data: Extract basic equipment information from the equipment asset management database, such as equipment name, number, model, manufacturer information, etc.;

[0075] (3) ES cluster communication logs: Obtain communication behavior logs from the ES (Elasticsearch) cluster of the power situation awareness system, and record information such as the communication status and content between devices;

[0076] (4) Data from the situation awareness scanning tool: Obtain vulnerability scanning and port scanning data of devices through the scanning tool to monitor potential security vulnerabilities or open ports of the devices;

[0077] (5) SCD file parsing: Parse standardized files (such as SCD files) to extract the configuration data and communication protocol information of devices, which serve as the basic data for constructing the knowledge graph;

[0078] The goal of data collection in this embodiment is to obtain power system data from multiple sources, ensure the diversity of data sources, and cover key contents such as device information, communication status, and security data.

[0079] S2: Data cleaning and fusion;

[0080] After collecting the data, for multi-source heterogeneous data, it is necessary to perform cleaning, parsing, extraction, and fusion to ensure data quality and consistency, specifically including:

[0081] (1) Data deduplication: Remove duplicate content in the collected data to ensure the uniqueness of each piece of data;

[0082] (2) Cleaning of invalid data: Clear data missing the device unique encoding (GUID) or other key fields to avoid incorrect inferences caused by data missing;

[0083] (3) Data normalization processing: Perform unified format conversion on data from different sources to make it conform to the standard format required by the knowledge graph, ensuring data consistency and accuracy;

[0084] (4) Data association: Associate the device information in the device asset ledger with communication message data, SCD standard file data, port scanning data, vulnerability scanning data, etc. to form a complete device information view;

[0085] The goal of this process is to construct an accurate and standardized device information set to provide reliable data support for the subsequent construction and verification of the knowledge graph;

[0086] S3: Extraction and analysis of typical communication data;

[0087] In large-scale power communication data, extracting typical communication data helps to better understand the communication behavior between devices and provides key data for the construction of the knowledge graph. This step includes statistical analysis of the data and extraction of typical samples, specifically including:

[0088] S31: The statistical analysis of communication data is carried out according to multiple fields (such as date, source substation, destination substation, source IP, destination IP, source port, destination port, communication protocol, etc.), and key indicators are statistically analyzed, such as the number of communications, the length of communication messages, etc.;

[0089] Among them, the calculation formula for the number of communications is expressed as:

[0090]

[0091] Among them, Ⅱ(comm_i) represents the indicator function of the communication message. If the message comm_i exists, then Ⅱ(comm i ) = 1, otherwise it is 0;

[0092] The calculation formula for the length of the communication message is expressed as:

[0093]

[0094] Among them, len(comm_i) represents the length of the i-th communication message, and n is the number of communication messages statistically analyzed;

[0095] S32: Data sampling and extraction of typical communication samples;

[0096] In a large amount of communication data, random sampling can select the most representative communication samples. Assume that k communication messages are randomly selected from all communication data to construct a communication group and extract typical samples.

[0097] Among them, the calculation formula for random sampling is expressed as:

[0098] Sample = {comm i |i ∈ Random(1, 2,..., n), |Sample| = k}

[0099] Among them, Random(1, 2,..., n) represents randomly selecting k messages from all communication messages, and n is the total number of communication messages;

[0100] S4: Construct a knowledge graph of secondary power equipment;

[0101] Based on the cleaned and analyzed data, start to construct a knowledge graph of secondary power equipment. This process includes the construction of the ontology layer and the data import of the instance layer, specifically including:

[0102] S41: Ontology layer construction;

[0103] Such as Figure 2As shown in the figure, the ontology layer defines the nodes and relationships of the knowledge graph of secondary power equipment. The node definitions include equipment nodes, substation nodes, manufacturer nodes, communication protocol nodes, etc., and the relationships represent the logical connections between nodes. The construction of nodes and relationships adopts the graph database model, on which reasoning and querying can be performed. To construct the ontology layer of the knowledge graph of secondary power equipment, a graph is established with equipment information as the center. The nodes designed in this embodiment include equipment nodes, sub-equipment nodes, substation nodes, manufacturer nodes, equipment type nodes, regional nodes, MAC nodes, security partition nodes, IP nodes, port nodes, and message nodes.

[0104] S42: Import of instance layer data and construction of relationships;

[0105] As Figure 3 shown, according to the result after data cleaning and fusion, the data is imported into the instance layer of the knowledge graph according to the structure of the ontology layer to construct the specific relationships between nodes. The specific steps are as follows:

[0106] (1) Data import: According to the design of the ontology layer, convert the data of each type of equipment, communication protocol, etc. into nodes of the knowledge graph and import their relevant attributes;

[0107] (2) Relationship construction: According to the actual connection relationships between devices, such as communication protocols and port information between devices, establish the connections between devices;

[0108] The import of instance layer data and the construction of relationships ultimately form a detailed and accurate knowledge graph of secondary power equipment, which can provide strong support for subsequent verification and reasoning;

[0109] Construct the instance layer of the knowledge graph of secondary power equipment, import the cleaned and associated data according to the design of the ontology layer of the knowledge graph of secondary power equipment, and construct the relationships between nodes;

[0110] S5: Verification of data assets and verification of secondary equipment names;

[0111] Using the constructed knowledge graph of secondary power equipment, two automatic verification methods, namely rule verification and deep learning model verification, are formed to achieve comprehensive verification of secondary equipment data. This step comprehensively verifies the data of secondary power equipment through the combination of rule verification and deep learning technology to ensure that the relationships and attributes between nodes are correct;

[0112] S51: Rule verification;

[0113] (1) Rule library construction: According to the empirical knowledge of power domain experts, construct a verification rule library, including equipment naming rules, communication relationship rules, equipment inter-logical relationship rules, etc. Each rule is defined through the query language of the graph database to form a template for verification rules;

[0114] (2) Check type;

[0115] Equipment name check: Check whether the equipment name conforms to the specification, and whether there are spelling mistakes or non-standard formats;

[0116] Attribute check: Check whether the key information of the equipment (such as manufacturer, equipment type, protocol, etc.) is complete, and whether there are missing attributes;

[0117] Logical conflict check: Detect whether the connection relationship between devices meets the expectation, for example, whether the communication protocols between devices match, and whether the ports conform to the regulations, etc.;

[0118] (3) Checking process: Check the nodes and relationships in the knowledge graph according to the rules, identify the abnormal, missing or logically conflicting nodes, and generate a detailed checking report;

[0119] Such as Figure 4 shown, the rule check mainly includes the following aspects:

[0120] Step S511: Equipment naming rule check;

[0121] Check for abnormal names. The equipment name does not conform to the naming specification or there are spelling mistakes, such as inconsistent spelling of the manufacturer name, non-standard abbreviation of the equipment type, etc. Secondary electrical equipment usually has certain naming rules. For example, the equipment name should include the manufacturer name, equipment type, model, area code, etc. Through regular expressions and graph database queries, equipment names that do not conform to the naming rules can be quickly identified;

[0122] 1. Regular expression check;

[0123] Assume that the specification of the equipment name is: <manufacturer name>-<equipment type>-<equipment model>, where the manufacturer name consists of capital letters and the equipment model consists of letters and numbers. The regular expression can be expressed as follows

[0124] \text{RegEx}(\text{DeviceName})=\text{^[A-Z]+-[A-Za-z0-9]+-[A-Za-z0-9]+$};

[0125] Among them, ^[A-Z]+ means that the manufacturer name consists of one or more capital letters, [A-Za-z0-9]+ means that the equipment type and model can contain letters and numbers, and $ means the end of the string. Through the regular expression, it can be verified whether the equipment name conforms to this format;

[0126] 2. When the equipment name does not conform to the specification, record the abnormality and generate a revision suggestion;

[0127] For example, the device name ABB-Transformer-ModelX complies with the specification, while the spelling mistake ("Transfomer") in ABB-Transfomer-ModelX will be marked as a naming error. The system will automatically record the exception and generate a revision suggestion;

[0128] Step S512: Attribute missing check;

[0129] Check for missing attributes. Missing key attribute information of the device (such as manufacturer, device type, communication protocol, etc.) may lead to incomplete node or edge information in the knowledge graph. Through rule checking, nodes with missing attributes can be identified and marked to prompt the user to supplement. The key attribute information of the device (such as manufacturer, device type, communication protocol, etc.) may be missing, resulting in incomplete node or edge information in the knowledge graph. Rule checking can detect these missing attributes, mark them, and prompt for supplementation. The checking rules for key attributes are as follows:

[0130] Attribute checking formula: For each device node, check whether necessary attributes (such as manufacturer, model, communication protocol, etc.) exist. Missing attributes will be marked as missing and the user will be prompted to supplement them. Specifically, it is expressed as:

[0131]

[0132] where Node i is the i-th device node, and Attribute j is the attribute that the node should have. If an attribute is missing, the result will be the list of this attribute;

[0133] For example, if the device Device1 is missing the manufacturer attribute, then:

[0134] Missing_Attributed(Device1) = {manufacturer}

[0135] The system will automatically prompt "Device Device1 lacks manufacturer information";

[0136] Step S513: Logical conflict check;

[0137] Check for logical conflicts. There are logical problems in the connection relationships between devices that do not meet expectations. For example, the communication protocols between two devices do not match, or there are unexplainable many-to-many relationships. Rule checking can clearly point out the specific location of the logical conflict and provide optimization suggestions.

[0138] There may be conflicts in the logical relationships between devices. For example, the communication protocols between two devices do not match, or there are unexplainable many-to-many relationships. Rule checking will check these potential logical errors to ensure the consistency of the communication relationships and configurations between devices.

[0139] Communication protocol verification: Assume that devices should communicate with each other following a specific protocol. If the communication protocol of a certain device is inconsistent with the preset rules, the system will mark a logical conflict.

[0140] Protocol conflict verification formula:

[0141] Protocol_Conflict(i,j) = Ⅱ(Protocol i ≠Protocol j )

[0142] where Ⅱ() represents the indicator function. If the protocols between device i and device j are inconsistent, the result is 1, indicating a protocol conflict.

[0143] For example, if the Modbus protocol should be used between Device1 and Device2, but in fact Device1 uses the DNP3 protocol, then:

[0144] Protocol_Conflict(Device1,Device2) = 1

[0145] The system will automatically mark this conflict and provide repair suggestions.

[0146] Step S514: Verification feedback and reporting;

[0147] Verification result feedback. The rule verification results are generated in the form of a report, including details of the abnormal data, missing attributes, and logical conflicts found. The verification results can be used for manual review and further adjustment of the rules or data to ensure the accuracy and comprehensiveness of the verification. The report content includes:

[0148] 1. Identification of abnormal data

[0149] 2. Devices with missing attributes

[0150] 3. Location and solution of logical conflicts

[0151] S52: Implement inference and completion of knowledge graph data based on deep learning technology;

[0152] Combined with deep learning technology, using the characteristics of graph neural networks and deep learning neural networks, the nodes, edges, and attributes in the knowledge graph are verified and completed. Deep learning technology, especially graph neural networks (GNN) and BERT models, play a crucial role in verification and inference completion. Through these methods, complex graph-structured data can be processed, and intelligent inference can be performed on missing device information, relationships, or attributes.

[0153] Such asFigure 5 As shown in Figure 5 , the main steps of this process include the inference completion of the graph neural network and the text information completion based on BERT, specifically including:

[0154] S521: Inference completion of the graph neural network (GNN);

[0155] The graph neural network (GNN) is a type of deep learning model that can directly operate on graph-structured data. It propagates features between the nodes of the graph through a message passing mechanism, thereby achieving the completion of missing data. In the knowledge graph of secondary power equipment, GNN can infer the potential connections and missing information between devices through the connection relationships between devices.

[0156] 1. Initialization of node features: The initial node features can be directly obtained from the attribute information of the device, such as device type, manufacturer, model, etc., or feature vectors can be extracted from the text description of the device;

[0157] 2. Information passing and node update: The graph neural network passes information through the adjacency matrix A and node features . In the t-th round, the feature update formula of node v i is:

[0158]

[0159] where N(i) represents the set of neighbor nodes of node v j , A ij represents the adjacency relationship between node v i and node v j (for example, 0 or 1, indicating whether there is a connection). C ij represents the normalization factor, which is used to balance the influence of each neighbor node on the update result, b i represents the bias term of node v i , and σ represents the activation function, usually using ReLU or other non-linear activation functions;

[0160] 3. Feature propagation and convergence;

[0161] After multiple rounds of information passing, the features of the nodes will gradually fuse the information of the neighbor nodes, and finally learn the global representation of the node. After T rounds of propagation, the final representation of the node is:

[0162]

[0163] 4. Inference and completion:

[0164] In the device knowledge graph, missing attributes or relationships can be complemented through the inference of node feature vectors. Based on the final feature vectors of the nodes the type, attributes of the device, or potential relationships between devices can be predicted;

[0165] (1) Inference of missing attributes: If a certain attribute of a device node (such as device type, manufacturer) is missing, the missing attribute can be inferred through its final feature vector For example, if the attribute representing the device type of a node is missing, the final node representation can be input into a classifier for inference. The attribute inference formula is expressed as:

[0166]

[0167] where, represents the predicted device type;

[0168] (2) Inference of potential relationships between devices: GNN can also be used to infer the relationships between devices.

[0169] For example, there may be a certain communication protocol or data flow relationship between device v i and device v j GNN infers these relationships through the connection information between devices. Finally, the edges between nodes can represent the relationships between devices, such as communication protocols, connection status, etc. If the communication protocol between device v i and device v j is missing, based on their node representations, the type of communication protocol they should use can be inferred. The relationship inference formula is expressed as:

[0170]

[0171] where, represents the predicted type of communication protocol;

[0172] S522: The BERT model performs node attribute completion;

[0173] In addition to graph neural networks, text encoding models (such as BERT) also play an important role in processing device description information (such as text information about device types, manufacturers, models, etc.) in the power device knowledge graph. BERT can generate semantic vectors of device information through the understanding of text context, thereby filling in the missing attributes;

[0174] 1. Extraction of device text information;

[0175] In the power secondary equipment knowledge graph, each equipment node usually comes with certain text information, such as equipment name, manufacturer, model, function description, etc. This text information is usually missing or inconsistent, so it can be used as input data for the BERT model to complete;

[0176] Device text information representation: Extract text information T from the device node i , such as descriptive texts like the name, manufacturer, model of the device. Assume the device v i has a descriptive text of:

[0177] T i = "″Siemens 7SJ600 protection relay"

[0178] Among them, T i represents the descriptive text of device v i ;

[0179] 2. Text information encoding;

[0180] The BERT model encodes the text information. The text input of each device is fed into the pre-trained BERT model to generate the context representation of the device text. The input of BERT needs to be tokenized first. Usually, the WordPiece tokenizer is used for splitting, and then the split text is represented as Token IDs and input into the BERT model;

[0181] Text input representation: Convert the device descriptive text T i into a token sequence {t1, t2,..., t n} through a tokenizer (such as WordPiece), where n represents the length of the text. Then, these tokens are passed into the pre-trained BERT model to obtain the context representation corresponding to each token. BERT processes the context information of each token through a multi-layer Transformer structure.

[0182] BERT encoding output: The BERT model encodes each token through the self-attention mechanism of Transformer to obtain the context representation vector of the token. Assume the encoding result of the text corresponding to device v i is:

[0183]

[0184] Among them, represents the encoding representation obtained by the text information T i in the BERT model, which contains the context information of the device text.

[0185] 3. Generate the semantic vector of the device node

[0186] The output of BERT is a set of context representation vectors, and each token has a corresponding vector. To obtain the semantic vector of the entire device, the vectors of all tokens output by BERT can be pooled. Usually, the vector corresponding to the [CLS] token is used as the overall representation of the device, or global pooling (such as average pooling) is used to obtain the semantic vector of the device;

[0187] Node representation generation: For the device node v i , the semantic vector generated by BERT can be expressed as:

[0188]

[0189] where Pool represents the pooling operation, which can be taking the [CLS] vector, or taking the mean or weighted average of each token vector, so as to obtain the semantic vector of device v i which contains the context information, function description, manufacturer name, etc. of the device; It contains the context information, function description, manufacturer name, etc. of the device;

[0190] 4. Use the MLP model for node attribute classification;

[0191] After obtaining the semantic vector of the device, it is input into a multi-layer perceptron (MLP) classifier to infer the missing attributes of the device. For example, attributes such as device type, manufacturer, and model may be missing. This method uses MLP to complete these missing attributes;

[0192] Attribute classification: Set an attribute of the device node (such as device type, manufacturer) as the classification target, and input the semantic vector of the device into a multi-layer perceptron (MLP) for classification. Suppose it is necessary to predict the attribute y i of device v i such as device type), and use the following classification formula:

[0193]

[0194] where, represents the predicted value of the attribute of device v i such as device type), and is the model for attribute classification based on the device semantic vector;

[0195] Structure of the MLP model: MLP usually includes multiple fully connected layers and non-linear activation functions. Suppose has a dimension of d, the input of MLP is and the output is k categories (for example, there are k possible categories for device type), and the output layer of MLP is the softmax layer:

[0196]

[0197] Among them, W represents the weight matrix of the fully connected layer in the MLP, b represents the bias term, softmax represents the activation function, and the predicted probability distribution of the output device attributes;

[0198] 5. Attribute Completion and Verification;

[0199] Through the above process, the missing device attributes (such as manufacturer, device type, etc.) can be predicted. If the attributes of a device node (such as the manufacturer) are missing, the relevant attributes of the device can be inferred through the BERT model. For example: Assume that the manufacturer attribute of device v i is missing. The BERT and MLP models will predict the manufacturer through the device text information and output the predicted value y i . The missing device attributes will be completed and fed back to the device node to update the device knowledge graph.

[0200] S523: Input this representation into a node classifier based on a multi-layer perceptron to obtain a multi-classification result for node type judgment;

[0201] S524: Continuously optimize the encoder parameters through the cross-entropy loss function;

[0202] The cross-entropy loss function is a commonly used loss function for classification problems in deep learning, and its formula is:

[0203]

[0204] Among them, N represents the number of categories in the node classifier, y i represents the true label (represented by one-hot encoding, where only the position corresponding to the target category is 1 and other positions are 0), represents the predicted probability of the model;

[0205] The cross-entropy loss function measures the accuracy of the model output by measuring the difference between the predicted distribution and the true distribution.

[0206] S6: Feedback and Update of Verification Results. The results of rule verification and deep learning verification are fed back to the knowledge graph in real time. For the discovered problems, the system generates detailed recommended update content, including name revision, attribute completion, and relationship adjustment. After confirmation through manual review or automatic verification, the revised results are updated to the knowledge graph; at the same time, the system records all verification historical data to provide reference for subsequent rule optimization and deep learning model update, forming a closed-loop mechanism of verification - feedback - optimization to ensure the continuous improvement and dynamic update of the knowledge graph.

[0207] In this embodiment, the information verification, inference completion, and dynamic optimization of secondary power equipment are realized by combining deep learning technology and rule verification. Data from different sources are collected through various means, including asset ledger data, communication message data, SCD standard file data, port scan data, vulnerability scan data, and port opening data.

[0208] Based on the standardized data model, the classification of entity nodes is defined, and an ontology layer knowledge graph including device entities, manufacturer entities, and communication entities is constructed. The constructed knowledge graph is imported into the graph database to form a structured data set.

[0209] Based on data verification, combined with rule verification and deep learning technology, a comprehensive verification and completion of the nodes and relationships in the knowledge graph are realized. Rule verification relies on the experience of domain experts and industry standards, defines verification logics including naming rules, communication relationship rules, and logical constraint rules, identifies abnormal device naming, missing attribute information, and logical conflicts, and realizes the dynamic optimization of rules through the rule management module, achieving accurate verification and dynamic optimization of device information. The deep learning technology, through graph neural networks and text models, conducts intelligent reasoning on the data in the knowledge graph based on deep learning technology, obtains the semantic and structural features of nodes through text encoding and structure encoding models, and inputs these multi-modal features into the model classifier for reasoning calculation. Record historical verification data and continuously optimize the deep learning model to form a closed-loop mechanism of inference completion, feedback update, and model optimization. Screen and optimize the verification and inference results through model calculation and comprehensive threshold determination, significantly improving the accuracy and consistency of secondary power equipment information, enhancing the practicality of the knowledge graph, and providing efficient data support for the intelligentization, real-time monitoring, fault diagnosis, and resource optimization configuration of power grid operation.

[0210] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for checking network asset information of a power situation awareness system based on knowledge graph, characterized in that: The steps include: Collect power secondary equipment data from multiple sources; Preprocess the data of power secondary equipment; Count key indicators based on multiple preset fields, select communication messages from communication data, build communication groups and extract typical samples; Construct the knowledge graph of power secondary equipment. Centered on equipment information, construct nodes and relationships of the knowledge graph of power secondary equipment in the ontology layer, and import the pre-processed power secondary equipment data into the instance layer of the knowledge graph according to the structure of the ontology layer. Build a verification rule library, verify the nodes and relationships in the knowledge graph according to the rules, and verify and complete the nodes, edges and attributes in the knowledge graph based on deep learning; Feed the results of rule verification and deep learning verification back to the knowledge graph to update the knowledge graph.

2. According to the knowledge graph-based power situation awareness system network asset information verification method of claim 1, it is characterized in that: Collect power secondary equipment data from multiple sources, including: Call the API interface to obtain real-time or historical power system equipment data; Extract basic equipment information from the equipment asset management database; Obtain communication behavior logs from the ES cluster of the power situation awareness system; Obtain device vulnerability scan and port scan data through scanning tools; Parse the standardized files to extract the device configuration data and communication protocol information.

3. According to the knowledge graph-based power situation awareness system network asset information verification method of claim 1, it is characterized in that: Data preprocessing of power secondary equipment data includes: Remove duplicate content from collected data, clear missing invalid data, convert data from different sources into a unified format, and associate device information in the device asset ledger with communication message data, SCD standard file data, port scan data, and vulnerability scan data.

4. According to the knowledge graph-based power situation awareness system network asset information verification method of claim 1, it is characterized in that: Key indicators are counted based on multiple preset fields. The fields include date, source plant, destination plant, source IP, destination IP, source port, destination port, and communication protocol. Key indicators include number of communications and length of communication messages.

5. According to the knowledge graph-based power situation awareness system network asset information verification method of claim 1, it is characterized in that: Construct nodes and relationships of the knowledge graph of power secondary equipment in the ontology layer, the nodes include: equipment node, sub-equipment node, plant node, manufacturer node, equipment type node, region node, MAC node, security partition node, IP node, port node, message node; The relationship is the connection relationship between devices, including the communication protocol and port information between the devices.

6. The method for verifying network asset information of a power situation awareness system based on knowledge graph according to claim 1 is characterized in that: The rule base includes device naming rules, communication relationship rules, and logical relationship rules between devices.

7. The method for verifying network asset information of a power situation awareness system based on knowledge graph according to claim 1 is characterized in that: Check the nodes and relationships in the knowledge graph according to the rules, including: Verify device naming rules and identify abnormal device names that do not comply with the naming rules through regular expressions and graph database queries; Check for missing attributes, identify and mark nodes with missing attributes; Check for logical conflicts and identify nodes with logical conflicts.

8. The method for verifying network asset information of a power situation awareness system based on knowledge graph according to claim 1 is characterized in that: Based on deep learning, the nodes, edges and attributes in the knowledge graph are checked and completed, including: Missing data is completed based on graph neural network reasoning, node attributes are completed based on the text encoding model, and multi-classification results of node type judgment are obtained based on the node classifier of the multi-layer perceptron. The encoder parameters are optimized through the cross entropy loss function.

9. The method for verifying network asset information of a power situation awareness system based on knowledge graph according to claim 8 is characterized in that: Completing missing data based on graph neural network reasoning, the specific steps include: Obtaining initial node characteristics based on attribute information of power secondary equipment; The graph neural network transmits information through the adjacency matrix and node features. The node feature update formula is: Where N(i) represents the node v i The neighbor node set A ij Represents node v i With node v j The adjacency relationship between ij represents the normalization factor, which is used to balance the influence of each neighbor node on the update result, b i Represents node v i The bias term, σ represents the activation function; After T rounds of propagation, the final representation of the node is: Among them, GNN represents graph neural network; The final node representation is input into the classifier for inference of missing attributes, expressed as: in, Indicates the predicted device type, and Classifier indicates the classifier; Reasoning Node v i With node v j The potential relationship between them is expressed as: in, Indicates the predicted communication protocol type, The nodes v i With node v j final expression of .

10. The method for verifying network asset information of a power situation awareness system based on knowledge graph according to claim 8 is characterized in that: Node attribute completion is performed based on the text encoding model, including: Extract text information from device nodes, encode the text information, input the text of each device into the pre-trained text encoding model, and generate the contextual representation of the device text, which is expressed as: in, Indicates text information T i The encoded representation obtained in the BERT model, Represents the device node v i The semantic vector Pool represents pooling operation; The semantic vector of the device node is input into the multi-layer perceptron for classification, and the device attribute classification is output, which is expressed as: in, Represents the device node v i The attribute prediction value of , softmax represents the activation function, W represents the weight matrix of the fully connected layer in MLP, and b represents the bias term; The missing device attributes are completed based on the attribute prediction values ​​and fed back to the device node to update the knowledge graph.

Citation Information

Cited By

  • Method and system for constructing Neo4j knowledge graph based on LLM natural language

    CN120744136A

  • Power plant relay protection setting value checking method and system fused with deep learning

    CN120851903A

  • Data management method, model, system, product and equipment based on large language model and adaptive knowledge graph

    CN121009083A

  • Multi-source property right data verification method

    CN121188439A

  • Point table generation method and device, equipment and medium

    CN121279422A