A knowledge graph construction method for power secondary equipment and related equipment
By constructing a multi-source data fusion mechanism for power secondary equipment and a knowledge conflict identification mechanism driven by a large model, the problems of heterogeneity and semantic conflict of multi-source data in power secondary equipment are solved, a high-precision knowledge graph is constructed, and the level of intelligence and application efficiency of power grid asset management are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2026-03-09
- Publication Date
- 2026-07-21
Smart Images

Figure CN122433862A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and information verification technology, and in particular to a method for constructing a knowledge graph for secondary power equipment and related equipment. Background Technology
[0002] With the continuous advancement of smart grid construction, the number of secondary equipment in the power grid, such as protection devices, monitoring and control devices, and communication equipment, is rapidly increasing. Key data on the operating status, configuration parameters, topology connections, and historical fault information of these devices are widely distributed across multiple independent business systems, including power grid resource management systems, production management systems, power monitoring and data acquisition systems, and network security audit systems. Due to the lack of unified data standards and sharing mechanisms in the initial design of these systems, this data exhibits significant structural heterogeneity and semantic inconsistency, forming serious "information silos" and posing significant challenges to the full lifecycle management and intelligent operation and maintenance of equipment.
[0003] To achieve effective organization and in-depth utilization of multi-source information from secondary power equipment, knowledge graph technology has been gradually introduced into this field in recent years. Existing knowledge graph-based solutions mainly revolve around two directions: On the one hand, through the semantic modeling capabilities of knowledge graphs, the static attributes (such as model, manufacturer, and installation location), dynamic operational data (such as alarm information and communication behavior), and relationships (such as physical connections and logical affiliations) of secondary power equipment are uniformly expressed, realizing the integration and association of multi-source heterogeneous data; on the other hand, a standardized data governance system is constructed, establishing a full-process management mechanism covering data collection, storage, verification, and sharing, aiming to break down "data silos," improve data consistency and availability, and support upper-level intelligent applications.
[0004] However, existing technologies still have significant shortcomings when applied to power secondary equipment scenarios. First, the quality of data sources is low. Internal enterprise data is widely distributed across multiple databases, including structured, semi-structured, and even unstructured data types. Many key fields rely on manual input and maintenance, easily leading to issues such as missing equipment attribute information, input errors, or inconsistent formats. This results in inconsistent overall data quality, severely impacting subsequent knowledge extraction and fusion. Second, the dimensions of data related to power secondary equipment are complex, including not only the inherent attributes of the equipment itself but also dynamic information such as communication behaviors, access control policies, and security events generated during operation. These diverse dimensions make it difficult for traditional methods to effectively integrate and model their multi-dimensional relationships. Furthermore, the phenomenon of "synonyms with different names" is common in multi-source data. Due to different naming conventions, these data appear in different forms in different systems, making it impossible for the system to automatically identify their equivalence. Solving such problems often heavily relies on the experience and judgment of domain experts, requiring significant manpower for rule formulation and semantic alignment, resulting in low efficiency and poor scalability.
[0005] Therefore, for the typical multidimensional and heterogeneous scenario of power secondary equipment, there is an urgent need for a knowledge graph construction method that can automatically handle semantic conflicts, improve attribute consistency, and support knowledge fusion and ablation, so as to realize the automated fusion of multi-source data, semantic consistency calibration, and accurate expression of equipment knowledge, thereby providing reliable knowledge support for the intelligent operation and maintenance, asset verification, and risk warning of the power grid. Summary of the Invention
[0006] The main objective of this application is to propose a knowledge graph construction method, electronic device, storage medium, and program product for power secondary equipment. This addresses the problems in existing technologies, such as difficulties in asset information fusion, inaccurate entity alignment, and ineffective integration of expert knowledge due to the strong heterogeneity of multi-source data, low data quality, and severe semantic conflicts. By constructing a multi-source data fusion mechanism and a knowledge conflict identification and attribute value consistency mechanism driven by a large model, the application achieves automated integration of data related to power secondary equipment, semantic consistency calibration, and high-precision modeling of the knowledge graph. This improves the intelligence and reliability of power grid asset verification, supporting the application needs of power monitoring systems in scenarios such as situational awareness, security assessment, and intelligent operation and maintenance.
[0007] To achieve the above objectives, one aspect of this application proposes a method for constructing a knowledge graph for secondary power equipment, the method comprising: Collect multi-source data related to secondary power equipment; The multi-source data is preprocessed and aggregated to obtain standardized data; Based on the standardized data, a unified asset information database is constructed through progressive multi-dimensional association rules. The progressive multi-dimensional association rules include IP-MAC binding association, IP association, device name fuzzy matching and feature hierarchical matching executed sequentially within the same site. This associates multi-source data with a unique asset identifier ID, forming a standardized asset file with "one ID and one file". Based on the unified asset information database, a knowledge graph with secondary power equipment as the core entity is constructed, defining entity types, relationship types, and attribute sets; Knowledge divergence ablation is performed on the multi-source data conflicts in the knowledge graph. The divergence ablation process at least partially involves identifying and calibrating conflict attributes through a large model and recording ablation logs. Output the knowledge graph optimized by divergence resolution.
[0008] In some embodiments, the collection of multi-source data related to secondary power equipment specifically includes: Collect traffic monitoring data, which includes source IP address, source port, destination IP address, destination port, and protocol type; Collect asset scanning data, which includes asset information obtained by the situational awareness system, component identification results generated by vulnerability scanning tools and nmap tools, port scanning data, vulnerability scanning results, substation configuration description file data, dispatch management system equipment data, and asset verification data.
[0009] In some embodiments, the preprocessing and aggregation of the multi-source data specifically includes: The collected multi-source data is cleaned, formatted, and deduplicated. For traffic monitoring data, data aggregation is performed based on the five-tuple to identify and eliminate duplicate communication records and extract communication behavior characteristics between devices; The asset scanning data is normalized, and the field naming conventions and data representation formats are standardized to establish a standardized data model covering device IP address, manufacturer, model, service type, open port, and known vulnerability attributes.
[0010] In some embodiments, the construction of a unified asset information database further includes: By accurately associating with site numbers and / or fuzzy matching with site names, the site range to which the device belongs can be clearly defined; Generate a globally unique asset identifier ID containing site information based on the physical and logical attributes of the device, and record the source data and collection time of each attribute; Within the scope of the site, multi-source data is integrated and mapped to the unique asset identifier ID according to the progressive multi-dimensional association rules.
[0011] In some embodiments, the feature hierarchical matching includes: First priority matching: Matching is based on the combination of equipment model and rated parameters; Second priority matching: Matching is based on the combination of logical nodes and associated devices; Third priority matching: Matching is based on a combination of manufacturer, commissioning time, and equipment type.
[0012] In some embodiments, the knowledge disagreement resolution process includes: Detect conflict attributes and distinguish between numerical conflicts and textual conflicts; For numerical conflicts, the conflict attributes and their context information, as well as objective resource data, are input into the large model. The large model evaluates the credibility of each candidate value according to the preset confidence calculation rules and outputs the ablation decision based on the credibility threshold. For text-based conflicts, the conflict attributes, their contextual information, and the standard rule base are input into the large model, which then performs semantic equivalence judgment and normalizes the candidate values that are semantically consistent. For conflicts that cannot be resolved by the large model, mark them as unresolvable conflicts and push them to the expert review module. Receive the correction results after expert review and write them back to the knowledge base. Record the input information, model output, objective resources and rule base version for each ablation, forming a traceable ablation log.
[0013] In some embodiments, the confidence calculation rules include: Perform format validation and standardization on each candidate value; Statistically analyze the number of times each candidate value matches across different objective resources and the time of the most recent match; Time weights are assigned based on the timeliness of the evidence; Calculate the weighted score for each candidate value and normalize it to obtain the confidence level; The confidence level is compared with a preset threshold. If the confidence level is greater than or equal to the first threshold, it is automatically adopted. If the confidence level is between the first threshold and the second threshold, manual review is recommended. If the confidence level is less than the second threshold, it is recommended to retain the conflict and mark it as non-resolvable.
[0014] In some embodiments, after the output optimized knowledge graph, the following is also included: The knowledge graph is deployed to the power system platform, providing a query interface to support applications such as power grid asset verification, security risk assessment, fault tracing and / or abnormal communication detection. Establish a dynamic maintenance mechanism for the knowledge graph, including daily incremental updates and a closed loop of expert operation and maintenance feedback. Optimize the knowledge graph based on data corrected by experts and regularly evaluate the quality of the graph.
[0015] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0016] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.
[0017] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0018] This application has the following advantages over the prior art: This application effectively addresses the core pain points of heterogeneous multi-source data, low quality, and semantic conflicts in the construction of knowledge graphs for power secondary equipment. Through standardized preprocessing and progressive multi-dimensional association rules (IP-MAC binding → IP association → name fuzzy matching → feature-level hierarchical matching), it achieves automated integration and centralized cross-source management of multi-source data such as traffic monitoring and asset scanning, significantly improving the efficiency of asset information fusion. In particular, the introduction of a priority strategy based on power industry patterns (model + rated parameters, logical node + associated equipment, manufacturer + commissioning time + equipment type) in feature-level hierarchical matching significantly improves the accuracy of entity alignment in cases of incomplete or conflicting information.
[0019] By leveraging a unique identifier system containing site information and a large-model-driven knowledge disagreement resolution mechanism, combined with objective resource calibration (traffic logs, nmap scan results, vulnerability scanning tool detection information, recently verified data, etc.) and expert review modules for manual verification and write-back of non-resolvable conflicts, this invention ensures the semantic consistency and modeling accuracy of the knowledge graph. The confidence calculation rules and automated decision thresholds introduced during the large-model resolution process make the resolution process quantifiable and auditable, significantly reducing the need for manual intervention.
[0020] The resulting high-quality knowledge graph can directly support the application of power monitoring systems in scenarios such as situational awareness, security assessment, and intelligent operation and maintenance, providing key technical support for the transformation of power grid operation and maintenance from traditional models to precision and intelligence. Simultaneously, by establishing an expert feedback loop and dynamic maintenance mechanism, the knowledge graph can be continuously optimized, serving the intelligent operation of the power system in the long term. Attached Figure Description
[0021] Figure 1 This is an overall flowchart of the knowledge graph construction method for secondary power equipment in the embodiments of this application.
[0022] Figure 2 This is a schematic diagram of the knowledge divergence resolution process in the embodiments of this application.
[0023] Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0026] like Figure 1 As shown, this embodiment provides a method for constructing a knowledge graph for secondary power equipment, specifically including the following steps: Step S101: Collect multi-source data related to secondary power equipment.
[0027] As a preferred technical solution, multi-source data related to secondary power equipment is collected, specifically including: collecting traffic monitoring data, which includes source IP address, source port, target IP address, target port and protocol type; collecting asset scanning data, which covers asset information obtained by the situational awareness system, component identification results generated by vulnerability scanning tools and nmap tools, port scanning data, vulnerability scanning results, SCD device data, OMS device data and asset verification data.
[0028] Step S102: Preprocess and aggregate the multi-source data.
[0029] As a preferred technical solution, preprocessing and aggregation of multi-source data specifically includes: The collected multi-source data is cleaned, standardized in format, and deduplicated. For traffic monitoring data, data aggregation is performed based on the five-tuple (source IP, source port, target IP, target port, protocol) to identify and eliminate duplicate communication records and extract communication behavior characteristics between devices. Asset scanning data is normalized, and field naming conventions and data representation formats are unified to establish a standardized data model covering attributes such as device IP address, manufacturer, model, service type, open port, and known vulnerabilities.
[0030] Step S103: Construct a unified asset information database.
[0031] As a preferred technical solution, a unified asset information database is constructed. First, the scope of equipment ownership is clarified by accurately associating site numbers and fuzzy matching site names. Then, a globally unique identifier ID containing site information is generated based on the physical and logical attributes of the equipment, and the source of the attributes is recorded to ensure traceability. Finally, within the same site, multi-source data is integrated using progressive multi-dimensional association rules of "IP-MAC binding → IP association → name fuzzy matching → feature hierarchical matching", and finally uniformly mapped to a unique asset ID to form a standardized asset file of "one ID, one file", realizing centralized and traceable management of cross-source data.
[0032] Step S104: Construct a knowledge graph with secondary power equipment as the core entity.
[0033] As a preferred technical solution, a knowledge graph with secondary power equipment as the core entity is constructed. Specifically, based on the asset information database, the entity types in the graph are defined as including equipment nodes, manufacturer nodes, model nodes, IP address nodes, port nodes, vulnerability nodes, and communication relationship nodes; the relationship types are defined as "belongs to", "located in", "running in", "communicating in", "vulnerable", "configured as", and "associated with"; and corresponding attribute sets are configured for each entity node to form the ontology layer and instance layer of the knowledge graph.
[0034] Step S105: Dissolving knowledge differences.
[0035] As a preferred technical solution, knowledge discrepancy resolution is performed, systematically resolving multi-source data conflicts through large-scale model technology: First, numerical and textual attribute conflicts are identified; for numerical conflicts, the large-scale model assesses credibility and selects the optimal value by combining objective resources (traffic logs, nmap scan results, vulnerability scanning tool detection information, recently verified data, etc.) and context; for textual conflicts, semantic understanding and rule base are used for normalization processing, and those that cannot be resolved are marked and submitted to the expert review module for manual review and write-back of unresolvable conflicts, forming an asset verification record; the entire process is recorded in the resolution log to ensure traceability of results and support continuous optimization and verification.
[0036] Step S106: Output the optimized knowledge graph.
[0037] As a preferred technical solution, the optimized knowledge graph is output. Specifically, after completing the knowledge divergence resolution process, a knowledge graph of power secondary equipment with complete structure, consistent semantics, and reliable quality is generated and output for subsequent training and inference of artificial intelligence models, supporting advanced applications such as power grid situation awareness, asset verification, safety risk assessment, and intelligent operation and maintenance.
[0038] The present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0039] This embodiment provides a method for constructing a knowledge graph for secondary power equipment, which specifically includes the following steps: Step S1: Multi-source data acquisition.
[0040] This embodiment, targeting the characteristics of secondary power equipment (such as relay protection devices, measurement and control devices, communication gateways, etc.), collects data from multiple channels, including internal systems, external tools, standardized documents, and asset verification data, to ensure that the data covers basic equipment information, operating status, communication behavior, and security attributes. Specifically, it includes: 1) Traffic monitoring data acquisition: By deploying traffic probes at the core nodes of the substation network, communication data packets between devices are captured in real time. Core fields such as source IP address, source port, destination IP address, destination port and protocol type (such as TCP, UDP, IEC 61850 MMS) are extracted to form a traffic monitoring dataset, ensuring the real-time dynamics of device communication are captured.
[0041] 2) Asset Scanning Data Acquisition: Utilize the situational awareness platform to obtain real-time asset information such as device online status, CPU utilization, and memory usage; use vulnerability scanning tools to batch query device component information for substation network segments (e.g., 192.168.2.0 / 24), identifying operating system types (e.g., embedded Linux 4.1) and application software versions; use nmap tools to perform port scanning (range 1-65535) and vulnerability scanning, recording open ports (e.g., 5000 / TCP, 22 / TCP) and corresponding vulnerabilities (e.g., CVE-2021-26855, high risk level).
[0042] 3) SCD file parsing: Use an XML parsing tool (such as the Python lxml library) to parse the substation SCD (substation configuration description) file, extract the model of secondary equipment (such as RCS-931), rated parameters (rated voltage 220kV, rated current 5A), logical nodes (such as LLN0 and PTRC as defined in the IEC 61850 standard), and the relationship between equipment (such as the signal mapping between line protection devices and merging units).
[0043] 4) OMS Ledger Data Extraction: Connect to the OMS (Dispatch Management System) database of the dispatch center through the interface, execute SQL query statements (such as SELECT Equipment Name, Installation Location, Commissioning Time, Maintenance Record FROM Secondary Equipment Ledger WHERE Plant Number='CS22001') to obtain the equipment name (such as 220kV Line B protection device), installation location (Screen No. 2, Relay Protection Room, 3rd Floor, Main Control Building), commissioning time (2021-08-15), and historical maintenance records (2023-08 Plug-in Replacement, 2024-03 Annual Verification).
[0044] 5) Asset verification data: Data in the unified asset information database after the asset attributes have been modified by experts, synchronizing their attribute values and modification time.
[0045] The goal of data collection in this embodiment is to cover all dimensions of information on secondary power equipment, solve the problem of incomplete information from a single data source, and provide a complete data foundation for subsequent knowledge graph construction.
[0046] Step S2: Multi-source data preprocessing and aggregation.
[0047] For the collected heterogeneous data (structured ledger data, semi-structured SCD files, and unstructured traffic logs), data quality is improved through data cleaning, standardization, and aggregation processes, specifically including: 1) Data Cleaning: Duplicate data is identified using hash value comparison, such as deleting duplicate records in the OMS ledger where the "equipment name + commissioning time" are exactly the same; data missing key fields (such as equipment IP, model) are marked as "to be completed" and temporarily stored in a temporary database; abnormal data is filtered based on power industry standards, such as removing invalid records where the port number is outside the range of 0-65535 or the IP address does not conform to the IPv4 format (such as 192.168.2.256); in addition, the rationality of business logic is checked, such as the equipment commissioning time must not be a future date, and key status fields must not be empty or filled with placeholders such as "unknown" or "NULL"; all records identified as invalid are isolated to the abnormal data area, a cleaning log is recorded, the specific error type is marked (such as "IP format error" or "port out of bounds"), and they are handled according to severity—for repairable format problems (such as extra spaces or incorrect capitalization), automatic correction is attempted.
[0048] 2) Traffic data aggregation: Using the five-tuple of "source IP + source port + destination IP + destination port + protocol" as the aggregation key, the Spark Streaming framework is used to aggregate traffic data in real time, and to count the communication frequency (e.g., a line protection device communicates with the dispatch center server 12 times every 5 minutes), the average message length (e.g., 350 bytes / message), and the communication success rate (e.g., 99.8%) within 5 minutes. A device communication behavior feature table is generated and stored in Hive.
[0049] 3) Data Standardization: Unify field naming conventions, merge "Manufacturer", "Supplier" from different data sources into the "Manufacturer" field, and unify "Equipment Number" and "Asset Number" into "Unique Equipment Identifier"; standardize data formats, convert dates to "YYYY-MM-DD" format (e.g., change "2021.08.15" to "2021-08-15"), unify IP addresses to dotted decimal format, and organize vulnerability numbers according to the CVE standard format (e.g., CVE-2021-26855).
[0050] Step S3: Construct a unified asset information database.
[0051] The unified asset information repository serves as the foundational data platform for the knowledge graph, storing cleaned and normalized multi-source asset information. Based on situational awareness asset data, it first clarifies the ownership scope of devices through "site association," then constructs a unique identifier system and progressive multi-dimensional association rules based on "IP-MAC binding → IP association → name fuzzy matching → feature-layered matching," simultaneously recording multi-source data traceability information to achieve traceable centralized management of cross-source data. Specifically, this includes: 1) Site Association: Clarify the scope of device ownership. Prioritize associating devices from each data source with specific sites to narrow down the scope for subsequent unique identifier generation and data association, avoiding cross-site confusion. The specific process is as follows: 1.1) Precise association of site numbers: Extract the site number (such as "Sub-220-01" or "Station-110-05") of the device from various data sources (traffic monitoring data, asset scanning data, SCD files, OMS ledgers, etc.) and directly bind the device to the corresponding site to ensure clear ownership.
[0052] 1.2) Fuzzy Association of Site Names: For data sources that do not record site numbers, extract the site name to which the equipment belongs (e.g., "Chengdong 220kV Substation" and "Xihu 110kV Substation"). Use fuzzy matching rules that include core fields (e.g., "Chengdong", "Xihu", and "220kV Substation") and combine them with a preset site name thesaurus (e.g., "Chengdong Substation" corresponds to "Chengdong 220kV Substation") to complete the association and binding between the equipment and the site.
[0053] After all devices have completed site association, a preliminary "site-device" mapping relationship is formed according to site category, laying the foundation for subsequent operations.
[0054] 2) Unique Identifier Generation: Construct a globally unique identifier for the device. Based on the associated site information and device attributes, a unique asset identifier ID is generated to ensure normalized device identification across data sources. The specific process is as follows: 2.1) Attribute extraction: Extract attributes such as MAC address, hardware serial number, factory code, IP address, and device name of the device, and synchronously record the source data of each attribute (such as MAC address from traffic monitoring data) and collection time, forming a "attribute value - data source - collection time" binding unit.
[0055] 2.2) UUID Generation: For devices that have completed site association, a 36-bit device-level unique identifier (UUID) is generated using the UUID v4 algorithm to ensure that there is no risk of duplication among different devices within the same site.
[0056] 2.3) Identifier Combination: The “Site Number (Site Identifier)” and “UUID” are concatenated in the format “[Site Identifier]-[UUID]” to form a globally unique asset identifier ID, which is also associated with the attribute binding unit of the device to ensure attribute traceability.
[0057] 3) Establish multi-dimensional association rules: Achieve accurate integration of multi-source data. Using the unique asset identifier ID as the core association index, within the defined site scope, multi-source data is integrated according to the progressive logic of "IP-MAC binding → IP association → multi-feature association", as follows: 3.1) IP-MAC Binding Association within the Same Site Pool: Within the same site, the unique correspondence between "device IP address - device MAC address" is prioritized to match the communication characteristics (source / target IP, port, communication frequency, etc.) in traffic monitoring data with the IP-MAC binding records in asset scanning data (situational awareness, nmap, vulnerability scanning tools); after a successful match, the data from both sides are associated with the same temporary index to form an "IP-MAC - temporary index" mapping, which solves the direct association between network layer and asset layer data.
[0058] 3.2) IP Association within the Same Site Pool: For data sources not associated via IP-MAC binding (such as SCD files, OMS ledgers), extract their device IP addresses and match them with the above "IP-MAC-Temporary Index"; if the IPs match, add the device model, name, rated parameters and other attributes to the corresponding temporary index; for data sources without IP addresses, proceed to the next step.
[0059] 3.3) Fuzzy matching association of names within the same site pool: For data sources without IP addresses, a fuzzy matching rule for device names containing core fields (such as "220kV Line B" and "Line Protection Device") is used to compare them with the device names already associated with the temporary index; if a match is successful, attribute information is added to the temporary index.
[0060] 3.4) Feature-based hierarchical matching within the same site pool: For devices that cannot be associated even with fuzzy name matching, extract the "basic features + scenario features" of the unassociated devices to construct matching dimensions, and perform matching according to priority: The first priority is matching the combination of "equipment model + rated parameters" (e.g., matching "RCS-931 + rated voltage 220kV" in the SCD file with the same combination of devices in the temporary index, utilizing the strong correspondence between power equipment models and parameters to achieve accurate matching); the second priority is matching "logical node + associated device" (for the SCD file, extract IEC 61850). The standard defines standardized logical nodes such as LLN0 and PTRC, as well as associated equipment identifiers such as merging units and measurement and control devices, which are compared with the equipment logical relationship database in the temporary index. The third priority is the matching of the combination of "manufacturer + commissioning time + equipment type", which is mainly adapted to old equipment data sources without key information such as model or logical nodes (such as early OMS ledgers). The core is to extract the three-dimensional combination of "manufacturer + commissioning time + equipment type" from the unassociated data source and compare it with the same dimension combination of the equipment already associated with the temporary index. The association is achieved by relying on the industry rule that the same type of equipment in the same power plant is often purchased in batches from the same manufacturer and put into centralized operation.
[0061] 3.5) Temporary unique identifier processing: For devices that are still not associated after the above matching (such as new models or devices with incomplete information), a temporary unique identifier containing site information is generated.
[0062] After all data sources are associated through the above rules, the temporary index and the successfully matched temporary identifier are uniformly mapped to a unique asset identifier ID, forming a "one ID, one file" structure, which is stored in Hive to provide a complete asset information foundation for subsequent knowledge graph construction.
[0063] Step S4: Construction of a knowledge graph for secondary power equipment.
[0064] The structured representation of knowledge based on a unified asset information database is divided into two parts: ontology layer design and instance layer import. Specifically, it includes: 1) Ontology Layer Design: Defines the entity types, relation types, and attribute constraints of the knowledge graph. Entity types include: 1.1) Equipment Nodes: These are further subdivided into sub-types such as line protection devices, busbar protection devices, and measurement and control devices. The core attributes are the unique identifier of the equipment, its model, manufacturer, and commissioning time.
[0065] 1.2) Related nodes: including vendor nodes (attributes: vendor name, region, contact information), IP nodes (attributes: IP address, subnet mask, gateway), port nodes (attributes: port number, service type, status), and vulnerability nodes (attributes: vulnerability number, risk level, scope of impact).
[0066] 1.3) The relationship types are defined as: “belongs to” (device-manufacturer), “deployed at” (device-plant), “use” (device-IP), “open” (device-port), “exists” (device-vulnerability), and “communicates” (device-device), and attributes are added to the relationship (such as the communication frequency and message length of the “communicates” relationship).
[0067] 2) Instance Layer Import: Convert the associated datasets of the unified asset information repository into a format supported by the graph database, store them in Arango, and import them in batches according to "entity table - relationship table": 2.1) Entity import: First import basic associated nodes such as vendor, IP, and port, then import device nodes, ensuring that the attributes of device nodes match those of associated nodes (e.g., the "vendor" attribute of the device node is consistent with the "vendor name" of the vendor node).
[0068] 2.2) Relationship Import: Generate relational data based on association rules, such as establishing a "Device-Belongs to-Manufacturer" relationship based on the attribute matching of device nodes and manufacturer nodes, and establishing a "Device-Communication-Device" relationship based on the communication records of traffic data and supplementing relevant attributes.
[0069] Step S5: Dissolving knowledge differences.
[0070] like Figure 2 As shown, a three-tiered mechanism of "conflict detection - disagreement resolution - expert repair" is used to resolve semantic conflicts in knowledge graphs, specifically including: 1) Conflict detection: Compare the attribute values of the same entity from multiple sources, identify conflicting fields, and classify them into two categories based on field type: numeric and text.
[0071] 2) Numerical conflict resolution (based on a large-scale knowledge model specific to the power industry) 2.1) The system extracts conflicting fields and their contextual information (data source name, collection time, associated entities, etc.) and constructs a unified ablation input.
[0072] 2.2) Input the conflict attributes and objective resources (traffic logs, nmap scan results, vulnerability scanning tool detection information, recently verified data, etc.) into the large model.
[0073] 2.3) Call the numerical ablation prompt word template, and let the large model determine which candidate value is most consistent with the objective facts and output the confidence distribution.
[0074] 2.4) Select the value with the highest confidence as the ablation result.
[0075] 3) Initial resolution of text-based conflicts (based on semantic understanding prompt templates) 3.1) When the conflict attribute is text type, input the candidate text and context features into the large model; 3.2) The system loads the standard rule base (model standard base, etc.); 3.3) Call the text ablation prompt word template to allow the large model to perform semantic comparison and synonym judgment; 3.4) If the model determines that the texts are semantically consistent or highly similar (similarity ≥ threshold), then they are uniformly normalized to the same attribute value.
[0076] 4) Text-related conflict retention and expert repair When the model determines that the text attributes differ significantly (low similarity or ambiguity), the system marks the conflict as "unresolvable," generates an explanatory report, and pushes it to the expert review module.
[0077] Experts manually repair the model based on the reasoning and context information provided by the model, and the repair results are written back to the knowledge base for subsequent model training and optimization.
[0078] 5) Results recording and traceability management The input prompts, model outputs, confidence distribution, objective resources used (traffic logs, nmap scan results, vulnerability scanning tool detection information, recently verified data, etc.) and rule base version for each ablation are recorded to form a traceable ablation log for subsequent verification and model tuning.
[0079] Numerical hint templates: You are a data ablation expert (large model agent) with knowledge in power grid and network security. Your task is to cross-validate conflicting numerical attributes (such as IP, MAC, port, version number, etc.) based on provided raw evidence (such as traffic logs, nmap scans, vulnerability / port detection tool results, and recently manually verified data), and output a credibility score and recommended decision. Please strictly follow the following input format to parse the evidence, perform format validation and standardization (e.g., IP and MAC format validation), calculate the credibility of each candidate value based on evidence frequency, time, source reliability, and consistency, and provide adoption / retention / human review recommendations based on specified thresholds. The output must include a detailed reasoning chain (evidence → conclusion) for auditing and expert review.
[0080] Field type: {IP / MAC} Candidate values: [Candidate value 1, Candidate value 2, Candidate value 3, ...] Context information: {Other attributes and attribute values of the knowledge graph} Original evidence: {{source: traffic data, original traffic data packets,} {source: nmap, nmap raw message}, {source: Asset verification, Asset attributes after verification}} Require: 1. Format Validation and Standardization: First, perform format validation on all candidate values (e.g., use the regular expression ^\d{1,3}(\.\d{1,3}){3}$ for IP addresses and ensure they are within the range of 0–255; normalize MAC addresses to lowercase and standardize delimiters, etc.). Provide a "format error" explanation for invalid values and exclude them from the confidence comparison.
[0081] 2. Evidence Matching Rules: Perform exact and approximate matching (e.g., port and service pairing, MAC / vendor OUI association). Count the number of matches for each candidate value in different sources and the most recent time.
[0082] 3. Timeliness considerations: Recent evidence has a greater weight than older evidence (e.g., evidence from the last 7 days is multiplied by a timeliness factor of 1.0, evidence from 7–30 days is multiplied by 0.8, and evidence from more than 30 days is multiplied by 0.6).
[0083] 4. Conflict resolution strategy / score summary: For each candidate value, calculate score = Σ(time_factor × match_strength) and perform Min-Max normalization to obtain confidence ∈ [0,1].
[0084] 5. Automatic adoption threshold (can be directly used for automated system decision-making): a) Confidence ≥ 0.8 → Automatic adoption (output action: AUTO_ADOPT) b) 0.5 ≤ confidence < 0.8 → Manual review recommended (output action: REVIEW) c) If confidence < 0.5, it is recommended to retain the conflict and mark it as non-resolvable (output action: HOLD_FOR_EXPERT). Text-based prompt templates: You are a power network semantic alignment and knowledge dissolution expert, responsible for identifying and unifying textual attribute conflicts (such as device model, operating system, manufacturer name, geographical location, etc.) in multi-source knowledge graphs.
[0085] Your task is to determine whether candidate texts are semantically equivalent based on candidate texts, contextual attribute information, and a standard rule base. If they are equivalent, they should be unified into a unified expression based on the standard rule base.
[0086] Please strictly adhere to the following principles: 1. Perform semantic comparison on each candidate text, taking into account professional terms, synonyms, abbreviations, version numbers, unit differences, symbol differences, and language differences (mixed Chinese and English writing) in the power industry. 2. If the text semantics are consistent, map it to terms from the standard rule base. 3. If the differences are significant or there is ambiguity (such as different models or different brands), it shall be judged as "non-synonymous" and the source of the difference shall be indicated; Field type: {} Candidate values: [Candidate value 1, Candidate value 2, Candidate value 3, ...] Context information: {Other attributes and attribute values of the knowledge graph} Standard rule lexicon: [List of candidate standard rule lexicons for attributes] Step S6: Optimized knowledge graph output and application.
[0087] After resolving knowledge disagreements, a standardized knowledge graph is output and integrated with intelligent application scenarios in the power system, specifically including: 1) Standardized output of graphs: Supports exporting in multiple formats (such as formats for knowledge sharing in the upper-level dispatch center and formats for cross-system data interaction). The files contain complete information on entity IDs, types, attributes, and relationships. Provides a graph visualization interface that supports filtering and viewing by "plant" and "equipment type" and displays entity relationships through node operations.
[0088] 2) Application Integration and Deployment: Deploy the knowledge graph to the power system's private cloud platform, providing query interfaces to support multiple applications. 2.1) Power Grid Asset Verification: The asset verification system queries the "equipment-attribute" relationship in the map through the interface, compares the ledger data with the map data, finds missing equipment or attribute errors, and generates supplementary / correction work orders.
[0089] 2.2) Security Risk Assessment: The security management system queries devices with high-risk vulnerabilities based on the "device-vulnerability" relationship, analyzes the vulnerability propagation path in conjunction with the "device-communication" relationship, and pushes priority suggestions for vulnerability remediation.
[0090] 2.3) Fault tracing: When a device alarms, the fault diagnosis system assists in locating the root cause of the fault by querying the relationship between "alarm device - associated device - historical fault" (such as a vulnerability in an associated device causing communication abnormalities).
[0091] 2.4) Abnormal Communication Detection: The network security monitoring system identifies abnormal communication patterns that deviate from the baseline by querying the dynamic relationship of "Device A - Communication Behavior - Device B" in the graph and combining it with real-time traffic aggregation data (such as communication frequency, message length, and time patterns). This includes high-frequency communication during non-working hours, sudden changes in target ports, and the addition of unregistered devices to the communication object. Security alarms are triggered and associated with the device roles and permission attributes in the graph to determine whether there is any risk of unauthorized access or lateral movement.
[0092] 3) Dynamic Knowledge Graph Maintenance: A daily incremental update mechanism is established to automatically synchronize the latest data from steps S1–S5, updating entity attributes and relationships in the knowledge graph in real time. Simultaneously, a closed-loop feedback mechanism based on expert daily operations and maintenance is introduced. Equipment change information (such as replacing device motherboards, modifying IP addresses, and adding communication links) entered by maintenance personnel during work order processing, on-site inspections, and fault handling is used as a trusted data source. After verification, the corresponding nodes and edges in the graph are automatically updated, and the data source is marked as "expert maintenance." The system regularly analyzes the types and frequency of attribute corrections made by experts, identifies weak links in data collection, and optimizes data source weights and model completion strategies accordingly. A full-scale evaluation is performed quarterly, statistically analyzing the entity completeness (e.g., equipment coverage) and attribute accuracy (e.g., the percentage of aligned attributes), outputting maintenance reports, and iterating the data governance process to ensure the knowledge graph remains accurate and reliable, serving the intelligent operation of the power system in the long term.
[0093] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0094] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0095] Please see Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 302 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called and executed by the processor 301 using the methods described above in the embodiments of this application. Input / output interface 303 is used to implement information input and output; The communication interface 304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 305 transmits information between various components of the device (e.g., processor 301, memory 302, input / output interface 303, and communication interface 304); The processor 301, memory 302, input / output interface 303 and communication interface 304 are connected to each other within the device via bus 305.
[0096] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0097] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0098] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0099] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0100] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented in the embodiments of this program product are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments. The executable computer program code or "code" used to perform the various embodiments can be written in high-level programming languages such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0101] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0102] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0103] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0104] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0105] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0106] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0107] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0108] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0109] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0110] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for constructing a knowledge graph for secondary power equipment, characterized in that, The method includes the following steps: Collect multi-source data related to secondary power equipment; The multi-source data is preprocessed and aggregated to obtain standardized data; Based on the standardized data, a unified asset information database is constructed through progressive multi-dimensional association rules. The progressive multi-dimensional association rules include IP-MAC binding association, IP association, device name fuzzy matching and feature hierarchical matching executed sequentially within the same site. This associates multi-source data with a unique asset identifier ID, forming a standardized asset file with "one ID and one file". Based on the unified asset information database, a knowledge graph with secondary power equipment as the core entity is constructed, defining entity types, relationship types, and attribute sets; Knowledge divergence ablation is performed on the multi-source data conflicts in the knowledge graph. The divergence ablation process at least partially involves identifying and calibrating conflict attributes through a large model and recording ablation logs. Output the knowledge graph optimized by divergence resolution.
2. The method according to claim 1, characterized in that, The collection of multi-source data related to secondary power equipment specifically includes: Collect traffic monitoring data, which includes source IP address, source port, destination IP address, destination port, and protocol type; Collect asset scanning data, which includes asset information obtained by the situational awareness system, component identification results generated by vulnerability scanning tools and nmap tools, port scanning data, vulnerability scanning results, substation configuration description file data, dispatch management system equipment data, and asset verification data.
3. The method according to claim 1, characterized in that, The preprocessing and aggregation of the multi-source data specifically includes: The collected multi-source data is cleaned, formatted, and deduplicated. For traffic monitoring data, data aggregation is performed based on the five-tuple to identify and eliminate duplicate communication records and extract communication behavior characteristics between devices; The asset scanning data is normalized, and the field naming conventions and data representation formats are standardized to establish a standardized data model covering device IP address, manufacturer, model, service type, open port, and known vulnerability attributes.
4. The method according to claim 1, characterized in that, The construction of a unified asset information database further includes: By accurately associating with site numbers and / or fuzzy matching with site names, the site range to which the device belongs can be clearly defined; Generate a globally unique asset identifier ID containing site information based on the physical and logical attributes of the device, and record the source data and collection time of each attribute; Within the scope of the site, multi-source data is integrated and mapped to the unique asset identifier ID according to the progressive multi-dimensional association rules.
5. The method according to claim 1 or 4, characterized in that, The feature hierarchical matching includes: First priority matching: Matching is based on the combination of equipment model and rated parameters; Second priority matching: Matching is based on the combination of logical nodes and associated devices; Third priority matching: Matching is based on a combination of manufacturer, commissioning time, and equipment type.
6. The method according to claim 1, characterized in that, The knowledge disagreement resolution process includes: Detect conflict attributes and distinguish between numerical conflicts and textual conflicts; For numerical conflicts, the conflict attributes and their context information, as well as objective resource data, are input into the large model. The large model evaluates the credibility of each candidate value according to the preset confidence calculation rules and outputs the ablation decision based on the credibility threshold. For text-based conflicts, the conflict attributes, their contextual information, and the standard rule base are input into the large model, which then performs semantic equivalence judgment and normalizes the candidate values that are semantically consistent. For conflicts that cannot be resolved by the large model, mark them as unresolvable conflicts and push them to the expert review module. Receive the correction results after expert review and write them back to the knowledge base. Record the input information, model output, objective resources and rule base version for each ablation, forming a traceable ablation log.
7. The method according to claim 6, characterized in that, The confidence level calculation rules include: Perform format validation and standardization on each candidate value; Statistically analyze the number of times each candidate value matches across different objective resources and the time of the most recent match; Time weights are assigned based on the timeliness of the evidence; Calculate the weighted score for each candidate value and normalize it to obtain the confidence level; The confidence level is compared with a preset threshold. If the confidence level is greater than or equal to the first threshold, it is automatically adopted. If the confidence level is between the first threshold and the second threshold, manual review is recommended. If the confidence level is less than the second threshold, it is recommended to retain the conflict and mark it as non-resolvable.
8. The method according to claim 1, characterized in that, Following the optimized knowledge graph output, the following is also included: The knowledge graph is deployed to the power system platform, providing a query interface to support applications such as power grid asset verification, security risk assessment, fault tracing and / or abnormal communication detection. Establish a dynamic maintenance mechanism for the knowledge graph, including daily incremental updates and a closed loop of expert operation and maintenance feedback. Optimize the knowledge graph based on data corrected by experts and regularly evaluate the quality of the graph.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.