Power knowledge processing method and device based on multi-source data fusion, medium and equipment

By acquiring graph structure data from various data sources in the power system, performing functional mapping and similarity algorithm verification, the problem of inconsistent device identifiers in different systems was solved. A multi-data source fusion power knowledge graph was constructed, achieving accuracy in device matching and data consistency, and supporting intelligent analysis of the power system.

CN121095008BActive Publication Date: 2026-05-12GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG POWER GRID CO LTD
Filing Date
2025-08-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The lack of a unified equipment identification system among various business systems in the power system has resulted in the same physical equipment being assigned different identifiers in different systems, forming information silos, hindering the integration and unified management of multi-source data, and making it difficult to support equipment status assessment and fault diagnosis.

Method used

通过获取各预设数据源的图结构数据,进行功能映射和节点标识符统一,利用预设图相似度算法结合网络结构和节点属性相似度确定目标设备节点,并进行多维度校验,最终构建多数据源融合电力知识图。

Benefits of technology

It achieves the retention of a unique node for the same physical device in the fusion graph, eliminates data redundancy, provides a foundation for intelligent power system analysis with data consistency and topology integrity, and reduces the misjudgment rate of device matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095008B_ABST
    Figure CN121095008B_ABST
Patent Text Reader

Abstract

The application provides a multi-source data fusion power knowledge processing method and device, medium and equipment, relates to the field of power data processing, and comprises the following steps: acquiring graph structure data corresponding to each preset data source of a target power system; performing function mapping on each node in each graph structure data to obtain standard function graph structure data corresponding to each preset data source; determining each target device node in different preset data sources according to each standard function graph structure data and a preset graph similarity algorithm; and fusing the graph structure data corresponding to each preset data source with each target device node as a fusion node to obtain a multi-source data fusion power knowledge graph. The multi-source data fusion power knowledge graph after fusion has the characteristics of data consistency, semantic unification and topological completeness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power data processing, and in particular to power knowledge processing methods, devices, media, and electronic equipment involving multi-source data fusion. Background Technology

[0002] In the daily operation and management of power systems, multiple independent professional management systems are typically deployed to achieve refined control over different business processes. For example, the Enterprise Asset Management System (EAM) focuses on the full lifecycle management of equipment, covering information such as equipment procurement, ledgers, and maintenance plans; the Supervisory Control and Data Acquisition System (SCADA) focuses on real-time monitoring of the operating status of power equipment, recording dynamic operating data such as voltage, current, and power; the maintenance work order system focuses on the dispatch, execution, and closed-loop management of equipment maintenance tasks, storing process information such as work order number, maintenance content, and processing results; in addition, there are Production Management Systems (PMS) and Enterprise Resource Planning Systems (ERP), which record corresponding data for specific business scenarios such as production scheduling and resource allocation.

[0003] These systems each focus on different aspects of power system operation, and the types of information they record vary significantly: some systems primarily use static attribute data (such as equipment model and manufacturer in EAM), some focus on dynamic operational data (such as real-time monitoring values ​​in SCADA), and others emphasize process data (such as task progress in maintenance work order systems). For critical power equipment such as transformers, while the maintenance-related knowledge in each system includes some basic common content (such as the physical location and core functions of the equipment), it is more reflected in the different data content generated by the differences in business scenarios, forming a multi-dimensional information system around the same equipment.

[0004] However, because these business systems often follow their own technical standards and data specifications during construction, they lack a unified equipment identification system. This results in the same physical equipment being assigned different identifiers or naming methods in different systems, a phenomenon known as "same entity, different names." For example, a transformer might be numbered "T1023-2020" in the EAM system, marked as "#B12 transformer" in the SCADA system, and recorded as "Dongcheng District Main Transformer A" in the maintenance work order system. This inconsistency in identification creates "information silos" for equipment data in different systems, severely hindering the integration and unified management of multi-source data on the same equipment. It makes it impossible to efficiently integrate information such as static attributes, dynamic operation, and maintenance records of the same equipment from different systems, and it also makes it difficult to support advanced applications such as equipment status assessment and fault diagnosis based on complete data. This restricts the in-depth mining of the value of power system data and the improvement of business collaboration efficiency. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a method, apparatus, medium, and electronic device for processing power knowledge through multi-source data fusion, which at least partially solves the problems existing in the prior art.

[0006] In a first aspect of this application, a method for processing power knowledge through multi-source data fusion is provided, the method comprising the following steps:

[0007] S100, acquire graph structure data corresponding to each preset data source of the target power system; wherein, each graph structure data contains at least one target device node; the node identifier of the target device node contained in each graph structure data is different.

[0008] S200: Perform function mapping on each node in each graph structure data to obtain standard function graph structure data corresponding to each preset data source; where each node in the standard function graph structure data is a function node; a function node is a node named according to its function.

[0009] S300, based on each standard functional diagram structure data and a preset graph similarity algorithm, determines each target device node in different preset data sources; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional diagram structure data.

[0010] S400 uses each target device node as a fusion node to fuse the graph structure data corresponding to each preset data source, so as to obtain a multi-data source fused power knowledge graph.

[0011] In a second aspect of this application, a power knowledge processing apparatus for multi-source data fusion is provided, the apparatus comprising:

[0012] The acquisition unit is used to acquire graph structure data corresponding to each preset data source of the target power system; wherein each graph structure data contains at least one target device node; and the node identifier of the target device node contained in each graph structure data is different.

[0013] The mapping unit is used to perform functional mapping on each node in each graph structure data to obtain the standard functional graph structure data corresponding to each preset data source; wherein, each node in the standard functional graph structure data is a functional node; a functional node is a node named according to its function.

[0014] The determination unit is used to determine each target device node in different preset data sources based on each standard functional graph structure data and a preset graph similarity algorithm; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional graph structure data.

[0015] The fusion unit is used to fuse the graph structure data corresponding to each preset data source with each target device node as the fusion node, so as to obtain a multi-data source fused power knowledge graph.

[0016] In a third aspect of this application, a non-transitory computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored in the storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to realize the aforementioned power knowledge processing method of multi-source data fusion.

[0017] In a fourth aspect of this application, an electronic device is provided, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0018] This application has at least the following beneficial effects:

[0019] The multi-source data fusion power knowledge processing method provided in this application obtains graph structure data corresponding to each preset data source. Then, it performs functional mapping on each node to obtain standard functional graph structure data, unifying all nodes into nodes named by function. This functional naming eliminates semantic differences between heterogeneous data sources, providing comparable functional benchmarks for device nodes from different data sources. This resolves device matching barriers caused by different description dimensions and provides a unified semantic carrier for subsequent graph similarity calculations. Next, based on a preset graph similarity algorithm, it simultaneously considers network structure similarity and node attribute similarity to determine target device nodes, achieving multi-dimensional verification of device matching and significantly reducing the misjudgment rate of single-attribute matching. Finally, it fuses the graph structure data with each target device node as a fusion node, ensuring that only one node for the same physical device is retained in the fused graph. Simultaneously, it inherits the topological relationships of each data source, eliminating data redundancy from the bottom layer and constructing a complete cross-data source topology network. This results in a multi-data source fused power knowledge graph with consistent data, unified semantics, and complete topology, providing high-quality underlying data support for intelligent power system analysis. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart of a power knowledge processing method based on multi-source data fusion provided in this application embodiment;

[0022] Figure 2 This is a structural block diagram of a power knowledge processing device for multi-source data fusion provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0025] It should be noted that the following description covers various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0026] Please refer to Figure 1 As shown, embodiments of this application provide a method for processing power knowledge through multi-source data fusion, the method comprising the following steps:

[0027] S100, acquire graph structure data corresponding to each preset data source of the target power system; wherein, each graph structure data contains at least one target device node; the node identifier of the target device node contained in each graph structure data is different.

[0028] Specifically, the preset data sources can be different management systems within the target power system. For example, these could include an Enterprise Asset Management (EAM) system, a Supervisory Control and Data Acquisition (SCADA) system, and a maintenance work order system. The power knowledge and data corresponding to each preset data source are defined as graph-structured data, which includes nodes, edges, and node attributes. For example, a target equipment node could be a transformer. In EAM, a transformer node identifier is "#1 Main Transformer." In the SCADA system, this transformer node identifier is "T1001." In the maintenance work order system, this transformer node identifier is "Main Transformer A," and the transformer's manufacturer's nameplate number is "TR-2023-001." For example, the graph-structured data could also be a knowledge graph.

[0029] S200: Perform function mapping on each node in each graph structure data to obtain standard function graph structure data corresponding to each preset data source; where each node in the standard function graph structure data is a function node; a function node is a node named according to its function.

[0030] Specifically, due to the autonomous management nature of each management system, in addition to the different node identifiers of the target device node, the node identifiers of other nodes connected to the target device node may also differ. For example, a substation (transformers are typically located within a substation and are part of the power transmission network) would have the node identifier "Substation A" in the EAM system.

[0031] In the SCADA system, the node identifier for this substation is "SubstationAlpha". Circuit breaker (used to protect the transformer from overload or short circuits, typically directly connected to the high-voltage or low-voltage side of the transformer). In the EAM system, the node identifier for this circuit breaker is "High-voltage circuit breaker CB-001"; in the SCADA system, the node identifier for this circuit breaker is "Breaker_HV_01". Maintenance plan / record (contains information about the transformer's periodic maintenance, including schedules, historical records, etc.). Example: In the EAM system, the node identifier for this maintenance plan / record is "Annual Maintenance Plan P1"; in the maintenance work order system, the node identifier for this maintenance plan / record is "Work Order W1". Location information (the specific geographical or logical location of the transformer). In the EAM system, the node identifier for this location information is "Location_ID:12345"; in the SCADA system, the node identifier for this location information is "GeoTag:XYZ"; in the maintenance work order system, the node identifier for this location information is "Substation A-25-X11".

[0032] Here, nodes in different systems are mapped and merged according to their functional roles. For example, TRX-001 in the EAM system and TransformerA in the SCADA system are both mapped as "core power transmission components". Monitoring devices such as TempSensor_003 in EAM and TS-01 in SCADA are both mapped as "monitoring devices".

[0033] S300, based on each standard functional diagram structure data and a preset graph similarity algorithm, determines each target device node in different preset data sources; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional diagram structure data.

[0034] Specifically, based on a preset graph similarity algorithm, the target device node is determined by considering both network structure similarity and node attribute similarity, thereby achieving multi-dimensional verification of device matching and significantly reducing the misjudgment rate of single attribute matching.

[0035] It should be noted that: the preset device node is a node with the same function as the target device node; and the target device node can be any one of the preset device nodes. That is, the target device node may be transformer A, which exists in every management system, while the preset device node may be transformer F, which may only exist in a certain management system.

[0036] In one exemplary embodiment of this application, step S300 includes:

[0037] S310, Obtain the initial similarity matrix; where the initial similarity matrix is ​​n×n dimensional; n is the total number of nodes contained in all graph structure data; the similarity of key elements in the initial similarity matrix is ​​1, and the similarity of non-key elements is 0; the key element is the similarity between nodes in the i-th row and i-th column; the non-key element is the similarity between nodes in the i-th row and j-th column; i=1, 2, ..., n; j=1, 2, ..., n; i≠j.

[0038] Here, the initial similarity matrix is ​​an n×n dimensional matrix, where n is the total number of nodes contained in all graph structure data corresponding to all preset data sources; the order of rows and columns is consistent. The key element is the similarity between nodes in the i-th row and i-th column. For example, the key element is the similarity between nodes in the i-th row (transformer A) and i-th column (transformer A), which is a self-matching of nodes, that is, the similarity between transformer A and transformer A is 1.

[0039] The initial similarity matrix setup in this embodiment essentially starts from the underlying logic of data association, establishing a baseline framework for subsequent node matching and similarity calculation. The similarity of key elements is set to 1, meaning the similarity between a node and itself is 1. This is based on the logic of node self-matching, where any node must be completely identical to itself in terms of attributes, functions, and network location, thus having a maximum similarity of 1. This setting provides a benchmark reference value for subsequent similarity calculations between other nodes. For example, when comparing transformer A with transformer A', their similarity is naturally 1, an absolute match relationship that can be determined without additional calculation.

[0040] The initial similarity of non-critical elements (i, j, i ≠ j) is set to 0, based on the principle of defaulting to no match when there is no prior information. Before any comparison calculation is performed, it is impossible to determine whether two different nodes are related or similar. Therefore, their similarity is temporarily set to the minimum value of 0. Subsequently, the values ​​of these elements can be updated and corrected through a preset similarity algorithm (such as combining network structure features) to gradually obtain the true similarity. This initial setting simplifies the initial matrix construction process and reserves adjustment space for subsequent dynamic calculations.

[0041] S320, determine the current similarity value of each non-key element according to the initial similarity matrix and the preset calculation rules to obtain the intermediate similarity matrix; wherein, the preset calculation rules determine the current similarity value according to the similarity between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column corresponding to the non-key element.

[0042] Here, step S320 includes:

[0043] S321, obtain the sum of the similarities between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column corresponding to each non-key element; where the similarity between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column is the similarity value at the corresponding position in the initial similarity matrix.

[0044] S322. Based on the sum of the similarities between each node adjacent to the i-th row node and each node adjacent to the j-th column node corresponding to each non-key element, the preset decay factor, and the product of the number of nodes adjacent to the i-th row node and the number of nodes adjacent to the j-th column node of the non-key element, the current similarity value of each non-key element is obtained to obtain the intermediate similarity matrix.

[0045] The current similarity value of each non-key element is proportional to the sum of the similarities between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column of the non-key element; proportional to the preset decay factor; and inversely proportional to the product of the number of nodes adjacent to the node in the i-th row and the number of nodes adjacent to the node in the j-th column of the non-key element.

[0046] Here, the current similarity value is determined based on the similarity between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column corresponding to the non-key element. Since the influence of neighboring nodes' similarity on the current node is not completely propagated but weakens with increasing association level, a preset attenuation factor is set to avoid similarity distortion due to over-propagation. The preset attenuation factor is typically set to 0.8 or 0.9.

[0047] For a non-key element (i,j), its corresponding current similarity value S i,j The following conditions must be met:

[0048] ;

[0049] Where C is the preset attenuation factor; L i L represents the number of neighboring nodes of the i-th node; j Let a be the number of neighboring nodes of the j-th node; a = 1, 2, ..., L i c = 1, 2, ..., L j S i-a,j-c Let be the similarity between the a-th neighbor of the i-th node and the c-th neighbor of the j-th node in the initial similarity matrix.

[0050] In this embodiment, adjacent nodes are those directly adjacent to the target node, excluding nodes with multiple layers of association. Here, since different neighbors have slightly different influences on the target node, the preset attenuation factor can flexibly adjust the "weight" of the similarity contributions of these neighbors with different degrees of closeness in the calculation, making the similarity result more consistent with the real physical association logic.

[0051] S330, obtain the average similarity difference of the intermediate similarity matrix.

[0052] Here, if the intermediate similarity matrix is ​​from the first iteration, meaning the average similarity difference is obtained from the initial similarity matrix, then the average similarity difference is the average of the differences between the value of each element in the intermediate similarity matrix and the value of the corresponding element in the initial similarity matrix. If the intermediate similarity matrix is ​​not from the first iteration, meaning the average similarity difference is obtained from the previous intermediate similarity matrix, then the average similarity difference is the average of the differences between the value of each element in the currently updated intermediate similarity matrix and the value of the corresponding element in the intermediate similarity matrix before the previous update.

[0053] S340, if the average similarity difference is equal to or greater than the preset difference threshold, then determine the current similarity value of each non-key element according to the intermediate similarity matrix and the preset calculation rules to update the intermediate similarity matrix; and jump to step S330 until the average similarity difference is less than the preset difference threshold, then determine the current intermediate similarity matrix as the target similarity matrix.

[0054] If the average similarity difference is equal to or greater than the preset difference threshold, it means that further iteration is needed. In this case, continue iterating until the average similarity difference is less than the preset difference threshold.

[0055] In this embodiment, similarity is calculated iteratively to gradually approximate the true and stable topological association similarity between power system nodes, solving the problem that a single calculation cannot accurately characterize complex associations. The iteration, through a cycle of calculation, difference judgment, and recalculation, allows the similarity value to gradually converge towards a stable state of topological association. In each iteration, the similarity of neighbors' neighbors and multi-layered transmission is activated and propagated to the current node.

[0056] S350: Determine the target device node based on the similarity between the target similarity matrix and the node attributes.

[0057] Specifically, step S350 includes:

[0058] S351, obtain the similarity between any two preset device nodes in the target similarity matrix; wherein, the preset device node is a node with the same function as the target device node; and the target device node is any one of the preset device nodes.

[0059] Here, the nodes in each of the above graph structure data are of multiple types, including preset device nodes and other types of nodes. Although we have placed all nodes in the matrix, this is because for a target device node, it is more likely to be associated with non-target device nodes or non-preset device nodes. Therefore, placing all nodes in the matrix is ​​to fully obtain the graph structure characteristics (associations with other nodes) of each preset device node (the target device node can be any of the preset device nodes).

[0060] After obtaining the target similarity matrix, the similarity between any two preset device nodes in the target similarity matrix is ​​obtained. For example, the preset device nodes can be transformer A, transformer B, etc. The preset device nodes should be the main power equipment in the power system, which exist in different preset data sources.

[0061] S352, if the similarity between any two preset device nodes is greater than the first preset similarity threshold, then the corresponding two nodes are determined to be a key device node pair.

[0062] Specifically, if the similarity between any two preset device nodes is greater than the first preset similarity threshold, it means that the two nodes may be the same device distributed in different systems.

[0063] S353, obtain the similarity between node attributes of each key device node pair.

[0064] Furthermore, in order to further obtain the similarity between key device node pairs, this embodiment, after determining the similarity of graph structure features, also obtains the similarity between node attributes to further determine whether the associated device node pairs are the same devices distributed in different systems.

[0065] S354. Based on the similarity of each key equipment node pair in the target similarity matrix and the similarity between node attributes, the key similarity between each key equipment node pair is obtained; wherein, the key similarity is proportional to the similarity of the key equipment node pair in the target similarity matrix and the key similarity is proportional to the similarity between node attributes.

[0066] Specifically, in one embodiment, P u,v =(1+W u,v )×S u,v Among them, P u,v W represents the key similarity between node u and node v in a key device node pair. u,v The similarity between the node attributes of node u and node v contained in a critical equipment node pair; 0 ≤ W u,v ≤1; S u,vThis represents the similarity between node u and node v in the target similarity matrix for the key equipment node pair.

[0067] Here, the original structural similarity is adjusted by attribute similarity to make the final similarity more accurate and avoid inaccurate results caused by a single factor.

[0068] S355, if the key similarity between any key node pair is greater than the preset key similarity threshold, then the two preset device nodes contained in the key node pair are determined to be the same target device node.

[0069] Specifically, if the key similarity between any pair of key nodes is greater than a preset key similarity threshold, then the two preset device nodes contained in that key node pair are determined to be the same target device node. As an example: if P... u,v If the similarity exceeds the preset threshold, it means that node u and node v are the same target device node distributed in different management systems (e.g., both are transformer F).

[0070] In one exemplary embodiment of this application, after step S351, the method further includes:

[0071] S356, If the similarity between any two preset device nodes is less than the first preset similarity threshold and the second preset similarity threshold, then the corresponding two nodes are determined to be an intermediate device node pair.

[0072] S357, if the absolute value of the difference in the number of adjacent nodes between the two preset device nodes contained in the intermediate device node pair is greater than the preset threshold for the difference in the number of adjacent nodes, then obtain the similarity between the node attributes of the intermediate device node pair.

[0073] S358, based on the similarity of each intermediate device node pair in the target similarity matrix and the similarity between node attributes, the intermediate similarity between each intermediate device node pair is obtained; wherein, the intermediate similarity is obtained by weighted summation of the similarity of the intermediate device nodes in the target similarity matrix and the similarity between the node attributes of the intermediate device nodes, and the weight of the similarity between the node attributes of the intermediate device nodes is greater than the weight of the similarity of the intermediate device nodes in the target similarity matrix.

[0074] In this embodiment, if the similarity between any two preset device nodes is less than both the first and second preset similarity thresholds (i.e., the similarity between any two preset device nodes is not very high), the number of adjacent nodes of the two preset device nodes included in the intermediate device node pair is first obtained. Because there exists a special case where transformer A has many adjacent nodes in management system A, including common nodes α and β, while transformer A has very few adjacent nodes in management system B, also including common nodes α and β, the similarity between two preset device nodes is inversely proportional to the product of the number of adjacent nodes of transformer A in management system A and the number of adjacent nodes of transformer A in management system B. This leads to a lower similarity of the obtained structures. Therefore, to avoid missing any nodes, the final similarity is further determined based on the similarity between node attributes.

[0075] In this embodiment, the method for determining the final similarity, compared to the previous embodiment, increases the influence of the similarity between node attributes on the final result (expanding the weight), which is more reasonable for the special case of this embodiment.

[0076] S359, if the intermediate similarity between any pair of intermediate nodes is greater than the preset intermediate similarity threshold, then the two preset device nodes contained in the intermediate node pair are determined to be the same target device node.

[0077] In one exemplary embodiment of this application, the similarity between the node attributes is determined according to the following steps:

[0078] S301, determine the first node attribute similarity between each key device node pair or each intermediate device node pair according to the first node attribute similarity algorithm; wherein, the first node attribute similarity algorithm is determined according to the number of operations required for string conversion between the attribute values ​​of preset attributes of two key device nodes or two intermediate device nodes.

[0079] Specifically, the first-node attribute similarity algorithm can be the Levenshtein distance algorithm. This algorithm only focuses on the literal meaning, not the semantics, and measures the minimum number of single-character editing operations (insertion, deletion, or replacement) required to transform one string into another. It is suitable for evaluating the similarity between two strings, especially for handling spelling errors or minor formatting changes. It is suitable for string-type attributes, such as names, model numbers, etc. It measures the minimum number of editing operations required to transform one string into another.

[0080] As an example: There is a transformer named Transformer#1, with the following records in the EAM and SCADA systems: EAM system: Model: TRX-2023-001; Manufacturer: Siemens; Installation location: Substation A; Commissioning date: 2015-06-01; SCADA system: Model: TRX-2023-01; Manufacturer: Siemens; Installation location: Substation A; Commissioning date: 2015-06-01; Following the steps above, the similarity score of the attribute value of each preset attribute can be calculated separately, and then the results can be combined (which can be directly summed) to obtain the final similarity score, and to determine whether the two records refer to the same transformer.

[0081] S302, if the first node attribute similarity is less than the preset node attribute similarity threshold, then the second node attribute similarity between each key device node pair or each intermediate device node pair is determined according to the second node attribute similarity algorithm; wherein, the second node attribute similarity algorithm is determined based on the semantic similarity between two key device nodes or between two intermediate device nodes and the attribute values ​​of the preset attributes.

[0082] Specifically, the Levenshtein distance primarily assesses the literal similarity of two strings. In power systems and other industrial sectors, inconsistent terminology frequently arises due to varying data entry habits among different personnel and systems. Examples include abbreviations vs. full names: "transformer" vs. "TRX", "high voltage" vs. "HV", and high-voltage switch vs. high-voltage circuit breaker vs. high-opening vs. high-closing. Therefore, it's crucial to consider semantics rather than just the literal meaning. Thus, for strings with low Levenshtein distance similarity, a secondary semantic similarity assessment is performed.

[0083] In this embodiment, if the similarity of the first node attribute is less than a preset node attribute similarity threshold, it indicates that the literal similarity between the two is low. In this case, the similarity of the second node attribute between each key device node pair or each intermediate device node pair is determined according to the second node attribute similarity algorithm. For example, the second node attribute similarity algorithm can be the Sentence-BERT algorithm. Sentence-BERT supports encoding short text (attribute values ​​and / or attribute names) to obtain the semantic vector of each segment attribute value, and then calculates the similarity.

[0084] Sentence vector models (such as Sentence-BERT) can be used to calculate the similarity between different attribute values, even if these attribute values ​​differ in language, expression, or structure. This method is particularly suitable for handling situations in power systems where "semantically the same but literally different" results are caused by issues such as non-standard naming, mixed use of Chinese and English, and mixing of abbreviations / full names.

[0085] Traditional string matching (such as Levenshtein) only considers whether the characters are similar, while sentence vector models such as Sentence-BERT focus on whether the characters are semantically similar.

[0086] First, treat the attribute values ​​as either "sentences" or "phrases".

[0087] Even a short attribute value (such as oil leak) can be considered a "semantic fragment". Sentence-BERT supports encoding short texts.

[0088] In one embodiment, the Sentence-BERT algorithm can use concatenated attribute names and values ​​(enhancing semantics) as input, allowing the model to better understand "what type of attribute this is." As an example, Sentence-BERT also works for short texts (such as "oil leak"), but the richer the information, the better: it is recommended to concatenate the context (such as "fault: oil leak").

[0089] S303, the similarity of the second node attribute is determined as the similarity between node attributes.

[0090] In one exemplary embodiment of this application, after step S301, the method further includes:

[0091] S304, if the similarity of the first node attribute is equal to or greater than the preset node attribute similarity threshold, then the similarity of the first node attribute is determined as the similarity between node attributes.

[0092] Here, if the similarity of the first node attribute is equal to or greater than the preset node attribute similarity threshold, it means that the two may be different due to missing letters or different date formats, and the literal similarity is very high. In this case, it is directly determined as the similarity between node attributes.

[0093] This embodiment first uses a literal algorithm with lower computational power to obtain the similarity between node attributes. When the similarity obtained by the literal algorithm is low, in order to further improve the accuracy of the similarity between node attributes, a semantic similarity algorithm with better computational power but higher accuracy is used to determine the accuracy of the similarity between node attributes, thus maintaining a good balance between saving computational power and ensuring accuracy.

[0094] S400 uses each target device node as a fusion node to fuse the graph structure data corresponding to each preset data source, so as to obtain a multi-data source fused power knowledge graph.

[0095] Specifically, firstly, a unified fusion node (such as "10kV Transformer-102") is created for each group of target device nodes as the unique identifier of that entity in the fusion knowledge graph.

[0096] Next, the associations related to the target device node in each data source are integrated: for the fusion node, all its neighbor nodes and connection relationships in each data source are retained (such as the connection between "T102" and "Line L1" in data source 1, and the connection between "Transformer-102" and "Sensor S5" in data source 2, which are all integrated into the connection between the fusion node and "Line L1" and "Sensor S5").

[0097] Finally, redundant information is removed (if the same connection is recorded repeatedly in multiple data sources, it is only kept once) to form a multi-data source fused power knowledge graph that contains valid information from all data sources, with no duplicate nodes and complete associations.

[0098] Please refer to Figure 2 As shown, an embodiment of this application provides a power knowledge processing device 100 for multi-source data fusion, the device comprising:

[0099] The acquisition unit 110 is used to acquire graph structure data corresponding to each preset data source of the target power system; wherein each graph structure data contains at least one target device node; and the node identifier of the target device node contained in each graph structure data is different.

[0100] The mapping unit 120 is used to perform functional mapping on each node in each graph structure data to obtain standard functional graph structure data corresponding to each preset data source; wherein, each node in the standard functional graph structure data is a functional node; a functional node is a node named according to its function.

[0101] The determining unit 130 is used to determine each target device node in different preset data sources based on each standard functional diagram structure data and a preset graph similarity algorithm; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional diagram structure data.

[0102] The fusion unit 140 is used to fuse the graph structure data corresponding to each preset data source with each target device node as the fusion node, so as to obtain a multi-data source fused power knowledge graph.

[0103] Embodiments of this application also provide a computer program product including program code that, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above according to various exemplary embodiments of this application.

[0104] Furthermore, although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0105] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0106] In an exemplary embodiment of this application, an electronic device capable of implementing the above-described method is also provided.

[0107] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0108] An electronic device according to this embodiment of the present application. The electronic device is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0109] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and buses connecting different system components (including memory and processor).

[0110] The memory stores program code that can be executed by a processor, causing the processor to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of this application.

[0111] The storage may include readable media in the form of volatile storage, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0112] The storage may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more applications, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0113] A bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus architectures.

[0114] The electronic device can also communicate with one or more external devices (such as keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (such as routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. As shown in the figure, the network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0115] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this application.

[0116] In exemplary embodiments of this application, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this application may also be implemented as a program product including program code, which, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this application described in the "Exemplary Methods" section above.

[0117] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0118] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0119] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0120] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0121] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this application, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0122] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0123] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for processing power knowledge through multi-source data fusion, characterized in that, The method includes: S100, Obtain graph structure data corresponding to each preset data source of the target power system; wherein, each graph structure data contains at least one target device node; the node identifier of the target device node contained in each graph structure data is different; S200: Perform function mapping on each node in each graph structure data to obtain standard function graph structure data corresponding to each preset data source; where each node in the standard function graph structure data is a function node; a function node is a node named according to its function. S300, based on each standard functional diagram structure data and a preset graph similarity algorithm, determines each target device node in different preset data sources; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional diagram structure data. S400 uses each target device node as a fusion node to fuse the graph structure data corresponding to each preset data source to obtain a multi-data source fused power knowledge graph. Step S300 includes: S310, Obtain the initial similarity matrix; where the initial similarity matrix is ​​n×n dimensional; n is the total number of nodes contained in all graph structure data; the similarity of key elements in the initial similarity matrix is ​​1, and the similarity of non-key elements is 0; the key element is the similarity between nodes in the i-th row and i-th column; the non-key element is the similarity between nodes in the i-th row and j-th column; i=1, 2, ..., n; j=1, 2, ..., n; i≠j; S320, determine the current similarity value of each non-key element based on the initial similarity matrix and preset calculation rules to obtain the intermediate similarity matrix; S330, obtain the average similarity difference of the intermediate similarity matrix; S340, if the average similarity difference is equal to or greater than the preset difference threshold, then determine the current similarity value of each non-key element according to the intermediate similarity matrix and the preset calculation rules to update the intermediate similarity matrix; and jump to step S330 until the average similarity difference is less than the preset difference threshold, then determine the current intermediate similarity matrix as the target similarity matrix. S350: Determine the target device node based on the similarity between the target similarity matrix and the node attributes.

2. The power knowledge processing method based on multi-source data fusion according to claim 1, characterized in that, Step S320 includes: S321, obtain the sum of similarities between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column corresponding to each non-key element; where the similarity between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column is the similarity value of the corresponding position in the initial similarity matrix. S322. Based on the sum of the similarities between each node adjacent to the i-th row node and each node adjacent to the j-th column node corresponding to each non-key element, the preset decay factor, and the product of the number of nodes adjacent to the i-th row node and the number of nodes adjacent to the j-th column node of the non-key element, the current similarity value of each non-key element is obtained to obtain the intermediate similarity matrix.

3. The power knowledge processing method based on multi-source data fusion according to claim 2, characterized in that, The current similarity value of each non-key element is directly proportional to the sum of the similarities between each node adjacent to the node in the i-th row and each node adjacent to the node in the j-th column corresponding to the non-key element; directly proportional to the preset decay factor; and inversely proportional to the product of the number of nodes adjacent to the node in the i-th row and the number of nodes adjacent to the node in the j-th column of the non-key element.

4. The power knowledge processing method based on multi-source data fusion according to claim 2, characterized in that, Step S350 includes: S351, obtain the similarity between any two preset device nodes in the target similarity matrix; wherein, the preset device node is a node with the same function as the target device node; and the target device node is any one of the preset device nodes; S352, if the similarity between any two preset device nodes is greater than the first preset similarity threshold, then the corresponding two nodes are determined to be a key device node pair; S353, obtain the similarity between node attributes of each key device node pair; S354, Based on the similarity of each key equipment node pair in the target similarity matrix and the similarity between node attributes, obtain the key similarity between each key equipment node pair; wherein, the key similarity is proportional to the similarity of the key equipment node pair in the target similarity matrix; the key similarity is proportional to the similarity between node attributes of the key equipment node pair. S355, if the key similarity between any key node pair is greater than the preset key similarity threshold, then the two preset device nodes contained in the key node pair are determined to be the same target device node.

5. The power knowledge processing method based on multi-source data fusion according to claim 4, characterized in that, After step S351, the method further includes: S356, If the similarity between any two preset device nodes is less than the first preset similarity threshold and the second preset similarity threshold, then the corresponding two nodes are determined to be an intermediate device node pair. S357, If the absolute value of the difference in the number of adjacent nodes between two preset device nodes contained in the intermediate device node pair is greater than the preset threshold for the difference in the number of adjacent nodes, then obtain the similarity between the node attributes of the intermediate device node pair. S358, Based on the similarity of each intermediate device node pair in the target similarity matrix and the similarity between node attributes, the intermediate similarity between each intermediate device node pair is obtained; wherein, the intermediate similarity is obtained by weighted summation of the similarity of the intermediate device nodes in the target similarity matrix and the similarity between the node attributes of the intermediate device nodes, and the weight of the similarity between the node attributes of the intermediate device nodes is greater than the weight of the similarity of the intermediate device nodes in the target similarity matrix; S359, if the intermediate similarity between any pair of intermediate nodes is greater than the preset intermediate similarity threshold, then the two preset device nodes contained in the intermediate node pair are determined to be the same target device node.

6. The power knowledge processing method based on multi-source data fusion according to claim 4 or 5, characterized in that, The similarity between the node attributes is determined according to the following steps: The first node attribute similarity between each key device node pair or each intermediate device node pair is determined according to the first node attribute similarity algorithm; wherein, the first node attribute similarity algorithm is determined according to the number of string conversion operations required between the attribute values ​​of the preset attributes of two key device nodes or two intermediate device nodes. If the first node attribute similarity is less than the preset node attribute similarity threshold, then the second node attribute similarity between each key device node pair or each intermediate device node pair is determined according to the second node attribute similarity algorithm; wherein, the second node attribute similarity algorithm is determined based on the semantic similarity between two key device nodes or between two intermediate device nodes and the attribute values ​​of the preset attributes. The similarity of the second node attributes is defined as the similarity between node attributes.

7. A power knowledge processing device that integrates multi-source data, characterized in that, The device includes: The acquisition unit is used to acquire graph structure data corresponding to each preset data source of the target power system; wherein each graph structure data contains at least one target device node; and the node identifier of the target device node contained in each graph structure data is different; The mapping unit is used to perform functional mapping on each node in each graph structure data to obtain the standard functional graph structure data corresponding to each preset data source; wherein, each node in the standard functional graph structure data is a functional node; a functional node is a node named according to its function. The determination unit is used to determine each target device node in different preset data sources based on each standard functional graph structure data and a preset graph similarity algorithm; wherein, the preset graph similarity algorithm determines whether the preset device node is a target device node based on the similarity between the network structures and the similarity between the node attributes in different standard functional graph structure data. The fusion unit is used to fuse the graph structure data corresponding to each preset data source with each target device node as the fusion node, so as to obtain a multi-data source fused power knowledge graph. The determining unit is also used to perform the following steps: S310, obtain the initial similarity matrix; wherein, the initial similarity matrix is ​​n×n dimensional; n is the total number of nodes contained in all graph structure data; the similarity of key elements in the initial similarity matrix is ​​1, and the similarity of non-key elements is 0; the key element is the similarity between nodes in the i-th row and i-th column; the non-key element is the similarity between nodes in the i-th row and j-th column; i=1, 2, ..., n; j=1, 2, ..., n; i≠j; S320, determine each non-key element according to the initial similarity matrix and the preset calculation rules. S330: Obtain the current similarity value of the elements to obtain an intermediate similarity matrix; S340: If the average similarity difference is equal to or greater than a preset difference threshold, determine the current similarity value of each non-key element according to the intermediate similarity matrix and preset calculation rules to update the intermediate similarity matrix; and jump to step S330 until the average similarity difference is less than the preset difference threshold, then determine the current intermediate similarity matrix as the target similarity matrix; S350: Determine the target device node according to the similarity between the target similarity matrix and node attributes.

8. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the method as described in any one of claims 1-6.

9. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 8.