Power Grid Data Penetration and Matching Method Based on Multi-Source Data

By layering the nodes in the power grid and building a similarity evaluation model, the problem of difficulty in collecting and analyzing multi-source data in the existing technology is solved, efficient data integration and analysis is achieved, and the efficiency and security of power grid data analysis are improved.

CN119884214BActive Publication Date: 2025-06-10INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD +1

Patent Information

Application Number
CN202510369495.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-10
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art is difficult to collect and process different data through nodes at different levels, and manage and analyze the data, and it is not convenient to manage the data vertically through the data table, and at the same time, it is convenient to query horizontally.

Method used

By obtaining the topological structure data of the power grid in the station area, collecting the topological data, monitoring data, metering data and geographic information data of each node, and adding data labels to the collected data. Divide nodes into high-level, medium-level and low-level nodes, and set up edge computing task proxy nodes and data warehouses on high-level nodes. Build a similarity evaluation model among low-level nodes and establish an associated link between node topological data of high-similar nodes.

Benefits of technology

It realizes the integration and analysis of multi-source data, improves the consistency and availability of data, improves the efficiency and security of power grid data analysis, optimizes resource allocation, and improves the flexibility and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884214B_ABST
    Figure CN119884214B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for grid data penetration and matching based on multi-source data, which relates to the technical field of smart grids, and solves the technical problems that it is difficult to collect and process different data through nodes at different levels, manage and analyze the data, and it is not convenient to vertically manage the data through data tables. At the same time, it is convenient for horizontal query of data. By collecting topology data, monitoring data, metering data and geographic information data and adding tags to them, the integration of multi-source data is realized. Edge computing task proxy nodes and data warehouses are set at high-level nodes, which is convenient for data processing and analysis of the distribution network data of the substation area to each high-level node. By constructing a similarity evaluation model between low-level nodes, it is convenient to more accurately identify nodes with similar functions, performances or characteristics, establish an association link between the node topology data between similar nodes, and realize the interconnection and sharing of data between similar nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart grids, and specifically relates to a method for grid data penetration and matching based on multi-source data. Background Art

[0002] With the rapid development of smart grid technology, the operation and management of power systems have become increasingly complex. There are various types of data sources in the power grid, including but not limited to: data from devices such as sensors, smart meters, protection devices, etc., and the structural information of the power grid, including the connection relationships of devices such as nodes, lines, transformers, etc. Therefore, it is necessary to effectively integrate and match data from different sources to achieve comprehensive monitoring and analysis of the grid state. The method for grid data penetration and matching based on multi-source data needs to collect, integrate, and process data from different sources to achieve a comprehensive understanding and analysis of the grid state.

[0003] Most grid data penetration and matching solutions only standardize the data, making it difficult to collect and process different data through nodes at different levels, manage and analyze the data, and it is not convenient to manage the data longitudinally through data tables. At the same time, it is convenient for horizontal query of the data. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a method for grid data penetration and matching based on multi-source data, which is used to solve the technical problems of being difficult to collect and process different data through nodes at different levels, manage and analyze the data, and it is not convenient to manage the data longitudinally through data tables. At the same time, it is convenient for horizontal query of the data.

[0005] To solve the above problems, the first aspect of the present invention provides a method for grid data penetration and matching based on multi-source data, including the following steps:

[0006] Obtain the topological structure data of the distribution network area, collect the topological data, monitoring data, metering data, and geographical information data of each node, and add data tags to the collected data;

[0007] According to the topological structure data of the distribution network area, as well as the monitoring data and metering data of the nodes, divide the nodes into high-level nodes, middle-level nodes, and low-level nodes, set edge computing task proxy nodes and data warehouses at the high-level nodes, and set data warehouses at the middle-level nodes;

[0008] Based on the topological data and monitoring data of the nodes, group the middle-level nodes and low-level nodes into different high-level node management groups. Establish communication channels from the low-level nodes in each high-level node management group to the middle-level nodes, store the data of the low-level nodes and this node in the data warehouse. Establish communication channels from the middle-level nodes in each high-level node management group to the high-level nodes, and store the data of the middle-level nodes and this node in the data warehouse;

[0009] The edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes links between the data collected by the nodes;

[0010] According to the topological data, monitoring data, measurement data, and geographical information data of the low-level nodes, construct a similarity evaluation model between the low-level nodes. Evaluate the similarity between each node through the similarity evaluation model between the low-level nodes, and establish an association link between the topological data of the nodes with high similarity.

[0011] As a further solution of the present invention: Obtain the topological structure data of the distribution network area, collect the topological data, monitoring data, measurement data, and geographical information data of each node, and add data labels to the collected data, including the following steps:

[0012] Obtain the authorized topological structure data of the distribution network area, and collect the topological data, monitoring data, measurement data, and geographical information data of each node;

[0013] Among them, the topological data includes the degree value of the node, the weight of the node edge, and the closeness centrality. The monitoring data includes: SCADA data and equipment failure data; The measurement data includes: smart meter data and DAS distribution automation system data; The geographical information data includes: GIS data, node meteorological data;

[0014] Number the nodes in the topological structure of the distribution network area, add a unique device ID to each device in the node, and add data labels to the data collected by the node. The data labels include: timestamp, device ID, and node number.

[0015] As a further solution of the present invention: According to the topological structure data of the distribution network area, as well as the monitoring data and measurement data of the nodes, divide the nodes into high-level nodes, middle-level nodes, and low-level nodes, including the following steps:

[0016] S1: Standardize the topological structure data of the distribution network area, as well as the monitoring data and measurement data of the nodes by the Z-score standardization method;

[0017] S2: Compose the standardized data into a data group in the order of the degree value of the node, the weight of the node edge and the closeness centrality, SCADA data and equipment failure data, and smart meter data and DAS distribution automation system data;

[0018] S3: Through the K-means clustering algorithm, set the number of clusters k = 3 according to the goal of dividing the nodes into three levels: high, medium, and low;

[0019] S4: Randomly select k nodes as the initial cluster centers;

[0020] S5: Calculate the Euclidean distance between the data group of each node and the data group of the cluster center, and assign each node to the cluster center with the Euclidean distance;

[0021] S6: Calculate the ratio of the mean of the Euclidean distances between the data groups of the nodes in the cluster and the data group of the cluster center to the average value of the data in the data groups of all the nodes in the cluster as the offset distance ratio, and reselect the center of each cluster;

[0022] S7: Repeat steps S5 and S6 until the offset distance ratio reaches the preset threshold.

[0023] As a further solution of the present invention: According to the topological data and monitoring data of the nodes, group the middle-level nodes and low-level nodes into different high-level node management groups, including the following steps:

[0024] According to the topological data and monitoring data of the nodes, set the quantity thresholds of the middle-level nodes and low-level nodes for each high-level node management group. Based on the principle of the closest distance, group the middle-level nodes into the high-level node management groups. If the quantity of the middle-level nodes in the corresponding high-level node management group reaches the threshold, re-group the middle-level nodes until all the middle-level nodes are grouped; According to the same grouping principle of the middle-level nodes, group the low-level nodes into different high-level node management groups;

[0025] As a further solution of the present invention: Establish a communication channel from the low-level nodes in each high-level node management group to the middle-level nodes, including the following steps:

[0026] Take the middle-level nodes in each high-level node management group as the hubs of their communication channels, set the quantity threshold of the low-level nodes connected to the middle-level nodes, and based on the principle of the closest distance, establish communication connections between the low-level nodes and the middle-level nodes in the group;

[0027] If the quantity of the low-level nodes connected to the middle-level nodes reaches the threshold, re-establish the communication channels of the low-level nodes until all the low-level nodes complete establishing communication channels with the middle-level nodes;

[0028] Among them, a communication channel is established through the TCP communication protocol, and the data received and collected by the intermediate-level nodes are all converted into the same data format for storage.

[0029] As a further solution of the present invention: the edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes a link between the data collected by the nodes, including the following steps:

[0030] Establish a library of data standardization rules, and convert the collected data into the same format and naming specification;

[0031] Set up an edge computing task proxy node in the high-level node to establish data lists for monitoring data, metering data, and geographic information data, and establish a mapping relationship between each data list according to the data tags of the node data;

[0032] Construct key-value pairs of the topological data of the node and the data in the corresponding data list according to the mapping relationship between the data lists, assign the primary key to the topological data of the node, and assign the foreign key to the corresponding data in the data list.

[0033] As a further solution of the present invention: set up an edge computing task proxy node in the high-level node to establish data lists for monitoring data, metering data, and geographic information data, and establish a mapping relationship between each data list according to the data tags of the node data, including the following steps:

[0034] Set up an edge computing task proxy node in the high-level node to establish data lists for monitoring data, metering data, and geographic information data, and regularly detect the establishment time of each data list, transfer the number list whose establishment time exceeds the threshold to the data warehouse for storage, and replace the specific data in the data list with the storage address;

[0035] Create an index for each type of data according to the data tags of the node data, and the index is: device ID + node number + timestamp;

[0036] The mapping relationship established between each data list is:

[0037] Traverse each record in the data lists of monitoring data, metering data, and geographic information data, and search whether there is a record corresponding to the index in each data list according to the data tags of the data in the list;

[0038] If it exists, combine the found data together to form a new comprehensive record;

[0039] Store all the combined records in a new list to form a data set for querying data, including monitoring, metering, and geographic information.

[0040] As a further solution of the present invention: According to the topology data, monitoring data, metering data, and geographical information data of the low-level nodes, construct a similarity evaluation model between the low-level nodes, including the following steps:

[0041] According to the topology data of the low-level nodes, evaluate the coincidence similarity between the nodes, through the following formula:

[0042]

[0043] Where Sto(A,B) is the coincidence similarity between the nodes, N(A) is the set of neighbor nodes connected to nodes A and B, and N(B) is the set of neighbor nodes connected to node B respectively;

[0044] According to the monitoring data and metering data of the low-level nodes, form a feature data set with the monitoring data and metering data, which contains a total of n feature data, and evaluate the data similarity between the nodes, through the following formula:

[0045]

[0046]

[0047] Where Smo(A,B) is the data similarity between the nodes, Dmo(A,B) is the Euclidean distance of the i-th feature data in the feature data sets of nodes A and B, XAi is the value of the i-th feature data in the feature data set of node A, and XBi is the value of the i-th feature data in the feature data set of node B;

[0048] According to the geographical information data of the low-level nodes, evaluate the position similarity between the nodes, through the following formula:

[0049]

[0050] Where Sgo(A,B) is the position similarity between the nodes, d(A,B) is the distance between nodes A and B, and μ(A,B) is the mean value of the distances between all pairs of nodes in the group where nodes A and B are located;

[0051] According to the coincidence similarity between the low-level nodes, the data similarity between the nodes, and the position similarity between the nodes, through the following formula, construct a similarity evaluation model between the low-level nodes:

[0052]

[0053] Among them, S(A,B) is the similarity evaluation value between low-level nodes, and w1, w2, and w3 are the importance weights of the coincidence similarity, data similarity, and position similarity between nodes, respectively.

[0054] As a further solution of the present invention: evaluate the similarity between each node through the similarity evaluation model between low-level nodes, and establish an association link between the node topology data of high-similarity nodes, including the following steps:

[0055] Evaluate the similarity between each node through the similarity evaluation model between low-level nodes. If the similarity evaluation value between nodes is greater than the threshold, it is determined that the nodes are high-similarity nodes; otherwise, the nodes are low-similarity nodes.

[0056] Establish an association link between the node topology data of high-similarity nodes.

[0057] Detect the middle-level nodes stored in the topology data of high-similarity nodes. If the data of two high-similarity nodes are stored in different middle-level nodes, query the foreign key of the data according to the data labels of the other party's nodes respectively for the two high-similarity nodes, and query the topology data of the corresponding low-level nodes according to the primary key corresponding to the queried foreign key, establish a communication link between the middle-level nodes stored in the node topology data of high-similarity nodes, and associate the node topology data of high-similarity nodes.

[0058] If the data of two high-similarity nodes are stored in the same middle-level node, query the foreign key of the data according to the data labels of the other party's nodes respectively for the two high-similarity nodes, and query the topology data of the corresponding low-level nodes according to the primary key corresponding to the queried foreign key, and associate the node topology data of high-similarity nodes.

[0059] Compared with the prior art, the beneficial effects of the present invention are:

[0060] The present invention integrates multi-source data by collecting topology data, monitoring data, metering data, and geographic information data and adding labels to them, ensuring that data from different sources can be analyzed within the same framework, and improving the consistency and availability of data. Edge computing task proxy nodes and data warehouses are set at high-level nodes, facilitating the data processing and analysis of the data of the distribution network in the area dispersed to each high-level node, and improving the operation efficiency and security of the power grid data analysis.

[0061] The present invention ensures the consistency of data from different sources and formats by standardizing the collected data, making subsequent data analysis more accurate and reliable. At the same time, by constructing a similarity evaluation model between low-level nodes, it is convenient to more accurately identify nodes with similar functions, performances, or characteristics, establish the association links between the node topology data of similar nodes, realize the interconnection and sharing of data between similar nodes, and also facilitate the subsequent unified analysis of the data of similar nodes, providing support for subsequent data analysis and decision-making.

[0062] By establishing association links between high-similarity nodes, the present invention also helps to optimize resource allocation, can quickly locate problems when a failure occurs, and use the data of other normal nodes for fault analysis and recovery, improving the reliability of the power grid. At the same time, it is convenient for subsequent collaborative analysis between high-similarity nodes, enhancing the flexibility and response speed of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0066] Please refer to Figure 1 , the first aspect embodiment of the present invention provides a power grid data penetration and matching method based on multi-source data, including the following steps:

[0067] Obtain the topological structure data of the distribution network area, collect the topological data, monitoring data, metering data, and geographic information data of each node, and add data tags to the collected data;

[0068] According to the topological structure data of the distribution network area, as well as the monitoring data and metering data of the nodes, the nodes are divided into high-level nodes, middle-level nodes, and low-level nodes. An edge computing task proxy node and a data warehouse are set at the high-level nodes, and a data warehouse is set at the middle-level nodes;

[0069] According to the topological data and monitoring data of the nodes, group the middle-level nodes and low-level nodes into different high-level node management groups. Establish communication channels from the low-level nodes in each high-level node management group to the middle-level nodes, store the data of the low-level nodes and this node in the data warehouse. Establish communication channels from the middle-level nodes in each high-level node management group to the high-level nodes, and store the data of the middle-level nodes and this node in the data warehouse;

[0070] The edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes links between the data collected by the nodes;

[0071] According to the topological data, monitoring data, measurement data, and geographic information data of the low-level nodes, construct a similarity evaluation model between the low-level nodes. Evaluate the similarity between each node through the similarity evaluation model between the low-level nodes, and establish an association link between the topological data of the nodes with high similarity.

[0072] Specifically, in this embodiment, obtain the topological structure data of the distribution network area, collect the topological data, monitoring data, measurement data, and geographic information data of each node, and add data labels to the collected data; according to the topological structure data of the distribution network area, as well as the monitoring data and measurement data of the nodes, divide the nodes into high-level nodes, middle-level nodes, and low-level nodes. Set edge computing task proxy nodes and data warehouses at the high-level nodes, and set data warehouses at the middle-level nodes; by collecting topological data, monitoring data, measurement data, and geographic information data and adding labels to them, realize the integration of multi-source data, ensure that data from different sources is analyzed under the same framework, and improve the consistency and availability of data. Setting edge computing task proxy nodes and data warehouses at high-level nodes facilitates the data processing and analysis of dispersing the distribution network area data to each high-level node for processing, and improves the operation efficiency and security of power grid data analysis. Dividing the nodes into high, middle, and low levels helps to establish a hierarchical management mechanism. High-level nodes are responsible for global coordination and decision-making, middle-level nodes are responsible for data storage and processing, and low-level nodes focus on specific monitoring and measurement work. This hierarchical structure can optimize resource allocation, improve the system response speed. It is also convenient for analyzing the real-time monitoring data of each node, timely discovering potential problems, reducing the probability of failures, and improving the reliability of power grid operation. Through the edge computing and the setting of the data warehouse at the middle-level node, the data of the low-level nodes is aggregated and integrated through the middle-level node, which is convenient for flexibly adjusting task allocation and resource use according to actual needs, and improving the adaptability of the system.

[0073] According to the topological data and monitoring data of the nodes, group the middle-level nodes and low-level nodes into different high-level node management groups. Establish a communication channel from the low-level nodes to the middle-level nodes through the low-level nodes in each high-level node management group, store the data of the low-level nodes and this node in the data warehouse. Establish a communication channel from the middle-level nodes to the high-level nodes through the middle-level nodes in each high-level node management group, and store the data of the middle-level nodes and this node in the data warehouse;

[0074] By establishing communication channels from low-level nodes to middle-level nodes and from middle-level nodes to high-level nodes, it facilitates fast and stable data transmission between low-level nodes and middle-level nodes, reduces latency, increases the data update frequency, and also facilitates data transmission to high-level nodes.

[0075] Grouping middle-level and low-level nodes into different high-level node management helps to achieve distributed management, improve the scalability and flexibility of the system, and facilitate coping with future expansion requirements. By real-time monitoring the data of low-level nodes and reporting to middle-level and high-level nodes, problems can be quickly identified and responses can be made, improving the reliability and security of power grid operation.

[0076] The edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes links between the data collected by the nodes; according to the topological data, monitoring data, metering data and geographical information data of the low-level nodes, construct a similarity evaluation model between low-level nodes, evaluate the similarity between each node through the similarity evaluation model between low-level nodes, and establish an association link between the node topological data of high-similarity nodes;

[0077] By standardizing the collected data, it ensures the consistency of data from different sources and formats, making subsequent data analysis more accurate and reliable. At the same time, by constructing a similarity evaluation model between low-level nodes, it is convenient to more accurately identify nodes with similar functions, performances or characteristics, establish an association link between the node topological data of similar nodes, realize data intercommunication and sharing between similar nodes, and also facilitate subsequent unified analysis of similar node data, providing support for subsequent data analysis and decision-making.

[0078] By establishing an association link between high-similarity nodes, it also helps to optimize resource allocation, can quickly locate problems in case of failures, and draw on the data of other normal nodes for fault analysis and recovery, improving the reliability of the power grid. At the same time, it is convenient for subsequent collaborative analysis between high-similarity nodes, enhancing the flexibility and response speed of the overall system.

[0079] In one embodiment of the present invention, obtaining the topological structure data of the distribution network area, collecting the topological data, monitoring data, metering data, and geographic information data of each node, and adding data tags to the collected data, including the following steps:

[0080] Obtaining the authorized topological structure data of the distribution network area, and collecting the topological data, monitoring data, metering data, and geographic information data of each node;

[0081] Among them, the topological data includes the degree value of the node, the weight of the node edge, and the closeness centrality. The monitoring data includes: SCADA data and equipment failure data; The metering data includes: smart meter data and DAS distribution automation system data; The geographic information data includes: GIS data, node meteorological data;

[0082] Numbering the nodes in the topological structure of the distribution network area, adding a unique device ID to each device in the node, and adding data tags to the data collected by the node. The data tags include: timestamp, device ID, and node number.

[0083] Specifically, in this embodiment, the node number is set in the following manner:

[0084] According to the topological structure of the distribution network area, assign a unique number to each node. For example, assuming there are 50 nodes, numbered Node_001, Node_002,..., Node_050 respectively.

[0085] Device ID assignment:

[0086] For each device in the node, assign a unique device ID. The format of "node number + device serial number" can be used to generate the device ID. For example:

[0087] The devices in Node_001 can be marked as Device_001-01, Device_001-02;

[0088] The devices in Node_002 can be marked as Device_002-01, Device_002-02; and so on.

[0089] When collecting data, add corresponding data tags to each data record, including timestamp, device ID, and node number.

[0090] The timestamp can adopt a standard format, such as ISO 8601, for example, "2023-10-01T12:00:00Z".

[0091] Data tag example:

[0092] timestamp: 2023-10-01T12:00:00Z, device_id: Device_001-01, node_id: Node_001。

[0093] In one embodiment of the present invention, according to the topological structure data of the distribution network area, as well as the monitoring data and metering data of the nodes, the nodes are divided into high-level nodes, middle-level nodes, and low-level nodes, including the following steps:

[0094] S1: Standardize the topological structure data of the distribution network area, as well as the monitoring data and metering data of the nodes by the Z-score standardization method;

[0095] S2: According to the order of the degree value of the node, the weight of the node edge and closeness centrality, SCADA data and device fault data, smart meter data and DAS distribution automation system data, form a data group with the standardized data;

[0096] S3: Through the K-means clustering algorithm, according to the goal of dividing the nodes into three levels: high, middle, and low, set the number of clusters k = 3;

[0097] S4: Randomly select k nodes as the initial cluster centers;

[0098] S5: Calculate the Euclidean distance between the data group of each node and the data group of the cluster center, and assign each node to the cluster center with the Euclidean distance;

[0099] S6: Calculate the ratio of the mean of the Euclidean distance between the data group of the nodes in the cluster and the data group of the cluster center to the average value of the data in the data groups of all the nodes in the cluster as the offset distance ratio, and reselect the center of each cluster;

[0100] S7: Repeat steps S5 and S6 until the offset distance ratio reaches the preset threshold.

[0101] In this embodiment, collect relevant data, including:

[0102] Node degree value: The number of connections of each node, reflecting its importance in the network.

[0103] Edge weight: The weight of each edge, indicating the importance of the connection, such as current capacity, impedance, etc.

[0104] Closeness centrality: Calculate the average distance between each node and other nodes, reflecting its central degree in the network.

[0105] SCADA data: Real-time monitoring data, such as current, voltage, power, etc.

[0106] Equipment failure data: Historical failure records, including failure frequency and scope of impact.

[0107] Smart meter data: Electricity demand and electricity consumption pattern information.

[0108] DAS data: Status information in the distribution automation system, such as switch status and fault detection.

[0109] In one embodiment of the present invention, according to the topological data and monitoring data of nodes, grouping middle-level nodes and low-level nodes into different high-level node management groups includes the following steps:

[0110] According to the topological data and monitoring data of nodes, set the quantity thresholds of middle-level nodes and low-level nodes for each high-level node management group. Based on the principle of the closest distance, group the middle-level nodes into the high-level node management groups. If the quantity of middle-level nodes in the corresponding high-level node management group reaches the threshold, re-group the middle-level nodes until all middle-level nodes are grouped; according to the same grouping principle of middle-level nodes, group the low-level nodes into different high-level node management groups;

[0111] In one embodiment of the present invention, establishing a communication channel from low-level nodes to middle-level nodes through each high-level node management group includes the following steps:

[0112] Take the middle-level nodes in each high-level node management group as the hubs of their communication channels, set the quantity threshold of low-level nodes connected to the middle-level nodes, and based on the principle of the closest distance, establish communication connections between the low-level nodes and the middle-level nodes in the group;

[0113] If the quantity of low-level nodes connected to the middle-level nodes reaches the threshold, re-establish the communication channels of the low-level nodes until all low-level nodes complete establishing communication channels with the middle-level nodes;

[0114] Among them, establish a communication channel through the TCP communication protocol, and all the data received and collected in the middle-level nodes are converted into the same data format for storage.

[0115] In one embodiment of the present invention, the edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes links between the data collected by the nodes, including the following steps:

[0116] Establish a library of data standardization rules, and convert the collected data into the same format and naming specification;

[0117] Set up edge computing task proxy nodes at high-level nodes to establish data lists for monitoring data, metering data, and geographic information data, and establish mapping relationships between the data lists according to the data tags of the node data;

[0118] Construct key-value pairs of the topological data of the nodes and the data in the corresponding data lists according to the mapping relationships between the data lists, assign the primary key to the topological data of the nodes, and assign the foreign key to the corresponding data in the data lists.

[0119] In one embodiment of the present invention, setting up edge computing task proxy nodes at high-level nodes to establish data lists for monitoring data, metering data, and geographic information data, and establishing mapping relationships between the data lists according to the data tags of the node data includes the following steps:

[0120] Set up edge computing task proxy nodes at high-level nodes to establish data lists for monitoring data, metering data, and geographic information data, regularly detect the establishment time of each data list, transfer the quantity list whose establishment time exceeds the threshold to the data warehouse for storage, and replace the specific data in the data list with the storage address;

[0121] Create an index for each type of data according to the data tags of the node data. The index is: device ID + node number + timestamp;

[0122] The mapping relationship established between the data lists is:

[0123] Traverse each record in the data lists of monitoring data, metering data, and geographic information data, and find whether there is a record corresponding to the index in each data list according to the data tags of the data in the list;

[0124] If it exists, combine the found data together to form a new comprehensive record;

[0125] Store all the combined records in a new list to form a data set for querying data, including monitoring, metering, and geographic information.

[0126] In one embodiment of the present invention, constructing a similarity evaluation model between low-level nodes according to the topological data, monitoring data, metering data, and geographic information data of low-level nodes includes the following steps:

[0127] Evaluate the coincidence similarity between nodes according to the topological data of low-level nodes through the following formula:

[0128]

[0129] Among them, Sto(A,B) is the coincidence similarity between nodes, N(A) is the set of neighbor nodes connected to nodes A and B, and N(B) is the set of neighbor nodes connected to node B respectively;

[0130] According to the monitoring data and measurement data of low-level nodes, the monitoring data and measurement data are combined into a feature data set, which contains a total of n feature data, and the data similarity between nodes is evaluated through the following formula:

[0131]

[0132]

[0133] Among them, Smo(A,B) is the data similarity between nodes, Dmo(A,B) is the Euclidean distance of the i-th feature data in the feature data sets of nodes A and B, XAi is the value of the i-th feature data in the feature data set of node A, and XBi is the value of the i-th feature data in the feature data set of node B;

[0134] According to the geographical information data of low-level nodes, the position similarity between nodes is evaluated through the following formula:

[0135]

[0136] Among them, Sgo(A,B) is the position similarity between nodes, d(A,B) is the distance between nodes A and B, and μ(A,B) is the mean value of the distances between all pairs of nodes in the group where nodes A and B are located;

[0137] According to the coincidence similarity between low-level nodes, the data similarity between nodes, and the position similarity between nodes, the following formula is used to construct a similarity evaluation model between low-level nodes:

[0138]

[0139] Among them, S(A,B) is the similarity evaluation value between low-level nodes, and w1, w2, and w3 are the importance weights of the coincidence similarity between nodes, the data similarity between nodes, and the position similarity between nodes respectively.

[0140] Specifically, in this embodiment, based on the calculation and evaluation of the similarity evaluation values between a large number of low-level nodes, the importance weight w1 of the coincidence similarity is set to 0.2, the importance weight w2 of the data similarity between nodes is set to 0.6, and the importance weight w3 between nodes is set to 0.2.

[0141] In one embodiment of the present invention, the similarity between each node is evaluated through a similarity evaluation model between low-level nodes, and an association link is established between the node topology data of high-similarity nodes, including the following steps:

[0142] The similarity between each node is evaluated through a similarity evaluation model between low-level nodes. If the similarity evaluation value between nodes is greater than the threshold, the node is determined to be a node between high-similarity nodes; otherwise, the node is a node between low-similarity nodes.

[0143] An association link is established between the node topology data of high-similarity nodes.

[0144] The middle-level nodes stored in the topology data of the nodes between high-similarity nodes are detected. If the data of two high-similarity nodes are stored in different middle-level nodes, the foreign keys of the data are queried according to the data labels of the other party's nodes by the two high-similarity nodes respectively. According to the primary keys corresponding to the queried foreign keys, the topology data of the corresponding low-level nodes are queried, a communication link is established between the middle-level nodes stored in the node topology data of the high-similarity nodes, and the node topology data of the high-similarity nodes are associated.

[0145] If the data of two high-similarity nodes are stored in the same middle-level node, the foreign keys of the data are queried according to the data labels of the other party's nodes by the two high-similarity nodes respectively. According to the primary keys corresponding to the queried foreign keys, the topology data of the corresponding low-level nodes are queried, and the node topology data of the high-similarity nodes are associated.

[0146] Specifically, in this embodiment, the similarity between each node is evaluated through a similarity evaluation model between low-level nodes. If the similarity evaluation value between nodes is greater than 0.5, the node is determined to be a node between high-similarity nodes; otherwise, the node is a node between low-similarity nodes.

[0147] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A power grid data connection and matching method based on multi-source data, characterized in that: The following steps are involved: Obtain the topological structure data of the power grid in the substation area, and collect the topological data, monitoring data, metering data and geographic information data of each node, and add data labels to the collected data; According to the topological structure data of the substation power grid, as well as the monitoring data and metering data of the nodes, the nodes are divided into high-level nodes, middle-level nodes and low-level nodes. Edge computing task agent nodes and data warehouses are set up at high-level nodes, and data warehouses are set up at middle-level nodes. According to the topological data and monitoring data of the nodes, the middle-level nodes and the low-level nodes are grouped into different high-level node management groups, a communication channel to the middle-level nodes is established through the low-level nodes in each high-level node management group, and the data of the low-level nodes and the current node are stored in the data warehouse. A communication channel to the high-level nodes is established through the middle-level nodes in each high-level node management group, and the data of the middle-level nodes and the current node are stored in the data warehouse; The edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes links between the data collected by the node; Based on the topological data, monitoring data, metering data and geographic information data of low-level nodes, a similarity evaluation model between low-level nodes is constructed, and the similarity between each node is evaluated through the similarity evaluation model between low-level nodes, and the association link between the node topological data of high-similarity nodes is established; Among them, based on the topological data, monitoring data, metering data and geographic information data of the low-level nodes, a similarity evaluation model between low-level nodes is constructed, including the following steps: According to the topological data of the low-level nodes, the coincidence similarity between nodes is evaluated using the following formula: ; Among them, Sto(A,B) is the coincidence similarity between nodes, N(A) is the set of neighbor nodes connected to nodes A and B, and N(B) is the set of neighbor nodes connected to node B; According to the monitoring data and metering data of the low-level nodes, the monitoring data and metering data are combined into a feature data set, which contains a total of n feature data. The data similarity between nodes is evaluated by the following formula: ; ; Among them, Smo(A,B) is the data similarity between nodes, Dmo(A,B) is the Euclidean distance of the i-th feature data in the feature data sets of nodes A and B, XAi is the value of the i-th feature data in the feature data set of node A, and XBi is the value of the i-th feature data in the feature data set of node B; According to the geographic information data of the low-level nodes, the location similarity between nodes is evaluated using the following formula: ; Among them, Sgo(A,B) is the position similarity between nodes, d(A,B) is the distance between nodes A and B, and μ(A,B) is the mean distance between all nodes in the group where nodes A and B are located; According to the overlap similarity between low-level nodes, the data similarity between nodes, and the position similarity between nodes, a similarity evaluation model between low-level nodes is constructed using the following formula: ; Among them, S(A, B) is the similarity evaluation value between low-level nodes, w1, w2 and w3 are the importance weights of the coincidence similarity between nodes, the data similarity between nodes and the position similarity between nodes, respectively.

2. The power grid data connection and matching method based on multi-source data according to claim 1 is characterized in that: Obtaining the topological structure data of the power grid in the substation area, collecting the topological data, monitoring data, metering data and geographic information data of each node, and adding data labels to the collected data, including the following steps: Obtain the authorized area power grid topology data, and collect the topology data, monitoring data, metering data and geographic information data of each node; Among them, topological data includes node degree, node edge weight and proximity centrality; monitoring data includes SCADA data and equipment failure data; metering data includes smart meter data and DAS distribution automation system data; geographic information data includes GIS data and node meteorological data; The nodes in the topological structure of the substation power grid are numbered, and a unique device ID is added to each device in the node. Data tags are added to the data collected by the node. The data tags include: timestamp, device ID and node number.

3. The power grid data connection and matching method based on multi-source data according to claim 1 is characterized in that: According to the topological structure data of the power grid in the substation area, as well as the monitoring data and metering data of the nodes, the nodes are divided into high-level nodes, middle-level nodes and low-level nodes, including the following steps: S1: The topological structure data of the substation power grid, as well as the monitoring data and metering data of the nodes are standardized by the Z-score standardization method; S2: The standardized data are organized into data groups according to the order of node degree, node edge weight and closeness centrality, SCADA data and equipment failure data, smart meter data and DAS distribution automation system data; S3: Using the K-means clustering algorithm, the goal is to divide the nodes into three levels: high, medium, and low, and set the number of clusters k=3; S4: Randomly select k nodes as initial cluster centers; S5: Calculate the Euclidean distance between the data group of each node and the data group of the cluster center, and assign each node to the cluster center of the Euclidean distance; S6: Calculate the ratio of the mean of the Euclidean distance between the data group of the nodes in the cluster and the data group of the cluster center to the mean of the data in the data groups of all the nodes in the cluster as the offset distance ratio, and reselect the center of each cluster; S7: Repeat steps S5 and S6 until the offset distance ratio reaches a preset threshold.

4. The power grid data connection and matching method based on multi-source data according to claim 1, characterized in that: According to the topological data and monitoring data of the nodes, the middle-level nodes and the low-level nodes are grouped into different high-level node management groups, including the following steps: According to the topological data and monitoring data of the nodes, the management grouping of each high-level node sets the threshold for the number of middle-level nodes and low-level nodes. The middle-level nodes are grouped into the high-level node management grouping based on the principle of the closest distance. If the number of middle-level nodes in the management grouping of the corresponding high-level node reaches the threshold, the middle-level nodes are grouped again until all the middle-level nodes are grouped; according to the same grouping principle for the middle-level nodes, the low-level nodes are grouped into different high-level node management groups.

5. The power grid data connection and matching method based on multi-source data according to claim 1 is characterized in that: Establishing a communication channel to a middle-level node through a low-level node in each high-level node management group includes the following steps: The middle-level nodes in each high-level node management group are used as the hub of its communication channel. A threshold value for the number of low-level nodes connected to the middle-level nodes is set. Based on the principle of the closest distance, the low-level nodes in the group are connected to the middle-level nodes. If the number of low-level nodes connected to the middle-level node reaches the threshold, the communication channel of the low-level nodes is re-established until all the low-level nodes have completed the establishment of communication channels with the middle-level nodes; Among them, a communication channel is established through the TCP communication protocol, and the data received and collected in the middle-level nodes are converted into the same data format for storage.

6. The method for connecting and matching power grid data based on multi-source data according to claim 1, characterized in that: The edge computing task proxy node of the high-level node standardizes the data collected by the node and establishes a link between the node collected data, including the following steps: Establish a library of data standardization rules to convert the collected data into the same format and naming conventions; Set up edge computing task agent nodes at high-level nodes to establish data lists of monitoring data, metering data, and geographic information data, and establish mapping relationships between various data lists based on data labels of node data; According to the mapping relationship between the data lists, the key-value pairs of the topological data of the node and the data in the corresponding data list are constructed, the primary key is assigned to the topological data of the node, and the foreign key is assigned to the corresponding data in the data list.

7. The method for connecting and matching power grid data based on multi-source data according to claim 6, characterized in that: Setting edge computing task agent nodes at high-level nodes to establish data lists of monitoring data, metering data, and geographic information data, and establishing mapping relationships between various data lists based on data labels of node data, including the following steps: Set up edge computing task agent nodes at high-level nodes to establish data lists of monitoring data, metering data, and geographic information data, and regularly detect the creation time of each data list, transfer the number of lists whose creation time exceeds the threshold to the data warehouse for storage, and replace the specific data in the data list with the storage address; According to the data label of the node data, create an index for each type of data. The index is: device ID + node number + timestamp; The mapping relationship between each data list is established as follows: Traverse each record in the data lists of monitoring data, metering data, and geographic information data, and find out whether there is a record corresponding to the index in each data list according to the data label of the data in the list; If it exists, the found data are combined together to form a new comprehensive record; All combined records are stored in a new list to form a dataset for query data, including monitoring, measurement and geographic information.

8. The power grid data connection and matching method based on multi-source data according to claim 1 is characterized in that: The similarity between each node is evaluated by a similarity evaluation model between low-level nodes, and an association link between node topology data of high-similarity nodes is established, including the following steps: The similarity between each node is evaluated through the similarity evaluation model between low-level nodes. If the similarity evaluation value between nodes is greater than the threshold, the node is judged to be a high-similarity node, otherwise, the node is a low-similarity node; Establishing association links between node topology data of highly similar nodes; Detect the middle-level nodes where the topological data of nodes between high-similarity nodes are stored. If the data of two high-similarity nodes are stored in different middle-level nodes, query the foreign keys of the data according to the data labels of the other nodes respectively. According to the primary keys corresponding to the queried foreign keys, query the topological data of the lower-level nodes of the corresponding nodes, establish the communication links between the middle-level nodes where the node topological data of the high-similarity nodes are stored, and associate the node topological data of the high-similarity nodes. If the data of two highly similar nodes are stored in the same middle-level node, the two highly similar nodes will query the foreign keys of the data based on the data labels of the other nodes, and query the topological data of the corresponding low-level nodes based on the primary keys corresponding to the queried foreign keys, and associate the node topological data of the highly similar nodes.

Citation Information

Patent Citations

  • Topology intelligent equipment management method and system based on graph database

    CN117453959A

  • Power grid topology automatic identification and construction method and system

    CN118133068A

Cited By

  • Generative adversarial network-based power grid monitoring signal data sample enhancement method and system

    CN121117605A