Automatic error correction method and system for marketing metering box archive topological graph based on RPA

By using RPA and machine learning technologies, combined with OCR and DTW distance, we trained a fault classification model and optimized the topology map of the marketing meter box archives, solving the problem of inaccurate correction in existing technologies and achieving efficient and intelligent topology map management.

CN120687914AActive Publication Date: 2025-09-23MARKETING SERVICE CENT OF STATE GRID JILIN ELECTRIC POWER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511174698.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-23
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

The existing correction scheme for the marketing meter box archive topology map fails to effectively consider the probability of node failure type, spatiotemporal correlation and characteristic data, resulting in inaccurate correction.

Method used

An RPA robot is used to automatically log in to the system, identify topology information through OCR and computer vision, calculate the DTW distance based on historical traffic data, train the fault classification model, and use the improved GAPSO algorithm for iterative calculation to optimize the topology structure.

Benefits of technology

It improves the accuracy and efficiency of topology map correction, realizes intelligent and adaptive marketing meter box file management, reduces human errors, and enhances the digital transformation capabilities of power marketing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687914A_ABST
    Figure CN120687914A_ABST
Patent Text Reader

Abstract

The invention discloses an RPA-based marketing metering box archive topological graph automatic deviation rectification method and system, relates to the technical field of electric power analysis, and solves the technical problem that the reasonability analysis of a metering box archive topological graph and the deviation rectification accuracy are not considered by multi-dimensional data such as node fault type probability, space-time correlation and feature data. By calculating the DTW distance between the nodes, a correlation matrix is generated, dynamic correlation hidden in a topological graph is revealed, and the spatial-temporal correlation of the nodes can be recognized. According to the topological structure after node rectification, rectification is carried out on an original digital topological structure, through multi-dimensional risk quantification, hybrid algorithm optimization and hierarchical rectification strategies, precision, intelligence and self-adaption of marketing metering box archive topological graph rectification are achieved, the node comprehensive risk, topological importance and algorithm efficiency are deeply fused, and the method has the advantages of being high in practicability and easy to popularize. And the correction efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power analysis, and specifically relates to an automated deviation correction method and system for a marketing meter box archive topology map based on RPA. Background Art

[0002] RPA (Robotic Process Automation) automates repetitive and regular tasks by mimicking human workflow on a computer. By simulating manual operations, RPA robots, combined with image recognition (OCR / CV) and a rules engine, automatically compare the topology of marketing meter box archives with standard templates, identifying offset nodes, incorrect connections, or missing data, and triggering a correction process, achieving a detection-analysis-correction process. Marketing meter box archives contain critical data such as device information, connection relationships, and operating status, requiring a more automated system.

[0003] The existing automated correction scheme for marketing meter box archive topology maps only corrects the meter box archive topology maps through historical traffic data, without considering the probability of node failure types, spatiotemporal correlation, and characteristic data and other multi-dimensional data to analyze the rationality of the meter box archive topology maps and perform corrections accurately. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes an RPA-based automated correction method and system for the marketing meter box archive topology map, which is used to solve the technical problem of the accuracy of rationality analysis and correction of the meter box archive topology map without considering multi-dimensional data such as node failure type probability, spatiotemporal correlation, and feature data.

[0005] To solve the above problems, the first aspect of the present invention provides an automated correction method for a marketing meter box archive topology map based on RPA, comprising the following steps: The RPA robot automatically logs into the marketing management system, GIS system, and design drawing library, extracts the topology map and associated data of the meter box archive, extracts text information from the topology map through OCR text recognition, and uses computer vision to identify node locations, connection line directions, and equipment icon types to build a digital topology structure. Obtain historical flow data of each node to identify meter box nodes, user nodes and transformer nodes, and set the level of each node according to historical flow data and node historical data; Calculate the DTW dynamic time warping distance between nodes, identify spatiotemporal correlations, and perform feature extraction on node data for nodes of different spatiotemporal correlations and levels; Based on the extracted feature data, a fault classification model is trained for nodes with different spatiotemporal correlations, and the probability of node fault type is output; According to the probability analysis results of each node failure type, node hierarchical data, and node degree centrality, the hierarchical coefficient is introduced to calculate the comprehensive risk of the node. Through the improved GAPSO algorithm, the fitness function and iteration termination conditions are set. After iterative calculation, the topological structure of each level node after correction is obtained. The original digital topological structure is corrected according to the topological structure of the node after correction.

[0006] Optionally, in an example of the above aspect, obtaining historical traffic data of each node to identify a meter box node, a user node, and a transformer node, and obtaining meter box data, business system data, and external environment data of the node include the following steps: Obtain historical flow data of meter box nodes, user nodes, and transformer nodes, label the corresponding node types of the flow data, and input it into the LSTM model to train the model to identify node types; Obtain historical traffic data for each node and input it into the trained LSTM model to identify the node type.

[0007] Optionally, in an example of the above aspect, setting each node level according to historical traffic data and node historical data includes the following steps: Based on the node type data, user nodes are divided into access layer nodes; Based on historical traffic data and node historical data, select nodes that meet the following conditions: the average daily traffic is greater than the 80th percentile of all nodes; the traffic standard deviation is less than the 20th percentile of all nodes; the number of directly connected nodes is greater than 10; The selected nodes are used as core layer nodes; The remaining unclassified nodes are divided into aggregation layer nodes.

[0008] Optionally, in an example of the above aspect, calculating the DTW dynamic time warping distance between nodes and identifying spatiotemporal correlation includes the following steps: Obtain the historical traffic data of each node, convert the node's historical traffic data into equal time series X and Y, with lengths of |X| and |Y| respectively, and ensure that the timestamps of all time series are aligned; Calculate the Euclidean distance D(i,j) between any two nodes to form an m×n distance matrix, where m and n are the sequence lengths; Create a (m+1)×(n+1) cumulative distance matrix D, D(0,0)=0, and set the first row and first column to infinity; All nodes are grouped into a node set, and the Euclidean distance between any two nodes is calculated according to the recursive formula D(i,j) = d(i,j) + min(D(i-1,j), D(i,j-1),D(i-1,j-1)), where d(i,j) is the traffic difference between node i and node j within the time series range, D(i-1,j) is the Euclidean distance between node i-1 and node j, D(i,j-1) is the Euclidean distance between node i and node j-1, and D(i-1,j-1) is the Euclidean distance between node i-1 and node j-1. Use DTW distance as edge weight to build a spatial relationship matrix between nodes; ‌Calculate the DTW distance at different time granularities of day, week, and month respectively to obtain the multi-scale DTW distance feature; The multi-scale DTW distance matrix is ​​fused through weighted summation or attention mechanism to form a spatiotemporal correlation DTW distance matrix; Use the spatiotemporal correlation DTW distance matrix to build a hierarchical tree and obtain clustering results through pruning; Directly use DTW distance as the similarity metric to divide the nodes into K clusters; Nodes are regarded as graph vertices, the inverse of DTW distance is used as edge weight, and the Louvain algorithm is used for node clustering; The nodes that are clustered into a unified cluster are regarded as spatiotemporal correlation nodes.

[0009] Optionally, in an example of the above aspect, for nodes of different spatiotemporal correlations and levels, feature extraction is performed on node data, including the following steps: According to the historical traffic data of each node in the spatiotemporal correlation node, a time-shift sliding window W is introduced, and the feature data within the window W is calculated, including: mean, variance, slope and frequency domain features.

[0010] Based on the node hierarchy, the transmission delay time of the upstream node connected to the node and the load of the downstream node connected to the node are counted; Calculate the ratio of the transmission delay time of the upstream node to the average transmission delay time in the historical data of the entire meter box archive topology map to form the upstream node propagation delay coefficient matrix; The ratio of the downstream node load connected to the calculation node to the node load average in the historical data of the entire meter box archive topology diagram is used to form the downstream node load influence coefficient matrix; The obtained coefficient matrix and characteristic data are used as the characteristic data of the node.

[0011] Optionally, in an example of the above aspect, training a fault classification model for nodes with different spatiotemporal correlations based on the extracted feature data and outputting node fault type probabilities includes the following steps: The feature data of the node are fused at different time granularities, such as day, week, and month. The attention mechanism is used to dynamically assign weights to obtain the fused feature data. The feature matrix of the node is obtained: (Nn, Twindows, G), where Nn is the node number, Twindows is the feature sequence, and G corresponds to the fused feature data. Extract the connection relationship topology diagram of nodes with different spatiotemporal correlations, number the nodes with different spatiotemporal correlations, and obtain the adjacency matrix of the connection relationship of each node based on the node number information adjacent to the nodes with different spatiotemporal correlations; Build a GCN-LSTM fault classification model: Establish a GCN layer with the input being the node feature matrix + adjacency matrix; Aggregate the node feature matrix and adjacency matrix information through the GCN layer and output the node's spatiotemporal feature matrix; Connect the output of the GCN layer to the input of the LSTM layer, and input the spatiotemporal feature matrix output by the GCN into the LSTM layer; The LSTM layer outputs time coding features, which are input into the fully connected layer + Softmax layer to output the fault type probability.

[0012] Optionally, in an example of the above aspect, calculating the node comprehensive risk based on the probability analysis results of each node failure type, node hierarchy data, and node degree centrality, and introducing a hierarchy coefficient, includes the following steps: Obtain the probability analysis results of each node's fault type, node-level data, and node degree centrality to quantify node risk; Calculate node risk value , where,Ri is the comprehensive failure probability; Li is the layer coefficient, the core layer coefficient = 3, the aggregation layer coefficient = 2, and the access layer coefficient = 1; Di′ is the normalized degree centrality = number of node connections / maximum possible number of connections, and β is the weight controlling the influence of centrality; Calculate the link risk value Sij from node i to connected node j: , where Delayij is the delay time from node i to connected node j, Max_Delay is the maximum value of the transmission delay time in the historical data of the entire meter box archive topology map, Lossij is the energy loss of power transmission from node i to connected node j, Max_Loss is the maximum value of the energy loss of power transmission between nodes in the historical data of the entire meter box archive topology map, wloss and wdelay are the corresponding weights, wloss=0.3, wdelay=0.7; The average of the obtained link risk values ​​of the node and the connected nodes is then weighted averaged with the node risk value to obtain the node comprehensive risk value.

[0013] Optionally, in an example of the above aspect, the improved GAPSO algorithm is used to set a fitness function and an iteration termination condition, and perform iterative calculations to obtain a topological structure of nodes at each level after correction, including the following steps: The improved GAPSO algorithm is used to perform particle encoding, where each particle represents a topological structure and its dimension is the number of node connections. Set the fitness function: , where Fitness is the fitness value corresponding to the corresponding topological structure; Z_Risk is the mean of the comprehensive risk values ​​of each node under the corresponding topological structure; Connectivity is the reciprocal of the mean number of connected components of each node under the corresponding topological structure; Set dynamic neighbor rules: core nodes are forced to retain three redundant paths; The node degree limit of the aggregation layer is 8. If it exceeds 8, high-risk links will be randomly deleted; Set the termination condition: 100 iterations or the fitness value changes by <2% for 10 consecutive times; The particles with the highest fitness values ​​are selected and output the optimal topology as the topology structure after correction.

[0014] Optionally, in an example of the above aspect, correcting the original digitized topology structure according to the topology structure after node correction includes the following steps: A1: The node comprehensive risk value of each node in the statistically corrected topological structure, as well as the node comprehensive risk value of each node in the original digital topological structure; A2: Compare the comprehensive risk values ​​of each corresponding node and calculate the difference between the comprehensive risk values ​​of the node before and after correction. If the difference is greater than the preset threshold, modify the topological connection of the corresponding node based on the topological structure of the node after correction. A3: For nodes that are missing connections after the topology structure is modified, supplementary connections are made based on the topology structure of the nodes after correction, and the process returns to step A1 until the difference between the comprehensive risk values ​​of each node before and after correction is less than the preset threshold.

[0015] According to another aspect of the present disclosure, an RPA-based automated correction system for marketing meter box archive topology maps is provided. The system adopts the above-mentioned RPA-based automated correction method for marketing meter box archive topology maps to realize automated correction of marketing meter box archive topology maps.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This method calculates the DTW distance between nodes to generate a correlation matrix, revealing hidden dynamic associations in the topology and helping to identify the spatiotemporal correlations of nodes. For nodes with different spatiotemporal correlations, fault classification models are trained for different node types. Input features include traffic characteristics and hierarchical coefficients. The model outputs the probability of node fault type, improving the accuracy of fault diagnosis.

[0017] This method corrects the original digital topology based on the corrected node topology. Through multi-dimensional risk quantification, hybrid algorithm optimization, and a hierarchical correction strategy, it achieves precise, intelligent, and adaptive correction of marketing meter box archive topology maps. Its core advantage lies in the deep integration of node comprehensive risk, topological importance, and algorithm efficiency. This not only addresses the one-sided risk assessment and low correction efficiency of traditional methods, but also drives business optimization through a closed-loop data model, providing a replicable technical paradigm for the digital transformation of power marketing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 Schematic diagram of the process of the present invention; Figure 2 The figure is a flow chart of the method for correcting the original digital topological structure according to the present invention. DETAILED DESCRIPTION

[0020] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] See also Figure 1-Figure 2 The first embodiment of the present invention provides an automated deviation correction method for a marketing meter box archive topology map based on RPA, comprising the following steps: The RPA robot automatically logs into the marketing management system, GIS system, and design drawing library, extracts the topology map and associated data of the meter box archive, extracts text information from the topology map through OCR text recognition, and uses computer vision to identify node locations, connection line directions, and equipment icon types to build a digital topology structure. Obtain historical flow data of each node to identify meter box nodes, user nodes and transformer nodes, and set the level of each node according to historical flow data and node historical data; Calculate the DTW dynamic time warping distance between nodes, identify spatiotemporal correlations, and perform feature extraction on node data for nodes of different spatiotemporal correlations and levels; Based on the extracted feature data, a fault classification model is trained for nodes with different spatiotemporal correlations, and the probability of node fault type is output; Based on the probability analysis results of each node's fault type, node hierarchical data, and node degree centrality, the hierarchical coefficient is introduced to calculate the comprehensive risk of the node. Through the improved GAPSO algorithm, which integrates the crossover and mutation mechanism of the genetic algorithm (GA) and the global search capability of the particle swarm optimization (PSO), the fitness function and iteration termination conditions are set, and after iterative calculation, the corrected topological structure of the nodes at each level is obtained. The original digital topological structure is corrected according to the corrected topological structure of the nodes.

[0022] Specifically, in this embodiment, RPA robots automatically log into the marketing management system, GIS system, and design drawing library to extract the topology map and associated data of meter box archives. Operating 24 / 7, RPA robots can complete cross-system data extraction, topology analysis, and error correction tasks without human intervention, completely eliminating the time constraints of manual operations. Meter box archive verification, which traditionally takes hours, can be completed in minutes by RPA, increasing efficiency by dozens of times. The robots can simultaneously log into multiple platforms, including the marketing management system, GIS system, and design drawing library, extracting data and building topology structures in parallel, eliminating the inefficient manual switching between systems.

[0023] By automating repetitive manual tasks, companies can reduce their reliance on dedicated personnel and reallocate human resources to higher-value tasks. Manual operations are prone to data errors or topological deviations due to fatigue and negligence. RPA, however, executes according to pre-set rules, resulting in a near-zero error rate and avoiding the costs of rework and misjudgment caused by inaccurate data.

[0024] RPA generates detailed logs for each operation, including login time, data extraction content, and correction results, enabling full process traceability. The robot can also incorporate a built-in compliance rule library to verify data compliance in real time during the correction process, ensuring that records adhere to industry standards and regulatory requirements.

[0025] OCR text recognition accurately extracts text information from topology diagrams (such as meter box numbers and equipment parameters), eliminating the error-prone nature of manual entry. A digital topology is constructed through intelligent recognition of node locations, connection line directions, and equipment icon types.

[0026] By obtaining the historical flow data of each node, the meter box node, user node and transformer node are identified, and the node hierarchy is set according to the historical flow data and node historical data; the DTW dynamic time warping distance between nodes is calculated to identify spatiotemporal correlations, and nodes with different spatiotemporal correlations and hierarchies are targeted; Traditional topology correction relies solely on static information such as node location and connectivity. This method, however, incorporates historical traffic data to upgrade nodes from "spatial symbols" to "spatiotemporal entities," revealing their operational patterns. For example, transformer node traffic typically exhibits periodic fluctuations, while user node traffic may experience nighttime lows due to lifestyle factors. These characteristics can help further distinguish node types.

[0027] By calculating the DTW distance between nodes and generating a correlation matrix, the dynamic associations hidden in the topological graph are revealed, which helps to identify the spatiotemporal correlations of nodes.

[0028] For nodes with different spatiotemporal correlations, fault classification models for different node types are trained. The input features include flow characteristics, hierarchical coefficients, etc. The model outputs the probability of node fault type to improve the accuracy of fault diagnosis.

[0029] By integrating historical traffic data, DTW spatiotemporal correlation analysis, and hierarchical fault models, the system has achieved a leap from "static topology correction" to "dynamic fault prediction." Its core advantage lies in its data-driven node cognitive upgrades and algorithmic revelation of hidden connections, providing an intelligent, forward-looking solution for power marketing meter box management.

[0030] According to the probability analysis results of the failure type of each node, the comprehensive risk of the node is calculated. Through the improved GAPSO algorithm, the fitness function and iteration termination conditions are set. After iterative calculation, the topological structure of the nodes at each level after correction is obtained. The original digital topological structure is corrected according to the topological structure of the nodes after correction.

[0031] Reflecting the health status of the node itself, the output is a classification model trained by historical traffic data; high-level node failures have a wide impact range and need to be given higher weights; considering the node degree centrality, failures of nodes with a large number of connected edges (such as hub meter boxes) are prone to trigger chain reactions, and their risks need to be considered in a superimposed manner; through the improved GAPSO algorithm, the fitness function and iteration termination conditions are set, and after iterative calculation, the topological structure of nodes at each level is obtained after correction. The original digital topological structure is corrected according to the topological structure of the nodes after correction. Through multi-dimensional risk quantification, hybrid algorithm optimization, and hierarchical correction strategies, the correction of the marketing meter box archive topology map is achieved in a precise, intelligent, and adaptive manner. Its core advantage lies in the deep integration of node comprehensive risk, topological importance, and algorithm efficiency. It not only solves the problems of one-sided risk assessment and low correction efficiency in traditional methods, but also drives business optimization through data closed loops, providing a replicable technical paradigm for the digital transformation of power marketing.

[0032] In an optional embodiment, obtaining historical flow data of each node to identify a meter box node, a user node, and a transformer node, and obtaining meter box data, business system data, and external environment data of the node include the following steps: Obtain historical flow data of meter box nodes, user nodes, and transformer nodes, label the corresponding node types of the flow data, and input it into the LSTM model to train the model to identify node types; Obtain historical traffic data for each node and input it into the trained LSTM model to identify the node type.

[0033] In an optional embodiment, setting each node level according to historical traffic data and node historical data includes the following steps: Based on the node type data, user nodes are divided into access layer nodes; Based on historical traffic data and node historical data, select nodes that meet the following conditions: the average daily traffic is greater than the 80th percentile of all nodes; the traffic standard deviation is less than the 20th percentile of all nodes; the number of directly connected nodes is greater than 10; The selected nodes are used as core layer nodes; The remaining unclassified nodes are divided into aggregation layer nodes.

[0034] In an optional embodiment, calculating the DTW dynamic time warping distance between nodes and identifying spatiotemporal correlations includes the following steps: Obtain the historical traffic data of each node, convert the node's historical traffic data into equal time series X and Y, with lengths of |X| and |Y| respectively, and ensure that the timestamps of all time series are aligned; Calculate the Euclidean distance D(i,j) between any two nodes to form an m×n distance matrix, where m and n are the sequence lengths; Create a (m+1)×(n+1) cumulative distance matrix D, D(0,0)=0, and set the first row and first column to infinity; All nodes are grouped into a node set, and the Euclidean distance between any two nodes is calculated according to the recursive formula D(i,j) = d(i,j) + min(D(i-1,j), D(i,j-1),D(i-1,j-1)), where d(i,j) is the traffic difference between node i and node j within the time series range, D(i-1,j) is the Euclidean distance between node i-1 and node j, D(i,j-1) is the Euclidean distance between node i and node j-1, and D(i-1,j-1) is the Euclidean distance between node i-1 and node j-1. Use DTW distance as edge weight to build a spatial relationship matrix between nodes; ‌Calculate the DTW distance at different time granularities of day, week, and month respectively to obtain the multi-scale DTW distance feature; The multi-scale DTW distance matrix is ​​fused through weighted summation or attention mechanism to form a spatiotemporal correlation DTW distance matrix; Use the spatiotemporal correlation DTW distance matrix to build a hierarchical tree and obtain clustering results through pruning; Directly use DTW distance as the similarity metric to divide the nodes into K clusters; Treat nodes as graph vertices, use the inverse of the DTW distance as the edge weight, and use the Louvain or Label Propagation algorithm to cluster nodes; The nodes that are clustered into a unified cluster are regarded as spatiotemporal correlation nodes.

[0035] In an optional embodiment, for nodes of different spatiotemporal correlations and levels, feature extraction is performed on node data, including the following steps: According to the historical traffic data of each node in the spatiotemporal correlation node, a time-shift sliding window W is introduced, and the feature data within the window W is calculated, including: mean, variance, slope and frequency domain features.

[0036] Based on the node hierarchy, the transmission delay time of the upstream node connected to the node and the load of the downstream node connected to the node are counted; Calculate the ratio of the transmission delay time of the upstream node to the average transmission delay time in the historical data of the entire meter box archive topology map to form the upstream node propagation delay coefficient matrix; The ratio of the downstream node load connected to the calculation node to the node load average in the historical data of the entire meter box archive topology diagram is used to form the downstream node load influence coefficient matrix; The obtained coefficient matrix and characteristic data are used as the characteristic data of the node.

[0037] In an optional embodiment, training a fault classification model for nodes with different spatiotemporal correlations based on the extracted feature data and outputting node fault type probabilities includes the following steps: The feature data of the node are fused at different time granularities, such as day, week, and month. The attention mechanism is used to dynamically assign weights to obtain the fused feature data. The feature matrix of the node is obtained: (Nn, Twindows, G), where Nn is the node number, Twindows is the feature sequence, and G corresponds to the fused feature data. Extract the connection relationship topology diagram of nodes with different spatiotemporal correlations, number the nodes with different spatiotemporal correlations, and obtain the adjacency matrix of the connection relationship of each node based on the node number information adjacent to the nodes with different spatiotemporal correlations; Build a GCN-LSTM fault classification model: Establish a GCN layer with the input being the node feature matrix + adjacency matrix; Aggregate the node feature matrix and adjacency matrix information through the GCN layer and output the node's spatiotemporal feature matrix; Connect the output of the GCN layer to the input of the LSTM layer, and input the spatiotemporal feature matrix output by the GCN into the LSTM layer; The LSTM layer outputs time coding features, which are input into the fully connected layer + Softmax layer to output the fault type probability.

[0038] In an optional embodiment, based on the probability analysis results of each node failure type, node hierarchy data, and node degree centrality, and by introducing a hierarchy coefficient, the node comprehensive risk is calculated, including the following steps: Obtain the probability analysis results of each node's fault type, node-level data, and node degree centrality to quantify node risk; Calculate node risk value ,where ,Ri is the comprehensive failure probability, for example, hardware failure, software vulnerability, configuration error, etc., which is obtained through historical data statistics or expert scoring, such as Ri = 0.8 for the core router; Li is the layer coefficient, the core layer coefficient = 3, the aggregation layer coefficient = 2, and the access layer coefficient = 1; Di′ is the normalized degree centrality = number of node connections / maximum possible number of connections, β is the weight controlling the influence of centrality, in this example, β=0.2; Calculate the link risk value Sij from node i to connected node j: , where Delayij is the delay time from node i to connected node j, Max_Delay is the maximum value of the transmission delay time in the historical data of the entire meter box archive topology map, Lossij is the energy loss of power transmission from node i to connected node j, Max_Loss is the maximum value of the energy loss of power transmission between nodes in the historical data of the entire meter box archive topology map, wloss and wdelay are the corresponding weights, wloss=0.3, wdelay=0.7; The average of the obtained link risk values ​​of the node and the connected nodes is then weighted averaged with the node risk value to obtain the node comprehensive risk value.

[0039] In an optional embodiment, an improved GAPSO algorithm is used to integrate the crossover and mutation mechanism of the genetic algorithm (GA) with the global search capability of the particle swarm optimization (PSO), set the fitness function and iteration termination condition, and perform iterative calculations to obtain the topological structure of the nodes at each level after correction, including the following steps: The improved GAPSO algorithm is used to encode particles. Each particle represents a topological structure, and its dimension is the number of node connections. For example, if the core layer has 4 nodes and each node is connected to 3 neighboring nodes, the particle dimension is 4×3=12. Set the fitness function: , where Fitness is the fitness value corresponding to the corresponding topological structure; Z_Risk is the mean of the comprehensive risk values ​​of each node under the corresponding topological structure; Connectivity is the reciprocal of the mean number of connected components of each node under the corresponding topological structure, which is 1 when fully connected; Set dynamic neighbor rules: core nodes are forced to retain three redundant paths; The node degree limit of the aggregation layer is 8. If it exceeds 8, high-risk links will be randomly deleted; Set the termination condition: 100 iterations or the fitness value changes by <2% for 10 consecutive times; The particles with the highest fitness values ​​are selected and output the optimal topology as the topology structure after correction.

[0040] In this embodiment, the deviation correction step is performed: Initialize the particle swarm and randomly generate 50 particles, each of which represents an initial topology.

[0041] The initial number of connections for core layer nodes is set to 3, and for aggregation layer nodes is set to 5.

[0042] Calculate fitness: For each particle, calculate Z_Risk and Connectivity; Set the fitness function: , where Fitness is the fitness value corresponding to the corresponding topological structure; Z_Risk is the mean of the comprehensive risk values ​​of each node under the corresponding topological structure; Connectivity is the reciprocal of the mean number of connected components of each node under the corresponding topological structure; Set dynamic neighbor rules: core nodes are forced to retain three redundant paths; The node degree limit of the aggregation layer is 8. If it exceeds 8, high-risk links will be randomly deleted; Speed ​​update: Introducing inertia weight , where t is the current iteration number, T is the maximum iteration number, and T=100.

[0043] Example: At the 20th iteration, ω=0.8, maintaining a large search range.

[0044] Set the termination condition: 100 iterations or the fitness value changes by <2% for 10 consecutive times; Output optimal topology: The particle with the highest fitness is selected as the topology after correction, ensuring that there are three redundant paths between core layer nodes and the degree of aggregation layer nodes is ≤8.

[0045] In an optional embodiment, correcting the original digitized topological structure according to the topological structure after node correction includes the following steps: A1: The node comprehensive risk value of each node in the statistically corrected topological structure, as well as the node comprehensive risk value of each node in the original digital topological structure; A2: Compare the comprehensive risk values ​​of each corresponding node and calculate the difference between the comprehensive risk values ​​of the node before and after correction. If the difference is greater than the preset threshold, modify the topological connection of the corresponding node based on the topological structure of the node after correction. A3: For nodes that are missing connections after the topology structure is modified, supplementary connections are made based on the topology structure of the nodes after correction, and the process returns to step A1 until the difference between the comprehensive risk values ​​of each node before and after correction is less than the preset threshold.

[0046] In this embodiment, the comprehensive risk values ​​of each corresponding node are compared, and a preset threshold value of the difference between the comprehensive risk values ​​of the node before and after correction is calculated. The comprehensive risk values ​​of the nodes of various topological connections are statistically analyzed based on the historical data of each node in the entire meter box archive topology diagram, and the average is calculated, and 1 / 2 of the average is used as the preset threshold value.

[0047] In another embodiment of the present invention, through an RPA-based marketing meter box file topology map automatic correction system, the system adopts the above-mentioned RPA-based marketing meter box file topology map automatic correction method to realize the marketing meter box file topology map automatic correction.

[0048] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. The RPA-based automated deviation correction method for the marketing meter box archive topology map is characterized by: The following steps are involved: The RPA robot automatically logs into the marketing management system, GIS system, and design drawing library, extracts the topology map and associated data of the meter box archive, extracts text information from the topology map through OCR text recognition, and uses computer vision to identify node locations, connection line directions, and equipment icon types to build a digital topology structure. Obtain historical flow data of each node to identify meter box nodes, user nodes and transformer nodes, and set the level of each node according to historical flow data and node historical data; Calculate the DTW dynamic time warping distance between nodes, identify spatiotemporal correlations, and perform feature extraction on node data for nodes of different spatiotemporal correlations and levels; Based on the extracted feature data, a fault classification model is trained for nodes with different spatiotemporal correlations, and the probability of node fault type is output; According to the probability analysis results of each node failure type, node hierarchical data, and node degree centrality, the hierarchical coefficient is introduced to calculate the comprehensive risk of the node. Through the improved GAPSO algorithm, the fitness function and iteration termination conditions are set. After iterative calculation, the topological structure of each level node after correction is obtained. The original digital topological structure is corrected according to the topological structure of the node after correction.

2. The RPA-based automated deviation correction method for the marketing meter box archive topology map according to claim 1 is characterized in that: Obtaining historical traffic data of each node to identify meter box nodes, user nodes, and transformer nodes, and obtaining meter box data, business system data, and external environment data of the node, including the following steps: Obtain historical flow data of meter box nodes, user nodes, and transformer nodes, label the corresponding node types of the flow data, and input it into the LSTM model to train the model to identify node types; Obtain historical traffic data for each node and input it into the trained LSTM model to identify the node type.

3. The RPA-based automated correction method for the marketing meter box archive topology map according to claim 1 is characterized in that: Setting up each node level based on historical traffic data and node historical data includes the following steps: Based on the node type data, user nodes are divided into access layer nodes; Based on historical traffic data and node historical data, select nodes that meet the following conditions: the average daily traffic is greater than the 80th percentile of all nodes; the traffic standard deviation is less than the 20th percentile of all nodes; the number of directly connected nodes is greater than 10; The selected nodes are used as core layer nodes; The remaining unclassified nodes are divided into aggregation layer nodes.

4. The RPA-based automated correction method for the marketing meter box archive topology map according to claim 1 is characterized in that: Calculating the DTW dynamic time warping distance between nodes and identifying spatiotemporal correlations involves the following steps: Obtain the historical traffic data of each node, convert the node's historical traffic data into equal time series X and Y, with lengths of |X| and |Y| respectively, and ensure that the timestamps of all time series are aligned; Calculate the Euclidean distance D(i,j) between any two nodes to form an m×n distance matrix, where m and n are the sequence lengths; Create a (m+1)×(n+1) cumulative distance matrix D, D(0,0)=0, and set the first row and first column to infinity; All nodes are grouped into a node set, and the Euclidean distance between any two nodes is calculated according to the recursive formula D(i,j) = d(i,j) + min(D(i-1,j), D(i,j-1),D(i-1,j-1)), where d(i,j) is the traffic difference between node i and node j within the time series range, D(i-1,j) is the Euclidean distance between node i-1 and node j, D(i,j-1) is the Euclidean distance between node i and node j-1, and D(i-1,j-1) is the Euclidean distance between node i-1 and node j-1. Use DTW distance as edge weight to build a spatial relationship matrix between nodes; ‌Calculate the DTW distance at different time granularities of day, week, and month respectively to obtain the multi-scale DTW distance feature; The multi-scale DTW distance matrix is ​​fused through weighted summation or attention mechanism to form a spatiotemporal correlation DTW distance matrix; Use the spatiotemporal correlation DTW distance matrix to build a hierarchical tree and obtain clustering results through pruning; Directly use DTW distance as the similarity metric to divide the nodes into K clusters; Nodes are regarded as graph vertices, the inverse of DTW distance is used as edge weight, and the Louvain algorithm is used for node clustering; The nodes that are clustered into a unified cluster are regarded as spatiotemporal correlation nodes.

5. The RPA-based automated deviation correction method for marketing meter box archive topology maps according to claim 1 is characterized in that: For nodes with different spatiotemporal correlations and levels, feature extraction is performed on node data, including the following steps: According to the historical traffic data of each node in the spatiotemporal correlation node, a time-shift sliding window W is introduced to calculate the feature data of the data in the window W, including: mean, variance, slope and frequency domain features; Based on the node hierarchy, the transmission delay time of the upstream node connected to the node and the load of the downstream node connected to the node are counted; Calculate the ratio of the transmission delay time of the upstream node to the average transmission delay time in the historical data of the entire meter box archive topology map to form the upstream node propagation delay coefficient matrix; The ratio of the downstream node load connected to the calculation node to the node load average in the historical data of the entire meter box archive topology diagram is used to form the downstream node load influence coefficient matrix; The obtained coefficient matrix and characteristic data are used as the characteristic data of the node.

6. The RPA-based automated deviation correction method for marketing meter box archive topology maps according to claim 1 is characterized in that: Based on the extracted feature data, a fault classification model is trained for nodes with different spatiotemporal correlations to output the node fault type probability, including the following steps: The feature data of the node are fused at different time granularities, such as day, week, and month. The attention mechanism is used to dynamically assign weights to obtain the fused feature data. The feature matrix of the node is obtained: (Nn, Twindows, G), where Nn is the node number, Twindows is the feature sequence, and G corresponds to the fused feature data. Extract the connection relationship topology diagram of nodes with different spatiotemporal correlations, number the nodes with different spatiotemporal correlations, and obtain the adjacency matrix of the connection relationship of each node based on the node number information adjacent to the nodes with different spatiotemporal correlations; Build a GCN-LSTM fault classification model: Establish a GCN layer with the input being the node feature matrix + adjacency matrix; Aggregate the node feature matrix and adjacency matrix information through the GCN layer and output the node's spatiotemporal feature matrix; Connect the output of the GCN layer to the input of the LSTM layer, and input the spatiotemporal feature matrix output by the GCN into the LSTM layer; The LSTM layer outputs time coding features, which are input into the fully connected layer + Softmax layer to output the fault type probability.

7. The RPA-based automated deviation correction method for marketing meter box archive topology maps according to claim 1 is characterized in that: Based on the probability analysis results of each node failure type, node hierarchy data, and node degree centrality, and by introducing the hierarchy coefficient, the node comprehensive risk is calculated, including the following steps: Obtain the probability analysis results of each node's fault type, node-level data, and node degree centrality to quantify node risk; Calculate node risk value , where,Ri is the comprehensive failure probability; Li is the layer coefficient, the core layer coefficient = 3, the aggregation layer coefficient = 2, and the access layer coefficient = 1; Di′ is the normalized degree centrality = number of node connections / maximum possible number of connections, and β is the weight controlling the influence of centrality; Calculate the link risk value Sij from node i to connected node j: , where Delayij is the delay time from node i to connected node j, Max_Delay is the maximum value of the transmission delay time in the historical data of the entire meter box archive topology map, Lossij is the energy loss of power transmission from node i to connected node j, Max_Loss is the maximum value of the energy loss of power transmission between nodes in the historical data of the entire meter box archive topology map, wloss and wdelay are the corresponding weights, wloss=0.3, wdelay=0.7; The average of the obtained link risk values ​​of the node and the connected nodes is then weighted averaged with the node risk value to obtain the node comprehensive risk value.

8. The RPA-based automated deviation correction method for marketing meter box archive topology diagrams according to claim 1 is characterized in that: The improved GAPSO algorithm is used to set the fitness function and iteration termination condition. After iterative calculation, the topological structure of nodes at each level after correction is obtained, including the following steps: The improved GAPSO algorithm is used to perform particle encoding, where each particle represents a topological structure and its dimension is the number of node connections. Set the fitness function: , where Fitness is the fitness value corresponding to the corresponding topological structure; Z_Risk is the mean of the comprehensive risk values ​​of each node under the corresponding topological structure; Connectivity is the reciprocal of the mean number of connected components of each node under the corresponding topological structure; Set dynamic neighbor rules: core nodes are forced to retain three redundant paths; The node degree limit of the aggregation layer is 8. If it exceeds 8, high-risk links will be randomly deleted; Set the termination condition: 100 iterations or the fitness value changes by <2% for 10 consecutive times; The particles with the highest fitness values ​​are selected and output the optimal topology as the topology structure after correction.

9. The RPA-based automated correction method for the marketing meter box archive topology map according to claim 1 is characterized in that: Correcting the original digital topology structure according to the topology structure after node correction includes the following steps: A1: The node comprehensive risk value of each node in the statistically corrected topological structure, as well as the node comprehensive risk value of each node in the original digital topological structure; A2: Compare the comprehensive risk values ​​of each corresponding node and calculate the difference between the comprehensive risk values ​​of the node before and after correction. If the difference is greater than the preset threshold, modify the topological connection of the corresponding node based on the topological structure of the node after correction. A3: For nodes that are missing connections after the topology structure is modified, supplementary connections are made based on the topology structure of the nodes after correction, and the process returns to step A1 until the difference between the comprehensive risk values ​​of each node before and after correction is less than the preset threshold.

10. The RPA-based marketing meter box file topology map automatic correction system is characterized by: The system adopts the RPA-based marketing meter box file topology map automatic correction method as described in any one of claims 1 to 9 to realize the automatic correction of the marketing meter box file topology map.

Citation Information

Patent Citations

  • Automatic error correction method and device for network equipment topological graph, equipment and storage medium

    CN114140557A

  • Power grid topology identification method and system, terminal and storage medium

    CN119382125A

  • Error correction method and system for topological structure of low-voltage transformer area

    CN119622349A

  • Power distribution network line transfer identification method, system and device based on CatBoost

    CN119646678A

  • Designing a network

    WO2009040385A1