Distributed data consistency checking method and system based on graph neural network
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2026-03-23
- Publication Date
- 2026-07-03
AI Technical Summary
Traditional distributed data consistency verification methods cannot capture the topological relationships and temporal fluctuation characteristics between nodes, resulting in low accuracy of anomaly diagnosis, inaccurate root cause tracing, and a lack of targeted repair strategies, which affects the stable operation of the power system and accurate decision-making.
Data probes are deployed at all nodes in the power grid meter reading data flow to construct a global data status map sequence. Graph neural networks are used to identify time-series fluctuation patterns, perform anomaly detection and root cause tracing, and generate consistency repair strategies.
It enables accurate detection and root cause quantification of anomalies across the entire power grid meter reading data chain, outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, ensures data consistency and reliability, and supports power dispatching, electricity billing and power supply quality assessment.
Smart Images

Figure CN121880464B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data verification, and in particular to a distributed data consistency verification method and system based on graph neural networks. Background Technology
[0002] As core data for power dispatching, electricity billing, and power supply quality assessment, power grid meter reading data involves multiple distributed nodes in its flow process. During transmission, storage, and synchronization, the data is susceptible to factors such as communication link interruptions, synchronization task delays, and cache not being refreshed, leading to data inconsistencies between distributed nodes, which in turn affects the stable operation of the power system and accurate decision-making.
[0003] However, traditional distributed data consistency verification methods mostly adopt a centralized verification mode, which only focuses on the data integrity of a single node and cannot capture the topological relationships and temporal fluctuation characteristics between nodes. This results in problems such as low accuracy of anomaly diagnosis, inaccurate root cause tracing, and lack of targeted repair strategies. Summary of the Invention
[0004] This invention addresses the technical problems of low accuracy in anomaly diagnosis, inaccurate root cause tracing, and lack of targeted repair strategies in existing distributed data consistency verification technologies. It provides a distributed data consistency verification method and system based on graph neural networks.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] In a first aspect, the present invention provides a distributed data consistency verification method based on graph neural networks, comprising:
[0007] Data probes are deployed at all nodes in the power grid meter reading data flow to synchronously collect and acquire the power data status distribution sequence within historical time periods, and to construct a global data status map sequence.
[0008] Identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence;
[0009] Based on the graph fluctuation deviation sequence, an anomaly diagnostic tool built on graph neural network is used to perform anomaly detection and root cause tracing on the global data state graph sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating a consistency repair strategy to perform consistency closed-loop governance on power grid meter reading data.
[0010] Secondly, the present invention provides a distributed data consistency verification system based on graph neural networks, comprising:
[0011] The data acquisition and graph construction module is used to deploy data probes at all nodes in the power grid meter reading data flow, synchronously collect the power data status distribution sequence within historical time periods, and construct a global data status graph sequence.
[0012] The fluctuation identification module is used to identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence;
[0013] The consistency verification module is used to perform anomaly detection and root cause tracing on the global data state graph sequence based on the graph fluctuation deviation sequence and a data anomaly diagnostic tool built on graph neural network. It outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generates a consistency repair strategy to perform closed-loop consistency management of power grid meter reading data.
[0014] The beneficial effects of this invention are:
[0015] Compared to existing technologies, this application first deploys data probes at all nodes in the power grid meter reading data flow chain to synchronously collect and acquire power data status distribution sequences over historical periods, and constructs a global data status map sequence. This transforms discrete data into a structured global data status map sequence containing node topological relationships and temporal characteristics, providing comprehensive and accurate basic data support for subsequent time-series fluctuation analysis, anomaly detection, and root cause tracing. Secondly, it identifies the time-series fluctuation patterns of the global data status map sequence and obtains the map fluctuation deviation sequence, achieving a precise characterization of the degree of data status fluctuation. This provides a reliable quantitative basis for subsequently distinguishing between normal fluctuations and abnormal deviations and triggering targeted anomaly diagnosis. Finally, based on the graph fluctuation deviation sequence, a data anomaly diagnostic tool built on graph neural network is used to perform anomaly detection and root cause tracing on the global data state graph sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating a consistency repair strategy to perform closed-loop governance of power grid meter reading data. This achieves accurate detection and root cause quantification of anomalies across the entire power grid meter reading data based on quantitative fluctuation criteria, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating targeted consistency repair strategies to form a closed-loop governance, effectively ensuring the consistency and reliability of power grid meter reading data.
[0016] Through the above technical solutions, this application effectively solves the technical problems of low accuracy in anomaly diagnosis, vague root cause tracing, and lack of targeted repair strategies in traditional methods, ensuring the consistency, reliability, and timeliness of power grid meter reading data among distributed nodes, and providing high-quality data support for power dispatching, electricity billing, and power supply quality assessment. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the distributed data consistency verification method based on graph neural networks provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the structure of the distributed data consistency verification system based on graph neural networks provided by the present invention.
[0019] In the attached diagram, the components represented by each number are as follows:
[0020] Data acquisition and map construction module 11, fluctuation identification module 12, consistency verification module 13. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0023] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0024] Example 1, as Figure 1 As shown, this embodiment of the invention provides a distributed data consistency verification method based on graph neural networks, including:
[0025] S10: Deploy data probes at all nodes in the power grid meter reading data flow to synchronously collect and obtain the power data status distribution sequence within historical time periods, and construct a global data status map sequence.
[0026] The flow of power grid meter reading data involves multiple distributed nodes such as terminal metering units and edge acquisition devices. During the transmission, storage and synchronization process, the data is easily affected by factors such as communication interruption and synchronization delay, resulting in inconsistencies. Traditional verification methods mostly focus on the data integrity of a single node and lack the ability to capture the topological association and temporal fluctuation characteristics of nodes across the entire link, making it difficult to achieve comprehensive perception and structured modeling of data status.
[0027] To address the aforementioned issues, this application deploys data probes at all nodes in the power grid meter reading data flow chain to synchronously collect and acquire the power data status distribution sequence within historical time periods, and constructs a global data status map sequence.
[0028] Specifically, step S10 in the method includes:
[0029] Data probes are deployed at all nodes in the power grid meter reading data flow in the target area to build a real-time data status monitoring network. The all-link nodes include at least terminal metering units, edge acquisition devices, communication aggregation nodes, data acquisition servers, business databases, and cache service nodes.
[0030] Using the real-time data status monitoring network, the power data status distribution sequence of the target area in historical time periods is synchronously collected according to the preset monitoring frequency and preset monitoring indicators. The preset monitoring indicators include at least data version identifier, timestamp sequence, node topology relationship, cache effective status and communication link quality.
[0031] In this embodiment, data probes are first deployed at all nodes of the power grid meter reading data flow in the target area to construct a real-time data status monitoring network. These all nodes include at least terminal metering units, edge acquisition devices, communication aggregation nodes, data acquisition servers, business databases, and cache service nodes. Terminal metering units are hardware devices installed on the user side or power supply side to collect energy consumption data and voltage / current parameters, such as smart meters. Edge acquisition devices are devices deployed at the edge of the power grid, responsible for aggregating data from terminal metering units and performing preliminary preprocessing, such as data format conversion and outlier filtering. Communication aggregation nodes are communication relay devices used to realize data transmission between edge acquisition devices and data acquisition servers, such as base stations and gateways. Data acquisition servers are servers that centrally receive and store data from all nodes, possessing data processing and distribution functions. Business databases are database systems used to store power business-related data, such as MySQL and Oracle. Cache service nodes are node devices used to temporarily store frequently accessed data and improve data query efficiency, such as Redis cache servers.
[0032] For example, if the target area is the core area of a city, the full-link nodes in the target area may include: smart meters in each residential community, community power distribution rooms, 5G base stations, regional power business hall computer rooms, etc. By deploying data probes at the full-link nodes, a real-time data status monitoring network covering the full-link nodes of the power grid meter reading data flow in the target area can be constructed.
[0033] Secondly, a real-time data status monitoring network is used to synchronously collect the power data status distribution sequence of the target area within a historical period according to a preset monitoring frequency and preset monitoring indicators. Among them, the preset monitoring indicators include at least the data version identifier, timestamp sequence, node topology relationship, cache effectiveness status, and communication link quality: the data version identifier is a unique identifier used to distinguish different data update versions; the timestamp sequence refers to the time record of data in each stage of collection, transmission, and storage, which can be used to trace the data flow trajectory; the node topology relationship refers to the physical connection and data transmission relationship between nodes in the entire link; the cache effectiveness status refers to the effectiveness status of the data stored in the cache service node, which can be used to determine whether the cached data is available; and the communication link quality refers to the performance indicators of the data transmission link between nodes, which can be used to evaluate the stability of data transmission.
[0034] Among them, the preset monitoring frequency refers to the time interval for data probes to collect data, which can be dynamically configured according to the power grid scale and data sensitivity of the target area; the historical period can be dynamically selected according to actual needs, such as the past 7 days, 30 days, etc.; the power data status distribution sequence refers to the preset monitoring indicator data set of all nodes in the entire link at each time point, arranged in chronological order.
[0035] For example, if the preset monitoring frequency is set to 1 minute / time, and the historical period can be set to the past 30 days, if the preset monitoring index data of terminal metering unit A at a certain moment is: data version identifier V2.3, collection timestamp 2024-05-20 08:00:00, transmission timestamp 2024-05-20 08:00:01, storage timestamp 2024-05-20 08:00:02, belonging to edge acquisition device B, cache active status active, packet loss rate of communication link with edge acquisition device B is 0.1%, and delay is 5ms, then this data is a data record in the power data status distribution sequence.
[0036] Furthermore, the "construction of a global data state map sequence" includes:
[0037] Using each power data state distribution in the power data state distribution sequence as a vertex, and using the data version identifier, timestamp sequence and cache effective status collected at the same time as vertex attributes, an initial data state map sequence is constructed.
[0038] Based on the node topology, edges are established between corresponding vertices, and the communication link quality is used as the basic attribute of the edges to structurally supplement the initial data state graph sequence, generating a global data state graph sequence.
[0039] In this embodiment, each power data state distribution in the power data state distribution sequence is first used as a vertex, and the data version identifier, timestamp sequence, and cache validity status collected at the same time are used as vertex attributes to construct an initial data state graph sequence. Here, vertex attributes refer to the feature information attached to a vertex, used to describe the specific state of the vertex. Specifically, each power data state distribution in the power data state distribution sequence is used as a vertex, and the data version identifier, timestamp sequence, and cache validity status collected at that time are used as the attributes of the corresponding vertex. The graphs at each time point are arranged in chronological order to form the initial data state graph sequence.
[0040] For example, using the data from the previous example, based on the data record at 08:00:00 on 2024-05-20 in the power data status distribution sequence, the terminal metering unit A and the edge acquisition device B can be regarded as two vertices respectively. The vertex attributes are their respective data version identifier, timestamp sequence and cache effective status at that time. The initial data status map of each time is constructed in the same way and arranged in chronological order to form the initial data status map sequence.
[0041] Secondly, edges are established between corresponding vertices based on the node topology, and the communication link quality is used as the basic attribute of the edges to structurally supplement the initial data state graph sequence, generating a global data state graph sequence. Here, the basic attribute of an edge refers to the feature information attached to the edge, used to describe the communication transmission state between nodes; the global data state graph sequence is a complete temporal graph set containing vertices and vertex attributes, edges and edge attributes, which can comprehensively reflect the topological relationships and data state characteristics between nodes.
[0042] Specifically, the node topology relationship clarifies the connection relationship between each node. For example, if node X and node Y have a direct data transmission relationship, an edge is established between vertex X and vertex Y. Then, the communication link quality is used as the basic attribute of the edge and associated with the corresponding edge to realize the structured supplement of the initial data state graph sequence, and finally form a global data state graph sequence that can fully reflect the node topology relationship and data state of the entire link.
[0043] For example, using the data from the previous example, based on the node topology, terminal metering unit A and edge acquisition device B have a direct communication connection. The communication link quality indicators between them are a packet loss rate of 0.2%, a latency of 8ms, and a bandwidth of 10Mbps. Therefore, an edge is established between the vertex corresponding to terminal metering unit A and the vertex corresponding to edge acquisition device B in the initial graph, and the packet loss rate of 0.2%, the latency of 8ms, and the bandwidth of 10Mbps are used as the basic attributes of this edge. The same method is used to complete the edge construction and attribute supplementation between all nodes, generating a global data state graph sequence.
[0044] In summary, compared to existing technologies, this application deploys data probes at all nodes in the power grid meter reading data flow chain to synchronously collect and acquire the power data status distribution sequence within historical time periods, and constructs a global data status map sequence. This achieves multi-dimensional synchronous acquisition and blind-spot-free coverage of power data across all nodes in the power grid meter reading data flow chain, transforming discrete data into a structured global data status map sequence containing node topological relationships and temporal characteristics. This provides comprehensive and accurate foundational data support for subsequent time-series fluctuation analysis, anomaly detection, and root cause tracing.
[0045] S20: Identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence.
[0046] The global data state map sequence can reflect the temporal evolution characteristics of node topology and data attributes. However, traditional data consistency verification methods lack systematic analysis of such temporal characteristics and cannot capture abnormal fluctuation trends of data state through quantitative methods, making it difficult to accurately distinguish between normal data fluctuations and abnormal data deviations.
[0047] To address the aforementioned issues, this application identifies the temporal fluctuation patterns of the global data state map sequence and obtains the map fluctuation deviation sequence.
[0048] Specifically, step S20 in the method includes:
[0049] For any two adjacent global data state maps in the global data state map sequence, perform structural similarity comparison and attribute similarity comparison, and determine the similarity of multiple maps by weighting.
[0050] The graph similarity is subtracted from 1 as the graph fluctuation deviation, and multiple graph fluctuation deviations are calculated based on the multiple graph similarities;
[0051] Based on the sequential order of the graphs in the global data state graph sequence, the fluctuation deviations of the multiple graphs are sorted to generate a graph fluctuation deviation sequence.
[0052] In this embodiment, structural similarity and attribute similarity comparisons are first performed on any two adjacent global data state graphs in the global data state graph sequence, and multiple graph similarities are determined by weighting. Adjacent global data state graphs refer to two graphs that are temporally consecutive in the global data state graph sequence, such as the graphs corresponding to time t and time t+1. Structural similarity comparison compares whether the number of vertices and the connection relationships of edges are consistent between two adjacent graphs. Attribute similarity comparison compares whether the corresponding vertex attributes and corresponding edge attributes are consistent between two adjacent global data state graphs. The graph similarity value ranges from [0,1], with a similarity closer to 1 indicating a higher similarity.
[0053] For example, the calculation process of graph similarity is as follows: First, select two temporally consecutive global data state graphs from the global data state graph sequence; second, structural similarity comparison can use graph edit distance, which quantifies the minimum number of editing operations (vertex addition / deletion, edge connection changes, etc.) required to transform one global data state graph into another, accurately capturing major anomalies such as topological mutations in power grid scenarios, and then normalizing the calculation results to the [0,1] interval to obtain structural similarity; third, attribute similarity comparison can use cosine similarity or Euclidean similarity. The distance algorithm encodes the attributes of each vertex into a vertex feature vector of a uniform dimension and the attributes of each edge into an edge feature vector. Then, it calculates the average similarity of the feature vectors of all corresponding vertices and the average similarity of the feature vectors of all corresponding edges in the adjacent global data state graph. The two are summed with equal weights and normalized to the interval [0,1] to obtain the attribute similarity. Finally, the graph similarity is determined by weighting: graph similarity = a × structural similarity + b × attribute similarity, where a and b are weight coefficients, and a + b = 1, which can be dynamically set according to the actual situation and business focus.
[0054] Secondly, the map fluctuation deviation is calculated by subtracting the map similarity from 1. Multiple map fluctuation deviations are obtained based on the similarity of multiple maps. Specifically, the calculation formula for map fluctuation deviation is: Map fluctuation deviation = 1 - Map similarity. The calculation logic for map fluctuation deviation is: the higher the map similarity, the smaller the difference between adjacent maps and the smaller the fluctuation deviation; conversely, the lower the map similarity, the larger the fluctuation deviation.
[0055] For example, if the similarity between two adjacent maps is 0.92, then the corresponding map fluctuation deviation is 1 - 0.92 = 0.08.
[0056] Finally, based on the chronological order of the maps in the global data state map sequence, the fluctuation deviations of multiple maps are sorted to generate a map fluctuation deviation sequence. This map fluctuation deviation sequence is an ordered set containing all adjacent map fluctuation deviations, arranged chronologically according to the global data state map sequence, used to reflect the temporal fluctuation characteristics of the data state as a whole.
[0057] In summary, compared to existing technologies, this application identifies the temporal fluctuation patterns of the global data state map sequence and obtains a map fluctuation deviation sequence. Thus, the temporal evolution characteristics of node topology and data attributes in the global data state map sequence are transformed into a quantifiable map fluctuation deviation sequence, achieving a precise characterization of the degree of data state fluctuation. This provides a reliable quantitative basis for subsequently distinguishing between normal fluctuations and abnormal deviations, and triggering targeted anomaly diagnosis.
[0058] S30: Based on the graph fluctuation deviation sequence, use a data anomaly diagnostic tool built on graph neural network to perform anomaly detection and root cause tracing on the global data state graph sequence, output the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generate a consistency repair strategy to perform consistency closed-loop governance on power grid meter reading data.
[0059] Data inconsistency in the entire power grid meter reading data flow is easily caused by various factors. Traditional anomaly detection methods are difficult to integrate the topological correlation and temporal fluctuation characteristics of the global data status map, resulting in inaccurate anomaly location and unclear root cause tracing. Furthermore, they lack targeted repair strategies and closed-loop governance mechanisms that are compatible with the diagnostic results, and cannot fundamentally solve the data consistency problem.
[0060] To address the aforementioned issues, this application utilizes a data anomaly diagnostic tool built on a graph neural network to perform anomaly detection and root cause analysis on the global data state graph sequence based on the aforementioned graph fluctuation deviation sequence. It outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generates a consistency repair strategy to perform closed-loop consistency management of power grid meter reading data.
[0061] Specifically, step S30 in the method includes:
[0062] Based on historical power grid meter reading data monitoring records, several global data status map sequence samples were collected, and the historical inconsistent node distribution and historical anomaly diagnosis probability distribution corresponding to different global data status map sequence samples were labeled to obtain several inconsistent node distribution samples and several anomaly diagnosis probability distribution samples. The historical anomaly diagnosis probability distribution includes historical anomaly probability distributions corresponding to multiple inconsistent node historical anomaly diagnosis types. The anomaly diagnosis types include at least node data loss due to communication link interruption, data version lag due to synchronization task delay, temporary data expiration due to cache not being refreshed, and transmission data corruption due to data verification errors.
[0063] The training data consists of several global data state map sequence samples, several inconsistent node distribution samples, and several abnormal diagnosis probability distribution samples. J-fold cross-partitioning is performed to obtain J sample training sets, where J is an integer greater than 5.
[0064] The graph neural networks are trained using the J sample training sets respectively, J data anomaly diagnosis branches are constructed, and the data anomaly diagnostic tool is built by integrating them.
[0065] In this embodiment, firstly, based on historical power grid meter reading data monitoring records, several global data status map sequence samples are collected. The historical inconsistency node distribution and historical anomaly diagnosis probability distribution corresponding to different global data status map sequence samples are then labeled, resulting in several inconsistency node distribution samples and several anomaly diagnosis probability distribution samples. The historical anomaly diagnosis probability distribution includes historical anomaly probability distributions corresponding to multiple inconsistency node historical anomaly diagnosis types. Anomaly diagnosis types include at least node data synchronization failure due to communication link interruption, data version lag due to synchronization task delay, temporary data expiration due to cache not being refreshed, and transmission data corruption due to data verification errors. Specifically, during sample collection, global data status map sequences under different seasons, different power grid loads, and different anomaly types can be selected as samples to cover as many typical operating scenarios of power grid meter reading data flow as possible, ensuring the diversity and representativeness of the samples. The historical anomaly diagnosis probability distribution labeling process combines historical anomaly event records, locates inconsistent nodes in each sample through data comparison, determines the anomaly diagnosis type corresponding to each inconsistent node based on the root cause analysis results, and finally, for each inconsistent node, calculates the frequency ratio of each type of anomaly diagnosis type in the same anomaly scenario as the anomaly probability of the corresponding anomaly type.
[0066] For example, in a certain sample, the inconsistent node is terminal metering unit X. According to historical records, this node has 20 anomalies in the same power grid load scenario. Among them, there are 18 cases of data version lag due to synchronization task delay and 2 cases of node data loss due to communication link interruption. Then the probability distribution of the anomaly diagnosis of this node is: data version lag due to synchronization task delay = 18 / 20 = 90%, node data loss due to communication link interruption = 2 / 20 = 10%.
[0067] Secondly, several global data state map sequence samples, several inconsistent node distribution samples, and several anomaly diagnosis probability distribution samples are used as training data. J-fold cross-partitioning is performed to obtain J training samples, where J is an integer greater than 5. J-fold cross-partitioning refers to randomly dividing the training data into J equal parts, selecting J-1 parts as the training set and 1 part as the validation set each time, and repeating this process J times. This method is used to improve the stability and generalization ability of the model training.
[0068] Specifically, the value of J in the J-fold cross-split can be dynamically determined based on the number of training data samples. The more samples there are, the larger the value of J can be to make full use of the data information. J must also satisfy the integer constraint greater than 5. This setting can ensure the diversity of the distribution of each training set through multiple rounds of splitting, and reduce the risk of overfitting caused by a single split, thereby improving the generalization ability of the model.
[0069] For example, if the number of training data samples is 1000, J can be 10. The training data is randomly and evenly divided into 10 parts by 10-fold cross-validation, with 100 samples in each part. Each time, 9 parts are selected as the training set and the remaining part is selected as the validation set. Therefore, a total of 10 training sets and 10 validation sets are obtained. Each training set contains 900 samples and each validation set contains 100 samples.
[0070] Finally, graph neural networks are trained using J sample training sets to construct J data anomaly diagnosis branches, which are then integrated to build a data anomaly diagnostic tool. Each data anomaly diagnosis branch is a model trained on a single sample training set and possesses independent anomaly diagnosis capabilities. Integrating the J data anomaly diagnosis branches utilizes collaborative reasoning to improve diagnostic accuracy and robustness.
[0071] For example, each sample training set is used to independently train a data anomaly diagnosis branch. Taking the data anomaly diagnosis branch built based on graph attention network as an example, the specific process is as follows: 1. Model construction: The model structure of the data anomaly diagnosis branch mainly consists of an input layer, a graph attention feature extraction module, a multi-task fusion layer, and a dual output layer. Among them, the input layer is used to receive global data state graph sequence samples and transform the vertex attributes and edge attributes of each global data state graph into high-dimensional feature vectors; the graph attention feature extraction module contains three stacked graph attention layers. Each layer calculates the attention weight between nodes, adaptively aggregates the feature information of neighboring nodes, and deeply mines the coupling features of node topology association and data attributes. The first graph attention layer outputs local node features, the second layer outputs global topology features, and the third layer fuses the features of the first two layers to obtain comprehensive features; the multi-task fusion layer performs dimensional mapping and feature reorganization on the comprehensive features through a fully connected network to adapt to the dual-task learning requirements; the dual output layers respectively use the Sigmoid activation function to output the probability of each node being an inconsistent node, and use the Softmax activation function to output the probability distribution of each inconsistent node belonging to various anomaly diagnosis types.
[0072] 2. Model Training: The global data state map sequence samples in the training set are used as model inputs, and the corresponding inconsistent node distribution samples and anomaly diagnosis probability distribution samples are used as training labels. The prediction accuracy of inconsistent node distribution and the cross-entropy loss of anomaly diagnosis probability distribution are used as joint optimization objectives, with each having a weight of 0.5. The initial learning rate is set to 0.001, the maximum number of iterations is 500, and the early stopping threshold is 20. All parameters of the model are iteratively updated using the Adam optimizer and backpropagation algorithm. When the joint loss function value of the validation set does not decrease for 20 consecutive iterations or reaches the maximum number of iterations, the model is considered to have converged, and a data anomaly diagnosis branch that has been trained and has stable performance is obtained.
[0073] Furthermore, step S30 of the method also includes:
[0074] The average value of the spectral fluctuation deviation sequence is obtained by averaging the spectral fluctuation deviation sequence.
[0075] The ratio of the standard deviation of the spectral fluctuation deviation in the spectral fluctuation deviation sequence to the mean of the spectral fluctuation deviation is set as the deviation fluctuation index.
[0076] The data state fluctuation coefficient is calculated based on the mean fluctuation deviation and the deviation fluctuation index of the aforementioned spectrum.
[0077] Based on the data state fluctuation coefficient, the data anomaly diagnostic tool is invoked to perform anomaly detection and root cause tracing on the global data state map sequence, and the distribution of inconsistent nodes and anomaly diagnosis probability distribution are output.
[0078] In this embodiment, the average value of the spectral fluctuation deviation sequence is first calculated to obtain the mean spectral fluctuation deviation. The mean spectral fluctuation deviation is the arithmetic mean of all values in the spectral fluctuation deviation sequence, used to reflect the overall fluctuation level of the global data state spectral sequence.
[0079] For example, if the spectrum fluctuation deviation sequence is [0.08, 0.06, 0.09, 0.07, 0.08], then the mean spectrum fluctuation deviation = (0.08 + 0.06 + 0.09 + 0.07 + 0.08) / 5 = 0.076.
[0080] Secondly, the ratio of the standard deviation to the mean of the spectral fluctuation deviations in the spectral fluctuation deviation sequence is defined as the deviation volatility index. The standard deviation describes the dispersion of spectral fluctuation deviations in the sequence; a larger standard deviation indicates a more dispersed distribution of fluctuation deviations. The deviation volatility index is a quantitative indicator describing the dispersion of fluctuation deviations, reflecting the stability of data state fluctuations. Specifically, the standard deviation is calculated by first calculating the sum of the squares of the differences between each spectral fluctuation deviation and the mean of the spectral fluctuation deviations, then dividing by the number of fluctuation deviations and taking the square root. The formula for calculating the deviation volatility index is: Deviation Volatility Index = Standard Deviation of Spectral Fluctuation Deviations / Mean of Spectral Fluctuation Deviations. A larger deviation volatility index indicates more unstable data state fluctuations and a greater likelihood of abnormal fluctuations.
[0081] Next, the data state fluctuation coefficient is calculated based on the mean deviation of the spectral fluctuation and the deviation fluctuation index. The data state fluctuation coefficient = mean deviation of the spectral fluctuation × (1 + deviation fluctuation index). A larger data state fluctuation coefficient indicates more severe data state fluctuations, poorer stability, and a higher probability of data inconsistency. Therefore, the data state fluctuation coefficient can comprehensively reflect the intensity and stability of data state fluctuations, and can be used to determine whether anomaly diagnosis needs to be initiated.
[0082] For example, if the average deviation of the spectrum is 0.076 and the deviation fluctuation index is 0.158, then the data state fluctuation coefficient = 0.076 × (1 + 0.158) = 0.088.
[0083] Finally, based on the data state fluctuation coefficient, the data anomaly diagnostic tool is invoked to perform anomaly detection and root cause analysis on the global data state graph sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis. Specifically, the data state fluctuation coefficient is positively correlated with the anomaly diagnosis intensity; the larger the data state fluctuation coefficient, the more data anomaly diagnostic branches need to be invoked to improve diagnostic accuracy; conversely, the smaller the data state fluctuation coefficient, the fewer data anomaly diagnostic branches can be invoked to improve diagnostic efficiency. Specifically, after randomly invoking a corresponding number of data anomaly diagnostic branches in the data anomaly diagnostic tool, each data anomaly diagnostic branch receives the global data state graph sequence, locates inconsistent nodes, and outputs the probability distribution of each inconsistent node belonging to various anomaly diagnosis types.
[0084] Furthermore, the step of "calling the data anomaly diagnostic tool based on the data state fluctuation coefficient to perform anomaly detection and root cause tracing on the global data state map sequence, and outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis" includes:
[0085] The optimal number of branches to be selected is obtained by multiplying the ratio of the data state fluctuation coefficient to the historical maximum data state fluctuation coefficient recorded in the target area within a preset time range by J and rounding it down, where H is greater than or equal to 2 and less than or equal to J.
[0086] H data anomaly diagnosis branches are randomly selected from the J data anomaly diagnosis branches, and anomaly detection and root cause tracing are performed on the global data state map sequence respectively, outputting H initial inconsistent node distributions and H initial anomaly diagnosis probability distributions.
[0087] The nodes whose abnormal frequency is greater than or equal to a preset frequency threshold among the H initial inconsistent node distributions are taken as the final inconsistent nodes, and an inconsistent node distribution is generated, wherein the preset frequency threshold is less than 0.2H.
[0088] Based on the inconsistent node distribution, the mean of the H initial anomaly diagnosis probability distributions is fitted to obtain the anomaly diagnosis probability distribution.
[0089] In this embodiment, the optimal number of branches H is first obtained by multiplying the ratio of the data state fluctuation coefficient to the historical maximum data state fluctuation coefficient recorded in the target area within a preset time range by J and rounding it down. Here, H is greater than or equal to 2 and less than or equal to J. The preset time range refers to the time interval used to statistically analyze the historical maximum data state fluctuation coefficient, such as the past 3 months. Specifically, the calculation logic for the optimal number of branches H is as follows: the closer the current data state fluctuation coefficient is to the historical maximum data state fluctuation coefficient recorded within the preset time range, the higher the risk of current data anomalies, requiring more data anomaly diagnosis branches to be called for collaborative reasoning; conversely, fewer data anomaly diagnosis branches can meet the diagnostic needs, thereby saving computing resources and improving computational efficiency.
[0090] For example, during the calculation process, the ratio of the data state fluctuation coefficient to the historical maximum data state fluctuation coefficient recorded in the target area within a preset time range is first multiplied by the total number of data anomaly diagnosis branches J, and then rounded to obtain H. H must satisfy the constraint condition 2≤H≤J.
[0091] For example, if the preset time range is the past 3 months, the maximum historical data state fluctuation coefficient recorded in the past 3 months is 0.35, the current data state fluctuation coefficient is 0.088, and J=10, then H=(0.088 / 0.35)×10≈3 (rounded up), that is, the optimal number of branches to select is 3.
[0092] Secondly, H data anomaly diagnosis branches are randomly selected from the J data anomaly diagnosis branches. These branches perform anomaly detection and root cause analysis on the global data state map sequence, outputting H initial inconsistent node distributions and H initial anomaly diagnosis probability distributions. Specifically, randomly selecting H data anomaly diagnosis branches avoids diagnostic bias from a single branch and improves the objectivity of the results. In the actual anomaly detection and root cause analysis process, each selected data anomaly diagnosis branch independently performs anomaly detection and root cause analysis on the global data state map sequence, extracting features and inferring diagnosis based on its own model parameters, and outputting the corresponding initial inconsistent node distribution and initial anomaly diagnosis probability distribution.
[0093] For example, three data anomaly diagnosis branches are randomly selected to perform anomaly detection and root cause tracing on the global data state map sequence, respectively. Data anomaly diagnosis branch 1 outputs the initial inconsistent node distribution as: terminal metering unit G, edge acquisition device H; data anomaly diagnosis branch 2 outputs the initial inconsistent node distribution as: terminal metering unit G; data anomaly diagnosis branch 3 outputs the initial inconsistent node distribution as: terminal metering unit G, edge acquisition device H, communication aggregation node I. At the same time, each data anomaly diagnosis branch outputs the corresponding initial anomaly diagnosis probability distribution.
[0094] Next, nodes whose anomaly frequency in the H initial inconsistent node distributions is greater than or equal to a preset frequency threshold are selected as the final inconsistent nodes, generating an inconsistent node distribution. The preset frequency threshold is less than 0.2H. Here, anomaly frequency refers to the number of times a node appears in the H initial inconsistent node distributions; the preset frequency threshold is the critical value for determining whether a node is a final inconsistent node. Setting it to less than 0.2H ensures that only nodes identified by the majority of data anomaly diagnostic branches are included in the final result, improving accuracy.
[0095] Specifically, the frequency of each node in the distribution of H initial inconsistent nodes is first counted as the corresponding abnormal frequency. Then, it is compared with a preset frequency threshold. If the abnormal frequency is greater than or equal to the preset frequency threshold, then the node is the final inconsistent node. All the final inconsistent nodes constitute the inconsistent node distribution.
[0096] For example, if H=3, the preset frequency threshold = 0.15×3=0.45, if the terminal metering unit G appears 3 times in the 3 initial distributions, the edge acquisition device H appears 2 times, and the communication aggregation node I appears 1 time, then the final inconsistent node distribution is: terminal metering unit G, edge acquisition device H, and communication aggregation node I.
[0097] Finally, based on the inconsistent node distribution, the mean of the H initial anomaly diagnosis probability distributions is fitted to obtain the final anomaly diagnosis probability distribution. Specifically, for each node in the final inconsistent node distribution, the anomaly type probability corresponding to that node is extracted from the H initial anomaly diagnosis probability distributions, and the average value of each anomaly type is calculated as the final anomaly diagnosis probability for that node. The anomaly diagnosis probability distributions of all inconsistent nodes constitute the final anomaly diagnosis probability distribution.
[0098] For example, for terminal metering unit G, the node data out-of-step probability output by data anomaly diagnosis branch 1 is 0.85, and the data version lag is 0.10; the node data out-of-step probability output by data anomaly diagnosis branch 2 is 0.80, and the data version lag is 0.15; and the node data out-of-step probability output by data anomaly diagnosis branch 3 is 0.82, and the data version lag is 0.13. Therefore, after mean fitting, the node data out-of-step probability = (0.85 + 0.80 + 0.82) / 3 ≈ 0.823, and the data version lag probability = (0.10 + 0.15 + 0.13) / 3 ≈ 0.127. Similarly, the probabilities of other anomaly types can be calculated using the same method to obtain the anomaly diagnosis probability distribution of this node.
[0099] Furthermore, the "generation of a consistency repair strategy" includes:
[0100] Configure a consistency repair mechanism based on the aforementioned data state fluctuation coefficient;
[0101] Establish a root cause-action mapping knowledge base, generate a consistency repair strategy based on the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and execute the consistency repair strategy according to the consistency repair mechanism.
[0102] In this embodiment, a consistency repair mechanism is first configured based on the data state fluctuation coefficient. This consistency repair mechanism refers to the specific repair range determined by the data state fluctuation coefficient, used to achieve differentiated and efficient repair. Specifically, the smaller the data state fluctuation coefficient, the more concentrated the data anomaly range, and a precise repair mechanism, such as targeted repair, is prioritized; conversely, the larger the data state fluctuation coefficient, the wider the potential anomaly range, requiring a large-scale repair mechanism, such as local or global repair, to ensure the comprehensiveness and efficiency of the repair.
[0103] Secondly, a root cause-action mapping knowledge base is established. Consistency repair strategies are generated based on the distribution of inconsistent nodes and the probability distribution of anomaly diagnoses, and these strategies are executed according to the consistency repair mechanism. The root cause-action mapping knowledge base is a structured database storing anomaly diagnosis types (root causes) and corresponding repair actions. Each root cause corresponds to at least one specific repair operation, ensuring the relevance of the repair strategies. The construction of the root cause-action mapping knowledge base can incorporate historical repair cases and expert experience. For example, the repair actions for node data out of sync might include triggering a communication link reconnection command, synchronizing the latest data from adjacent normal nodes, and verifying the integrity of the synchronized data; the repair actions for temporary data expiration might include sending a forced cache refresh command, verifying the consistency between the refreshed data and the business database, and recording the repair results.
[0104] For example, the consistency repair strategy is executed according to the consistency repair mechanism, and the specific process is as follows: 1. Target aggregation and root cause classification: Using the distribution of inconsistent nodes as the initial target set, the target nodes are grouped and aggregated according to the repair scope defined by the currently effective repair mechanism, such as aggregation by preset power supply area under local repair. For each aggregation unit, the anomaly diagnosis probability distribution of all nodes within it is analyzed, and the dominant root cause of the unit is determined through weighted voting (weight is the anomaly probability of each node) or threshold screening mechanism.
[0105] 2. Strategy Matching and Task Generation: The root cause-action mapping knowledge base pre-stores the association between various abnormal root causes and corresponding repair action templates, such as the routing switching process corresponding to the root cause of communication link interruption. For each aggregation unit and its dominant root cause, a suitable repair action template is matched from the root cause-action mapping knowledge base, and the template is instantiated into a specific parameterized repair task.
[0106] 3. Task orchestration and output: Taking into account the data status fluctuation coefficient, node business priority and inter-task dependencies, all generated repair tasks are prioritized and their execution sequences are orchestrated, and a structured repair strategy execution plan is finally output.
[0107] Furthermore, the "configuration of a consistency repair mechanism based on the data state fluctuation coefficient" includes:
[0108] Configure a first data state fluctuation coefficient scalar and a second data state fluctuation coefficient scalar, wherein the first data state fluctuation coefficient scalar is smaller than the second data state fluctuation coefficient scalar;
[0109] If the data state fluctuation coefficient is less than the first data state fluctuation coefficient scalar, the consistency repair mechanism is configured as fixed-point repair, wherein fixed-point repair only repairs inconsistent nodes;
[0110] If the data state fluctuation coefficient is greater than or equal to the first data state fluctuation coefficient scalar and less than the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as local repair, wherein local repair is to repair all nodes within the preset power supply area where the inconsistent node is located;
[0111] If the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as global repair, wherein global repair is to repair all nodes in the target area.
[0112] In this embodiment, a first data state fluctuation coefficient scalar and a second data state fluctuation coefficient scalar are configured, wherein the first data state fluctuation coefficient scalar is smaller than the second data state fluctuation coefficient scalar. The configuration of the first and second data state fluctuation coefficient scalars is based on statistical analysis of historical operational data from the target area's power grid meter readings. The first data state fluctuation coefficient scalar can be taken as the 95th percentile of the data state fluctuation coefficients under historical normal operating scenarios. This percentile means that during normal power grid operation, 95% of the fluctuation coefficients are below this value; exceeding it indicates that data fluctuations exceed the normal stable range, posing a risk of local anomalies. The second data state fluctuation coefficient scalar can be taken as the minimum data state fluctuation coefficient corresponding to a major data inconsistency fault event in the target area's history. This value is the critical threshold for major faults; exceeding it indicates severe data fluctuations, with a high probability of a full-link or large-scale data inconsistency problem. Simultaneously, the constraint that the first data state fluctuation coefficient scalar is smaller than the second data state fluctuation coefficient scalar must be met. This setting method ensures that the classification criteria for the repair mechanism align with the actual operating patterns of the power grid, improving the scientific nature of the consistency repair mechanism configuration and the reliability of the solution.
[0113] Secondly, if the data state fluctuation coefficient is less than the first data state fluctuation coefficient scalar, the consistency repair mechanism is configured as fixed-point repair. Fixed-point repair only repairs inconsistent nodes and does not involve other normal nodes. It is suitable for scenarios with small data anomaly range and low fluctuation degree, which can reduce the impact of repair operations on normal nodes and improve repair efficiency.
[0114] Furthermore, if the data state fluctuation coefficient is greater than or equal to the first data state fluctuation coefficient scalar and less than the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as local repair. Local repair involves repairing all nodes within the preset power supply area where the inconsistent node is located. The preset power supply area refers to the smallest power supply management unit divided according to the power grid plan. Each power supply area contains several terminal metering units, 1-2 edge acquisition devices, and corresponding communication aggregation nodes, facilitating precise management of local areas. Local repair refers to the mechanism of performing repair operations on all nodes within the preset power supply area where the inconsistent node is located, suitable for scenarios with a wide range of data anomalies and potential anomalies.
[0115] Specifically, the logic of local repair is as follows: Within the preset power supply area where the inconsistent node is located, there may be unidentified potential abnormal nodes. By repairing all nodes within the preset power supply area, data inconsistency issues within the area can be eliminated, preventing the spread of anomalies. For example, if the data state fluctuation coefficient is greater than or equal to a first data state fluctuation coefficient scalar and less than a second data state fluctuation coefficient scalar, and the preset power supply area where the inconsistent node terminal metering unit G and edge acquisition device H are located is area 5 of XX community, then a local repair mechanism is configured to repair all terminal metering units, edge acquisition devices, and communication aggregation nodes within that area.
[0116] Finally, if the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as global repair. Global repair involves repairing all nodes in the target area. Specifically, global repair refers to a mechanism that performs repair operations on all nodes in the entire power grid meter reading data flow chain within the target area. This mechanism is suitable for scenarios with a wide range of data anomalies and severe fluctuations, and can completely eliminate data inconsistency issues across the entire area.
[0117] Specifically, the global repair is initiated when the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar. This indicates that the data state within the target area is extremely unstable and there may be large-scale data inconsistencies, requiring comprehensive verification and repair of all nodes. For example, if the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar, a global repair mechanism is configured to repair all terminal metering units, edge acquisition devices, communication aggregation nodes, data acquisition servers, business databases, and cache service nodes within the target area.
[0118] In summary, compared to existing technologies, this application, based on the aforementioned graph fluctuation deviation sequence, utilizes a data anomaly diagnostic tool constructed based on a graph neural network to perform anomaly detection and root cause tracing on the global data state graph sequence. It outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generates a consistency repair strategy to perform closed-loop governance of power grid meter reading data. In this way, it can achieve accurate detection and root cause quantification of anomalies across the entire power grid meter reading data chain based on quantified fluctuation criteria, output the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generate targeted consistency repair strategies to form a closed-loop governance, effectively ensuring the consistency and reliability of power grid meter reading data.
[0119] In summary, the embodiments of this application have at least the following technical effects:
[0120] Compared to existing technologies, this application first deploys data probes at all nodes in the power grid meter reading data flow chain to synchronously collect and acquire the power data status distribution sequence within historical time periods, and constructs a global data status map sequence. This achieves multi-dimensional synchronous acquisition and blind-spot-free coverage of power data across all nodes in the power grid meter reading data flow chain, transforming discrete data into a structured global data status map sequence containing node topological relationships and temporal characteristics. This provides comprehensive and accurate basic data support for subsequent time-series fluctuation analysis, anomaly detection, and root cause tracing.
[0121] Secondly, this application identifies the temporal fluctuation pattern of the global data state map sequence and obtains the map fluctuation deviation sequence. In this way, the temporal evolution characteristics of node topology and data attributes in the global data state map sequence are transformed into a quantifiable map fluctuation deviation sequence, achieving a precise characterization of the degree of data state fluctuation. This provides a reliable quantitative basis for subsequently distinguishing between normal fluctuations and abnormal deviations, and triggering targeted anomaly diagnosis.
[0122] Finally, based on the aforementioned graph fluctuation deviation sequence, this application utilizes a data anomaly diagnostic tool constructed based on a graph neural network to perform anomaly detection and root cause tracing on the global data state graph sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating a consistency repair strategy to perform closed-loop governance of power grid meter reading data. In this way, accurate detection and root cause tracing of full-link anomalies in power grid meter reading data based on quantitative fluctuation criteria can be achieved, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating targeted consistency repair strategies to form a closed-loop governance, effectively ensuring the consistency and reliability of power grid meter reading data.
[0123] Through the above technical solutions, this application effectively solves the technical problems of low accuracy in anomaly diagnosis, vague root cause tracing, and lack of targeted repair strategies in traditional methods, ensuring the consistency, reliability, and timeliness of power grid meter reading data among distributed nodes, and providing high-quality data support for power dispatching, electricity billing, and power supply quality assessment.
[0124] Example 2, as Figure 2 As shown, based on the same inventive concept as the distributed data consistency verification method based on graph neural networks provided in Embodiment 1, this embodiment of the invention also provides a distributed data consistency verification system based on graph neural networks, comprising:
[0125] The data acquisition and graph construction module 11 is used to deploy data probes at all nodes in the power grid meter reading data flow, synchronously collect and obtain the power data status distribution sequence within historical time periods, and construct a global data status graph sequence.
[0126] The fluctuation identification module 12 is used to identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence;
[0127] The consistency verification module 13 is used to perform anomaly detection and root cause tracing on the global data state graph sequence based on the graph fluctuation deviation sequence and a data anomaly diagnostic tool built on graph neural network, output the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generate a consistency repair strategy to perform closed-loop consistency management of power grid meter reading data.
[0128] The data acquisition and map construction module 11 is specifically used for:
[0129] Data probes are deployed at all nodes in the power grid meter reading data flow in the target area to build a real-time data status monitoring network. The all-link nodes include at least terminal metering units, edge acquisition devices, communication aggregation nodes, data acquisition servers, business databases, and cache service nodes.
[0130] Using the real-time data status monitoring network, the power data status distribution sequence of the target area in historical time periods is synchronously collected according to the preset monitoring frequency and preset monitoring indicators. The preset monitoring indicators include at least data version identifier, timestamp sequence, node topology relationship, cache effective status and communication link quality.
[0131] Furthermore, the "construction of a global data state map sequence" includes:
[0132] Using each power data state distribution in the power data state distribution sequence as a vertex, and using the data version identifier, timestamp sequence and cache effective status collected at the same time as vertex attributes, an initial data state map sequence is constructed.
[0133] Based on the node topology, edges are established between corresponding vertices, and the communication link quality is used as the basic attribute of the edges to structurally supplement the initial data state graph sequence, generating a global data state graph sequence.
[0134] Specifically, the fluctuation recognition module 12 is used for:
[0135] For any two adjacent global data state maps in the global data state map sequence, perform structural similarity comparison and attribute similarity comparison, and determine the similarity of multiple maps by weighting.
[0136] The graph similarity is subtracted from 1 as the graph fluctuation deviation, and multiple graph fluctuation deviations are calculated based on the multiple graph similarities;
[0137] Based on the sequential order of the graphs in the global data state graph sequence, the fluctuation deviations of the multiple graphs are sorted to generate a graph fluctuation deviation sequence.
[0138] The consistency verification module 13 is specifically used for:
[0139] Based on historical power grid meter reading data monitoring records, several global data status map sequence samples were collected, and the historical inconsistent node distribution and historical anomaly diagnosis probability distribution corresponding to different global data status map sequence samples were labeled to obtain several inconsistent node distribution samples and several anomaly diagnosis probability distribution samples. The historical anomaly diagnosis probability distribution includes historical anomaly probability distributions corresponding to multiple inconsistent node historical anomaly diagnosis types. The anomaly diagnosis types include at least node data loss due to communication link interruption, data version lag due to synchronization task delay, temporary data expiration due to cache not being refreshed, and transmission data corruption due to data verification errors.
[0140] The training data consists of several global data state map sequence samples, several inconsistent node distribution samples, and several abnormal diagnosis probability distribution samples. J-fold cross-partitioning is performed to obtain J sample training sets, where J is an integer greater than 5.
[0141] The graph neural networks are trained using the J sample training sets respectively, J data anomaly diagnosis branches are constructed, and the data anomaly diagnostic tool is built by integrating them.
[0142] Furthermore, the consistency verification module 13 is also specifically used for:
[0143] The average value of the spectral fluctuation deviation sequence is obtained by averaging the spectral fluctuation deviation sequence.
[0144] The ratio of the standard deviation of the spectral fluctuation deviation in the spectral fluctuation deviation sequence to the mean of the spectral fluctuation deviation is set as the deviation fluctuation index.
[0145] The data state fluctuation coefficient is calculated based on the mean fluctuation deviation and the deviation fluctuation index of the aforementioned spectrum.
[0146] Based on the data state fluctuation coefficient, the data anomaly diagnostic tool is invoked to perform anomaly detection and root cause tracing on the global data state map sequence, and the distribution of inconsistent nodes and anomaly diagnosis probability distribution are output.
[0147] Specifically, the phrase "calling the data anomaly diagnostic tool based on the data state fluctuation coefficient to perform anomaly detection and root cause tracing on the global data state map sequence, and outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis" includes:
[0148] The optimal number of branches to be selected is obtained by multiplying the ratio of the data state fluctuation coefficient to the historical maximum data state fluctuation coefficient recorded in the target area within a preset time range by J and rounding it down, where H is greater than or equal to 2 and less than or equal to J.
[0149] H data anomaly diagnosis branches are randomly selected from the J data anomaly diagnosis branches, and anomaly detection and root cause tracing are performed on the global data state map sequence respectively, outputting H initial inconsistent node distributions and H initial anomaly diagnosis probability distributions.
[0150] The nodes whose abnormal frequency is greater than or equal to a preset frequency threshold among the H initial inconsistent node distributions are taken as the final inconsistent nodes, and an inconsistent node distribution is generated, wherein the preset frequency threshold is less than 0.2H.
[0151] Based on the inconsistent node distribution, the mean of the H initial anomaly diagnosis probability distributions is fitted to obtain the anomaly diagnosis probability distribution.
[0152] Specifically, the "generation of consistency repair strategy" includes:
[0153] Configure a consistency repair mechanism based on the aforementioned data state fluctuation coefficient;
[0154] Establish a root cause-action mapping knowledge base, generate a consistency repair strategy based on the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and execute the consistency repair strategy according to the consistency repair mechanism.
[0155] Furthermore, the "configuration of a consistency repair mechanism based on the data state fluctuation coefficient" includes:
[0156] Configure a first data state fluctuation coefficient scalar and a second data state fluctuation coefficient scalar, wherein the first data state fluctuation coefficient scalar is smaller than the second data state fluctuation coefficient scalar;
[0157] If the data state fluctuation coefficient is less than the first data state fluctuation coefficient scalar, the consistency repair mechanism is configured as fixed-point repair, wherein fixed-point repair only repairs inconsistent nodes;
[0158] If the data state fluctuation coefficient is greater than or equal to the first data state fluctuation coefficient scalar and less than the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as local repair, wherein local repair is to repair all nodes within the preset power supply area where the inconsistent node is located;
[0159] If the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as global repair, wherein global repair is to repair all nodes in the target area.
[0160] In summary, the embodiments of this application have at least the following technical effects:
[0161] Compared to existing technologies, this application firstly deploys data probes at all nodes in the power grid meter reading data flow through a data acquisition and graph construction module. This synchronously collects and acquires the power data status distribution sequence over historical periods and constructs a global data status graph sequence. This transforms discrete data into a structured global data status graph sequence containing node topological relationships and temporal characteristics, providing comprehensive and accurate foundational data support for subsequent time-series fluctuation analysis, anomaly detection, and root cause tracing. Secondly, through a fluctuation identification module, it identifies the time-series fluctuation patterns of the global data status graph sequence and obtains the graph fluctuation deviation sequence. This enables precise characterization of the degree of data status fluctuation, providing a reliable quantitative basis for distinguishing between normal fluctuations and abnormal deviations and triggering targeted anomaly diagnosis. Finally, through the consistency verification module, based on the graph fluctuation deviation sequence, a data anomaly diagnostic tool built on a graph neural network is used to perform anomaly detection and root cause tracing on the global data state graph sequence. This outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generates a consistency repair strategy to perform closed-loop governance of power grid meter reading data. This achieves accurate detection and root cause quantification of anomalies across the entire power grid meter reading data chain based on quantitative fluctuation criteria, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating targeted consistency repair strategies to form a closed-loop governance, effectively ensuring the consistency and reliability of power grid meter reading data. In this way, the technical problems of low anomaly diagnosis accuracy, fuzzy root cause tracing, and lack of targeted repair strategies in traditional methods are effectively solved. This ensures the consistency, reliability, and timeliness of power grid meter reading data among distributed nodes, providing high-quality data support for power dispatching, electricity billing, and power supply quality assessment.
[0162] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0167] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0168] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A distributed data consistency verification method based on graph neural networks, characterized in that, The methods include: Data probes are deployed at all nodes in the power grid meter reading data flow to synchronously collect and acquire the power data status distribution sequence within historical time periods, and to construct a global data status map sequence. Identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence; Based on the graph fluctuation deviation sequence, an anomaly diagnostic tool built on graph neural network is used to perform anomaly detection and root cause tracing on the global data state graph sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generating a consistency repair strategy to perform consistency closed-loop governance on power grid meter reading data. Among them, the data anomaly diagnostic tool based on graph neural networks includes: Based on historical power grid meter reading data monitoring records, several global data status map sequence samples were collected, and the historical inconsistent node distribution and historical anomaly diagnosis probability distribution corresponding to different global data status map sequence samples were labeled to obtain several inconsistent node distribution samples and several anomaly diagnosis probability distribution samples. The historical anomaly diagnosis probability distribution includes historical anomaly probability distributions corresponding to multiple inconsistent node historical anomaly diagnosis types. The anomaly diagnosis types include at least node data loss due to communication link interruption, data version lag due to synchronization task delay, temporary data expiration due to cache not being refreshed, and transmission data corruption due to data verification errors. The training data consists of several global data state map sequence samples, several inconsistent node distribution samples, and several abnormal diagnosis probability distribution samples. J-fold cross-partitioning is performed to obtain J sample training sets, where J is an integer greater than 5. The graph neural networks are trained using the J sample training sets respectively, J data anomaly diagnosis branches are constructed, and the data anomaly diagnostic tool is built by integrating them. Specifically, based on the spectrum fluctuation deviation sequence, a data anomaly diagnostic tool is used to perform anomaly detection and root cause tracing on the global data state spectrum sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, including: The average value of the spectral fluctuation deviation sequence is obtained by averaging the spectral fluctuation deviation sequence. The ratio of the standard deviation of the spectral fluctuation deviation in the spectral fluctuation deviation sequence to the mean of the spectral fluctuation deviation is set as the deviation fluctuation index. The data state fluctuation coefficient is calculated based on the mean fluctuation deviation and the deviation fluctuation index of the aforementioned spectrum. Based on the data state fluctuation coefficient, the data anomaly diagnostic tool is invoked to perform anomaly detection and root cause tracing on the global data state map sequence, and the distribution of inconsistent nodes and anomaly diagnosis probability distribution are output. Specifically, based on the data state fluctuation coefficient, the data anomaly diagnostic tool is invoked to perform anomaly detection and root cause tracing on the global data state map sequence, outputting the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, including: The optimal number of branches to be selected is obtained by multiplying the ratio of the data state fluctuation coefficient to the historical maximum data state fluctuation coefficient recorded in the target area within a preset time range by J and rounding it down, where H is greater than or equal to 2 and less than or equal to J. H data anomaly diagnosis branches are randomly selected from the J data anomaly diagnosis branches, and anomaly detection and root cause tracing are performed on the global data state map sequence respectively, outputting H initial inconsistent node distributions and H initial anomaly diagnosis probability distributions. The nodes whose abnormal frequency is greater than or equal to a preset frequency threshold among the H initial inconsistent node distributions are taken as the final inconsistent nodes, and an inconsistent node distribution is generated, wherein the preset frequency threshold is less than 0.2H. Based on the inconsistent node distribution, the mean of the H initial anomaly diagnosis probability distributions is fitted to obtain the anomaly diagnosis probability distribution.
2. The distributed data consistency verification method based on graph neural networks according to claim 1, characterized in that, Data probes are deployed at all nodes in the power grid meter reading data flow to synchronously collect and acquire the power data status distribution sequence within historical time periods, including: Data probes are deployed at all nodes in the power grid meter reading data flow in the target area to build a real-time data status monitoring network. The all-link nodes include at least terminal metering units, edge acquisition devices, communication aggregation nodes, data acquisition servers, business databases, and cache service nodes. Using the real-time data status monitoring network, the power data status distribution sequence of the target area in historical time periods is synchronously collected according to the preset monitoring frequency and preset monitoring indicators. The preset monitoring indicators include at least data version identifier, timestamp sequence, node topology relationship, cache effective status and communication link quality.
3. The distributed data consistency verification method based on graph neural networks according to claim 2, characterized in that, Construct a global data state map sequence, including: Using each power data state distribution in the power data state distribution sequence as a vertex, and using the data version identifier, timestamp sequence and cache effective status collected at the same time as vertex attributes, an initial data state map sequence is constructed. Based on the node topology, edges are established between corresponding vertices, and the communication link quality is used as the basic attribute of the edges to structurally supplement the initial data state graph sequence, generating a global data state graph sequence.
4. The distributed data consistency verification method based on graph neural networks according to claim 1, characterized in that, Identifying the temporal fluctuation patterns of the global data state map sequence and obtaining the map fluctuation deviation sequence includes: For any two adjacent global data state maps in the global data state map sequence, perform structural similarity comparison and attribute similarity comparison, and determine the similarity of multiple maps by weighting. The graph similarity is subtracted from 1 as the graph fluctuation deviation, and multiple graph fluctuation deviations are calculated based on the multiple graph similarities; Based on the sequential order of the graphs in the global data state graph sequence, the fluctuation deviations of the multiple graphs are sorted to generate a graph fluctuation deviation sequence.
5. The distributed data consistency verification method based on graph neural networks according to claim 1, characterized in that, Generate a consistency repair strategy, including: Configure a consistency repair mechanism based on the aforementioned data state fluctuation coefficient; Establish a root cause-action mapping knowledge base, generate a consistency repair strategy based on the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and execute the consistency repair strategy according to the consistency repair mechanism.
6. The distributed data consistency verification method based on graph neural networks according to claim 5, characterized in that, Based on the data state fluctuation coefficient, a consistency repair mechanism is configured, including: Configure a first data state fluctuation coefficient scalar and a second data state fluctuation coefficient scalar, wherein the first data state fluctuation coefficient scalar is smaller than the second data state fluctuation coefficient scalar; If the data state fluctuation coefficient is less than the first data state fluctuation coefficient scalar, the consistency repair mechanism is configured as fixed-point repair, wherein fixed-point repair only repairs inconsistent nodes; If the data state fluctuation coefficient is greater than or equal to the first data state fluctuation coefficient scalar and less than the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as local repair, wherein local repair is to repair all nodes within the preset power supply area where the inconsistent node is located; If the data state fluctuation coefficient is greater than or equal to the second data state fluctuation coefficient scalar, the consistency repair mechanism is configured as global repair, wherein global repair is to repair all nodes in the target area.
7. A distributed data consistency verification system based on graph neural networks, characterized in that, The method for performing the distributed data consistency verification method based on graph neural networks according to any one of claims 1-6 includes: The data acquisition and graph construction module is used to deploy data probes at all nodes in the power grid meter reading data flow, synchronously collect the power data status distribution sequence within historical time periods, and construct a global data status graph sequence. The fluctuation identification module is used to identify the temporal fluctuation pattern of the global data state map sequence and obtain the map fluctuation deviation sequence; The consistency verification module is used to perform anomaly detection and root cause tracing on the global data state graph sequence based on the graph fluctuation deviation sequence and a data anomaly diagnostic tool built on graph neural network. It outputs the distribution of inconsistent nodes and the probability distribution of anomaly diagnosis, and generates a consistency repair strategy to perform closed-loop consistency management of power grid meter reading data.
Citation Information
Patent Citations
Fault root cause positioning method and system driven by dynamic knowledge graph
CN120950284A
Distributed electric power measurement anomaly detection method and system based on graph neural network
CN121388838A