A power data quality evaluation method and system based on anomaly detection
By constructing a static graph structure and dynamic adjacency matrix, combining adaptive weight allocation and multi-head attention mechanism, the problem of variable dependence in power data changes over time is solved, and more accurate power data quality evaluation and abnormal detection are achieved, improving the adaptability and safety of the power system.
Patent Information
- Application Number
- CN202510290737.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing power data quality evaluation methods fail to effectively capture the dynamic characteristics of variable dependencies in power data over time, resulting in insufficient abnormal detection accuracy in rapidly changing or non-stationary environments, affecting the safe operation and decision-making of the power system.
A data preprocessing model, dynamic and static data calculation model, anomaly detection model and data quality evaluation model are constructed, and a static graph structure and dynamic adjacency matrix are generated. The dynamic and static anomaly patterns of power data are captured through adaptive weight allocation and multi-head attention mechanisms, and the power data quality evaluation is carried out in combination with hierarchical analysis method.
It improves the accuracy and robustness of power data abnormal detection, enhances the model's adaptability to rapidly changing or non-stationary environments, provides a more scientific and systematic power data quality evaluation system, and supports real-time monitoring and decision-making of power systems.
Smart Images

Figure CN119782966B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a power data quality evaluation method and system based on anomaly detection, belonging to the technical field of power data processing. Background Art
[0002] One of the key indicators affecting the quality of power data is outlier data, which has an important impact on the accuracy, integrity, and self-consistency of power data quality. For example: 1. The existence of abnormal power consumption data will reduce the accuracy of data quality assessment. 2. When data mining is performed on a dataset containing a large number of outliers, the results cannot accurately reflect the characteristics of the data. 3. Outlier data will affect the judgment and decision-making of power system dispatchers and even threaten the safe operation of the system. 4. If outlier data is used as modeling data, it will interfere with the change rules and affect the training accuracy. Therefore, the anomaly detection of power data has become the core issue in the research of data quality assessment.
[0003] Furthermore, Chinese Patent (Publication No.: CN118378023B) proposes a power data anomaly detection method based on graph computing, including extracting and preprocessing power big data, constructing a power data graph model, extracting structural and attribute features, using graph computing algorithms for anomaly detection, and visualizing the output results. By optimizing the preprocessing operation, improving the graph model construction method, and improving the application of graph computing in anomaly detection, the accuracy, comprehensiveness, and efficiency of power data anomaly detection are improved.
[0004] The above solution uses a data graph model to extract structural and attribute features, but does not consider the variable dependence relationship in power data, which will change over time. For example, factors such as seasonal fluctuations have a greater impact on power data. Therefore, the above graph computing method has limitations in capturing these dynamic changes and cannot adapt to the rapid changes or non-stationary environment in the power system, resulting in large errors in power data anomaly detection and affecting the quality evaluation of power data.
[0005] The information disclosed in this background art is only used to understand the background of the inventive concept of the present invention, and therefore it may include information that does not constitute prior art. Summary of the Invention
[0006] In view of the above problems or one of the above problems, one object of the present invention is to provide a power data quality evaluation method and system based on anomaly detection. By constructing a data preprocessing model, a static and dynamic data calculation model, an anomaly detection model, and a data quality evaluation model, a static graph structure and a dynamic adjacency matrix are generated to identify abnormal user data, and the power data quality evaluation based on anomaly detection is completed. Therefore, the long-term trend can be captured by using the static graph structure, the local dependence structure that changes with time can be flexibly adapted by using the dynamic adjacency matrix, the short-term dynamic changes can be captured, and a variable interaction picture can be provided. Thus, the variable dependence relationship in the power data can be fully considered to change with time, effectively enhancing the adaptability of the model to the rapidly changing or non-stationary environment in the power system. Therefore, accurate power data anomaly detection results can be obtained, and then accurate quality evaluation of the power data can be carried out.
[0007] In view of the above problems or one of the above problems, another object of the present invention is to provide a power data quality evaluation method and system based on anomaly detection. Through adaptive weight allocation and multi-head attention mechanism, the dynamic and static abnormal patterns of power data can be effectively captured, the robustness and generalization ability of the model can be improved, overfitting can be reduced, manual threshold setting is not required, and the human influence can be reduced, thereby improving the accuracy of anomaly detection and the adaptability to complex scenarios; and through the analytic hierarchy process, complex decision-making problems are decomposed into a hierarchical structure, combined with abnormal user information, a judgment matrix is constructed and a weight vector is calculated, so as to obtain a more scientific, systematic and reliable power data quality evaluation system, providing theoretical support for power system data management.
[0008] To achieve one of the above objects, the first technical solution of the present invention is as follows:
[0009] A power data quality evaluation method based on anomaly detection, comprising the following steps:
[0010] Step 1, through a pre-constructed data preprocessing model, clean the power data to be evaluated to obtain the power consumption data of multiple users;
[0011] Step 2, using a pre-constructed static and dynamic data calculation model, based on the electricity consumption similarity between users, process the power consumption data of multiple users to construct a static graph structure and a dynamic adjacency matrix;
[0012] The static graph structure has several nodes and connection edges, which are used to smooth short-term fluctuations and capture long-term trends; the nodes represent users, and the connection edges represent the similarity between users;
[0013] The dynamic adjacency matrix is used to flexibly adapt to the local dependence structure that changes with time, capture short-term dynamic changes, and provide a variable interaction picture;
[0014] Step 3: Use the pre-constructed anomaly detection model to process the static graph structure and the dynamic adjacency matrix, capture the similarity among users, and identify abnormal user data;
[0015] Step 4: Use the pre-constructed data quality evaluation model to evaluate the power data based on the abnormal user data, obtain the quality evaluation result, and complete the power data quality evaluation based on anomaly detection.
[0016] By constructing a data preprocessing model, a static and dynamic data calculation model, an anomaly detection model, and a data quality evaluation model, generating a static graph structure and a dynamic adjacency matrix, identifying abnormal user data, and completing the power data quality evaluation based on anomaly detection, the present invention can utilize the static graph structure to capture long-term trends, utilize the dynamic adjacency matrix to flexibly adapt to the time-varying local dependence structure, capture short-term dynamic changes, and provide a variable interaction picture. Thus, it can fully consider that the variable dependence relationship in power data changes over time, can timely capture the dynamic changes of data, effectively enhance the adaptability of the model to the rapidly changing or non-stationary environment in the power system, so as to obtain accurate power data anomaly detection results, and further can accurately evaluate the quality of power data.
[0017] Furthermore, the dynamic adjacency matrix of the present invention not only retains the capture of long-term dependence relationships, but also can sensitively respond to short-term dynamic changes, providing more reliable support for the real-time monitoring and decision-making of the power system. Furthermore, by regularly updating the static graph structure and the dynamic adjacency matrix, the present invention can more accurately reflect the instantaneous interaction pattern between variables, ensure high sensitivity to the current situation, thereby improving the accuracy of anomaly detection and the adaptability to complex scenarios.
[0018] As a preferred technical measure:
[0019] Step 1: The method for cleaning the power data to be evaluated through the pre-constructed data preprocessing model to obtain the power consumption data of multiple users is as follows:
[0020] Obtain the power data to be evaluated, where the power data is data involving multiple users sampled by day;
[0021] Clean and standardize the power data, remove the power consumption exceeding the set range, and obtain the power standard data;
[0022] Fill in the missing data in the power standard data to obtain the power complete data;
[0023] Split the power complete data according to user attributes to obtain the power consumption data of multiple users.
[0024] As a preferred technical measure:
[0025] Step 2: Using the pre-constructed dynamic and static data calculation model, based on the electricity consumption similarity among users, process the electricity consumption data of multiple users. The method for constructing a static graph structure is as follows:
[0026] Sum up the electricity consumption data of each user over a one-week time span to smooth short-term fluctuations and capture long-term trends, obtaining the weekly total electricity consumption of several users;
[0027] Normalize the weekly total electricity consumption of several users to construct multiple user electricity consumption vectors, ensuring that the electricity consumption data among different users is on the same scale;
[0028] Based on multiple user electricity consumption vectors, calculate the included angle value between any two user electricity consumption vectors to obtain included angle information;
[0029] According to the included angle information, calculate the cosine similarity between each user and all the remaining users to obtain similarity data;
[0030] Based on the similarity data, select the top M users with the highest cosine similarity as the neighbor nodes of this user, and use the cosine similarity as the weight of the connection edge to represent the strength of the similarity;
[0031] Use the connection edge to connect this user and the neighbor nodes to obtain the static initial graph structure of this user;
[0032] Stitch together the static initial graph structures of all users to obtain a complete graph structure;
[0033] Use the weekly electricity consumption data of users as feature vectors to establish a feature matrix and a static adjacency matrix;
[0034] Based on the feature matrix and the static adjacency matrix, assign values to the complete graph structure to obtain a static graph structure.
[0035] As an optimal technical measure:
[0036] Step 2: Using the pre-constructed dynamic and static data calculation model, based on the electricity consumption similarity among users, process the electricity consumption data of multiple users. The method for constructing a dynamic adjacency matrix is as follows:
[0037] According to the electricity consumption characteristics of different users, set different sampling time lengths;
[0038] The sampling time length is half a month or one month or one quarter;
[0039] According to the sampling time length, divide the electricity consumption data of users to obtain a window length, that is, the electricity consumption data of users within a time period;
[0040] According to the window length, represent the daily power consumption data of each user as a vector to obtain a power consumption matrix regarding the window length;
[0041] According to the power consumption matrix and based on the power consumption similarity among users, set an adjacency matrix regarding the window length, thereby obtaining a dynamic adjacency matrix among users.
[0042] As a preferred technical measure:
[0043] Step 3: Use a pre-constructed anomaly detection model to process the static graph structure and the dynamic adjacency matrix, capture the similarity among users, and identify abnormal user data as follows:
[0044] Obtain the static graph structure, which includes a feature matrix and a static adjacency matrix;
[0045] According to the feature matrix and the static adjacency matrix, obtain the feature vector of a certain node and the feature vectors of its neighbor nodes;
[0046] Use the trained graph attention network to perform weighted averaging on the feature vector of a certain node and the feature vectors of its neighbor nodes to obtain the static feature vector of a certain node;
[0047] In order to capture the dynamic dependencies that change over time in the power system, based on the dynamic adjacency matrix, obtain the power consumption matrix corresponding to the window length;
[0048] Use the graph attention network to perform graph attention network convolution on the power consumption matrix to obtain several feature vectors of each user;
[0049] Average the several feature vectors to obtain a comprehensive dynamic feature vector:
[0050] Concatenate the static feature vector and the dynamic feature vector together to form a new concatenated feature vector, which fuses the static attributes and dynamic attributes of the user;
[0051] Input the concatenated feature vector into a multi-layer perceptron and perform binary classification through an activation function to obtain the final classification probability of this user; the anomaly detection probability score of each user is a decimal between 0 and 1, indicating the probability that this user is abnormal;
[0052] If the final classification probability is greater than the set anomaly threshold, then determine that this user is an abnormal user and count the number of abnormal users; otherwise, determine as a normal user.
[0053] As a preferred technical measure:
[0054] Step 4. Using the pre-constructed data quality evaluation model, evaluate the power data based on the abnormal user data. The method for obtaining the quality evaluation result is as follows:
[0055] Select several indicators for power data quality evaluation, including integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator;
[0056] Adopt the analytic hierarchy process, and construct the corresponding weight matrix according to the influence degrees of the integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator on the power data quality;
[0057] Based on the weight matrix, set up the judgment matrix;
[0058] Standardize each column of the judgment matrix, sum the standardized elements by row, and finally standardize the summation result to obtain the weight vector;
[0059] Calculate the specific values of the integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator to obtain the dimension scores of the four indicators;
[0060] Multiply the dimension scores of the four indicators by the weight vector to calculate the power data quality score;
[0061] According to the power data quality rating table, rate the power data quality score to obtain the quality evaluation result.
[0062] As an optimal technical measure:
[0063] The method for calculating the specific value of the integrity indicator is as follows:
[0064] Obtain the power consumption data of multiple users and get the total number of data records;
[0065] Based on the data record integrity judgment mechanism, judge whether each record in the power consumption data contains the necessary field information to obtain the first complete record count;
[0066] The data record integrity judgment mechanism refers to whether each record in the data set contains the necessary field information; if a certain record lacks a certain necessary field information, it is considered an incomplete data record; the necessary field information includes timestamp, voltage, current, and power;
[0067] Based on the time series frequency integrity judgment mechanism, judge whether the records in the power consumption data are sampled and recorded at the expected time interval to obtain the second complete record count;
[0068] The time - series frequency integrity judgment mechanism refers to whether the records in the dataset are sampled and recorded at the expected time intervals; if the data at some time points is missing, or the recorded time intervals are uneven, it will be regarded as the frequency of the time - series being incomplete;
[0069] Add the number of the first complete records and the number of the second complete records to obtain the number of non - empty records that meet the conditions;
[0070] Divide the number of non - empty records by the total number of data records to obtain the specific value of the integrity index;
[0071] Or / and, the method for calculating the specific value of the uniqueness index is as follows:
[0072] Obtain the power consumption data of multiple users and get the total number of data records;
[0073] Based on the numerical uniqueness judgment mechanism, judge whether each record in the power consumption data is unique in some key fields to obtain the number of data uniqueness;
[0074] The numerical uniqueness judgment mechanism means that each record is not repeated in some key fields; some key fields include timestamp, device identifier, and user identifier;
[0075] Based on the number of data uniqueness and the total number of data records, calculate the specific value of the uniqueness index;
[0076] Or / and, the method for calculating the specific value of the timeliness index is as follows:
[0077] Obtain the power consumption data of multiple users and get the total number of data records;
[0078] Based on the timeliness judgment mechanism, judge whether the timestamp of each record in the power consumption data increases monotonically to obtain the number of timeliness errors;
[0079] The timeliness judgment mechanism is used to measure whether the data is recorded in a timely manner and at the expected time intervals, that is, the time of each record needs to be later than the time of the previous record; if it is found that the timestamp does not conform to the monotonicity principle, it is considered as a time - series error;
[0080] Based on the total number of data records and the number of timeliness errors, calculate the specific value of the timeliness index;
[0081] Or / and, the method for calculating the specific value of the accuracy index is as follows:
[0082] Obtain the number of abnormal users and the power consumption data of multiple users, and get the total number of data records;
[0083] Based on the accuracy index judgment mechanism, find the number of abnormal data;
[0084] The abnormal data includes invalid data, data with format errors, and data with unreasonable content;
[0085] Add the number of abnormal data and the number of abnormal users to obtain the number of abnormal records;
[0086] Based on the number of abnormal records and the total number of data records, obtain the specific value of the accuracy index.
[0087] As a preferred technical measure:
[0088] Based on the weight matrix, the method for setting the judgment matrix is as follows:
[0089] Step 11, based on the weight matrix, construct an initial judgment matrix for the integrity index, uniqueness index, timeliness index, and accuracy index;
[0090] Step 12, calculate the maximum eigenvalue of the initial judgment matrix;
[0091] Step 13, based on the maximum eigenvalue of the initial judgment matrix, calculate the consistency index;
[0092] And obtain the random consistency index by referring to the random consistency index table;
[0093] Step 14, compare the consistency index and the random consistency index to obtain the consistency ratio;
[0094] Step 15, compare the consistency ratio with the ratio threshold, which specifically includes the following content:
[0095] When the consistency ratio is less than the ratio threshold, it indicates that the initial judgment matrix meets the consistency check, and execute Step 16;
[0096] When the consistency ratio is greater than or equal to the ratio threshold, reconstruct the initial judgment matrix and execute Step 12;
[0097] Step 16, use the initial judgment matrix as the final judgment matrix.
[0098] The ratio threshold is 0.1.
[0099] To achieve one of the above purposes, the second technical solution of the present invention is:
[0100] A power data quality evaluation method based on anomaly detection, including the following content:
[0101] Clean and transform the power data to be evaluated to obtain the power consumption data of multiple users;
[0102] Based on the graph attention network's adaptive weight allocation and multi-head attention mechanism, process the power consumption data of multiple users, and construct a static graph structure and a dynamic adjacency matrix for capturing user similarity;
[0103] Process the static graph structure and the dynamic adjacency matrix to identify abnormal users and count the number of abnormal users;
[0104] Through the analytic hierarchy process, combined with the abnormal users and the number of abnormal users, construct a judgment matrix and calculate the weight vector, evaluate the power data, and obtain the quality evaluation result, completing the power data quality evaluation based on anomaly detection.
[0105] The present invention effectively captures the dynamic and static abnormal patterns of power data through adaptive weight allocation and multi-head attention mechanism, improves the robustness and generalization ability of the model, reduces overfitting, eliminates the need for manual threshold setting, reduces human influence, thereby improving the accuracy of anomaly detection and the adaptability to complex scenarios; and decomposes complex decision-making problems into a hierarchical structure through the analytic hierarchy process, combines abnormal user information, constructs a judgment matrix and calculates the weight vector, thus obtaining a more scientific, systematic, and reliable power data quality evaluation system, providing theoretical support for power system data management.
[0106] To achieve one of the above purposes, the third technical solution of the present invention is:
[0107] A power data quality evaluation system based on anomaly detection, which includes:
[0108] One or more processors;
[0109] A storage device for storing one or more programs;
[0110] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned power data quality evaluation method based on anomaly detection.
[0111] Compared with the prior art solutions, the present invention has the following beneficial effects:
[0112] The present invention generates a static graph structure and a dynamic adjacency matrix by constructing a data preprocessing model, a static and dynamic data calculation model, an anomaly detection model, and a data quality evaluation model, identifies abnormal user data, and completes the power data quality evaluation based on anomaly detection. Therefore, the variable dependence relationship in power data can be fully considered to change over time. At the same time, the present invention uses the static graph structure to capture long-term trends, uses the dynamic adjacency matrix to flexibly adapt to the locally dependent structure that changes over time, captures short-term dynamic changes, and can provide a picture of variable interaction, thereby effectively enhancing the adaptability of the model to the rapidly changing or non-stationary environment in the power system, obtaining accurate power data anomaly detection results, and then effectively performing accurate quality evaluation on power data.
[0113] Furthermore, the dynamic adjacency matrix of the present invention not only retains the capture of long-term dependence relationships, but also can sensitively respond to short-term dynamic changes, providing more reliable support for the real-time monitoring and decision-making of the power system. Furthermore, by regularly updating the static graph structure and the dynamic adjacency matrix, the present invention can more accurately reflect the instantaneous interaction patterns between variables, ensure a high sensitivity to the current situation, and thus improve the accuracy of anomaly detection and the adaptability to complex scenarios.
[0114] Even further, the present invention effectively captures the dynamic and static abnormal patterns of power data through adaptive weight allocation and the multi-head attention mechanism, improves the robustness and generalization ability of the model, reduces overfitting, eliminates the need for manual threshold setting, reduces human influence, and thus improves the accuracy of anomaly detection and the adaptability to complex scenarios; and decomposes complex decision-making problems into a hierarchical structure through the analytic hierarchy process, combines abnormal user information, constructs a judgment matrix and calculates a weight vector, thereby obtaining a more scientific, systematic, and reliable power data quality evaluation system, providing a theoretical support for power system data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 is a schematic flow chart of a method for evaluating the quality of power data of the present invention;
[0116] Figure 2 is a schematic framework diagram for constructing a power data quality evaluation system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0117] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0118] On the contrary, the present invention covers any alternatives, modifications, equivalent methods, and solutions that are defined by the claims and fall within the essence and scope of the present invention. Further, in order to enable the public to better understand the present invention, in the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention even without the description of these details.
[0119] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0120] As Figure 1 shown, the first specific embodiment of the power data quality evaluation method based on anomaly detection of the present invention:
[0121] A power data quality evaluation method based on anomaly detection, comprising the following steps:
[0122] Step 1, through a pre-constructed data preprocessing model, clean the power data to be evaluated to obtain the power consumption data of multiple users;
[0123] Step 2, using a pre-constructed static and dynamic data calculation model, based on the electricity consumption similarity between users, process the power consumption data of multiple users to construct a static graph structure and a dynamic adjacency matrix;
[0124] The static graph structure has several nodes and connection edges, which are used to smooth short-term fluctuations and capture long-term trends; the nodes represent users, and the connection edges represent the similarity between users;
[0125] The dynamic adjacency matrix is used to flexibly adapt to the locally dependent structure that changes over time, capture short-term dynamic changes, and can provide a picture of variable interactions;
[0126] Step 3, using a pre-constructed anomaly detection model, process the static graph structure and the dynamic adjacency matrix, capture the similarity between users, and identify abnormal user data;
[0127] Step 4, using a pre-constructed data quality evaluation model, evaluate the power data based on the abnormal user data to obtain a quality evaluation result, and complete the power data quality evaluation based on anomaly detection.
[0128] The second specific embodiment of the power data quality evaluation method based on anomaly detection of the present invention:
[0129] A power data quality evaluation method based on anomaly detection, comprising the following steps:
[0130] Step 1: Through a pre-constructed data preprocessing model, clean and transform the power data to be evaluated to obtain the power consumption data of multiple users;
[0131] Step 2: Use a pre-constructed static and dynamic data calculation model to process the power consumption data of multiple users based on the electricity consumption similarity between users. Each user is regarded as a node, and the similarity between users is represented by an edge connection, and a graph structure for capturing the similarity between users is constructed;
[0132] Step 3: Use a pre-constructed anomaly detection model to process the graph structure, identify abnormal users, and count the number of abnormal users;
[0133] Step 4: Use a pre-constructed data quality evaluation model to generate an accuracy index based on the abnormal users and the number of abnormal users, evaluate the power data, and complete the power data quality evaluation based on anomaly detection.
[0134] The third specific embodiment of the power data quality evaluation method based on anomaly detection of the present invention:
[0135] A power data quality evaluation method based on anomaly detection, including the following:
[0136] Clean and transform the power data to be evaluated to obtain the power consumption data of multiple users;
[0137] Based on the adaptive weight allocation and multi-head attention mechanism of the graph attention network, process the power consumption data of multiple users, and construct a static graph structure and a dynamic adjacency matrix for capturing user similarity;
[0138] Process the static graph structure and the dynamic adjacency matrix, identify abnormal users, and count the number of abnormal users;
[0139] Through the analytic hierarchy process, combine the abnormal users and the number of abnormal users, construct a judgment matrix and calculate the weight vector, evaluate the power data, obtain the quality evaluation result, and complete the power data quality evaluation based on anomaly detection.
[0140] In this embodiment, the method for constructing the static graph structure is as follows:
[0141] Before constructing the static graph structure of the graph attention network, it is necessary to calculate the cosine similarity between node i and all the remaining nodes. Specifically, when calculating the similarity between two user nodes using cosine similarity, first sum the power consumption data of each user over a one-week time span to smooth short-term fluctuations and capture long-term trends.
[0142] Then, the total weekly electricity consumption is normalized to ensure that the electricity consumption data among different users is on the same scale, eliminating the influence of dimensional differences. Each user is finally represented as a vector of length 147, containing the normalized weekly electricity consumption data of the user within 147 weeks.
[0143] Next, by calculating the cosine similarity between each user and all other users, the top 5 users with the highest similarity are selected as neighbor nodes, and these nodes are connected with edges. The weight of the edge is set to the value of the cosine similarity.
[0144] The graph structure constructed in this way not only reflects the long-term electricity consumption patterns of users, but also can flexibly adapt to the complexity and variability over time in the power system, providing a solid foundation for subsequent graph attention network model training and anomaly detection.
[0145] For the weekly vectors of two users, the cosine similarity measures their similarity by calculating the cosine value of the angle between the two vectors , and is specifically defined as:
[0146] (1)
[0147] where and are the norms (lengths) of vectors and vector respectively.
[0148] Here, each node is regarded as a vector. For node i, after calculating the cosine similarity between it and all the remaining nodes, we get , and the top 5 nodes with the highest numerical values are taken as neighbor nodes to form the initial static graph structure of the nodes. The weight of the edge is the value of the cosine similarity between the nodes.
[0149] In this embodiment, the method for constructing the dynamic adjacency matrix is as follows:
[0150] To further improve the ability of the graph attention network to capture the complex and variable user electricity consumption patterns in the power system, this application introduces a dynamic adjacency matrix on the basis of the static adjacency matrix. The dynamic adjacency matrix can flexibly adapt to the time-varying local dependence structure, capture short-term dynamic changes, and thus provide a more detailed and accurate picture of variable interactions.
[0151] Specifically, instead of averaging the data set according to a one-week time span, the daily data is directly divided into multiple time windows, and the total number of windows is . Each window contains data of size , where represents the total number of users, denotes the window length, and is generally a relatively long window length, such as one month or longer. In this way, multiple time windows can be obtained, and each window corresponds to the user's electricity consumption data within a certain time period.
[0152] Next, for each window, the daily electricity consumption data of each user is represented as a vector. Suppose the th window contains days of data, then the electricity consumption of user i within this window can be represented as a vector of length where each element represents the electricity consumption of user i on a certain day within the window. Let the electricity consumption matrix of the th window be where represents the total number of users. Then, using the similarity of data among users, the adjacency matrix within each window is defined, and its expression is as follows:
[0153] (2)(2)
[0154] (3)(3)
[0155] where denotes the total number of users, denotes the length (number of days) of the time window, denotes the dimension of the feature space, denotes a real number, is the dynamic adjacency matrix of the t-th time window, with size representing the connection relationship between users, is the reset value of the dynamic adjacency matrix of the t-th time window; denotes the activation function (such as Sigmoid or ReLU) used to perform a non-linear transformation on the similarity score; and denote the mapping matrices, with size used to map the user's electricity consumption data to the feature space; and denote the results after mapping the user's electricity consumption data matrix and to the feature space through the mapping matrices with size , denotes the user's electricity consumption data matrix of the th time window, with size where each row represents the daily electricity consumption vector of a user. Specifically, first use two different mapping functions and , map to the same high-dimensional space .
[0156] Next, use the inner product similarity to obtain the similarity matrix, that is and . The two similarity matrices can respectively represent the dependence of the source node on the target node and the dependence of the target node on the source node, and the sum of the two represents bidirectional modeling. To maintain the sparsity of the adjacency matrix and prevent overfitting, a threshold is used to screen the adjacency matrix. If the similarity score at a certain position is greater than or equal to the threshold, the value at that position is 1, otherwise it is 0. Usually, the threshold can be set to the average value of all similarity scores in all adjacency matrices. Finally, the obtained is the dynamic adjacency matrix between users in the th window.
[0157] The fourth specific embodiment of the power data quality evaluation method based on anomaly detection of the present invention:
[0158] A power data quality evaluation method based on anomaly detection, including the following:
[0159] According to the data evaluation criteria of "Information Technology - Data Quality Evaluation Index" GB / T 36344-2018, a power system data quality evaluation dimension framework is formulated, and evaluations are carried out from four dimensions of accuracy, integrity, uniqueness, and timeliness respectively, and the analytic hierarchy process is introduced. By constructing a judgment matrix and calculating the weight vector, a complete power data quality evaluation system is obtained.
[0160] Among them, the index of data accuracy is obtained from the anomaly detection model based on the graph attention network. The anomaly detection model based on the graph attention network is constructed based on the graph attention network, and it at least includes a data preprocessing model, a static and dynamic data calculation model, an anomaly detection model, and a data quality evaluation model. The processing method is as follows:
[0161] First, by calculating the cosine similarity, construct the original static graph structure and dynamic adjacency matrix of user nodes; through the static graph structure and dynamic adjacency matrix, identify the abnormal situations in the power consumption data to obtain the accuracy index; then, through the adaptive weight allocation and multi-head attention mechanism of the graph attention network, effectively capture the complex patterns and anomalies in the daily power consumption data of users; by weighted summing the features of neighbor nodes, update the representation of each node, and learn higher-level feature representations through the hidden layer. At the output layer, the graph attention network can generate anomaly scores and perform classification, and identify anomalies by comparing the difference between the predicted value and the actual value.
[0162] In this embodiment, the method for processing power data using a data preprocessing model is as follows:
[0163] First, obtain the power data collected by the smart grid meters. This power data is data involving multiple users sampled daily. Then, clean and standardize the power data, remove the power consumption beyond the set range, and obtain the power standard data. Furthermore, fill in the missing data in the power standard data to obtain the complete power data. Split the complete power data according to user attributes to obtain the power consumption data of multiple users.
[0164] In this embodiment, the method for processing power data using a dynamic and static data calculation model based on a graph attention network is as follows:
[0165] Graph Attention Networks (abbreviated as GAT) is a type of graph neural network (GNN) used to process graph-structured data. The main feature of the graph attention network is to aggregate neighbor information by adaptively assigning different weights to different neighbor nodes, which can capture more complex dependencies in the graph. The graph attention network utilizes the multi-head attention mechanism to enhance the model's expressive power, thereby better capturing the interactions between nodes. In the scenario of this application, the input of the graph attention network is the feature matrix and adjacency matrix of users. For user node i, its feature vector is a normalized weekly vector of size 147. The output of the graph attention network is the anomaly detection probability score for each user node, and the anomaly detection probability score for each user is a decimal between 0 and 1, representing the probability of the user being abnormal.
[0166] In the graph attention network, the representation of each node is calculated by taking a weighted average of the features of its neighbor nodes. The weights are automatically generated by the attention mechanism, which takes into account the edges between nodes and the features of the nodes themselves. Specifically, for node i, its representation can be calculated by the following formula:
[0167] (4)
[0168] where, is the representation of node i at the th layer, is the weight matrix, is the representation of neighbor node j at layer l, the set of neighbor nodes of node i, is the activation function; represents the attention weight of node i to node j, and for each node i, all the attention weights are normalized, which can ensure that the sum of the weights is 1.
[0169] Attention weight The normalization calculation method is as follows:
[0170] (5)
[0171] (6)
[0172] Among them, is a single-layer feedforward neural network used to calculate edge ; represents the similarity between node and node ; and represent the weekly power consumption feature vectors of node i and node j, is an activation function.
[0173] In formula (6), the numerator part generates an original attention score by calculating the similarity between the features of node i and its neighbor node j. The denominator part normalizes all neighbor nodes to ensure that the sum of all attention coefficients is 1, which can then be interpreted as a probability distribution.
[0174] After layers of graph attention network, the final representation vectors of all user nodes are obtained. This user vector representation is calculated based on the static adjacency matrix. Next, to capture the time-varying dynamic dependencies in the power system, this application introduces a dynamic adjacency matrix to form a static and dynamic data calculation model. The specific processing method is as follows:
[0175] A dynamic adjacency matrix is constructed for each time window, where represents the th window, and there are a total of windows. For each window, the user power consumption data within the window is used for graph attention network convolution according to formulas (4)-(6) to obtain the feature vector of each user. In this way, B feature vectors of each user on B windows are obtained.
[0176] Next, these B feature vectors are averaged to obtain a comprehensive dynamic feature vector , and its calculation formula is as follows:
[0177] (7)
[0178] Then, the feature vector obtained from the static adjacency matrix and the dynamic feature vector Concatenated together to form a new feature vector , which contains the static and dynamic attributes of a user.
[0179] In this embodiment, the method for anomaly detection using the anomaly detection model is as follows:
[0180] Input the concatenated feature vector into a multi-layer perceptron (MLP), and perform binary classification through the activation function Sigmoid to obtain the final classification probability of the user , and its calculation formula is as follows:
[0181] (8)
[0182] where and are both parameter matrices.
[0183] The activation function Sigmoid maps the output value to the interval [0,1], representing the probability that the user belongs to an abnormal user. If is greater than the set anomaly threshold, and at this time the anomaly threshold can take the value of 0.5, then it is determined that the user is an abnormal user; otherwise, it is determined to be a normal user.
[0184] In this way, the model proposed in this application can not only capture the global stable dependence relationship using the static adjacency matrix, but also capture the time-varying local dependence structure through the dynamic adjacency matrix, thereby improving the accuracy and robustness of anomaly detection. This method enables the model to better adapt to the complex and changing environment in the power system and provide more reliable anomaly detection results.
[0185] In this embodiment, according to the data quality evaluation model, the method for evaluating the power data quality is as follows:
[0186] According to the "Information Technology - Data Quality Evaluation Index" GB / T 36344-2018, four indexes for evaluating the power data quality are established, including integrity, uniqueness, timeliness, and accuracy, and the quantification standards are given. Through these dimensional indexes, the data quality can be evaluated.
[0187] Integrity is divided into data record integrity and time series frequency integrity.
[0188] Data record integrity refers to whether each record (i.e., a row of data) in the dataset contains all necessary information. For example, in a power dataset, each record may include multiple fields such as timestamp, voltage, current, power, etc. If a certain record lacks the data of a necessary field, it will be considered incomplete.
[0189] The frequency integrity of a time series refers to whether the records in a dataset are sampled and recorded at the expected time intervals. If data is missing at certain time points, or the recording time intervals are uneven, this will be regarded as incomplete frequency of the time series. For example, if data should be recorded hourly, but data for some hours is missing, this will affect the data integrity.
[0190] Integrity The calculation formula is as follows:
[0191] (9)
[0192] where is the number of non-empty records meeting the above conditions, and N is the total number of data records.
[0193] Uniqueness is numerical uniqueness, which requires that each record is not repeated on certain key fields (such as timestamp, device identification ID, user identification ID, etc.).
[0194] Uniqueness The calculation formula is as follows:
[0195] (10)
[0196] where is the number of duplicate values of attribute i, N is the total number of data records, and k is the number of key fields in the dataset.
[0197] The timeliness index based on time points calculates timeliness, which is used to measure whether the data is timely and whether it is recorded at the expected time intervals. The timestamps in the data should be monotonically increasing, that is, the time of each record should be later than the time of the previous record. If it is found that the timestamps do not conform to the monotonicity principle, this is considered a timing error.
[0198] Timeliness The calculation formula is as follows:
[0199] (11)
[0200] where represents the number of timing verification error records meeting the above conditions, and N is the total number of data records.
[0201] The accuracy index reflects the correctness and credibility of the data, and it is calculated based on the number of abnormal data.
[0202] Abnormal data is divided into the following types:
[0203] a) Abnormal user data detected by the graph attention network GAT. The remaining grouped data in b, c, and d are the data remaining after removing abnormal users;
[0204] b) Completely invalid data, such as numerical values with obvious errors like -99999;
[0205] c) Data with incorrect formats;
[0206] d) Data with unreasonable numerical content.
[0207] The calculation formula for accuracy Acc is as follows:
[0208] (12)
[0209] where represents the number of abnormal records that meet all the above conditions, and N is the total number of data records.
[0210] In this embodiment, based on the analytic hierarchy process, the method for constructing a power data quality evaluation system is as follows:
[0211] Regarding the four-dimensional indicators of integrity, uniqueness, timeliness, and accuracy as the secondary indicators (benchmark layer) of the evaluation model, the analytic hierarchy process (AHP) is used to determine the weights of the secondary indicators of the data quality model. The power data quality score of the user is set as the total target layer of the AHP. See Figure 2 .
[0212] For the data quality score Q, its calculation formula is as follows:
[0213] (13)
[0214] (14)
[0215] (15)
[0216] where is the weight vector, is the score of each dimension, is the integrity weight value, is the uniqueness weight value, is the timeliness weight value, is the accuracy weight value.
[0217] The scale method Saaty 1-9 is currently the most accurate method verified by experiments for obtaining the AHP judgment matrix. However, during the solution process, consistency verification is required. By introducing the consistency index (CI for short) and the random consistency index (RI for short), the judgment matrix D is tested by judging the consistency ratio (CR for short). The calculation formula is as follows:
[0218] (16)
[0219] (17)
[0220] Among them, is the largest eigenvalue of the judgment matrix D, is the total number of data records, can be obtained by referring to the random consistency index table.
[0221] When CR < 0.1, it indicates that the judgment matrix D meets the consistency verification, and the eigenvector solved from the judgment matrix D is the weight vector of the corresponding index. Otherwise, it is necessary to reconstruct the judgment matrix until it passes.
[0222] In this embodiment, the method for setting the data quality rating is as follows:
[0223] The availability of power data quality determines the reliability of power data mining. Through research, only data rated as good or above can support subsequent data analysis work. Data rated as general or below needs to immediately perform offline verification and inspection on the power acquisition system. Otherwise, low-quality power data cannot play a role in data mining. The classification of power data quality rating is shown in Table 1.
[0224] Table 1: Power Data Quality Rating Table
[0225]
[0226] A specific embodiment of applying the present invention to evaluate the quality of a certain power data:
[0227] The method for evaluating the quality of a certain power data by applying the quality evaluation method of the present invention is as follows:
[0228] Step 1: Obtain the dataset SGCC collected by smart meters of the power grid. This dataset contains the daily power consumption data of 42,372 users. The data records the power consumption of each user per day for approximately 147 weeks from January 2014 to October 2016. In the dataset, a total of 3,615 users are marked as electricity theft users, while the remaining 38,757 are non-electricity theft users, and the ratio of electricity theft users to non-electricity theft users is approximately 1:11. In addition, there is a large amount of missing data in the dataset. The possible reasons include instrument damage or electricity theft behavior. The behavior of such users may affect the stability and reliability of the power system. Then, normalize the power consumption data of each user to obtain standard power data.
[0229] To facilitate obtaining different data quality scores in the end, divide the total dataset into four sub-datasets, denoted as Region A, Region B, Region C, and Region D respectively. The ratios of abnormal to non-abnormal users are: 1:3, 1:5, 1:8, and 1:10 respectively, so as to simulate the actual situations in different regions. Take Region A as an example below, and the experimental methods for the other regions are the same.
[0230] Denote the total dataset of Region A as DataA. The ratio of abnormal users to non-abnormal users is 1:3. Extract 80% of the dataset as the training set, denoted as TrainA, and still maintain the abnormal ratio of 1:3. The remaining data is evenly divided into the validation set and the test set, denoted as ValA and TestA respectively. For each dataset, it contains data information such as user identification ID and daily power consumption. Since the amount of power data in units of days is too large, sum it up according to the time span of one week later, and denote each user node as a vector , and its expression is as follows:
[0231]
[0232] where k is the number of users, i is the number of weeks of power consumption, is the power consumption per week.
[0233] Step 2: The parameter settings of the Graph Attention Network (GAT) are crucial for the accuracy and efficiency of anomaly detection. The relevant parameters include the number of hidden layers, the number of attention heads, the dimension of constructing the dynamic adjacency matrix, the number of iterations, the regularization term, and the loss function. The methods for setting the relevant parameters are as follows:
[0234] The Graph Attention Network model usually contains multiple hidden layers to learn the high-level representation of node features. In the present invention, considering the characteristics of power data and to ensure the generalization ability of the model, 2 hidden layers are set.
[0235] The multi-head attention mechanism allows the model to learn node features from different perspectives. In the present invention, 4 attention heads are set to capture complex patterns in power data.
[0236] To construct a dynamic adjacency matrix, it is necessary to map the data within the window to a high-dimensional space. In the present invention, this dimension is set to 64.
[0237] The learning rate determines the speed at which the model updates its parameters during each iteration. Setting an appropriate learning rate is crucial for preventing oscillations caused by overly fast learning or slow convergence due to overly slow learning. In the present invention, the initial learning rate is 0.005, and a learning rate decay strategy is adopted during training.
[0238] The number of iterations refers to the number of times of repeated training on the entire training set. To ensure that the model fully learns the patterns in the data, the number of iterations set in the present invention is 200 times.
[0239] The regularization term is used to prevent overfitting. In the present invention, an L2 regularization term is set, and its coefficient is 0.0005.
[0240] The loss function is used to measure the gap between the predicted value and the actual value of the model. In the present invention, the cross-entropy loss function is adopted to guide the learning of the model.
[0241] Step 3, the method for anomaly detection using the graph attention network is as follows:
[0242] Step 1, clean and standardize the power data, removing obvious error data such as electricity consumption beyond a reasonable range. At the same time, reasonably fill in the missing data so that the subsequent detection algorithm can run smoothly. In addition, the original data is sampled by day. To adapt to longer-term trend analysis and pattern recognition, it is necessary to convert the data sampled by day into a weekly summary form. This process involves summing up the electricity consumption data within seven days to form the total weekly electricity consumption. Such a conversion helps to smooth short-term fluctuations, reveal more stable long-term trends, and also facilitates the identification of abnormal patterns, because data with a long time span is more likely to reveal regular anomalies.
[0243] Step 2, based on the similarity of electricity consumption patterns among users, construct a graph structure. Each user is regarded as a node, denoted as , and the similarity between users is represented by edge connections. To accurately capture the similarity among users and construct a high-quality graph structure, first normalize the electricity consumption data of each user to ensure that the data of each user is on the same scale, thereby eliminating the influence of dimensionality. The method for normalizing user data is as follows:
[0244] First, find the mean value in the weekly electricity consumption data of this user With variance , and then for any weekly electricity consumption data , normalize it to .
[0245] Next, according to the weekly electricity consumption sequence vectors of each user node, use formula (1) to calculate the cosine similarity between each user and all other users, and put it into the array . Select the top 5 users with the highest similarity to it. These users will become the neighbor nodes of this user, and connect them with edges in the graph. The weight of the edge is set to the magnitude of the cosine similarity to represent the strength of the similarity. The initial feature of the nodes in the graph is the weekly electricity consumption data.
[0246] Up to this point, the construction of the initial graph structure is completed. At the same time, in order to construct a dynamic graph structure, divide the original daily electricity consumption data by each quarter and divide the dataset into multiple windows.
[0247] Step 3, input the user node features, user labels (0 or 1) of the training set TrainA, and the graph composed of cosine similarity into the graph attention network model. The specific method is as follows:
[0248] First, input the weekly electricity consumption data of all users and the constructed initial graph structure into the graph attention network model, and use formulas (4)-(6) to calculate the representations of all users based on the static matrix.
[0249] Next, use formula (3) to construct the dynamic adjacency matrices of multiple windows, and then use formulas (4)-(6) again to calculate the dynamic representations of each user under multiple windows.
[0250] Then, use formula (7) to average and fuse the dynamic representations of each user, and splice them with the static representations to obtain the final representations of all users.
[0251] Finally, use formula (8) to calculate the probability scores of all users, calculate the binary cross-entropy loss function using the true labels of the users in TrainA and their probability scores, and use the gradient descent method to optimize all the parameters mentioned in the above formulas. It should be noted that during the training process, it is also necessary to calculate the loss function of all users on the validation set ValA. If the loss on the validation set does not decrease for multiple consecutive rounds, the training will be terminated.
[0252] Step 4, after the model training is completed, input the test set TestA to judge the model performance. Finally, input the total electricity data DataA of region A. The graph attention network generates node classifications at the output layer. By comparing the differences between the predicted values and the actual values, identify and label abnormal users, count the number of abnormal users, and generate one of the relevant accuracy Acc metrics.
[0253] Step 4: Use the Analytic Hierarchy Process (AHP) to evaluate the data quality. The specific process is as follows:
[0254] According to Formulas (9) to (12), calculate the values of integrity (Com), uniqueness (Unq), timeliness (Tim), and accuracy (Acc) for the dataset DataA respectively. Then, obtain the scores for each dimension according to Formula (15). 。
[0255] The Analytic Hierarchy Process (AHP) is a method for multi-criteria decision-making problems. It evaluates the importance of various factors by constructing a judgment matrix and calculating the weight vector. In this embodiment, AHP is used to comprehensively evaluate the quality of power data. Statistical analysis of the power data for each evaluation dimension is performed to obtain the basis for comparing importance, and the importance of the four indicators is ranked as shown in Table 2.
[0256] Table 2: AHP Weight Matrix for Power Data Quality Evaluation
[0257]
[0258] Thus, the judgment matrix D is obtained, and its expression is as follows:
[0259] (14)
[0260] Then, for the judgment matrix D, first standardize each column of the matrix, sum the standardized elements by row, and finally standardize the sum results. Then, obtain the weight vector according to Formula (14), and its calculation formula is as follows:
[0261]
[0262] Among them, is the weight value of integrity Com, is the weight value of uniqueness Unq, is the weight value of timeliness Tim, is the weight value of accuracy Acc.
[0263] Calculate according to Formulas (16) and (17) to obtain the consistency index CR = 0.07 < 0.1, so the consistency requirement is met.
[0264] Finally, obtain the power data quality score of Region A according to Formula (13). Similarly, the data quality scores of Regions B, C, and D can be obtained. Then, obtain the data quality ratings of all regions according to Table 1, thereby completing the quality evaluation of the power data in all regions.
[0265] According to the above calculation results, the present invention can evaluate the quality of power data through the graph attention network anomaly detection model and the analytic hierarchy process, which can improve the accuracy and robustness of anomaly detection: by introducing the graph attention network, it effectively overcomes the high dependence of traditional neural networks on data quality. By adopting the adaptive weight allocation and multi-head attention mechanism, it can capture complex anomaly patterns in the power system, reduce overfitting, and significantly improve the accuracy and robustness of anomaly detection.
[0266] At the same time, the present invention can effectively reduce human interference and achieve automated detection. The automatic learning mechanism of the graph attention network GAT makes the anomaly detection process no longer rely on artificially set thresholds, thereby reducing subjective interference and the possibility of missed detection and false detection, and improving the objectivity and consistency of the detection results.
[0267] Furthermore, the present invention can effectively enhance the adaptability to complex scenarios. When processing graph-structured data, the graph attention network can efficiently aggregate the information of neighbor nodes. Especially in the power system with a large number of nodes and complex data, it improves the generalization ability of the model, enabling it to better handle the anomaly detection tasks in complex scenarios.
[0268] Further, based on the analytic hierarchy process, the present invention establishes a scientific and systematic data quality evaluation system, provides a scientific method combining qualitative and quantitative aspects, and unifies the evaluation criteria for power data quality. This system is more systematic and reliable, providing theoretical support for the data management of the power system and subsequent in-depth data mining.
[0269] Even further, by improving anomaly detection and data quality evaluation, the present invention guarantees the high quality and reliability of power data, provides strong support for data mining and management of the smart grid, and further improves the security and operation efficiency of the entire power system.
[0270] An equipment embodiment applying the method of the present invention:
[0271] An electronic device, which includes:
[0272] One or more processors;
[0273] A storage device for storing one or more programs;
[0274] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for evaluating the quality of power data based on anomaly detection.
[0275] A computer medium embodiment applying the method of the present invention:
[0276] A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the above-described method for evaluating power data quality based on anomaly detection.
[0277] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.
[0278] The present application is described with reference to the flowcharts and / or block diagrams of the methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows or / and blocks Figure 1 one or more blocks.
[0279] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one or more flows or / and blocks Figure 1 one or more blocks.
[0280] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more flows or / and blocks Figure 1 one or more blocks.
[0281] The model in this application is an object that constitutes an objective descriptive morphological structure with the help of physical or virtual representations. The object is not equal to an object, not limited to physical and virtual. It can be a data processing function, software program, processing mode, usage method, operation mode, work process, application process, electronic hardware, circuit module, processing system, system imitation or simulation object.
[0282] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art can still modify or equivalently replace the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A power data quality evaluation method based on anomaly detection, characterized in that: It includes the following steps: Step 1, through a pre-constructed data preprocessing model, clean the power data to be evaluated to obtain the power consumption data of multiple users; Step 2, use a pre-constructed static and dynamic data calculation model to process the power consumption data of multiple users based on the electricity consumption similarity between users, and construct a static graph structure and a dynamic adjacency matrix; The static graph structure has several nodes and connection edges, which are used to smooth short-term fluctuations and capture long-term trends; The nodes represent users, and the connection edges represent the similarity between users; the method for constructing the static graph structure is as follows: Sum the power consumption data of each user according to a one-week time span to smooth short-term fluctuations and capture long-term trends, and obtain the weekly total power consumption of several users; Normalize the weekly total power consumption of several users to construct multiple user power consumption vectors to ensure that the power consumption data between different users is on the same scale; Based on multiple user power consumption vectors, calculate the included angle value between any two user power consumption vectors to obtain included angle information; According to the included angle information, calculate the cosine similarity between each user and all the remaining users to obtain similarity data; Based on the similarity data, select the top M users with the highest cosine similarity as the neighbor nodes of the user, and use the cosine similarity as the weight of the connection edge to represent the strength of the similarity; Use the connection edge to connect the user and the neighbor nodes to obtain the static initial graph structure of the user; Stitch the static initial graph structures of all users to obtain a complete graph structure; Use the user's weekly power consumption data as the feature vector to establish a feature matrix and a static adjacency matrix; Based on the feature matrix and the static adjacency matrix, assign values to the complete graph structure to obtain the static graph structure; The dynamic adjacency matrix is used to flexibly adapt to the time-varying local dependence structure, capture short-term dynamic changes, and provide a variable interaction picture; the method for constructing the dynamic adjacency matrix is as follows: According to the power consumption characteristics of different users, set different sampling time lengths; The sampling time length is half a month, one month, or one quarter; According to the sampling time length, divide the power consumption data of the users to obtain a window length, that is, the user power consumption data within a time period; According to the window length, represent the daily power consumption data of each user as a vector to obtain a power consumption matrix regarding the window length; According to the power consumption matrix and based on the electricity consumption similarity between users, set an adjacency matrix regarding the window length to obtain the dynamic adjacency matrix between users; Step 3, use a pre-constructed anomaly detection model to process the static graph structure and the dynamic adjacency matrix, capture the similarity between users, and identify abnormal user data; the method is as follows: Obtain the static graph structure, which includes a feature matrix and a static adjacency matrix; According to the feature matrix and the static adjacency matrix, obtain the feature vector of a certain node and the feature vectors of the neighbor nodes; Using the trained graph attention network, perform weighted averaging on the feature vector of a certain node and the feature vectors of its neighbor nodes to obtain the static feature vector of the certain node; To capture the time-varying dynamic dependencies in the power system, based on the dynamic adjacency matrix, obtain the power consumption matrix corresponding to the window length; Using the graph attention network, perform graph attention network convolution on the power consumption matrix to obtain several feature vectors for each user; Average the several feature vectors to obtain a comprehensive dynamic feature vector: Concatenate the static feature vector and the dynamic feature vector together to form a new concatenated feature vector, which integrates the static and dynamic attributes of the user; Input the concatenated feature vector into a multi-layer perceptron and perform binary classification through an activation function to obtain the final classification probability of the user; The anomaly detection probability score for each user is a decimal between 0 and 1, indicating the probability of the user being abnormal; If the final classification probability is greater than the set anomaly threshold, determine that the user is an abnormal user and count the number of abnormal users; otherwise, determine as a normal user; Step four, use the pre-constructed data quality evaluation model, based on the abnormal user data, evaluate the power data to obtain the quality evaluation result, and complete the power data quality evaluation based on anomaly detection.
2. A power data quality evaluation method based on anomaly detection as described in claim 1, characterized in that: Step one, through the pre-constructed data preprocessing model, clean the power data to be evaluated to obtain the power consumption data of multiple users, and the method is as follows: Obtain the power data to be evaluated, and the power data is data involving multiple users sampled by day; Clean and standardize the power data, remove the power consumption beyond the set range to obtain the power standard data; Fill in the missing data in the power standard data to obtain the power complete data; Split the power complete data according to user attributes to obtain the power consumption data of multiple users.
3. A power data quality evaluation method based on anomaly detection as described in claim 1, characterized in that: Step four, use the pre-constructed data quality evaluation model, based on the abnormal user data, evaluate the power data to obtain the quality evaluation result, and the method is as follows: Select several indicators for power data quality evaluation, including integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator; Adopt the analytic hierarchy process, and construct the corresponding weight matrix according to the influence degrees of the integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator on the power data quality; Based on the weight matrix, set the judgment matrix; Standardize each column of the judgment matrix, sum the standardized elements by row, and finally standardize the summation result to obtain the weight vector; Calculate the specific values of the integrity indicator, uniqueness indicator, timeliness indicator, and accuracy indicator to obtain the dimension scores of the four indicators; Multiply the dimension scores of the four indicators by the weight vector and calculate to obtain the power data quality score; According to the power data quality rating form, rate the power data quality score to obtain the quality evaluation result.
4. A power data quality evaluation method based on anomaly detection as claimed in claim 3, characterized in that: The method for calculating the specific value of the integrity index is as follows: Obtain the power consumption data of multiple users and get the total number of data records; Based on the data record integrity judgment mechanism, judge whether each record in the power consumption data contains the necessary field information, and obtain the number of the first complete records; The data record integrity judgment mechanism refers to whether each record in the data set contains the necessary field information; if a certain record lacks a certain necessary field information, it is considered an incomplete data record; the necessary field information includes timestamp, voltage, current and power; Based on the time series frequency integrity judgment mechanism, judge whether the records in the power consumption data are sampled and recorded at the expected time interval, and obtain the number of the second complete records; The time series frequency integrity judgment mechanism refers to whether the records in the data set are sampled and recorded at the expected time interval; if the data at some time points is missing, or the recorded time interval is uneven, it will be regarded as the frequency of the time series is incomplete; Add the number of the first complete records and the number of the second complete records to obtain the number of non-empty records that meet the conditions; Divide the number of non-empty records by the total number of data records to obtain the specific value of the integrity index; Or / and, the method for calculating the specific value of the uniqueness index is as follows: Obtain the power consumption data of multiple users and get the total number of data records; Based on the numerical uniqueness judgment mechanism, judge whether each record in the power consumption data is unique in some key fields, and obtain the number of data uniqueness; The numerical uniqueness judgment mechanism means that each record is not repeated in some key fields; some key fields include timestamp, device identifier and user identifier; Based on the number of data uniqueness and the total number of data records, calculate the specific value of the uniqueness index; Or / and, the method for calculating the specific value of the timeliness index is as follows: Obtain the power consumption data of multiple users and get the total number of data records; Based on the timeliness judgment mechanism, judge whether the timestamp of each record in the power consumption data increases monotonically, and obtain the number of timeliness errors; The timeliness judgment mechanism is used to measure whether the data is recorded in a timely manner and at the expected time interval, that is, the time of each record needs to be later than the time of the previous record; If it is found that the timestamp does not conform to the monotonicity principle, it is considered a timing error; Based on the total number of data records and the number of timeliness errors, calculate the specific value of the timeliness index; Or / and, the method for calculating the specific value of the accuracy index is as follows: Obtain the number of abnormal users and the power consumption data of multiple users, and get the total number of data records; Based on the accuracy index judgment mechanism, find the number of abnormal data; The abnormal data includes invalid data, format error data, and content unreasonable data; Add the number of abnormal data and the number of abnormal users to obtain the number of abnormal records; Based on the number of abnormal records and the total number of data records, obtain the specific value of the accuracy index.
5. A method for evaluating power data quality based on anomaly detection according to claim 4, characterized in that: Based on the weight matrix, the method for setting the judgment matrix is as follows: Step 11, based on the weight matrix, construct an initial judgment matrix for integrity index, uniqueness index, timeliness index and accuracy index; Step 12, calculate the maximum eigenvalue of the initial judgment matrix; Step 13, based on the maximum eigenvalue of the initial judgment matrix, calculate the consistency index; And obtain the random consistency index by referring to the random consistency index table; Step 14, compare the consistency index and the random consistency index to obtain the consistency ratio; Step 15, compare the consistency ratio with the ratio threshold, specifically including the following: When the consistency ratio is less than the ratio threshold, it indicates that the initial judgment matrix meets the consistency check, and execute step 16; When the consistency ratio is greater than or equal to the ratio threshold, reconstruct the initial judgment matrix and execute step 12; Step 16, use the initial judgment matrix as the final judgment matrix.
6. A power data quality evaluation system based on anomaly detection, characterized in that: It includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a method for evaluating power data quality based on anomaly detection according to any one of claims 1-5.
Citation Information
Patent Citations
Power data anomaly detection method, system, device and medium based on graph computing
CN118378023B
Multivariable time series data anomaly detection method and system based on dynamic graph learning and long and short term convolution
CN117251731A
Semiconductor wafer manufacturing anomaly detection method and device based on space-time diagram neural network
CN118411589A