Power system abnormal data recovery method and system considering time and accuracy priority

By combining the Analytic Hierarchy Process (AHP) and weight matrix scoring with the Spatiotemporal Convolutional Network (STCN) and the Low-Rank Matrix Factorization (LMF) algorithm, the problem of high-precision and fast response in abnormal data recovery in power systems was solved, enabling stable operation and real-time decision support for power systems.

CN119577333BActive Publication Date: 2026-04-07ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID NINGXIA ELECTRIC POWER COMPANY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously meet the demands for high precision and rapid response in power system anomaly data recovery, especially in the long-term planning and business management of power management data. The low accuracy of existing methods affects the reliability of analysis results and the accuracy of decision-making.

Method used

The weight matrix is ​​determined by the analytic hierarchy process (AHP). Based on the comprehensive score of security protection level and response recovery level, a high-precision recovery algorithm or a fast recovery algorithm is selected. The spatiotemporal convolutional network STCN and the low-rank matrix factorization algorithm are used to recover data to ensure the stable operation of the power system and real-time decision-making.

Benefits of technology

It achieves an efficient balance between the accuracy and response speed of data recovery in the power system, provides reliable data support, and ensures the stable operation of the power system and the accuracy of real-time decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The application provides a power system abnormal data recovery method and system considering time and precision priority, and belongs to the technical field of power data security. The method comprises the following steps: obtaining power data to determine whether there is an abnormal value; obtaining the quantized value of each abnormal value in two dimensions based on a security protection level index and a response recovery level index to form a row vector, and further establishing a decision index matrix; determining a weight matrix by using an analytic hierarchy process; calculating the comprehensive score of each abnormal value based on the decision index matrix and the weight matrix; determining a priority label according to the comprehensive score and a threshold value, performing label processing on each abnormal value in the power data set, and obtaining a data set with a priority label; recovering each abnormal value in the data set by using a recovery algorithm corresponding to the priority label to obtain a normal data set; the recovery algorithm comprises a high-precision recovery algorithm and a fast recovery algorithm; and the normal data set is sent to a data control center to provide a data source for subsequent real-time decision analysis of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of power data security, in particular to a power system abnormal data recovery method and system considering time and precision priority. BACKGROUND

[0002] As an indispensable infrastructure in modern society, the power system supports the normal operation of the economic society. With the continuous expansion and increase in complexity of the power system, the importance of power data in business forecasting, power generation scheduling and system operation is increasingly prominent. The network architecture of the new type of power system can be divided into two categories according to the type of business and the security level: a network for production data transmission and a network for management data transmission.

[0003] In the power system, historical data is often used as the basis for analysis to conduct real-time monitoring or prediction of the future. The length of the time interval of the historical data involved can be seconds, minutes, hours or days. For some fine monitoring, even several time windows of data within 1 second are extracted for analysis. Therefore, data delay or loss during transmission may directly affect the safe and stable operation of the system, leading to large prediction errors and inaccurate monitoring information. Therefore, abnormal data recovery is a common operation in the power system, and the requirements for data recovery operations are also very high. The existing technology proposes to use linear interpolation method for data recovery. This method responds quickly and can ensure the continuity and safety of the system, but the accuracy is not high. It is more suitable for real-time operation and control of the power system, such as scheduling automation, equipment monitoring and fault handling, which require high data recovery capability. However, power management data is mainly used for long-term planning, data analysis and business management, etc. These data play a crucial role in power system planning, decision support and optimal scheduling, and the data accuracy directly affects the reliability of the analysis results and the accuracy of the decision. The existing technical solution cannot support the high precision requirement of power management data for data recovery. SUMMARY

[0004] Therefore, the application provides a power system abnormal data recovery method and system considering time and precision priority. According to the demand characteristics of different types of power data recovery, the corresponding recovery method is adopted, which meets the accuracy and response speed requirements in the process of power system abnormal data recovery, and provides reliable data support for the stable operation and real-time decision of the power system.

[0005] The technical solution adopted by the embodiment of the application to solve the technical problems is:

[0006] A power system abnormal data recovery method considering time and precision priority, comprising:

[0007] Step S1: Obtain the power dataset and determine if there are any outliers; the outliers include remote control information exchanged between the data center SCADA-EMS system and the power plant, data information exchanged between the data center SCADA and EMS systems, power metering information, and other power market information, including generation cost, market clearing price, transaction volume, and user load;

[0008] Step S2: Based on the security protection level index and the response recovery level index, quantified values ​​of each outlier in two dimensions are used to form a row vector A. ij Furthermore, a decision index matrix A is established for all outliers, where A ij middle element a ij This represents the quantified value of the i-th outlier under the j-th level index;

[0009] Step S3: Use the analytic hierarchy process (AHP) to determine the weight matrix W = [w1, w2]. T This includes the security protection level weight w1 and the response recovery level weight w2;

[0010] Step S4: Calculate the comprehensive score Z for each outlier based on the decision index matrix A and the weight vector W. i ;

[0011] Step S5, based on the comprehensive score Z i And the priority label is determined by the threshold Z′. i Each outlier in the power dataset is labeled to obtain a dataset with priority labels; the priority labels correspond to the outlier recovery algorithm.

[0012] Step S6: The outliers in the dataset with priority labels are recovered using the recovery algorithm corresponding to the priority labels to obtain the normal dataset after data recovery. The recovery algorithm includes a high-precision recovery algorithm and a fast recovery algorithm. The fast recovery algorithm uses a low-rank matrix factorization algorithm for data recovery, and the high-precision recovery algorithm uses a spatiotemporal convolutional network STCN for data recovery. The spatiotemporal convolutional network STCN is composed of a temporal convolutional network TCN, a graph convolutional network GCN, and a fully connected layer.

[0013] Step S7: Send the normal dataset to the data control center to provide a data source for subsequent power system analysis.

[0014] The preferred quantitative standard for the security protection level index is as follows: when the data type of the outlier belongs to the remote control information exchanged between the data center SCADA / EMS system and the power plant, the quantitative value is 4; when it belongs to the data information exchanged between the data center SCADA / EMS system, the quantitative value is 3; when it belongs to the power energy metering information, the quantitative value is 2; and when it belongs to other information in the power market, the quantitative value is 1.

[0015] The quantification standard for the response recovery level index is as follows: the quantification value is 5 when the data recovery response time requirement for the outlier is immediate; the quantification value is 4 when the data recovery response time requirement is less than 1 second; the quantification value is 3 when the data recovery response time requirement is less than 1 minute; the quantification value is 2 when the data recovery response time requirement is less than 5 minutes but greater than 1 minute; and the quantification value is 1 when the data recovery response time requirement is greater than 5 minutes.

[0016] Preferably, step S2 includes:

[0017] Step S21: Based on the security protection level index, derive the quantified value a of the i-th outlier in the security protection level index dimension. i1 Based on the response recovery level index, the quantified value 'a' of the i-th outlier in the response recovery level index dimension is obtained. i2 ;

[0018] Step S22, let row vector A ij =[a i1 ,a i2 The decision index matrix A is formed by combining the i row vectors.

[0019]

[0020] Where i∈[1,n];

[0021] Step S23: Standardize matrix A to obtain decision index matrix A′:

[0022]

[0023] In the formula, A j Min(A) represents the column vectors in matrix A; j ) and max(A j ) represent the minimum and maximum values ​​in column j, respectively.

[0024] Preferably, step S3 includes:

[0025] Construct the judgment matrix P:

[0026]

[0027] Where, p 11 This indicates that the importance of the safety protection indicator relative to itself is 1, p 12 This indicates the importance of security protection relative to response and recovery, p 12 Let p ∈[1,9] 21 =1 / p 12 p 22 =1, p 21 This indicates the importance of response recovery relative to security protection, p 22 This indicates that the importance of the response recovery relative to itself is 1;

[0028] Perform column normalization on each element of the judgment matrix P to obtain the normalized judgment matrix P′:

[0029]

[0030] Among them, element p′ ij The calculation method is as follows:

[0031]

[0032] Average each row of the judgment matrix P′:

[0033]

[0034] We obtain the security protection level weight w1 and the response recovery level weight w2, as well as the weight matrix W = [w1, w2]. T .

[0035] Preferably, step S4 includes:

[0036] The comprehensive score Z for each outlier is calculated by multiplying the decision index matrix A′ with the weight matrix W. i :

[0037]

[0038] Among them, Z i Indicates the first i A comprehensive score for each outlier.

[0039] Preferably, the priority label in step S5 i The generation rules are as follows:

[0040]

[0041] Among them, Priority i A value of 1 indicates that the fast recovery algorithm is selected, Priorityi A value of 0 indicates that a high-precision recovery algorithm is selected.

[0042] Preferably, step S6 includes:

[0043] For the priority label Priority i Outliers with a value of 1 can be quickly recovered using a low-rank matrix factorization algorithm.

[0044] Let data matrix Y represent the power dataset, and establish a mask matrix W for data matrix Y, where the elements w in the mask matrix W are... ij This represents the element y in the data matrix Y. ij The validity of is expressed as:

[0045]

[0046] Establish low-rank matrices U and V, and adjust the reconstruction matrix X = UV using an optimization objective function to make the reconstruction matrix X approximate the effective data in Y; the optimization objective function is:

[0047]

[0048] In the formula, w ij (Y ij -(UV) ij ) 2 This indicates that the error is calculated only at valid data locations; λ represents the regularization term; λ represents the regularization parameter.

[0049] By introducing auxiliary variables Z and the Lagrange multiplier matrix Λ, the objective function is decomposed into two subproblems using the Alternating Direction Multiplier Method (ADMM), which optimizes X and Z respectively. X, Z, and Λ are then iteratively updated.

[0050] First, fix Z and Λ, optimize the value of X, and simultaneously consider the current values ​​of Z and Λ. The iterative formula for X is:

[0051]

[0052] In the formula, k is the number of iterations, and ρ is the penalty parameter.

[0053] Then optimize the auxiliary variable Z. The optimization formula for Z is:

[0054]

[0055] The Lagrange multiplier matrix Λ is updated using the following formula:

[0056] Λ (k+1) =Λ k +ρ(X (k+1) -Z(k+1) )

[0057] Calculate the error between the valid data locations in the reconstruction matrices X and Y ||(X)|| (k+1) -Y)⊙W‖ F When the condition is satisfied ||(X) (k+1) -Y)⊙W‖ F The iteration stops when the threshold value is less than ε, where ε is a preset threshold.

[0058] Using the final reconstructed matrix X=UV, a low-rank inference method is employed to fill in missing data and repair anomalous data based on valid data in Y.

[0059] Furthermore,

[0060] For the priority label Priority i Outliers with values ​​of 0 are recovered with high precision using the Spatiotemporal Convolutional Network (STCN).

[0061] Assume the input time series of the power data in the power dataset is Q = {q1, q2, q3, ..., q}. t}, where q t This represents the power data collected at time t;

[0062] The convolution operation of TCN is defined as follows:

[0063]

[0064] In the formula, z t This represents the temporal features extracted by TCN at time t; w i d represents the weights of the convolution kernel; k is the size of the convolution kernel, representing the number of time points involved in a single convolution operation; i It is the expansion rate; b is the bias term; t > d k-1 After undergoing dilated convolution operations across multiple TCN layers, the output feature sequence Z is obtained. t ={z t1 ,z t2 ,z t3 ,...,z tn};

[0065] Using Z t Further construct the initial feature matrix H (0) =[z t1 ,z t2 ,z t3 ,...,z tn As input to the Graph Convolutional Network (GCN);

[0066] Spatial features of the data are extracted using a Graph Convolutional Network (GCN). The topological structure of the power system is abstractly represented as a graph G = (V, E), where V is the set of nodes representing measurement points in the power system; E is the set of edges representing the physical connections between measurement points; the adjacency matrix A describes the connection relationships between nodes: if nodes are connected, A = 1; otherwise, A = 0; the degree matrix D represents the number of connections for each node; the formula for calculating the graph convolution of the (l+1)th layer is:

[0067]

[0068] In the formula, H (l+1) Let H be the node feature matrix of the (l+1)th layer, initially... (0) The temporal feature matrix extracted by TCN; W (l) Here, σ is the weight matrix of the l-th layer; σ is the activation function; through layer-by-layer graph convolution, node features are propagated and fused in the topology of the power system, generating a node feature matrix H containing spatiotemporal information. (L) The final output matrix H of the graph convolutional network contains the temporal and spatial features of each node.

[0069] Finally, the final output matrix H is mapped to the final recovery result of the power data through a fully connected layer; the calculation formula for the fully connected layer is:

[0070] R = HW h +b h

[0071] In the formula, R is the recovery matrix, which contains the complete estimation results for all nodes in the power system; W h This is the weight matrix of the fully connected layer, used to project the feature matrix H onto the output space of the data reconstruction; b h It is a bias term that adjusts the output result.

[0072] As can be seen from the above technical solution, the power system abnormal data recovery method and system that prioritizes time and accuracy provided in this invention first acquires power data to determine whether outliers exist; based on security protection level indicators and response recovery level indicators, it derives row vectors composed of the quantified values ​​of each outlier in two dimensions, and further establishes a decision indicator matrix; it uses the analytic hierarchy process (AHP) to determine the weight matrix; it calculates the comprehensive score of each outlier based on the decision indicator matrix and the weight matrix; it determines priority labels based on the comprehensive score and threshold, and performs label processing on each outlier in the power data set to obtain a dataset with priority labels; it uses the recovery algorithm corresponding to the priority labels to recover each outlier in the dataset to obtain a normal dataset; the recovery algorithm includes a high-precision recovery algorithm and a fast recovery algorithm; and it sends the normal dataset to the data control center to provide a data source for subsequent real-time decision analysis of the power system. This invention, by decomposing the data matrix, quickly reconstructs abnormal data, enabling effective data recovery in a short time. The recovered data is sent to the data control center to support further analysis and decision-making in the power system. This invention effectively balances accuracy and response speed in the power system abnormal data recovery process, providing reliable data support for the stable operation and real-time decision-making of the power system. Attached Figure Description

[0073] Figure 1 This is a flowchart illustrating the power system anomaly data recovery method that prioritizes time and accuracy, as described in this invention.

[0074] Figure 2 This is a structural diagram of a power system anomaly data recovery system that prioritizes time and accuracy, as presented in this invention. Detailed Implementation

[0075] The technical solution and effects of the present invention will be further described in detail below with reference to the accompanying drawings.

[0076] Power data typically includes production data and management data. Production data transmission networks are primarily used for operations directly related to the real-time operation and control of the power system, such as dispatch automation, equipment monitoring, and fault handling. These networks have high requirements for data transmission recovery speed, as data delays or loss can directly affect the safe and stable operation of the system. To ensure the system can quickly return to normal operation in the event of data anomalies or transmission interruptions, data recovery in production data transmission networks must prioritize rapid response, minimizing recovery time to rebuild a reliable data stream in the shortest possible time, ensuring system continuity and security. In contrast, management data transmission networks are mainly used for long-term planning, data analysis, and business management. The data in these networks often involves applications such as load forecasting, energy efficiency assessment, and historical data analysis, with lower requirements for data recovery speed. However, this data plays a crucial role in power system planning, decision support, and optimized dispatch, and its accuracy directly affects the reliability of analysis results and the accuracy of decisions. Therefore, for management data transmission networks, data recovery must prioritize data integrity and accuracy. This means that during the recovery process, meticulous processing and multiple verifications must be used to ensure that the recovered data meets high-precision requirements, supporting long-term data analysis and planning. Therefore, the security protection requirements for data in the two types of networks differ significantly, and their respective characteristics also place different demands on data recovery technologies.

[0077] This invention proposes a strategic data recovery method, employing a power system anomaly data recovery technology architecture that prioritizes time and accuracy. First, a priority assessment module quantifies the data's security and real-time requirements based on its security protection level and response recovery level, generating a comprehensive score and determining priority labels. The data recovery module selects appropriate recovery algorithms based on these priorities. The high-precision recovery process utilizes Temporal Convolutional Networks (TCNs) and Graph Convolutional Networks (GCNs) to extract spatiotemporal features, while the fast response process employs low-rank matrix factorization and ADMM optimization to achieve rapid recovery and output complete results.

[0078] refer to Figure 1 and Figure 2 As shown, the power system abnormal data recovery method that prioritizes time and accuracy provided by this invention is implemented by [the following entity / entity]. Figure 2 The specific implementation of the data recovery method in the aforementioned system includes:

[0079] Step S1: Obtain the power data set and determine if there are any outliers. Outliers include remote control information exchanged between the data center SCADA-EMS system and the power plant, data information exchanged between the data center SCADA and EMS systems, power metering information, and other power market information, including generation costs, market clearing prices, trading volume, and user load. SCADA (Supervisory Control and Data Acquisition) and EMS (Energy Management System) are usually used together to jointly undertake power grid regulation and management.

[0080] Step S2, based on the security protection level index U r and response recovery level index U p The quantized values ​​of each outlier in two dimensions are used to derive a row vector A. ij Furthermore, a decision index matrix A is established for all outliers, where A ij middle element a ij This represents the quantified value of the i-th outlier under the j-th level index;

[0081] Step S3: Use the analytic hierarchy process (AHP) to determine the weight matrix W = [w1, w2]. T This includes the security protection level weight w1 and the response recovery level weight w2;

[0082] Step S4: Calculate the comprehensive score Z for each outlier based on the decision index matrix A and the weight vector W. i ;

[0083] Step S5, based on the comprehensive score Z i And the priority label is determined by the threshold Z′. i Labeling is performed on each outlier in the power dataset to obtain a dataset with priority labels; the priority labels correspond to the outlier recovery algorithm.

[0084] Step S6: Use the recovery algorithm corresponding to the priority label to recover each outlier in the dataset with priority label, and obtain the normal dataset after data recovery. The recovery algorithm includes a high-precision recovery algorithm and a fast recovery algorithm. The fast recovery algorithm uses a low-rank matrix factorization algorithm for data recovery, and the high-precision recovery algorithm uses a spatiotemporal convolutional network STCN for data recovery. The spatiotemporal convolutional network STCN is composed of a temporal convolutional network TCN, a graph convolutional network GCN, and a fully connected layer.

[0085] Step S7: Send the normal dataset to the data control center to provide a data source for subsequent power system analysis.

[0086] The quantitative standard for the better security protection level index is as follows: when the data type of the outlier belongs to the remote control information exchanged between the data center SCADA / EMS system and the power plant, the quantitative value is 4; when it belongs to the data information exchanged between the data center SCADA / EMS system, the quantitative value is 3; when it belongs to the power energy metering information, the quantitative value is 2; and when it belongs to other information in the power market, the quantitative value is 1.

[0087] The quantitative standards for response recovery level indicators are as follows: for data recovery of outliers, the allowable response time requirement is 5 when responding immediately, 4 when the allowable response time requirement is less than 1 second, 3 when the allowable response time requirement is less than 1 minute, 2 when the allowable response time requirement is less than 5 minutes but greater than 1 minute, and 1 when the allowable response time requirement is greater than 5 minutes.

[0088] Table 1 Safety Protection Level Indicators

[0089]

[0090] Table 2 Recovery Response Level Indicators

[0091]

[0092]

[0093] One possible implementation of step S2 is as follows:

[0094] Step S21: Based on the security protection level index, derive the quantified value a of the i-th outlier in the security protection level index dimension. i1 Based on the response recovery level index, the quantified value 'a' of the i-th outlier in the response recovery level index dimension is obtained. i2 ;

[0095] Step S22, let row vector A ij =[a i1 ,a i2 Let i row vectors be used to form a decision index matrix A:

[0096]

[0097] Where i∈[1,n]; this matrix can be used to centrally represent the scores of each data point on these two decision indicators;

[0098] Step S23: To ensure the comparability of different indicators in the calculation, matrix A is standardized to obtain the decision indicator matrix A′:

[0099]

[0100] In the formula, A j Min(A) represents the column vectors in matrix A; j ) and max(A j ) represent the minimum and maximum values ​​in column j, respectively. The purpose of standardization is to map the scores of different indicators to the same numerical range (usually 0 to 1), eliminating the dimensional differences between different indicators. For example, the security protection score may be between 1 and 5, while the response recovery score may be between 1 and 10.

[0101] After standardization, the Analytic Hierarchy Process (AHP) is used to determine the weight of each indicator, namely, the security protection level weight w1 and the response recovery level weight w2. AHP is a structured decision-making method that compares the importance of each indicator by constructing a judgment matrix, thereby obtaining a reasonable weight allocation. An optional implementation of step S3 is as follows:

[0102] Construct a judgment matrix P based on expert opinions or experience:

[0103]

[0104] Where, p 11 This indicates that the importance of the safety protection indicator relative to itself is 1, p 12 This indicates the importance of security protection relative to response and recovery, p 12 ∈[1,9], the larger the value, the more important the security protection. Let p 21 =1 / p 12 p 22 =1, p 21 This indicates the importance of response recovery relative to security protection, p 22 This indicates that the importance of the response recovery relative to itself is 1;

[0105] Column normalization is performed on each element of the judgment matrix P to calculate the relative weight of each indicator, resulting in the normalized judgment matrix P′:

[0106]

[0107] Among them, element p′ ij The calculation method is as follows:

[0108]

[0109] Here, let n = 2;

[0110] The weight of each indicator is obtained by averaging each row of the judgment matrix P′:

[0111]

[0112] We obtain the security protection level weight w1 and the response recovery level weight w2, as well as the weight matrix W = [w1, w2]. T .

[0113] The specific implementation of step S4 includes: calculating the comprehensive score Z of each outlier by multiplying the decision index matrix A′ with the weight matrix W. i :

[0114]

[0115] Among them, Z i Z represents the overall score of the i-th outlier, and the overall score Z provides the basis for the selection of subsequent recovery algorithms.

[0116] Priority label in step S5 i The generation rules are as follows:

[0117]

[0118] Among them, Priority i A value of 1 indicates that the fast recovery algorithm is selected, Priority i A value of 0 indicates that a high-precision recovery algorithm is selected.

[0119] Finally, the anomaly dataset with priority labels is exported, generating a large anomaly dataset with clear priority indicators. This dataset provides clear decision support for the system in the subsequent anomaly data recovery process, thereby ensuring the real-time nature and accuracy of data recovery.

[0120] In data recovery, different algorithms affect the accuracy and speed of the recovery. Temporal Convolutional Networks (TCNs) are deep learning methods specifically designed for processing time-series data, capable of capturing long-term dependencies and complex temporal patterns. Building upon this, Graph Convolutional Networks (GCNs) further model the spatial dependencies of the data, making them particularly suitable for data with topological structures. GCNs utilize the connections between nodes, capturing these dependencies through graph convolution operations, further improving the accuracy of the recovery. Combining TCNs and GCNs can simultaneously enhance the accuracy and spatial consistency of time-series data, but it also introduces higher computational complexity and longer processing time. Therefore, while this combination can achieve extremely high recovery accuracy, its time overhead is significant.

[0121] In contrast, matrix factorization is an optimization method based on linear algebra that approximates the recovery of missing or outlier data by decomposing the data matrix into low-rank forms. Matrix factorization has lower computational complexity and converges quickly through iterative optimization, thus enabling data recovery in a shorter time. Although low-rank matrix factorization may be slightly less accurate than deep learning models, its fast computational characteristics make it very suitable for scenarios with high time requirements for data recovery.

[0122] The specific implementation of step S6 includes:

[0123] For the priority label Priority i Outliers with a value of 1 can be quickly recovered using a low-rank matrix factorization algorithm.

[0124] Let data matrix Y represent the original data (including outliers and missing values) of the power dataset. Construct a mask matrix W with respect to data matrix Y, where each element w in the mask matrix W represents a specific value. ij Represents the elements y in the data matrix Y ij The validity of is expressed as:

[0125] This information may contain anomalies or be missing. To handle missing and anomalous data, a mask matrix W is introduced to mark the validity of the data:

[0126]

[0127] The mask matrix W ensures that the objective function focuses only on valid data during decomposition, ignoring missing or outlier data. The method for determining whether data is normal is described in existing technical solutions and will not be elaborated upon here.

[0128] Establish low-rank matrices U and V, and adjust the reconstructed matrix X = UV using the optimization objective function to make the reconstructed matrix X approximate the effective data in Y; the optimization objective function is:

[0129]

[0130] In the formula, w ij (Y ij -(UV) ij ) 2 This indicates that data is only valid at the location of the data (i.e., W). ij =1 position) Calculate the error, limit the size of U and V, and avoid overfitting; λ represents the regularization term; λ represents the regularization parameter, which controls the weight of the regularization term.

[0131] To simplify the solution process, auxiliary variables Z and the Lagrange multiplier matrix Λ are introduced. The objective function is decomposed into two subproblems using the Alternating Direction Multiplier Method (ADMM), which optimizes X and Z respectively. X, Z, and Λ are then iteratively updated.

[0132] First, fix Z and Λ, optimize the value of X, and simultaneously consider the current values ​​of Z and Λ. The iterative formula for X is:

[0133]

[0134] In the formula, k is the number of iterations, and ρ is a penalty parameter used to control the consistency between X and Z;

[0135] Then optimize the auxiliary variable Z to make Z more consistent with the current X. The optimization formula for Z is:

[0136]

[0137] g(Z) makes Z better satisfy the current optimization conditions;

[0138] Update the Lagrange multiplier matrix Λ to satisfy the asymptotic convergence of the constraints. The update formula for Λ is:

[0139] Λ (k+1) =Λ k +ρ(X (k+1) -Z (k+1) (17)

[0140] Through multiple iterations, X and Z will gradually converge, and the final reconstructed matrix X = UV will approximate the effective data portion of the original matrix Y, filling in missing data and repairing outlier data. In the final reconstructed matrix X, the missing and outlier data positions will be appropriately filled, with the filling values ​​inferred from the low-rank of the effective data in Y.

[0141] Calculate the error between the valid data positions in the reconstructed matrices X and Y ||(X) (k+1) -Y)⊙W∥ F When the condition is satisfied that ∥(X) (k +1) -Y)⊙W∥ F Iteration stops when the value is less than ε, where ε is a preset threshold; ⊙ represents element-wise multiplication; W is a mask matrix, containing only valid data positions (w... ij =1) Able to calculate errors;

[0142] Using the final reconstructed matrix X=UV, a low-rank inference method is employed to fill in missing data and repair anomalous data based on valid data in Y.

[0143] Due to its unique characteristics, power system data exhibits significant temporal and spatial variations. It not only changes systematically over time but is also influenced by the interactions between nodes in different geographical locations and with different physical connections within the system. This characteristic necessitates the use of deep learning structures capable of simultaneously capturing the features of data in both temporal and spatial dimensions to improve the accuracy of recovering from anomalies. Therefore, this invention introduces a Temporal Convolutional Network (TCN) to capture the temporal dependencies of the data and combines it with a Graph Convolutional Network (GCN) to extract the structured features of the data in the spatial dimension.

[0144] For the priority label Priority i Outliers with values ​​of 0 are recovered with high precision using a spatiotemporal convolutional network (STCN). First, the time-series data is processed using TCN to capture its long-term dependencies and extract key temporal features. TCN then expands the receptive field through dilated convolution operations, ensuring the modeling of long-term patterns in the data.

[0145] Suppose the input time series of the power data in the power dataset is Q = {q1, q2, q3, ..., q t}, where q t This represents the power data collected at time t;

[0146] The convolution operation of TCN is defined as follows:

[0147]

[0148] In the formula, z t This represents the temporal features extracted by TCN at time t, used to capture the temporal dependencies of the input sequence; w i is the weight of the convolution kernel, controlling the contribution of each input point in the convolution; k is the size of the convolution kernel, representing the number of time points involved in a single convolution operation; d i The dilation rate controls the stride of the convolutional kernel in the time series, allowing it to skip several time points to capture dependencies over a longer time span; b is the bias term, adjusting the output of the convolution; the above formula compares the input data at the current time t with the input data at the previous time td. i Perform convolution operations on the data; t > d k-1 After undergoing dilated convolution operations across multiple TCN layers, the output feature sequence Z is obtained. t ={z t1 ,z t2 ,z t3 ,...,z tn} can comprehensively characterize the temporal properties of the input time series; let the set Represents a set composed of elements selected from set Q. Taken from a set

[0149] Using Z t Further construct the initial feature matrix H (0) =[z t1 ,z t2 ,z t3 ,...,z tn As input to the Graph Convolutional Network (GCN);

[0150] Graph Convolutional Networks (GCNs) are used to extract spatial features from data to reflect the physical connections between nodes in a power system. The topology of a power system is abstractly represented as a graph G = (V, E), where V is the set of nodes, representing measurement points (such as substations and power plants) in the power system, with each node corresponding to an observation location; E is the set of edges, representing the physical connections between measurement points; the adjacency matrix A describes the connections between nodes: if nodes are connected, A = 1; otherwise, A = 0; the degree matrix D represents the number of connections for each node. The graph convolution operation of GCN extracts spatial features by fusing information from adjacent nodes layer by layer. The formula for calculating the graph convolution of the (l+1)th layer is:

[0151]

[0152] In the formula, H (l+1) Let H be the node feature matrix of the (l+1)th layer, initially... (0) The temporal feature matrix extracted by TCN; W (l) σ is the weight matrix of the l-th layer, used to learn the feature relationships between nodes; σ is the activation function, introducing a nonlinear transformation so that the model can more effectively express complex spatial structural relationships; through layer-by-layer graph convolution, node features are propagated and fused in the topology of the power system, generating a node feature matrix H containing spatiotemporal information. (L) As the final output matrix H of the graph convolutional network, the final output matrix H contains the temporal and spatial features of each node; through spatial feature extraction of GCN, the spatiotemporal dependence of the data is enhanced, thereby improving the accuracy of the recovery.

[0153] Finally, the final output matrix H is mapped to the final recovered power data through a fully connected layer; the calculation formula for the fully connected layer is:

[0154] R = HW h +b h (20)

[0155] In the formula, R is the recovery matrix, which contains the complete estimation results for all nodes in the power system; Wh This is the weight matrix of the fully connected layer, used to project the feature matrix H onto the output space of the data reconstruction; b h This is the bias term, which adjusts the output. The output of the fully connected layer is the result of high-precision reconstruction.

[0156] This invention optimizes resource allocation by prioritizing data recovery based on security and real-time requirements. For data with high real-time requirements, matrix factorization is used for rapid recovery, shortening response time and ensuring system continuity and stability. For data with high accuracy requirements, a spatiotemporal convolutional network (STCN) is used to combine temporal and spatial features for high-precision recovery, ensuring data accuracy and integrity, thereby improving data recovery quality.

[0157] The terms "multiple," "multi-layered," and "multiple times" mentioned in this plan refer to a number greater than 1, and the specific number is determined based on actual needs.

[0158] The above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the invention. Those skilled in the art will understand that implementing all or part of the above-described embodiments and making equivalent changes in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A method for recovering abnormal power system data that prioritizes both time and accuracy, characterized in that, include: Step S1: Obtain the power dataset and determine if there are any outliers; The outliers include remote control information exchanged between the data center SCADA-EMS system and the power plant, data information exchanged between the data center SCADA and EMS systems, electricity metering information, and other electricity market information, including generation costs, market clearing prices, trading volume, and user load. Step S2: Based on the security protection level index and the response recovery level index, quantified values ​​of each outlier in two dimensions are used to form a row vector A. ij Furthermore, a decision index matrix A is established for all outliers, where A ij middle element a ij This represents the quantified value of the i-th outlier under the j-th level index; Step S3: Use the analytic hierarchy process (AHP) to determine the weight matrix W = [w1, w2]. T This includes the security protection level weight w1 and the response recovery level weight w2; Step S4: Calculate the comprehensive score Z for each outlier based on the decision index matrix A and the weight vector W. i ; Step S5, based on the comprehensive score Z i And the priority label is determined by the threshold Z´. i Each outlier in the power dataset is labeled to obtain a dataset with priority labels; the priority labels correspond to the outlier recovery algorithm. Step S6: The outliers in the dataset with priority labels are recovered using the recovery algorithm corresponding to the priority labels to obtain the normal dataset after data recovery. The recovery algorithm includes a high-precision recovery algorithm and a fast recovery algorithm. The fast recovery algorithm uses a low-rank matrix factorization algorithm for data recovery, and the high-precision recovery algorithm uses a spatiotemporal convolutional network STCN for data recovery. The spatiotemporal convolutional network STCN is composed of a temporal convolutional network TCN, a graph convolutional network GCN, and a fully connected layer. Step S7: Send the normal dataset to the data control center to provide a data source for subsequent power system analysis; Step S3 includes: Construct the judgment matrix P: ; Where, p 11 This indicates that the importance of the safety protection indicator relative to itself is 1, p 12 This indicates the importance of security protection relative to response and recovery, p 12 Let p ∈[1,9] 21 =1 / p 12 p 22 =1, p 21 This indicates the importance of response recovery relative to security protection, p 22 This indicates that the importance of the response recovery relative to itself is 1; Perform column normalization on each element of the judgment matrix P to obtain the normalized judgment matrix P': ; Among them, element p´ ij The calculation method is as follows: ; Average each row of the judgment matrix P': ; We obtain the security protection level weight w1 and the response recovery level weight w2, as well as the weight matrix W=[w1,w2]. T ; Step S6 includes: For the priority label Priority i Outliers with a value of 1 can be quickly recovered using a low-rank matrix factorization algorithm. Let data matrix Y represent the power dataset, and establish a mask matrix W for data matrix Y, where the elements w in the mask matrix W are... ij This represents the element y in the data matrix Y. ij The validity of is expressed as: ; Establish low-rank matrices U and V, and adjust the reconstruction matrix X=UV using an optimization objective function to make the reconstruction matrix X approximate the effective data in Y; the optimization objective function is: ; In the formula, w ij (Y ij -(UV) ij ) 2 This indicates that the error is calculated only at valid data locations; λ represents the regularization term; λ represents the regularization parameter. By introducing auxiliary variables Z and the Lagrange multiplier matrix Λ, the objective function is decomposed into two subproblems using the Alternating Direction Multiplier Method (ADMM), which optimizes X and Z respectively. X, Z, and Λ are then iteratively updated. First, fix Z and Λ, optimize the value of X, and simultaneously consider the current values ​​of Z and Λ. The iterative formula for X is: ; ; In the formula, k is the number of iterations, and ρ is the penalty parameter. Then optimize the auxiliary variable Z. The optimization formula for Z is: ; ; The Lagrange multiplier matrix Λ is updated using the following formula: ; Calculate the error between the valid data locations in the reconstruction matrices X and Y ||(X)|| (k+1) -Y)⊙W‖ F When the condition is satisfied ||(X) (k+1) -Y)⊙W‖ F The iteration stops when the threshold value is less than ε, where ε is a preset threshold. Using the final reconstructed matrix X=UV, a low-rank inference method is employed to fill in missing data and repair anomalous data based on valid data in Y. Furthermore, For the priority label Priority i Outliers with values ​​of 0 are recovered with high precision using the Spatiotemporal Convolutional Network (STCN). Assume the input time series of the power data in the power dataset is Q={q1,q2,q3,...,q t }, where q t This represents the power data collected at time t; The convolution operation of TCN is defined as follows: ; In the formula, z t This represents the temporal features extracted by TCN at time t; w i d represents the weights of the convolution kernel; k is the size of the convolution kernel, representing the number of time points involved in a single convolution operation; i It is the expansion rate; b is the bias term; t > d k-1 After undergoing dilated convolution operations across multiple TCN layers, the output feature sequence Z is obtained. t ={z t1 ,z t2 ,z t3 ,...,z tn }; Using Z t Further construct the initial feature matrix H (0) =[z t1 ,z t2 ,z t3 ,...,z tn As input to the Graph Convolutional Network (GCN); Spatial features of the data are extracted using a Graph Convolutional Network (GCN). The topological structure of the power system is abstractly represented as a graph G=(V,E), where V is the set of nodes representing measurement points in the power system; E is the set of edges representing the physical connections between measurement points; the adjacency matrix A describes the connection relationships between nodes: if nodes are connected, A=1; otherwise, A=0; the degree matrix D represents the number of connections for each node; the formula for calculating the graph convolution of the (l+1)th layer is: ; In the formula, H (l+1) Let H be the node feature matrix of the (l+1)th layer, initially... (0) The temporal feature matrix extracted by TCN; W (l) Here, σ is the weight matrix of the l-th layer; σ is the activation function; through layer-by-layer graph convolution, node features are propagated and fused in the topology of the power system, generating a node feature matrix H containing spatiotemporal information. (L) The final output matrix H of the graph convolutional network contains the temporal and spatial features of each node. Finally, the final output matrix H is mapped to the final recovery result of the power data through a fully connected layer; the calculation formula for the fully connected layer is: ; In the formula, R is the recovery matrix, which contains the complete estimation results for all nodes in the power system; W h This is the weight matrix of the fully connected layer, used to project the feature matrix H onto the output space of the data reconstruction; b h It is a bias term that adjusts the output result.

2. The power system anomaly data recovery method considering time and accuracy priority as described in claim 1, characterized in that: The quantitative standard for the security protection level index is as follows: when the data type of the outlier belongs to the remote control information exchanged between the data center SCADA / EMS system and the power plant, the quantitative value is 4; when it belongs to the data information exchanged between the data center SCADA / EMS system, the quantitative value is 3; when it belongs to the power energy metering information, the quantitative value is 2; and when it belongs to other information in the power market, the quantitative value is 1. The quantification standard for the response recovery level index is as follows: the quantification value is 5 when the data recovery response time requirement for the outlier is immediate; the quantification value is 4 when the data recovery response time requirement is less than 1 second; the quantification value is 3 when the data recovery response time requirement is less than 1 minute; the quantification value is 2 when the data recovery response time requirement is less than 5 minutes but greater than 1 minute; and the quantification value is 1 when the data recovery response time requirement is greater than 5 minutes.

3. The power system anomaly data recovery method considering time and accuracy priority as described in claim 2, characterized in that, Step S2 includes: Step S21: Based on the security protection level index, derive the quantified value a of the i-th outlier in the security protection level index dimension. i1 Based on the response recovery level index, the quantified value 'a' of the i-th outlier in the response recovery level index dimension is obtained. i2 ; Step S22, let row vector A ij =[a i1 ,a i2 The decision index matrix A is formed by combining the i row vectors. ; Where i∈[1,n]; Step S23: Standardize matrix A to obtain the decision index matrix A': ; ; In the formula, A j Min(A) represents the column vectors in matrix A; j ) and max(A j ) represent the minimum and maximum values ​​in column j, respectively.

4. The power system anomaly data recovery method considering time and accuracy priority as described in claim 3, characterized in that, Step S4 includes: The comprehensive score Z of each outlier is calculated by multiplying the decision index matrix A' with the weight matrix W. i : ; ; Among them, Z i Indicates the first A comprehensive score for each outlier.

5. The power system anomaly data recovery method considering time and accuracy priority as described in claim 4, characterized in that, The priority label in step S5 i The generation rules are as follows: ; Among them, Priority i A value of 1 indicates that the fast recovery algorithm is selected, Priority i A value of 0 indicates that a high-precision recovery algorithm is selected.

6. A power system anomaly data recovery system that prioritizes both time and accuracy, characterized in that, Used to implement the method according to any one of claims 1-5.

7. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Power enterprise self-service big data multistage closed-loop repair method and device

    CN112783681A

  • Comprehensive energy toughness evaluation method and system based on multi-source factor coupling

    CN118735135A